10,000 Matching Annotations
  1. Jul 2026
    1. eLife Assessment

      This valuable study provides novel evidence that congenital aphantasia is associated with structural differences in frontotemporal and cingulate systems, with relative sparing of early visual regions and major visual pathways. The multimodal structural imaging approach is carefully implemented and will be of interest to researchers studying mental imagery and aphantasia. However, the strength of evidence is incomplete because the data cannot adjudicate between alternative cognitive interpretations, and the multiple discovery streams make the findings better viewed as key exploratory evidence, rather than as establishing a definitive structural phenotype of aphantasia.

    2. Reviewer #1 (Public review):

      Summary

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (3) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

    3. Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths:

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (4) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

      We will add two recent in-press Cortex papers to the Discussion. One provides lesion-based double-dissociation evidence against V1 as a necessary causal substrate of visual imagery. The other shows that aphantasic individuals can display visualizer-like oculomotor patterns during mental map exploration despite reporting little or no imagery vividness. Together, these studies help clarify our interpretation of our null V1 findings and structural effects in higher-order brain regions, which are consistent with aphantasia involving altered integration or access rather than a primary V1-dependent imagery deficit.

      Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

      We thank the reviewer for this thoughtful and constructive comment. Lack of introspective report of voluntary imagery is arguably the defining signature of aphantasia. This motivated us to primarily interpret our anatomical findings in a broader cognitive context of higher-order control, internal monitoring, and conscious access in aphantasia. We expect that a reliable behavioural test measuring imagery sensitivity and accessibility would allow us to direct link these findings to individual imagery ability. Nevertheless, to our best knowledge, this kind of test on imagery is still missing. Instead, our findings point to some plausible structural signature or brain regions that may be related to conscious imagery, which motivate future studies to examine their direct or causal roles. We agree with the reviewer, future studies should test the relationship between these anatomical structures and the accessibility to internal representation, together with related aspects of internal monitoring. We will therefore add a dedicated paragraph to discuss the plausible cognitive mechanisms during the revision.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

      We are grateful to the Reviewer for the constructive and thoughtful assessment of our manuscript. In response to the reviewer’s comments, we will revise the manuscript to clarify the dependency between diffusion-derived analysis streams, to state more explicitly the biological limits of diffusion MRI metrics, to add a null-network sensitivity analysis for the clustering coefficient findings, to include confidence intervals for reported effect sizes, and to temper the interpretation of the cortical thickness result. We will also revise the Abstract and Discussion to better reflect the exploratory nature of the study and to frame the findings as hypotheses requiring replication in larger independent samples. We believe that these revisions will make the manuscript more balanced, transparent, and appropriately cautious, while preserving the central conclusion that congenital aphantasia is associated with structural differences centered on higher-order frontotemporal and cingulate systems.

    1. eLife Assessment

      This valuable study identifies Gcn5 as a regulator of blood cell development in the Drosophila lymph gland, with links to autophagy and nutrient-sensing mTORC1 signalling. The evidence is solid that altering Gcn5, autophagy genes and mTORC1 activity perturbs blood cell homeostasis, and the revised manuscript adds helpful genetic and quantitative analyses. However, the evidence for a clean linear Gcn5-mTORC1-TFEB/autophagy pathway is insufficient, because several cell-type-specific phenotypes remain difficult to reconcile and the pathway logic relies on different genetic tools, cell populations and pharmacological perturbations.

    2. Reviewer #1 (Public review):

      In their manuscript Arjun et al. investigate the role of the histone acetyl transferase Gcn5 in controlling drosophila blood cell homeostasis in the larval lymph gland. Using gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland cell populations, they show that these manipulations impact (but in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation, DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and show that reducing the expression of the autophagy machinery affect blood cell differentiation. By using drugs as well as genetic approaches to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR, but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      Overall, the main conclusions are sound but interpreting several lines of experiments and results remain complicated. Consequently, the overall picture of the role of Gcn5 in Drosophila larval lymph gland development, and its relationship to mTOR and autophagy, remains unclear.

    3. Reviewer #2 (Public review):

      Summary:

      Drosophila haematopoiesis has been shown to be governed by a number of signalling pathways such as JAK/STAT and Dpp. This important study shows a role for nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it TFEB activity thus acting as a fine-tuning mechanism which maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy has a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weakness:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections do not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. And the modulations of Gcn5 in PSC also has variable effects across hemocyte subpopulations which is not explored in the manuscript. Interestingly, also the domain deletion constructs show differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained. While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4. The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      Comments on the revised manuscript:

      Overall, the revisions make the narrative more coherent. The authors have also added data which substantiates their conclusions.

      However, in some instances, the authors are not clearly able to explain the discrepancies in the data (Gen-5 depletions under the Hml-Gal4 in the whole larval lysates remove p62 completely) which is not ideal.

      A query regarding the discrepancies in the immunofluorescence data: The authors have removed the IF data which suggested that there could be differences in the shuttling of Gcn5 between the nucleus and cytoplasm. The authors suggest that immunofluorescence issues are at the root of these variable results, but the reviewer wonders whether there could be further unexplored mechanisms re: shuttling that is unexplored here and would have been potentially novel.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In their manuscript, Arjun et al. investigate the role of the histone acetyltransferase Gcn5 in the control of drosophila blood cell homeostasis in the larval lymph gland. They use gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland populations to show that these modulations impact (in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation or DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and they show that decreasing the expression of the autophagy machinery increases blood cell differentiation. Using drugs to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      While the authors did a lot of experiments and good quantifications of the blood cell phenotypes, many results do not make much sense or do not bring valuable information about Gcn5 mode of action. Several conclusions of the manuscripts are not backed by solid data (e.g. that Gcn5 action is mediated by TFEB and the autophagy machinery) and different aspects of the literature are not well taken into consideration. Some results (such as the validation of the knockdown and overexpression of Gcn5) seem flawed. There are some concerns about the results obtained with gcn5 zygotic mutants and an interpretation of the phenotypes observed upon manipulation of Gcn5 expression in different cell types is missing.

      We have now performed several experiments to address the comments raised by the reviewer and have also provided possible explanation of the phenotypes in cases where it was lacking.

      Important revisions are needed to improve the quality of the manuscript and confirm the authors' findings.

      Reviewer #2 (Public Review):

      Summary:

      Drosophila hematopoiesis has been shown to be governed by a number of signaling pathways such as JAK/STAT and Dpp. This important study shows the role of nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it, TFEB activity thus acting as a fine-tuning mechanism that maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy have a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub-sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weaknesses:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections does not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. The modulations of Gcn5 in PSC also have variable effects across hemocyte subpopulations which is not explored in the manuscript.

      We have now explained and discuss why Gcn5 modulation could be affecting the PSC size. Please check Discussion section Paragraph 1 line 10 onwards.

      Interestingly, also the domain deletion constructs show a differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained.

      Currently, with our observations all that we can comment about that data is that expression of domain deletion mutants causes aberrant hematopoiesis indicating a dominant negative phenotype since they are expressed in the wild type genetic background. Beyond this, we will be exploring mechanistically how these domains are functioning during hematopoiesis in future studies. We have already described the dominant negative effect in the text: Discussion Section Paragraph 3.

      While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4.

      We have now revised Figure 1 and have only included the images with Collier-Gal4 and Dome-Gal4 which clearly shows both the niche cells, Dome-positive progenitors and Dome-negative cells of the primary LG lobe essentially showing that Gcn5 is expressed throughout the primary LG lobe. In Fig. 1C-F’, Gcn5 expression is both in the nucleus and cytoplasm as this molecule shuttles between cytoplasm and nucleus. The staining pattern with the other Gal4 could be due to problems in the immunofluorescence protocol and acquisition parameters. We have now removed those images from Figure 1. Please check revised Figure 1.

      The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      We have used a pan-hemocyte promoter for the autophagy analysis to investigate if Gcn5 regulation over autophagy is a hemocyte specific effect which we indeed see. We have removed the western blot data now in the revised manuscript where we looked at Atg8 and p62 levels in whole larval lysates when Gcn5 was perturbed using hemocyte driver as the results were puzzling and difficult to comprehend given the complete absence of a p62 band in Gcn5 knockdown conditions. Also, it’s worth noting that Hml-Gal4 is also active in the LG hemocytes. We did 2 alternate promoters for prohemocytes to cross-validate some of our results and the chemical modulators experiment was done since effects like mTOR inhibition/nutrient sensing effects are systemic and hence such modalities were employed.

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      It is a possibility that Gcn5 perturbation could be affecting lymph gland size although we have not seen any consistent trend that would point towards this phenotype either upon knockdown or over-expression. We believe Gcn5 controls blood cell differentiation phenotypes strongly via mTORC1. But in order to answer reviewer’s comment we have now re-analyzed our crystal cell differentiation data particularly and quantitated it and represented it as crystal cell differentiation index for dome-Gal4 specific Gcn5 modulation and for the data with genetic modulation of mTORC1 pathway. Please see Fig 3P and S10J for the revised analysis.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      We thank the reviewer for their useful critique. We have now addressed this concern and we have genetically perturbed the mTORC1 pathway in the progenitors using both abrogation of TORC1 via depletion of Tor or Raptor or by activation using over-expression of Rheb. We have now included this data as Supplementary figures – Fig S10 and S11 and have described it in the results section. Please see results section “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The abstract could clearly be improved. It does not make a clear presentation of what is new in the manuscript. The conclusions that Gcn5 function in the lymph gland is mediated by the autophagy machinery and the acetylation of its non-histone target TFEB are not grounded and purely circumstantial. The implication of mTOR and nutrition in drosophila larval blood cell homeostasis has already been studied but not mentioned here. Most of the time the authors do not provide any possible explanation about the phenotypes they observe and how they fit with the current literature. Several pieces of results are of serious concern.

      We would like to thank the reviewer for their feedback. We have revised the abstract and have incorporated the insights obtained from our study. We have now included relevant literature that talks about the implication of mTOR and nutrition in Drosophila larval blood cell homeostasis (Please see Introduction section Paragraph 2 in the manuscript). We have also noted the input of the reviewer on many phenotypes lacking any description of a possible explanation. We have worked on results section to provide possible explanation and speculation wherever relevant.

      In the introduction, the authors do not provide an up-to-date and accurate presentation of the field. For example, they could use much more recent and comprehensive reviews since Evans et al. 2003. (eg. MID: 30733377 or 35887113). Their choice for signaling pathways involved in Drosophila blood cell progenitors seems very much biased for lead author self-citation rather than more directly related citations. It is surprising too that the authors failed to mention a series of publications on Akt/mTOR and nutrient sensing impact on drosophila larval blood cells (PMID: 22951642 ; 22911822 ; 22407365 ; 22510984). Along the same line, there are already several reports on autophagy genes implicated in Drosophila hematopoiesis and blood cell functions (PMID: 23406899; : 33560224 ; 20498061 ; 37623416). The introduction on GCN5 is a bit of a catalogue and should be streamlined- citing a recent review would be useful (PMID: 32735945). Again, the authors fail to cite publications showing that Gcn5 levels can be modulated by nutrition (PMID 27022023; 27874008) and they do not mention that amino acids starvation or mTOR inhibition leads to a decrease in GCN5 activity / TFEB acetylation (ref 40). Taking into account all the missing information, the novelty of the present manuscript is strongly decreased.

      We would like to thank the reviewer for the detailed suggestions on including the relevant literature that are appropriate and relevant to be mentioned in the context of the observations in our manuscript. We have now included these references and have cited them as per reviewer’s suggestions. Please check Introduction section paragraph 2.

      Results

      While there is little doubt that Gcn5 is expressed in the entire primary lobes based on Fig 1C-F, the quality of the staining in G-J (especially H, J) is really poor and essentially looks like non-specific background with no clear signal in the nuclei. Better images should be presented. The conclusion of the paragraph ("all cellular populations of the LG") and title of Fig.1 are not fully accurate as the authors do not provide evidence that Gcn5 is also expressed in posterior lobes.

      As per reviewer’s suggestions, since the images in Fig 1A-D’ clearly show that Gcn5 is expressed in the entire primary LG lobe in PSC cells, MZ and CZ; we have removed panels E-H’ which lacked clear nuclear signal. Fig1A-D’ clearly show the nuclear staining pattern of Gcn5. We have also modified the conclusion of the paragraph to say that Gcn5 is expressed in cellular populations of the primary lymph gland lobe accordingly.

      Concerning, Fig S1 and Fig 2, while the analysis seems technically sound, the results are puzzling. The lack of P1 differentiation in gcn5 null heterozygotes is very surprising. The authors should check that this stock does not carry a mutation in nimC1 (for details see: PMID: 23899817) and use other plasmatocyte differentiation markers to confirm their observation (also with the different allelic combinations). I'm also concerned by the levels of plasmatocyte differentiation and crystal cell number in the control line (notably in S1H), which seem very low (and quite variable for P1 as there is a notable difference between S1H and Fig 2H). Moreover, the analysis of the allelic combinations gives rather incoherent results: PCSC cell numbers are affected only in null/hypomorph, whereas differentiation (NimC1 and Hnt), as well as DNA damage, was only increased in hypomorph homozygotes. The authors propose no hypothesis to explain these observations.

      We have now repeated these experiments with the E333st null allele by placing it on a different balancer and we observe homozygotes that are alive till late third instar/early pupal stage as shown before by Carre et al., 2005. We have now included these revised results on the plasmatocyte differentiation status of the E333st heterozygotes and homozygotes (See Fig 2 and Fig S1). We do find P1 positive cells in the E333St heterozygotes unlike earlier. Plasmatocyte and crystal cell numbers in the control line always shows some level of heterogeneity. We have included the wild type control individually with those respective mutants during the experiment hence drawing a cross comparison across two different experiments would not be appropriate. We have now explained the observations obtained on PSC cell numbers (Discussion section paragraph 1). Experiments to check all hematopoietic aspects of the gcn5 null have been done after changing the balancer line and the null mutants overall show a decrease in PSC size and a widespread increase in hemocyte differentiation which could be due to a systemic effect due to various signalling pathways being affected which needs to be investigated and is beyond the scope of this study. This has also been discussed in the Discussion section Paragraph 1.

      Although a side-by-side comparison would have been better suited, it seems that the homozygotes or trans-heterozygotes do not have stronger phenotypes than the heterozygotes as far as crystal cell and DNA damage are concerned, which is rather unexpected. Besides the authors should introduce why they look at DNA damage.

      We agree with the reviewer that for the crystal cell and DNA damage phenotype the homozygotes or trans-heterozygotes do not have a stronger phenotype as compared to the heterozygotes alone but since these are whole animal mutants there could activation/inactivation of various signalling pathways and systemic effects that would be difficult to account for and comprehend here which needs to be investigated further. The only conclusion that we draw from these observations is that Gcn5 is required for maintaining blood cell homeostasis. Regarding DNA damage, we have now included the rationale and supporting literature for why we have studied DNA damage in the context of Gcn5. Please see result section 2 paragraph 1.

      Importantly too, the authors failed to obtain gcn5 E333st/E333st (null) larvae, whereas Carre et al. originally reported that E333st/E333st individuals are viable until the late third instar larvae. I suspect that the stock they use carries additional mutations that need to be eliminated by back-crossing it to control flies for several generations. Of note too, a recent report showed that a deletion of gcn5 (generated by CRISPR) does not prevent adult emergence, challenging the conclusion that gcn5 expression is absolutely required for fly development (PMID: 37545086).

      The reviewer is right in pointing out that E333st homozygotes survive until late third instar as reported by Carre et al.,2005. We have procured the null allele again and used another balancer to obtain homozygotes and we were able to get homozygotes that survived till late third instar as reported earlier. We have now included new data from these homozygotes for all hematopoietic aspects and heterozygotes particularly for plasmatocyte differentiation Please see Fig 2 and Fig S1 and corresponding results section 2 of the manuscript.

      Concerning the validation of Gcn5 knock-down and overexpression: the results are highly dubious. In Fig S2B (hml>Gcn5 RNAi), there is virtually no Gcn5 signal in the primary lobes but hml is normally expressed only in the cortical zone. How is it possible? Similarly, the western blot (which is really too much cropped around the bands of interest) does not show any signal in the hml>Gcn5 RNAi lane (not even some background. According to the Methods section, the western was performed on whole larvae extracts; hml-mediated knock-down can not wipe out its expression in all the tissues. As for the overexpression, flag immunostaining in hml>Gcn5-flag is mostly cytoplasmic (S2E), which doesn't make sense and does not fit with S2C (Gcn5 immunostaining).

      Hml-Gal4 is a pan hemocyte driver and its expression is not limited to the CZ of the primary lymph gland lobe (Banerjee et al., 2019) and recent single cell sequencing data corroborate this that Hml domain is not limited to the cortical zone (Yarikipati and Bergmann, 2026). GFP driven by Hml-Gal4 is spread out across the primary LG lobe which could explain the phenotype of no Gcn5 signal obtained in the immunofluorescence experiment. Regarding the western blotting experiment which was performed on whole larval extracts, we were also puzzled by lack of Gcn5 bands in these lysates upon depleting Gcn5 using Hml-Gal4. We need to systematically probe further to understand expression of Gcn5 in other tissues and organs. We have now removed the western blot data as the data obtained cannot be comprehended at the moment. Regarding the FLAG staining experiment – the staining gave us a cytoplasmic pattern and since Gcn5 is known to shuttle between the cytoplasm and nucleus it is possible that the anti-FLAG staining detected the Gcn5 localizing in the cytoplasm. It is difficult to draw a direct comparison here between the images S2C and S2E as both are different antibodies.

      The initial analysis of Gcn5 level modulation in the prohemocytes, PSC or Hml+ cells is mainly descriptive and the authors do not elaborate on possible explanations based on the current literature.

      We have added a possible explanation wherever required for these respective results on Gcn5 modulation in prohemocytes, PSC and Hml positive hemocytes. Please see result section 3 where we elaborate on possible explanation for the phenotypes observed.

      The structure/function analysis of Gcn5 is based on overexpression of truncated mutants in the prohemocytes using the tep4-GAL4 driver and monitoring PSC cell, prohemocyte maintenance, plasmatocyte and crystal cell differentiation as well as DNA damage. As the overexpression of the full-length protein was made with a different driver (Dome), it is difficult to interpret the data. Nevertheless, no clear message emerges from this analysis and the authors do not reach any conclusion. Thus, the interest of these experiments remains limited.

      The structure-function analysis was largely done to understand which of the domains of Gcn5 upon over-expression results in a dominant negative like phenotype and our analysis shows that expression of some of these domain mutants results in a dominant negative phenotype in the wild type genetic background which we have now stressed upon in the text. However, further mechanistic understanding and in-depth analysis of each of these domains of Gcn5 warrants further separate investigation and is beyond the scope of this study. Please see the end of result section 4 for conclusion and possible explanation.

      The authors then analyze autophagy markers (in hml>Gcn5 LOF or GOF). Contrary to their say, hml-GAL4 is not a pan-hemocyte marker. It would have been interesting to ensure that the effects observed on Atg8 and Ref(2)P in the lymph gland are cell-autonomous- as expected for a direct role of Gcn5 on this pathway. Again, it is very surprising that p62 is not detected in the western blot on whole larval extracts when Gcn5 is knocked down in Hml+ cells only (Fig 5D). Moreover, quantifications on multiple samples will be needed to validate the increase/decrease of p62 and Atg8 as detected by western blot. As for the RT-qPCR (Fig S5), according to the Methods sections, they were made on adult blood cells but this is not explicit in the result section.

      We have corrected the text and mentioned Hml-Gal4 as a hemocyte specific Gal4 shown earlier as Gal4 marking both embryonic and larval hemocyte population (Goto et al., 2003, Yarikipati and Bergmann, 2026). Regarding the Atg8 and Ref (2)P blots – yes, it is surprising to us too that the p62 is not detected in the larval lysates when Gcn5 is depleted using Hml-Gal4. However, this result was consistent over the replicates performed and needs to be further studied. Since this phenotype of complete absence of p62 in larval lysates upon Gcn5 depletion cannot be comprehended and explained, we have removed the western blot data from the figure and have just retained the immunofluorescence data and have also quantified the Atg8 and p62 puncta per cell and included this data in Figure 5, Graphs D and E. For the qRT-PCR we have now included a description in the corresponding results section. Please see result section – result 5 under “Autophagic flux in the Drosophila blood cells is negatively regulated by Gcn5”.

      The knock-down of TFEB or several autophagy genes in the prohemocytes (tep4-GAL4) leads to a rather convincing increase in plasmatocyte and crystal cell differentiation. It would have been interesting though to quantify prohemocyte maintenance, PSC cell number, and DNA damage. Also, the authors should have performed Gcn5 GOF/LOF experiments with the same driver (they present tep>Gcn5 RNAi in Fig 8 but without the proper controls).

      We have now included data for prohemocyte index (Figure S8M) upon knockdown of TFEB and other autophagy genes along with PSC cell number, DNA damage (Supple Fig S8) in the revised manuscript. Please see corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe” for the description of the results.

      The use of chloroquine should be better described. How long was the treatment? Did the authors observe an effect on autophagy in the lymph gland? Chrorloquine also affects lysosomal pH, so it remains to be demonstrated that the effects observed here are only autophagy-related.

      We have now written a detailed protocol for the treatment in the methods section and also mentioned the treatment time which is 16 hours in the results. We have included data to validate the effect of Chloroquine on autophagy by p62 and Atg8 staining in the LG and have quantitated the data (Refer Supple Fig S9) and the corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe”

      Similarly, the use of drugs to activate (3BDO) or inhibit (Rapamycin) mTOR should be better controlled. More generally, given the promiscuous roles of mTOR (and autophagy) in the larvae, tissue-specific manipulations would be better suited.

      We have now perturbed mTOR pathway genetically by activation and in-activation and have studied the effect on blood cell differentiation. Please see Figure S10 and the corresponding result section titled “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” where we discuss the results of genetic perturbation of mTOR pathway.

      Actually, as pointed out above, it has already been shown that modulation of Akt/TOR in hemocytes or amino-acid deprivation affects blood cell homeostasis (see above). The authors should definitely discuss how their results fit with the literature on this subject.

      We have added relevant literature in the introduction section and have also discussed how Gcn5 could fit into this context of nutritional sensing and control of hematopoiesis. Please check revised Introduction section paragraph 2. Also, check discussion section in last paragraph where we have discussed role of Gcn5 in nutrient sensing.

      Again, Gcn5 levels need to be quantified using multiple samples (Fig 7M, N) before concluding.

      Sorry for not including the quantitation earlier but we have now included the quantitation for the blots presented in Fig. 7 M and N.

      Finally, the authors show that 3BDO still induces an increase in blood cell differentiation when gcn5 is knocked-down in tep4+ cells and that Rapamycin still represses differentiation when Gcn5 is overexpressed in Dome+ cells. They conclude that mTORC1 overrides the effect of Gcn5. This seems a far-reaching conclusion given the available evidence.

      We have now toned down the conclusion that we make to accommodate other possibilities which we have been unable to test here currently.

      In particular, in the conditions used, the authors do not necessarily assess the activity/requirement for Gcn5 and mTORC1 in the same cell population.

      Other comments and suggestions:

      The discovery of the SAGA complex is not Grant 1999 but 1997 (PMID: 9224714).

      Ref 30 is not appropriate -nothing to do with HAT.

      GCN5 not only acetylates TFEB but also Atg7 (PMID: 28594263) to limit autophagy.

      Thank you so much for these suggestions. We have made the necessary amendments in the references.

      In the results section, the first paragraph is largely a repetition of the introduction. The same is true for most paragraphs in this section. A shorter (hypothesis-driven) introductory sentence would be more adequate.

      We have now taken the suggestion into consideration and made the necessary change in the results section throughout the manuscript.

      Fig 1: it seems that there is a higher accumulation of Gcn5 in a few cells in the cortical zone. This may correspond to crystal cells and could be easily confirmed.

      We have now checked this aspect. Please see supple fig S5 where we co-stain lozenge-GFP cells containing LG with Gcn5 to check for the accumulation. However, we do not see any accumulation in the Lozenge-positive crystal cells.

      Figure 3: the authors should also quantify the proportion of progenitors (dome>GFP+) in the different conditions.

      We have now done this and added it to the Figure. Please see panel N in Figure 3 and Figure S8M.

      Figure S3: how do the authors explain that Gcn5 knockdown in the PSC reduces plasmatocytes differentiation (but does not affect PSC cell number or crystal cell differentiation)? What could be the origin of the increase in DNA damage (essentially in CZ)? How do they explain that Gcn5 over-expression increases PSC size but does not affect (reduce?) blood cell differentiation?

      These observations need to be investigated further. We currently have no answer to these comments. The signals that are produced by the PSC could be affected due to which we observe these phenotypes like an effect on plasmatocyte differentiation and an increase in DNA damage whereas no effect on PSC cell numbers or crystal cell numbers which needs to be studied further. Also, in the case of Gcn5 over-expression in PSC we do not know how the increased size of PSC controls differentiation. This would need further experimentation and since this paper is not about the role of Gcn5 in PSC exclusively, we will look into this in our future studies. These aspects will be studied in our future follow-up studies as it is beyond the scope of the current manuscript.

      Figure S4: how do the authors explain the non-cell autonomous increase in PSC cell number upon Gcn5 KD/GOF in hml+ cells? How do they explain the increase in crystal cell number in Gcn5 GOF? Is it really cell-autonomous (i.e. all the Hnt+ cells are Hml+?)?

      We have discussed how Gcn5 depletion or over-expression in HmlΔ cells could affect PSC cell numbers. Please see discussion section, paragraph 1. Regarding the crystal cell phenotype - We have now tested if the increase in crystal cell numbers is cell autonomous by driving Gcn5 over-expression using a crystal cell specific driver and we find that the increase is cell-autonomous. Please refer to Supple Fig S5.

      The discussion is lengthy and should be reduced. It does not appropriately consider the current literature.

      We have tried to reduce the length of the discussion and have also added relevant references as per recommendations of the reviewer.

      Reviewer #2 (Recommendations For The Authors):

      (1) In general, it is not clear why in some of the experiments Tep-Gal4 is used to modulate proteins in prohemocytes while in others Dome-Gal4 is used.

      There is no particular reason. These Gal4’s have been used interchangeably as both label the hematopoietic progenitor population. Although recent single cell sequencing data has identified subsets within the progenitors namely core progenitors marked by tep4 largely and dome being a distal progenitor marker (Cho et al.,2020, Girard et al.,2021), in our study perturbations in Gcn5 using either of the Gal4’s results in a similar phenotype.

      (2) Considering alteration in lymph gland size (Figure 2), the number of positive cells should be analysed in relation to total cell numbers or s4ize.

      Although we do not find any visible differences or defects in the overall LG size in various genetic conditions discussed in this manuscript, we have done so for the plasmatocyte differentiation where we have represented it as plasmatocyte differentiation index (relative to the size of primary LG lobe) throughout the manuscript. We have now done this for crystal cell numbers too for critical genotypes in this manuscript and have represented it is as crystal cell index for example please see Figure 2O, 3P, S5G, S10J where these graphs have now been added.

      (3) Figure 1A G-I' does not look like mCD8 GFP expression, but rather cytoplasmic GFP.

      We have made the change in the figure and the corresponding text accordingly.

      (4) One of the main conclusions in the manuscript is that Gcn5 affects autophagy (Figure 5). Here, the puncta need to be quantified (relative to total cell numbers).

      Thank you for the suggestion. We have now quantitated the p62 and Atg8 positive puncta per cell and have represented it as panel D and E in Figure 5.

      (5) Figure 5 D and E show p62 and Atg8 total protein levels in the larvae when Gcn5 is modulated only in the hemocytes. It is surprising that there is a complete reduction in p62 levels across the whole larvae when Hml gal4 is used for the knockdown.

      Yes, we observe a complete absence of p62 in whole larval lysates when Gcn5 is depleted using Hml-Gal4 and we see this across replicates. This result is indeed puzzling to us and difficult to comprehend as to why a hemocyte specific driver would result in such a dramatic change hence we have decided to remove the western blot data as it is difficult to draw a solid conclusion from. We have retained the immunofluorescence data which shows a consistent alteration in autophagy upon Gcn5 perturbation using Hml-Gal4 and we have now included the quantification for the number of p62 and Atg8 positive puncta per cell for the IF data.

      (6) The beta-actin levels in the western blots in Figure 5 are highly oversaturated and do not represent loading control adequately. Also, it looks like there is substantially more total protein in 5D 3rd lane where Gcn5 is overexpressed.

      Thank you for pointing this out. We have loaded equal amount of protein in all the wells so we are unsure why the actin bands look over-saturated. We have now removed the western blot data from this figure as the data is puzzling and difficult to comprehend given a total absence of p62 in whole larval lysates in Gcn5 depletion conditions using Hml-Gal4. Hence, we are just retaining the immunofluorescence data.

    1. eLife Assessment

      This study uses convincing modeling methods and analyses of rich behavioral datasets to investigate the role of attention in value-based decision making; for instance, as when choosing between two snacks. The results are important, as they challenge existing theories that assume that paying attention to an available option biases the eventual choice toward that option. The results suggest that the correlation between attention and decision-making is formed largely after rather than before the (internal) choice process has terminated, a finding that offers an intuitively appealing rethinking of how attention and decision-making processes interact during value-based choices.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the weaknesses raised in the previous round of review.]

      Summary:

      This study examines whether gaze direction actively shapes choice during food preference decisions or whether gaze and choice evolve largely independently until the moment of commitment. The established framework in this context, the aDDM, assumes that gaze causally biases the accumulation of evidence in favour of the fixated item. The authors show convincingly that this model fails to fit key behavioural patterns across several datasets, as do other published models that make the same assumption. The authors propose an alternative model (Post-Decision-Gaze or PDG) in which gaze and decision formation are decoupled: gaze does not influence the decision process, nor is it drawn toward the ultimately chosen item, until after the decision threshold is reached. Only during the motor execution period (after commitment) is gaze directed to the chosen option. They demonstrate that this model fits several observed patterns better than the aDDM and related variants.

      Strengths:

      The work thoroughly considers multiple models and datasets. It advances an interesting alternative perspective on gaze-decision interactions and highlights meaningful shortcomings in existing models. The authors take the time to explain how modelling assumptions produce specific patterns in the data, which is certainly insightful to readers interested in the modelling of value-based decision making.

      Weaknesses:

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

    3. Reviewer #2 (Public review):

      Summary:

      Zylberberg et al. reanalyze eye-tracking and behavioral data to test two predictions of the attentional Drift Diffusion Model, finding that these predictions are not met. Similarly, predictions of normative models (inspired by rational inattention) are not in line with the data, and the authors propose a post-choice model of attention. This model better accounts for the two effects but also does not account for all patterns, so the authors conclude that eye movements most likely reflect both pre- and post-decisional processes.

      Strengths:

      A clear strength is the systematic falsification-based approach of the paper, establishing (partially) new predictions and testing to what extent these are met by extant models and by a newly developed theory. The authors do a good job in providing intuitions behind the effects and the reasons why models such as the aDDM predict them. The paper is of substantial relevance for the field, as it shows that effects pertaining to the last fixation(s) should be interpreted with caution. Another strength is the paper's transparency as the authors clearly acknowledge that their new model does not do a perfect job either.

      Weaknesses:

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

    4. Reviewer #3 (Public review):

      Summary:

      In this study, the authors reanalyzed choice, RT and gaze datasets collected from human subjects performing a food-choice task. They show that models that posit a causal role for attention in shaping the decision-making process fail to account for empirical observations in the data. These include the attentional drift diffusion model (aDDM) and models that derive attention-choice associations from an optimal policy. The authors show that a model that assumes that gazes are directed towards the chosen option after decision commitment captures more (but not all) empirical findings, suggesting that attention may reflect decisions once they are made instead of contributing to their formation. However, this post-decision-gaze (PDG) model failed to capture all aspects of the data, suggesting that gaze may reflect both decisional and post-decisional operations, and existing models are still missing some features of the gaze-directing process. The authors provide convincing evidence that post-decision gaze explains a number of empirical findings in this task.

      Strengths:

      (1) The analyses are generally appropriate, and the conclusions are supported by the data.

      (2) The study was rigorous, as the authors considered a number of alternative possible models for behavior, and evaluated their performance based on a wide range of qualitative predictions (as opposed to exclusively relying on model comparison).

      (3) The proposal that gaze may largely reflect post-decisional processes is interesting, and as far as I am aware, novel.

      Weaknesses:

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

      In particular, the "saccadic execution" parameter appears far too long and too variable to reflect merely execution; instead, it likely includes decisional components. This would make more sense since manual and saccadic planning essentially rely on distinct brain areas, hence it seems unrealistic that crossing a single threshold would trigger both manual and saccadic execution. Similarly, recovered manual non-decision times are substantially longer (though not more variable) than expected motor execution durations for button presses. These patterns suggest that parts of what the model treats as non-decision time are likely decisional in nature, although perhaps related to "action decision" rather than the "value-based decision" of interest to the authors. To what extent these two processes neatly follow each other or overlap could be usefully considered.

      We have added a paragraph to the Discussion explaining how our model’s estimates of sensory and motor latencies relate to corresponding values inferred from physiology or behavioral manipulations (e.g., Bompas et al., 2024). Specifically, we write:

      “The key assumption of the PDG model is that there is a delay between the moment a choice is internally committed and the moment it is externally reported with a key press. Because eye movements are typically faster than manual responses (𝜏<sub>e</sub> < 𝜏<sub>m</sub> in our simulations), this delay creates a window during which gaze can already be directed toward the covertly chosen item before the response is formally registered. We do not interpret these non-decision latencies as irreducible physiological minima for moving the eyes or pressing a button (Bompas et al., 2025). Rather, they are inferred indirectly by fitting an additive non-decision-time parameter to the behavioral data, which we decompose into a sensory delay (𝜏<sub>s</sub>) and a manual execution delay (𝜏<sub>m</sub>). Values of 𝜏<sub>e</sub> are then chosen so that the model reproduces the observed magnitude of the behavioral effects. This estimation procedure has important limitations. Some participants show relatively “flat” chronometric functions: response times vary little with value despite otherwise normal psychometric performance. Such patterns likely reflect processes not explicitly represented in the model, including procrastination, reduced motivation, task-unrelated thought, or noise in item ratings. Within a drift-diffusion framework, however, these cases are accommodated by assigning a long non-decision time together with a short evidence-accumulation period (Table S1). Consequently, some estimated non-decision times are substantially longer than would be expected if they represented only sensory and motor delays. A further limitation is conceptual. We model non-decision time as occurring either before or after evidence accumulation, whereas in reality decisional and non-decisional components are likely temporally interleaved (Graziano et al., 2011). This simplification may also inflate the recovered latency estimates. With these caveats in mind, sensory and oculomotor delays on the order of 300 ms remain broadly plausible, although they likely lie near the upper end of a realistic range. The estimated eye-movement latency is especially long. For instance, in monkeys trained to report simple perceptual decisions with a saccade, roughly 100 ms elapses between the threshold-crossing signal in parietal cortex (or the superior colliculus) and the executed eye movement (Roitman and Shadlen, 2002; Stine et al., 2023). Crucially, however, varying the assumed non-decision latencies across a reasonable range does not alter the qualitative predictions of the model (Fig. 8).”

      Further, we have added a parameter sensitivity analysis. Importantly, although the magnitude of the predicted effects depend on the non-decision latencies, the qualitative aspect of these predictions do not (new Figure 8). Specifically, (i) the increasing tendency to look at the ultimately chosen item as time elapses (new Fig. 8A), (ii) the lack of an interaction between the last-fixation bias and overall value (Fig. 8B), and (iii) the absence of an effect of choice consistency on Δdwell (Fig. 8C) are all findings that are independent of 𝜏<sub>e</sub>.

      Reviewer #2 (Public review):

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Following this suggestion (and the reviewer’s elaboration in the private comments to the authors), we have substantially restructured the manuscript. Both aDDM predictions are now presented together (new Fig. 2), and Figs. 3–4 test these predictions across multiple food-choice datasets. In doing so, we no longer treat the data from Krajbich et al. (2010) separately, and we extend the analysis of the last-fixation–choice association (MELFB) to additional datasets. We note that the same datasets could not be used in both Figs. 3 and 4, as some lack information on the final fixation required for the MELFB analysis. Nevertheless, results are highly consistent across datasets and align with findings from a recent study by Ting & Gluth (2025), which independently identified and examined one of our key predictions; this work is now cited in the revised manuscript. Finally, to reduce redundancy, we have consolidated all aDDM variants and optimal models into a single figure (new Fig. 10).

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

      The key predictions of the model depend on the difference between the manual (𝜏<sub>m</sub>) and eye-movement-related (𝜏<sub>e</sub>) latencies. We have now added a parameter-sensitivity analysis to show how the model predictions depend on this difference. The new analysis shows that while the quantitative predictions do depend on the precise latency values, the results are qualitatively similar across values of 𝜏<sub>e</sub> (new Figure 8).

      Reviewer #3 (Public review):

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

      Thank you for this suggestion. We added a new paragraph to the discussion (paragraph #2), where we argue that it is sensible for a decision maker to direct the gaze to the chosen item once a covert choice commitment has been made, as the benefits of attending to a stimulus do not end with the decision itself. Specifically we now write:

      “Instead, these observations are better explained by a post-decision account of the gaze-choice association that is, one in which gaze shifts to the selected item after a covert commitment to a choice. We argue that directing gaze to the chosen item after a covert choice commitment is sensible, as the benefits of attending to a stimulus do not end with the decision itself. In naturalistic settings, for instance, selecting a food item is typically followed by the action of reaching toward it, where visual attention supports spatial localization and motor planning for the upcoming action. Although participants in our computerized task did not physically act on their choices, these sensorimotor processes are likely highly automatized and may still be engaged by default, even when not strictly required. Beyond motor preparation, post-decisional attention may also serve additional functions, such as facilitating sensory anticipation of the reward, supporting metacognitive evaluation of the decision, and contributing to value updating for future choices. From this perspective, a degree of attentional “stickiness” whereby the chosen item remains preferentially attended after commitment could emerge as an effectively optimal policy once these post-decisional processes are taken into account. Moreover, a specific feature of the task design may further reinforce this tendency: in the snacks paradigm, the unchosen item typically disappears from the screen immediately after a response is registered. It is therefore plausible that directing gaze to the chosen item after commitment partly reflects anticipation of the imminent disappearance of the unchosen option. To disentangle these mechanisms, it would be interesting for future work to test whether this attentional bias persists when the chosen item, rather than the unchosen one, is the stimulus that disappears upon response.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments:

      (1) Framing of the modelling approach

      The manuscript would benefit from acknowledging the known limitations of DDM-based frameworks, especially given that the entire study is conducted within these constraints. The introduction highlights successes of the DDM, but the manuscript does not mention any of its conceptual or empirical limitations.

      We are unsure about what specific limitations the reviewer has in mind, but we have added a paragraph to discussion mentioning some limitations, like the inflation of the non-decision times and the difficulty of interpreting the fit parameters (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (2) Dependence on non-decision time assumptions

      The alternative model's explanatory power appears to rely heavily on assumptions regarding the decomposition of non-decision time: fixed visual encoding (𝜏<sub>s</sub>= 0.3 s), manual non-decision time (𝜏<sub>m</sub>; two free parameters), and saccadic execution (𝜏<sub>e</sub>; fixed parameters μ<sub>e</sub> = 0.35, σ<sub>e</sub> = 0.11).

      - 𝜏<sub>e</sub> is substantially longer and more variable than typical saccadic execution times, suggesting it likely incorporates decisional components.

      - Estimated 𝜏<sub>m</sub> values are approximately twice as long as known manual execution durations.

      - σnd is more plausible, implying that variability is captured correctly but mean durations are not.

      Together, these points raise the possibility that portions of what the model treats as non-decision time are in fact part of a (action) decision process. Only then does it make sense to assume that Tm is usually larger than Te. If Tm and Te were truly execution delays, then Tm would always be larger than Te.

      You may find it helpful to consider the framework in Bompas et al. Psych Review (2024), which discusses in detail what non-decision time is likely to comprise across effectors.

      Thank you we have added (i) a sensitivity analysis showing that our results are robust to changes in the specific value used for the eye movement related latencies (new Fig. 8), and (ii) a new paragraph in Discussion addressing the issue of the mismatch between our parameter estimates and the manual and saccadic execution times (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (3) Code availability.

      The authors should consider sharing all relevant code and data publicly.

      We agree, we now share the code and data on GitHub and indicate so in the revised manuscript.

      Minor Comments:

      (1) Lines 74-77. These are not worded as predictions but as questions; one tests predictions, but answers questions. I feel it would be clearer to stick to predictions (like in the abstract), and the introduction could benefit from explaining these predictions in a bit more detail (I found it difficult to get my head around these predictions from the intro text only).

      We rewrote the section in the introduction where we provide a gist of the model predictions (last paragraph of Introduction). We agree with the reviewer that the previous explanation was not clear.

      (2) It is confusing that panel B appears to the left of panel A in Figure 2.

      We agree. We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed and the panels follow a more logical order.

      (3) Figure 3C - remove MATLAB toggles.

      Yes, thanks.

      (4) Figure 5A shows the proportion of left choices, but the text and legend refer to right choices.

      Good catch, thank you.

      Reviewer #2 (Recommendations for the authors):

      This may appear self-serving, but the authors seem to be unaware of some highly relevant work from our group. Most importantly, in a recent publication (Ting & Gluth, 2024, JEP General), we have already looked at the dependency of the last- (or final-) fixation bias on overall value in value-based (VB) and perceptual (P) decisions. In VB, we found a negative effect; in P we did not find a significant effect. This is largely consistent with the current results, showing a negative but not significant trend. Another relevant work is Gluth et al. (2020, Nat Hum Behav), where we extended the aDDM by assuming that the probability to fixate on an option is a function of the accumulated evidence for that option. It would be interesting to know whether this assumption changes the predictions of the aDDM. Finally, we just published a new theory on how people search for information to make efficient value-based decisions (Gluth et al., in press, Psychol Rev; https://osf.io/preprints/psyarxiv/3qzak_v2). Although this theory focuses on multi-attribute choices, it can be applied to "simple" choices, too (by assuming that there is only one attribute = value). Interestingly, while the model also mispredicts a (slight) increase of the last-fixation bias with overall value, it correctly predicts the independency of the dwell-time advantage effect on choice consistency as well as the small increase of the effect with RT (attached here is a figure to show this: [https://elife-rp.msubmit.net/elife-rp_files/2026/01/22/00149589/00/149589_0_attach_9_477122. pdf], and the match with the empirical data shown in Figure 3B and 12 is striking). In general, the model shares many features of the Callaway and Jang models, but does not need to assume a biased value prior, which the authors suggest is responsible for the misprediction of the second effect. I leave it up to the authors to discuss this new theory, but I wanted to point this out.

      Thank you for pointing this out; these are all relevant points and studies.

      We now note that the first of our predictions has recently been identified and tested by Ting and Gluth (2025).

      We also considered extending the manuscript with a variant of the model proposed by Gluth et al. (Psychological Review, 2026). In fact, we attempted to fit this model to the Krajbich et al. (2010) dataset under the assumption that the duration of each sampling epoch is a free parameter. We find this model very interesting. However, in our current implementation it appears to make the same qualitative prediction as the aDDM, namely that ΔDwell depends on choice consistency (see Author response image 1).

      Given this, we have decided not to include these results in the manuscript. It remains possible that with further development particularly with a more realistic specification of fixation durations (e.g., allowing them to depend on value) the model could account for the full set of observed effects. We think this would be best addressed in a separate study.

      That said, we do find the model promising, as it provides a better account than most of the alternative models we explored for the patterns shown in panels D, H, and I.

      Author response image 1.

      Fits of a variant of the MACS model (Gluth et al. 2026) to the data of Krajbich et al. (2010).

      The paper would benefit substantially from restructuring. The aDDM's predictions are provided first, together with the empirical data, and then the optimal models are discussed. But Figure 2 shows all of this together. Later, the new (PDG) model is elaborated, and its predictions are shown. Towards the end of the results, variations of the aDDM and combinations of aDDM and PDG are shown in a series of figures (8-11), followed by a last figure showing one of the tested effects in other datasets. All of this feels pretty much thrown together without a clear structure. For instance, the aDDM and the optimal models could be described together (or the optimal models get a separate figure). The additive variants could be described earlier. And some figures could be put into the supplement. And the empirical results of the different studies could be shown together.

      We fully agree with this suggestion. We have now restructured the manuscript along the lines proposed by the reviewer (see the more detailed explanation of the restructuring in our response to the public comments).

      I strongly suggest avoiding the term "influence" in the y-axis of Figure 2, upper row, as it implies causality. Similarly, in line 182, the term "causal influence" is used in the context of the Callaway model, but as far as I know, this is not what the model assumes.

      We replaced the y-axis label with “Association of last dwell with choice (β)”

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 2 - Panel labels for A and B are reversed?

      We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed.

      (2) Does 3C include a .pdf screenshot?

      Thank you, it’s a Matlab bug on Mac. I guess they want us to switch to Python -:)

      (3) Figure 4 - It would be helpful if the green line were defined in the figure legend.

      Added

      (4) The effect size in 5B looks much more dramatic than in 2B(A?) - Is this for one example subject as opposed to all subjects? Please clarify what is different about the data.

      We are no longer showing the psychometric functions in Figure 2.

      (5) Line 252 - they say they compared the probability of choosing the right item (Fig. 5B) by the y-labels of that figure, which are all p(choose left).

      Yes, corrected now.

      (6) In general, they reference the subpanels of Figure 5 out of order, which causes the reader to jump around. They might consider reordering the panels of the figure so they follow the ordering of descriptions in the text.

      We agree, we have rearranged the figure panels to follow the ordering of the descriptions in the text.

    1. eLife Assessment

      This important study measures single-unit activity in area MT of awake-behaving monkeys to test the idea that sensory adaptation contributes to flexible evidence accumulation during decision making. The authors provide compelling evidence that adaptation to different temporal contexts shapes both perceptual judgements and neural responses. Although the precise computational mechanisms underlying these effects remain uncertain, the results support the conclusion that recent sensory history influences the temporal dynamics of decision formation. This work will be of interest to researchers studying visual perception, sensory adaptation, and decision making.

    2. Reviewer #1 (Public review):

      McGaughey and Gold ask where in the decision process the flexibility of evidence accumulation arises, proposing that it is not solely a property of downstream integrators but is also supported by stimulus-specific sensory adaptation in the middle temporal area (MT). Recording single-unit activity in rhesus macaques during a motion direction-discrimination task in which an adapting stimulus of varying temporal stability precedes an identical test stimulus, they find that more rapidly changing contexts produce weaker and less discriminable MT responses to the test stimulus, which they argue accounts in part for context-dependent changes in decision-making behavior. Through session-level correlations they further identify pupil-linked arousal as a parallel, apparently separable contributor.

      The main strength is the shift of perspective toward the encoding stage: rather than treating MT as a static input to flexible downstream integrators, the authors show that early sensory cortex can itself contribute adaptive, context-dependent signals that shape behavior. The conceptual advance is supported by a well-designed paradigm-total exposure to each motion direction is matched across conditions and the test stimulus is held identical-together with single-unit recordings and simultaneous pupillometry. The behavioral effect is consistent across three animals, and the fact that context-dependent differences emerge over repeated stimulus presentations within a trial, rather than as a sustained baseline offset across blocks, ties the effect convincingly to stimulus-specific adaptation.

      The behavioral effect constrains the temporal dynamics of decision formation but does not uniquely identify its algorithmic basis: a leak, a saturating non-linearity, or a reduction in the gain of integration are all compatible with a shallower rise of accuracy with viewing time, and the reduced MT discriminability is itself an encoding-stage efficiency effect of this kind. The manuscript appropriately treats the algorithmic basis as unresolved, noting that distinguishing these accounts would require analyses not available here, such as reverse-correlation or motion-energy kernels with lower-coherence test stimuli.

      The inference that the adaptation- and arousal-related signals operate independently rests on the absence of session-wise correlations between the neural and pupil measures and their behavioral contributions. Given the noise in the trial-wise estimates, this is best read as consistent with, rather than demonstrating, true independence, as the authors note.

      Overall, the authors largely achieve their aim of showing that sensory adaptation in MT shapes the evidence available for time-dependent perceptual decisions. The evidence for a sensory-encoding contribution is convincing, while the claim of independence between adaptation and arousal is more tentative and is framed as such.

    3. Reviewer #2 (Public review):

      McGaughey and Gold trained rhesus macaque monkeys to perform a motion-direction discrimination task in which a behaviorally irrelevant adapting stimulus with either fast or slow direction alternations preceded a variable-duration test stimulus, while simultaneously recording single-unit activity in area MT and pupil diameter. They report that adaptation to the more rapidly changing stimulus was associated with reduced behavioral sensitivity, attenuated test-evoked MT responses, and larger pupil-linked arousal signals. The authors interpret these behavioral changes as evidence for context-dependent adjustments to the temporal dynamics of decision formation and argue that these adjustments are supported by both sensory adaptation in MT and arousal-related mechanisms. More broadly, they conclude that flexible evidence accumulation in dynamic environments arises from distributed adjustments across sensory encoding and neuromodulatory systems rather than solely from changes within a downstream accumulator. If correct, this interpretation has important implications not only for our understanding of perceptual decision making, but also for broader theories concerning the functional role of sensory adaptation.

      The conclusions of the paper are generally supported by the data. Evidence for adaptation-induced changes in sensory encoding, behavior, and pupil dynamics is convincing, and the revised manuscript substantially strengthens the connection between the behavioral findings and the proposed decision-making framework.

      Comments on revised version.

      The revised manuscript provides a clearer account of how recent stimulus history influences behavioral performance. In the original version, aspects of the psychometric functions were interpreted as evidence for a more leaky evidence-accumulation process, although some of these effects could potentially have reflected alternative mechanisms, including influences of the adapting stimulus on short-duration trials. The additional analyses and discussion included in the revision clarify that information from the adapting stimulus contributes to behavior at short viewing durations and appropriately temper claims regarding the specific computational mechanism underlying the observed behavioral effects. While the data do not uniquely identify whether these effects arise from changes in leak, other nonlinearities, or related decision processes, they provide convincing evidence that recent temporal context influences the temporal dynamics of decision formation.

      My original review also noted that different sections of the manuscript relied on different behavioral metrics and analytical approaches when relating behavioral changes to neural and pupil-linked measures. The revised manuscript now provides a clearer rationale for these choices, including distinctions arising from the different trial types and time windows used in the neural and pupil analyses.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) Alternative mechanisms for performance differences.

      The authors assume that the difference in performance between the low-switch (LS) and high-switch (HS) frequency conditions is explained by a change in the "leakiness" of integration. However, several other mechanisms could potentially explain this effect:

      (1) Temporal Uncertainty: Integration might start later in the HS condition, leading to lower performance.

      (2) Reduced Efficiency: Integration could be less efficient in the HS condition (i.e., lower signal-to-noise ratio) without a change in the leak parameter itself.

      (3) Evidence Contamination: Motion information from the adapting stimulus in the HS condition may be integrated rather than ignored, which might be the case since the transition from the adapting to the test stimulus is not externally cued.

      To distinguish between these alternatives, I suggest two possible analyses. First, a formal model comparison could be performed, though I acknowledge this may be inconclusive in the absence of response-time data. Second, an analysis of motion energy kernels could be revealing; the leak hypothesis makes the specific prediction that for long test stimuli, early samples should contribute more to the choice in the LS condition than in the HS condition, relative to late samples.

      We thank the reviewer for raising these important points. We agree that we cannot definitively identify the algorithmic underpinnings of the behavioral effects we report and have made substantial revisions to the manuscript to be clearer about what is supported and what is speculative in our claims. Most importantly, we agree that we do not know if the context-dependent differences in how accuracy depends on viewing time are based on adjustments to a leak or to something else (e.g., a saturating non-linearity, as we identified in Glaze et al, 2015, that is separate from the leak itself), which we cannot resolve with this dataset, even with more formal model comparisons. We therefore:

      Changed the wording throughout the manuscript to refer to changes in leakiness as just one of several possible sources of the behavioral differences. We also added this point to the list of “limitations” (and possible future directions, including using motion-energy kernels, which would require us to use lower-coherence test stimuli) in the Discussion (L487-493).

      Added a new figure panel (Fig. 2D), a new Extended Data figure (Extended Data Fig. 3), and additional explanatory text (L168-175) that collectively describe the behavior in more detail, including quantifying a “crossover” dynamic similar to what we reported previously (Glaze et al, 2015).

      Added new explanations (L152-163) and analyses (Extended Data Fig. 9) indicating that the monkeys used some information from the end of the adapting stimulus to inform their decisions, which accounts for the patterns of choices at the shortest viewing durations.

      Indicate that the context-dependent differences in the slopes of the psychometric functions (and complementary analyses based on “raw” accuracy measures as a function of binned viewing duration) rule out the temporal uncertainty and evidence contamination explanations, but are consistent with effects on the temporal dynamics of the decision process (L175-179).

      (2) Independence of neural and pupil-linked signals.

      The authors take the lack of session-wise correlation between context-dependent contributions from neural and pupil terms as evidence that these two signals provide independent contributions to the behavioral effect. However, could this lack of correlation simply be a result of high variability or noise in these estimates? The data shown in Figure 7B suggests that measurements are very noisy, which might obscure a potential relationship.

      We agree that the lack of session-wise correlation between neural and pupil terms cannot be taken as definitive evidence of independence. We have both softened the language around the claim (L368) and added a sentence to the Discussion (L464-468) acknowledging that this lack of correlation may reflect underlying noise and/or variability rather than true independence of the underlying mechanisms.

      Reviewer #1 (Recommendations for the authors):

      (3) The neural data analyses rely fundamentally on "switch" trials (Figures 3-5). It might be informative to also examine "non-switch" trials to see if there are specific neural markers indicating the exact moment the motion stimulus becomes behaviorally relevant. Given that this may fall outside the primary focus of the paper, it is up to the authors whether to pursue this line of inquiry.

      We thank the reviewer for this suggestion. We agree and have added new analyses of data from non-switch trials (Extended Data Fig. 9), which show some effects of stimulus information from the adapting epoch on the monkeys’ choices, as we detail below in response to related comments from the other reviewers.

      Reviewer #2 (Public review):

      Aspects of the behavioral analysis would benefit from a tighter connection between theoretical claims about evidence accumulation and the empirical features of the psychometric functions. For example, the rightward shifts observed across adapting conditions are interpreted as consistent with a reset of accumulation on switch trials, but similar patterns could also arise from failures to detect the test stimulus on a subset of trials, leading responses to default to the final adaptor direction. Likewise, changes in psychometric slope and asymptote are attributed to differences in evidence accumulation without explicit modelling or consideration of alternative explanations.

      Clarifying how specific features of the psychometric functions map onto distinct components of the decision process will strengthen the link between the theoretical framework and the behavioral data.

      We agree and have made substantial revisions to address these important points. Specifically, we added a new figure panel (Fig. 2D), new Extended Data Figures (3 and 9), and several lines of explanatory text (L152-179) that collectively describe the behavior in more detail, including clarifying that: 1) for the shortest viewing durations, the monkeys’ decisions were informed by information from the adapting stimulus, which accounts for generally lower accuracy on LSF (longer exposure to the final adapting direction, thus more accumulated evidence for that direction before processing the switch) vs. HSF (shorter exposure to the final adapting direction, thus less accumulated evidence for that direction before processing the switch) switch trials; and 2) as viewing duration increased, the rate of rise of accuracy versus viewing duration was higher for LSF vs. HSF trials, implying differences in the process of evidence accumulation. As detailed in our response to a similar comment from Reviewer 1, above, we are now careful to temper our claims about the specific computational basis (e.g., a leak or other form of nonlinearity) for these differences.

      We also de-emphasized our treatment of the asymptotes of the psychometric functions. In principle, these regimes could give insights into leakiness (which can limit the total amount of information that can be accumulated) and lapses (which are measured at the asymptotes). In practice, however, the long-duration trials that constitute the asymptotes were relatively under sampled (to promote the unpredictability of the offset of the stimulus, which we believed was the more important consideration when designing the experiment), yielding unreliable estimates.

      A slight concern is the lack of a consistent analytical approach for relating behavioral changes to neural and pupil-linked measures. Different sections of the manuscript rely on different behavioral metrics-such as differences in accuracy within a selected stimulus-duration range (e.g., Figure 5C) or psychometric slope differences (Figure 6C) without clear justification for these choices. The analytical approach likewise varies between simple correlational analyses (Figure 5C, Figure 6C), pseudo-experimental group comparisons (Figures 5D, E), and the inclusion of neural or pupil terms in the behavioral psychometric regression model (Figure 7B). While each metric and approach may be defensible in isolation, adopting a more consistent framework will help convince readers that the reported effects are robust and not contingent on the selective choice of metric or analysis.

      We thank the reviewer for this thoughtful critique and agree that the rationale for our choice of behavioral metrics and analytical approaches could be stated more clearly. We have added text to the relevant sections of the Results (L247-251) clarifying these choices. In particular:

      The neural analyses (Figures 3D-E, Figure 4, Figure 5D-E) focused on preferred-motion switch trials, because: 1) low switch-frequency non-switch trials provide an additional 800 ms of exposure to the final adapting-stimulus motion direction relative to high switch-frequency non-switch trials, which confounds comparisons of context-dependent evidence encoding between conditions, and 2) MT neurons exhibit minimal responses to null motion (although note that we also included analyses based on ROC area, which is computed from both preferred- and null-motion switch trials, to account for possible contributions of null-motion responses; Figure 5A-C). Thus, to ensure a meaningful comparison between neural and behavioral measures, we used behavioral accuracy on switch trials as the relevant metric in Figure 5C-E, rather than psychometric slope, which is estimated across both switch and non-switch trials.

      The pupil analyses (Figure 6) focused on a time window preceding test-stimulus onset, representing the arousal state around when the decision process started, and included both switch and non-switch trials. Thus, for these analyses we used psychometric slope, which is estimated across both switch and non-switch trials.

      We used several different analyses to compare and contrast the neural-behavioral and pupil-behavioral relationships because they provide complementary and useful insights. The correlational analyses in Figures 5C and 6C characterize session-level relationships between neural/pupil signals and behavior. The group comparisons in Figures 5D–E provide a complementary visualization of the same relationship. The model-based approach in Figure 7 then allows direct quantification of the trial-wise contributions of each signal to behavior within a common framework. Importantly, the conclusions drawn from each approach converge on the same interpretation, which we believe speaks to the robustness of the reported effects.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 2 legend. Description of 'running average (5-trial window)' is unclear - presumably this is a running average in stimulus space rather than across trials.

      We thank the reviewer for flagging this ambiguity. We have updated the legend (L136-137) to clarify that the running average is computed across trials sorted by test-stimulus duration.

      (2) L158. Difficult to establish an asymptotic performance level for HSF conditions within the stimulus duration range tested.

      We have removed the reference to asymptotic performance and replaced it with a discussion of performance on longer-duration switch trials in the context of the newly added Figure 2D.

      (3) L515 Equation 1. While this is a standard formulation of lapse rate in psychometric functions, the construction here in terms of switch probability is not standard. Given the task and training, it seems more likely that on lapse trials, the animal will respond according to the last adapted direction (rather than randomly switch/stay with equal probability).

      We thank the reviewer for this point. We agree that it is possible that on at least some of the “lapse” trials the monkeys may respond according to the final adapting-stimulus direction rather than choosing randomly. However, we cannot distinguish those alternatives using this task design. We include a statement to this effect in Methods (L569-571).

      To explore the idea further, we refit the behavioral data using separate upper and lower asymptotes corresponding to lapse rates on switch and non-switch trials, respectively. Across monkeys, there were no significant differences between upper and lower lapse rates for either low (Wilcoxon signed-rank test for equal medians: p = 0.15, Cohen's d = -0.13) or high switchfrequency (p = 0.07, Cohen's d = -0.16) conditions. So, at the very least, there was no evidence for lapse-like errors driven by switch- (or non-switch-) specific defaults to the final adapting direction.

      (4) L256. Statistical significance of attenuation is not directly tested here.

      We have replaced "were attenuated" with "we did not identify any reliable context-stability differences" (L297) to accurately reflect what was directly tested without implying a statistical comparison between groups of sessions that was not performed.

      (5) L429. Does the increase in explanatory power warrant the increased complexity of the model here?

      We thank the reviewer for raising this important point. We used Tjur's pseudo-R<sup>2</sup> because it does not increase by default with added model complexity, making it more conservative than other R<sup>2</sup> measures in this respect. Tjur's pseudo-R<sup>2</sup> is a coefficient of discrimination, and as such its value increases only when additional terms improve the model's ability to separate predicted probabilities across response outcomes. Thus, the observed increases in explanatory power when adding neural or pupil terms reflect real improvements in discriminability rather than an artifact of model complexity. We have added a brief clarification of this point to the Methods (L662-664).

      Reviewer #3 (Public review):

      The task design may not be optimal. While the amount of time the monkey is exposed to each motion direction during the adapting stimulus is matched, it's hard to know if the reduced MT responses to the test stimulus are truly due to the greater frequency of switches during the HSF adapting stimulus or because the monkeys have been exposed to more repetitions of the stimulus. It's increased sensory adaptation in either case, but it makes it problematic to interpret this as temporal context-dependent adaptation specifically. I think this could potentially be partially addressed by an analysis that is in the paper, but could potentially be emphasized/fleshed out more, specifically the results shown in Figure 4D that seem to show that most of the reduction in neural response for adapting units occurs between the first and second stimuli.

      The reviewer raises an important point. The number of stimulus repetitions and switch frequency are confounded in the experimental design, making it difficult to attribute context-dependent differences in MT responses to the temporal pattern of switches rather than to accumulated repetitions. We also note, as the reviewer acknowledges, the observed differences reflect sensory adaptation either way. Figure 4D does offer relevant evidence, suggesting that a majority of the change in neural response occurred with just one stimulus repetition. This finding complicates an interpretation where adaptation scales with the number of stimulus repetitions. We have added several lines to the Results about these points (L231-233).

      The pupillometric analysis seems to be an indirect way of assessing whether the accumulator itself might be modulated by temporal context, but the link could be made clearer. The authors show that context-dependent behavior is related to pupil size, which is related to arousal/neuromodulation, but it would be helpful to have some idea of what neural mechanisms underlying adaptive decision-making are actually impacted by this neuromodulation. Lacking neural data to address this question (e.g., from a brain region proposed to be involved in the accumulation process), at least more discussion of this would be helpful. Essentially, I'm unsure of how to interpret the pupil results: the argument that temporal context affects instantaneous evidence encoding in MT that then drives the accumulator is very clear, but I am a bit confused about what, mechanistically, I should think about the effect of neuromodulation doing.

      We thank the reviewer for this thoughtful comment and agree that the mechanistic interpretation of the pupil results could be made clearer. We acknowledge that we cannot directly identify the neural mechanisms underlying the arousal-related contributions to adaptive evidence accumulation from pupil data alone, given that pupil size is an indirect and imperfect proxy for neural (e.g., LC-NE system) activity. However, we can offer some informed conjecture and have added to the Discussion (L469-482) in an effort to elaborate on possible mechanisms.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract could be retooled - does not emphasize the pupillometry/arousal results very much, and they are presented more as a control than an independent result.

      We agree and have revised the Abstract accordingly.

      (2) Do all neural/pupil analyses use only switch trials? Sometimes the figure captions do specify only switch trials, but not everywhere. It would be helpful to specify either in the Methods or at the beginning of each figure caption that all subplots show switch trial results. Also, if you do always use switch trials, it would be useful to see in the Supplement how the non-switch trial results differ from switch trials. It seems like they may in interesting ways based on the behavioral results (supporting a reset of evidence accumulation on switch but not non-switch trials).

      We thank the reviewer for flagging these important points. We have added a justification for switch trials (L186-190) as well as clarification about which trial types were used for which analyses (L246-249) and information about trial types to relevant figure captions. We have also added a new Extended Data figure (Extended Data Fig. 9) examining relationships between neural activity and behavior on non-switch trials. As inferred by the reviewer, behavior on non-switch trials is consistent with the use of information from the adapting stimulus.

      (3) In Figure 3C, 5B, etc, when computing firing rate for the test stimulus (50-500 ms), are differently sized windows used to compute the rate for different test stimulus durations (since some will be <500 ms)? Or are only trials where the test stimulus duration is > 500 ms used for this analysis?

      We thank the reviewer for raising this point. To clarify, the 50–500 ms window does not reflect a fixed window applicable for all trials. Rather, neural activity from 50 ms after test-stimulus onset through test-stimulus offset was included for each trial, with 500 ms serving as the upper bound for trials with longer durations (> 500 ms). We have clarified this in the Methods (L607-610) to avoid ambiguity.

      (4) I think it might be better to be consistent with the time windows used for analysis; specifically, to choose either the 50-500 ms window used in Figures 3, 4, and 5B, or the 200- 400 ms window used for the remaining analyses in Figure 5.

      We agree that using the same window for all of the analyses would improve consistency, but not doing so provides advantages that we believe take precedent and now describe in more detail. The broader 50–500 ms window used for Figures 3, 4, and 5B was chosen to characterize MT neural activity over a relatively large a time window, ensuring that every trial contributes to each estimate. Because test-stimulus durations were drawn from a truncated exponential distribution (100–1200 ms), restricting these analyses to the 200–400 ms window would have excluded the substantial proportion of trials with durations <200 ms (but would yield similar figures and conclusions). The narrower window used in subsequent analyses allows us to focus on the conditions that exhibited the biggest modulations of neural activity when comparing them to behavior.

      (5) Similarly, provide justification for using only trials ending 375-600 ms after test stimulus onset for the behavioral correlations. It seems reasonable to choose a subset of test stimulus durations where the monkeys' behavior is greater than chance but less than ceiling, but it would be good to specify this so that it doesn't seem arbitrary.

      We agree and have added text to make this important point (L249-251).

    1. eLife Assessment

      By investigating spine nanostructure and dynamics across multiple genetic mouse models for neurodevelopmental disorders, this important study has the potential to uncover convergent or divergent synaptic phenotypes that may be specifically associated with autism versus schizophrenia risk. The imaging and overall breadth of the methods are convincing. The purely in vitro nature of the study slightly limits the generalisability of the findings, though these limitations are acknowledged and discussed in the manuscript.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Kashiwagi et al. undertook a population analysis of dendritic spine nanostructure applied to the objective grouping of 8 mouse models of neuropsychiatric disorders. They report that spine morphology in cultured hippocampal neurons shows a higher similarity among schizophrenia mouse models (compared with autism spectrum disorder (ASD) mouse models) and identify an effect of Ecrg4 (encoding small secretory peptides) on spine dynamics and shape in these models.

      Strengths:

      The study developed a method for objectively comparing spine properties in primary hippocampal neuron cultures from 8 mouse models of psychiatric disorders at the population level using high-resolution structured illumination microscopy (SIM) imaging. This novel technique identified two distinct groups of mouse models according to the population-level spine properties: those with ASD-related gene mutations and those with schizophrenia-related gene mutations. Functional studies, including gene knockdown and overexpression experiments, identified an effect of Ecrg4 on the spine phenotype of the schizophrenia model mice.

      Weaknesses:

      The main weakness is that the study is wholly in vitro, using cultured hippocampal neurons. The authors present this as an advantage, however, arguing that spine morphology as measured in a reduced culture system can demonstrate direct effects of gene mutations on neuronal phenotypes in the absence of indirect influences from nonneuronal cells or specific environments.

    3. Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for timelapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. My major concerns from the initial manuscript, especially regarding cherry picking and circularity have been addressed with revised analytical approaches. I have some remaining minor comments.

      (1) The comparison between two wild-type samples versus wild-type-mutant samples is helpful - I think this could be added to the manuscript.

      As suggested, we added the figure comparing two wild-type samples against wild-type mutant samples as Supplementary Figure 2. 

      (2) For results of time-lapse imaging - please spell out in the results section the direction of change (lines 270 - 277).

      As suggested, we added the direction of change (an increase in the turnover rate) to the text (page 12, lines 270-271).

      (3) Using linear mixed effect models for statistical analysis is a significant improvement. While a sample size (n) of mice = 3 is not ideal, I think given the multiple different mouse lines used and intensity of analysis, this is probably the best that can be done, although further validation in larger samples eventually is to be hoped for.

      We appreciate the reviewer for recognizing the effort required to collect data across multiple mouse lines.

      (4) The revised text is much improved, but I still think the authors should be upfront somewhere in the text that the schizophrenia-associated genes can only confer biased risk for schizophrenia (and that the clinical phenotype can also include autism). As I said before, I think this is the best we can do and I agree with their choices, but it is important not to overstate the link. The differences they see make it clear that these are still relevant distinctions.

      As suggested by the reviewer, we further modified the discussion related to the comparison between ASD- and schizophrenia-associated mouse models (pages 23-24, lines 508-522).

      “The nanoscale features of dendritic spines in mouse models of Nlgn3<sup>R451C/(y or R451C)</sup>, Syngap1<sup>+/−</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, which we classified as being related to ASD, are highly heterogeneous. This heterogeneity may reflect the broad clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. Accordingly, these four mouse models may represent distinct subgroups characterized by different degrees or forms of hippocampal dysfunction. Notably, among the ASD-related models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those found in the 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> mouse models. Although we classified 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> as schizophrenia-related models, both 22q11.2 deletion syndrome and Setd1a haploinsufficiency in humans are also associated with ASD, suggesting substantial overlap in the genetic risk factors underlying ASD and schizophrenia. Further systematic analyses linking rare genetic variants to synaptic phenotypes in mouse models may provide important insights into the mechanisms underlying both shared and disorder-specific synaptic alterations in neurodevelopmental and psychiatric disorders.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I would suggest that it might be preferable to use the word 'neuropsychiatric' rather than 'mental' in the title.

      As suggested, we modified the manuscript title.

      (2) I think it would be clearer to say that DEGs are listed if present 'in three or more models' rather than >2 (I appreciate the latter is mathematically clear, but can easily be read as 2 or more if reading fast). This is changed in the figure legend, but I suggest it is also changed in the main text (line 352-3)

      As suggested, we changed the main text to incorporate "in three or more models" (page 16, line 352).

      (3) Please add to Methods (line 557) that 'control cultures were prepared from littermate embryos....'

      As suggested, we added the phrase "control cultures were prepared from littermate embryos" (page 26, line 559).

      (4) Sorry to add something, but please could the authors add a definition of how they calculate spine turnover (and add units to the y axis of Figure 5A-C)?

      As suggested, we modified the y-axis of Figure 5A-C (% as unit) and added the method of calculating spine turnover rate in the text (page 36, lines 808-811).

    1. eLife Assessment

      This useful study combines experiments and mathematical modeling to show that antibiotic protection provided by resistant cells can extend across both surface-associated and freely growing bacterial populations. Notably, they show that treatment efficacy depends on population composition and density. The evidence supporting the main conclusions is incomplete, primarily because the biofilm context is not adequately characterized and demonstrated, raising the concern that it might represent only an aggregate of cells on the surface (rather than a biofilm) under the studied experimental conditions.

    2. Reviewer #1 (Public review):

      Summary:

      This important study examines how antibiotic-resistant bacterial cells can protect neighboring sensitive cells in mixed populations that occupy both surface-associated and freely growing states. Using experiments in Enterococcus faecalis together with a mathematical model, the authors test the hypothesis that protection would be stronger in biofilm-associated populations, but instead find that resistance-mediated protection extends broadly across both population types. The work provides evidence that antibiotic efficacy depends strongly on community composition, population density, and density-dependent detoxification dynamics.

      Strengths:

      A major strength of the study is the close integration of experimental measurements with a relatively simple quantitative model that captures many of the observed population dynamics. In particular, the work highlights how interactions between antibiotic detoxification, cellular growth, and saturation at carrying capacity can generate nonintuitive behavior, including the reported population inversion effect. The agreement between the well-mixed model and the experimental observations is convincing, and the spatial analyses suggest that cells within the biofilm are sufficiently intermixed that large-scale spatial segregation is unlikely to dominate the observed behavior.

      Weaknesses:

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The study also raises interesting questions regarding the role of spatial structure and exchange between planktonic and biofilm-associated populations. It would be informative to explore whether biofilm-specific protection becomes more pronounced at lower antibiotic concentrations, where local detoxification may compete more directly with antibiotic penetration into the biofilm, and in this context, the dynamics of exchange between biofilm and planktonic populations would be interesting to understand. Overall, the evidence supporting the central conclusions is convincing, and the study will likely be of broad interest to researchers studying microbial communities, antibiotic resistance, and collective population dynamics.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Martins et al. examined the cooperative response of E. faecalis cells to beta-lactams, in both planktonic culture and in biofilm. They found that the competition outcome between the susceptible and resistant strains is frequency dependent; they have also quantified how the competition curves change with inoculation OD and antibiotic concentration. To the authors' surprise, the competition dynamics are not that different in biofilm and in planktonic culture, which the author attributed to the unstructured nature of the thus-grown E. faecalis biofilms, quantified through correlation analysis. Using a well-mixed model capturing growth, death, and drug degradation by the resistant cells, the authors were able to quantitatively capture the experimental observation.

      Strengths:

      Overall, the data presented are solid. Although there is not much surprise after the understanding that the E. faecalis biofilm is unstructured, the manuscript still provides a useful "null case", so to speak, for researchers in the field when considering antibiotics in the context of biofilm. The theoretical model presented and the procedure of fitting the experimental data are useful to the research community.

      Weaknesses:

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

    4. Reviewer #3 (Public review):

      Summary:

      The authors studied social aspects of antibiotic resistance by co-cultivating antibiotic-resistant and sensitive Enterococcus faecalis (an important pathogen) as biofilms to assess the extent to which sensitive cells can take advantage of the protection provided by resistant cells against both a beta-lactam antibiotic and in the presence of a B-lacatamase inhibitor. By quantifying the proportion of each cell type using fluorescence microscopy, they conclude that protection is provided equally in the biofilm and planktonically, and that the biofilm is completely unstructured with regard to the locations of the two cell types. A mathematical model is then used to show that no spatial information is needed to recapitulate the results and that the protective effect can be described completely by the growth rates of the two cell types and the affinity of the β-lactamase to the antibiotic and inhibitor. The strength of evidence is difficult to assess due to unclear descriptions of some methods, and the significance of the findings is limited by the experimental setup, where antibiotics were added very close to the time of inoculation.

      Strengths:

      The co-cultivation of antibiotic-resistant and sensitive bacteria allows for exploration of the social aspects of antibiotic resistance. Fluorescently-tagged strains allow for unambiguous tracking of the two cell types. The simultaneous analysis of biofilm and planktonic cells enables insight into whether these different growth modalities are influenced by social aspects of antibiotic resistance. In analyzing the structure of the biofilm, the use of a null model with randomized cell positions allows for an accurate determination of whether the observed data are due to some effect; however, as noted below, there is a caveat to this analysis. The broad observation that biofilm and planktonic populations are linked is generally supported by the data; however, this result is closely tied to the experimental setup used. The development of a mathematical model that can recapitulate results from a second set of data with values obtained from fitting a different set of data shows robustness of the model for using it to explain the results.

      Weaknesses:

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. The described 'population inversion' effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed. Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account. The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case. The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

    5. Author response:

      We would like to thank the editors for their interest in our work and the three referees for their time and careful reading of the manuscript. The reviewers have provided a series of helpful suggestions that we discuss in this provisional reply and will seek to address in the revised version of the manuscript.

      The main concern raised is that the bacterial community we refer to as a biofilm may instead correspond to a cell aggregate. Following the passing of Prof. Kevin Wood, in whose lab the experimental work was carried out, our ability to perform additional experiments is limited. Nevertheless, we plan to wash and fluorescently stain the extracellular matrix before imaging to measure the extent to which the observed bacterial community is an attached biofilm. In the meantime, we would like to highlight the work of Wen Yu et al. [1], in which E. faecalis biofilms were grown in 96-well plates under antibiotic stress. In particular, one of the strains of E. faecalis used in this article was OG1RF, the same strain used in our study. Crystal violet staining was used to quantify biofilm biomass, providing evidence for biofilm formation under those conditions. While we recognize that the experimental setup differs from ours and that the OG1RF sample used did not contain fluorescent and resistance plasmids, these results nevertheless support the expectation that OG1RF will readily form biofilms.

      Reviewer #1 (Public review):

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The timescale of drug degradation is an important system metric. We appreciate the referee’s suggestion to quantify antibiotic concentration over time. We plan to perform experiments in which samples are collected from the culture at fixed time intervals. After removing the bacteria from the samples via centrifugation, serial dilutions of the supernatant will then be spotted on a lawn of sensitive cells to measure the antibiotic efficacy at each time point.

      Reviewer #2 (Public review):

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

      We agree with the referee that additional evidence would strengthen our study. As noted above, we will perform additional experiments to demonstrate the presence of an attached biofilm.

      Reviewer #3 (Public review):

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. [...] The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

      The experiment was designed to address how coupled planktonic and biofilm populations develop in the presence of antibiotics, which we will more explicitly discuss in the revised manuscript. We do agree that investigating how mature biofilms and their planktonic populations respond to antibiotic stress is an exciting direction for future studies. However, we believe that is beyond the scope of the our study on the development of coupled populations. We will be sure to explicitly identify this limitation in our revisions.

      The described ‘population inversion’ effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed.

      The reviewer is correct that this ‘population inversion’ is a frequency-dependent (perhaps also density-dependent) effect, and we should have situated it within that broader ecological framework. We will use this terminology in our revisions. We do want to acknowledge that the late Dr. Kevin Wood was fond of this phrasing to describe the reversal of the dominant strain, which is not necessarily true for frequency-dependent effects. Although we do not know for certain, we suspect this was a play on the ‘population inversion’ term used in quantum physics, used to describe a system in which its excited state (high energy) population unexpectedly outnumbers its ground state (low energy) population.

      The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case.

      If the reviewer is indeed discussing Figures 1a and 1b, these are not comparable as the starting fractions differ. On the other hand, if the reviewer was talking about Figures 2a and 2b (which is more clearly discussed by looking at Figures 2c and 2f), we agree that it appears that planktonic communities tend to have a slightly greater frequency of sensitive cells than the biofilms. We will be sure to highlight this observation and possible explanations in our revisions. However, given the uncertainty in these observations, we do not believe the differences are sufficient to alter our overall conclusion that final resistant fractions in biofilm and planktonic populations are quantitatively similar. Furthermore, the no drug treatment shows the same trend, which suggests it’s an effect of different growth dynamics of these two strains at high density rather than driven by the protective effects of resistance cells.

      Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account.

      The entirety of the Z stack was used to measure the final resistant fraction in the biofilm. On the other hand, we used only the densest slice of the Z stack to calculate the correlations. The correlations follow the same trend when calculated over less dense slices, but as density decreases, noise increases, so such plots did not bring more clarity to our conclusions and were not included in the manuscript. Additionally, only horizontal correlations (over a slice) were calculated because consecutive Z-stack slices were imaged with a 2.5 µm spacing. Given that the average cell diameter is approximately 1 µm, calculating vertical correlations may miss neighboring cells located between imaged slices, making such measurements unreliable. We will clarify the points raised by the reviewer in the results section and add more detail to the imaging methods section in the revised manuscript.

      References

      (1) Wen Yu, Kelsey M. Hallinen, and Kevin B. Wood. “Interplay between Antibiotic Efficacy and Drug-Induced Lysis Underlies Enhanced Biofilm Formation at Subinhibitory Drug Concentrations”. In: Antimicrobial Agents and Chemotherapy 62.1 (Dec. 2017), 10.1128/aac.01603–17. doi: 10.1128/aac.01603-17. url: https://journals.asm.org/doi/10.1128/aac.0160317 (visited on 01/11/2026).

    1. eLife Assessment

      This important study provides the first in vivo evidence that nonsense-mediated mRNA decay (NMD) in mature astrocytes regulates astrocyte function, synaptic plasticity, and anxiety-related behavior. Using a broad range of approaches, the authors show that conditional deletion of Upf2 alters astrocyte morphology and calcium signaling while impairing synaptic transmission and plasticity, providing solid support for the central conclusion that astrocytic NMD influences neural circuit function. Some key mechanistic claims remain incompletely supported, including whether phenotypes reflect astrocyte remodeling versus loss, the interpretation of synaptic engulfment data, the link between NMD targets and calcium signaling, and the extent to which calcium dysregulation explains the observed synaptic and behavioral effects.

    2. Reviewer #1 (Public review):

      Summary:

      Lituma and colleagues investigate the role of NMD in astrocytes, an underexplored question given that prior work on NMD in the brain has focused exclusively on neurons. Using a tamoxifen-inducible, astrocyte-specific Upf2 conditional knockout (cKO) mouse, they report that loss of astrocytic NMD causes: (1) reductions in astrocyte cell volume and surface area across hippocampus, visual cortex, and prefrontal cortex; (2) decreased excitatory synapse density, reduced dendritic spine density, and impaired synaptic engulfment; (3) deficits in basal synaptic transmission and LTP, with selective impairment of mGluR-LTD; (4) elevated spontaneous calcium transients in astrocytes; and (5) anxiety-like behavior in the elevated plus maze (EPM) and contextual fear conditioning paradigms. Transcriptomic analysis of FACS-isolated astrocytes identifies 277 differentially expressed genes, ~40% of which carry canonical NMD-inducing features, implicating pathways linked to calcium signaling, phagosome formation, and glial development. A rescue experiment using the CalEx calcium extrusion pump demonstrates partial restoration of synaptic strength and anxiety behavior when astrocytic calcium is normalized.

      The study addresses an important gap in our understanding of RNA regulation in glial cells, and the overall conceptual framework is well described. The experimental design is generally appropriate, and the multi-pronged approach lends the main claims a degree of validity.

      Strengths:

      (1) Novelty: This is the first study to systematically examine NMD function in astrocytes in vivo. The identification of astrocytic NMD targets via RNA-seq combined with an NMD-inducing feature classifier is a meaningful methodological contribution.

      (2) Multi-method approach: The authors combine morphological analysis (Imaris 3D reconstruction), synaptic markers (PSD-95, LAMP2 engulfment assay), spine density measurements, acute slice electrophysiology, two-photon calcium imaging, behavioral testing, and transcriptomics. The convergence across these methods strengthens confidence in the claims.

      Weaknesses:

      (1) While the transcriptomic analysis is a valuable addition, the connection between specific NMD targets and the observed calcium phenotype remains largely correlational. The authors identify Gabbr2 and Adora1 as upregulated candidates with canonical NMD features and speculate that their elevated expression drives aberrant calcium signaling. However, no validation (e.g., qRT-PCR or protein-level confirmation) of these candidates is presented. The mechanistic pathway between NMD disruption and elevated calcium is thus inferred from pathway analysis rather than demonstrated. This is a significant gap between the transcriptomic and physiological arms of the study, and the authors should be more explicit about this limitation or, ideally, provide at least one validated target.

      (2) The reduction in astrocyte surface area in cKO mice is interpreted as contributing to reduced synapse contact and engulfment capacity. This is a reasonable hypothesis, but the study does not directly demonstrate that reduced astrocyte territory correlates with reduced synaptic coverage at the level of individual cells or brain regions. The temporal sequence of these events is unknown. Do morphological deficits precede synaptic changes? Clarification and qualification of this causal chain in the Discussion would strengthen the manuscript.

      (3) LFS-induced LTD is unaffected, while mGluR-LTD is reduced. This is intriguing and potentially informative about astrocyte contributions to distinct LTD mechanisms, but the difference receives limited discussion. Given the relevance of mGluR signaling to calcium dynamics and the identified pathway enrichments (GPCR signaling), this specificity deserves more attention.

      (4) The CTRL + CalEx condition is included in the EPM experiment but not in the electrophysiology or calcium imaging experiments, making it difficult to fully assess whether CalEx itself has off-target effects on synaptic transmission or anxiety in wild-type animals. The CTRL + CalEx EPM data (Figure 7F) appears to show a modest reduction in open arm time relative to CTRL, which, if robust, would suggest that excessive calcium reduction in astrocytes is also anxiogenic. This finding would be physiologically relevant and deserves comment.

    3. Reviewer #2 (Public review):

      Astrocytes are highly responsive to their environment and play a range of critical roles in brain function. Lituma et al. theorize that one mediator of that responsiveness is the regulation of RNA stability. They therefore undertake an assessment of astrocytes missing Upf2, a protein required for mRNA degradation via nonsense-mediated decay. This is an interesting study, approaching astrocyte biology from a novel angle. The authors take on an ambitious set of experiments, spanning morphological assessment, synaptic engulfment, electrophysiology, behavior, and calcium imaging.

      The authors show convincing data that knocking out Upf2 in astrocytes impairs synaptic plasticity, affects behavior, and changes the complement of astrocytic mRNA. These results, in and of themselves, are intriguing and suggest that NMD is an important biological process in astrocytes, warranting further study.

      My primary concern is whether the authors may be largely studying dying cells. The idea that NMD disruption has a dramatic effect on astrocyte morphology is an intriguing idea, but it is not fully established here. The nuclei in the example cKO morphology images appear small and/or fragmented. This raises concerns that the authors did not ensure that they had the full 3D morphology of the astrocyte in the section, and the cell is in part cut off, which would compromise any data on the morphology. The authors state that the tissue was sectioned at 70 um. The diameter of an astrocyte in the adult mouse brain is typically between 50 and 70 um. Unless astrocytes are perfectly positioned in the center of the slice, at this thickness, the majority of astrocytes will almost certainly be partially cut off. More detail on how cells were chosen and what quality control metrics were implemented would alleviate concerns here. An alternative possible explanation for these small/fragmented nuclei is that cKO astrocytes may be unhealthy to the point that they are actively dying. Using the transgenic ZsGreen label, the authors state that they observe a size change (Figure S4); this is not readily apparent and is not quantified in any way. It does appear from these images that there may be a loss of some astrocytes; cell death, which would also be an interesting finding, is a fundamentally different process than morphologic restructuring in living cells. The authors do attempt to count astrocytes (Figure S6B), but do so with GFAP. This is a fundamentally flawed approach. Because GFAP is not readily detectable in most healthy astrocytes in most gray matter regions, GFAP should not be used to quantify astrocyte numbers; this experiment should be repeated with a better marker, such as Aldh1l1, Sox9, etc.

      Synaptic engulfment: This is an extraordinarily high degree of engulfment in the control animals compared to many published studies, leading to concern as to the technical approach. Indeed, the overall low level of PSD-95 signal in control conditions in adult mice is concerning as to the technical accuracy of the approach. It is unclear exactly how the investigators labeled the astrocytes; presumably via the ZsGreen label, but it is never stated, and the only images shown are the highly processed Imaris renderings. The small astrocytic processes, or leaflets, that make up the vast majority of the astrocytic arbor are on the order of 100nm in diameter. The processes shown in Figure 2B are, according to the scale bar, at least 20x that size. It is difficult to have much faith in these results as currently presented.

      The signal-to-noise ratio of the GCaMP experiments is worryingly low, likely responsible for the abnormally low dF/F in all conditions and the lack of significant change between control and CalEx, when control astrocytes should show a much higher GCaMP signal than any CalEx-expressing astrocyte. That said, the higher Ca++ in Upf2 KO astrocytes is intriguing. Given the roles of elevated calcium in cell death, this may reflect cells that are unhealthy to the point that they are starting to die.

      The authors conduct a FACS-based analysis of astrocytic mRNA from control vs Upf2-KO, with intriguing results. An important caveat, though, is that a large amount of astrocytic mRNA is in the processes. If mRNA stability is being actively and rapidly regulated, it seems likely that the mRNA in the processes would be the most relevant population of regulated mRNA. FACS-based approaches to astrocyte purification will, as robustly shown elsewhere, strip off those processes. Particularly given that the authors have shown that the processes may be the most actively changing astrocytic compartment with Upf2 KO, this is a strange choice of technique vs. something like Ribotag that would preserve the mRNA in processes. At least, there should be some discussion regarding using FACS for this analysis and the consequences for profiling mRNA in astrocytic processes.

      Minor points:

      (1) The use of the Aldh1l1-CreER mouse is a strong choice and has been shown to be highly astrocyte-specific. Combining that transgenic mouse with viruses driven by different forms of the GFAP promoter is quite bizarre in several ways. First, GFAP-dependent AAVs have been shown repeatedly to have significant neuronal leak. Second, these mice are, in all cases, receiving two different viruses, driven by different forms of the GFAP promoter, and the non-Cre virus is not Cre-dependent (vs. a much more standard approach of using a Cre-dependent second virus to ensure that all analyzed cells received both viruses). The authors mention that "this experimental design ensures that phenotypes are not caused by an acute effect of tamoxifen." It is certainly true that tamoxifen is not a biologically neutral molecule. However, the mice still receive tamoxifen, both in these morphology virus experiments and in almost all other experiments. This experimental approach is not inherently bad, nor does it necessarily invalidate the data (although the near-certain neuronal contamination due to the GFAP promoter-driven viruses is a concern). It is, however, convoluted in ways that appear unnecessary. If there is a strong rationale for this approach beyond the tepid explanation already present, it should be explicitly mentioned.

      (2) The characterization of the knockout is incomplete. While the authors should be applauded for their attempts to phenotype the cells in which they observe Cre-mediated recombination, there are issues with their technical approach. Most importantly, and an issue that affects other analyses in the paper as well: the vast majority of astrocytes in the healthy cortex do not express GFAP. Therefore, using GFAP to claim high astrocyte specificity and efficiency is a fundamentally flawed approach. Second, MBP is a myelin marker, not a cytoplasmic marker, and would not successfully colocalize with a cytoplasmic marker like ZsGreen even if recombination in oligodendrocytes did occur. Third, recombination at one set of LoxP sites is not a reliable indicator of recombination at other sites. Recombination efficiency is highly dependent on the spacing between the LoxP sites and cannot be reliably extrapolated to other floxed genes without validation. Finally, the most likely culprit for off-target recombination with Aldh1l1-CreERT2 (or other astrocyte-selective Cres, and certainly the GFAP-based viral promoters) is neurons, which the investigators did not test for. Neuronal Aldh1l1-CreERT2 leak is most likely to occur in the hippocampus. With the images shown in Fig S3, it is unclear whether it is possible to convincingly colocalize Upf2 staining with a cytosolic marker of all astrocytes, such as Aldh1l1 or S100b, but such data would be more appropriate. An alternative approach to validation would be in situ hybridization.

      (3) Supplementary Table 2 should include gene IDs, not just Ensemble IDs.

      (4) It is not fully clear what the investigators are denoting as a spine in Figure 2E; the two images do not appear to have the large degree of difference that the quantification suggests. The oversaturation of the signal complicates assessment.

      (4) A more detailed discussion of the rationale behind the timeline would be helpful. What is the half-life of Upf2, and how rapidly do NMD genes build up upon Upf2 disruption? In particular, in the case of virus experiments, the timeline is quite fast: ~2.5 weeks from injection to analysis. ssAAV expression takes over a week to reach appreciable levels.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigate mRNA targets of the nonsense-mediated decay (NMD) pathway in astrocytes and link the dysfunction of NMD in astrocytes to aberrant synaptic transmission that has downstream effects on behavior. Specifically, they find a link between the aberrant synaptic transmission with elevated spontaneous calcium signaling in astrocytes, and functionally they demonstrate that manipulating astrocyte calcium signaling with CalEx modulates astrocyte calcium signaling towards wildtype levels and improves anxiety behavior. They investigate the astrocyte calcium signaling changes in Upf2 conditional knockout mice in several brain regions that have been linked to anxiety behavior, including the hippocampus and prefrontal cortex. They also observe aberrant astrocyte calcium signaling in the visual cortex, demonstrating that dysfunction of the NMD pathway in astrocytes has widespread effects on synaptic transmission in various brain regions. This work identifies, through RNA-Sequencing, potential mRNA targets of NMD in astrocytes, and shows that pathway enrichment of these targets highlights calcium signaling. Altogether, this work highlights the importance of the basic cellular process of NMD in astrocytes, which are known to have extensive local translation of proteins in their perisynaptic processes. NMD may be particularly important in astrocytes due to their intimate association of processes with neuronal synapses, and the authors suggest that alterations to NMD function in astrocytes may be an important avenue for future investigation in neurodevelopmental disorders.

      Strengths:

      Altogether, this work is a critical foundation for future research into astrocyte contributions to neurodevelopmental disorders. The authors do a thorough characterization of astrocyte conditional Upf2 knockout mice in several brain regions. They present a complete story that connects molecular events (NMD pathway regulation of mRNA degradation) to astrocyte regulation of circuit activity to organismal behavior. The electrophysiological analysis is thorough, and the manipulation of calcium activity ties astrocyte calcium activity to anxiety behavior. The RNA-sequencing dataset is useful to the scientific community and provides a resource of candidate molecules that might be dysregulated in neurodevelopmental disorders.

      Weaknesses:

      The study suffers from some overstated claims and a lack of statistical rigor in some experiments, as detailed below.

      (1) The title states that "Astrocytic Nonsense-mediated mRNA decay regulates calcium signaling to support synapse function and restrain anxiety". The term "restrain anxiety" implies that the NMD pathway has a direct effect on a molecular switch to control anxiety. Anxiety behavior is a complicated process, controlled by many biological phenomena and synaptic transmission in the circuit as a whole, and is not directly linked to a specific NMD mRNA target. This title is overstating the findings of the study.

      (2) In general, the first figures (1-2) suffer from low power (N = 3) and statistical rigor. The statistics are inflated by analyzing individual fields of view and per-cell data rather than performing the statistics on the average of biological replicates. It is preferable to show the biological replicate data so that readers can observe the natural biological variability between replicates.

      (3) The claim that astrocytes have decreased engulfment of synapses in the Upf2 conditional knockout mice is not strongly substantiated by the data. The resolution of confocal microscopy and the static nature of histological images make it difficult to measure synaptic engulfment as an active process. Additionally, the metric of quantifying the % occupancy of PSD95 puncta within the total astrocyte volume may be skewed due to overall differences in cell size (shown in Figure 1). There is not much discussion of how a decrease in astrocyte engulfment of synapses may lead to decreased synapse number. To the contrary, one might expect decreased engulfment to result in increased synapse density.

      (4) The authors use Gfap as a marker to count astrocyte cell number and assess if there are changes in cell number between genotypes (Figure S6). However, Gfap does not label all astrocytes in the cortex and, in fact, is rather an aberrantly expressed marker in conditions of inflammation, as opposed to the hippocampus, where Gfap is basally expressed in all astrocytes. In the cortex, there seems to be a trend for reduced Gfap in the conditional knockout mice, which may suggest differences in astrocyte molecular signatures rather than cell numbers. Another astrocyte marker, like Aldh1L1, will be more accurate to assess this question histologically.

      (5) The authors state that "Preventing abnormally high basal calcium activity in NMD-deficient astrocytes restores normal excitatory synapse function...". However, this claim is not substantiated by the data. CalEx manipulation certainly shifts the input-output curve but does not restore to wildtype baseline levels (Figure 6E). Additionally, synapse number does not appear to be restored to wildtype levels (Figure 6D - although the p-value for this comparison is now shown). The investigators do observe improvements in anxiety phenotypes, suggesting there is some modulation of circuit activity, but the claim that CalEx manipulation restores baseline synaptic transmission is not supported.

    5. Author response:

      We thank the reviewers for their careful reading of our manuscript and for providing positive, constructive feedback. In particular, we thank the reviewers highlighting the several strengths of our study.

      To address the reviewers’ major concerns, we will revise the presentation of our main findings (specifically data/animal vs data/ROI), provide more clarity in the Results, Methods, and Discussion sections, and modify the title to better reflect these nuances.

      Additionally, we will perform the following new experiments:

      (1) Astrocytic Marker Validation: To further confirm comparable astrocyte cell counts between the CTRL and Upf2-cKO conditions, we will perform immunostainings using Aldh1L1, Sox9, or S100b instead of GFAP.

      (2) NMD Candidate Validation: To validate top candidate NMD target transcripts, we will perform immunostainings or qRT-PCR for Gabbr2, Adora1, S100b, or Cldn9.

      (3) Sample Size Expansion: To strengthen the morphological and PSD-95 quantifications, we will increase the sample size (N) by incorporating additional animals.

      (4) Mechanistic Timeline & Phenotype Linkage: We value the reviewer’s comment regarding the timeline of morphological and Ca<sup>2+</sup> phenotypes. To gai insight into whether these phenotypes are independent or linked, we will perform 3D reconstructions in CalEx conditions to assess whether Ca<sup>2+</sup> restoration rescues astrocyte morphology in Upf2-cKO mice. This will allow us to determine if increased Ca<sup>2+</sup> activity is upstream of the morphological alterations. Taken together, we believe that incorporating these manuscript revisions will strengthen the clarity and conclusions of our work. We thank the reviewers for their time and careful evaluation of our study.

    1. eLife Assessment

      This is a valuable paper looking at nanoscale organization of the membrane associated periodic cytoskeleton in mouse sciatic nerve axons. Despite previous studies, the precise organisation of the structure remains unclear, especially in vivo, and this manuscript significantly adds to this knowledge with solid data. An unexpected observation is the presence of discrete nanoscale clusters, regularly distributed around sections of axons. However, the paper misses a description of these clusters along the longitudinal axis of the axons.

    2. Reviewer #1 (Public review):

      Summary:

      The article "Nanoscale organization of beta-II spectrin within segments of the membrane-associated periodic skeleton in mouse sciatic nerve axons" by Gazal et al. looks into the organization of the spectrin scaffold in mouse sciatic nerves using super-resolution microscopy. It is now well established that axons, across species, contain a membrane-associated periodic scaffold mainly composed of circumferential actin filaments and longitudinally arranged spectrin tetramers. While super-resolution imaging of neurons in cell culture is relatively easy, exploring the ultrastructure of myelinated axons in intact nerve fibers is a daunting task. Nevertheless, the authors have attempted this by fixing and preparing cross-sections of sciatic nerves. They have then tried to quantify the fluorescence intensity patterns of specific components, especially that of labeled beta-II spectrin and have analysed its distribution.

      One of the main findings is that spectrin is distributed along the axonal periphery and along the outer part of the myelin sheath. By labelling multiple cellular components and using intensity analysis, the authors show the sequence of structural organization of a few key components. They see that, unlike in the case of axons in culture, the axonal cross-sections within the sciatic nerve deviate significantly from a circular shape. They then use 3D-dSTORM to investigate the distribution of beta-II spectrin along the axonal circumference. They see that this distribution is very heterogeneous, both in the sizes of spectrin puncta and their arrangement along the periphery. The amount of spectrin scales linearly with axonal circumference.

      Strengths:

      Super-resolution imaging of axons of intact nerve fibers to investigate the organization of beta-II spectrin.

      Weaknesses:

      While most of the findings, like the spatial distribution of spectrin and related components, are reasonably well supported by data, I have concerns regarding the subsequent claims made in the article. The detection of axial periodicity based on the observation of a peak in the inter-tetramer spacing distribution is not very convincing, and a 3D representation (or a video of 3D reconstruction) would have been better. And so are the claims on characteristic spectrin spacing of 200 nm along the axonal circumference. A peak in the distribution does not imply a periodic arrangement.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting paper by the Unsain lab looking at the nanoscale organization of the membrane-associated periodic cytoskeleton in mouse sciatic nerve axons. The precise organization of the structure remains unclear, especially in vivo, and this manuscript significantly adds to our knowledge of this important structure. While some of the findings in the study are somewhat expected (though still valuable to see in an in vivo setting), an interesting observation is the presence of discrete nanoscale clusters that scale up with the size of the axon, which challenges previous assumptions.

      Strengths:

      Strong, convincing data; clever combination of imaging and analytical tools to make novel points; well written; excellent composition of figures.

      Weaknesses:

      (1) Figure 2A/3A: The large and small clusters of spectrin, as seen in cross sections, are unexpected and novel. The authors have done a clever job of combining imaging and analyses, but some things are still unclear. First, the authors should be consistent in their language when they talk about the spectrin clusters. Recommend precise language to define the small and large clusters when they first appear in the text, and then use the definitions consistently throughout the text. Second, based on the data shown, one does not get a clear idea of how the small and large clusters are organized along the longitudinal axis of the axon. In that context, are Figure 2B and C from imaging along the longitudinal axis? If not, it's unclear how the authors can conclude that the spectrin assemblies have a distance of ~170 nm along the linear axis. In general, a perceived limitation of this study is that while the authors have done a good job looking at cross sections, there is no information on the longitudinal distribution of spectrin in these axons. Looking at both cross- and longitudinal sections would also clarify details about the large spectrin clusters. For instance, are they small sausage-like structures, or long rods of spectrin running along the length of the axon? One assumes that all the analyses in Figures 3 and 4 are from the small clusters. Can the authors do a similar analyses of the large clusters? Finally, a schematic model showing both cross- and longitudinal- sections would make things clearer, but the authors would need to show the longitudinal data for that.

      (2) It is interesting to think that the larger spectrin accumulations may be similar to the condensate-like structures seen by Boyer et al., as the authors mention in the discussion. In that context, it is possible that these focal accumulations are local reservoirs of spectrin that are also seen in mature axons (indeed, these accumulations were also seen in mature axons in the Boyer et al. paper, and they also speculated that these accumulations may be local reservoirs). Can the authors check if actin/adducin is also present in these larger spectrin accumulations?

      (3) While talking about the nanoscale clusters, it is important to specify that the authors are talking about circumferential clusters. Though the writing is excellent, one still does not get the precise definition of "clusters" from just reading the abstract, and it would be good if the authors could work on that more (I recognize that this is not easy to do).

    4. Reviewer #3 (Public review):

      Summary:

      In the presented work, the authors investigate spectral staining in axons of the sciatic nerve, where the MPS has been detected before using STED microscopy. They employ 3D-dSTORM in tissue sections and analyze the data, measuring localization of clusters on the axon perimeter and the relative distribution of those. From these data the conclude that large gaps in spectrum localizations exist and that clusters around the axon exist that are spaced at 200nm.

      Major Comments:

      (1) The presented data are at times overinterpreted, and the discussion lacks a critical view of the data. For example, the statement "...Unlike previous suggestions from qualitative evidence in cultured neurons (REfs), βII‑spectrin distribution in MPS segments of peripheral nerves is discontinuous, with extensive stretches of the perimeter lacking βII‑spectrin." is quite strong, given it is based on immunofluorescence staining and dSTORM microscopy in tissue. Absence of evidence of staining is not evidence of absence.

      (2) The authors claim in the abstract that "The number of these clusters scales linearly with the axonal perimeter, maintaining a constant membrane occupancy of ~20% across varying axon diameters." Again, this is from a cut through an axon, while measuring the density of clusters on the perimeter. If they claim area occupancy, an area should be imaged, and the dots (clusters) should be measured in surface coverage in a 2D projection of the axonal surface.

      (3) In general, this reviewer suggests being a bit more moderate in statements such as: "These findings challenge simplified models of the MPS based on cultured systems and demonstrate that the MPS in peripheral nerves is composed of discrete structural units." These statements are bold from the relatively few measurements in a single method and a single viewpoint. Especially when considering that techniques such as dSTORM depend extremely highly on labeling density, and apparent clustering of localization is highly prone to misinterpretation. If the authors desire to make such statements, working with endogenously labeled protein would be warranted. The authors should at least hedge such statements.

      (4) If the authors want to make statements about general organization, why do they not compare adjacent cuts through the axon? If there are continuous spectrin filaments, the clusters should appear at the same site across repeated cuts through the axon.

      Besides this, this reviewer welcomes the effort that has been made to establish dSTORM in tissue sections and to investigate the MPS in native tissue.

    5. Author response:

      We sincerely thank the editors and reviewers for their overall positive assessment and constructive feedback on our manuscript detailing the nanoscale organisation of βII-spectrin of the membrane-associated periodic skeleton (MPS) in mouse sciatic nerve axons. Their perspective and comments will help refining the manuscript.

      A common comment by the reviewers relates to the description of the characteristic longitudinal periodicity of the MPS. We value these comments, which we believe are motivated by the fact that the longitudinal periodicity of the MPS is undoubtedly the most studied and prominent feature of the MPS in cultured neurons. However, the main goal of the present project was to describe how βII-spectrin is organised in the transverse axis of individual segments of the MPS in nerve tissue. This is why we utilised cross-sections of the sciatic nerve, hence achieving the best resolution possible in that plane, at the expense of the resolution in the axial axis. Furthermore, this study clearly shows that the transverse morphology of axons, and thus of the MPS, of neurons in the tissue is highly irregular, in comparison to cultured neurons. This imposes an extra challenge to observe correlated longitudinal structures when the observation length is limited, as in our studies. Nonetheless, to improve this aspect of the manuscript, we will revise our data and previous evidence, clarify the methodological trade-offs made, and make our interpretations more accurate.

      Additionally, we will clarify several imaging- and definition-related inquiries, including tests for insufficient staining, the interpretation of βIII-tubulin staining, the assessment of axon–glia boundaries, and consistency in the use of terms like ‘clusters’ and ‘periodicity’, among others.

      We believe these and other revisions will substantially strengthen the manuscript and comprehensively address the reviewers' feedback.

    1. eLife Assessment

      This valuable study shows the impact of the metabolic state of bacteria on phage infection. The experimental results, based on various phages infecting E. coli, are convincing and consistent with a two-step adsorption mathematical model. This study should be of interest to the communities working on cell metabolism and on host-pathogen interactions.

    2. Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. The mechanistic interpretation offered by the model highlights an interesting trade-off between adsorbing to a metabolically active host and discriminating between active and inactive hosts that, somehow, a phage has to optimize. It would be interesting, in the future, to investigate how different phages optimize this task.

      Comments on revised version.

      I am happy with how the authors have addressed all the comments. The manuscript is much clearer and more readable and the previous overstated claims have been removed/clarified.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phage that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well designed and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication and the paper shows that host physiology has to be taken into account at all steps of phage infection.

    4. Reviewer #3 (Public review):

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, 𝜙80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      As I have stated in the first submission of this manuscript, the data presented by the authors clearly supported the principal conclusion of the study. The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully use a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Comments on revised version.

      In this revised version, the authors have successfully resolved all of my comments. I appreciate that the main text has been majorly revamped, which greatly helps the readers follow the motivation behind the experiment and analyses, and interpret the data. I agree that the revised terminology choice "commitment to infection", instead of the previous interchangeably used "adsorption"/"entry", is much more logical, considering the experimental data. I also commend the authors for writing the modeling part in a very clear, pedagogical, and instructive manner. Overall, I believe that this manuscript will be valuable to those who are interested in phage-bacterial interactions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource-limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. While the model is sensible from a mechanistic perspective, the experimental evidence that supports how each model's rate is affected by the cell metabolic state is weak, as only ratios of these rates can be inferred from the data.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phages that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well-designed, and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication, and the paper shows that host physiology has to be taken into account at all steps of phage infection.

      Weaknesses:

      There are some concerns about the experimental setup and which conclusions can be drawn from it:

      Before phage infection, bacterial cultures are grown to exponential growth, washed, and then resuspended with glucose or arsenate-azide for 10min. It is however, questionable that 10 minutes is enough to simulate high and low metabolic states realistically. 10 minutes seems to be quite short to go from exponential growth to a low metabolic state, given the transcriptional memory of previous environments. It seems more likely that the population will be quite heterogeneous, with cells in various states of transition towards low metabolic states.

      While we agree with the reviewer that during metabolic transitions there may be a period in which the population is heterogeneous, with cells in different stages of transition toward a low metabolic state, the 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyper diffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027). We have also corrected the DOI for this reference in the manuscript. Furthermore, the ATP pool of log-phase E. coli turns over several times per second (Holms et al., Arch. Mikrobiol. 1972, http://dx.doi.org/10.1007/BF00425016). We therefore assumed the bacteria were energy depleted after 10 minutes. We have clarified this point in the revised manuscript.

      Given that arsenate and azide inhibit cellular metabolism, i.e., have antimicrobial effects, cells might not just downregulate metabolism but also activate the stress response, and this causes some of the observed effects on phage adsorption. Therefore, the 'low metabolic state' of the cells in this paper could mean that cells are starved or that they are stressed or both.

      The reviewer is correct. We don’t exclude indirect effects. However, as nutrients were removed from the bacteria by washing and energy metabolism was inhibited by the addition of arsenate and azide, we assumed a stress response requiring biosynthesis would be unlikely to occur.

      The abundance of receptors could change between the high and low metabolic media conditions and contribute to the observed differences in adsorption, while the authors seem to assume in their model that the initial adsorption rate always remains the same.

      We do not think that the observed differences in adsorption are explained by a change in receptor abundance. In a previous study using the same experimental protocol as in the present work, phage λ was compared to the metabolically insensitive mutant λh (Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119). If the lower adsorption in the low-metabolic condition were caused by a reduced number of receptors, then λh should also have shown a lower adsorption rate under the same condition. Instead, no measurable effect on λh adsorption rate was observed. We therefore conclude that the effect is not explained by changes in receptor number on the timescale of the experiment. We have clarified this point in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, ɸ80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy-depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      Strengths:

      The data presented by the authors clearly supported the principal conclusion of the study ("Viral commitment to infection depends on host metabolism"). The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully used a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Weaknesses:

      (1) The authors isolated and measured the numbers of free phages in the medium after infection of bacteria under different treatments. These measurements were analyzed in two different ways: (1) simply as ratios (corrected/normalized using different controls), and (2) fitted using a simple mathematical model. I have concerns regarding both analyses.

      (1.1) For the first method, having different time points at which the sample of each phage is collected critically complicates data interpretation. As one incubates the phage-bacteria mixture for a longer time, more infection occurs, and the number of phages collected from the mixture decreases. Therefore, the different incubation time forfeits the goal of "a systematic and quantitative comparison across different phages [...]", just as the authors self-criticized. Conceivably, the authors could have used the shortest measurement time for all phages (i.e., 10 minutes, as for phage λ). Alternatively, the authors could have applied a systematic criterion such as half (or any other fraction) of the latent period of each phage, which would still "maximize the incubation period while ensuring that manipulations were completed before the first infection cycle concluded". In my view, the seemingly arbitrary measurement time for each phage renders the entire first analysis very challenging to interpret. It also goes against the author's proposition that the protocol was "standardized" or "consistent". It is not clear what the readers are supposed to take away from this first analysis, or rather, which evidence, finding, or conclusion the manuscript would lose if the authors only presented the modeling-based analysis.

      (1.2) The second method of analysis sought to remove the dependence of the measurements on time. I completely agree with this goal, and the findings extracted from this analysis significantly contributed to the merits of this manuscript. However, the authors achieved this goal using a single time point for each phage to calculate the infection rate (η). As shown in Figure S3, each of the phage depletion curves is anchored by only one data point (note that the P(t)/P(0) = 1 at t = 0 is assumed, not measured). This goes against the typical way this collision model is used in the literature, where a time series is measured and used to fit the model (e.g., DOI 10.1007/978-1-60327-164-6 18, or more recently, PMID 39700139). This practice in the current manuscript reduced the robustness of the inferred η values. This problem is exacerbated by assumptions used by the authors in formulating this model. For instance, the authors used a constant value for the bacterial concentration, B, because "bacterial growth and lysis were negligible" (lines 135-136). However, considering that the bacteria were cultured at 37oC in a very rich medium (first in YT broth, then in 2% glucose), the measurement times of 20, 30, and 55 minutes are most likely one or a few generations of bacterial growth and division.

      Related note: I suggest that one of the panels in Figure S3 should be moved to the main text, since it is critical to the second method of analysis.

      We would like to clarify that the manuscript does not present two separate methods, but rather one method presented in two steps: a first step with results that are directly tied to the experimental measurements and show whether the effect is present for each phage, followed by a second, analytical step that makes the results comparable across phages.

      The first step presents the ratios because they directly reflect the measurements performed in the experiment and allow the reader to see the effect of the metabolic state for each phage in contrast to its control. We agree that these ratios are time-dependent and therefore not suitable for quantitative comparison between phages. Their purpose is to illustrate the experimental outcome and to show that the effect is present (or absent) on a per-phage basis not to compare magnitudes across phages.

      We then follow this with the second step, allowing the reader to follow the logic of the analysis. The analytical step that follows does not represent a second method, but a continuation of the same analysis. Here, we remove the time-dependence specifically in order to make comparison of the effect across phages possible, by connecting our results to standard measures such as the adsorption rate η. Importantly, P(0) is measured for every phage in every experiment. The only modeling assumption used (a standard one in the field) is the exponential form for the decay in free phage number, which naturally yields P(t)/P(0) = 1 at t = 0.

      Regarding the reviewer’s concern that bacterial growth may not have been negligible over the relevant time window, we note that recent work on rich-to-minimal growth lags in E. coli reports substantial delays before growth resumes after nutrient downshift. One 2023 study (Wu et al., Nature Microbiology 2023, https://doi.org/10.1038/s41564-022-01310-w) considering wild-type E. coli shows in Fig. 2c a lag of up to about 2 hours after a shift from MOPS minimal medium with 0.2% glucose plus 18 amino acids to the same medium without amino acids. Another 2023 study (Zhu and Dai, Nature Communications 2023, https://doi.org/10.1038/s41467-023-36254-0) examining both rel+ and rel− strains reports a growth lag of about 49 minutes for rel+ and more than 5 hours for the relA deletion strain. While these conditions are not identical to ours, they support the general point that growth does not immediately resume after such shifts. We therefore think it is unlikely that, following transfer from YT, the cells underwent one or a few full generations during the time window of our adsorption measurements.

      On the related note: Following the comments of all reviewers on Figure S3, we have decided to remove it to avoid confusion.

      (2) The data were able to distinguish phages that successfully infected bacteria and those that remained free in the medium, and the authors appropriately interpreted the data as such throughout the Results section. However, in the Discussion (starting from the very first sentence, line 172), the authors used terms that include "adsorption" and "entry" more interchangeably (for example, see the three sentences in lines 310-313, for "viral entry efficiency is shaped by [...]", then "adsorption kinetics modeling"). I do not see how the authors' data could distinguish between adsorption (the phage particles attaching to the outside of the cell) and entry (the phage DNA being injected into the cell). Conceivably, any phage particles that irreversibly attach to a cell but do not yet inject their genome into the cell would still be removed from the medium and therefore not quantified. Another example: in lines 189-191, the authors interpreted that "[...] when the bacterium is in a low metabolic state, the phage does not bind irreversibly to the host", but how do the authors eliminate the case of no phage binding (i.e., the reversible step) to begin with?

      We agree with the reviewer that our use of the terms adsorption, entry, and infection should have been more careful. Our experiment can only identify the irreversible commitment of phage to a host cell. We have therefore revised the text to refer consistently to phage commitment.

      Similarly, in lines 283-293, how do the authors delineate whether energy depletion would increase the k_off term or decrease the k_inj term, because either would result in more free phages in the medium as observed in the data? I believe that the writing of the Discussion, as it stands now, is doing a disservice to the conclusions presented in the Results section.

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: if energy depletion leads to reduced commitment, this can arise either because k_off increases, because k_inj decreases, or because both change, as long as k_off/k_inj becomes larger in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also leads to the trade-off now discussed in the revised manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become significantly larger in the inactive case. Depending on whether this is achieved through changes in k_off or k_inj, the cost of discrimination appears either as slower commitment or as additional energy dissipation. We agree that the previous wording overstated the mechanistic interpretation, and we have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (3) The authors presented an argument that performing infection of all five phages in the same condition is an advantage, allowing for comparison across different phages. While this goal is a completely valid one, it is difficult to reconcile that with the fact that different phages require different optimal conditions for successful infection. For instance, phage T5 famously requires Ca2+ for successful infection into the host bacterium (and later successful replication); see PMID 13174489. However, all infections were performed in TMG, which lacks Ca2+. Perhaps the absence of T5 dependence on the host metabolism is because the infection condition used by the authors was not optimal for T5 to begin with? Similar arguments could be made for other phages.

      Our study alone cannot eliminate that possibility. However, we have cited multiple previous studies, for example references citing Braun et al., showing that T5 remains insensitive to the host metabolic state under different buffer conditions. We therefore believe it is unlikely that the lack of metabolic dependence we observe for T5 is simply due to suboptimal infection conditions.

      (4) Whereas the manuscript examined five coliphages, only phage T5 and phage λ were discussed extensively. I believe some discussion points for these two phages need clarification.

      We focused our discussion on the phages T5, λ and φ80 because these are the phages for which similar effects have been reported previously in the literature. This allowed us to connect our findings directly to existing work and to discuss mechanistic hypotheses in a meaningful comparative framework. For the remaining phages, to our knowledge no prior studies have examined their behavior under comparable metabolic conditions, and therefore a similarly detailed discussion would have been speculative. Nevertheless, all five phages are treated equally in the presentation of the experimental results and in the quantitative comparison of adsorption rates.

      (4.1) Phage T5: The data obtained by the authors show that the infection rate of phage T5 is not impacted by the metabolic state of the host cell. Considering that the authors used the terms "infection", "adsorption", and "entry" interchangeably to refer to the irreversible commitment of a phage to a host cell (see point 2), this discussion regarding phage T5 lacks one critical literature context: DNA entry of phage T5 is known to occur in two phases (first-step transfer and second-step transfer). Critically, the second step can only occur if phage proteins encoded by the phage DNA transferred in the first step are expressed (see PMID 10577483 and the cited papers therein). In that context, metabolic poisoning of the host bacteria should have impeded T5 infection. The authors should comment on this point.

      As the reviewer pointed out, our usage of the terms infection, adsorption, and entry should have been more careful. Our experiment can only identify irreversible commitment of phage to a host cell. For T5, we expect that this irreversible commitment already occurs upon first-step transfer of phage DNA. As a result, even if second-step transfer is impeded under metabolic poisoning, our method would not resolve that effect. We have added this clarification to the revised manuscript.

      (4.2) Phage λ: The experiment using phage λ in this current study shares many resemblances to that in Brown et al. 2022. That feature alone is not a problem, but at many places in the text, the writing is ambiguous as to whether it is discussing the results in Brown et al. 2022 or in the current manuscript. I am giving three examples below, but this is not exhaustive: (i) Lines 67-69, there is no Brown et al. 2022 reference immediately after "a mutant phage variant (λh) could bypass this dependency [...]" (not just in the previous sentence); (ii) Line 228 should clearly say "Our previous findings suggested that phage λ is capable of [...]", since it concerns Brown et al., 2022, not the current study; and (iii) Lines 245-246, there is no Brown et al., 2022 reference immediately after "we observed that a mutant variant [...] even energy-depleted host" (without a reference, it reads like the authors "observed" that finding in this current manuscript).

      The reviewer is right. In those places, the text was ambiguous as to whether it referred to the present study or to Brown et al. (2022). We have now inserted the reference at the relevant points and revised the wording where needed to make this distinction explicit.

      Also, regarding phage λ: The discussion between line 230 and line 249 is very interesting, but since it concerns the differences between λ PaPa and Ur-λ, the authors should consider mentioning and discussing a very relevant recent study, PMCID: PMC6312755.

      We agree that the study by Guan et al. is very relevant and interesting. However, our point in this part of the Discussion is only to clarify that we used λ PaPa and not the originally isolated λ strain. We have therefore limited the discussion here to that distinction.

      (5) Control experiments, or references to prior studies, are needed to support that the As/Az treatment at this concentration and duration (at least 10 minutes) is sufficient to deplete the metabolic state of the cell. For instance, this can be shown by impeded or null cell growth, arrested motility (using a standard swimming assay), or a fluorescent reporter for the energetic state of the cell.

      The 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyperdiffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027) where the effect was assessed by monitoring the rate of movement of the λ receptor on the bacterial surface. We have clarified this point in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As mentioned earlier, I found the paper interesting and addressed an important and significant knowledge gap.

      My biggest concern is about the interpretation of the experimental data in light of the two-step model. In particular, around line 286, it is stated "k_inj is more sensitive to metabolic state than k_off". Assuming k does not depend on metabolic state, which is a fair assumption, the equation for eta only depends on the ratio between k_inj and k_off and not on the individual parameters separately. Consequently, there is no way of saying which one of the two is more affected by metabolic state, unless the model already assumes that k_off is not influenced by metabolic state. The results could equally be explained by k_inj decreasing in metabolically depleted cells, or k_off increasing in such cells. If this is an assumption of the model, this should be clearly stated and not reported as a consequence of the data, as it is at the moment. Also, how does this mathematical model connect to the fitting function used in Figure 2b?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: discrimination requires that k_off/k_inj be larger for inactive hosts than for active hosts, such that commitment is specifically reduced in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This introduces a trade-off: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination requires this ratio to become significantly larger for inactive hosts. If this is achieved through changes in k_off, discrimination comes at the cost of slower commitment by allowing more time to leave; if it is achieved through changes in k_inj, it can preserve fast commitment to active hosts but requires additional energy dissipation in order to actively modulate commitment. We have therefore revised the text accordingly to frame the argument in terms of this trade-off, rather than attributing the effect specifically to k_inj. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      I have a related experimental criticism. The kinetic model presented assumes an exponential decay of free phage, which is a commonly used assumption in the phage literature. Given that the phage types used in this study lyse relatively slowly, it would be good to actually see adsorption curves, in which free phage is measured at different time points between inoculation and lysis. This data would not only provide useful evidence for the kinetic model, but it should also replace what is now in Figure S3, which consists of fitting one experimental point with one line. As it currently stands, Figure S3 is not useful actually misleading.

      We appreciate the reviewer’s point. We agree that adsorption curves, in which free phage is measured at different time points between inoculation and lysis, would provide a stronger basis for evaluating the kinetic model. However, we do not have the resources to perform these additional experiments within the scope of the present study. Following the comments of all reviewers on this point, we have therefore decided to remove Figure S3 to avoid confusion.

      Finally, it is not clear to me why the quantity "Ratio" has been chosen to be presented in Figure 1, rather than the ratio of estimated adsorption rates eta'/eta, which is much more intuitive for a phage study and contains the same information. I would recommend switching to this choice, unless there is a clear rationale for why the quantity "Ratio" is more useful/effective. Showing eta'/eta would also increase the readability of Figure 1, as it would move the y-axis to a logarithmic scale and better visualize values around 1.

      We used “Ratio” in Figure 1 to illustrate the experimental design, controls, and measured quantities directly, as it more transparently reflects the data collected. In the second part of the analysis, where we compare time-independent adsorption rate estimates, we have presented the corresponding values of η′/η as suggested.

      Minor comments:

      (1) Introduction

      Line 31: "... such as nutrient limitation, fluctuating temperatures, and variable energy availability" - if drawing a distinction between energy availability and nutrient limitation, please make explicit what this distinction is. Energy availability seems like a natural consequence of nutrient availability.

      While energy and nutrient availability are often linked in E. coli, they represent distinct physiological constraints. Nutrient limitation refers to the lack of essential biosynthetic precursors such as nitrogen, phosphorus, or amino acids. Energy availability, in contrast, reflects the cell’s ability to generate ATP and reducing equivalents through metabolic processes. For example, under anaerobic conditions, E. coli may have ample nutrients but limited energy production due to the lower efficiency of fermentation compared to aerobic respiration. Thus, energy limitation can occur independently of nutrient limitation.

      (2) Results

      (a) Whole Section: Please label equations.

      All equations have now been labelled in the revised manuscript.

      (b) Lines 105 to 114: As stated in Major Comments, I think the clarity of the paper would be improved by introducing the relative adsorption rate here and dropping the concept of Ratio entirely. However, if the authors wish to use Ratio, I would recommend the following:

      Lines 105 to 109 are confusing to read because of the number of connectives: "... ratio of free viruses from permissive AND resistant hosts respectively TO the free viruses in buffer under energy-depleted AND energy competent conditions". This would be clearer if each quantity were given an algebraic symbol, and RP, RR, and Ratio were defined through formal algebra, rather than mixed mathematical and sentence notation.

      This section has been rewritten for clarity. We now introduce explicit algebraic symbols and define the quantities formally, which removes the ambiguity present in the sentence-only description while retaining the intended meaning.

      The chemical names "arsenate" and "azide" should appear in the body of the text before they appear abbreviated in an equation. Please state at this point that these are both metabolic inhibitors, as it is not immediately clear what role they play or why you are using them.

      The text has been updated to introduce arsenate and azide by name before the abbreviations are used, and we now explicitly note that they act as metabolic inhibitors.

      On line 114, the authors helpfully provide an interpretation of Ratio = 1. It would be useful to provide at the same time interpretations of Ratio >1 and <1, perhaps 2 and 0.5 specifically?

      We have added brief explanations illustrating the interpretation of Ratio values greater than and less than 1, including examples of 2 and 0.5.

      I would consider giving this quantity a more interpretable name than Ratio. This quantity represents how much a bacteriophage preferentially adsorbs to metabolically active cells, so perhaps "Selectivity" or "Adsorption Bias"?

      We intentionally retained the generic term “Ratio”, as this quantity reflects an intermediate experimental measure used to describe the process rather than a newly defined metric. Its purpose is to bridge the experimental observations and the subsequent quantification of effects on the adsorption rate (η).

      (c) Lines 117 to 122: the authors sometimes refer to ratios explicitly, "average ratio of around 1.6" and other times say e.g., "a greater than 3 times increase in viral particles". Using more consistent language (saying "Ratio" every time) would be clearer.

      We have standardized the terminology in this section and now refer to all fold-changes consistently using “Ratio” to avoid ambiguity.

      (d) Figure 1

      Phages λ and T6 look like they have ratios less than 1 for resistant cells? If this is true / if the ratio is statistically significantly below 1, please comment.

      The ratios for λ and T6 are not statistically different from 1. The apparent deviation is within the standard error of the mean. To make this clearer, we have added the corresponding p-values to Table S2 in the Supplementary Information.

      Ratios near 1 are difficult to distinguish from 1, especially in panels A and D. Using a logarithmic scale on the y-axis would make the plots more readable.

      Because the values in these panels are not statistically different from 1, changing to a logarithmic scale would not alter the interpretation. We therefore retained the current axis scaling to reflect that there is no meaningful deviation from 1 in these cases.

      The data corresponding to individual experiments have no error bars. Given that the number of free virions was determined by plaque assay, which carries an intrinsic sampling error, this uncertainty should be reflected in the plots.

      We thank the reviewer for this important comment. Because plaque assays have compound sources of stochastic variation, assigning a per-measurement error bar would risk implying false precision. For this reason, we present the values from each biological replicate directly, and the uncertainty is represented in the statistical summary across replicates. Specifically, for each phage and condition we show the three independent experimental measurements and report the mean along with the standard error of the mean. This approach allows us to represent biological variability without implying a precision that cannot be accurately quantified at the level of single plaque counts.

      Similarly, the average value does show error bars, but it is not stated what these error bars correspond to: standard error in the mean, standard deviation of the sample, or combined uncertainty?

      The caption has been updated to state that the error bars represent the standard error of the mean.

      The resistant bacteria seemed to have ratios close to 1 in all cases. Is this because very few virions adsorbed under both energy conditions?

      Resistance is commonly associated with a lack of a surface receptor for the phage (or generally an entry pathway). We use the resistant bacteria as a control group for the effect of the conditions on adsorption. For resistant bacteria, the Ratio should be 1 since virions do not adsorb under both energy conditions. Any slight variations from 1 should come from sampling errors or small heterogeneity in the population.

      (e) Figure 2

      Please comment on what the error bars here represent. Error bars in Figure 2 A seem to permit negative (or at least zero) values of relative adsorption rate for phages m13 and T6, possibly implying an overestimate of the error? If it is the case that multiple values used to calculate the mean are far apart, possibly showing the values individually through a superimposed swarm plot would be clearer.

      This point is now addressed in the Supplementary Information, where we clarify how the error bars were calculated.

      (3) Discussion

      (a) Line 189: "high metabolic state" is imprecise. Say "energy-competent" to be consistent with earlier language.

      To maintain continuity with earlier terminology, we now include “energy-competent” in parentheses alongside “high metabolic state,” while retaining the original phrasing for readability.

      (b) Figure 3, population level

      Show adsorbed virions physically attached to bacteria, rather than removing them completely from the image, as currently, the implication is that at a high metabolic state, there are fewer virions total, not fewer virions remaining in solution because more are adsorbed. You could go as far as to add a third "after centrifuging" row, showing the adsorbed phages stuck in the pellet and the unadsorbed phages remaining in solution.

      Thank you for this suggestion. Figure 3 has been updated to depict adsorbed virions attached to bacterial cells, clarifying that the decrease represents adsorption rather than loss of total particles. This change improves the accuracy and interpretability of the schematic.

      (4) Methods and Materials

      (a) Figure 5

      The step "estimate cell numbers from OD" appears to follow incubating plates overnight. If the cells you are counting come from the pellet produced by centrifuging 3 steps prior, you could add a fork into the black line connecting the steps, with one branch corresponding to the supernatant and phages, and the other to the pellet and cells?

      Thank you for pointing this out. The order in the figure has been corrected: cell numbers are estimated from OD before overnight incubation. This resolves the confusion without the need for branching in the workflow diagram.

      (a) Line 332

      You allow as much time as possible for adsorption without the possibility of lysis. Did you determine the lysis times / latent periods of these phages through one-step-growth-curves, or use published results, in which case please cite? Having obtained the lysis time by either method, what fraction of the lysis time did you allow for adsorption? Also, please add supplementary tables with lysis times used for the different phages.

      We thank the reviewer for this comment. We used published latent-period values as guides and verified compatibility with our own system when selecting incubation times. We have clarified this in the text and added the relevant citations. We did not use a common fixed fraction of the lysis time for all phages; instead, incubation times were chosen to allow sufficient time for adsorption but not for completion of the first lytic cycle. For λ, productive lytic development was blocked in the host background used, as in Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119. For ϕ80 and T5, we used published latent-period values as guides and verified their compatibility with our own system (De Paepe and Taddei, PLoS Biology 2006, http://dx.doi.org/10.1371/journal.pbio.0040193). M13 is a chronic filamentous phage and therefore does not have a standard lytic latent period; in our host–phage combination, it required more than 1 h before phage release. For T6, we relied primarily on the kinetics observed in our own system, since adsorption was unusually slow for this phage–host pair under our assay conditions. Although literature reports describe shorter T6 latent periods under specific assay conditions (Foster and Johnson, Journal of General Physiology 1951, http://dx.doi.org/10.1085/jgp.34.5.529), this is consistent with published work showing that adsorption and infection kinetics can vary substantially with host background, surface structure, and experimental conditions (Heller and Braun, Journal of Bacteriology 1979, http://dx.doi.org/10.1128/jb.139.1.32-38.1979; Storms et al., Biochemical Engineering Journal 2012, http://dx.doi.org/10.1016/j.bej.2012.02.010).

      (5) Supplementary

      Figure S1

      This data is useful in understanding the main body of the paper, and I think this should form part of a main figure (possibly with the individual experimental data points superimposed over the bars). This could come before or as part of Figure 1?

      We thank the reviewer for this suggestion. We have explored including these data directly in the main figure but found that doing so substantially reduced the readability of the figure, as the underlying table is visually dense. For this reason, we chose to summarize the results in Figure 1 and present the detailed data separately in Figure S1 of the Supplementary Material, along with the Ratio analysis, which more effectively conveys the trends without overloading the main figure.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) L16-18: This sentence could be made more accessible as 'error correction' is not an intuitive term in the phage field.

      We have updated the overall theory section including the terminology. Instead of error correction, we now refer to it as a discrimination process.

      (2) L96-98: Does this potentially indicate a trade-off where evolution for stronger binding cannot evolve at the same time as responsiveness to metabolic activity?

      We agree that this sentence made a stronger evolutionary claim than our data support. Since we only tested four laboratory phages, we cannot conclude that there is an evolutionary trade-off between stronger binding and responsiveness to host metabolic activity. We have therefore removed this sentence to avoid making an unsupported evolutionary interpretation.

      (3) L102: What does 'post-cellular' mean?

      Postcellular supernatant is simply the liquid that remains after cells have been removed. During centrifugation, the cells pellet at the bottom, and the liquid above (which can contain viruses) is the postcellular supernatant.

      (4) L105-107: Worth splitting into two sentences as it is a bit unclear if ratios are built between permissible and resistant hosts or between buffers or both.

      Thank you for the suggestion. We have rewritten this section into two sentences to clarify how the ratios are constructed, and we hope the revised wording improves readability.

      (5) L110-122: Figures S1 and S2 could be referenced here.

      References to Figures S1 and S2 have now been added in this section.

      (6) L137: As P(0) is the viral concentration in buffer, I am assuming that the phage lysate has been diluted in buffer and phages have been added to cultures from the same dilution tube to guarantee equal starting numbers, but I couldn't find this in the methods.

      This clarification has been added to the Methods and Media section of the Supplementary Information.

      (7) L243: It would be worth defining what 'hyperdiffusion' means.

      We have added a brief definition of “hyperdiffusion”.

      (8) L253-256: I do not entirely follow this explanation.

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #3. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (9) L284: Why is k_inj necessarily more sensitive to the metabolic state than k_off? Could membrane changes under stress increase k_off?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: reduced commitment in inactive cells can arise through an increase in k_off, a decrease in k_inj, or both, as long as k_off/k_inj becomes larger in the inactive case. What matters is therefore not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also underlies the trade-off now discussed in the manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become much larger in the inactive case. We have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (10) Figure 1: There seems to be more variation between replicates in phage Lambda than in other phages. Is this caused by receptor number heterogeneity in the population?

      Unfortunately we do not have a way to compare receptor number heterogeneity across the different phage receptors in our experiments. We therefore cannot conclude that the larger variation observed for phage λ is caused by receptor number heterogeneity in the population.

      (11) Figure S1: There seems to be a significant difference between phage Lambda viability in the two buffers - do the authors have an idea where this comes from?

      There is no difference in λ viability between the two buffers. The apparent difference in the figure is due to sampling variability.

      (12) Figure S3: Last sentence of the legend probably shouldn't say 'upper'.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      Reviewer #3 (Recommendations for the authors):

      (1) The text reads as incomplete in some places. Can the authors please provide clarifications on the following points?

      (1.1) Lines 235-256: How do the authors draw a conclusion that "a phage can detect host metabolic status" from a study that used purified LamB receptors (i.e., no live cells with any metabolism) extracted from two different bacterial species (i.e., not a difference in metabolic states)?

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #2. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (1.2) Line 270, in the abstract, and in the caption of Figure 4: The authors described the model using terms such as "an error-correction mechanism" or "standard error correction", but there is little explanation. Can the authors clarify what kind of "error" is discussed here, and how it is "corrected"? In the "standard error correction" model, what determines which method of correction is "standard"? If "error correction" is a standard term in phage-bacterial interaction modeling, please provide references.

      We agree with the reviewer that our use of the term error correction was not appropriate in this context. The proper term is discrimination process rather than error correction. We have now corrected this terminology throughout the manuscript and clarified the underlying logic in the relevant sections.

      (1.3) Line 301: The authors speculated that phage T5 is "better suited to ecological niches", but I am not sure how that is consistent with their data showing T5 is more rampant, that they infect both energy-competent and energy-depleted cells, not just depleted cells. Why "niches", and why are T5 better suited to environments "where energy-limited cells dominate", not just any environment?

      We agree that this point was not stated clearly enough. What we intended to convey is that T5 would be at a net disadvantage in a niche containing a mixture of energy-competent and energy-deficient hosts. We have updated the main text accordingly.

      (1.4) Line 303, and related to point 6.3. above: Phage λ can also infect and replicate in "starved bacterial cells" (shown in Kourilsky 1974 and Geng et al. 2024, both of which were cited in this manuscript). How do the authors reconcile these reports with the discussion point in line 303, and their data that only phage T5, but not λ, shows insensitivity to the host metabolic state?

      Our data do not imply that phage λ is unable to infect starved bacteria. As shown in Kourilsky (1974) and Geng et al. (2024), λ can indeed infect and replicate in nutrient-limited cells. Our results specifically indicate that λ infection under starvation proceeds with a reduced adsorption rate, while T5 maintains the same adsorption rate even when the host is starved. Thus, our conclusion is that T5 is insensitive to the host metabolic state at the level of adsorption, whereas λ is not. We acknowledge that the wording in line 303 may have unintentionally led to confusion, and we have revised this part of the text to avoid that.

      (2) The following comments relate to the text and figures in the manuscript. There are many places in the manuscript that could use fine proofreading and copy-editing for clarity and consistency. For example:

      (2.1) If I understand it correctly, the equation in between lines 109 and 110 should be clarified using terms such as "Free viral particles after mixing with bacteria in Arsenate and Azide" and "Free viral particles in bacteria-free buffer with Arsenate and Azide". As it stands, it is not clear which terms correspond to conditions where bacteria are present.

      The equation has been updated to explicitly indicate which terms refer to mixtures containing bacteria and which refer to bacteria-free controls, so that the correspondence between conditions is now clear.

      (2.2) Equations in between line 276 and 283, and elsewhere: Some concentration terms are enclosed in brackets ("[BP]"), while most are not.

      This notation has been clarified. We now use “[PB]” specifically to denote the transient phage–bacterium complex, distinguishing it from the product P⋅B. All other concentration terms are written without brackets for consistency.

      (2.3) Figure 4 and in equations: "BP" or "PB"?

      The notation has been made consistent throughout; we now use “PB” exclusively to denote the phage–bacterium complex.

      (2.4) Line 284 and line 286: The "inj" in "k_inj" is sometimes italicized, sometimes not.

      The notation has been standardized so that k_inj is now formatted consistently throughout the manuscript, without italicizing “inj.” Also we have replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (2.5) Figure 5: Was the step "Estimate cell numbers from OD" really performed on the next day after the experiment (i.e., >12 hours after infection and phage plating), not immediately after cell washing?

      Thank you for pointing this out. The figure has been updated to reflect the correct order of steps: cell numbers are estimated from OD immediately after washing, followed by overnight incubation of the plates.

      (2.6) Figure S1: As it stands now, the x-axis of each panel can be read either as "Permissive, Resistant bacteria, Buffer" (missing "bacteria" for the first pair of bars), or "Permissive (bacteria), Resistant (bacteria), Buffer (bacteria)" (extra "bacteria" for the last pair of bars).

      The intended interpretation is the second one (permissive bacteria, resistant bacteria, buffer).

      (2.7) Figure S3: The panel letters "A" and "B" are missing in the figure. Also, it is not clear why the legend for the five phages and the legend for the measurement times are not combined.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      (2.8) Strain table in the Methods and Materials: Please write genotypes with italicization, and consistently indicate mutations and deletions with the minus sign superscript or the Δ prefix. Also, for the S3222 strain: Is it really the entire Mal regulon mutated ("Mal-"), or just lamB-? In Brown et al. 2022, it was only the latter.

      Genotypes have been reformatted with consistent notation. For S3222, the correct designation is Mal-, as in the SI of Brown et al. 2022. In this case, Mal- is intended as a phenotypic designation rather than a specific genotype, and we have therefore formatted it accordingly, i.e. neither italicized nor written in lower case.

    1. eLife Assessment

      This study makes a valuable contribution by broadening the range of eukaryotic model systems and establishing Blastocystis, the most prevalent microeukaryote in the human gut, tractable for reverse-genetics investigations. The presented imaging data are convincing and informative, although confirmation by molecular methods would further strengthen the study. The work should interest readers studying host-microbe interactions in the human gut, as well as those developing new systems for eukaryotic research.

    2. Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in Blastocystsis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did - so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      Comments on revised version.

      The authors have revised their manuscript to clarify that molecular analyses have not yet occurred and have resolved the technical/publication issues with the figures. I look forward to seeing these tools used in future publications to answer important questions in Blastocystsis research.

    3. Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide a solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to visual presentation of the data would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested "23 promoter-terminator pairs" to better reflect that only promoters varied.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, the part A of Figure 2 is split into two panels with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of the part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible and it draws attention to a distracting design feature rather to the data themselves.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      (4.2) In the parts B-D, the legend should explain more clearly what each image shows and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images as the current distortions complicate interpretation.

      Comments on revised version.

      The revised version provides sufficient clarity and appropriate visual presentation. Some confusion evidently arose due to my misunderstanding, so I thank the authors for their comprehensive clarifications and patience.

    4. Author response:

      The following is the authors’ response to the original reviews.

      In revising the manuscript, we have focused on three main priorities raised during review: (1) improving precision around evidential claims, particularly concerning vector maintenance and P2A-mediated protein separation; (2) substantially improving figure quality, accessibility, and legend clarity; and (3) correcting inconsistencies and expanding methodological detail where requested.

      This study was intended as a foundational genetic toolkit and methodological framework for Blastocystis ST7-B, establishing practical workflows for DNA delivery, endogenous regulatory-element benchmarking, antibiotic-selected recovery, clonal propagation, and reporter-based analysis in a genetically challenging anaerobic microbial eukaryote. The central evidence presented is therefore functional in nature: reproducible transgene delivery, selectable recovery and propagation of colony-derived transgenic lines, and detectable reporter expression using multiple anaerobic-compatible reporter systems.

      We agree with the reviewers that several additional experiments, including Western blot analysis of P2A-containing constructs, outward-facing PCR, plasmid rescue assays, and selection-withdrawal experiments, would further strengthen the mechanistic interpretation of the system and help distinguish episomal persistence from genomic integration. We have therefore revised the manuscript throughout to clearly separate what is directly demonstrated from what remains a plausible working interpretation or important future direction.

      Importantly, the revised manuscript no longer presents episomal maintenance or complete P2A-mediated protein separation as demonstrated conclusions. Instead, these are now discussed explicitly as unresolved mechanistic questions requiring future molecular analysis. Nevertheless, the central methodological conclusion remains unchanged: stable selectable transgene expression, recovery of colony-derived transgenic lines, and reporter-positive Blastocystis ST7-B transformants can now be reproducibly obtained.

      Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in the Blastocystis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did, so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      We thank the reviewer for the positive assessment of the manuscript and for recognising both the technical difficulty and broader utility of establishing genetic tools in Blastocystis and other experimentally challenging microbial eukaryotes. We also appreciate the reviewer’s identification of the main evidential limitations in the original manuscript, particularly regarding vector maintenance and P2A-mediated protein separation.

      First, regarding the question of vector topology and the interpretation of episomal maintenance.

      We agree that the original manuscript presented episomal persistence too strongly relative to the evidence currently available. We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working model rather than a directly demonstrated conclusion.

      The transfection system used here was adapted from Li et al. (2019), including use of the pXS2-P<sub>Legumain</sub>-derived plasmid framework. Importantly, the construct used in the present study does not contain the original Trypanosoma brucei tubulin-targeting region associated with homologous integration in the original pXS2 system. Complete plasmid sequencing confirmed that the constructs function here as heterologous expression plasmids carrying Blastocystis ST7-B regulatory elements and transgenes. While this does not demonstrate episomal persistence, it also means that genomic integration cannot be inferred from the historical pXS2 vector architecture alone.

      We further note that comparative genomic analyses by Gentekaki et al. (2017) suggest that Blastocystis lacks components of the canonical non-homologous end-joining (NHEJ) machinery, implying that homologous recombination is likely to represent the principal route for double-stranded DNA repair. Because the constructs used here did not contain Blastocystis homology arms, there is currently no obvious mechanism favouring targeted homologous integration. Nevertheless, we fully agree that genomic integration cannot presently be excluded.

      To reflect this appropriately, the revised manuscript now explicitly separates the demonstrated functional outcomes from unresolved mechanistic questions concerning vector maintenance. We also identify several future approaches that would help distinguish episomal persistence from genomic integration, including outward-facing PCR, plasmid rescue followed by full plasmid sequencing, Southern blotting, FISH, selection-withdrawal experiments, and long-read sequencing approaches.

      We have revised the manuscript throughout to remove statements implying demonstrated episomal maintenance and now present episomal persistence only as a plausible working interpretation.

      In the Methods section under Cloning, the following text has been added:

      Lines 202–206: “The constructs used in this study were derived from the pXS2-P<sub>Legumain</sub> vector described by Li et al. (2019), which adapted a heterologous expression-vector backbone for transient plasmid-based expression in Blastocystis ST7-B. Here, the same molecular backbone was used as a plasmid scaffold carrying Blastocystis-derived regulatory elements and transgenes.”

      In the Discussion, the following text has been added/edited:

      Lines 665–673: “The molecular maintenance state of the introduced constructs remains unresolved: episomal maintenance is a plausible working model, but genomic integration cannot be formally excluded. The constructs used here lack Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components (Gentekaki et al., 2017), making targeted integration by standard repair routes unlikely but not impossible. Direct assays such as outward-facing PCR, plasmid rescue followed by full plasmid sequencing, FISH, or selection-withdrawal experiments will be required to distinguish episomal persistence from integration.”

      Second, regarding P2A-mediated protein separation.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. We have therefore revised the manuscript to avoid implying that complete protein-level separation was directly demonstrated.

      The revised manuscript now states only what is directly supported by the current data: that P2A-containing bicistronic constructs supported antibiotic-selected recovery of transgenic lines together with detectable downstream reporter expression. The microscopy data therefore support functional downstream reporter expression, but do not by themselves exclude residual uncleaved fusion products.

      We selected P2A because it is a compact and well-characterised peptide with high reported separation efficiency across multiple eukaryotic systems, including microbial eukaryotes. However, we agree that P2A performance can be context-dependent, and we now explicitly identify biochemical validation of P2A cleavage efficiency as an important future direction.

      Importantly, these revisions do not alter the central methodological conclusion of the study, namely that selectable transgene expression, propagation of reporter-positive lines, and recovery of colony-derived Blastocystis ST7-B transformants can now be reproducibly achieved.

      Text inserted in the Results:

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.”

      Lines 403–404: “However, protein-level separation was not directly tested, and the extent of any residual uncleaved fusion product remains unresolved.”

      Text inserted in the Discussion:

      Lines 619–629: “P2A was selected because it is a well-characterised peptide with high reported separation efficiency in human cell lines, zebrafish embryos, and mice (Kim et al., 2011). It also has precedent across microbial eukaryotes, including the protest Dictyostelium discoideum (Zhu et al., 2023), the fungi Aspergillus niger (Schuetze and Meyer, 2017) and Ustilago maydis (Müntjes et al., 2020), and the apicomplexan parasites Toxoplasma gondii (Markus et al., 2019) and Plasmodium falciparum (Dans et al., 2024). However, P2A performance is context-dependent, and the evidence presented here is functional rather than biochemical. P2A-containing constructs support antibiotic-selected recovery and downstream reporter expression in Blastocystis ST7-B, but ribosomal skipping efficiency and any residual uncleaved product will require direct protein-level validation.”

      Reviewer #1 (Recommendations for the authors):

      (1) Please could you show a Western blot to confirm if P2A has worked? It could be that the proteins are being expressed as a polyprotein.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. This is an important point, and we have revised the manuscript accordingly to avoid implying that complete protein-level separation was directly demonstrated.

      The current study was designed as a first-generation functional genetic toolkit for Blastocystis ST7-B, focused primarily on establishing reproducible workflows for selectable transgene expression, reporter recovery, and propagation of transgenic lines in this experimentally challenging anaerobic microbial eukaryote. The toolkit is therefore validated here through functional outcomes, including antibiotic-selected survival, stable propagation through extended passaging (>15 passages) and cryopreservation, and detectable reporter fluorescence above wild-type autofluorescence.

      P2A was selected because it is a compact and well-characterised peptide with high reported ribosomal skipping efficiency across multiple eukaryotic systems, including microbial eukaryotes, as discussed above. Nevertheless, we fully agree that direct biochemical validation would strengthen the mechanistic interpretation of the bicistronic system in Blastocystis ST7-B. We therefore now explicitly identify Western blot analysis, ideally using epitope-tagged upstream and downstream products, as an important future direction for quantitative assessment of P2A cleavage efficiency and any residual uncleaved fusion products.

      Relevant manuscript revisions are described above under the general response to Reviewer 1.

      (2) Something has gone wrong with figure formatting. Figure 2 is nearly illegible and I cannot read the text in section A. Sections B, C, and D have lost their labels and are fuzzy and surrounded by black. A similar issue affects Figure 3. Everything is just black with a few cells. It is illegible when printed.

      We thank the reviewer for highlighting these presentation issues and agree that the submitted figure quality significantly impaired readability and interpretation. The problems appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, contrast, and panel labelling in the review PDF.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved typography, panel separation, colour scaling, and legend clarity throughout. In response to additional reviewer suggestions, individual data points have now been added to Figures 2B and 2C to improve transparency and interpretability of the underlying data distributions.

      Figures 2 and 3 have been replaced with fully revised high-resolution versions with improved panel labelling, accessibility, typography, and figure legends.

      (3) The data from Figure 2B would be better placed in Table 1 with a column for robust/moderate/intermediate/weak/very weak. This would be much easier for the reader.

      We thank the reviewer for this helpful suggestion. We believe the comment refers to the promoter activity data shown in Figure 2A rather than the voltage optimisation data in Figure 2B. To improve readability and accessibility of these data, we have revised Figure 2A extensively to make the promoter activity tiers more legible and easier to interpret directly from the heat map and accompanying box plots.

      We considered incorporating simplified activity classifications into Table 1. However, activity patterns were construct-specific rather than simply locus-specific. In several cases, multiple promoter fragments derived from the same locus produced substantially different reporter outputs, and activity did not scale monotonically with promoter fragment length. We therefore felt that assigning a single categorical activity label at the locus level would oversimplify the dataset and reduce the construct-level resolution that is central to the toolkit value of the study.

      Instead, we addressed the reviewer’s concern by substantially improving the presentation and readability of Figure 2A, allowing readers to identify robust, moderate, intermediate, weak, and very weak expression constructs more directly while preserving the underlying construct-specific information.

      Figure 2A has been revised to improve clarity, accessibility, and legibility of the promoter activity tiers, allowing construct-level expression classes to be interpreted more directly from the heat map and accompanying boxplots.

      (4) How do you know if the constructs are maintained as episomes? Have you done an outward-facing PCR?

      We agree that direct molecular evidence distinguishing episomal persistence from genomic integration is currently lacking, and we appreciate the reviewer highlighting this important limitation. We have therefore revised the manuscript throughout to avoid presenting episomal maintenance as a demonstrated conclusion and now describe it only as a plausible working interpretation based on the current evidence and vector design.

      We have not performed outward-facing PCR in the present study. As discussed in the general response above, we now explicitly identify outward-facing PCR, plasmid rescue followed by full plasmid sequencing, selection-withdrawal assays, FISH, and long-read sequencing approaches as important future directions for resolving the molecular maintenance state of the constructs.

      The revised manuscript now clearly separates the demonstrated functional outcomes, including selectable transgene expression, recovery of colony-derived transgenic lines, and stable reporter-positive propagation under selection, from the unresolved mechanistic question of vector topology.

      This issue has been addressed throughout the revised manuscript, including in the Methods and Discussion sections, where episomal maintenance is now presented as a plausible but unconfirmed interpretation rather than a demonstrated conclusion.

      Minor Comments

      Line 66: is this one to two billion individuals with Blastocystis, or one to two billion Blastocystis cells per gut?

      The intended meaning was colonised individuals globally. We agree that the original phrasing was ambiguous and have corrected it for clarity.

      Lines 66–67 revised to: “…microorganisms in the human gut, and is estimated to colonise approximately one to two billion people globally (Scanlan and Stensvold, 2013).”

      Line 148: Supplier of IMDM?

      The supplier information was already present in the original manuscript as IMDM L0191 (Biowest).

      No additional manuscript change required.

      Line 157: Who annotated the dataset, the 2017 paper or the present study?

      The dataset annotation derives from Armengaud et al. (2017). We agree that the original wording was unclear and have revised this section substantially to improve clarity regarding the rationale and workflow used for promoter and terminator candidate selection.

      “The relevant Methods section has been extensively revised for clarity and expanded detail” (Lines 156–189).

      Line 166: Who predicted the 3′ UTR, the 2017 paper?

      This information derives from the NCBI annotation associated with the Blastocystis ST7-B genome based on Denoeud et al. (2011). This has now been clarified in the Methods section.

      Clarified in revised Methods section.

      Line 237: How long did it take in days?

      Approximately 2 days.

      Line 270 revised to: “…turned yellow without drug treatment, usually within 2 days post-transfection.”

      Line 325: Typo, missing gap between Figure and 1A.

      Corrected in revised manuscript.

      Reviewer #2 (Public review):

      This manuscript presents a substantial technical advance for the genetic manipulation of Blastocystis by establishing an integrated workflow for stable episomal transgenesis, antibiotic selection, clonal recovery, and reporter-based imaging in the ST7-B subtype. The study is particularly valuable because it combines multiple previously fragmented approaches into a coherent and practically applicable toolkit, including endogenous regulatory elements, optimized electroporation conditions, selectable markers, and anaerobic compatible fluorescent reporters. This methodological work greatly expands the molecular toolbox and future studies focused on both basic and infection biology can now build on the ability to express and localize proteins in fixed as well as live cells.

      The microscopy data are convincing and clearly demonstrate functional reporter expression and successful recovery of stable transgenic lines. Nevertheless, because this is primarily a methodological paper, the study would be further strengthened by the inclusion of Western blot validation of reporter expression and bicistronic constructs. In particular, biochemical analysis of the P2A-containing constructs would help assess the efficiency of ribosomal skipping and exclude the possible presence of uncleaved fusion proteins, thereby providing stronger support for the interpretation of the imaging data and the functionality of the expression system.

      We thank the reviewer for this thoughtful and positive assessment of the manuscript and for recognising the value of integrating previously fragmented approaches into a coherent and practically usable genetic toolkit for Blastocystis ST7-B. We particularly appreciate the reviewer’s recognition that the system expands the currently available molecular toolbox for both cell biological and infection-related studies in this experimentally challenging anaerobic microbial eukaryote.

      We also appreciate the reviewer’s comments regarding biochemical validation of the P2A-containing bicistronic constructs. We agree that Western blot analysis would strengthen the mechanistic interpretation of the reporter system by directly assessing ribosomal skipping efficiency and the possible presence of residual uncleaved fusion products. In response, we have revised the manuscript throughout to ensure that the conclusions remain appropriately evidence-based and do not imply that complete protein-level separation was directly demonstrated.

      The revised manuscript now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, stable propagation of reporter-positive lines, and detectable downstream reporter expression, from unresolved mechanistic questions concerning P2A cleavage efficiency and vector maintenance state. We now also identify biochemical validation of P2A-mediated protein separation as an important future direction for further refinement of the system.

      Relevant manuscript revisions addressing these points are described above under the response to Reviewer 1.

      Reviewer #2 (Recommendations for the authors):

      The quality of images could be better. The figures lacked resolution — possibly a conversion artefact.

      We agree that the figure quality in the submitted review PDF significantly reduced readability and visual interpretation. The issues appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, typography, panel labelling, and contrast rendering.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved panel separation, typography, colour scaling, contrast settings, and figure legends to improve accessibility and interpretability both on screen and in print. In addition, the export workflow and file formatting have been updated to improve compatibility with journal production requirements and reduce the likelihood of compression-related rendering artefacts during manuscript compilation.

      Figures 2 and 3 have been replaced with revised high-resolution versions with improved typography, panel labelling, contrast settings, and accessibility.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for the propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design, to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to the visual presentation of the data, would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      We thank the reviewer for this important point and agree that the original manuscript presented episomal persistence too strongly relative to the currently available evidence. In particular, we agree that the title and several sections of the manuscript implied a level of mechanistic certainty that was not directly demonstrated.

      We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working interpretation rather than a demonstrated conclusion. The revised text now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, recovery and propagation of colony-derived transgenic lines, and stable reporter-positive maintenance under selection, from the unresolved mechanistic question of vector topology.

      As discussed in our response to Reviewer 1, the constructs used here do not contain Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components, making targeted integration by standard repair routes less strongly supported mechanistically, although genomic integration cannot presently be excluded.

      We agree that selection-withdrawal experiments would provide useful additional evidence regarding construct persistence and have now explicitly identified such assays, together with outward-facing PCR, plasmid rescue, FISH, and long-read sequencing approaches, as important future directions for resolving the molecular maintenance state of the transgenes.

      The manuscript has been revised throughout to remove wording implying demonstrated episomal maintenance. Episomal persistence is now discussed only as a plausible working interpretation pending direct molecular validation.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      We thank the reviewer for this careful reading and for identifying this inconsistency. We agree that the distinction between candidate loci and successfully generated promoter constructs was not sufficiently clear in the original manuscript and could create uncertainty regarding the dataset.

      The original candidate set comprised 14 loci selected for promoter and terminator discovery. However, only 11 loci yielded successfully cloned and experimentally tested promoter constructs. The remaining three loci were retained in Table 1 for completeness and transparency, as repeated cloning attempts were unsuccessful despite two independent efforts.

      We have revised the manuscript to make this distinction explicit and to clarify that the reported NanoLuc benchmarking experiments were ultimately performed using constructs derived from 11 successfully cloned loci.

      Lines 361–364: “To expand the available regulatory parts, we screened 23 NanoLuc reporter constructs containing putative endogenous promoter–terminator pairs from 11 of 14 candidate loci; three loci could not be cloned after two independent attempts and are indicated in Table 1.”

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested “23 promoter-terminator pairs” to better reflect that only promoters varied.

      We thank the reviewer for the opportunity to clarify this point. We agree that the original presentation may have created the impression that promoter and terminator regions were independently varied and benchmarked, whereas the experimental design was primarily focused on construct-level comparison of endogenous regulatory modules.

      As described in the Methods, each construct contained a candidate endogenous upstream promoter region together with the corresponding endogenous downstream terminator region derived from the same locus. For consistency and to keep the cloning and screening strategy experimentally tractable, a fixed 500 bp downstream terminator fragment was used for each locus rather than systematically varying terminator length or independently testing terminator activity.

      We therefore retain the description “endogenous promoter–terminator pairs,” since each construct contains both endogenous upstream and downstream regulatory regions from the same genomic locus. However, we agree that the assay was not designed to independently dissect promoter versus terminator contributions to reporter output. We have revised the manuscript accordingly to make this distinction explicit and avoid ambiguity regarding the scope of the benchmarking analysis.

      Lines 365–368: “Each construct paired a candidate upstream promoter region with the corresponding downstream terminator region from the same locus, defined here as the native 500 bp sequence immediately downstream of the stop codon. Where multiple promoter lengths were tested for the same locus, the terminator fragment was kept constant (Table 1; Figure 1A).”

      This design allowed construct-level benchmarking of paired promoter–terminator modules but did not test promoter strength or terminator activity independently.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      We thank the reviewer for these important methodological points. We agree that the original manuscript did not sufficiently clarify the transient nature of the NanoLuc benchmarking assay or the rationale underlying the assay design and timing.

      The promoter benchmarking assay was designed as an early transient-expression screen adapted from the NanoLuc-based workflow of Li et al. (2019), with modifications, rather than as a stable-maintenance assay. No selectable marker was included because the objective was to compare relative early reporter output across constructs shortly after DNA delivery, before prolonged culture effects became dominant.

      The 16–18 h post-electroporation time point was selected based on the NanoLuc expression kinetics reported by Li et al. (2019) and empirical optimisation during assay development. This window allowed robust transient reporter detection while limiting confounding effects arising from prolonged plasmid loss, differential outgrowth, variable recovery, or later culture-level changes.

      We agree that, in the absence of selection, the observed NanoLuc signal cannot be interpreted as an absolute measure of promoter strength independent of DNA uptake efficiency, early plasmid retention, or post-transfection recovery dynamics. We have therefore revised the manuscript to clarify that Figure 2A reports relative transient reporter output under standardized early post-transfection conditions rather than isolated promoter activity alone.

      We now also explicitly state that the data shown in Figure 2A derive from three independent electroporation experiments per construct, each assayed in technical duplicate.

      Lines 241–248: “Promoter–terminator activity was assessed 16–18 h after electroporation using a transient NanoLuc assay adapted from Li et al. (2019), with modifications. This early time point was selected to capture reporter output within the transient-expression window after DNA delivery, before prolonged plasmid loss, differential outgrowth, or culture-level changes could dominate the readout. Because the constructs did not contain a selectable marker, the measured NanoLuc signal reflects early transient reporter output rather than promoter strength independent of DNA uptake, early plasmid retention, or post-transfection recovery.”

      Additional clarification added to Figure 2 legend stating that measurements derive from three independent electroporation experiments, each assayed in technical duplicate.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, part A of Figure 2 is split into two panels, with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible, and it draws attention to a distracting design feature rather than the data themselves.

      We thank the reviewer for these detailed comments regarding figure design and visual interpretation. We agree that the original presentation of Figure 2 introduced unnecessary visual ambiguity through inconsistent colour usage, the use of a diverging colour scale for strictly positive values, and poor readability associated with the dark background and low-resolution export.

      In response, Figure 2 has been extensively redesigned to improve clarity, accessibility, and interpretability. The previous blue–white–red diverging heatmap has been replaced with a sequential colour palette appropriate for positive-only expression data. Boxplot colouring has also been simplified and harmonised with the heatmap scheme to avoid implying unsupported categorical distinctions. In addition, panel organisation, typography, scaling, and legend structure have all been revised to improve readability and reduce ambiguity regarding interpretation of the plotted values.

      We also agree that the black background distracted from the data presentation and impaired visibility of panel labels and image boundaries. The revised figures therefore use white backgrounds together with clearer panel separation and improved label visibility throughout.

      Figure 2 has been completely reformatted using a sequential colour scale in panel A, simplified and harmonised boxplot colouring, larger typography, improved panel separation, revised legends, and white backgrounds throughout. Corrected high-resolution source figures have been provided.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      We thank the reviewer for this helpful comment regarding figure composition and panel separation. We agree that the original presentation made it difficult to distinguish intentional panel boundaries from imaging artefacts, particularly in the low-resolution review PDF generated during manuscript compilation.

      To improve clarity, Figure 3 has been reformatted using white backgrounds, clearer panel spacing, and more explicit separation between individual snapshots and imaging channels. High-resolution source images have also been provided to ensure that fluorescence patterns, image boundaries, and panel organisation remain clearly interpretable both on screen and in print.

      Figure 3 has been reformatted with improved panel separation, white backgrounds, clearer image boundaries, and revised high-resolution source figures.

      (4.2) In parts B-D, the legend should explain more clearly what each image shows, and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      We thank the reviewer for these helpful suggestions regarding figure annotation and legend clarity. We agree that the original presentation did not sufficiently explain the composition of the imaging panels, particularly under the low-resolution conditions of the review PDF.

      To improve interpretability, the revised Figure 3 now includes clearer panel organisation, improved annotations, and expanded figure legends explicitly identifying the individual imaging channels and staining conditions shown in each subpanel. The leftmost panels in parts B–D are now more clearly identified in both the figure and legend, together with the corresponding fluorescence or staining conditions used in each experiment.

      As mentioned in the Methods sections we used Hoechst 33342 to visualise DNA; but we agree that the Hoechst 33342-labelled structures required additional clarification. The revised legend section now explains that multiple Hoechst 33342-positive regions are commonly observed because Blastocystis cells can contain multiple nuclei depending on cell stage and subtype-specific morphology.

      In addition, high-resolution source images have been provided to ensure that fluorescent signals, panel boundaries, and imaging features remain clearly interpretable both on screen and in print.

      Figure 3 legends and annotations have been revised to clarify imaging channels, staining conditions, and panel organisation. The figure caption was also edited to include: “DNA was visualised using Hoechst 33342. Most cells contained two nuclei, and smaller Hoechst 33342-positive signals consistent with mitochondrial DNA were also observed in some instances.”

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images, as the current distortions complicate interpretation.

      We thank the reviewer for this important observation and agree that the apparent morphological differences between the UnaG/smURFP and SNAP-tag panels required additional clarification.

      The images shown for the different reporter systems were acquired under different imaging conditions and microscope configurations following an instrument-related issue during part of the imaging workflow, as noted in the Methods section. As a result, direct visual comparison of cell morphology between reporter systems is not appropriate. The primary purpose of these panels is instead to demonstrate reporter detectability, live-cell labelling capability, and the characteristic fluorescence patterns obtained with the different anaerobiosis-compatible reporter systems.

      In particular, the SNAP-tag panels were included to demonstrate successful live-cell labelling without permeabilisation together with the expected increase in fluorescence signal at higher substrate concentrations, rather than to support quantitative comparison of cell morphology across imaging conditions.

      We considered replacing the affected images. However, equivalent replacement datasets acquired under directly comparable conditions are not currently available. We have therefore retained the original images but revised the figure legend to clarify the intended interpretation and limitations of these panels explicitly.

      Figure 3 legend revised to include:

      “Because images for the different reporter systems were acquired under different imaging conditions, they are presented to demonstrate reporter detectability and labelling pattern and should not be used for quantitative comparison of cell morphology across reporter systems.”

      Reviewer #3 (Recommendations for the authors):

      The reader may find the current order confusing starting with construct design before testing which drug to use for selection. The narrative would work better if it started with antibiotic selection as the first logical step for generating stable cell lines.

      We thank the reviewer for this thoughtful suggestion regarding narrative structure and agree that multiple organisational strategies are possible for presenting a methodological workflow of this type.

      We considered reorganising the Results section to begin with antibiotic selection and drug sensitivity profiling. However, we ultimately retained the overall structure because the manuscript is organised as a toolkit-development framework rather than as a strictly chronological experimental protocol. The Results therefore begin with regulatory-element discovery and construct design, which form the conceptual and experimental foundation of the toolkit, before progressing to DNA delivery optimisation, drug sensitivity profiling, clonal recovery, and reporter validation.

      We felt that this structure most clearly reflects the dependency relationships within the system: regulatory elements are required before constructs can be assembled, constructs are required before electroporation conditions can be evaluated, and selectable constructs are required before stable selection and clonal recovery can be meaningfully assessed.

      (2) The text states that the screen 'focused on the 1,000 most abundant proteins to establish a preliminary library capable of supporting varying levels of transcription.' Since the genome has ~6,000 protein-coding genes, the top 1,000 cover the most abundant proteins — not a wide expression range.

      We thank the reviewer for this important clarification. We agree that the original wording could incorrectly imply that the screen was intended to sample broadly across the full transcriptional range of the Blastocystis genome. This was not the case, and we have revised the manuscript accordingly.

      Our strategy was instead designed to enrich for candidate loci with a higher prior likelihood of supporting detectable transgene expression. Because no genome-wide promoter map, transcription start site dataset, or experimentally validated regulatory annotation was available for Blastocystis ST7-B at the inception of this work, we used the abundance-ranked Blastocystis ST4-WR1 proteomic dataset of Armengaud et al. (2017) as a practical starting point for candidate discovery.

      Importantly, the Blastocystis ST4-WR1 proteome is highly skewed, with 193 proteins contributing approximately 50% of the detected proteome and the 13 most abundant proteins contributing approximately 10% (Armengaud et al., 2017). We therefore selected the top 1,000 proteins not as a representation of the genome-wide expression range, but as a proteomics-guided enrichment strategy to identify loci more likely to contain active endogenous regulatory regions suitable for initial toolkit development.

      We have revised the relevant Methods section substantially to clarify both the rationale and the workflow used for candidate selection, homolog identification, and promoter/terminator definition.

      The Methods section (Lines 156–189) has been extensively revised to clarify the rationale underlying candidate regulatory-element selection. The revised text now explicitly states that the strategy was designed to enrich for likely active loci for toolkit development rather than to systematically survey the full range of promoter strengths across the Blastocystis genome.

      Additional methodological detail has also been added regarding:

      Use of the Armengaud et al. (2017) proteomic and proteogenomic datasets,

      Homolog identification in Blastocystis ST7-B,

      Locus selection criteria,

      Promoter boundary definition,

      And operational definition of candidate terminator regions.

      (3) The Methods contain an inconsistency: cells were left in 0.5 mL, then 1 mL was added, but then only 0.5 mL is apparently used for transfection. What happened to the 1 mL?

      We thank the reviewer for identifying this ambiguity in the transfection workflow description. The apparent inconsistency arose because the protocol description moved from bulk cell resuspension to preparation of individual electroporation reactions without explicitly stating how the intermediate suspension was used.

      After washing, approximately 0.5 mL of cytomix buffer remained above the pellet, and 1 mL of complete cytomix buffer was then added to generate an approximately 1.5 mL cell suspension. Cells were counted from this pooled suspension, after which the volume corresponding to 5 × 10<sup>7</sup> cells was transferred into each individual electroporation reaction. Following addition of DNA, each electroporation reaction was adjusted to a final volume of 500 µL with complete cytomix buffer. The remaining cell suspension was retained for additional transfections or control reactions.

      We agree that the original wording could be misinterpreted and have revised the Methods section to clarify the sequential handling steps more explicitly.

      Lines 225-229 revised to read: “The resulting approximately 1.5 mL pooled cell suspension was used for total viable cell counting using a hemacytometer.”

      “After counting, the volume corresponding to 5 x 10<sup>7</sup> cells was transferred to each electroporation reaction and combined with 25 µg of plasmid DNA. The total electroporation volume was adjusted to 500 µL with complete cytomix buffer.”

      (4) Figures 2 and 3 are too low-resolution for the font size used and for clearly viewing the microscopy images.

      We thank the reviewer for highlighting these readability issues. As noted in our responses above regarding Figures 2 and 3, the low-resolution appearance primarily resulted from manuscript compilation and PDF export artefacts affecting typography, image rendering, and panel clarity in the review version.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions featuring improved typography, panel labelling, contrast, accessibility, and image clarity for both on-screen viewing and print reproduction.

      Revised high-resolution versions of Figures 2 and 3 have been provided as described above. No additional manuscript changes were required beyond the figure revisions already outlined.

      (5) Figure 4 is confusing because the left and right panels appear inconsistent, with much higher concentrations required for growth inhibition in the culture-based assay than the resazurin assay indicated. The rationale for the resazurin assay should be explained, and the complete growth inhibition (CGI) concentration should be highlighted in the right panel.

      We thank the reviewer for highlighting this potential source of confusion. We agree that the distinction between the two assay endpoints was not sufficiently emphasised in the original figure presentation and legend.

      The apparent discrepancy arises because the two assays measure different biological endpoints under different assay conditions. The resazurin assay was used to estimate IC<sub>50</sub> values, corresponding to the concentration at which metabolic activity was reduced by approximately 50% under the assay conditions. In contrast, the small-culture assay was designed to determine complete growth inhibition (CGI), defined operationally as the concentration at which no detectable culture outgrowth occurred after incubation, using phenol red acidification as a culture-level readout.

      Because these assays measure partial metabolic inhibition versus complete suppression of detectable culture outgrowth, the corresponding concentration ranges are not expected to coincide directly. The higher concentrations observed in the right-hand panels therefore reflect the more stringent endpoint associated with complete growth inhibition rather than inconsistency between the assays.

      We agree that this distinction should have been explained more clearly in the original manuscript. We have therefore substantially revised the Figure 4 legend to clarify the rationale underlying both assays, explicitly distinguish IC<sub>50</sub> and CGI endpoints, and explain how the CGI values were used to guide subsequent antibiotic selection conditions for Blastocystis ST7-B transformants. The CGI transition range has also been made more visually explicit in the revised figure presentation.

      Figure 4 caption revised to: “Antibiotic potency and selection-window determination in Blastocystis ST7-B. Dose–response curves for puromycin, trimethoprim, and WR99210 were estimated from a resazurin-based viability assay (n = 3 independent replicates per drug per concentration). Points show mean ± SD, and the insets list the estimated IC50 values with R<sup>2</sup>-values > 0.75 for all fitted curves. The IC<sub>50</sub> estimates represent the drug concentrations that reduced resazurin-based metabolic activity by 50% under the assay conditions.”

      Right panels: “small-culture complete growth inhibition assay using 1 × 10<sup>7</sup> WT Blastocystis ST7-B cells per culture, assayed in triplicate across a wide range of concentrations. Cultures were incubated for 2 days, and outgrowth was assessed using phenol red acidification of the medium as a culture-level readout, with yellow indicating growth and red indicating no detectable growth. The yellow-to-red transition was used to estimate the concentration required for complete growth inhibition and to guide the subsequent antibiotic selection strategy for Blastocystis ST7-B transformants.”

      “IC<sub>50</sub> and CGI represent distinct assay endpoints: the former measures partial reduction in metabolic activity, whereas the latter identifies the concentration at which no detectable culture outgrowth occurs under the small-culture assay conditions.”

      (6) In Figure 3B, the unexpected UnaG fluorescence pattern could be due to protein sequestration because the protein is mildly toxic to the cell. This should be discussed in addition to the reasons already provided.

      We thank the reviewer for this thoughtful suggestion and agree that protein sequestration or reporter-associated cellular stress represent plausible alternative interpretations of the observed UnaG fluorescence pattern.

      We considered the possibility of UnaG-associated toxicity during interpretation of these data. However, under the conditions tested, we did not observe clear evidence of a substantial toxic effect: UnaG-expressing Blastocystis ST7-B cells could be recovered as stable lines, maintained under antibiotic selection, and propagated through continued culture. We therefore felt that direct attribution of the observed fluorescence pattern to reporter toxicity would currently remain speculative.

      At present, we consider the biochemical properties of the UnaG system itself to provide a more parsimonious explanation for the observed localisation pattern. In particular, unconjugated bilirubin is highly hydrophobic and would be expected to partition preferentially into lipid-rich cellular environments. This interpretation is consistent with the lipid-rich peripheral and intracellular structures previously reported in Blastocystis ST7-B (Liao et al., 2023).

      We have therefore revised the Discussion to acknowledge that the observed UnaG fluorescence pattern may reflect a combination of reporter-specific biochemical behaviour, bilirubin partitioning, local intracellular environment, or possible sequestration phenomena. At the same time, we avoid assigning toxicity as a demonstrated mechanism in the absence of direct measurements of cell fitness, reporter abundance, or bilirubin distribution. Such experiments would be required to evaluate this possibility rigorously.

      Lines 642-648: “Consistent with this, lipid-rich peripheral and intracellular structures have been reported in Blastocystis ST7-B, potentially providing favourable microenvironments for BR partitioning and contributing to the punctate UnaG fluorescence pattern (Liao et al., 2023). An alternative possibility is that the observed signal pattern reflects reporter sequestration or reporter-associated cellular stress. However, because UnaG-expressing lines were recovered, maintained under selection, and propagated through continued culture, toxicity remains a possible but untested explanation rather than a demonstrated mechanism.”

      Minor Comments

      Figure 2: Parts B and C should also show individual datapoints for better reader assessment.

      We agree that inclusion of individual data points improves transparency and interpretability of the underlying data distributions.

      Individual data points have now been overlaid on the boxplots in Figures 2B and 2C.

      Figure 3A: Separate channels (fluorescence, bright-field, merge) should be shown rather than only the merge. The current overlay is difficult to interpret, especially for colour-blind readers.

      We appreciate the reviewer’s concern regarding accessibility and interpretability. We considered separating the fluorescence, bright-field, and merged channels for Figure 3A. However, this panel was intended primarily as an overview demonstrating reporter detectability within the bicistronic construct context, while the detailed fluorescence distribution is explored more extensively in the subsequent UnaG panels. We therefore retained the merged presentation for Figure 3A. Importantly, the image is not dependent on red–green discrimination, as it combines a greyscale bright-field background with a high-contrast green/cyan fluorescence signal that remains distinguishable through brightness and contrast differences. In addition, colour-blind-friendly lookup tables (LUTs) were used throughout the revised figure set.

      To further improve accessibility, the original red annotation arrow has been replaced with a colour-blind-friendly annotation colour.

      Briefly define system components (P2A, UnaG, smURFP, SNAP-tag) and add an abbreviation list.

      We agree that brief contextual definitions improve accessibility for readers less familiar with these reporter systems. Rather than adding a separate abbreviation list, we have added short explanatory descriptions at the points where these components are first introduced in the manuscript.

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.” Line 515–516: “UnaG, a bilirubin-binding fluorescent protein originally isolated from the muscle of the Japanese eel (Kumagai et al., 2013)…” Line 527: “smURFP (small ultra-red fluorescent protein)…”

      Abstract: “among the most prevalent microbial eukaryote” should be “eukaryotes”.

      Corrected in revised manuscript.

      Conclusion (2nd sentence): unclear what “endogenous regulatory part discovery” means.

      We agree that this phrase required clarification. The intended meaning was the identification and benchmarking of native Blastocystis ST7-B promoter and terminator elements for construct design and toolkit development. We have clarified this directly in the revised Conclusion section.

      Lines 682–683 revised to: “By bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements, …”

      Author contributions: “critical advise” should be “advice”.

      Corrected in revised manuscript.

      Again, we thank the reviewers for their careful evaluation, constructive criticism, and thoughtful feedback on the manuscript. The review process has substantially strengthened the manuscript by helping us clarify the distinction between what is directly demonstrated experimentally and what remains mechanistically unresolved.

      The central methodological conclusions of the study remain unchanged: the toolkit enables selectable transgene expression, recovery of colony-derived lines, and propagation of reporter-positive transgenic Blastocystis ST7-B lines, extending genetic accessibility in this organism substantially beyond the previous transient transfection framework.

      At the same time, the revised manuscript now more explicitly acknowledges important unresolved mechanistic questions, including vector topology, P2A-mediated protein separation efficiency, and persistence in the absence of selection. These are now discussed transparently together with the future experimental approaches that will be required to address them directly.

      We believe the revised manuscript now presents a clearer, more rigorous, and more accessible description of a practical genetic toolkit for Blastocystis ST7-B and hope that the revisions and clarifications satisfactorily address the reviewers’ concerns.

    1. eLife Assessment

      This important study explores whether complex structures that are lost during evolution can re-evolve, which is a long-standing debate in evolutionary and developmental biology. The authors demonstrate that re-evolution can occur if the gene regulatory network that underlies the development of complex traits is maintained. The evidence supporting its conclusions is convincing and the work will be of interest to those studying the evolution and development of complex traits.

    2. Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype re-evolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      Comments on revised version.

      The authors have addressed the concerns in the revision. No further comments.

    3. Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species, and in different ant castes within these species. The Authors show that this is linked to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve as the key genes have earlier developmental roles, such that CRISPr knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      As the authors note in their response this is very difficult to achieve. While the addition of this data would raise this manuscript to an outstanding one, I think the data presented is solid, well-presented and provides novel insight.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Vasquez-Correa and colleagues describes the expression pattern of the ocelli (simple eye) gene regulatory network in ants. They correlate the expression pattern of these genes with the presence and absence of ocelli in different classes and species of ants. The presence of ocelli is a polyphenic trait in ants - understanding the molecular and developmental underpinnings of polyphenic traits is of significant interest to evolutionary biologists, developmental biologists, and ecologists. The authors propose that the presence of the latent expression of the ocellar network in classes of ants that do not display ocelli in the adults may underlie the re-evolution of ocelli within the ant lineage.

      Strengths:

      The strengths of the manuscript are that it is well written, the images are of the highest quality, and the data support the conclusions of the authors.

      We thank Reviewer 1 for their positive comments.

      Weaknesses:

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants. A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are, in fact, responsible for ocelli development in the ant.

      The reproductive caste in ants is typically composed of both winged males and winged queens. We agree with Reviewer 1 that the queen caste, which develop 3 fully functional ocelli, is an important point of comparison in our study to the wingless minor workers and soldiers. Unfortunately, however, laboratory colonies rarely produce reproductive queens, and in the field, queen production in colonies of C. floridanus occurs within a narrow seasonal window, making the collection of queen larvae particularly challenging for developmental work. In contrast, the winged males, which also develop 3 functional ocelli like the queens for help during mating flights, can be readily generated in the lab throughout the year. Therefore, we use males as a proxy for characterizing ocelli development and GRN in queens and the winged reproductive caste as a whole. Given the deeply conserved gene regulatory networks underlying this trait across insects, we believe this is a reasonable assumption.

      We also agree with Reviewer 1 that using RNAi to knock down genes in the ocelli GRN would improve the study. For completeness of the scientific record, we would like reviewers and readers to know that we actually did, in fact, invest significant effort trying to knock down otd-1 (ortholog of the Drosophila otd gene), which functions as key upstream regulator of ocellar development. In Drosophila, RNAi knockdown of otd disrupts the development of all three ocelli as well as fine morphological features on the anterior of the head. In C. floridanus, otd -1 is expressed in the head capsule and brain (see Author response image 1 in this response). Injection of dsRNA of otd-1into whole soldier-destined larvae, significantly reduced otd -1 expression in the brain relative to its control, while in the head capsule, otd -1 expression remained largely unchanged relative to its control (see Author response image 1 in this response). This indicates that in the same individual, the injected otd -1 dsRNA was able to penetrate and significantly reduce otd -1 expression in the brain, but, was unable to penetrate the head capsule, where otd -1 expression remained largely unchanged. No ocellar phenotypes could be observed in pupae or adults. Therefore, for technical (not biological) reasons, we were unable to knockdown genes in the ocelli GRN in the head capsule. We hope to solve this technical problem in the coming years to add a mechanistic explanation for the latent expression and maintenance of the ocelli GRN in workers that completely lack ocelli as adults.

      Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype revolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      We thank Reviewer 2 for their positive comments.

      Weaknesses:

      Although the molecular pathway is conserved, the mechanism underlying the lack of ocelli in workers remains unclear. In C. floridanus, it could be explained by the evidence of no expression of certain developmental genes, but in other species, e.g. Polyrachis rastellata, is their expression intact, or reduced? There is no control male.

      In addition, HCR in species with partial re-evolution (if their genomes have been sequenced) would be useful to understand the mechanism. For example, there might be differential spatial expression between medial and lateral ocelli.

      We agree with Reviewer 3 that investigating the mechanisms underlying the lack of specific ocelli in these and other species is the next step for this research. Here, our main focus was instead on trying to explain the mechanisms underlying partial reversion of ocelli through the persistence of ocelli GRN expression in adult workers lacking ocelli. We therefore focused on the latent expression of the ocelli GRN in Polyrachis rastellata, a species that completely lack ocelli in adult workers, and how it may have facilitated the partial reversion of a single ocellus in its congener Polyrachis bihamata. Therefore, although we did not reveal specific interruption points in the ocelli GRN in Polyrachis rastellata, our results showing that this species expresses three genes of the ocelli GRN, offers sufficient evidence that this network is conserved and likely facilitated the partial reversion to a single ocellus in P. bihamata.

      We also agree with Reviewer 3 regarding the male control in P. rastellata and obtaining the species in our study that have undergone partial re-evolution. Unfortunately, these ants occur in Southeast Asia and are very difficult to collect. For males in P. rastellata, our colony died before we could try to induce male development. However, given the deep conservation of the network in the males of a genus within the same subfamily (Camponotini), we feel it is reasonable to assume that the network would also be conserved in the males of P. rastellata, especially since the genes we sampled are conserved in workers that do not develop ocelli as adults. As am sure the Reviewer may know that this is a continual challenge of working with emerging models in evodevo.

      Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species and in different ant castes within these species. The authors show that this is linked to to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstatrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      We thank Reviewer 3 for their positive comments.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve, as the key genes have earlier developmental roles, such that CRISPR knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      We agree with Reviewer 3 that functional experiments in species where workers both have ocelli present and absent is a key next step in this research. Please see our response to Reviewer 1 on our failed attempts to achieve this. We are therefore grateful to Reviewer 3 for acknowledging the challenges in trying to establish RNAi and CRISPR in the head capsules of developing workers in these ants. Also, please see our response to Reviewer 2 on the difficulty of finding and collecting these ants, which occur mainly in Southeast Asia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants.

      A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are in fact responsible for ocelli development in the ant.

      Please see our response to Reviewer 1 above.

      Reviewer #2 (Recommendations for the authors):

      For the questions below, if there is no experimental evidence, consider addressing them in the Discussion.

      Do sizes of ocelli different between castes? For example, even workers have 1-3 ocelli, their sizes are smaller than those of males/queens, especially in workers with 1 ocellus. If so, might it be continuous (not binary) changes in downstream gene expression that control ocellus size, with no ocellus below threshold? Does this favor the hypothesis of threshold but not switch?

      We thank Reviewer 3 for highlighting an important point about the size and development of ocelli. Observations suggest that ocelli tend to be larger in queens and males than in workers and soldiers in species with ocelli. However, we lack quantitative data to test this conclusively. We now include a sentence on the Discussion stating that an important avenue of future work should investigate whether threshold or switch mechanisms influencing the presence/absence, as well as size, of ocelli between queens and workers.

      For the species whose workers have a single ocellus, are there variations, e.g. spanning from 0, 1 to 2? If 2, always one medial plus one of the two laterals? If always one, it would be a good control for staining to see up- vs down-regulation of downstream gene expression within the same individual.

      We agree with Reviewer 2 that this is a fascinating approach to our question. We have not observed natural wild-type variation in the number of developing ocelli in the same-sized individuals in the worker caste. However, in a distantly related leaf-cutting ant species (Atta cephalotes) belonging to different subfamily (the Myrmicinae) individuals with different head-to-body scaling within the same colony can vary in the number of ocelli. For example, soldiers of Atta cephalotes include individuals developing one, two, or three ocelli. These configurations can appear as only the median ocellus, only the two lateral ocelli, or even the median plus a single lateral ocellus. Interestingly, these correlations vary with changes in the size and head-to body scaling, suggesting that each ocellus can undergo different degrees of development, with one or more remaining vestigial or completely absent. On the other hand, workers in other species consistently develop a single ocellus, like in workers of Polyrachis bihamata, with no correlation to size or head-to-body scaling. These cases highlight how evolutionarily labile this trait is among workers of different ant species, which supports our proposal that the underlying gene regulatory network remains latent, thereby facilitating the emergence of novel trait combinations. We therefore agree on the importance of comparing the developmental mechanisms underlying these patterns temporally across larval stages and between individuals within a colony. We have now incorporated 2 sentences into the discussion, stating that this will be an important avenue for future work.

      Is there any function of a single ocellus in workers, or just a consequence of incomplete down-regulation of gene expression?

      Thank you again for highlighting these important points that help us to elaborate on the discussion of our study. The functional role of ocelli in species that develop these structures remains largely understudied. However, for some species particularly within the Formicinae clade the function of the three ocelli in workers has been investigated, revealing that they serve as a celestial compass that facilitates navigation. We reference these findings in our Introduction and Discussion to illustrate that the presence of three ocelli in workers can represent an adaptive trait. In contrast, the functional significance of a single ocellus or of partially developed ocelli remains an important question. This knowledge gap presents a promising avenue for future research to understand the adaptive value of reduced, partially suppressed ocellar development. We have now added a sentence in the discussion stating this.

      In previous studies, JH treatment can increase the number of ocelli in workers, consistent with its role in promoting reproductive development. In the ocellus developmental pathway, what causes the reduction of downstream gene expression in C. floridanus? Does JH directly regulate their expression?

      We thank Reviewer 2 for proposing yet another interesting question for future investigation, which we have added to the Discussion.

      The only current evidence available in C. floridanus is a recent study (MacMillan et al. 2025), in which minor workers and soldiers were treated with JH at different developmental stages. Unfortunately, no evidence of ocelli induction was observed in JH-treated individuals, suggesting that the mechanisms of ocelli development in C. floridanus might be highly canalized, especially in species that exhibit worker polymorphism (inter-individual variation in size and head-to-body scaling within the worker cate). However, more studies are required to understand why in Monomorium pharonis (no worker polymorphism) ocelli development can be readily induced by JH, while in another C. floridanus (with worker polymorphism) it appears quite difficult.

      "In D. melanogaster, the head develops from the eye-antenna disc" This statement is not correct. The brain does not belong to the eye-antennal disc.

      We thank Reviewer 2 for catching the misspelling. We have changed the name to eye-antenna disc in the sentence.

      Reviewer #3 (Recommendations for the authors):

      It is hard to see the developing ocelli in Figure 7 - could the authors increase the contrast to make them more visible?

      We have made the suggested changes to Figure 7 in the main article, and it has indeed improved the figure.

      Author response image 1.

      RNAi knockdowns in developing soldiers of Camponotus floridanus show a reduction of otd -1 expression in the brain, but no effect on otd -1 expression in the eye-antenna disc. A. HCR revealing otd -1 expression in the brain B. qPCR of otd -1 expression after RNAi knockdown shows significantly reduced otd -1 expression in the brain, C. HCR revealing otd -1 expression in the eye-antenna disc D. qPCR of otd -1 expression after RNAi knockdown shows no significant affect on otd -1 expression in the eye-antenna disc.

    1. eLife Assessment

      Building on earlier studies, this manuscript reports a role for pol kappa in cisplatin resistance in the very specific scenarios of head and neck squamous cell carcinoma, providing evidence that the PIP box of Pol kappa is critical for cisplatin resistance in these cells. The findings are of a highly focused relevance and will be useful in the field, but the conclusions are limited to very specific cancer cells. Conclusions cannot be generalized to all cisplatin resistance mechanisms and cell types and are based on incomplete evidence that presents uncertainties and discrepancies that need to be resolved.

    2. Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA-interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data. USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism. The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Overinterpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polk in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014). Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

    3. Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

    4. Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Thank you so much for appreciating our efforts to demonstrate role of Polκ mediated axes in cisplatin resistance in head and neck cancer cells.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      We completely agree with you, and the presented data only support Polκ's role in HNSCC as demonstrated in both acute cisplatin exposure as well as the cisplatin-resistant HNSCC models. Other cell and cancer types may adopt different strategies for cisplatin resistance.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      Since H357-S and SCC9-S cells are highly sensitive to cisplatin, knocking down of Polκ unlikely will alter the phenotype, as other TLS DNA polymerases like Polκ and Polκ play critical role in such lesion bypass. Since no other DNA polymerase was upregulated in these cells upon cisplatin exposure and in the cisplatin-resistant cells, it was intriguing to demonstrate a direct role of Polκ in chemoresistance and that has been proven in this study. Since the catalytic activity of Polκ is not required to induce chemoresistant in these cells, we strongly believe that Polκ-dependent mutagenesis play minimal or no role in adapting cells to tolerate cisplatin. Nevertheless, we will knock down Polκ in these cells and determine cisplatin sensitivity

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data.

      We appreciate the Reviewer's critical and insightful comment. In our view, the interaction between Polκ and USP18 is very specific as USP2 and IgG alone do not pull down Polκ. Similarly, we also show that both Polκ and USP18 interact with PCNA. We agree with the reviewer that the existence of a stable complex of PCNA-Polκ-USP18 has not been fully demonstrated in the current version. We will perform additional experiments to strengthen our finding: a) Co-IP experiments with Polκ PIP mutants (wild-type vs. mutant) should be performed to determine whether USP18 loses its ability to bind PCNA in the absence of Polκ-PCNA interaction. b) Mapping the domain in Polκ that is involved in USP18 binding and their Co-IP experiment. Additionally, high resolution co-localisation images including quantified data will be provided.

      USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism.

      We find the point raised by the reviewer is very intriguing, however, as it will require a significant amount of time and effort to demonstrate ISGylation of DDR proteins and deISGylation by UPS18, and the insight that we may gain is unlikely to add to the central theme of this paper, we will expand this in our subsequent related study. Thank you for the suggestion.

      The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Over interpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polκ in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014).

      Thank you very much for the suggestions. We will take care of the portions and modify as suggested. The necessary reference will be added as appropriate.

      Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

      Thank you for pointing this out and we will modify the portion for better clarity.

      Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Thank you for appreciating our finding that the PIP box of Polκ is critical for cisplatin response

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      We are extremely sorry for the lack of clarity. Please note that two cisplatin-resistant models (H357 and SSC9) have been used and the results were very consistent in both cells. Fig. 1C clearly demonstrates about the generation of these resistant models and the original reference has been already cited.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

      Thank you for pointing this out. In our view, the increased G2 arrest is due to more fork stalling or collapsed than the new origin firing. Also, in our assay we observed less than 10% of new origin fired DNA fibres, and that could be the reason of no significant change in new origin firing among various cells.

      Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      Thank you very much for finding our study interesting and the pending concerns will be addressed as suggested.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      In this study, we have explored eight different cell types (breast, brain, liver, head and neck, pancreatic, prostrate, lungs, and kidney) to check the expression of Polκ upon cisplatin exposure, and HNSCC cells only showed Polκ up-regulation. Therefore, we went ahead for further demonstration of the role of Polκ in cisplatin resistance in OSCC using four different cell models (H357-S, H357-R, SSC9-S, and SSC9-R). By adding more cell lines to study will unlikely change the central theme of the paper. Yes, by acquiring and analysing clinical samples from the cisplatin responder and non-responders would have strengthen our finding.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      It’s an interesting point and we will check whether overexpression of Polκ in H357-S cells could induce resistance to cisplatin and alters IC<sub>50</sub>. Thank you for the suggestion.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      We appreciate your suggestion. As suggested by Reviewer #1 also, we will check the phenotype of Polκ knockdown H357-S cells.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      Thank you for the suggestion, we will modify the text accordingly.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      Please note that GFP-Polκ has been overexpressed in H357 Polκ knockout cells to nullify the effect of endogenous Polκ, otherwise we will not be able to test the role of various Polκ mutants.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      We are extremely sorry for the lack of clarity. The inference has been derived from two sets of experiments as shown in Fig. 4C and Fig. 4D (and is with HU).

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      It has already been demonstrated in Fig. 9C where by knocking down USP18, the DDR proteins like Chk1, Chk2, CtIP, and Artemis can be recovered for ubiquitin-mediated proteasomal degradation. The same results are also obtained when its interacting partner Polκ is deleted. In our view, the presented results have sufficiently demonstrated the role of Polκ-Usp18 in the repair of cisplatin adducts through DDR proteins.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

      We appreciate your suggestion. Since the Usp18 protein is not readily available, we will not be able to show; however, we believe the interaction is direct, and we will be able to map the binding site in Polκ.

    1. eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

    3. Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and postmortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodents vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

    5. Author response:

      eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

      We thank the Editor and the Reviewers for taking the time to carefully review our manuscript and for providing constructive comments and suggestions, as well as the opportunity to revise our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

      We sincerely thank the reviewer for the positive evaluation of our study and for highlighting the significance of examining synaptic tagging and capture following theta-burst stimulation (TBS) in rodents and non-human primates.

      We agree that TBS is a physiologically relevant induction paradigm and that differences in inhibitory circuit dynamics may also contribute to the species-specific effects observed in our study. As highlighted by Larson and Munkácsy (2015), repeated bursts delivered at theta frequency (~5 Hz) can transiently suppress feed-forward inhibition through GABAB receptor-mediated mechanisms, thereby enhancing postsynaptic depolarization and facilitating NMDA receptor activation. We therefore agree that species differences in inhibitory regulation and burst-evoked depolarization may contribute to the distinct expression of synaptic tagging and capture observed between rats and non-human primates.

      We further agree that analysis of the “depolarization envelope” during sequential bursts may provide additional insight into the relative strengths of early induction mechanisms. We will therefore perform these analyses using the existing recordings and compare the depolarization envelope between rodents and NHPs in the revised manuscript. Following the reviewer’s suggestion, we will expand the Discussion section to acknowledge the potential contribution of inhibitory circuit dynamics and depolarization envelope differences during sequential bursts.

      Importantly, however, we believe that differences in downstream molecular maintenance mechanisms also contribute to these species-specific effects. In support of this, our molecular analyses revealed enhanced recruitment of plasticity-related proteins and transcriptional pathways in NHP hippocampus following TBS, including increased expression of BDNF and PKCζ. These findings suggest that both induction-related network properties and downstream molecular stabilization mechanisms may collectively contribute to the enhanced associative plasticity observed in NHPs.

      We also thank the reviewer for the important point regarding PKMζ antisense experiments and the study by Mei et al. (2011). We agree that the absence of an effect of PKMζ antisense oligodeoxynucleotides does not necessarily exclude a role for persistently elevated PKMζ in the maintenance of theta-burst late-LTP. As demonstrated by Mei et al., BDNF together with theta-burst stimulation can maintain late-LTP in the absence of protein synthesis, potentially through stabilization of PKMζ protein levels by reducing degradation rather than through de novo synthesis. However, these findings are not directly comparable to our study, since our experiments involved theta-burst stimulation alone without exogenous BDNF application. Interestingly, our results suggest species-specific differences in the interaction between BDNF and PKMζ signaling pathways. In rats, TrkB/Fc-mediated blockade of BDNF impaired TBS-LTP maintenance, whereas PKMζ inhibition alone had no significant effect. In contrast, in NHP hippocampal slices, inhibition of either BDNF signaling or PKMζ alone failed to abolish late-LTP, whereas simultaneous inhibition of both pathways disrupted LTP maintenance.

      These findings suggest that endogenous BDNF signaling and PKMζ may operate through partially redundant or compensatory mechanisms, particularly in the primate hippocampus. Therefore, although our findings indicate that de novo PKMζ synthesis may not be strictly required under the present experimental conditions, we cannot fully exclude the possibility that protein synthesis-independent stabilization or maintenance of PKMζ contributes to theta-burst late-LTP maintenance in rodents or NHPs. We will now clarify this point in the revised Discussion section.

      Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      We thank the reviewer for this important and thoughtful comment regarding statistical interpretation and biological replication. We agree that, particularly for electrophysiological experiments where multiple slices may originate from the same animal, the effective sample size for species-level conclusions should be considered at the animal level rather than solely at the slice level.

      In the revised manuscript, we will clearly indicate the number of biological replicates (animals) together with the number of slices contributing to each electrophysiological experiment, as well as the biological replicates used for qPCR and Western blot analyses. We will also clarify whether multiple slices from the same NHP/rat contributed to the same experimental condition. These details will be incorporated into the figures and figure legends wherever appropriate.

      In addition, we will perform animal-level analyses by averaging slice responses within each animal prior to statistical comparison and, where appropriate, apply hierarchical or mixed-effects statistical models to account for the nested structure of slices within animals.

      We acknowledge that the number of non-human primates (NHPs) available for this study was inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with primate electrophysiology and tissue collection. Consequently, achieving sample sizes comparable to rodent studies is often not feasible in NHP research. Nevertheless, to further strengthen the biological robustness of the findings, we are currently in the process of obtaining additional NHP brain samples and plan to repeat key experiments in an additional 3-4 animals. We believe these revisions and additional experiments will substantially strengthen the statistical rigor and overall interpretation of the study.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      We thank the reviewer for this insightful comment regarding the mechanistic interpretation of the STC findings. In the present study, we selected the 30 min interval based on well-established classical STC paradigms in rodents, where this interval reliably falls within the effective tagging and capture window. Using this experimentally validated interval allowed us to directly compare whether TBS is sufficient to support STC in primates versus rats under equivalent experimental conditions. Accordingly, the primary objective of this study was to determine whether TBS-induced STC varies across species, rather than to comprehensively define the temporal dynamics of the tagging window.

      We agree, however, that the current experiments do not distinguish whether the primate-specific effect reflects prolonged tag persistence, enhanced plasticity-related protein (PRP) synthesis, altered capture efficiency, or a shifted temporal window. Addressing these possibilities would indeed require systematic temporal interval analyses (e.g., ±15, ±30, ±60, and ±90 min), which represent important future directions. Such experiments are particularly challenging in non-human primates because the availability of primate tissue and experimental resources for large-scale electrophysiological studies remains limited and is currently beyond our experimental capacity due to substantial ethical, logistical, financial, and technical constraints.

      Nevertheless, we fully agree with the reviewer that these experiments are important for advancing the mechanistic interpretation of the findings. Similar temporal analyses have recently proven informative in our rodent studies (Chong YS, Ang SR, Sajikumar S. Commun Biol. 2025;8:553). Importantly, we are currently in the process of obtaining additional non-human primate samples and plan to extend the present work by examining an additional 60 min temporal interval to further characterize the temporal properties of synaptic tagging and capture in non-human primates.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and post-mortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      We thank the reviewer for raising this important and insightful point. We agree that differences in developmental stage between the experimental groups represent an important consideration when interpreting potential species-dependent effects. In the present study, rat experiments were performed in 5-7 week-old animals, whereas non-human primate (NHP) tissues were obtained from 5-7-year-old monkeys. This difference largely reflects the practical, ethical, and logistical constraints associated with NHP research and tissue availability. We acknowledge that these ages are not developmentally equivalent and that maturation state may influence BDNF signaling, protein synthesis capacity, synaptic plasticity thresholds, and transcriptional responses relevant to late-LTP and STC mechanisms.

      We also recognize that differences in euthanasia procedures, tissue extraction, slice preparation, and postmortem handling between rodent and primate tissues may influence tissue physiology and electrophysiological properties. Although extensive care was taken to optimize tissue viability and maintain stable recordings within each species, these variables cannot be completely excluded as contributing factors to the observed differences.

      Accordingly, we will revise the Discussion section to more explicitly acknowledge these limitations and clarify that our findings support potential species-dependent differences under the present experimental conditions, rather than definitive intrinsic species-specific mechanisms. Nevertheless, despite the inherent challenges associated with NHP electrophysiological studies, we believe that the present findings provide an important initial framework for understanding the translational relevance of synaptic tagging and capture mechanisms across species.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      We thank the reviewer for this important comment regarding statistical reporting and interpretation. We agree that the repeated occurrence of identical exact p-values in several nonparametric analyses reflects the relatively small sample sizes and the discrete nature of the statistical distributions. This issue is particularly relevant for the NHP experiments, where biological replication is inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with obtaining and processing primate tissue.

      In the revised manuscript, we will provide exact n values for all comparisons, including the number of biological replicates (animals) and slices where applicable. We will also include additional statistical details, including effect sizes and confidence intervals where appropriate, to improve transparency and facilitate interpretation of the reported findings. Furthermore, we are currently in the process of obtaining additional NHP samples and will attempt to include more biological replicates in the revised version to further strengthen the robustness of the analyses.

      We also agree that the issue of multiple testing should be addressed more explicitly, particularly because multiple genes and proteins were examined. In the revised manuscript, we will clearly state the statistical correction methods applied for multiple comparisons where appropriate. For analyses in which corrections were not applied, we will provide justification, noting that several experiments were based on hypothesis-driven candidate targets rather than exploratory large-scale screening analyses. These statistical considerations will be clarified in the Methods and Results sections.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

      We thank the reviewer for this thoughtful and balanced assessment of our work. We agree that the present data primarily support the conclusion that, under the specific experimental conditions examined, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than that observed in rat slices. We also agree that broader interpretations regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and potentially redundant BDNF/PKMζ-related mechanisms require additional mechanistic investigation and experimental validation.

      Accordingly, we will moderate these interpretations throughout the revised manuscript and clearly state that these conclusions remain preliminary. We will further emphasize that additional experiments, including increased biological replication, expanded temporal analyses, and further mechanistic investigations, will be necessary to more conclusively define the basis of the observed species-dependent differences. Within our current experimental capacity, we are actively working to obtain additional non-human primate samples and plan to incorporate additional biological replicates and key follow-up experiments in the revised version to further strengthen the robustness of the findings.

      At the same time, we believe the present study provides an important initial contribution to an understudied area by directly examining synaptic tagging and capture mechanisms in the primate hippocampus. Given the limited availability of non-human primate electrophysiological data in the field, these findings may offer a valuable framework for future studies investigating the translational and evolutionary relevance of associative synaptic plasticity mechanisms across species.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodent’s vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

      We thank the reviewer for this thoughtful and insightful comment, as well as for the encouraging appreciation of our long-duration plasticity recordings and associative plasticity experiments, which are both technically demanding and time-intensive. We fully agree that interpretation of cross-species differences in synaptic plasticity requires careful consideration of multiple biological and environmental variables, including circadian state, enrichment conditions, strain differences, diet, lighting conditions, and species-specific behavioral ecology.

      Regarding the specific concern related to circadian phase and sleep-wake state, the reviewer raises an important point. Rats are nocturnal animals, whereas macaques are diurnal, and hippocampal plasticity mechanisms are known to be influenced by circadian rhythms and sleep-dependent regulation of synaptic proteins and signaling pathways. Previous studies have demonstrated modulation of LTP, synaptic tagging and capture and protein synthesis in rats across normal sleep-wake cycles. We therefore agree that these factors may influence plasticity outcomes and should be carefully considered in comparative studies.

      Studies have further shown that theta frequency is highly sensitive to sleep-related manipulations. Specifically, theta frequency decreases immediately after sleep, remains elevated during sleep deprivation, and rapidly declines following recovery sleep. In aged animals, these effects appear comparatively attenuated, suggesting reduced sleep-dependent modulation of theta dynamics with aging. Therefore, disruption of normal circadian or sleep-wake patterns may significantly alter theta activity and associated plasticity mechanisms within a species and may not accurately reflect physiological baseline states (Utku Kaya et al., 2026).

      In our experiments, recordings from rats and macaques were performed during their respective active phases under standardized laboratory housing conditions, and we will further clarify these details in the revised Methods section. Nevertheless, we acknowledge that circadian state and related physiological variables cannot be completely excluded as contributing factors to the observed differences between species.

      More broadly, we agree with the reviewer that the present study does not permit definitive conclusions regarding universal “rodent versus primate” rules of synaptic plasticity. Our intention was not to propose a generalized dichotomy between rodents and primates, but rather to report that, under the experimental conditions used here, SC-CA1 TBS-LTP and associated synaptic tagging mechanisms differed between rats and macaques. We agree that broader evolutionary or cognitive interpretations would require systematic comparative analyses across multiple species, including both nocturnal and diurnal rodents as well as diverse primate species. Such studies would provide a stronger framework for distinguishing conserved versus species-specific mechanisms of plasticity.

      At the same time, we believe the present findings remain important because they provide one of the first direct experimental comparisons of SC-CA1 TBS-LTP-associated plasticity mechanisms between rodents and non-human primates under controlled ex vivo conditions. Although the interpretation should be done cautiously, the observed differences raise the possibility that certain metaplastic or protein synthesis-dependent mechanisms may not be fully conserved across species. Accordingly, we will revise the Discussion section to better emphasize the exploratory and comparative nature of the study, while explicitly acknowledging the limitations and potential confounding factors highlighted by the reviewer.

    1. eLife Assessment

      This important study assessed the replicability of a selection of lab-based biomedical experiments in papers published by authors based in Brazil. The study adds a unique perspective to the literature on replication, and provides rich data on the approach taken, the outcomes, and the challenges involved in conducting large-scale crowd-sourced research. The evidence supporting the claims is convincing, but there is scope for clarifying the presentation of the results and extending the discussion section.

    2. Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      (2) The article appears to oscillate between:

      i) a description of the approach and the inherent challenges of such a multicenter replication program.

      ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections. Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers

    3. Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was under represented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX). In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this. If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications. I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean. An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made. In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

    4. Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      We thank the reviewer for the comments.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      We can try to simplify the description of protocols at specific points for example, by providing an overarching description of the study design in the beginning of the Methods, rather than citing our previous eLife paper (Amaral et al., 2019), as suggested below. The methods are indeed quite extensive, but the this may be inevitable in a large-scale project such as this and we note that Reviewer #2 thought that part of the supplementary material should be incorporated back in the main text, which is a suggestion in the opposite direction. It may thus be hard to strike a balance between readability and comprehensibility that can address both reviewers’ opinions.

      (2) The article appears to oscillate between:

      (i) a description of the approach and the inherent challenges of such a multicenter replication program

      (ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      There is a bit of redundancy between tables and text, but this was intentional to make both of them self-explanatory. We also think stating the results in the text can allow us to make each of the replication criteria clearer, a concern that was also mentioned by the reviewer.

      As for requiring particular expertise in statistics for understanding, we mostly disagree. The main results (Tables 1 and 2, Figure 2) are expressed as percentages, and the only statistical concepts needed for interpreting these results are understanding prediction and confidence intervals. For this, we could provide a bit more guidance on their interpretation in the Methods section. Beyond that, most of the secondary results (e.g. Figure 3 and Figure 4) involve linear correlations, which is about as simple as statistical analysis gets.

      Of the results presented in the main manuscript, only Table 3 contains anything beyond percentages and correlations. We do agree that the meaning of each ratio in this table could be more clearly described, but there are essentially no expert-level statistics involved in their calculations.

      Other than that, the main statistical issues are the ideal way to aggregate the results from different replications for which we use different strategies for robustness purposes. However, all of these results are already in the supplementary material, so we don’t feel they interfere to much with the readability of the main manuscript.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections.

      This is indeed a good idea, and we plan to include an initial overarching description of the project in the Methods section of the revised manuscript.

      Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      Again, this is the opposite of what was suggested by Reviewer #2, so we would rather keep the Methods section more or less at its current level of detail.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      The reviewer is correct in his interpretation. Evaluating the main findings of articles or cleaning a field of wrong statements was never a goal of our study (and we were clear about this from the start). Our aim with the project was metascientific (i.e. evaluate the reproducibility of biomedical experiments with a set of common methods) rather than driven by a particular interest in the findings themselves. This is reflected by our choice of selecting experiments from a random sample of articles from multiple fields, rather than filtering by area of interest or importance. It also underlies our choice to evaluate experiments rather than claims, as this was more statistically tractable and potentially more objective as a meta-research goal.

      To be clear, we don’t feel this approach is inherently better or worse than evaluating claims in the literature, as in the Drosophila immunity article case (i.e. Westlake et al., 2026), which is also an important goal. They are merely approaches that answer different questions. Ultimately, we probably made our choice based on (a) our expertise/interest in meta-research rather than in the fields the replications stemmed from and (b) an attempt to engage Brazilian researchers in the project in a way that was non-confrontational and minimized backlash from their peers. We feel this was valuable for many of the lessons learned, although it also meant learning less about the research findings in question.

      Even though this was not a goal of the study, there is some knowledge obtained about the findings that is indeed largely absent from the current manuscript. We do not feel the current format allows for much discussion of 45 different findings, but we do have plans to address these in future articles (as outlined in our response to point 5). In the meantime, qualitative descriptions of each experiment can be found at https://osf.io/w5z9a. This is already mentioned in the Methods but could be reiterated in the results as well.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      We did not go too deep into that finding because we are publishing a separate article focused on the prediction project, which should look into factors that correlate with prediction accuracy, both at the level of predictors (e.g. research field, career level) and of individual predictions (e.g. information taken into account for each answer). We also feel that, given the multiplicity of predictors in the prediction analyses, these findings are a bit tentative, as the strongest predictors may be subject to effect size inflation from the “winner’s curse” effect (as outlined by Reviewer #2). We can try to emphasize it a little more in the discussion (although it already merits a whole paragraph on pages 23-24), but we feel we would be able to discuss it more critically in a follow-up article.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers.

      Indeed, some of our results did not fit this overarching analysis and were left for future publications. One of them is already available as a preprint, while the others are currently in preparation. Specifically, other results from the project should be spread about across five different articles.

      (a) A narrative article focused on challenges and lessons learned with the project, already published as a preprint at https://osf.io/preprints/metaarxiv/8y3tg_v1 (Amaral et al., 2026).

      (b) An article analyzing the prediction survey and markets results in detail (following the pre-analysis plan detailed in https://osf.io/6av7k/files/pjhgd and adding some exploratory analyses on prediction rationales).

      (c) Three articles describing the results of specific experiments with each experimental method (MTT, PCR, elevated plus maze) along with a discussion of aspects inherent to the method that seem to influence reproducibility.

      We can add this information more explicitly to the Methods section, including the links to the papers that have already been published at the time the manuscript is revised.

      Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was underrepresented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Thanks!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      We once more thank the reviewer for the compliments.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      We do acknowledge that the article currently includes a lot of supplementary material. This includes both supplementary figures/tables relating to the paper and many supplementary methods files (mostly hosted at the Open Science Framework). However, we also note that this is already a rather long paper as it stands and that Reviewer #1 has made the opposite suggestion of simplifying it. Thus, it may be hard to strike a balance that will suit all preferences, and we feel that maybe our attempt has landed somewhere in the middle of both reviewers’ ideal versions of the paper.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      We can try to engage with some of the above-mentioned literature in more depth in particular replication studies from other fields (some of which have appeared after our preprint (e.g. Tyner et al., 2026) and with the risk of bias and transparency literature (e.g. Serghiou et al., 2021). That said, we note once more that the article (and the Discussion section) are already quite long, and that analyzing each of these articles in depth is likely to be unfeasible.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Our original plan was to use p values as a predictor (see protocol at https://osf.io/9rnuj), but we later realized this was inadequate as it did not account for effect direction (i.e. significant effects in the opposite direction as the original may yield low p values, but this should not count as replication success). We thus switched to t values to be able to assign positive and negative signs depending on effect size direction. We note that, as we are using non-parametric Spearman coefficients (in which the module of t correlates negatively with the p value), the two approaches are effectively equivalent when original and replication effects have the same direction. This change was accounted for and justified in our list of protocol deviations at https://osf.io/9hj7t.

      Effect size (in relative terms) is already being used in the second predictor in the analysis (i.e. effect size decrease), as our idea was to use one significance-based predictor and one effect size-based predictor, to match what was done for the replication rates). We feel that using relative effects (e.g. response ratios) by themselves may not be as adequate, as for experimental methods with large coefficients of variation and/or low sample sizes (especially PCR ones), one can find large relative effects that are nevertheless far from statistical significance. This also makes relative effects not very commensurable between methods.

      We do believe there is a fair argument, however, to use standardized effect sizes as an alternative to t values (i.e. difference measured in standard errors of the mean) to measure significance/evidence strength. As some replications ended up underpowered, low t values may sometimes be due to insufficient statistical power/low sample size rather than replication failures. Using standardized effect sizes is not devoid of pitfalls (e.g. they can be quite variable when sample size is low), but it is worth doing as a robustness analysis.

      That said, there are a few statistical issues to be decided on how to calculate this (e.g. whether studies should be meta-analyzed using standardized mean differences rather than relative ones for this purpose, or whether an analog of the standardized effect size should be calculated for the log ratio of means). We would have to look more carefully into the multiple possibilities to decide on the best approach (and we do accept suggestions!).

      In the meantime, we note that running the prediction analysis using only experiments with ≥80% power yields a slightly higher correlation of t scores with researcher predictions (ρ = 0.49, p = 0.005), so we do not think that these underpowered experiments affect the trend too much. If anything, they could be masking a higher correlation between researcher predictions and replicability.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      We agree that we should make the description more precise (e.g. “reaching the same results when analyzing a set of data in the same way” for reproducibility and “finding similar results with new data collected under similar conditions” for replicability). We will update these definitions in the revised manuscript.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      We would argue that the reader would be correct in this case: the argument is a bit speculative. It does go in the direction of what is generally accepted within the field (i.e. that publication pressure can lead to lower reproducibility for a range of factors), but we’re not sure this connection has been demonstrated empirically, except for indirect evidence (such as the lower reproducibility in papers stemming from top institutions and “trophy journals” in, the higher frequency of positive results in US states with more researchers in Fanelli, 2010, or the higher number of problematic images for highly productive researchers in some countries in Fanelli et al., 2022. We could cite this evidence in the introduction and make the speculated connection more explicit, perhaps adding modeling work as well (e.g. Ioannidis, 2005; Smaldino & McElreath, 2016) to explain why this could be the case. But essentially, our opinion is that the connection remains a speculation.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      We can offer a more detailed description of the recruitment process (e.g. number and distribution of lectures, social media strategy used, etc.), although we would rather do this in a supplementary document so as not to make the Methods section even lengthier. We note, however, that we never aimed to recruit a “representative sample” of labs from the country: we were busy enough trying to get enough labs for the project to happen, and aware that the call would be inevitably biased by our own communication capabilities and personal networks.

      That said, the response rates for different regions of Brazil do generally match the distribution of research labs and graduate programs within the country (with some distortions likely caused by our personal networks, such as the large number of labs in Rio de Janeiro state), and seem to indicate a rather wide dissemination of the call. One way to visualize this would be to present the distribution of corresponding articles from the original studies selected for the replication (or even from the whole sample of articles obtained for experimental selection) along with the distribution of labs at different stages of the project in Figure S3, which generally show similar patterns. This would actually lend support to our statement that “the population of labs that performed replications was largely similar to the one that produced the original results” in the discussion.

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      We tried to find the combination of methods that would maximize the number of labs that would be included in the project. This is explicitly stated in our Methods Selection document at https://osf.io/qxdjt, but could be stated more explicitly in the paper as well.

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      We initially used single screening by three different reviewers (see https://osf.io/6av7k/files/u5zdq for criteria), as we were merely looking for a sample of experiments; thus, comprehensive inclusion of all eligible studies was not a priority. After this initial screening step, inclusions were confirmed in a consensus meeting with the three reviewers involved.

      Data extraction was also done by a single individual, but the resulting data led to a protocol that was later checked by two reviewers who had access to the paper and were explicitly oriented to judge whether the protocol consisted in a valid replication. Thus, discrepancies between what was in the paper and what was included in the protocol could potentially be flagged at these stages (as they were in many cases). We do note, however, that this is likely not as effective to prevent errors as having data extracted independently, as reviewers may overlook mistakes more easily when comparing two documents rather than extracting data anew. We did find that some errors in extraction slipped by, such as an MTT experiment where treatment concentration was inadvertently changed from mM to μM in a particular protocol step; this was picked up and corrected by 2 out of the 3 labs, but not by the third one, leading the latter replication to be invalidated.

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      As stated in the manuscript, we initially capped experiments at a predicted cost of R$ 5.000 (around USD 1336 at that time), considering reagent cost alone (as equipment and labor was provided by labs), as mentioned in the manuscript. Exclusion rates for that reason were 12/74 (16%) for MTT experiments, 36/132 (27%) for PCR ones and 4/40 (10%) for EPM ones. This is stated at

      This turned out to be an underestimation in many cases, especially as it did not account for pilot experiments, need for repetition, etc; thus, many experiments ended up costing considerably more than that ceiling. As we had included a contingency fund for those cases which we expected would occur , we avoided removing experiments from the sample for this reason as much as possible. Nevertheless, one elevated plus maze experiment ended up not being replicated for cost reasons, as the necessary rat strain was provided by a single facility in the country, meaning that a large number of rats would have to be acquired and transported to all labs at a cost that we were not able to cover.

      As these costs were covered by the coordinating team, we do not feel that this is likely to underlie the reduction in geographical coverage. Other reasons related to lab structure could have led to labs in less well-resourced regions to leave the project, but they probably has nothing to do with the experiments selected.

      That said, the cost cap does mean that the selection of experiments is not completely representative of the literature, but is enriched in relatively cheap and simple experiments which were able to perform (which was our next step for selecting the final sample of experiments. Exclusion rates due to lack of lab expertise and/or infrastructure to perform the experiment were 21/56 (37%) for MTT experiments, 67/89 (75%) for PCR ones and 7/34 (21%) for EPM experiments.

      We will try adding some of this information to the flowchart in Figure 1, as we agree it provides more context on the representativeness of the selected experiments.

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      The scale ranged from 1 (No relevant differences) to 5 (Very relevant differences that prevent considering the study as a direct replication). This scale was used for both the lab and the validation committee scores, and is described at https://osf.io/xgth2 (debriefing protocol) and https://osf.io/e3fjg (validation protocol).

      For the validation committee, we did use a threshold (any score of 4 or a sum of scores of 10 or more among 3 evaluators) to decide what had to be discussed to decide on inclusion, as mentioned on Page 7 of the Methods. For the labs, we used no threshold labs answered the protocol deviation question as a scale, but the decision of whether to consider the study a valid replication or not was not tied to this score.

      We can make both of these points (meaning of the scale and connection to lab’s decision to consider the replication valid) clearer in the Methods section.

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      For the initial analysis of justifications, one reviewer read all answers and flagged those that seemed to concern reproducibility of the methods (e.g. “we replicated the protocol exactly as planned”) rather than results reproducibility (e.g. “effects went in the opposite direction”). We then revised these answers among the whole coordinating team to decide whether we should contact the lab asking them to revise them. We can add this information to the Methods section.

      For classifications of the justification into categories (i.e. Table S7), justifications were classified by two independent reviewers based on categories created after an initial inspection of the data, and discrepancies were resolved by consensus. We can add this information to the table legend.

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      After we extracted data from the lab spreadsheets and summarized the results by code, labs received the results by e-mail and were asked to fill in a form on whether the results were in agreement with what they had found (see details at https://osf.io/nfr6y). Discrepancies in results at least 1 experiment were noted by 36% of the 53 (out of 56) labs that responded. Many of these stemmed from the coordinating team misunderstanding issues such as group identity or experimental unit identification in the spreadsheet. Others had to do with different ways to perform calculations (e.g. relative gene expression or % time spent in open arms). In some cases, simple errors in data transcription or typos caused the discrepancy.

      We were also surprised (and concerned) by the number of experiments in which we later found data errors that were not detected by this process (e.g. 18% of total). Our best understanding of this is that not every lab checked the results with the necessary care, as some errors were quite obvious, as in experiments in which sample size was different, or in which group labels were reversed. Ultimately, agreeing with a form that says “did you find any discrepancies?” may have been performed as a box-ticking exercise with little attention, and was probably not the ideal way to check data which led us to start reviewing results in live meetings afterwards. This is discussed in more detail in our challenges article (Amaral et al., 2026)

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX).

      R 4.5.1 was used for the analysis. We can add this information (which was present in the data repository in the R session info.txt file) and provide the R reference in the manuscript as well.

      In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this.

      If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Yes, we did use the escalc() function for this calculation (for both the replications and the original effect sizes). We can mention this in the manuscript.

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications.

      We calculated the coefficients of variation as the pooled SD divided by the mean of both group means. The reviewer is correct about the possibility of small-sample effects in this case (which we were not aware of). We will thus look into the possibility of implementing this via the escalc () function in the analysis of the revised manuscript.

      We also acknowledge that this could be a source of bias in the comparisons between original and replication CVs (albeit likely a minor one). That said, we note that sample sizes are not always larger in the replication for some experiments with large original effects, power calculations sometimes yielded lower sample sizes in the individual replication, albeit infrequently. On average, though, replication sample sizes were indeed larger.

      I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean.

      This is indeed the case; that said, the CV of the original effect also has random error relative to the true population CV and in that case, there is no way to estimate the uncertainty, as we have a single measure of that parameter. So there is probably no way around ignoring uncertainty in this case.

      We also note that we are looking for evidence of systematic CV inflation across all experiments (rather than for a statistically robust comparison between the CVs of any individual replication). For the sake of measuring this systematic inflation, the use of multiple experiments does allow us to estimate variability at the experiment level which should incorporate the lower-level variability between individual replications if this is not included in the model. Thus, we do not feel that our procedure introduced a systematic bias in the analysis at the experiment-level (although one could argue that it may lead to less precision).

      An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      We thank the reviewer for this suggestion, which indeed seems like an option in this case. We will look into this possibility, although we cannot guarantee at the moment that we will implement it, as we were not previously familiar with the method and will have to study it in more detail.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Yes, but there are likely different reasons for that. Labs leaving after being included may have been due to those in less privileged regions of Brazil (e.g. the northern and western regions of Brazil, generally speaking) having more difficulty in persisting in the project. That said, most of the “disappearance” happens between registration and inclusion which usually has to do with the labs not working with the methods that were ultimately included in the project. We also note that most of the states that lose representation were those that had a single lab to begin with, which may make the visual pattern more striking than the actual trend (as states in the South/Southeast also lose labs, but don’t disappear from the map).

      We note again that we never planned to achieve geographical representativeness when recruiting the labs on the contrary, we were aiming to maximize the number of available labs to run the project. That said, we do agree that for the sake of examining whether the population of labs is similar to the one that generated the original experiments (a claim that we do make in the discussion), this representativeness is important to assess. Once more, to allow the reader to evaluate this, we plan to add an additional map to Figure S3 to describe the Brazilian states where the original experiments came from (based on corresponding author affiliations) in which a similar bias towards the South and Southeast Region can be observed.

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      We agree that this would be useful information, and can experiment with the possibility, but our feeling is that the figure will likely become too noisy in cases where the 95% CIs overlap (which are quite frequent). If this is indeed the case, an option to allow the reader to examine this would be better to add an explicit link to the forest plots for each individual experiment (https://osf.io/sx9gv) in the figure legend.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      We selected the highest correlation values for each continuous outcome (t score and lnRR) and presented these separately in the text. This is a systematic way to perform the selection, but is obviously subject to the “winner’s curse” effect. We agree that adding both metrics for each predictor would be a fair way to keep this in perspective for the reader, but we would have to think about how to do this without sounding too confusing (as results for the two main outcomes are quite different).

      We do note, however, that the outcomes are indeed different and are expected to vary independently in some cases. For the correlation with replication probability predictions, for example, the effects in opposite directions would likely be expected, as larger original effect sizes will likely lead to larger probabilities to be assigned, but also to a higher possibility of effect size decrease. This low correlation between outcomes is probably something that should be pointed out and discussed in the revised manuscript.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made.

      We had given this some thought when writing the manuscript – but ultimately opted not to include confidence intervals for our replication percentages and to use the replication rates as descriptive measures only (as done in other replication studies such as (Errington et al., 2021).

      Even though we aimed for our sample of original experiments to be as systematic as possible, it is ultimately constrained by many factors (the choice of methods, the particular expertise of the labs, etc.) thus, adding confidence intervals represents the uncertainty around the replication rate of a very specific population of experiments, which is not directly comparable to those included in other replication efforts in any case.

      We will reconsider whether we should include confidence intervals for replication rates: although doing this for every replication rate in Table 1 and Table 2 may end up being too much information, it could probably be done at least for the replication rates of the main analysis in the text. We note that calculating confidence intervals for percentages is straightforward, requiring only the numbers that are in the table thus, any reader that wants to estimate uncertainty for those rates should be able to do it easily.

      We will also point out the uncertainty around the percentages mentioned in the discussion when comparing our replication rates with those of other studies, which we agree is an important issue to touch on.

      In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Both correlations in that case are non-parametric (e.g. Spearman’s ρ), so they cannot be directly transformed into Pearson’s r without making assumptions about the distribution (which we would probably avoid doing given the very marked outlier in our own). We can calculate a non-parametric confidence interval for our own correlation coefficient by resampling, but we will have to investigate whether this can be done using the available data from (Errington et al., 2021) (which is probably the case if effect sizes for all experiments have been shared).

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      It is indeed interesting, and seems to confirm an intuition that has long been present in the reproducibility field, but actually has little evidence to support it: if anything, there is evidence in the opposite direction in psychology (Youyou et al., 2023), although they looked at cumulative publication number, while we used number of publications in a fixed interval.

      We can expand a bit further on that finding: that said, we do note that the correlation is relatively weak and has a p value of 0.04. Thus, given the multiplicity of predictors would not be that unlikely to occur by chance, even though it seems intuitive. Thus, even though the relationship seems intuitive, we think it should be considered tentative at best and would refrain from discussing it in too much detail.

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Yes, that is what we mean by “incorrect labeling of error bars” (as can be grasped from the cited references).

      We can perform this regression, which seems relatively straightforward to do. That said, we note that another likely cause for outliers at least for cell line studies would be the use of different (and eventually inadequate) experimental units (e.g. having error bars that represent technical replicates of the same measurement rather than truly independent experiments). We suspect that this may have an even greater effect in terms of causing error bars not to express the same thing and the regression will not help in differentiating the two causes.

      We should also note that different types of experiments may be expected to have very different SDs, so the regression is likely to have a lot of error associated with it. In particular, it’s probably worth doing separate regressions for each method, to account for the likely difference in CVs between animal and cell line experiments, for example. This could also help tease apart the two causes above, as the experimental unit problem mentioned above will likely only be observed for cell experiments.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

      Thanks!

      Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      We thank the reviewer for the compliments. Again, a more extensive list of insights can be found in our challenges article (Amaral et al., 2026), which we will cite in the revised version.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

      We acknowledge the missing replications as a weakness, and we hope we have made that point clear in the discussion.

      Concerning the Brazilian research ecosystem, we could try to explore this in more detail in the introduction. In particular, we believe that a better understanding of the Brazilian academic system, including its regional disparities and the general composition of its workforce (which is largely composed of undergraduate and graduate students), can be useful in interpreting some of the findings.

      We can try to provide a bit more context at the end of the introduction (perhaps between the last 2 paragraphs, which would also address a point made by Reviewer #1), and also in different points of the discussion including those comparing replication rates with other studies or discussing infrastructural difficulties, some of which may be specific to the Brazilian context (such as difficulties in acquiring specific reagents or licenses). Still, we reiterate that, due to the lack of studies with comparable samples in other regions, we cannot tease apart the factors that are specific to Brazil from those affecting lab biology as a whole from the data alone.

      References:

      Amaral OB, Neves K, Wasilewska-Sampaio AP, Carneiro CF. 2019. The Brazilian Reproducibility Initiative. eLife 8:e41602. DOI: https://doi.org/10.7554/eLife.41602

      Amaral OB, Valério B, Carneiro CFD, Mota GPS, Neves K, Abreu M, Tan PB. 2026. Challenges for building up confirmatory science in lab biology: lessons learned from the Brazilian Reproducibility Initiative. MetaArXiv, DOI: https://doi.org/10.31222/osf.io/8y3tg_v1

      Errington TM, Mathur M, Soderberg CK, Denis A, Perfito N, Iorns E, Nosek BA. 2021. Investigating the replicability of preclinical cancer biology. eLife 10:e71601. DOI: https://doi.org/10.7554/eLife.71601

      Fanelli D. 2010. Do pressures to publish increase scientists’ bias? An empirical support from US states data. PLoS One 5:e10271. DOI: https://doi.org/10.1371/journal.pone.0010271

      Fanelli D, Schleicher M, Fang FC, Casadevall A, Bik EM. 2022. Do individual and institutional predictors of misconduct vary by country? Results of a matched-control analysis of problematic image duplications. PLoS One 17:e0255334. DOI: https://doi.org/10.1371/journal.pone.0255334

      Ioannidis jpa. 2005. why Most Published Research Findings Are False. PLoS Medicine 2. DOI: https://doi.org/10.1371/journal.pmed.0020124

      Serghiou S, Contopoulos-Ioannidis DG, Boyack KW, Riedel N, Wallach JD, Ioannidis JPA. 2021. Assessment of transparency indicators across the biomedical literature: How open is open? PLOS Biology 19:e3001107. DOI: https://doi.org/10.1371/journal.pbio.3001107

      Smaldino PE, McElreath R. 2016. The natural selection of bad science. R Soc Open Sci 3:160384. DOI: https://doi.org/10.1098/rsos.160384, PMID: 27703703

      Tyner AH, Abatayo AL, Daley M, Field S, Fox N, Haber NA, Hahn KM, Struhl MK, Mawhinney B, Miske O, Silverstein P, Soderberg CK, Stankov T, Abbasi A, Aberson CL, Aczel B, Adamkovič M, Albayrak N, Allen PJ, Andreychik M, Awtrey E, Axxe E, Azevedo F, Bader MD, Bago B, Bailey J, Bakker M, Banik G, Banks GC, Baskin E, Batruch A, Beatteay A, Behr SM, Berente N, Berry Z, Białkowski J, Bodroža B, Boeschoten L, Bognar M, Bokhove C, Bonfiglio D, Bouwman R, Brady TF, Braithwaite SR, Briceño Jiménez G, Brick C, Bricka T, Briker R, Brown AN, Brown GDA, van Aert RCM, Caldwell K, Capitan S, Capitán T, Chandler J, Charles T, Chartier CR, Chawdhary R, Cheng KJ, Chopik WJ, Clark B, Colvin VE, Comer CC, Costantini G, Coupé T, Cummins J, Czernatowicz-Kukuczka A, de Leeuw J, Dobolyi D, Druckman JN, Duan J, Dujmović M, Dunleavy DJ, Durkee PK, Emery C, Esterling KM, Evans TR, Fedor A, Fernández-Castilla B, Fiala N, Field JG, Fong N, Fonseca MA, Freeman ALJ, Freese J, Geiger SJ, Geng J, Getz LM, Geven LM, Gleibs IH, Gonzales DP, Gooty J, Gourdon-Kanhukamwe A, Greculescu C, Griffin SM, Grigoryan L, Grunow M, Gunby N, Hall B, Hanel PHP, Hannon EE, Harper S, Held MJ, Hickman L, Higgins NC, Hippel S, Hoeppner S, Hong S, Hostler TJ, Inzlicht M, Izydorczak K, Jaeger B, Jankowsky K, Jarke-Neuert J, Jensen M, Jokić B, Jolles D, Jolly P, Jones AM, Juanchich M, Kačmár P, Kapoor H, Keljanovic A, Koirala S, Kołczyńska M, Kouroupaki D, Kühnen U, Landgrave M, Larson MJ, Laulié L, Lawrence ACE, Le Forestier JM, Leahy KE, Lee S, Leslie J, Lewis SC, Limnios C, Lin H, Liu A-C, Lloyd JW, Ludvig EA, Lynott D, MacDonald J, Mallik P, Mallinson DJ, Marinazzo D, Martarelli CS, Matacotta J, McBride A, McHugh C, McMillan G, Méndez E, Metzger M, Michaelides MP, Michalak J, Micheli L, Miller JK, Milyavskaya M, Molden DC, Monjaras AG, Moreau D, Morrow A, Moya C, Mudrik L, Mulder LB, Munt KA, Nandi A, Nason K, Nast C, Nave G, Nax HH, Neubauer F, Nguyen PLL, Nichols AL, Nilsonne G, O’Boyle E, Oettinghaus J, Oh J, Oshana A, Ostermann T, Ostrowski RP, Oyebanjo A, Panczak R, Patrianakos J, Pavez I, Pavlov YG, Persson S, Perugini M, Peters K, Pieters C, Ponizovskiy V, Porter ND, Prenoveau JM, Purić D, Purol MF, Puthillam A, Quinn KA, Ramljak M, Reed WR, Ritchie M, Ritzau M, Roche SP, Rodela R, Röer JP, Ropovik I, Rothschild J, Saal J, Safadi H, Samaha J, Sanchez M, Sankaran S, Santos D, Sargent AC, Sauter M, Schmidt K, Schnabel L, Schroeder AN, Schuetz SW, Schuetze BA, Schulte-Mecklenbeck M, Schütz A, Sevigny EL, Shackleton E, Shafranek RM, Shaki S, Shakya S, Sirota M, Sisco MR, Sitnikov MM, Slevc LR, Smalarz L, Smith CT, Snyder JS, Sommet N, Sonmez F, Spellman BA, Stanulewicz-Buckley N, Stock G, Street CNH, Strømland E, Sundelin T, Syed M, Szabelska A, Szaszi B, Szumowska E, Tagat A, Täuber S, Tay L, Thapa S, Thatcher J, Tsaklakidou D, Tummers L, Turkovich E, Tutor MV, Urbanska K, van ’t Veer AE, van Assen M, van de Ven N, van den Goorbergh R, Vargo EJ, Vaughn LA, Vazire S, Vermeulen JM, Vo DTH, Volkman V, Wagenmakers E-J, Wagner D, Walasek L, Walter F, Warmelink L, Wei L, Weißflog MI, Weller N, Wichman AL, Wilbiks J, Williams JR, Wolfe K, Wort F, Wright R, Wulff JN, Xue X, Yan VX, Yang Y, Yoon S, Žeželj I, Zhang Y, Ziano I, Zogmaister C, Zupan Z, Zwaan RA, Nosek BA, Errington TM. 2026. Investigating the replicability of the social and behavioural sciences. Nature 652:143–150. DOI: https://doi.org/10.1038/s41586-025-10078-y

      Westlake H, David F, Tian Y, Krakovic K, Dolgikh A, Juravlev L, Bournonville TE de, Carboni A, Melcarne C, Shan T, Wang Y, Mu Y, Kotwal A, Pirko N, Boquete JP, Schüpfer F, Rommelaere S, Poidevin M, Liu Z, Kondo S, Ratnaparkhi GS, Chakrabarti S, Liu G, Masson F, Xiaoxue L, Hanson MA, Jiang H, Cara FD, Kurant E, Lemaitre B. 2026. Reproducibility of scientific claims in Drosophila immunity: A retrospective analysis of 400 publications. eLife 15. DOI: https://doi.org/10.7554/eLife.108404.1

      Youyou W, Yang Y, Uzzi B. 2023. A discipline-wide investigation of the replicability of Psychology papers over the past two decades. Proceedings of the National Academy of Sciences 120:e2208863120. DOI: https://doi.org/10.1073/pnas.2208863120

    1. eLife Assessment

      This valuable paper uses a mathematical model applied to a dataset of E coli / ESBL carriage and transmission to infer drivers of drug resistance in France. The strength of support for the study findings is incomplete. While the research question is of importance, and the mathematical model has structural and methodological integrity, numerous issues are noted: insufficient description of the data, lack of included equations and code, definitions of antibiotic use that are not complete, low sensitivity of assays for carriage, technical issues with statistical prior selection and parameter identification, and application of non-regional ECDC surveillance data to France.

    2. Reviewer #1 (Public review):

      Summary:

      The authors used a large dataset evaluating gut carriage of Enterobacterales and ESBL organisms from children aged 6-24 months as the basis for a modeling study to investigate what factors are most important for determining the prevalence of ESBL resistance. The modeling incorporated travel, a simple model of carriage duration (short and long), fitness cost of resistance on transmission and clearance, and antibiotic use. They found that antibiotic use is the primary driver of resistance prevalence, with transmissibility of resistant strains also important for setting the prevalence. Travel, while important when prevalence is very low, plays less of a role in maintaining prevalence once it is established (in keeping with other recent work). They estimated the fitness cost of resistance (terming a reduction of 14% on the rate of transmission and an increase of 23% on the rate of clearance as "low"). While the extent of assumptions and simplifications makes me skeptical of the quantitative conclusions, the qualitative ones seem reasonable and reinforce the long-held principles of the field--reducing antibiotic pressure and interrupting transmission--and highlight the importance of understanding the biological factors that shape the duration of carriage and the likelihood of colonization.

      Strengths:

      This study incorporates many of the factors that might influence the carriage prevalence of ESBL Enterobacterales. This builds on the work led by this group, both in primary data collection and in theory. Overall, it's such a tough problem that I commend the authors for trying to tackle it. The authors take a thoughtful, rigorous approach, acknowledging simplifications and assumptions where they need to, so as to evaluate the various factors shaping ESBL prevalence.

      Weaknesses:

      Part of the reason it's such a tough problem is that we have limited data to structure and parameterize a complex model.

      (1) The data are not sufficiently described.

      The primary data source for this modeling exercise comes from a study of 6-24-month-old children who underwent rectal swabs and evaluation of the carriage prevalence of Enterobacterales, and then whether these Enterobacterales were ESBL; moreover, the study included data on travel and on antibiotic use. Could the authors please direct us to these primary data? Could the authors also justify the parameters in their models from these data--for example, could they please provide the distribution of antibiotic use and the associated timing? Could they also explain why they decided to treat all Enterobacterales as if they were E. coli (line 307)? Is there evidence that all Enterobacterales occupy the same niche and compete with each other?

      (2) The model should be more fully described and the limitations explored/explained.

      - The authors should point to the code and the ODEs.<br /> - I understand the focus on the pediatric population; the authors argue that this is reasonable because ESBL colonization is similar across age groups. But presumably, antibiotic use differs across age groups, and there is colonization pressure from within households.<br /> - The authors only consider resistance to extended-spectrum beta-lactams and use of beta-lactam antibiotics, but ESBL Enterobacterales are often resistant to other antibiotics as well. How much does the use of other antibiotics also select for Enterbacterales that happen to carry ESBL resistance? "One bug/one drug" modeling, as done here, neglects the complexities of the actual patterns of resistance and range of antibiotic use.<br /> - Do the data support the T3 or S3 compartments, which, if I understand correctly, means no exposure to antibiotics can happen during three months after either treatment or travel? What do the data say about the patterns of antibiotic use? I'd imagine that the likelihood of antibiotic use is not homogenous, but instead, there are some who use repeated rounds of antibiotics.<br /> - Why do the authors exclude individuals who used antibiotics in the prior 7 days? What justifies that cutoff? The authors speculate that the impact of excluding these individuals is likely to be minimal; why exclude them, then? Did the authors evaluate the results if they were included?<br /> - What is the basis of "niche differentiation", as described starting on line 221? Why should clearance of one strain be slower when the strain co-occurs in a host with a strain of another type?

    3. Reviewer #2 (Public review):

      Overview:

      This study integrates several datasets into a unified modeling framework that incorporates several mechanisms thought to impact the spread of ESBL-resistant bacterial strains. The model accounts for tradeoffs between persistor and colonizer strains, travel rates, antibiotic treatment and strain clearance, direct competitive interactions, and, most importantly, a series of distinct costs associated with the carriage of ESBL resistance. The resulting 75-compartment model is internally consistent and structurally neutral. However, the parameter estimation is flawed in many ways, compromising the interpretations of the model.

      On the usage of the Swedish infant data set to estimate colonization and persistence:

      First, while other papers have taken similar approaches, the Swedish infant data set is fundamentally inadequate to estimate colonization and persistence rates. This is because very few colonies were typed per sampling event (2 to 6 colonies per event). The original authors themselves argued that strains of indistinguishable morphology would not be able to be differentiated by this method. They also provided data showing that strain identity was not directly related to colony morphology (same strain often displaying distinct morphologies).

      The consequence of this is that strains present in low abundance would be missed with a high likelihood. However, if they were to be stochastically sampled, this would count as a "colonization" event, and if they were missed in subsequent samplings, this would count as a "loss" event. In other words, the statistical methods described conflate within-host dynamics (which might lead to distinct within-host abundances) with between-host dynamics (colonization and loss).

      Beyond this conceptual issue, some technical aspects aren't particularly sound. The mean of the inferred posterior for the lambda and mu parameters are then used to calculate the beta, gamma, d, and epsilon parameters through a linear regression. The more technically correct way of doing this would be to directly infer these parameters from the data and obtain a full posterior for these parameters.

      This highlights another issue: these parameters are passed down to the next statistical model as point estimates, with no associated uncertainty. This artificially inflates the (already low) confidence of the estimates for the cost parameters.

      Finally, when this procedure generated parameters that were inconsistent with their expectations (clearance is too high to explain prevalence in France), they adjusted the parameters by discarding and recalculating their beta parameters to artificially enforce neutrality between their strains and enforce the expected prevalence. This is problematic because beta and gamma were jointly estimated, and there is no particular reason why some of them should be discarded. The more natural interpretation would be that parameters inferred from Swedish infants do not translate well to French adults, which should preclude their usage in this context.

      On the estimation of costs of ESBL resistance:

      The core of the second statistical model is to use prevalence data, travel data, and treatment data in conjunction with the previously inferred colonization and loss parameters to infer the costs of carrying antibiotic resistance. Therefore, the accuracy of this section is contingent on an accurate estimation of the previous parameters. However, these colonization and loss parameters are inherited with no uncertainty (just point estimates are passed down), which, as previously mentioned, generates an artificially precise posterior distribution for the resistance parameters.

      However, the most severe issue with the statistics lies in the choice of priors for the cost parameters. All of them are uniform in a positive range that implies a positive cost. Importantly, the average over a positive range will always be positive; therefore, this method will ALWAYS estimate a positive mean for the costs. Note that the posterior distribution of some cost parameters seems to peak around zero and abruptly decays with no mass to the left of zero. This is caused by the choice of prior. Had delta been allowed to be negative (i.e., antibiotic resistance carried a benefit, having the prior be uniform between -1 and 1), the posterior distribution would likely be much more symmetrical, and the confidence interval would have included 0.

      Restating, because the prior is a continuous function between 0 and 1, it contains infinitely more mass in the region that represents there being a cost (delta>0) than in the region representing no cost (delta=0). This means that it is a mathematical impossibility for this model to infer the absence of a cost.

      Therefore, the main finding of the paper ("We found that resistance is costly") is a mathematical artifact of the prior choice and of the model structure.

    4. Reviewer #3 (Public review):

      Cotto and colleagues integrated data analysis with mathematical modeling to examine extended-spectrum beta-lactamase (ESBL)-producing E. coli in France. While ESBL prevalence has risen globally, it has stabilized at approximately 6-8% across Europe. Established risk factors for ESBL carriage include prior antibiotic exposure and travel to high-prevalence regions, most notably South-East Asia. The dataset incorporated information on ESBL-producing E. coli and travel history in young children, and the model was calibrated to ECDC surveillance data on ESBL across Europe, supplemented by literature-derived parameters on antibiotic use, E. coli biology, and transmission dynamics. The authors report that ESBL-carrying strains exhibit a 14% fitness cost in community transmission relative to susceptible bacteria, yet are cleared 23% less frequently. ESBL carriage was strongly associated with factors that prolong gut colonization. Both antibiotic treatment rates and transmission efficiency were identified as key determinants of community-level ESBL prevalence.

      Strengths:

      The study addresses a clinically and epidemiologically important topic. The integrated modeling approach is methodologically sound and well-suited to disentangling the relative contributions of transmission and antibiotic selection pressure.

      Weaknesses:

      Several concerns regarding the data used in this study warrant consideration. First, model calibration relied on ECDC surveillance data pooled across multiple European countries, several of which have substantially lower antibiotic consumption than France (ECDC ESAC-Net Annual Epidemiological Report, 2024). Given that antibiotic use is a primary driver of ESBL selection, ESBL prevalence is likely to be heterogeneous across these settings. Calibrating to a geographically diverse dataset risks introducing systematic bias into parameter estimates that may not be representative of the French context. The authors should repeat the analysis using France-specific data, or, where this is not feasible, restrict the calibration dataset to countries with comparable antibiotic consumption profiles. Second, the travel exposure data may be insufficient to adequately capture importation dynamics from South-East Asia, as the cohort consisted exclusively of young children, a demographic less likely to travel to high-prevalence regions than older age groups. This may result in an underestimation of travel-associated importation as a contributor to community ESBL prevalence, and the generalizability of these findings to the broader population should be interpreted with caution.

    5. Author Response:

      We thank you for this assessment of our work and the positive assessment of the overall theoretical framework. We can fully answer the concerns, in particular regarding data quality, and will provide detailed answers in the following directions:

      “insufficient description of the data”: We will describe the data as much as possible and will share the data and the analysis code.

      “lack of included equations and code”:  We will share the mathematical equations in the supplementary material, and the full code on an online repository. The reason why our data repository (10.5281/zenodo.18480481) is not yet public is that it cannot be changed after publication. For the review process, we provide a github link to data and code here https://github.com/oliviercotto/eLife_epidR. We will ultimately share the link to the final version of the files on Zenodo.

      “definitions of antibiotic use that are not complete”: We will complete the definition of antibiotic use, which is the use of any antibiotic between 7 days and 3 months before sampling. Children who used any antibiotic 7 days before sampling were not included in the study. The type of antibiotic used is given in supplementary material S1: 93% of the antibiotics prescribed are beta-lactams (amoxicillin, amoxicillin/clavunalate, oral 3rd generation cephalosporins).

      “low sensitivity of assays for carriage”: the carriage study conducted in Sweden is used to get plausible estimates of carriage duration parameters in infants.

      • Strain definition is based mainly on randomly amplified polymorphic DNA (RAPD), not colony morphology. Strains with distinct morphology but the same RAPD profile are considered one strain. Conversely, it was checked that strains of the same timepoint with the same morphology most often had the same RAPD profile.

      • We did check that these data are not much affected by imperfect sampling: observations of a strain ‘disappearing’ from sampling then ‘reappearing’  at later timepoints are rare (14 out of 273 strains). This is why we did not correct these occurrences in the previous version of the analysis. In the revised version, we will add a description of these occurrences and correct them. This correction did not significantly alter the inferred parameters in our preliminary analyses.

      • At a broad level, the fact that E. coli clades vary in their carriage duration is very well established across multiple independent datasets; the precise value of carriage duration difference for “persistent” vs. “transient” that we inferred here (a two-fold difference, supplementary material S3) is actually relatively conservative, in the sense that other studies have detected more important differences. We will create a table summarising available evidence on colonization parameters of E. coli to show that the insights from the Swedish data are qualitatively robust.

      • Yet, we will conduct a range of sensitivity analyses to see how the inferred costs of ESBL resistance vary when varying differences in carriage durations, competition and niche differentiation.

      “technical issues with statistical prior selection and parameter identification”: We disagree there is a “technical issue” with prior selection: The fact that resistance is costly, hence that our priors are left-bounded at 0 for the cost parameters, is a prior expectation based on the observation that resistances do not go to fixation. If resistances only conferred an advantage in treatment, but zero cost, then they would quickly evolve to 100% frequency–contrary to what is observed in virtually all epidemiological studies of resistance. That said, we will relax the definition of these priors to test that the data is also compatible with a strong cost on some traits, and no cost or a “negative cost” on other traits.

      Regarding parameter identification and the specific comments on the inference of colonisation parameters: we re-inferred all colonisation parameters with direct inference assuming specific functional forms. This does not alter much the final colonisation parameters that we then use for our main inference. We will also conduct sensitivity analyses to examine how changing some of the colonisation parameters (carriage duration, competition and niche differentiation)  would alter the main inference.

      “application of non-regional ECDC surveillance data to France”: We will clarify our text, as there is a misunderstanding here: we do not use non-regional ECDC surveillance data for inference. We use ECDC data (i) for illustrative purposes, to show that trends in ESBL in France in this surveillance system are very similar to those observed in France in our focal dataset, thus showing the consistency and representativeness of our data. (ii) to give an overview of the weak and inconsistent association of ESBL with age across Europe, thus supporting the relevance of our approach even if our data concerns infants and children. We will make sure this is clarified in the updated version of the manuscript.

    1. eLife Assessment

      This study presents an important finding regarding the effect of Yoda molecules on PIEZO2 function, challenging the assumption that they selectively activate PIEZO1. The evidence supporting this claim is solid, but several methodological and conceptual issues need to be addressed. Overall, this work will be of broad interest to researchers working with PIEZO channels across various biological scales.

    2. Reviewer #1 (Public review):

      Summary:

      In this work, T. Wijerathne et al. investigated and reported the agonistic effect of Yoda1 and Yoda2 over PIEZO2 function using patch clamp electrophysiology, Ca2+ imaging, and molecular dynamics. They find that Yoda1 sensitizes PIEZO2 to membrane tension, can induce Ca2+ influx, and decreases its inactivation to a lesser degree than it does to PIEZO1 channels. Additionally, their data shows that Yoda2 sensitizes PIEZO2 channels to membrane indentation to a greater extent, but it has a weaker effect on channel inactivation than Yoda1. Interestingly, they report that a mutation in a conserved arginine between PIEZO channels can be used to abolish PIEZO1-mediated Ca2+ flux in response to Yoda molecules. As a whole, the results presented here should be put into perspective against previous and future works involving systems where both PIEZO1 and PIEZO2 might be expressed. This is especially true for works where Yoda1 has been used as a basis for determining the absence of PIEZO2.

      Strengths:

      The authors use multiple techniques to investigate how Yoda molecules affect the three most important biophysical aspects of PIEZO channels that, when changed, result in pathophysiological responses: a) sensitivity to mechanical stimuli, b) Ca2+ entry, and c) channel inactivation. Lastly, they find a specific amino acid/region that could be exploited for drug design and/or development.

      Weaknesses:

      The methods and discussion sections are lacking enough detail to fully evaluate the findings and put them into perspective, respectively.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript challenges the long-standing assumption that Yoda1 and Yoda2 are PIEZO1-selective activators. Using patch-clamp electrophysiology and calcium imaging in HEK293TΔPZ1 cells overexpressing PIEZO2, the authors demonstrate that Yoda1 potentiates PIEZO2 stretch-activated currents to a similar extent as PIEZO1 and slows PIEZO2 poking-current inactivation (albeit with lower efficacy). They further show that the more potent analog Yoda2 affects PIEZO2 at nanomolar concentrations and use mutagenesis and molecular dynamics simulations to propose that Yoda2's benzoic acid group forms a transient salt bridge with R1724 in the putative Yoda binding pocket, explaining its enhanced potency.

      Strengths:

      The authors are established Piezo/biophysics experts; the study is highly important, technically competent, and carries significant implications for the reinterpretation of prior work that used Yoda compounds as PIEZO1-selective probes.

      The core finding that Yoda1 modulates PIEZO2 stretch currents is convincing and important. However, several conceptual, methodological, and presentational issues need to be addressed before acceptance, as detailed below.

      Weaknesses:

      (1) The abstract states that Yoda1 potentiates PIEZO2 "as efficaciously as PIEZO1." This claim is accurate only for stretch currents and single-channel open probability, but the paper itself demonstrates important asymmetries: i) Yoda molecules slow PIEZO2 poking-current inactivation ~2-fold, versus ~5-10 fold for PIEZO1 (Figure 3b and ref #60). ii) Spontaneous Ca²⁺ entry via PIEZO2 requires non-physiological conditions (high extracellular Ca²⁺, hypertonic solutions) that are unlikely to occur in native cells.

      The abstract should be revised to clearly qualify where equivalence holds and where efficacy differences exist. IMO, the current wording risks overcorrecting the historical bias (PIEZO1-only) by going too far in the other direction.

      (2) Related concern: the PIEZO2 Ca²⁺ signal in Figure 2 is only detectable using a Ca²⁺-boosted solution (CBS ie 30 mM Ca²⁺). Physiological extracellular Ca²⁺ and cells normally do not experience sustained hypertonicity at these magnitudes. The authors should explicitly clarify that the practical implication of their findings is primarily for electrophysiological (patch-clamp) experiments and that the Ca²⁺ imaging caveat applies only under amplified conditions. Specifically, the authors should state that in standard Ca²⁺ imaging assays with physiological buffers, PIEZO2 is unlikely to confound Yoda1 results.

      Related point: Can cytochalasin D (CytoD) restore a Yoda1-dependent Ca²⁺ signal in physiological saline? This would help determine whether the weak PIEZO2 response is primarily a membrane tension issue (cytoskeletal tethering) versus intrinsically lower channel expression or permeability. The authors already have tagged PIEZO1/2 constructs and could, in principle, normalize by surface expression.

      (3) The mean inactivation tau values for wild-type PIEZO2 poking currents in both DMSO and Yoda1 conditions (Figure 3b, approximately 15-40 ms range) appear substantially higher than values reported in published literature (typically 5-10 ms; eg, PMID: 20813920). This discrepancy needs to be addressed.

      (4) The authors perform all MD simulations on a truncated PIEZO1 model and justify this choice by noting that the Yoda binding region is highly conserved between homologs. This is a reasonable and defensible starting point given the availability of well-validated PIEZO1 simulation set ups in their lab. A few points are nonetheless worth addressing: While PIEZO2 simulations are not strictly required, the authors are encouraged to briefly discuss whether any long-range structural differences between PIEZO1 and PIEZO2 (outside the binding site itself) could influence Yoda2 binding dynamics, particularly in light of the chimera data showing that PIEZO2 sequence in repeat A abolishes Yoda1 sensitivity. This reviewer still doesn't understand the reason behind this discrepancy despite it being acknowledged in the text.

      Another MD-related comment is that three simulation replicas (which is impressive for such a big system) show markedly different salt bridge occupancy (82.6%, 49.7%, 99.8%; stated in the text). This wide variation suggests incomplete sampling in at least one replica. The authors should provide RMSD plots for ligand and protein backbone to assess convergence and possibly discuss whether the 49.7% replica represents a genuinely distinct binding mode or incomplete equilibration.

      (5) The Discussion proposes that PIEZO2's weaker Ca²⁺ response to Yoda1 could partly reflect lower membrane expression. Since the authors already have fluorescently tagged PIEZO1 and 2 constructs, a simple fluorescence intensity comparison between the two (acknowledging it would reflect total rather than surface expression) could provide at least indirect support for this claim. Alternatively, if such a comparison is not feasible, the authors may consider removing membrane expression from the list of proposed explanations or explicitly acknowledging that this remains unsubstantiated speculation. The max poking currents may somewhat and roughly indicate the level expression difference too, if done exactly side by side.

      (6) The abstract or concluding remarks should highlight that Dooku1 is not PIEZO1-selective in its agonist-like action on PIEZO2, and that Cmpd15/Cmpd64 appear to be better PIEZO1-selective tools. This nuance is buried in the Results section.

      (7) The authors should not cite PMID 31015490. Clearly, any work on MCC13 is confounded by the overwhelming expression of PIEZO1 (PMID: 42084270). Instead, the authors should also cite the literature from others who have clearly recorded stretch currents from PIEZO2 before the cited studies (eg, PMID: 37590348).

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript reports that Yoda1 and Yoda2 agonize PIEZO2 in a manner similar to PIEZO1, increasing open probability and stretch sensitivity, but the mechanism underlying this sensitivity is incomplete. Mutagenesis was shown exclusively in PIEZO1, with no corresponding mutagenesis in PIEZO2, so the proposed mechanism in PIEZO2 is inferred by homology rather than directly tested. All experiments use mouse PIEZO2, and the human ortholog should be used before generalizing the proposed reinterpretation of the field.

      Strengths:

      The pressure-clamp electrophysiology demonstrating a shift in half-activation pressure for PIEZO2 is compelling evidence in support of the central claim.

      Weaknesses:

      (1) In the single-channel recordings (Figure 1a), it's unclear how many channels were present in those patches. After applying -60 mmHg pressure, multiple channels would be activated (as seen in Figure 1e). The number of channels in the patch and their inactivation rate could significantly influence the open probability in such experiments. To overcome this, in the original Yoda1 article (Syeda, Ruhma, et al. eLife 2015), no additional pressure was used. Additionally, the reported open probability comparison (n=7 Yoda1 vs n=17 DMSO patches) has an SEM nearly as large as the effect itself (0.30 {plus minus} 0.11), consistent with a small number of outliers driving this. The underlying mean open and shut times are reported without any statistical test; only the derived open probability receives a p-value. Additionally, in Figure 1a, the Yoda1 condition noise is different from the control. This should be stated if noise filtering was applied and how, given that this could affect open probability analysis.

      (2) The calcium imaging data in Figure 2 raise significant concerns regarding the chemical activation claim. The calcium-boosted solution (30 mM Ca2+) is not physiological and appears to be generally stressing cells rather than specifically activating PIEZO2: the control condition under CBS already shows an elevated signal, consistent with cells being unwell at this calcium concentration, and adding Yoda1 on top of this shifted baseline raises further questions about specificity rather than confirming it. Separately, it is unclear why DMSO alone produces measurable PIEZO2-associated calcium influx in HBSS, a result that is not addressed in the text. Figure 2 should clearly indicate when DMSO/Yoda1 perfusion was initiated, and y-axis labels are missing from panels A and B.

      (3) In the poke experiments, an activation threshold should be calculated and reported, and amplitude data (e.g., peak current versus indentation depth) should be shown rather than only inactivation tau values. It is also unclear why mClover3- and N-GFP-tagged constructs were used in these experiments, since electrophysiological recording already confirms channel expression without requiring a fluorescent tag.

      (4) For inactivation kinetics (Figure 3b), the authors use unpaired comparisons across separate cells, whereas the deactivation experiments (Figure 3c) use paired; it should be applied to the inactivation experiments as well. Deactivation kinetics for PIEZO2 itself should be shown. If the claim is that Yoda1 acts on PIEZO2 through the same mechanism proposed for PIEZO1, then a PIEZO1/2 chimera should be expected to show a corresponding effect on deactivation tau; instead, this chimera is reported as completely Yoda1-insensitive despite both parental channels being Yoda1-sensitive, as shown in this study.

      (5) Given that this reflects a different experimental paradigm for Yoda EC50, PIEZO1 should be included within Figure 4b. Additionally, EC50 bar plots should be present on this figure. The inactivation time constant for PIEZO2 without Yoda1 is inconsistent across figures, below 20 ms in Figure 3b but above 20 ms in Figure 4c.

      (6) Finally, the modeling is performed exclusively on PIEZO1, whereas the manuscript's central focus is PIEZO2. It is therefore unclear whether the proposed structural mechanism, including the basis for Yoda2's reduced efficacy on PIEZO2, can be directly extrapolated to PIEZO2.

    1. eLife Assessment

      In this manuscript, the authors describe a new member of the KCNE auxiliary subunits of potassium channels from a lamprey. This new subunit represents an early evolutionary member which confers new properties when expressed along with KCNQ channels. The authors present convincing evidence from several experimental approaches. The contents of this manuscript are important and should be relevant to understanding both the mechanism of modulation of KCNQ channels by KCNE subunits and the evolutionary history of these subunits, which this manuscript now extends to the divergence of early vertebrates.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

      Weaknesses:

      (1) No error bars visible for lamprey Q1 isoforms (open symbols) in Figure 2G. No statistical comparison was provided to indicate whether lamprey Q1 isoform V1/2s are significantly different (nor in Supplementary Table 1).

      (2) There is the same issue in Figures 3 and 4. No appropriate statistical comparison is made between V1/2s for different truncations of PmKCNE0 (Figure 3), or between KCNQ1 species isoforms with and without PmE0.

    3. Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

    1. eLife Assessment

      In this important study, Boudjema et al. use cell culture models and high quality advanced microscopic imaging to provide detailed analyses of the cellular processes underlying centriole amplification, apical migration, and assembly of hundreds of motile cilia in multi-ciliated cells. The authors present convincing evidence showing that in these cells all the molecular and cellular steps controlling centriole biogenesis that in cycling cells extend over almost two cell cycles, occur within a single cell cycle variant. This work provides a better understanding of the regulation and order of these processes and is of interest to all cell biologists and in particular researchers studying centrioles and cilia.

    2. Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      Comments on revised version.

      The authors have significantly improved the manuscript, by refocusing it, introducing text and figure changes, and by adding new data including functional analyses. The revised version now has convincing data that support the claims. All my remaining concerns have been addressed.

    3. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the Cen2-GFP; mRuby-Deup1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described Cen2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of Deup1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells(treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules and microtubule-based transport in different stages of differentiation in brain MCCs. Addressing the role of microtubules during different stages of centriole amplification required development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative spatiotemporal analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive and are further strengthened in the revised version through additional analysis and adoption of new methods. This comprehensive analysis advances our understanding of MCC biology regarding the involvement of microtubules.

      Comments on revised version.

      The revised manuscript is substantially improved, and given the scope, it is appropriate that it primarily establishes a detailed spatiotemporal framework. That said, a few points would further strengthen clarity and impact. First, several observations naturally raise follow-up mechanistic questions, for example whether additional cytoskeletal systems such as actin contribute to steps like centriole apical migration. A slightly more detailed framing of these open questions would help guide future work. Second, some terminology introduced to label observed microtubule-based structures (for example "nest") may not be essential. Finally, while the authors have increased quantification, some analyses would benefit from super plot-style displays with replicate-level comparisons, particularly for intensity-based readouts.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully addressed the insightful comments provided by the reviewers which thoroughly increased our comprehension of the dynamics of centriole amplification. The manuscript has been revised accordingly and put in the context of the two papers we published since our last submission, showing that MCC differentiation is a genuine cell cycle variant. A point by point answer to all reviewer comments is provided below.

      Briefly:

      We have streamlined terminology and nomenclature in text and figures / better define experimental conditions with nocodazole

      We have tested the role of dyneins in the dynamics of centriole amplification

      We have done correlative light and electron microscopy on the early stages of centriole amplification

      We have analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors

      Collectively, this allowed us to make a clearer parallel with what occurs during centriole duplication and to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles.

      We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Public Reviews:

      Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      We have streamlined terminology and nomenclature, clarified the description of the complex events, and test the role of dyneins in centriole amplification. Microtubules density in MCC does not allow to extract information from imaging. In addition, we have done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Altogether, our new data allowed to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles. We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Reviewer #2 (Public Review):

      This important work will be of interest to centriole and cilia cell biologists. It describes in detail how microtubules control multiple aspects of centriole amplification in brain multiciliated cells. This study provides a greater time-resolved and molecular proteomic mapping of the different steps involved, with or without microtubule disruption. Boudjema et al. show that microtubules are important throughout the centriole amplification process, from the early stages, where the procentrioles emerge from a pericentriolar "nest", through the growth stage where microtubules maintain the perinuclear localisation, to the detachment stage, where microtubules assist in perinuclear disengagement and apical migration. The results are generally well supported by the evidence, but the manuscript would benefit significantly from some heavy editing to introduce more niche terms, standardize abbreviations in text, and labels on figures to help bring the readers, especially non-specialists, along with them - increasing the accessibility of their work.

      We thank the reviewer for his/her enthusiasm. We have streamlined terminology and nomenclature and clarified the description of the complex events to increase the accessibility of our work. We also replied points by points to his/her specific comments.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical-basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, and MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the CEN2-GFP; mRuby-DEUP1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described CEN2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of DEUP1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM, and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells (treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules in centriole amplification. Addressing the role of microtubules during different stages of centriole amplification required the development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive. This comprehensive analysis advances our understanding of MCC biology.

      Weaknesses:

      The role of microtubules and other molecular players during different stages of centriole amplification in brain MCCs can be further studied and strengthened using the tools developed in the manuscript. A more quantitative description of some of the analysis performed in the manuscript is required to strengthen the conclusions.

      We thank the reviewer for his/her enthusiasm. We have tested the role of dyneins in the dynamics of centriole amplification, done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Recommendations for the authors:

      As you will see, all reviewers felt that the analyses of the involvement of microtubules should be strengthened by including controls and additional experiments. Also, they agree that significant text editing would help to improve the manuscript's accessibility and readability.

      Specifically, they would suggest (1) streamline terminology and nomenclature in text and figures; (2) better define experimental conditions with nocodazole (concentrations used, effect on microtubules, effect on canonical centriole duplication); and (3), in the absence of other complementary genetic perturbation experiments, add a limitations paragraph in the discussion about conclusions drawn from nocodazole treatment alone.

      Reviewer #1 (Recommendations For The Authors):

      Main issues:

      (1) The authors use variable terminology to describe the same or similar events/structures. For example, in Figure 1 they refer to "centrosome stage" where they observe a pericentrin "cloud", which they later refer to as a "nest". In all other figures the first stage is not referred to as the "centrosome stage" but as the "cloud stage". Again, they also describe the "cloud" as a "nest" occasionally, but not always. In the cartoon, the nest is termed "centrosome cradle". The variable and inconsistent use of terms is confusing and the authors do not provide any explanation for the use of one vs. another.

      The text is now corrected. The centrosome stage corresponds to the stage preceding the beginning of centriole amplification in MCC progenitor. The pericentrosomal cloud of centriole and deuterosome elements forms later on, during the amplification A-stage. The formation of this cloud marks the beginning of A-stage, and persists up to G-stage where it dissolves. When we show that the cloud hosts the first stages of centriole biogenesis, we defined it as a “nest”. We do not use anymore the term craddle.

      (2) What prompted the authors to use the term "nest"? It gives the impression that they describe aspecific physical entity/structure (also depicted in this way in Figure 3P, with microtubules outside of this structure), but what is the evidence for this?

      The cloud is the spatial entity and the term “nest” is used to define a function of this transient compartment. We decided to keep the term “nest” as we now identified it with correlative light and electron microscopy, in addition to U-ExM, and show that the accumulation of centriole and deuterosome elements is accompanied by the formation of immature procentrioles, deprived of MT walls, as well as immature and empty deuterosomes. The scheme with MT outside the cloud/nest is misleading as we see MT organized by the mother centriole. We have now changed this.

      (3) The "nest" may simply be a dynamic accumulation of precursor particles around the centrosome, similar to what has been described for centriolar satellites. Rather than proposing a new entity, I suggest testing whether the "nest" particles may colocalize with PCM1 and thus may be related to centriolar satellites. Based on the data, the nest would simply be the centrosomal MTOC that organizes a radial microtubule array on which particles move around its center. In the absence of other evidence, I am not convinced that a new term is needed.

      We totally agree with the reviewer: the centrosome, as MTOC, concentrates centriolar and deuterosome components. This cloud is consistently dissolved when MT are depolymerized or dyneins inhibited. So, the physical entity is a “cloud”. We used the term “nest” to propose one function for this cloud which is to form deuterosomes and centrioles, before they move away for maturation. In fact, deuterosome and centriole formation are hindered when the cloud is dissolved. We have tried to edit the text all over the manuscript to make it clearer.

      (4) Role of MTs: are microtubules required or do they just facilitate some of the investigated events?

      The reason why the role of MT has not been tested yet during centriole amplification is probably because MT not only constitute the cell cytoskeleton on which molecular motors ride to transport cargos or distribute forces, they are also the core component of the structures we are studying. This is why we have tested a range of nocodazole concentrations and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (Fig. 4 Supplementary 1A-B). This may lead to an underestimation of the role of MT but we cannot study the role of MT on centriole amplification if centrioles cannot be formed.

      Does multi-ciliation in these models eventually occur normally under the concentrations and treatment conditions used here? This should be tested and discussed in the context of whether microtubules are indeed required and at what step of the entire process (amplification, migration, ciliogenesis) they may be critical.

      We did both chronic and acute treatments.

      Chronic treatments were done to test the overall efficiency of centriole amplification when MT (or dyneins) are perturbed. Chronic treatments were used to assess the role of MT (or dyneins) on the global efficiency of centriole and deuterosome formation (number of cells able to amplify, number/size/loading of deuterosomes, final number of centrioles (Fig. 4H-I, Fig. 4 Supplementary 2 B-D). In these chronic treatment, we focused on centriole amplification and not ciliation since it was the scope of this study. Also, we did not take ciliation as a readout of amplification because ciliation is relying on MT polymerization.

      Then, we also did acute treatments to test the role of MT (or dyneins) at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration; Fig. 4, 5, 7, 8 and associated supplementary figures). Since one stage is dependent on the precedent one, this enabled us to decipher the direct role of MT (or dyneins) on each single stage. We have now edited text, methods, legends and pictograms to be clear on whether acute or chronic treatment was done.

      (5) Can the authors include control (non-amplifying) progenitors in their analyses? It would be useful to know what the signal and distribution of each specific marker are before differentiation begins (before the cloud stage).

      Non amplifying progenitors are analyzed and constitute the so-called “centrosome stage”. We have now precised it and called it the “progenitor stage”.

      (6) Figure 2: Again, the terminology is confusing, since the authors describe that DEUP1 forms a "cloud" with centrin during the A stage.

      Corrections have been done as explained in point 1.

      (7) Description Figure 3: the authors introduce yet another term: "halo" A-stage. Is this the early A stage? Again, this is not explained and confusing. More systematic and consistent description is needed.

      Corrections have been done as explained in point 1. The term halos is used un the lab as it was the first term we used in our Nature paper in 2014 in reference to the halo described by Erich Nigg when they overexpressed Plk4. It was an error to use it in the manuscript.

      (8) Nocodazole treatments: the used concentrations are quite high.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely (Fig. 4 Supplementary 1A-B).

      (a) To avoid non-specific effects the authors should test what the minimal concentration is that completely depolymerizes microtubules in their cell model and perform analyses at this concentration.

      We have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A-B), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). In case it was not clear, we refer to this now several time and more clearly in the text and methods.

      (b) They should demonstrate depolymerization of microtubules by microtubule staining in the acute and chronic noc treatments and at the different noc concentrations used.

      This is, and was, in supplementary material (same, Fig. 4 supplementary 1A).

      (c) The authors should demonstrate that the used nocodazole concentrations do not impair normal centriole biogenesis during the cell cycle in these cells; if so, impaired assembly of centriole wall MTs may contribute to the observed effects in Figure 4.

      As mentioned in point 8b, we have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). The ability of the cells to form centrioles during chronic treatments were always assessed using immunostainings of SAS6 and/or CEN2-GFP signals (now exemplified in Fig. 4 Supplementary 1B). We also did EM analysis on cells treated with the highest doses of nocodazole (Nocodazole 10 uM for 24h) and this showed that centrioles can form with, what seems to be MT walls, in cells totally deprived of cytoplasmic MT fibers (Fig. 4 Supplementary 3-4). However, this does not show that all the cells can, because the number of cells that can be analyzed by EM are not sufficient to conclude. Also, one cannot assess whether MT walls are properly polymerized. However, the absence of MT walls should not change the results of the Figure 4, which are based on DEUP1, SAS6 or CEN2-GFP signals for deuterosomes and centrioles. Also MT depolymerization affects the formation of deuterosomes, which should not be altered by MT wall defects as it is not affected, even when centriole formation is blocked (LoMastro et al., 2024). Last but not least, we now show that blocking dyneins, as a comparable and even greater effect, on the formation of the cloud, deuterosomes and centrioles (Fig. 4C-I and Supplementary Fig. 4), which confirms that MTOC function, rather that MT wall formation, explain the centriole biogenesis alteration shown in Figure 4.

      (9) The authors repeatedly refer to the centriole-to-centrosome conversion of amplified centrioles and how this resembles centriole-to-centrosome conversion during the cell cycle. However, they incorrectly claim that this occurs at the G2/M transition. PLK1-dependent modification occurs at this stage, but conversion and PCM recruitment only occur after mitosis (see original work by the Tsou lab, which needs to be cited here).

      We agree with the reviewer. We have now added additional data to show clearly that centriole biogenesis, which requires two cell cycles to proceed in cycling cells, is accelerated during the MCC cell cycle variant where the elongation and maturation cycles are superimposed. This is now clearly shown in Fig. 3, 5, 9 and discussed.

      (10) Figure 6H-J: the authors claim that at low noc concentration, more D-stage cells showed incomplete disengagement than in controls, but the effect is shown only for the highest 10 µM concentration. Do any eof the phenotypes in Figure 6 also occur at the lowest noc concentration (assuming it depolymerizes MTs)? Again, it is crucial to demonstrate this, to exclude unspecific effects not linked to MT depolymerization.

      An error was made on the figure (but not in the legend). In Figure 6, chronic treatments are at 1 or 5 µM. Only acute treatments were done using 10 µM. In both cases, MT are not entirely depolymerized in these experiments (Fig. 4 supplementary 1A).

      (11) Disengagement, Figure 7: The authors describe that DEUP1 signal spreads all over the cytoplasm and becomes diffuse during this process, but one cannot see a diffusive signal throughout cells in the figures.

      We pushed the contrast to make it clearer but the deuterosomes are still bright at this stage and it is difficult to have both signal clear (now in Fig. 6B). We have also changed the example in video (now video 19) to show it more clearly with DEUP1 channel alone.

      (12) Figure 7: localization of disengaged centrioles at microtubule "nodes" is not clear from the images. There are many centrioles and random colocalization may be expected simply based on the high number. Higher resolution and/or magnification and quantification would be needed.

      We have edited and now say that centrioles “colocalize” with MT which, since centrioles nucleate MT, seems normal. We agree that it could be random, but given the density of MT, and the number of centrioles, it does not seem opportune to us to quantify. We can just say that we never see centrioles is regions that are deprived of MT.

      (13) The term "diffusive" to describe slow centriole movements in Figure 8 suggests that it is not motor or force-dependent, but there is no evidence for that. Movement based on opposing forces could produce a similar result, but would not be considered diffusive.

      We agree. We have changed “diffusive” by “diffusive-like”.

      (14) The manuscript would greatly benefit from the analysis of some candidate motor activities that may drive the movement and migrations of centrioles in this system. This would support the importance of the microtubule network for the specific steps in these processes, and better define its role beyond "being required". Dynein may be a candidate or minus end-directed kinesins. Since chemical inhibitors are available, these types of experiments would be straightforward.

      We formerly tested ciliobrevin but had hard time because of the small stability of the drug. Since our submission to eLife, we tested dynapyrazol and dynarestin and found dynapyrazol very efficient in dissolving the Golgi, a good readout of dynein inhibition. We sought to test the role of dyneins, using dynapyrazol, on (i) the formation of the pericentrosomal cloud in A-stage, (ii) the oscillation of DEUP1+ structures during A-stage, (iii) the number, size, loading of deuterosome, (iv) the final number of centrioles, (v) the migration to the nuclear membrane and (vi) the final apical migration of centrioles. The results are now inserted in main and associated Fig. 4, 5, 7, 8, 9.

      (15) Discussion:

      "the role microtubules" lacks "of"

      This is now edited.

      "This lack is..." Lack of what?

      This is now edited.

      "reflexive link" - meaning of "reflexive" is not clear in this context

      We have removed it.

      In my opinion, the study does not identify a nest composed of DEUP1, PCNT, and Centrin2; it only shows that these components accumulate as particles around the centrosome, which functions as MTOC. Consequently, it seems that the "nest" does not exist when MT is depolymerized. One could consider the center of the centrosomal MT array as a nest in this context, but there is no evidence of a specific new structure as suggested by the way the term is used in the manuscript.

      This is what we want to say: the center of the MT array become a nest in this context. We do not state that there is a specific new structure. We just say that MT and dynein dependent concentration of centriole and deuterosome components exists and that this region nests the birth of centrioles and deuterosomes. Also, this compartment is restricted in time and space, which justifies to use a specific term. The MTOC exists in the progenitor cell, while this compartment, marked by DEUP1, Centrin, PCNT accumulation, appears at the beginning of amplification and grows during A-stage to be dissolved at G-stage when all the deuterosomes and centrioles have moved away.

      What is the evidence that "DEUP1 is a centrosomal protein before building deuterosome structures"? It would be good to refer to the specific experiment. Does DEUP1 localize at centrioles also in the absence of microtubules? If not, I would not consider it a centrosomal protein.

      We have removed this statement to avoid misinterpretation.

      "This reminds the centriole-to-centrosome conversion..." the sentence is missing an "of"; also, again the authors confuse the order of events during the cell cycle, where centrosome conversion occurs after completion of mitosis, not at G2/M transition.

      We have removed this statement to avoid misinterpretation. Also, see Point 9.

      "microtubule dependent nuclear migration" should be rephrased; it sounds as if the nucleus migrates.

      This has been changed

      The following discussion of disengagement being linked to association with the nuclear envelope and resembling the process in cycling cells is misleading. In cycling cells movement of centrioles along the nuclear envelope occurs at G2/M and drives centrosome separation (separation of centriole pairs) in preparation for mitosis, not centriole disengagement.

      We are now clearer. We compare centriole-loaded deuterosome organization around the nuclear membrane to the migration of new centrosomes during early prophase (Fig. 5F-H, Fig. 5 Supplementary 2G-K).

      Regarding the possibility that forces by microtubules generated by the daughter centriole drive disengagement also in cycling cells, I would argue that this is unlikely since the daughter centriole can only nucleate microtubules after disengagement has occurred (and conversion to centrosome/PCM recruitment). Once this happens, it may physically separate the disengaged centrioles, which is a different type of activity. Indeed, originally the term "disengagement" was coined to specifically describe the loss of the perpendicular engagement of daughter centrioles with their mothers (Tsou and Stearns, Nature, 2006).

      We have removed this statement to avoid misinterpretation. The perpendicular engagement is difficult to assess on deuterosomes but we do see by live imaging, that attachment changes during D-stage, before centrioles detach clearly from deuterosomes.

      "high resolutive" should be "high resolution"

      Edit done.

      "splitted" should be "split"

      Edit done.

      "Consistently, when the mitotic oscillator is dis-inhibited and cells enter pseudo-mitotic events, centrioles show clear and rapid cell-cycle like clustering" This sentence is not understandable without further explanation; what does mitotic oscillator refer to? What are pseudo-mitotic events? What is cell cycle-like clustering?

      We have removed this statement.

      Minor:

      (1) Abstract: "Centriole number must be restricted to two..." Since cells are born with two centrioles and have 4 centrioles (2 pairs) when they enter mitosis, this sentence is inaccurate.

      The sentence has changed.

      (2) Abstract: "reflexive link"; I am not sure what the term "reflexive" refers to?

      We have removed this statement to avoid misinterpretation.

      (3) Figure 1C, D: it should be described better that the larger magnification panels represent overlays of many cells and what marker they show. This is not obvious since the smaller single-cell panels always show two different markers. Also, it would be more useful to show also single cells in the magnified view. The overlay does not allow us to see if a marker forms a cloud or a single dot, which is as important as the cell-to-cell variation in distribution.

      We have clarified this in the text and the legend. The cell-to-cell variation cannot be estimated with the overlay, but the projection from several cells (number precised) allows to see that the signal is confined in a restricted region. Or not. Which is what we wanted to analyze.

      Related to the above, the authors say that pericentrin forms a cloud at the top left in panel D, but there is only one confined centrosomal dot in the single-cell panel.

      The sentence has changed.

      (4) Results, Figure 2F; video 4: The authors claim connection and disconnection of DEUP1 aggregates with centrosomal centrioles; can the authors comment on the spatial resolution including in z in this movie to support this claim? Can they exclude that the structures are in proximity of each other rather than "connected"?

      This is a single z-section of 500nm. The resolution in xy is 128nm/pixel. Given the sizes of deuterosomes and a mature centriole, and given the fact that we observed this dynamics in several cells in live, we can state that the structures are connected. This is consistent with deuterosomes frequently observed “kissing” the daughter centriole by EM in the present manuscript (Fig. 2D, Fig. 2 supplementary 3 and 4 and Fig. 4 Supplementary 3-4). One has to look carefully at the daughter centriole (marked “dc”) and span in on the serial sections to see the connected deuterosome (marked by a star): this is at very early stage and therefore it is small. We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both structures are frequently sticked to each others on tens of nanometers.

      (5) The term "dynamics" as used in the manuscript should be plural.

      It has been used plural, except when for “dynamic microtubules” and “dynamic attachment to the nucleus”, which we think is ok? We have not found any other singular uses in our manuscript.

      (6) Figure 5: what does "YL1/2 procentriole intensity" refer to in panel F? This should be the intensity of microtubule asters.

      This has been modified.

      (7) Figure 6 - supplement 1B: contrary to the claim in the text, one cannot see tight colocalization with the nuclear pore marker. This seems to be a very small subset of particles and even in those cases colocalization is not tight. Also, what is the relevance of nuclear pore colocalization?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. What we want to say is that there is a tight connection with the nuclear envelope as shown by the localization of NPC on the same z-section as centrioles. This is why we present a single z, to show that centrioles and NPC are on the same z-plane of 500nm. NPC are stained to outline the nuclear membrane. This is also clearly visible for G-stage centrioles in the XY plane. We have now added an entire z-stack on video 18.

      Reviewer #2 (Recommendations For The Authors):

      To improve accessibility of their manuscript, we would suggest making the following edits:

      (1) Define 'specialist' or 'niche' terms each time you introduce them, such as 'pericentrosomal nest', or 'flower-like structures'.

      This has been clarified.

      (2) Have a think about abbreviations, again ones that work for people outside the project- this paper uses 'PC' for 'procentriole' but for many 'PC' is 'Parental centriole' or Figure 6J talks about 'D total' or 'D partial', leaves readers confused.

      This has been clarified.

      (3) Standardize your abbreviations throughout particularly for your treatments- sometimes Noco sometimes, NOCO, or your imaging experiments sometimes Cen-GFP, sometime CEN2-GFP (Figure 7A, D vs. Figure 6) or DEUP1- mRuby, DEUP1-mRuby3 or mRuby3-DEUP1?

      We now use Nocodazole or Noco in the text and the figure respectively, CEN2-GFP and mRubyDEUP1.

      (4) About 10% of the population, including several key figures in this field, are red-green color blind. Although 4 colour fluorescence is difficult to get right for everyone, choosing palettes (especially for two colour panels) is inclusive. More so, greyscale or inverted monochrome images make it easier for everyone to visualize changes in localization, size, and intensity. Red on black small foci is particularly difficult to discern. For example, Figure 3 - more individual channels in grayscale with arrows to mc, dc, and cilia would be helpful - difficult to distinguish stainings.

      We thank the reviewer for this comment and for this recommendation of being more inclusive. We have done the changes.

      To improve the conclusions drawn, we suggest some revisions below:

      (1) Since the paper really hangs on it, a clearer description of the rationale for when, how long and how much nocodazole treatment was done is needed. The logic currently is difficult to follow seemingly random jumps 10x concentration are used. Microtubules control many aspects of cell biology and could be impacted. For example, I particularly found Figures 6D and H difficult to follow i.e. the timing for 6H seems off.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely. We have of course tested a range of nocodazole concentrations at the beginning of the study and shown the extent of MT depolymerization under each treatment. We used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4 reviewer 1). The level of perturbation of MT and consequences on centriole formation at the different timings and doses were done for each experiment and are exemplified in Fig. 4 supplementary 1A-B. This figure was already present in the first version of the manuscript but we have now edited text, methods and pictograms to clarify this.

      (2) Perhaps an extension of this point- in general how interdependent are the processes? If there is a defect at the nest stage, how much are the later defects secondary to this, or do MTs genuinely play direct roles at all stages or are these knock-on effects? How do the authors rule this out? Defects in the nest, lead to smaller and more DEUP1+ foci, with defects in concentrating procentriole factors and centrin, which lead to... For example, Figure 4B looks like centrin is reduced upon noco treatment? Does noco treatment affect Cetn2GFP levels globally? Individual channels grayscale would help visualise this better.

      See also our answer to reviewer 1 point 8c.

      The stages are indeed interdependent. This is why we did both chronic and acute treatments. Chronic treatments were done to test the overall efficiency of centriole amplification when MT are perturbed. We typically used low dose of 1µM because nocodazole remains 48h in the culture medium. Acute treatments were done to test the role of MT at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration). Most of the acute treatments were done live and nocodazole was applied after the first time point of live monitoring. We used 10µM to have a rapid effect, and because nocodazole remains only several hours in the culture medium. This allowed to monitor the stage “n”, in cells where the stage “n-1” was completed without any drug which allowed to analyze a stage without having perturbed the precedent one.

      We now also test the consequences of dynein inhibition using both acute and chronic dynapyrazole treatments. We show that except for centriole migration, dynein inhibition phenocopies MT depolymerization (centriole number, perinuclear organization and disengagement as well as deuterosome number/loading/size).

      Nocodazole chronic treatments do affect intensity of CEN2-GFP at G-stage centrioles suggesting an altered A-to-G transition. In D-stage, CEN2-GFP signal seems normal. We now mention this in the text and in the Fig. 4 Supplementary 1B.

      (3) The authors nicely show the importance of MTs in the structure of the nest from which procentrioles and DEUP1 positive structures emerge. They suggest this nest may be what supports procentriole generation in the absence of DEUP1 and parental centrioles. Firstly how does this nest look in the absence of DEUP1 and/or parental centrioles (centrinone treatment)? This may be what they are trying to show in Figure 5 Supplement 1 but it currently is very difficult to digest what it is showing relative to controls and whether this is significant in the way it is plotted.

      The nest is conserved in the DEUP1KO with or without centrosomal centrioles, as shown by accumulation of Centrin and PCNT at the center of the self-organised MT network (Mercey et al., 2019). This is in fact what motivated our study on the role of MT in centriole amplification. We have edited the legend to precise the quantification done, which is not related to this question. In this quantification, we show that the increased propensity to accumulate PCNT by centriole-loaded deuterosomes between A and G-stage is maintained in the absence of deuterosomes, indicating that centrioles themselves accumulate/recruit PCNT.

      (4) Can you do CLEM on DEUP1-Ruby and these early foci at the cloud stage to see if they are visible at the ultrastructural level, relative to procentrioles, microtubules, and other electron-dense structures?

      We thank the reviewer for this question. We have done CLEM on the pericentrosomal cloud during very early steps of centriole amplification. This showed that DEUP1 early accumulation at the centrosome corresponds to a region rich in fibro granular aggregates, suggesting that DEUP1 may be translated here, through locally concentrated centriolar sattelites, known to be involved in local translation. Then, small deuterosomes and immature centrioles are formed, within this cloud of sattelites, confirming that the pericentrosomal cloud is a nest for centriole biogenesis (Fig. 2C-D + Fig. 2 Supplementary 2-6 for control and Fig. 4 Supplementary 3-4 for nocodazole treated cells). This also shows that immature deuterosomes are not necessarily round shaped, and can be deprived of centriole loading.

      (5) Check the scale bars- see Fig 4E. Check throughout.

      Done.

      (6) Figure 3 Supplement 1 and 2 don't match the legend and are likely reversed - which one is right?

      Done.

      (7) Technical issue - I couldn't play videos 6 or 16? Check these work.

      Done.

      (8) Nomenclature mammalian proteins- mouse or human- should be all caps DEUP1, PLK4, SAS6,etc. Watch your units- space between number and unit.

      This has been done.

      (9) Many of the graphs involve three biological replicates but why not plot the mean of each of the three experiments and do stats? The number of events measured may conflate the significance. Try using Superplots.

      Here is how we proceed: we count the number of occurrence of the phenotype we monitor, and the total number of cells. We apply a X<sup>2</sup> to test whether there is a significative difference between our replicates in each condition. If not, we pool the number of occurrence of the phenotype we monitor and the total number of cells for the 3 replicates, and for each condition. Finally we apply a X<sup>2</sup> between the different conditions. This is how we usually proceed to avoid comparing a mean of percentages. This is now explained in the methods.

      Minor points:

      (1) "DEUP1 is a centrosomal protein and assembles deuterosomes in the pericentrosomal region in brain MCC". I am not sure you have evidence that DEUP1 is a centrosomal protein. You don't seem to study the relationship between centrosomes and DEUP1? Rewrite this title and tone down this claim.

      This has been modified.

      (2) Why the crossbow micropattern (versus some other shape) - seems very specific but not discussed?

      We wanted a shape where centrosome is not localized at the center of mass of the nucleus. Among the corresponding patterns, the crossbow was the one where differentiating cells had less propensity to detach.

      (3) Figure 2 - are the foci of DEUP1 at the cloud stage smaller than at A stage? How do they grow? Measure the diameter at cloud stage, just after they leave the cloud and then once they move away from centrosomal cloud and each other. If so, and they do indeed grow in size from the cloud stage to the growth stage which I think your images suggest - do you envision this happening with the gradual addition of DEUP1 rather than fusion?

      Early deuterosomes are not easy to detect by light microscopy, because of accumulation of DEUP1 in the cloud. We did CLEM on the cloud of early A-stage cells to resolve the earliest deuterosomes which are often very small (see Fig. 2D, Fig. 2 Supplementary 2-6) suggesting that they grow, either by fusion, which we never observe in our movies at later A-stage, or by accretion of DEUP1. However, by light microscopy, we can detect very early but big deuterosomes, which we see splitting later on into smaller ones. So, we cannot conclude on the mechanism that regulate deuterosome size. This is now discussed in the discussion of the manuscript.

      You say in the discussion:

      "Consistently, we never observed fusion events of DEUP1 condensates in our time-lapse experiments. More importantly, we did FRAP experiments on endogenously tagged mRuby-DEUP1 in cells at the different stages of centriole amplification, and did not find significant recovery, supporting that centrosomal DEUP1+ foci and deuterosomes are not liquid-like structures (Figure 8 Supplementary 2)." How do you prove there is no fusion of deuterosomes?

      It is always difficult to prove the absence of something, we agree! But we did tens of movies with high temporal resolution and never observed fusion events. But, as we say in the previous question, the very early deuterosomes can be very small and we do not distinguish them from the DEUP1+ cloud by live imaging. So at this stage, we cannot say. But later on, during A- or G-stage and when deuterosomes are outside the cloud to be easily observed, we very often observe deuterosomes bumping into each others and stay in close contact for minutes, but then moving away. This, for us, supports the lack of fusion properties. But the question remains open. We now explain this in the manuscript and have added an example in video 28.

      If they are getting bigger as I think your imaging suggests from cloud to growth stage, then how is this happening?

      MT depolymerisation and dynein inhibition leads to the formation of very small deuterosomes. Dynein inhibition can even lead to a block in the formation of new deuterosomes suggesting that DEUP1 concentration is a crucial parameter for condensation into deuterosomes. Deuterosome growth may happen through oligomerization of DEUP1 molecules allowed by their dyne-independent concentration. Sorokin in 1968 proposed that a supersaturation of deuterosome components may lead to their solid crystallization into deuterosomes. Deuterosome size can also be regulated by a more complex molecular cascade, involving post-translational modifications of DEUP1 or PCM, such as phosphorylations driven by the cell cycle machinery. This would be consistent with the fact that deuterosomes are very big in the absence of CCNO, a cyclin required for entering the MCC cell cycle variant. This will need further investigations.

      I'm not sure FRAP actually proves fusion doesn't happen.

      Agreed, this is not what we wanted to say, we clarified. The FRAP experiment just suggests that it is not liquid-like.

      It is technically difficult to laser ablate individual or only subsets of deuterosomes...

      This is what was done but anyway, FRAP does not firmly show that deuterosome compartments are not liquid-like as we now precise.

      (4) How do you fix your cells for expansion as you have no preservation of cytoplasmic microtubules? You are saying that there is a "nest" of MTs but beta tubulin ONLY stains the cilia and centriole - why is this? Tyrosinated tubulin on regular confocal shows strong cytoplasmic staining. See Figure 3.

      Cytoplasmic microtubules do not preserve well through the expansion process. We did try a few different fixations and pre-extraction methods but they come at a trade-off to preserving centrioles. i.e. we could either preserve cytoplasmic tubes or centrioles but not both with the same processing method.

      (5) "PCNT puncta partially overlap with centrin (Figure 3 Supplementary 2C). At this stage, PLK4, the master regulatory kinase, and SAS6, one of the first centriolar components are either absent or present as small foci within the cloud, often on the wall of the parent centrioles (Figure 3B-C)." some arrows to highlight this would be useful - difficult to see?

      We have tried to make arrows on what is now Fig. 3 Supplementary 1 G, but there is to many CENTRIN colocalizing with PCNT. We have enhanced the contrast of the merge to make it more visible.

      (6) Figure 3I legend - what are the arrows pointing at? Yellow and white on inserts? ". Around the same time as tubulin, centrin is also recruited to procentrioles (Figure 3I). This stage is probably the stage that we previously documented as A"

      However you see centrin at DEUP1 foci in D, and you don't show any eg. SAS6 or PLK4 positive DEUP1+ structures lacking centrin specifically, centrin seems to be present on all the procentrioles in Figure 3I. Did I miss it where you show centrin negative procentrioles in the cloud?

      Fig. 3I (now Fig. Supplementary 1J), yellow arrows are pointing at centrioles with non-acetylated MT while white arrows point at acetylated MT. This is now indicated in the legend.

      Regarding CENTRIN, it is present as a diffuse staining around the centrosome since the very beginning of amplification (now in Fig. 3 Supplementary 1A with different contrasts), in addition to compose the parental centrioles. This staining can therefore overlap with DEUP1 staining when DEUP1 appears (Fig. 3 Supplementary 1B, E) but not necessarily. In live we observe that CENTRIN and DEUP1 foci can move independently at early stages (Fig. 2 Supplementary 1B, video 2). This is later on, as shown now in Fig. 3 Supplementary 1J (previously Fig. 3I), that procentrioles are all strongly positive for CENTRIN.

      A new paper (Laporte et al., Cell 2024) recently showed that the recruitment of CENTRIN on duplicating procentrioles first occurs at the distal end, visible by a small dot, and then appears gradually at the level of the inner scaffold when procentriole reach 160nm, the stage where POC5 appears, which corresponds to the A-to-G transition in our MCC progenitors (Al Jord et al., 2014). One can therefore consider that the same is happening in our cells, and that, with the CENTRIN cloud, we have difficulties to detect the distal CENTRIN dot. We have changed the text to add this reference and discuss CENTRIN apparition in MCC procentrioles.

      (7) " The DEUP1 asymmetry previously described at the centrosomal daughter centriole (Al Jord etal., 2014) becomes visible in some cells during the cloud stage (Figure 3B, N; Figure 3 Supplementary 2B) and in a majority of cells" difficult to see - maybe enlarge and single channel from Figure 3F-H in the supplemental Figure 3 to emphasise this?

      We have either changed the pictures or the contrast to be more representative with the quantifications. This is visible in Fig. 3A, D, E, G; Fig3. Supplementary 1E and now using correlative light and EM in Fig. 2 Supplementary 2, 3, 4 and Fig. 4 Supplementary 3-4. One has to look carefully at the daughter centriole (marked “dc”). We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both strutures are frequently sticked to each others on tens of nanometers.

      (8) Do you have videos of DEUP1 oscillations with nocodazole to show a lack of oscillations?

      We have now added videos of DEUP1 oscillations under nocodazole and dynapyrazole treatments.

      (9) "In addition, co-staining of centrioles and nuclear pore proteins show a tight colocalization(Figure 6 Supplementary 1B)." I see the colocalisation in panel 1 but less obvious with panel 2 maybe have some more zoomed in panels and some quantification of the colocalization? Is it more striking at the G stage than the D stage?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. There is too many centrioles and NPC, they cannot do otherwise than colocalize… What we want to say is that there is a tight connexion with the nuclear envelope. This is why we present a single z, to show that centrioles and NPC are on the same z-plane. This is also clearly visible for centrioles that are loaded on deuterosomes that are around the nuclear membrane in the XY plane. We also added a video to show an entire z-stack of this kind of staining.

      (10) "Indeed, SAS6 normally disappears from procentrioles when centrioles are docked, just beforeciliation (Al Jord et al., 2014). This suggests that centrioles were able to degrade SAS6, a process also dependent on APC/C (Strnad et al., 2007), but failed to disengage from deuterosomes." Figure 6 Supplement 1E-F - are you sure it wasn't that Sas6 wasn't loaded correctly at the earlier stage and so is reduced recruitment rather than premature disengagement of Sas6? If it is indeed premature disengagement of Sas-6 - what about CP110 - does the CP110 get loaded and is it still present in noco treated cells arrested in the D phase?

      We do not observe SAS6-negative procentrioles on deuterosomes at G-stage but only on deuterosomes in D-stage cells (cells with partly disengaged procentrioles). This is why we hypothesize that, because of the long duration of D-stage and knowing that SAS6 is finally degraded at the end of amplification (Al Jord et al., 2014), we are in the presence of cells where SAS6 has been degraded but where centrioles did not manage to disengage. This is now clarified in the text.

      (11) Can you track deuterostome splitting live? Maybe not enough spatial or time resolution?

      One has to monitor in 3D (multiple z because deuterosomes move a lot), 2 colors, high temporal resolution (dt=2-5’; to be able to track a single deuterosome), and long duration (deuterosomes are sometimes touching each other and then moving away, giving the impression that they split). This eventually leads to the bleaching of the mRuby fusion protein… We have put an example of what we think is a deuterosome splitting in Fig. 6E (former Fig. 7D). But we decided to finally monitor with low temporal resolution (dt=40’) to avoid photobleaching, and analyze numerous deuterosomes and cells to quantify the number and size of deuterosomes over time in single cells.

      (12) The MT nodes - can you segment the tyrosinated MTs and define nodes and then quantify theDEUP1 presence on them?

      Please see answer to reviewer 1 regarding this point.

      (13) Figure 8 supp 1 (E): Representative XY distribution of CEN2-GFP+ centrioles at the end of migration (Sas6 negative) in brain MCCs treated with DMSO, Nocodazole 1µM and 5µM (48h). Scale bar, 5µm Bit more detail on how you define fully migrated vs still migrating centrioles in z. You say you are using Sas-6 negativity to define fully migrated cells in the legend, yet you say noco treatment leads to premature sas-6 negativity, and yet the apical migration takes longer upon noco treatment?

      Nocodazole does not lead to premature SAS6 negativity but to a partial disengagement which lead to SAS6 negative “mature” centrioles being still connected to deuterosomes. We define complete migration when all the centrioles are on the apical side of the nucleus. We now clearly define what “apical” migration stands for in the main text and changed the pictograms in Fig. 8G to clarify this.

      (14) Figure 8H and video 18 - it isn't obviously clear to me that the noco-treated cells are "more erratic" or how you decide what counts as apically migrated successfully. How do you control for drift in z? Can you track individual centrioles as you did in untreated and define what is "erratic about their movement?

      Erratic means that the centrioles are moving away from each others, and back, in a non-predictable way, instead of migrating up and gathering. The drift in z of the whole cell is visible because there is always some centrioles, that are apically located at the beginning, that remains on the apical membrane, probably because they are already docked.

      We have indeed followed the centrioles individually in the nocodazole condition. However, in the control, the XYZ coordinates of one of the centrioles of the centrosome, which normally don’t move, are substracted to the coordinates of all the other centrioles as explained in the method section. This allows to have a subcellular reference, and to circumvent the movements of the cell, which are non-negligible at all at this timescale. In the nocodazole treated cells, the centrosomal centrioles share the erratic movements of the other centrioles and can migrate up and down, which exclude them as a reference. Since the nucleus is also moving a lot, we were left with no reference point.

      (15) Figure 8 supplement 1E can you quantify the final area of centriole patch in XY upon noco treatment?

      It was in main Fig. 8J and is now in Fig. 8 Supplementary 1F.

      (16) Figure 8J legend- MBB is never defined as an acronym.

      Thank you for pointing this.

      (17) Define what is the frequency and how is it calculated - Figure 8J.

      This is the MBB patch area in µm<sup>2</sup>

      Text edits:

      (1) "Altogether, these results suggest that, in this non-tissue-specific proxy of MCC progenitors, microtubules organize the onset of centriole amplification in the pericentrosomal region."

      Sentences have changed.

      (2) "Increasing the temporal resolution to 5-15s reveals that DEUP1+ foci observe an exhibit oscillatory dynamics to at the centrosome (Figure 2E, colored arrows, Video 3, 5/10 cells observed for 1-4min)."

      Sentences have changed.

      (3) "stage procentrioles were involved in this perinuclear migration and distribution. In fact, this dynamic is reminiscent of the centrosome migration that occurs during the G2-to-M progression in cycling cells in preparation for mitotic spindle organization. In cycling cells, this" Grammar - maybe change to "stage procentrioles were involved in this perinuclear migration and distribution. This is reminiscent of the centrosome migration that occurs during the G2-to-M".

      Sentences have changed.

      (4) "We then wondered whether these microtubule-dependent dynamics was were required for an efficient subsequent centriole disengagement during the following D-stage."

      Sentences have changed.

      (5) "Then, monitoring tens of disengagement movies, we identified a transient stage during which disengaging procentrioles redistribute isotropically in the 3 dimensions, along the nuclear membrane (Figure 6A, 4:30, Video 7) before losing its contact to migrate to the apical surface (Figure 6A, 6:30 to 14:00)."

      Sentences have changed.

      (6) Discussion: "Since pioneer electron microscopy studies on basal body production in quail oviduct MCC 35 years ago (Boisvieux-Ulrich et al., 1987, 1990; Boisvieux-Ulrich et al., 1989), this work is the first to assess the role of microtubules in the now finely described centriole amplification process. This"

      Sentences have changed.

      (7) "Using live imaging on brain MCC, we highlight the existence of a nest composed of DEUP1, PCNT and Centrin2, pre-assembled before the onset of centriole amplification onset."

      Sentences have changed.

      (8) "Recently, formation of DEUP1 pure condensates in solution as well as FRAP experiments after overexpression of DEUP1 in MCC progenitors suggested that deuterosomes where are not liquidlike structures (Yamamoto & Kitagawa, 2019). Consistently, we never observed fusion events of DEUP1."

      Sentences have changed.

      (9) "This reminds is reminiscent of the centriole-to-centrosome conversion occurring at the G2-M transition followed by the associated microtubule dependent nuclear migration of new centrosomes at mitosis onset (Agircan et al., 2014)."

      Sentences have changed.

      (10) "Following individual trajectories requires high resolutive resolution spatio-temporal live imaging while avoiding excessive light exposure which disturbs centriole migration (Boudjema et al., 2024)."

      Sentences have changed.

      (11) "Using high temporal resolution microscopy, we further identify that individual dynamics is are complex and can be splitted between divided into the baso-apical migration, where centrioles move in a processive and more..."

      Sentences have changed.

      Reviewer #3 (Recommendations For The Authors):

      (1) Growing MEF-MCCs on micropatterns has successfully mimicked the dynamics of centriole amplification in brain MCCs, allowing the authors to study the spatial origin of procentrioles. Since this is a powerful system, a more quantitative description of the system will be informative and beneficial for future studies. For example: What is the efficiency of this system? Do the cilia that form in MEF-MCCs motile?

      The system of MEF-MCCs has been described in a previous paper from the Kintner lab. It seems that growing the MEF-MCCs on micropatterns did not ameliorate the ciliation which is partial, probably due to the absence of an apico-basal polarity.

      (2) Figure 2: The analogy drawn by the authors between DEUP1 oscillatory dynamics and centriolar satellites is intriguing. In early amplifying cells within the cloud, do these DEUP1 structures co-localize with the satellite marker PCM1?

      We have added immuno stainings of PCM1 in mRuby-DEUP1 / CEN2-GFP cells in Fig. Supplementary 2E. Within the centrosomal cloud, DEUP1 colocalizes with PCM1. Interestingly, this PCM1 concentration at the centrosome is dependent, at least in part, on dyneins. Then, PCM1 can localize around the deuterosomes, but it is never colocalized with deuterosomes (not shown). This is also showed by immuno-EM in Zhao et al., 2019. Although it was shown that PCM1 is a proximity interactor of DEUP1 (called ccdc67 at that time) by Firat-Karalar et al., 2014., absence of PCM1 staining on deuterosomes does not favor the hypothesis of PCM1 and DEUP1 being part of the same entities. One could hypothesizes that DEUP1 is transcribed locally within the satellites, explaining the colocalization of the 2 proteins and the + BioID results, and then form PCM1negative deuterosomes.

      (3) The authors propose a physical link between deuterosomes and centrosomes based on their oscillatory behavior. How are the oscillatory dynamics of DEUP1 affected by nocodazole treatment or inhibition of microtubule motors (i.e ciliobrevin treatment)?

      These oscillations are inhibited by nocodazole (Fig. 4D). They are also inhibited by dynapyrazole (Fig. 4D). We never succeeded in having a nice disruption of the Golgi apparatus with ciliobrevin and therefore we did not used it.

      (4) In addition to nocodazole treatment, it would be important to determine the consequences of microtubule stabilization by taxol and inhibition of microtubule motors during critical stages of centriole amplification where microtubules are reported to play a role for the first time in this manuscript. Another interesting area of investigation will be to study the extent to which microtubule PTMs contribute to these processes.

      We now blocks dyneins during the different stages of amplification. The results are in main and associated Fig. 4, 5, 7, 8. The role of microtubule PTM, is not in the scope of this manuscript.

      (5) Describing microtubule dynamics along with Centrin/DEUP1 dynamics will be informative in assessing whether these structures associate and/or move along microtubules? Have the authors performed their imaging experiments with SIR tubulin?

      Yes, we have tried hard! But we have encountered different obstacles:

      3-color video microscopy is phototoxic,

      siRTubulin is bleaching very rapidly

      The density of microtubules in MCC makes the observation hardly informative

      (6) Figure 5: The role of PLK1 in centriole-centrosome conversion and generation of multiple MTOCs can be tested with a PLK1 inhibitor for further confirmation.

      We have also tried but inhibiting Plk1 blocks the A-to-G and G-to-D transitions so it was not possible to uncouple the role of Plk1 in stage transitions versus centriole maturation.

      (7) Figure 6: The tight co-localization of nuclear pore proteins with centrioles poses questions about the role of nuclear pore proteins or other nuclear proteins that are associated with centrioles during centriole disengagement and migration. Considering the existing literature on centrosome-nucleus attachments, can there be a way to test this question within the scope of this manuscript?

      We have tried to deplete Nup133 but it’s killing the cells. Our additional experiments now show that the nuclear migration of centrioles during G-stage is dynein dependent, reinforcing the parallel with centrosome migration in prophase. We also added results from our scRNA sequencing (Fig. 5 Supplementary 1) showing that some key players of centriole migration to the nuclear membrane are conserved in the MCC cell cycle variant, and expressed with a comparable dynamics as to the canonical cell cycle.

      (8) Figure 8: Manually tracking a subset of migrating centrioles to define their dynamics during centriole migration and docking provides valuable analysis for determining the molecular mechanism of these processes. In addition to microtubules, does actin contribute to this process? Since centrioles eventually migrate to the apical side in nocodazole-treated cells, there should be other molecular players involved in this process.

      We did block actin polymerization but we found that the different stages were affected and that it would be better to dedicate a whole manuscript on the role of actin during each stage of amplification. We discuss the migration mechanism, and the putative role of actin, in the discussion.

      (9) The legends for Supplementary Figures 1 and 2 in Figure 3 are mixed and need correction.

      Figures have been remodelled.

      (10) In Figure 3P, the term "PLK4+" is labeled in bright green, which is not clearly visible. It maybe beneficial to change the color of this label for better visibility.

      We have tried to correct this.

      (11) Figure 6F quantifies "% tethered flowers" on the nuclear membrane. When quantifying, is the3D localization of DEUP1 flowers in both DMSO- and Noc-treated cells considered? A flower may appear to be on the nucleus in 2D, but it could be detached from the membrane in a 3D view.

      The quantifications are done in 3D. However, flowers that are below or above the nucleus are not quantified since the space is confined and the resolution in z to small to see whether they are connected or not. This is now precised in the legend.

      Before the editors proceed with an updated assessment, they've requested that we pass on some of the comments that have arisen as part of the evaluation of your revised manuscript. They feel that these concerns should be addressed before we proceed with issuing a formal assessment and publishing the revised Reviewed Preprint:

      We thank the reviewers and the editors for the corrections and insighfull comments. We apologize for our delayed answer and hope our corrections in the main text and some of the figures will give them satisfaction.

      The revised manuscript is greatly improved with nice new data regarding the role of microtubules. It also has changed quite a bit including the title. The new focus is on the cell and centriole cycle variants in MCC. While this helped to focus the study, there remains an important issue related to the interpretation of the data and the proposed 2-in-1 cycle model. Before providing the final updated assessment, we ask you to address the following points (which were raised already in the first round of review): The manuscript still contains statements that are not aligned with published work and the current view in the field regarding the timing of events during canonical centriole biogenesis. These timings are in conflict with your model that 2 centriole cycles are "superposed" in the MCC cell cycle variant, as currently presented. An alternative straightforward interpretation would be that multiciliogenesis uses an accelerated centriole duplication cycle where key steps occur concomitantly or in short succession instead of being separated by mitotic divisions as in the canonical cycle.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we think we have proposed. When correcting our confusions as regard to centriole-to-centrosome conversion (as explained below) and putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (corrected Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a superposition of events that; although driven by the same molecular machinery, are normally occuring in two consecutive cell cycle. We explain ourself briefly in two paragraphs, before answering point by point to the questions of the reviewers.

      As regard to centriole-to-centrosome conversion:

      We thank the reviewer for pointing out that we used “MTOC conversion” for what is normally called “centrosome maturation”. We have removed the term “centriole-to-centrosome conversion” during the first round of revision but we now realize that “MTOC conversion” leads to the same misinterpretation as regard to the literature on centriole duplication.

      The reviewer asks us to refer to the work of the Tsou lab (Wang 2011, reference now added in the manuscript) showing that daughter centrioles are “modified” (e.g. recruit PCM, become competent for MT nucleation and duplication) during late M/early G1. This “centriole-to-centrosome conversion” can’t occur for our procentrioles at this stage since they are not even born during the mitosis that precedes MCC differentiation. Also, in our cells, such modification does not include the capacity to become competent for duplication since we know that procentrioles become basal bodies without making any round of duplication (Al Jord et al., 2014).

      Also, we have not done the experiments to tackle the question on when our centriole become “modified-like”. What we can say is that during A-stage, they become progressively positive for PCM (Fig. 5 Supplementary 2) and a weak signal shows that some MT are seen emerging from them (Fig. 5 and Fig. 5 Supplementary 2, and see point by point answer).

      What we do see is that, at the A-to-G transition, they increase their PCM recruitment, show clear and strong MTOC ability (sometimes as strong as the centrosomal centrioles), and that this is associated with migration and separation of centrosome/deuterosomes around the nuclear membrane (Fig. 5). We therefore connect this to what occurs at the G2/M transition which is an increased recruitment of PCM protein, an increased ability to nucleate MT, associated with centrosome migration and separation at the nuclear membrane. Since this process in the canonical cell cycle is called “centrosome maturation”, we therefore should refer to this term in our study. However, centrioles in the MCC variants are not organized in centrosomes, so we now compare what we see to the “centrosome maturation” of the canonical cell cycle with an associated reference (Joukov et al., 2018), but name it “centriole maturation”.

      We have modified the text (track changes visibles) and the schemes (Fig. 5, Fig. 5 Supplementary 1 and 2, Fig. 9, Fig. 9 Supplementary S1; new versions uploaded) accordingly.

      As regard to 1.5 or 2 cell cycles

      Except for the “MTOC conversion” that we have now changed, as explained above, we think our work does suggest (depicted on Fig. 9) what the reviewer states for centriole duplication: “In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs. Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium)”.

      We feel that going from early S to a G1 phase, after 2 mitosis, is what one can call “2 cell cycles”. One of the paper that inspired us a lot when studying how the cell cycle machinery can drive centriole amplification in MCC is a paper from Jadranka Loncarek team (Kong et al., 2014) where they also state that “nascent centrioles gradually mature through 2 cell cycles”. Very interestingly, in this study they show that when they enhance Plk1 activation, they could erase centriole age and new procentrioles are able to recruit PCM and appendages within only 1 cell cycle, without mitotic progression, like what we see in MCC. We have added the reference in our discussion.

      Point by point answer

      (1) Original work on canonical centriole disengagement and centriole-to-centrosome conversion should be cited (e.g. PMID: 16862117, PMID: 21576395)

      As explained earlier, we used the wrong term since the begining. We do not speak about the centriole-to-centrosome (nor MTOC) conversion since we do not test when centriole modification (Wang et al., 2011) occurs in the MCC cell cycle variant. We know that PCNT is present on the procentrioles during A-stage (as shown in Fig. 5 Supplementary 2B), but we do not know when it is recruited (UExM did not work properly with this antibody). We quantify a weak MT staining in regrowth experiment during A-stage and see that procentrioles can be connected to MT in both brain MCC and MEFs (as shown in Fig. 5D, E for brain MCC and Fig. 5 Supplementary 2F for MEFs) , but we do not know when during A-stage they become competent for nucleation. We therefore did not speak about this process that we do not document. What we clearly document/quantify is the enhanced MT nucleation capacities at the A-to-G transition, concomitent with the nuclear migration (easily defined with Cen2-GFP or GT335 stainings) and that we compare to centrosome maturation occuring at the canonical G2/M transition.

      (2) The authors state in several places that canonical centriole formation and maturation takes two iterations of the canonical cell cycle. This is imprecise. Based on the above work and work by others, the broadly accepted view is that it takes 1.5 cell cycles. This difference matters for the final proposed model (see below). Reviewed e.g. here: PMID: 20869612; PMID: 30601682

      Our answer is in the preamble.

      (3) "Centriole maturation cycle superposes with centriole elongation cycle in the MCC cell cycle variant": Your description of the canonical cycle differs from the current view in the field. In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). All this occurs in 0.5 cycles. Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs (total of 1.5 cell cycles). Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium).

      (4) Fig 5A, B and Fig. 9

      (a) Are 2 separate figures needed for the model? They seem redundant.

      We find it easier not to wait Fig. 9 to have the first part depicted.

      (b) The model shows loss of SAS6 throughout G1, but this already occurs during M/early G1

      Thanks. It was already ok in Fig. 9, we have modified for Fig. 5.

      The model shows "MTOC capacity/conversion" during S phase, but this occurs during early G1

      Thanks a lot, as explained earlier, we used the term MTOC conversion occurring in G1 for what is normally called centrosome maturation occurring in G2/M, as explained earlier. We do not speak anymore of MTOC conversion since we have not tackled this question (explained above). We have therefore removed MTOC conversion in the texts and the schemes and replaced it by “centrosome maturation” for the duplication cycle, and by “enhanced MT nucleation capacity” for the MCC cycle. To be clearer and schematize that procentrioles are competent for MT nucleation before G2/M or A/G transitions, we have added some MT nucleated from G1 procentrioles during the canonical cycle, and from late A-stage procentrioles during the MCC cycle.

      The model shows disengagement only in the second M phase, but this occurs already at the first M phase, directly following centriole biogenesis, right before centosome conversion.

      This is a big edition error in both Fig. 5 and 9. Of course the daughter centriole disengage during the first M-phase. This has been changed. Thanks a lot for spotting it. This, however does not contradict the hypothesis of superposition.

      We also added the acquisition of distal appendage which was written in Fig. 5 but not in Fig.

      9 for duplication during the second M-phase.

      When the correct timings are incorporated in the figure, the proposed superposition of two cycles is not an accurate description of the events. Instead, your data seem consistent with a model where MCC incorporates all steps in one cell cycle variant that lacks mitoses, so that disengagement and MTOC conversion occur together with centriole elongation, followed immediately by acquisition of DAs and SDAs.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we tried to propose. When putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a super opposition of events that; although driven by the same molecular machinery, are normally occurring in two consecutive cell cycle. This is notably consistent with the findings of Kong et al., 2014 cited previously.

      (5) While all reviewers felt that there was no need to introduce the new term "nest", they leave it to the authors to keep it. However, the authors may want to consider that the term is still not introduced and explained properly, which may confuse readers. For example, while this section reads like an introduction to the term: "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC", the term is already used two times before without explanation. The first mentioning is at the beginning of the results section and is followed by citations, which gives the impression that these studies describe the nest, which is not the case.

      The first mention of “nest” is in the end of introduction resuming the findings of the paper where the term is in the following context: “we found that centriole amplification emerges in a pericentrosomal “nest” concentrating core centriole/deuterosome elements”. We looked at nest definition in the Collins Dictionnary : “a structure or other place where creatures, esp. birds, give birth or leave their eggs to develop”, we felt this was clear. We added quotation marks around the term nest.

      Then, the result section opens with this sentence: “The origin of amplified centrioles in MCC remains controversial. Some live imaging experiments and electron microscopy suggest that the centrosome could constitute a nest for centriole and deuterosome biogenesis (Al Jord et al., 2014; Kalnins et al., 1972; Mori et al., 2017), but others have proposed that procentriole-loaded deuterosomes emerge independently from the centrosome location, all over the cytoplasm (Nanjundappa et al., 2019; Sorokin, 1968; Zhao et al., 2013, 2019).”. Here, the term nest is again used as a place of birth for centrioles and deuterosomes which is what is actually proposed in these papers. First, Kalnins el al., in 1969 (we made an error on the reference date, this has been changed), resume in their abstract “This observation suggests that all of the clusters may form initially in close association with the diplosomal centrioles”. Then, not to mention Al Jord 2014 which comes from our lab, the title of Mori et al. is “Cytoplasmic E2f4 forms organizing centres for initiation of centriole amplification during multiciliogenesis”, and in the paper, they show that E2F4 accumulates at the centrosome. This is now also proposed by collaborators for MCIDAS (Lu et al., 2025). We feel that these references, which are often omitted, are appropriated at this location.

      Then we continue with: “To test whether microtubules drive the organization of a centrosomal nest from which procentrioles emerge”, which keeps the notion of the place of birth.

      Then the title "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC" arrives. In this section we first speak about a pericentriosomal cloud on which we zoom in using CLEM, to then conclude at the end of the section “Altogether live imaging mRuby-DEUP1/CEN2-GFP during early A-stage suggests that core deuterosome and centriole components are concentrated in a primordial cloud around the centrosome, which constitutes a nest where centrioles and deuterosomes concomitantly form before they move away from the centrosomal region (Fig. 2F)”.

      Finally, we begin the discussion section regarding the nest by: “We named this transitory compartment a “nest” since deuterosomes and procentrioles emerge specifically in this region and grow while moving away from it.”

      During the first revision, we tried to make it clearer. If this is still not the case after and the reviewer has another proposition of definitions/phrasing, we will be glad to consider it.

      As replied to the other reviewer, the term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the transitory region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      The following comments from Reviewer #3 may also provide further context regarding the editors' remaining concerns:

      The authors have done an excellent job addressing the points I raised overall, and the revision is substantially improved in focus and clarity. That said, some concerns raised by other reviewers, particularly regarding terminology and statistical analysis, could have been addressed more fully. One issue remains insufficiently resolved. Several quantitative analyses (for example Fig. 5C and 5E) still appear to rely on pooled single-event measurements collected across three independent experiments. This approach can overstate statistical significance. The authors indicate in their rebuttal that they use chi-square tests to compare proportions and to justify pooling across replicates. However, I am not convinced this addresses the issue for the intensity-based and single event distributions shown in the panels specified above. I recommend that these key analyses be represented with biological replicates shown explicitly (superplot-style, with replicates distinguished).

      Our reply was for the comparison of proportions and not the intensity-based and single event distributions shown in the panels Fig. 5C and Fig. 5E. We have now changed our plots to represent biological replicates explicitly (superplot-style, with replicates distinguished). As for the statistical analysis: we evaluated differences in marker intensity between A-stage and G-stage samples using a linear regression model, with stages as the main effect and replicate as a fixed covariate, to account for batch variation. Statistical significance was assessed using Type II ANOVA.

      Separately, I continue to feel that some newly introduced terminology (for example, the "nest") may not be necessary at this stage. It may be sufficient to describe these structures and focus on their spatiotemporal behavior, composition, and measurable features, rather than assigning new names. Having read the authors' response, I understand that they would like to retain this terminology, which is acceptable; however, it may not be readily adopted by the field.

      The term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      Minor correction (remove "in MCCs" part from the following sentence):

      In MCC, PCM1 depletion alters deuterosome formation and centriole production in brain and airway MCC (Hall et al., 2023; Zhao et al., 2021).

      Done

    1. eLife Assessment

      This manuscript provides valuable insight into how genome organization changes as cells progress through the cell cycle after mitotic exit, identifying two sharp genome remodeling events at G1-S and to a lesser extent, at S-G2 transitions. The conclusions are supported by solid, rigorous data, including sequencing and orthogonal imaging data. The use of sorted unsynchronized cells rather than cells treated with drugs is a particular strength.

    2. Reviewer #1 (Public review):

      This work convincingly shows that, rather than gradually "evolving" throughout interphase, global chromatin architecture undergoes unexpectedly sharp remodeling at G1-S (and to a lesser extent, S-G2) transitions. By applying "standard" Hi-C analyses on carefully sorted cells, the authors provide an excellent temporal view of how global chromatin architecture is changed throughout the cell cycle. They show a surprisingly abrupt increase in compartmentation strength (particularly interactions between the "active" A compartments) at G1-S transition, which is slightly weakened at S-G2 transition. Follow-up experiments show convincingly that the compartment "maturation" does not require the DNA synthesis accompanying S phase per se, but the authors have not identified the responsible factors (work for future publications). The possible biological ramifications of these architectural changes (setting up potential replication "factories", and/or facilitating transcription-replication conflict resolution, both more pertinent for the active A compartments, which are most affected) have been well discussed in the article, but still remain speculative at this stage.

      My major criticism of this article is aimed more at the state of the field in general, rather than this specific article, but it should be discussed to give a more balanced view: what actually is a chromatin compartment? Chromosomal tracing and live tracking experiments have shown that the majority of "structures" identified from Hi-C experiments are statistical phenomena, with even "strong" interactions only being infrequent and transient. A-B compartments are "built up" from multiple very low-frequency "interactions", so ascribing causal effects for genome functions is even tougher. As a result, I have very little confidence in the results of the authors' polymer simulations and their inferred "peninsula" A compartment structures without any other supporting experimental data.

      Comments on revised version.

      The authors have included orthogonal DNA FISH evidence to support their claims which greatly strengthens the manuscript. Their further precisions within the discussion have answered all of my previous concerns with the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript by Choubani et al presents a technically strong analysis of A/B compartment dynamics across interphase using cell-cycle-resolved Hi-C. By combining the elegant Fucci-based staging system with in situ Hi-C, the authors achieve unusually fine temporal resolution across G1, S, and G2, particularly within the short G1 phase of mESCs. The central finding that A/B compartment strength increases abruptly at the G1/S transition, stabilizes during S phase, and subsequently weakens toward G2 challenges the prevailing view that compartmentalization strengthens monotonically throughout interphase. The authors further propose that this "compartment maturation" is triggered by S-phase entry but occurs independently of active DNA synthesis, and that it involves a consolidation and large-scale reorganization of A-compartment domains.

      Strengths:

      Overall, this is a thoughtfully executed study that will be of broad interest to the 3D genome community. The data are of high quality, and the analyses are extensive, albeit not completely novel. In particular, previous work (Nagano et al 2017 and Zhang et al 2019) has shown that compartments are re-established after mitosis and strengthened during early interphase, and single-cell Hi-C studies have reported changes in compartment association across S phase. In particular, Nagano et al show that DNA replication correlates with a build-up of compartments, similar to what is presented here, with the authors' conclusion that compartment strength peaks in early S. The idea that it weakens toward G2, rather than continuing to strengthen, appears to be novel and differs from the prevailing framing in the literature.

      Comments on revised version.

      The authors have responded constructively to my major conceptual concerns. The distinction between DNA synthesis and replication initiation has been clarified appropriately. The additional insulation analysis substantially strengthens the argument that compartment maturation is not simply a consequence of changing loop extrusion dynamics, although I would encourage slightly more cautious wording regarding "independence" from cohesin-mediated extrusion. The peninsula model is now framed appropriately as a heuristic interpretation and supported by orthogonal imaging data. Finally, the discussion of conservation across cell types has been appropriately tempered. Overall, I believe the manuscript has been significantly improved.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work convincingly shows that, rather than gradually "evolving" throughout interphase, global chromatin architecture undergoes unexpectedly sharp remodeling at G1-S (and to a lesser extent, S-G2) transitions. By applying "standard" Hi-C analyses on carefully sorted cells, the authors provide an excellent temporal view of how global chromatin architecture is changed throughout the cell cycle. They show a surprisingly abrupt increase in compartmentation strength (particularly interactions between the "active" A compartments) at G1-S transition, which is slightly weakened at S-G2 transition. Follow-up experiments show convincingly that the compartment "maturation" does not require the DNA synthesis accompanying S phase per se, but the authors have not identified the responsible factors (work for future publications). The possible biological ramifications of these architectural changes (setting up potential replication "factories", and/or facilitating transcription-replication conflict resolution, both more pertinent for the active A compartments, which are most affected) have been well discussed in the article, but still remain speculative at this stage.

      We thank Reviewer #1 for their positive and constructive assessment of our work, and we agree that the questions of responsible factors and biological ramifications are important directions for future studies.

      My major criticism of this article is aimed more at the state of the field in general, rather than this specific article, but it should be discussed to give a more balanced view: what actually is a chromatin compartment? Chromosomal tracing and live tracking experiments have shown that the majority of "structures" identified from Hi-C experiments are statistical phenomena, with even "strong" interactions only being infrequent and transient. A-B compartments are "built up" from multiple very low-frequency "interactions", so ascribing causal effects for genome functions is even tougher. As a result, I have very little confidence in the results of the authors' polymer simulations and their inferred "peninsula" A compartment structures without any other supporting experimental data.

      We thank the reviewer for raising this important conceptual point. This issue extends beyond the scope of the present study but reflects an important ongoing discussion in the 3D genome field regarding the biological interpretation of chromatin compartments.

      We agree that Hi-C interactions should not be interpreted as stable pairwise contacts present in every cell. A growing body of evidence from chromatin tracing and live-cell imaging studies has demonstrated that many chromatin interactions identified by Hi-C are probabilistic and dynamic, with substantial cell-to-cell variability. Relatively speaking, however, A/B compartment organization represents a robust population-level property of genome organization that is highly reproducible across biological replicates and closely correlates with multiple independent genomic features. In particular, replication timing (RT) correlates very well with A/B compartment organization, with early and late RT domains corresponding to A and B compartment domains, respectively.

      Furthermore, single-cell DNA replication sequencing (scRepli-seq) analyses have revealed remarkably low cell-to-cell variability in RT, suggesting that RT profiles and A/B compartment organization reflect biologically meaningful and relatively stable features of nuclear architecture rather than purely statistical artifacts. Thus, while individual chromatin contacts may be transient and probabilistic, the megabase-scale compartment organization inferred from them appears sufficiently reproducible to support reproducible RT programs and other genome functions. Additional support comes from decades of work on DNA replication demonstrating that spatiotemporal replication patterns, visualized as replication foci following short EdU pulses, are remarkably reproducible between individual cells throughout S-phase progression. These patterns reveal clear spatial segregation between early-replicating A-compartment regions and late-replicating B-compartment regions even at the single-cell level.

      To directly address the reviewer’s concern that A/B compartment organization might represent only an ensemble-level statistical phenomenon without biological relevance at the single-cell level, we performed L1/B1-EdU DNA FISH on asynchronous mESCs and MC12 embryonic carcinoma cells. L1 elements are enriched in B compartment domains, while B1 elements are enriched in A compartment domains, allowing visualization of compartment segregation in individual nuclei across the cell cycle. This single-cell analysis confirmed our Hi-C findings: compartment segregation increased from G1 to early S, remained elevated throughout S phase with reduced cell-to-cell variability, and then weakened in G2. Thus, compartment segregation is detectable in single cells, and the temporal dynamics of compartment maturation identified by population Hi-C were independently recapitulated at single-cell resolution. We have added a new Results section describing these findings titled “Stepwise A/B compartment reorganization during interphase is conserved at single-cell resolution”, including new Figure panels 2D–H and Figure S5.

      Regarding the polymer simulations, we agree that these models should be interpreted with caution. We do not view them as direct representations of individual nuclei, but rather as heuristic models that help visualize structural trends present in the Hi-C data. To make this point explicit, we have added the following statement to the revised manuscript: “We note that these models are derived from population-averaged Hi-C data and should therefore be interpreted as a heuristic framework for understanding A/B compartment dynamics, rather than as definitive representations of individual nuclei.”

      That said, we did try to provide orthogonal experimental support for the "A peninsula" model by performing DNA FISH. In brief, we measured distances between probe pairs spanning two A domains on chromosomes 2 and 15 across different cell-cycle stages. We observed significant increases in inter-probe distances from G1 to early/mid S, with the most pronounced changes involving the central probes (i.e., probes located near the domain center), consistent with physical extension of the A domain during S phase. While these data do not prove the exact geometry depicted by the model, these findings provide independent experimental support for the peninsula model as a simplified but biologically grounded interpretation of the Hi-C data. These results are described in the Results section titled “A-compartment consolidation during S-phase involves enhanced long-range contacts and structural reorganization” and are presented in new Figure panels 5D–F and Figure S12.

      We thank the reviewer again for raising this important conceptual issue, which prompted us to better clarify both the biological interpretation and the limitations of our analyses.

      Specific minor points:

      (1) A better explanation for how Figure 1E was generated is required, because this figure could be very misleading. Figure 1F and all other cis-decay plots (and the Hi-C maps themselves) show that the strongest interactions are always at smaller genomic separations, so why should there be more "heat" at the megabase ranges in Figure 1E?

      We appreciate the reviewer's observation. The apparent discrepancy is simply due to the fact that the decay plot (Fig. 1E in the original submission, now Fig. S2C) does not include the shortest-range interactions. The lowest distance plotted is 25 kb, following the method originally described in Nagano et al. (Nature, 2017), which we used as a reference. The shortest-range interactions (below 25 kb) are indeed the most enriched, as seen on the diagonal of the Hi-C maps (Fig. 2A) and in the standard cis-decay plot (Fig. 1F in the original submission, now Fig. S2F). With the 25 kb cutoff in place, the "heat" observed at megabase distances (specifically 12–50 Mb) in early/mid G1 corresponds to the dark, non‑specific band around the diagonal visible in the Hi-C maps at the same time points. This is also reflected in the cis-decay plot (Fig. S2F), where distances in that range appear above the expected curve (a "bump" rather than a linear decay).

      To avoid confusion, we have updated the figure legend accordingly (Fig. S2C): “(C) Contact decay profiles for all cell cycle phases, plotted from 25 kb to 50 Mb, illustrating a continuum of cis-interactions and a progressive shift from long-range (> 12 Mb) to short-range (< 1 Mb) interactions during the G1-to-S phase transition.”

      We hope this explanation clarifies the figure.

      (2) An ultra-high-resolution Hi-C study (Harris et al., Nat Commun, 2023) identified very small A and B compartments, including distinctions between gene promoters and gene bodies, raising further questions as to what the nature of a compartment really is beyond a statistical phenomenon. It is unreasonable to expect the authors to generate maps as deep as this prior study, but how much do their conclusions change according to the resolution of their compartment calling? The authors should include a balanced discussion on the "meaning" of A/B compartments.

      We thank the reviewer for highlighting recent ultra-high-resolution work, such as Harris et al. (Nat Commun, 2023), which reveals compartment-like features at much finer genomic scales. We agree that these findings raise important questions regarding the scale-dependence and interpretation of A/B compartmentalization.

      In our study, we specifically focus on coarse-grained compartment organization, analyzed across multiple resolutions (from ~1 Mb to sub‑megabase scales). Importantly, the key conclusions, including the abrupt strengthening of compartmentalization at the G1/S transition, are robust across these resolutions.

      We also note that fine-scale compartment-like features likely operate under different rules than larger-scale compartments. Recent evidence suggests that these "micro‑compartments" are more dynamic and transient (Harris et al., Nat Commun, 2023; Goel et al., Nat Struct Mol Biol, 2025), whereas the large-scale compartments analyzed here capture more stable, global segregation patterns. Understanding how these two regimes relate to one another remains an important open question.

      We have added the following statement in the Discussion acknowledging the scale-dependent nature of compartmentalization: “At the same time, recent ultra-high-resolution Hi-C studies [36,37] have revealed compartment-like features at much finer genomic scales, emphasizing that A/B compartmentalization is, to some extent, inherently scale-dependent. Understanding how these fine-scale, often transient micro-compartments relate to the more stable, large-scale segregation patterns described here will be an important direction for future studies.”

      Reviewer #2 (Public review):

      Summary:

      This manuscript by Choubani et al presents a technically strong analysis of A/B compartment dynamics across interphase using cell-cycle-resolved Hi-C. By combining the elegant Fucci-based staging system with in situ Hi-C, the authors achieve unusually fine temporal resolution across G1, S, and G2, particularly within the short G1 phase of mESCs. The central finding that A/B compartment strength increases abruptly at the G1/S transition, stabilizes during S phase, and subsequently weakens toward G2 challenges the prevailing view that compartmentalization strengthens monotonically throughout interphase. The authors further propose that this "compartment maturation" is triggered by S-phase entry but occurs independently of active DNA synthesis, and that it involves a consolidation and large-scale reorganization of A-compartment domains.

      Strengths:

      Overall, this is a thoughtfully executed study that will be of broad interest to the 3D genome community. The data are of high quality, and the analyses are extensive, albeit not completely novel. In particular, previous work (Nagano et al 2017 and Zhang et al 2019) has shown that compartments are re-established after mitosis and strengthened during early interphase, and single-cell Hi-C studies have reported changes in compartment association across S phase. In particular, Nagano et al show that DNA replication correlates with a build-up of compartments, similar to what is presented here, with the authors' conclusion that compartment strength peaks in early S. The idea that it weakens toward G2, rather than continuing to strengthen, appears to be novel and differs from the prevailing framing in the literature.

      We thank Reviewer #2 for their thoughtful assessment and critique. We address their specific concerns below.

      Weaknesses:

      That said, several aspects of the conceptual framing and interpretation would also benefit from further clarification, and the mechanistic interpretation of the reported compartment dynamics requires more careful positioning relative to established models of genome organization. Specific concerns are outlined below:

      (1) One of the major conclusions of the study is that compartment maturation does not require ongoing DNA replication. However, the interpretation would benefit from more precise wording. Thymidine arrest still permits licensing, replisome assembly, and other S-phase-associated chromatin changes upstream of bulk DNA synthesis. Therefore, their data, as presented, demonstrate independence from DNA synthesis per se, but not necessarily from the broader replication program. Please clarify this distinction in the text and interpretations throughout the manuscript.

      We thank the reviewer for this important distinction. We agree with their point and have never claimed that compartment maturation is independent of the broader replication program. That is why we carefully used the term "active DNA synthesis" rather than "replication" throughout the manuscript.

      However, we acknowledge that one sentence in the text was ambiguous. The original sentence read: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a pre-replicative state where replication had not yet initiated, although cell-cycle markers indicated entry into S-phase.”

      We have now revised it to: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a state where the replication program (including origin licensing, replisome assembly, and helicase activation) has been initiated, as indicated by cell-cycle markers, but ongoing DNA synthesis (elongation) is blocked. ”

      This clarifies that compartment maturation is independent of active DNA synthesis (elongation) but not necessarily independent of upstream replication-associated processes. The change has been made in the manuscript.

      (2) A major conceptual issue that is not addressed at all is the well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartmentalization. Numerous studies have shown that loss of cohesin or reduced loop extrusion leads to stronger compartment signals, whereas increased cohesin residence or enhanced extrusion weakens compartmentalization. Given this framework, an obvious alternative explanation for the authors' observations is that the abrupt increase in compartment strength at G1/S, and its decline toward G2, could reflect cell-cycle-dependent modulation of cohesin activity rather than a compartment-intrinsic "maturation" program.

      The manuscript does not explicitly consider this possibility, nor does it examine loop extrusion-related features (such as loop strength, insulation, or stripe patterns) across the same cell-cycle stages. Without discussing or analyzing this widely accepted model, it is difficult to distinguish whether the reported compartment dynamics represent a novel architectural mechanism or an indirect consequence of known changes in extrusion behavior during the cell cycle. I strongly encourage the authors to analyze their data to determine if they observe anti-correlated loop changes at the same time they observe compartment changes. Ideally, the authors would remove loop extrusion during interphase using well-established cohesin degrons available in mESCs and determine if the relative differences in compartment dynamics persist.

      We thank the reviewer for raising this interesting point. We agree that there is a well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartment strength in the literature.

      To test whether cell cycle compartment dynamics, particularly compartment maturation at the G1/S transition, could be explained by changes in loop extrusion, we analyzed insulation at RAD21/CTCF sites (mESC data from Hansen et al., eLife, 2017) across the cell cycle. During normal cycling, we indeed observed an anti-correlation: insulation dropped as compartment strength increased at the G1/S transition. However, in G1/S-arrested cells, insulation did not drop compared to late G1 (it even slightly increased) even though compartment maturation still occurred, indicating that the two processes can be uncoupled. This is consistent with other studies showing that loop extrusion and compartment dynamics are driven by independent mechanisms (Nora et al., Cell, 2017; Zhang et al., Nat Commun, 2021), although we cannot fully rule out some contribution from loop extrusion dynamics without direct cohesin degron experiments.

      We have added a new Results section describing these findings titled “Compartment maturation is independent of cohesin-mediated loop extrusion”, including new Figure panels 3H, I, and Figure S7.

      (3) The proposed "peninsula-like" A-domain structures are inferred from ensemble Hi-C data and polymer modeling, rather than directly observed physical conformations. That is, single-cell imaging data clearly have shown that Hi-C (especially ensemble Hi-C) cannot uniquely specify physical conformations and that different underlying structures can produce similar contact patterns. The "peninsula" language, as written, risks being interpreted as a literal structural model rather than a conceptual visualization. Instead of risking this as just another nuanced Hi-C feature in the field, the authors could strengthen the manuscript by either (i) explicitly framing the peninsula model as a heuristic description of contact redistribution rather than a definitive physical architecture, or (ii) discussing alternative structural scenarios that could give rise to similar Hi-C patterns. Clarifying this distinction would improve the rigor and help readers better understand what aspects of A-compartment consolidation are directly supported by the data versus model-based extrapolations. For example, it would be useful to clarify whether the observed increase in long-range A-A contacts reflects spatial extension of internal A regions, changes in loop extrusion dynamics, increased compartment mixing within the A state, or population-averaged heterogeneity across alleles.

      We thank the reviewer for this important clarification. We agree that the "peninsula" model should be framed as a heuristic description. As detailed in our response to Reviewer #1 (see above), we have added a disclaimer to the manuscript and provided orthogonal DNA FISH support for physical extension of A-domains during S phase. We have also ensured that the language emphasizes the conceptual nature of the model.

      (4) The extension of the analysis to additional cell types using HiRES single-cell data is a valuable addition and supports the idea that compartment maturation is not unique to mESCs. However, the limitations of these data, in particular, the limited phase resolution, in addition to the pseudo-bulk aggregation and variable coverage, should be emphasized more clearly in the main text. Framing these results as evidence for conservation in principle, rather than definitive proof of identical dynamics across tissues, would be a more appropriate framing.

      We agree with the reviewer. We have already explicitly acknowledged the limited temporal resolution and variable coverage of the HiRES dataset in the main text. To better reflect its supporting role, we have moved the HiRES figure (previously Fig. 4) to Fig. S10 and merged the corresponding results section with the previous one titled: “Formation of a consolidated A compartment in S-phase”.

      We have also revised the language to avoid overstatement. The original conclusion read: “Together, these findings strongly indicate that compartment maturation and the accompanying A compartment consolidation represent a robust and universally observed feature across different developmental contexts.”

      This has been changed to: “Together, these findings support the notion that compartment maturation and the accompanying A-compartment consolidation are not unique to mESCs and may represent a broadly conserved feature of mammalian chromatin organization.”

      Similarly, the abstract has been adjusted from: “Moreover, compartment maturation was not limited to mESCs but was also observed across different developmental contexts in mice.” to: “Moreover, compartment maturation was not limited to mESCs but was also evident across different developmental contexts in mice.”

      These changes frame the results as evidence for conservation in principle rather than definitive proof of identical dynamics across tissues.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please address the minor points in the public review.

      In addition, on page 7, line 285: "In contrast, interactions showed minimal change across all distances though interphase". Do the authors mean "In contrast, B-B interactions..."?

      We thank the reviewer for catching this. The sentence has been corrected.

    1. eLife Assessment

      This important study identifies PRRT2 as an auxiliary regulator of Nav channel slow inactivation in vitro and in vivo, proposing that PRRT2 facilitates entry into, and delays recovery from, the slow-inactivated state. The revised manuscript has been substantially strengthened, providing compelling evidence that PRRT2 is relevant to normal brain physiology and disease pathophysiology, providing a mechanistic link between PRRT2 mutations and episodic neurological phenotypes. Overall, this study will be of interest to ion channel biophysicists and neurophysiologists, particularly those studying channelopathies.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Lu and colleagues demonstrate convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      Comments on revised version.

      The manuscript by Lu and colleagues has been revised sufficiently to address all my prior concerns.

      Experiments are conducted rigorously including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

    3. Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".<br /> PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last (20th) compound APs in panels B and C.

    4. Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      (4) The mechanistic separation between trafficking of PRRT2 and its gating effects is not clearly resolved.

      (5) Additional studies with Nav1.6 should be carried out.

      Comments on revised version.

      These comments have been addressed in the revised version.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Lu and colleagues demonstrates convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously, including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      We thank the reviewer for these positive comments and for the thoughtful evaluation of our work.

      Weaknesses:

      There are a few missing experiments and one place where data are over-interpreted.

      (1) An in vitro study of Nav1.6 is conspicuously absent. In addition to being a major brain Na channel, Nav1.6 is predominant in cerebellar Purkinje neurons, which the authors note lack PRRT2 expression. They speculate that the absence of PRRT2 in these neurons facilitates the high firing rate. This hypothesis would be strengthened if PRRT2 also enhanced slow inactivation of Nav1.6. If a stable Nav1.6 cell were not available, then simple transient co-transfection experiments would suffice.

      We thank the reviewer for raising this point. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms.

      We have now performed new heterologous expression experiments to test whether PRRT2 modulates Nav1.6 slow inactivation. Consistent with our findings for other Nav isoforms, PRRT2 significantly enhances the slow inactivation of Nav1.6. We have incorporated these data into the revised Results and Figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) To further demonstrate the physiological impact of enhanced slow inactivation, the authors should consider a simple experiment in the stable cell line experiments (Figure 1) to test pulse frequency dependence of peak Na current. One would predict that PRRT2 expression will potentiate 'run down' of the channels, and this finding would be complementary to the biophysical data.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we performed a pulse-train protocol in the stable Nav1.2 cell line and quantified the use-dependent attenuation (“run-down”) of peak sodium current across successive depolarizations (Figure 1-figure supplement 1C). Compared with control cells, PRRT2-expressing cells exhibited a larger decline in peak current during trains, indicating greater reduction in channel availability during repetitive depolarizations (Figure 1-figure supplement 1C). This pattern is consistent with our observations above showing that PRRT2 enhances Nav channel slow inactivation. These new data have been incorporated into the revised manuscript. Please refer to Page 5, Lines 133-140; Figure 1-figure supplement 1C.

      (3) The study of one K channel is limited, and the conclusion from these experiments represents an over-interpretation. I suggest removing these data unless many more K channels (ideally with measurable proxies for slow inactivation) were tested. These data do not contribute much to the story.

      We agree with the reviewer’s assessment. To avoid over-interpretation and to maintain focus on PRRT2-dependent regulation of Nav channel slow inactivation, we have removed the potassium channel dataset and the associated conclusions from the revised manuscript.

      (4) In Figure 2, the authors should confirm that protein is indeed expressed in cells expressing each truncated PRRT2 construct. Absent expression should be ruled out as an explanation for the enhancement of slow inactivation.

      We thank the reviewer’s concern regarding expression of the truncated PRRT2 constructs in the Nav1.2 stable cell line, particularly PRRT2(1-266), which shows little effect on slow inactivation of Nav1.2 channels. In the revised manuscript, we conducted western blot to verify expression of the PRRT2(1-266)-HA construct in the Nav1.2 stable cell line. We have added these results to the revised manuscript, please refer to Page 6, Lines 171-173; Figure 2-figure supplement 1A and B.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is primarily expressed in the nervous system and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 directly interacts with Nav1.2 and Nav1.6, modulating channel properties and neuronal excitability.

      In this study, Lu et al. reported that PRRT2 is a physiological regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect can be replicated by the C-terminal region (256-346) of PRRT2, and is highly conserved across species from zebrafish, mouse, to human PRRT2. TRARG1 and TMEM233, the other two DspB family members, showed similar effects on Nav1.2 slow inactivation. Co-IP data confirms the interaction between Nav channels and PRRT2. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared to WT mice.

      Strengths:

      (1) This study is well designed, and data support the conclusion that PRRT2 is a potent regulator of slow inactivation of Nav channels.

      (2) This study reveals similar effects on Nav1.2 slow inactivation by PRRT2, TMEM233, and TRARG1, indicating a common regulation of Nav channels by DspB family members (Supplemental Figure 2). A recent study has shown that TMEM233 is essential for ExTxA (a plant toxin)-mediated inhibition on fast inactivation of Nav channels; and PRRT2 and TRARG1 could replicate this effect (Jami S, et al. Nat Commun 2023). It is possible that all three DspB members regulate Nav channel properties through the same mechanism, and exploring molecules that target PRRT2/TRARG1/TMEM233 might be a novel strategy for developing new treatments of DspB-related neurological diseases.

      We thank the reviewer for careful evaluation and insightful suggestions.

      Weaknesses:

      (1) Previously, the authors have reported that PRRT2 reduces Nav1.2 current density and alters biophysical properties of both Nav1.2 and Nav1.6 channels, including enhanced steady-state inactivation, slower recovery, and stronger use-dependent inhibition (Lu B, et al. Cell Rep 2021, Fig 3 & S5). All those changes are expected to alter neuronal excitability and should be discussed.

      We thank the reviewer for this suggestion. Although the present study focuses on PRRT2-dependent regulation of slow inactivation, we agree that PRRT2 may influence excitability through additional Nav-dependent mechanisms, including reduced current density and shifts in the voltage dependence of channel inactivation (Fruscione et al., 2018; Lu et al., 2021; Valente et al., 2023). Notably, because PRRT2 facilitates entry of Nav channels into slow-inactivated states both from closed states and from open states during prolonged depolarization, some of these previously reported effects may partly reflect enhanced slow inactivation and the resulting reduction in Nav channel availability. We have expanded the Discussion to integrate these prior findings and to clarify that these additional PRRT2-dependent effects may converge to shape neuronal excitability. Please refer to Page 16, Lines 445-452.

      (2) In this study, the fast inactivation kinetics was examined by a single stimulus at 0 mV, which may not be sufficient for the conclusion. Inactivation kinetics at more voltage potentials should be added.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we expanded our analysis of Nav1.2 fast-inactivation kinetics to include a range of test potentials (-20, -10, 0, +10, +20 and +30 mV) in the presence and absence of PRRT2. These experiments showed that PRRT2 expression did not significantly affect Nav1.2 fast-inactivation kinetics under these conditions. We have incorporated these new results into the revised manuscript. Please refer to Page 4, Lines 100-103; Figure 1C.

      (3) It is a little surprising that there is no difference in Nav1.2 current density in axon-blebs between WT and Prrt2-mutant mice (Figure 7B). PRRT2 significantly shifts steady-state slow inactivation curve to hyperpolarizing direction, at -70 mV, nearly 70% of Nav1.2 channels are inactivated by slow inactivation in cells expressing PRRT2 when compared to less than 10% in cells expressing GFP (Figure supplement 1B); with a holding potential of -70 mV, I would expect that most of Nav channels are inactivated in axon-blebs from WT mice but not in axon-blebs from Prrt2-mutant mice, and therefore sodium current density should be different in Figure 7B, which was not. Any explanation?

      We thank the reviewer for raising this point. In our axonal bleb recordings, although the holding potential was -70 mV, sodium current density was measured after a hyperpolarizing pre-pulse to -110 mV, which was applied before the test depolarization to relieve inactivation as much as possible (as described in the Methods). Therefore, the current density measurement in Figure 7B reflects the available current after this recovery step, rather than the steady-state availability at -70 mV. The lack of a difference in Figure 7B does not contradict the PRRT2-dependent shift in steady-state slow inactivation. In the revised manuscript, we have clarified this point explicitly in the Results and figure legend to avoid confusion. Please refer to Page 10, Lines 294-295.

      (4) Besides Nav channels, PRRT2 has been shown to act on Cav2.1 channels as well as molecules involved in neurotransmitter release, which may also contribute to abnormal neuronal activity in Prrt2-mutant mice. These should be mentioned when discussing PRRT2's role in neuronal resilience.

      We thank the reviewer for this suggestion. In addition to the Nav-dependent mechanisms, previous studies have shown that PRRT2 also regulates synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018) and presynaptic surface expression of Cav2.1 channels (Ferrante et al., 2021). These effects are also expected to influence neurotransmitter release and, consequently, neuronal and network excitability. In the revised manuscript, we have expanded the Discussion to acknowledge that these additional PRRT2-dependent mechanisms may also contribute to cortical resilience. Please refer to Page 16, Lines 452-457.

      Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      We thank the reviewer for this positive evaluation of our work and for the constructive comments.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav channel interface—through approaches such as targeted mutagenesis, crosslinking, and structural determination—will be required to define the binding interface and establish the molecular basis of gating modulation. Please refer to Page 16, Lines 465-468.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      We agree with the reviewer. Impaired slow inactivation in Prrt2-mutant mice is one plausible contributor to reduced cortical resilience. PRRT2 has also been reported to regulate surface exposure of Nav and Cav2.1 channels (Ferrante et al., 2021), as well as neuronal synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018). Each of these PRRT2-associated processes could influence cortical excitability in vivo. We have therefore expanded the Discussion to clarify that the cortical phenotype likely reflects the combined contribution of multiple PRRT2-dependent mechanisms, rather than an isolated defect in slow inactivation alone. Please refer to Page 16, Lines 446-458.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      We thank the review for this comment regarding physiological relevance. In the revised manuscript, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Please refer to Page 14 and 15, Lines 414-416; Lines 429-430.

      (4) The mechanistic separation between the trafficking effect of PRRT2 and its gating effects is not clearly resolved.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. Previous studies in heterologous overexpression systems have shown that PRRT2 can influence Nav channel trafficking and surface expression, raising the possibility that the observed effects on slow inactivation regulation might be secondary to altered channel abundance or localization. However, slow inactivation develops on a timescale of tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer intervals (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). These distinct temporal profiles argue against trafficking as the primary basis for the effects of PRRT2 on Nav channel slow inactivation described here, although direct quantification of dynamic changes in Nav channel surface expression will be required to fully exclude such a contribution (Liu et al., 2022; Tyagi et al., 2025). We have incorporated this point into the Discussion section. Please refer to Pages 13, Lines 378-388.

      (5) Additional studies with Nav1.6 should be carried out.

      We thank the reviewer for this suggestion. We have performed experiments to directly examine the effects of PRRT2 on Nav1.6 slow inactivation and incorporated these new data into the revised Results and figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for future experiments (not for this paper)

      (1) Exploit the lower protein expression in V5-PRRT2 mice to examine the effects of a hypomorphic allele.

      We thank the reviewer for this insightful suggestion. We note that the V5 epitope knock-in reduced PRRT2 protein expression, which may functionally resemble a hypomorphic allele. Accordingly, in addition to its utility for biochemical experiments (e.g., co-immunoprecipitation), this line could serve as a genetic tool to interrogate PRRT2 dose-dependent effects in vivo. We have added this point to the revised manuscript, please refer to Page 9, Lines 265-267.

      (2) Examine disease-causing PRRT2 mutations.

      We thank the reviewer for this constructive suggestion. Testing disease-associated PRRT2 variants for their ability to regulate Nav channel slow inactivation would be an important next step to strengthen the disease relevance of the mechanism proposed here. Moreover, identifying missense variants that selectively disrupt slow-inactivation regulation could help pinpoint residues that are critical for PRRT2-Nav functional coupling and thereby inform future structure-function studies. We plan to pursue this direction in follow-up work.

      (3) Investigate spreading depolarization in PRRT2-deficient mice.

      We thank the reviewer for this suggestion. Although we have shown that PRRT2 deficiency facilitates spreading depolarization in the cerebellum, whether PRRT2 exerts similar control over spreading depolarization susceptibility in the cerebral cortex remains to be determined. We plan to address this in an independent study and to test how cortical spreading depolarization relates to other PRRT2-associated neurological disorders.

      Reviewer #2 (Recommendations for the authors):

      This study is, in general, well executed, and the manuscript is well written. However, I do have some questions.

      (1) The authors' previous works have shown that PRRT2 regulates both Nav1.2 and Nav1.6, considering the wide expression Nav1.6 in CNS and its role in neuronal activity, what makes the authors not include Nav1.6 in this study?

      We thank the reviewer for raising this question. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms. In response to reviewers’ concern, we have now performed new experiments to directly examine the effect of PRRT2 on Nav1.6 slow inactivation. These results have been incorporated into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) Please explain why you chose 0 mV rather than -70 mV (closer to membrane potential) in the slow inactivation protocol.

      We thank the reviewer for raising this question. Nav channels can enter into slow inactivation from both resting/closed states and activated/open states. In our steady-state slow-inactivation assays, we found that PRRT2 enhances Nav1.2 slow inactivation under both conditions (Figure 1-figure supplement 1A and B). In whole-cell recordings, Nav1.2 channels typically begin to activate at command voltages more depolarized than approximately -60 mV. Accordingly, a conditioning voltage of -70 mV predominantly probes entry into slow inactivation from closed states, whereas 0 mV drives channel activation and more effectively induces slow inactivation. We therefore chose 0 mV as the primary conditioning potential because it is widely used in conventional slow inactivation protocols and induces slow inactivation more robustly than conditioning voltages at -70 mV. We have added this explanation in Methods section of revised manuscript, please refer to Page 20, Lines 569-571.

      (3) The authors mentioned that the insertion of V5 markedly reduced the PRRT2 protein level; thus, Prrt2-V5 knock-in mice could be considered as PRRT2 knock-down mice. Is there any noticeable difference in phenotype between Prrt2-V5 knock-in mouse and Prrt2-mutant mouse? In other words, is PRRT2 knockdown sufficient to affect neuronal excitability, or is a complete PRRT2 ablation required?

      We thank the reviewer for raising this concern regarding the functional consequences of reduced PRRT2 expression in the Prrt2-V5 knock-in mice. Given that PRRT2 protein levels are markedly reduced in this line, and that cerebellar stimulation-induced dystonia is a characteristic phenotype of PRRT2 deficiency, we tested whether Prrt2-V5 knock-in mice also exhibit this phenotype. We found that electrical stimulation of the cerebellar cortex induced dystonia-like attacks in a subset of Prrt2-V5 knock-in mice. These dystonic behaviors resembled those previously observed in Prrt2-mutant mice, whereas no such behaviors were induced in wild-type mice (Figure 6-figure supplement 1). These findings indicate that a substantial reduction of PRRT2 expression (approximately 80%) is sufficient to impair neuronal function and elicit a disease-relevant phenotype in a subset of animals, supporting the interpretation that the V5 knock-in allele is hypomorphic. We have incorporated these results into the revised manuscript, please refer to Page 9, Lines 265-267; Figure 6-figure supplement 1.

      (4) In Discussion (Page 13, lines 358-361), the authors mentioned a putative interaction between PRRT2 and the Nav channel by modeling, while there is no related data. Please either add modeling data or remove those sentences.

      We thank the reviewer for this suggestion. To avoid over-interpretation, we have removed the statements regarding the AlphaFold-based interaction model from the revised manuscript. We agree that the interaction interface remains to be demonstrated experimentally, and we now discuss this point in the Limitations section. Please refer to Page 16, Lines 465-468.

      (5) Typo: Page 14, line 399, "TMEM232" should be "TMEM233".

      We thank the reviewer for pointing out this typo. We have corrected it in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Mechanistic depth: While the functional data show altered slow-inactivation kinetics, the mechanistic explanation remains superficial. The AlphaFold-based prediction of PRRT2 interaction with DIV-S3 is speculative. The authors should clarify their illustrative rather than evidential intent and avoid over-interpretation.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav interface, including targeted mutagenesis, crosslinking, and structural determination, will be necessary to elucidate the molecular basis of this interaction and its effect on channel gating. Please refer to Page 16, Lines 465-468.

      (2) Separation of trafficking vs. gating effects: Previous studies showed PRRT2 influences Nav trafficking and surface expression. Here, surface expression changes are not systematically quantified. Such an analysis would strengthen the argument that gating effects are not secondary to altered channel abundance or localization.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. We agree that direct analysis of Nav channel surface localization during prolonged depolarization and hyperpolarization would provide stronger evidence to distinguish gating effects from trafficking-dependent mechanisms. However, such experiments are technically challenging in this context: conventional surface biotinylation assays do not provide the temporal resolution required for these rapid protocols, and live-cell imaging approaches to monitor dynamic changes in Nav channel surface expression during slow-inactivation paradigms have not yet been established in our laboratory.

      Although PRRT2 has been reported to regulate Nav channel surface expression in heterologous systems, we consider it unlikely that trafficking is the major determinant of the slow-inactivation effects described here. Slow-inactivation develops on a timescale ranging from tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer timescales (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). We have expanded the Discussion in a revised manuscript. Please refer to Pages 13, Lines 378-388.

      (3) Isoform generalization: Data on other Nav channel subtypes are presented as evidence of a conserved mechanism. However, given tissue-specific expression of PRRT2, these findings may be of limited in vivo relevance. At the very least, additional studies with Nav1.6 should be carried out.

      We thank the review for this suggestion. In response, we conducted new experiments to examine the effect of PRRT2 on Nav1.6 slow inactivation. These results show that PRRT2 promotes entry of Nav1.6 channels into slow-inactivated states and delays their recovery, consistent with its effects on the other Nav isoforms examined in this study. We have incorporated these new data into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      Furthermore, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Pages 14 and 15, Lines 414-416 and 429-430.

      (4) In vivo functional link: The EEG after-discharge threshold assay suggests decreased cortical resilience, but causality between slow-inactivation impairment and hyperexcitability remains indirect. Complementary in vivo recordings would strengthen the physiological link.

      We thank the reviewer for this helpful suggestion. To further link impaired slow-inactivation to the hyperexcitability, we applied a repetitive stimulation protocol in corpus callosum slices, a white-matter region of brain enriched in both PRRT2 and Nav channels. During high-frequency stimulation (e.g., 20 Hz), the amplitude of the compound action potential progressively decreased over the course of the stimulus train. This phenomenon, often referred to as adaptation, reflects activity-dependent reduction in Nav channel availability (Fleidervish et al., 1996; Mickus et al., 1999; Kim et al., 2012). Compared with wild-type mice, Prrt2-mutant mice exhibited less adaptation during high-frequency stimulation, consistent with impaired slow inactivation during repetitive activity, which may contribute to hyperexcitability (Figure 7-figure supplement 2). We have added these results to the revised manuscript. Please refer to Pages 11, Lines 311-322; Figure 7-figure supplement 2.

      (5) Structural interaction: It remains unclear whether PRRT2 binds the α-subunit directly or through accessory proteins. Crosslinking or detergent-solubilization controls of different stringencies could clarify this.

      We thank the reviewer for raising this important issue. We agree that our co-immunoprecipitation data do not distinguish whether PRRT2 associates with the Nav channel α-subunit directly or through other components of the protein complex. To avoid over-interpretation, we have revised the relevant text in the manuscript to remove any implication of direct binding and now describe the result as an association between PRRT2 and Nav channels.

      We have also expanded the Limitations section to note that additional experiments, such as crosslinking and structural studies, will be required to define the interaction interface between PRRT2 and Nav channels. Please refer to Page 16, Lines 465-468.

      (6) Comparisons to other regulators: The paper positions PRRT2 as distinct from FHFs and β-subunits. The data support this, but the discussion could more critically assess whether PRRT2 acts by stabilizing a pore-based inactivated conformation, as suggested for other slow-inactivation modulators.

      We thank the reviewer for this insightful suggestion. At present, relatively few modulators have been characterized in detail with respect to their effects on Nav channel slow-inactivation kinetics. Moreover, even for compounds such as lacosamide, which has been proposed to act as a slow-inactivation modulator, the underlying mechanism remains under debate (Errington et al., 2008; Jo and Bean, 2017). Therefore, in the revised manuscript, we discussed the possible mechanism of PRRT2 in the context of current models of Nav channel slow inactivation.

      Previous studies suggest that entry into the slow-inactivated state involves at least two coupled processes: conformational changes in the voltage-sensing domains and structural rearrangements in the pore region, including the selectivity filter and intracellular activation gate (Catterall et al., 2024; Silva, 2014). During prolonged depolarization, voltage sensors become stabilized in the up-state, while the pore undergoes progressive rearrangements associated with slow inactivation (Balser et al., 1996; Vilin et al., 1999). Thus, mechanisms that further stabilize voltage sensors in the up-state and/or facilitate pore-based inactivated conformations could enhance slow inactivation.

      Within this framework, PRRT2 may enhance slow inactivation by facilitating one or both of these processes, although direct evidence is still lacking. We have incorporated this discussion in relative section of revised manuscript. Please refer to Page 14, Lines 389-404.

      Response references:

      Jo S, Bean BP. Lacosamide Inhibition of Nav1.7 Voltage-Gated Sodium Channels: Slow Binding to Fast-Inactivated States. Mol Pharmacol. 2017 Apr;91(4):277-286.

      Errington AC, Stöhr T, Heers C, Lees G. The investigational anticonvulsant lacosamide selectively enhances slow inactivation of voltage-gated sodium channels. Mol Pharmacol. 2008 Jan;73(1):157-69.

      (7) Behavioral/clinical link: Given the strong human genetics background of PRRT2 disorders, a brief analysis or reference to electrophysiological phenotypes in patient neurons would contextualize the cortical findings.

      We thank the reviewer for this suggestion. Previous studies showed that iPSC-derived excitatory neurons from a patient carrying a homozygous PRRT2 mutation exhibited increased sodium currents and neuronal hyperexcitability (Fruscione et al., 2018). Given that slow inactivation regulates Nav channel availability and thereby influences neuronal excitability, these electrophysiological abnormalities in patient-derived neurons may, at least in part, reflect impaired PRRT2-dependent regulation of Nav channel slow inactivation. We have added this point to the relative section of the revised manuscript. Please refer to Pages 15, Lines 432-437.

      Minor comments

      (1) Figures should include statistical sample sizes (n) and ideally overlay data points rather than only means {plus minus} SEM.

      We thank the reviewer for this suggestion. In the revised manuscript, we present both individual data points and mean ± SEM in the column graphs. For the line graphs, individual data points were not overlaid because of space and readability constraints, and these panels therefore display mean ± SEM only. Sample sizes for each group are provided in the corresponding figure legends.

      (2) The AlphaFold model should be provided as a supplementary figure with confidence scores indicated.

      We thank the reviewer for this suggestion. However, because the predicted Nav1.2-PRRT2 interaction interface has not yet been experimentally validated in our study, we chose to remove the AlphaFold-based model from the revised manuscript to avoid over-interpretation.

      (3) Clarify whether TTX sensitivity was verified in the axonal bleb preparation.

      We thank the reviewer for raising this point. We verified the identity of the sodium currents in the axonal bleb preparation by their sensitivity to TTX, and this information has now been added to Figure 7A in the revised manuscript. Please refer to Page 10, Line 290; Figure 7A.

    1. eLife Assessment

      This important study investigates how surface stickiness shapes whisker mechanics and peripheral neural responses during active touch. The biomechanical evidence that surface stickiness alters whisker mechanics and stick-slip dynamics is compelling, supported by a large and high-quality 3D dataset, while the electrophysiological evidence is solid but limited by a small sample size and insufficient validation of the sticky stimuli. The work will be of broad interest to sensory neuroscientists studying active touch.

    2. Reviewer #1 (Public review):

      Summary:

      This study offers a careful and technically strong look at how surface stickiness changes whisker-surface interactions and how that information reaches peripheral sensory neurons. The authors use 3D whisker tracking to capture bending, twisting, rolling, and tip motion during contact with surfaces that differ in stickiness, coarseness, and position. They show that sticky surfaces, especially silicone, broaden the range of whisker deformation, produce stronger but less frequent stick-slip events, and change firing rates in some trigeminal ganglion neurons. Overall, the study is valuable because it goes beyond standard 2D tracking and shows that out-of-plane motion and roll are important for understanding how whiskers encode texture.

      Strengths:

      The study is technically strong and well motivated. Its main strength is the use of 3D whisker tracking to show that surface stickiness affects whisker deformation in ways that standard 2D tracking would miss, including torsion, roll, out-of-plane motion, and stick-slip dynamics. The authors also connect these mechanical effects to TG activity, providing evidence that stickiness information is available in peripheral sensory responses. Overall, the work expands the study of whisker-based texture sensing beyond coarseness and provides a richer biomechanical framework for understanding tactile encoding.

      Weaknesses:

      The main weakness is that stickiness is not formally defined early in the manuscript, even though it is the central experimental variable. Several methodological choices also need clearer justification or validation, including the use of 2D measures as comparators for torsion and roll, the thresholds used for stick-slip detection, the degree-5 polynomial fit, the reference ROI, and aspects of the 3D surface reconstruction. The neural evidence should also be interpreted cautiously because the TG sample is small, only a subset of units discriminated silicone, and the correlation between strain sensitivity and silicone discrimination is suggestive rather than definitive.

    3. Reviewer #2 (Public review):

      The authors explore the sensation of stickiness from the point of view of whisker exploration and encoding in the trigeminal ganglion. In doing so, they develop methods for 3D whisker tracking to describe stick-specific parameters such as stick-slip rates and strain. Overall, the methods are strong, and the authors present the results appropriately. Overall, I think exploration of the sensation of stickiness is a great question.

      My main criticism is in relation to the chosen stimuli, and I wonder whether the authors may have room to explore more naturally sticky materials and what this may mean for the animal.

      (1) Chosen stimuli for stickiness:

      Four different materials are used, with the aim of presenting animals with graded measures of stickiness. The results show that silicone stands out against the others; it's less clear whether the intermediate textures (Delrin and resin) may be truly intermediate in stickiness.

      I wonder if the stimuli chosen were truly representative of the aim of providing a gradient of stickiness. Did the materials differ in other features, such as surface temperature, texture, etc., which could explain some results? The authors discuss this in terms of coefficients of friction and how these estimates are not quantified in relation to whiskers themselves.

      Measures of stick-slip and strain with silicone vs other materials make intuitive sense. Could the authors add additional naturally sticky stimuli to exemplify the results? For example, adhesive, glue, or a sugary substance.

      (2) Tracking methods and quantification:

      The 3D tracking methods, which incorporate whisker twists, strain, and other fine features of whisker exploration, present an advance in terms of analysis of how whiskers may explore more complex, natural features of environments. The analyses and quantifications are all solid and robust. The technical approaches are well-prepared to take the work a step further in terms of stimulus choice.

      (3) Peripheral coding of stickiness:

      The authors report that some units respond preferentially to whisking on silicone and that this has to do with strain on the whisker. Is there a possibility to understand the nature or anatomy of these units and why they might be preferential for the sticky sensation? Can the location in the follicle be assigned? And/or would the methodology allow for assignment of where the specifically sticky-tuned units project centrally?

      (4) Relationship to natural stimuli:

      A piece missing from the paper is more discussion and exploration of why stickiness may be important for sensory coding, as well as potentially more naturally sticky stimuli. One could imagine that a mouse navigating the world could find stickiness attractive, if it were a source of sweet food, for example, or it could potentially be a sensation the animal prefers to avoid. Stickiness could also indicate contamination or a sticky trap, to be avoided. If the authors are able to add naturally sticky stimuli, the whisker exploration and encoding could potentially provide further cues towards the valence of stickiness for mice.

    4. Reviewer #3 (Public review):

      This paper tackles an underexplored dimension of whisker-based texture sensing: while surface coarseness encoding has been extensively characterized in rodents, the mechanical and neural basis for stickiness sensing has not previously been examined. The authors make two intertwined contributions that together represent a substantial advance: a methodological one - a 3D whisker tracking pipeline operating at 4000 fps, capable of capturing torsion, roll, and out-of-plane whisker motion - and a scientific one - a first characterization of how whisker mechanics and primary trigeminal afferent responses differ between surfaces of high and low stickiness. The work is technically solid, the dataset is large, and the question is well motivated both by the multidimensional nature of tactile texture perception and by the practical advantages of the whisker system for studying touch mechanics.

      Strengths.:

      The 3D tracking system is a timely advance over existing tools, particularly in its handling of non-planar whisker shapes and the full automation required for the sub-millisecond resolution needed to detect stick-slip events. The mechanical dataset is extensive. The finding that whisking against silicone expands the sampled whisker strain space and produces stronger but less frequent stick-slip events is clearly demonstrated and internally consistent with the proposed mechanism of greater strain accumulation before frictional release - a physically intuitive result. The open release of the tracking code considerably increases the value of this work to the broader community.

      Weaknesses:

      A few aspects of the paper, if sharpened, would considerably strengthen the evidence and the clarity of the conclusions.

      The central claim - that "stickiness information is available to the whisker system" - does not capture the precision of what the paper demonstrates. As stated, the finding is close to guaranteed: any variation in surface friction will produce some change in whisker mechanics, so the presence of mechanical differences between materials is expected rather than surprising. The more valuable question the paper is well positioned to answer is which specific dimensions of the whisker mechanical response are most informative about surface stickiness. The paper reports effects on strain distribution breadth, stick-slip amplitude, and stick-slip rate, but does not synthesize which of these - or which sub-dimensions (bending, twisting, or rolling) - carry the most discriminating information. Identifying the salient dimensions of the mechanical response and relating them to the proposed frictional mechanism would sharpen the paper's conclusions substantially.

      A related but distinct limitation is the absence of direct force measurements during whisker-surface contact. The authors acknowledge this openly, and I recognize it is not easily remedied within the current experimental setup. It does, however, constrain interpretation: without knowing the actual forces generated at the whisker-surface interface, the assumed stickiness ordering of the tested materials cannot be validated, and - importantly - the relative contribution of surface friction and material compliance to the observed mechanical differences cannot be determined. This is an important direction for future work in this area.

      The paper argues carefully that 2D tracking is insufficient for capturing the full mechanical picture of whisker-surface interactions, and the figure currently in the supplementary material (Figure S2) makes this case convincingly through multiple analyses. This argument is the core justification for the paper's methodological contribution and deserves a place in the main manuscript. Furthermore, while the mechanical case for 3D over 2D tracking is well made, it has not yet been tested at the neural level: the regression model used to predict neural firing incorporates 3D variables, but its performance is not compared against an equivalent model restricted to 2D variables. Such a comparison would directly demonstrate whether torsion and roll - the signals inaccessible to 2D tracking - carry neural predictive value, and would elegantly unite the paper's methodological and scientific contributions.

      Finally, the three-dimensional plots in Figure 3 are the paper's primary representation of its main mechanical result, and there is a real opportunity to make them considerably more informative. The whisker deformation probability distributions (panel B) are rendered in 3D from a single viewing angle, making it difficult to assess the shape or anisotropy of the distributions - and in particular to see which dimensions expand most for silicone relative to the other materials. This is precisely the information needed to identify the most salient dimensions of the stickiness signal, and two-dimensional representations would make it directly readable.

    1. eLife Assessment

      This study presents a useful compendium of triangulated single-cell eQTLs, Mendelian randomisation and colocalization of genetic signals in prostate cancer datasets. Biological interpretation in the context of the aging prostate gland, the tumour microenvironment and immune cell specificity is incomplete, so this study is a starting point for further study, and would require validation of the resulting putative causal genes.

    2. Reviewer #1 (Public review):

      Summary:

      Using Mendelian randomisation on available GWAS data, the investigators identified eGenes associated with prostate cancer and applied the data to define relevant immune cell types involved. Additional analysis was performed to explore potential candidate targets and agents from licensed medicines.

      This is an interesting approach as the investigators have expertise in other research fields, applied here to prostate cancers. The use of three different datasets is significant, and the approach to further analyse implicated eGenes in drug target analysis is relevant and timely.

      A particular strength is taking putative genes from Mendelian randomisation analysis to target and potential drug agents.

      Some aspects of the study would need to be clarified to enable interpretation of the findings in the context of the prostate gland and prostate cancers: expanding the descriptions of the supporting Supplementary Data and Tables, explanations of the analysis for the general reader, and clarification of the selection of eGenes (Figure 5).

    3. Reviewer #2 (Public review):

      Summary:

      This study integrates bulk and single-cell transcriptomic-derived eQTLs from two separate consortia (PRACTICAL and Finngen) to identify immune-cell-specific therapeutic targets in prostate cancer. Mendelian randomization and Bayesian colocalization have been used to produce druggable eGene modules through STRING and DrugBank.

      This is an interesting study that is attempting to address risk-associated, immune-specific transcriptomic repertoires in prostate cancer. It is knitting together concepts of drug repurposing and prostate cancer immunogenicity. This is an entirely computational study, which would benefit from some wet lab experimental validation.

      It is very tricky to attribute cell-type-specific responses, especially when the majority of genes involved represent cytoskeletal or stress responses, which are ubiquitous throughout the prostate microenvironment. This point is relevant for the drug repurposing section: if these drugs are targeting immune cell-specific repertoires, what would the response be of the entire environment? It would be useful to contextualize the validity of each proposed therapy in a specific prostate cancer context and the involvement of AR antagonism or radiotherapy.

      Strengths and limitations of this study:

      Strengths:

      This is a scientifically interesting and potentially impactful study, particularly in its attempt to integrate immune-cell-specific transcriptomics, causal inference, and drug repurposing in prostate cancer. The methodology is well described, and the data (albeit limited) are well analyzed.

      Limitations:

      The central weakness is the overstatement of the conclusions regarding immune-cell-specific causality, without sufficiently contextualizing the biological meaning of the findings.

      Highlighted genes, such as LMNA, XBP1, histone-related genes, and stress-response markers, are ubiquitous regulators involved in fundamental cellular processes, including ageing, unfolded protein response (UPR), integrated stress response (ISR), chromatin remodeling, proliferation, and metabolism. It is unclear whether these signatures truly represent immune mechanisms, or instead reflect broader inflammatory and age-associated biology expected within an ageing glandular organ such as the prostate.

      Immune cell identity alone may not be sufficient to infer biological relevance because immune state characterization (e.g., exhausted versus functional T cells, or distinct macrophage/myeloid phenotypes) is largely absent from the current analysis. The assertion that specific immune populations are correlated with prostate cancer susceptibility is probably an overstatement unless the nature of these cells can also be characterized.

      The interpretation of "causal variants" is not always specified, i.e., what phenotype is being associated: prostate cancer susceptibility, recurrence, progression, or treatment response (e.g. is there direct causality from immune-cell variants to prostate cancer?).

      Overall, there is a need for stronger biological and translational contextualization: how do the identified pathways relate to ageing-associated inflammation, PIN, microbiome-driven inflammatory changes, and stress-response biology in the prostate gland? While the manuscript identifies network hubs and enriched pathways, it often stops short of explaining what these modules biologically represent or how they may influence prostate cancer development, progression, treatment resistance, or immune evasion.

      There are additional publicly available spatial transcriptomic or single-cell datasets which could be used to validate whether the purported immune-cell-specific genes are genuinely enriched in immune populations adjacent to tumour cells. In the drug repurposing analyses, the current study does not explicitly handle prostate cancer subtypes such as HSPC, CRPC, NEPC, or DNPC and co-treatment with androgen receptor antagonism or radiotherapy.

    1. eLife Assessment

      This useful study explores how macrophage cell-cycle state may influence endocytosis, Mycobacterium tuberculosis uptake, and the intracellular stress experienced by bacteria. While the question is interesting and the experimental approach has promise, the evidence for the central claim that endocytic capacity is specifically regulated by cell-cycle stage is incomplete. The main concern is that fluorescence-based sorting and total-fluorescence measurements likely covary with cell size, so the reported phenotypes could reflect biomass accumulation or other cell-cycle-associated changes rather than endocytic capacity as the causal determinant. As a result, whilst the study raises a hypothesis that is of importance, additional controls are required before the proposed mechanism can be considered well supported.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript- "Cell cycle-dependent variation in endocytosis drives phenotypic diversity in M. tuberculosis" by Subhash et al. demonstrates how host cell heterogeneity shapes intracellular pathogen phenotypes. The central and novel finding of this study (G2-phase cells have higher endocytic capacity and harbour more oxidised Mtb) highlights that a host cell cycle (interphase-driven) changes in endocytic capacity regulate bacterial redox states.

      Strengths:

      Overall, the study is well-executed and conceptually rich, establishing a causal link between host cell cycle progression, endocytic heterogeneity, and M. tuberculosis phenotypic diversity.

      The combination of multiple modalities, including live-cell imaging, flow cytometry, scRNA-seq, and redox-sensitive bacterial reporters, supports these findings and substantially strengthens the biological relevance of the work.

      The writing is generally clear, and the figures are well-organised.

      This work will be of interest to readers across cell biology, microbiology, and infection biology

      Weaknesses:

      However, several central claims are only partially supported, the mechanistic depth is limited, and several experimental and analytical concerns need to be addressed.

      Major Comments:

      (1) The authors demonstrate a correlation between the G2 phase and elevated endocytic capacity. However, the mechanistic link (upstream molecular mechanism) between the cell cycle and endocytic upregulation remains largely unaddressed. The authors speculate that membrane biogenesis during volumetric expansion may drive increased endocytosis and note that lipid biosynthesis genes are upregulated in high-endocytic cells. It would substantially strengthen the paper to test this directly, by examining whether inhibition of lipid biosynthesis (e.g., with fatostatin or cerulenin) selectively reduces the G2-associated increase in endocytic capacity. Alternatively, cyclin-CDK axis perturbations (e.g., CDK1 inhibition with RO-3306 to specifically block G2/M entry) could be used to ask whether cells arrested in G2 maintain elevated endocytosis, helping distinguish cell-cycle-position-dependent from cell-cycle-progression-dependent effects.

      (2) The current data show a clear association between high endocytic capacity and more oxidised Mtb, and the authors (consistent with their prior work) hint at lysosomal delivery as the likely mechanism. However, direct evidence for this in the current paper is limited. An experiment examining phagosomal pH or lysosomal fusion (e.g., using a pH-sensitive reporter or lysotracker) specifically in high- and low-endocytic-capacity cells after infection would help confirm this.

      (3) Temporal resolution of Mtb redox dynamics. The plasticity experiment (Figure 6C-D) is elegant and shows that Mtb redox states revert as host cells divide and daughters enter G1. However, the experiment compares day 0 and day 3 post-sorting, which spans multiple cell divisions. While a finer time resolution (spanning 24h) would establish the causal relationship, the authors could discuss the possibility and consequences of multiple cell divisions between day 0 and day 3 used in the present study.

      (4) Relevance of G2 percentages in differentiated macrophages. In Figure 7 and Supplementary Figure S7, only 4.4-5.7% of THP-1-derived macrophages and 5.7% of BMDMs are in G2. While the authors demonstrate statistically significant differences in Mtb redox states between G1 and G2 macrophages, the biological significance of such a small G2 fraction in a non-dividing population deserves discussion specifically with respect to: a) Are these cells re-entering the cycle? b) Is the G2 designation capturing a distinct functional state rather than active cycling? The authors should include additional markers (e.g., phospho-histone H3 for mitotic cells or BrdU incorporation to test for active S-phase) to characterise this population and clarify its identity and origin in differentiated macrophages, thereby meaningfully informing interpretation.

      In conclusion, this is an important mechanism-driven study that highlights an important link in host-driven bacterial phenotypic heterogeneity. The experiments are thorough, the model is well-supported, and the study has implications for infection biology.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors utilize a combination of techniques to show that macrophage endocytic capacity is partially dictated by cell-cycle stage, that Mycobacterium tuberculosis (Mtb) more readily infects macrophages that are in G2/M -phases, and that bacteria that are internalized by macrophages at different stages of the cell-cycle experience different levels of intracellular stress (as reported by the redox state of the bacteria). Furthermore, the authors provide evidence that terminally differentiated macrophages retain memory of the cell-cycle stage that they were in prior to differentiation, at least in the context of endocytic capacity.

      This work provides evidence for the growing idea that fundamental heterogeneity in both host and bacterial organisms can alter the host-pathogen relationship in important ways. However, based on the current data, I am not convinced that the manuscript establishes endocytic capacity as the causal link between macrophage cell-cycle stage and bacterial state. The main issue is that fluorescence-based sorting for cell-cycle stage is likely to covary with cell size. Larger cells, including those later in the cell cycle, may be more likely to fall into the "high" fluorescence gate, while smaller cells may be enriched in the "low" population. Therefore, the observed phenotypes may still be cell-cycle-associated, but the causal determinant could be a correlated feature of cell-cycle progression rather than endocytic capacity itself. This is a significant caveat because nearly all the data, including the live-cell imaging following individual cells, rely on 'total' fluorescence, which will scale strongly with cell size.

      If the authors' conclusion that endocytic capacity is cell-cycle regulated holds true after appropriate controls, this would significantly advance our understanding of the causal interplay between host cell-cycle state, endocytosis, and Mtb physiology. However, an alternative interpretation is that the observed differences in Mtb uptake and bacterial redox state are associated with cell-cycle stage but are not caused directly by differences in endocytic capacity. For example, they could instead reflect other cell-cycle-linked changes in macrophage physiology, such as cell size, intracellular volume, metabolic state, or some other mechanism important for Mtb pathogenesis. If the authors find that their data are best explained by cell-cycle stage independent of endocytic capacity, this would still represent an important advance. However, in that case, the manuscript should clearly distinguish the association with cell-cycle state from the downstream effector mechanisms, which would remain to be determined.

      Strengths:

      The authors utilize various macrophage models for their studies, which is important considering the variability in macrophage behavior, as well as the growing evidence that differences between mouse and human macrophages are relevant for Mtb infection.

      Weaknesses:

      The most important caveat is the covariance between fluorescence-based reporters and cell size. This concern applies to both the sorting experiments, which directly measure total fluorescence, and the time-lapse microscopy experiments, in which the authors show total fluorescence rather than mean, area-normalized fluorescence in Figure 3C. This could be explained by biomass accumulation alone, rather than by a specific cell-cycle-dependent increase in endocytic capacity. Without distinguishing total signal from concentration or activity per unit cell area/volume, it is difficult to conclude that endocytosis itself is regulated by cell-cycle stage rather than simply scaling with cell size.

      Although the authors provide some evidence that the mean GFP intensity, which more closely reflects concentration, differs between the sorted populations in Figure 3B, they do not report statistics for this comparison. Moreover, this control is not carried through the rest of the manuscript, including in key experiments such as Figure 2B. As a result, it remains difficult to determine whether the observed differences between "high" and "low" populations reflect cell-cycle state specifically or instead reflect differences in total reporter fluorescence driven entirely by cell size.

      The evidence for cell-cycle-dependent effects would be more convincing if the authors included additional controls. For example, they could:

      (1) Plot both mean GFP intensity and total GFP intensity in Figure 3B, ideally alongside an unrelated fluorescent reporter that does not vary across the cell cycle. This would help distinguish changes in reporter concentration from changes driven by cell size or total fluorescence.

      (2) Sort cells based on an unrelated fluorescent marker and test whether the same phenotypes - infectivity, dextran uptake, bacterial redox state, etc. - differ between high- and low-fluorescence populations. If these phenotypes are specific to the cell-cycle reporter and not observed with an unrelated marker, this would strengthen the conclusion that the effects are linked to cell-cycle state rather than to fluorescence intensity, cell size, or sorting artifacts.

    1. eLife Assessment

      This important submission from Ambler and colleagues brings new insights into how torpor conditions may confer resilience in cases of cardioprotection. It has novelty, which can be enhanced by additional in vivo support. The study is backed by solid evidence, and represents a unique interoceptive mechanism of interest.

    2. Reviewer #1 (Public review):

      Summary:

      Torpor can be induced by chemogenetic activation of the medial preoptic area. This activation leads to protection from myocardial infarction in an isolated heart preparation despite normalization of the ambient temperature, thus, in principle, uncoupling hypothermia from torpor-induced neuroprotection. Putative pathways of protection are suggested by proteomic studies.

      Strengths:

      (1) Elegant strategy for inducing torpor in rats.

      (2) Appropriate controls for verifying the neuron transducer.

      (3) Cardiac protection is significant and appears independent of hypothermia.

      (4) Interesting omic strategy to begin to find established and novel pathways mediating organ autonomous torpor-induced protection.

      Weaknesses:

      (1) The study would benefit from using inhibitory chemogenetics of the same neurons to demonstrate that this might make cardiac response to ischemia worse.

      (2) Infecting an area of the brain not known to be involved in torpor would be a useful control.

      (3) In vivo cardio protection seems essential as the validation of the strategy requires support that is in the intact animal.

      (4) The assumption that the positive effects of torpor are mediated via a phosphoproteomic change rather than a translational or transcriptional control mechanism is not established.

      (5) A 40 percent reduction in infarct size may work for genetically identical rats with no co-morbidities, but is unlikely to be significant enough to weather the variability that emerges in humans because of these differences and more. The question is not what the mechanism is, but how do we make it more robust? Overall, this is at best a preliminary data set that requires more experiments to deliver on its immense promise.

    3. Reviewer #2 (Public review):

      Summary:

      Elley and colleagues induced a synthetic torpor-like state in rats (a non-hibernating species) by chemogenetically activating neurons in the medial preoptic area of the hypothalamus. They show that this state substantially reduced cardiac infarct size in an ex vivo ischemia-reperfusion model. They further report that protection persisted when ambient temperature was raised to prevent hypothermia, and used exploratory phosphoproteomics to identify candidate cardioprotective signaling pathways.

      Strengths:

      This is the first demonstration that a torpor-like state is cardioprotective in a species that does not naturally enter torpor, which meaningfully advances the potential clinical utility of synthetic torpor. The experimental design is logical, and the controls are generally appropriate. The characterisation of the responsible neuronal population using ISH against QPLOT markers adds mechanistic depth and supports the cross-species conservation argument. The phosphoproteomic analysis, though exploratory, generates plausible and biologically coherent hypotheses grounded in the hibernation literature.

      Weaknesses:

      The primary weakness is that the central conclusion - that hypothermia is not necessary for cardioprotection - exceeds the evidence. The thermoneutral groups were not demonstrably normothermic (36.4 vs 37.05{degree sign}C, p=0.44 with n=6), core temperature telemetry was absent in the majority of control animals contributing to the infarct endpoint, and the decisive test, i.e., a correlation between individual nadir temperature and infarct size, was never performed. Additional weaknesses include the absence of sex-stratified analysis despite known estrogenic contributions to torpor

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Elley and colleagues describes experiments on the effects of synthetic torpor on ex vivo heart ischemia. The key aspect of the study was the use of viral-vector mediated manipulation of the hypothalamic medial preoptic area (MPA) in rats. They used AAV-CaMKIIa-hM3D(Gq). The authors report that chemogenetic activation of the MPA prior to an ex vivo heart ischemia-reperfusion insult induces cardio protection against infarct size that is independent of prior in vivo hypothermia. Phosphoproteomic analysis of cardiac tissue suggested changes in cell survival and death pathways.

      Strengths:

      This study has important strengths. The idea is novel. The experimental design is appropriately rationalized and fascinating. The manuscript is written and presented concisely.

      Weaknesses:

      The study has important weaknesses in the experimental design and validation of the model.

      (1) The study is based on the use of a DREADD-designed viral vector (AAV-CaMKIIa-hM3D(Gq) -mCherry) that is activated by 2 mg/kg IP injection of CNO. The rationale is to putatively activate the MPA. The authors show no evidence for chemogenetic activation of neurons in the MPA. This could be done using a variety of different approaches, even phosphoproteomics.

      (2) The stereotaxic injections are difficult to precisely and locally place, particularly bilaterally. Figure 2F is only a schematic. It would be better to show actual low magnification brain sections (bregma +0.12 to -0.48) from a representative rat to show the placement of the AAV.

      (3) The control rats were injected with AAV-CaMKIIa-EGFP. Why was EGFP used instead of mCherry for the control?

      (4) Ideally, a mutant non-activatable variant of AAV-CaMKIIa-hM3D(Gq) should have been used for a better control.

      (5) The authors should comment on whether there is any neurotoxicity in the MPA associated with the forced AAV expression of hM3D-Gq.

      (6) Is there any inflammatory pathology seen in the MPA with AAV transduction?

      (7) There are no experiments to show that the systemic torpor is specifically associated with the MPA region. Experiments should be done with injections of AAV-CaMKIIa-hM3D(Gq)-mCherry placed in other brain regions, for example, the nearby nucleus accumbens.

      (8) The mapping of the distribution of neurons responsible for synthetic torpor is not mechanistic enough and is not directly to the point. While excitatory and inhibitory markers are examined, a more interesting and deeper approach would have been to use glutamate receptor antagonists to manipulate the torpor response.

      (9) The ischemia and reperfusion aspects of the Lagendorff method need to be clarified. The isolated hearts are already ischemic after their removal from the rat. The reperfusion aspect is caused by reflow of blood to generate oxidative stress, but in the ex vivo model, is there really reperfusion injury?

      (10) The authors show that whole animal oxygen consumption is reduced in the torpor state. The measurement is crude and most likely reflects the inactivity of the animal's skeletal muscle in the torpor state. A more relevant and direct experiment would be to do oxygen consumption (or Seahorse) assays on extracts of the isolated hearts.

      (11) The authors report that the synthetic torpor induces bradycardia. There is no follow-up on this important observation. The MPA-heart connection is not analyzed. (A) Is the link through cardiovascular centers in the brainstem? (B) Is the torpor-induced bradycardia mediated through increased parasympathetic or decreased sympathetic autonomic tone? Pharmacological experiments could also be done.

    1. eLife Assessment

      This study addresses a recent discovery by others that electroconvulsive therapy (ECT) generates seizure activity and spreading depolarization (SD), reflected in large increases in calcium, which can be followed by imaging calcium fluctuations in neurons. This work is useful. However, the evidence to show that SD, rather than seizures, confers the neuroplastic and other therapeutic effects of ECT is incomplete.

    2. Reviewer #1 (Public review):

      The work corroborates the idea, recently suggested by Rosenthal et al. (2025), that spreading depolarization is involved in the mechanisms of electroconvulsive therapy. Using a mouse model of electroconvulsive therapy and various sophisticated approaches to visualize cortical activity, the authors provide an extensive description of traveling calcium waves induced by electroconvulsive stimulation. The study confirms that the calcium events have properties typical of cortical spreading depolarization and seeks to show that the calcium/SD waves mediate therapeutic and neuroplastic effects of electroconvulsive therapy. The authors find that after electroconvulsive stimulation associated with calcium/SD waves, Fos expression increases widely; in the cortex, this increase is localized to the hemisphere affected by calcium waves. They show that some EEG predictors of the beneficial effects of electroconvulsive therapy correlate with the occurrence of calcium/SD waves. Despite the solid methodology and the study's interesting, its conclusions are not fully supported by the data.

      In particular:

      (1)The title of the paper claims that "electroconvulsive stimulation drives cortical spreading depolarization dependent immediate early gene expression". However, immunohistochemical staining shows that Fos expression increases not only in the cortex but also in many subcortical regions, including the hippocampus and amygdala (Figure 5A). Really, conventional electroconvulsive therapy stimulates nearly the entire brain volume and induces generalized seizure activity that can trigger SD not only in the cortex but also in other brain sites. Therefore, regions beyond the cortex can also drive the effects of electroconvulsive therapy. Next, the authors use Fos staining as a marker of neuronal plasticity. However, Fos is also a marker of preceding neuronal activation. As electroconvulsive stimulation, seizures, and SD are associated with high neural activity, it is unclear whether the observed Fos upregulation results from the prior activation or heralds the subsequent plastic changes. Other markers of neuroplasticity (e.g., BDNF) should also be examined.

      (2) Postictal EEG suppression is one of the most promising correlates of positive clinical outcomes after electroconvulsive therapy. Cortical SD is also tightly coupled with suppression of neuronal activity in affected regions. Although the authors report that postictal suppression is stronger after stimulations with cortical SDs than without SDs, the cortices affected (ipsi) and unaffected (contra) by unilateral cortical calcium/SD events exhibit identical suppression (Figure 6F). The result contradicts established knowledge in the field. If the calcium events are cortical SDs, they should induce EEG suppression only in the affected hemisphere.

      (3) The study states a beneficial role of calcium/SD waves in ECS effects. However, SD alters numerous aspects of brain function, leading to a range of effects that can underlie side effects as well. Assessment of the behavioral effects of stimulation with and without calcium/SD waves can help clarify the issue.

      The results of the work suggest that cortical SD can contribute to electroconvulsive therapy-related mechanisms and help to optimize the stimulation parameters to achieve maximal therapeutic effect.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses the question of mechanisms underlying the therapeutic effects of electroconvulsive therapy (ECT). Clinical efficacy of ECT in major depression (and other disorders) is well established and has often been assumed to be a direct consequence of seizure activity generated by the current application. However, as the authors point out, this explanation is unsatisfactory. A recent study (Rosenthal et al., 2025) provided evidence that ECT generates a wave of cortical spreading depolarization (CSD) in mice, and initial evidence that similar events were generated in patients undergoing ECT. Based on their observations, Rosenthal et al. proposed that CSD, rather than seizure, may engage plasticity mechanisms that contribute to the brain's clinical response to ECT. The current study adds to that prior work by reporting other consequences of CSD, in addition to sustained Ca2+ elevations. The current study also links EEG characteristics immediately following the ECT with the likelihood of generating a CSD, which can help optimize ECT parameters.

      Strengths:

      An important research topic, linking a large set of rodent studies with a limited clinical EEG data set.

      The data acquisition and analyses appear to be of very high quality, and the main results are well illustrated.

      Association between EEG characteristics linked to good clinical outcome matched by mouse EEG data linked to CSD.

      Characterization of multiple consequences of CSD following ECT in the mouse brain.

      Weaknesses:

      The main characterization of CSD propagation comes from GCaMP Ca2+ measurements, as previously reported (Rosenthal et al., 2025). That prior study also provided key electrophysiological evidence of CSD with a DC shift after ECT in mice (supplemental data). Given the prior evidence for ECT-CSD, the additional measures shown in the current manuscript are fully expected. Thus, the 2-photon imaging of Ca2+ elevations following CSD (Figure 4) is consistent with prior 2-photon imaging studies of CSD, and the complex hemodynamic and pH changes are expected to contribute to propagation of EGFP fluorescence changes (Supplemental Figure 5). These data are well presented, but, contrary to the results section here, these results appear confirmatory rather than necessary to build a case that the key event generated by ECT is a CSD.

      The authors state that "our conclusion that CSD is the primary driver of plasticity is based on its role in driving Fos expression" (line 472). Related to the point above, there is already a very well-established literature showing that CSD leads to rapid and robust Fos expression in rodent cortex, so this is fully consistent with prior work. The prior work, CSD-fos work, should be summarized and/or cited more clearly in the manuscript. Showing that Fos increases only in the hemisphere where there is a large CSD-Ca2+ wave is a clear demonstration of this. While Fos increases can certainly be well linked to plasticity in some experimental paradigms, the implication that Fos increases underlie CSD-induced plasticity and possibly therapeutic effects of ECT is not appropriate. Fos increases after CSD are a reliable marker of the very strong neuronal activation that occurs, but Fos increases are not specific for plasticity and can be activated by challenges that do not generate synaptic plasticity. A range of other gene expression changes have been identified with CSD and may contribute to adaptive plasticity; these could be mentioned alongside speculation about Fos. To support the main conclusions of this paper about CSD driving plasticity via Fos, Fos knockout or knockdown studies are needed, as has been used in prior plasticity studies.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript combines widefield calcium imaging, electroencephalography, 2-photon imaging, and immunohistochemistry in mice to re-demonstrate that electroconvulsive stimulation (ECS) induces a seizure followed by cortical spreading depolarization, as previously shown. The putative novel finding - which is not unexpected - is that ECS is also correlated with increased expression of the immediate early gene cFOS, although this has also been shown previously. The authors speculate that CSD drives cFOS expression, which might contribute to the therapeutic effects of ECT; however, experiments performed do not provide causal evidence for this hypothesis. Instead, the authors use expression of cFOS - a nonspecific activity-dependent gene induced in various pathological and non-therapeutic contexts - as a proxy for plasticity and/or therapeutic effect. Hence, overall, the significance of the findings is limited and primarily serves to replicate prior work, with the evidence evaluated as incomplete.

      Strengths:

      The experiments are generally well executed from a technical perspective.

      Main Weaknesses to be addressed in revision:

      (1) The main findings of this paper are replication experiments of prior work, and thus, the novelty and significance of this manuscript are relatively limited.

      - It is already known that the mean frequency of ECT-induced seizures decays between peak and offset in humans (Stuiver et al. Clin Neurophysiol. 2026 Jan:181:2111439. doi: 10.1016/j.clinph.2025.2111439) and mice (Murakami et al. J Pharmacol Sci 2008 Jan;106(1):78-83. 10.1254/jphs.FP0071453), which the authors re-demonstrate in Figure 1.

      - It has already been demonstrated that ECT in mouse models induces lateralized CSD waves in a manner that depends on stimulation parameters and the initial evoked response during stimulation (Rosenthal et al. Nat Comm. 2025 May 18;16(1):4619. doi: 10.1038/s41467-025-59900-1); the authors replicate this in Figures 1, 2, 3, 6.

      - It is already widely established that EEG and calcium signals are highly concordant in mouse brain physiology, as shown in Figure 1. It is already known that CSD propagates from supragranular to granular and infragranular layers (Zakharov et al. Epilepsia. 2019 Dec;60(12):2386-2397. doi: 10.1111/epi.16390) as shown in Figure 4.

      - It is already known that CSD waves induce cFOS expression (e.g., Dell'Orco et al. Front Cell Neurosci. 2023 Dec 14:17:1292661. doi: 10.3389/fncel.2023.1292661; Hermann and Hossman. Neuroscience. 1999 Jan;88(2):599-608. doi: 10.1016/s0306-4522(98)00249-8) as the authors replicate in Figure 5.

      Minimally, the authors should revise claims regarding novelty, as the manuscript, as written, is misleading to a reader not familiar with the field. There is limited innovation in re-demonstrating that these events are seizures and that they involve spreading depolarization.

      (2) The authors frame their hypothesis that CSD could be a potential mediator of the therapeutic effects of ECT, but they do not measure therapeutic effects or directly test this hypothesis. The principal advancement of the paper is showing that ECT-induced CSD triggers hemisphere-specific cFOS expression as a proxy of plasticity. However, it is already known that CSD induces cFOS expression (as noted above). The observation that cFOS expression was induced only by CSD, not by the initial seizure, is likely a byproduct of the greater activity induced by CSD than by seizure. cFOS expression is nonspecific to plasticity or therapeutic effects and can be triggered by many non-therapeutic interventions. The cFOS data thus do not meaningfully measure therapeutic plasticity. The authors also selectively cite references suggesting that EEG metrics such as seizure duration predict positive therapeutic outcomes, but this link is controversial and not well established in the clinical literature.

      Minor Weaknesses:

      (3) For the n=3 mice used for concurrent 2P imaging with microprism implant, these animals also had ChrimsonR co-expression, but there are no optogenetic studies described in this paper, which is confusing. Yet, this co-expression introduces a significant confound, as GCaMP6 emission (525/50nm band in this study) will overlap substantially with the ChrimsonR excitation spectrum. Thus, the fluorescence emission used to image these neurons may be optogenetically activating them at the same time. Please explain.

      (4) Incision of the cortex for implantation of a prism is a significant cortical injury that likely induces CSD instantaneously and may change the propensity for CSD in subsequent recordings. Please comment on this limitation and address how much time elapsed after surgery before imaging.

      (5) Method details are missing or insufficiently described for location, titer, and injection strategy for 2-photon experiments.

      (6) Given the wide range of parameters used for ECS in mice and ECT in humans, the authors should provide tables for what stimulation parameters were used for each recording. These protocols were chosen manually rather than randomly or systematically, which introduces confounding factors into analyses that use parameters as an independent variable.

      (7) While much of the cFOS staining after unilateral CSD shows hemisphere-specific asymmetry, several regions (piriform cortex, amygdala, thalamus) do appear to have bilateral cFOS expression. Please comment on this.

      (8) The discussion states: "If CSD accounts for plasticity effects, triggering a CSD in a non-seizure context may be sufficient to elicit therapeutic effects. This is supported by the clinical success of ultra-brief stimulation treatments that do not cause seizures, such as rTMS with accelerated protocols, which achieves treatment efficacy on par with ECT for major depressive disorder". Are the authors implying that TMS induces CSD? What evidence supports this idea?

      (9) This statement - "Assuming psychosis is the result of thalamocortical coupling that is too weak in frontal areas of the cortex" (lines 583-585) - may be overly speculative.

    1. eLife Assessment

      This important study establishes a robust live-imaging toolkit to characterize excitatory and inhibitory synaptic dynamics during neuronal development, advancing our mechanistic understanding of synaptic homeostasis and neural circuit maturation. The core findings clarify how stable E/I balance is maintained despite persistent synaptic turnover, with broad implications for developmental neurobiology and neurodevelopmental disorders. The methodology and quantitative data are convincing and well validated, and this work represents a significant advance that will be of significant interest to researchers in synaptic biology, cell imaging and neuroscience.

    2. Reviewer #1 (Public review):

      [Editors' note: all three reviewers confirm that all initial concerns have been fully resolved through comprehensive revisions and supplementary analyses.]

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      Comments on revised version:

      The authors have addressed all my questions/comments. No further questions for this manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomato-Gephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HA-Homer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Comments on revised version:

      The authors addressed all my questions and comments. Their edits have made this paper significantly stronger. I believe that this is an important paper for the field.

    4. Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      Comments on revised version:

      The authors have fully addressed my concerns, and this is now a strong manuscript for the synaptic field.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this valuable study, the authors developed long-term imaging tools to simultaneously monitor the temporal and spatial dynamics of excitatory and inhibitory synapses and reported that excitatory and inhibitory synapses need to develop synergistically during synaptogenesis to maintain balance. While the analysis and quantification of the imaging data are incomplete, there is convincing evidence that the developed tools are feasible. If these tools can function stably in vivo, their applications will be much broader.

      We have completely overhauled our analysis and quantification methods and generated custom-made drift correction and tracking pipelines. Also, we have tested these tools ex vivo.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      We thank the Reviewer for their positive feedback and careful review of our study.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      We also thank the Reviewer for their insights and suggestions. We have added discussion of this important point to the Discussion section. Furthermore, we have tested the applicability of our tools ex vivo (new Figures 1, 4, and 6). While using these tools in vivo for live imaging is the eventual goal, we started in a reduced culture system given the relative simplicity. Our current study now provides a framework for future experiments applying these approaches in more complex in vivo systems.

      Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomatoGephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HAHomer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      We thank the Reviewer for their positive assessment of our study.

      Weaknesses:

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Minor weaknesses of the manuscript include:

      (1) The lack of a characterization of endogenous Homer1-positive excitatory synapses using TKIT.

      We attempted to perform live imaging of endogenous Homer1-positive synapses using the TKIT approach by tagging endogenous Homer1 with mClover3 but encountered low signal/noise while live imaging. This prompted us to focus our current study on live imaging endogenous Gephyrin. Future studies using more robust tags (e.g. StayGold, HaloTag) for TKIT tagging of endogenous Homer1 will likely help circumvent this issue.

      (2) Discussion about other approaches to study excitatory and inhibitory synapses using endogenous proteins (e.g., intrabodies - FingR or nanobodies) should be included.

      This important point was also raised by other Reviewers. We have now significantly expanded the Discussion section, including discussion of this point.

      (3) The activity state of a neuron and/or a synapse might alter the dynamic properties (formation, maintenance, and/or elimination). A discussion on whether the overexpression of Homer1 and/or gephyrin might alter synapse/neuron activity would provide greater interpretability of the results. A discussion of the potential limitations and benefits of the reporter and TKIT approaches would be beneficial.

      We agree and have added discussion of these points to the Discussion section.

      (4) A description and interpretation of the computational approach to calculate particle tracking would be helpful. I found that particle tracking figures, while elegant, are difficult to interpret.

      As discussed in more detail below, we have generated drift correction and particle tracking approaches for the revised manuscript. We now elaborate on these new approaches in the paper.

      We thank the Reviewer again for their very helpful input and suggestions.

      Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      We thank the Reviewer for their positive assessment of our study.

      However, a number of aspects remain to be addressed in order for the study to support the claims made by the authors. First, the novelty aspect of the development of the fluorescently tagged synaptic proteins is unclear, since reporters of this nature are in routine use in many labs. Second, the analysis of the acquired images often seems incomplete, with only example images but no quantification shown, or the distinction between spatial and temporal dynamics appearing unclear. Third, given this incomplete analysis, the interpretations of the authors are not always convincingly supported by the data presented. In conclusion, substantial improvements are required to render the main messages of the study clear and compelling.

      We agree and have incorporated all of the Reviewer’s suggestions in the revised manuscript (please see below).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This is an interesting study. This reviewer has the following questions/comments for the authors:

      (1) Please provide evidence that the gRNAs targeting each gene of synaptic protein have no offtarget effects.

      We now include analysis of off-target effects for the TKIT tools (new Figure S6).

      (2) While structural E/I balance is shown, functional electrophysiological validation (e.g., mEPSC/mIPSC ratios) is absent. It is interesting to know whether the balanced functional structural changes translate to functional?

      We thank the Reviewer for this insightful suggestion and now include these recordings in the revised paper (new Figure 8).

      (3) In lines 217-218, please define thresholds for "stable" vs. "dynamic" puncta (e.g., temporal and spatial criteria).

      We more clearly define our categorization parameters (e.g. new Figure 2).

      (4) In Figure 5B: The low co-localization between endogenously tagged Bassoon and antibodystained Bassoon is likely due to the low TKIT efficiency. Quite a few HA-tagged Basson signals are insensitive to Basson-antibody. The authors are suggested to explain those.

      We thank the Reviewer for identifying this and add discussion to the Results section.

      (5) For the data analysis. If each n represents an independent neuronal culture, should the authors are suggested to provide the number of neurons/dendrites analyzed for each independent culture?

      We have added these important details to the manuscript.

      (6) Regarding the title, the author used the term "coordinated dynamics". This reviewer finds it is a bit over-claim because the stable ratios of the number of excitatory synapses and inhibitory synapses are likely an association, not actively "coordinated". I suggest that the authors rephrase this.

      We agree that we cannot argue that excitatory and inhibitory synapses are causally coordinated in our current study. Their levels are likely associated by either association or direct coupling, which we now discuss further in the first paragraph of the Discussion. We have rephrased the title accordingly.

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions that I think will improve the manuscript:

      (1) Please define Syn1/2 on line 129.

      We have defined this in the revised paper.

      (2) For Figures 2B, C, and 4B, C: are the puncta in panel C from the dendrites in panels B? If so, it would be helpful to identify the ROIs selected in panels C.

      We now include this in new Figure 2.

      (3) For the particle tracking figures, while the ability to track all synaptic puncta is very impressive, it is sometimes difficult to clearly track the lifespan of a synaptic puncta from the current figures. I believe that it would be helpful if the authors selected specific examples of synapses formed, maintained, and eliminated.

      We agree and now include more examples.

      (4) I believe that more detail about the computational approach and analysis for the particle tracking (Figs 2E and 4E) would help the interpretability of the figure.

      This important point was also raised by the other Reviewers. We generated custom tools during the revision that significantly expand the capabilities of our tracking approaches and more clearly describe them in the revised manuscript.

      (5) Similar to the rigorous gephyrin TKIT analysis (Fig. 6), did the authors perform a similar analysis for Homer1c TKIT? This might be valuable to confirm that overexpression of the Homer1 reporter does not indirectly alter synapse dynamics.

      We attempted to perform live imaging of mClover3 TKIT-tagged endogenous Homer1 but encountered low signal/noise with live imaging. We now add discussion that optimization of more robust tags (e.g. StayGold, HaloTag) will likely be necessary for live imaging of different target proteins.

      (6) The tools developed by Garbett et al. have the potential to be broadly utilized in the field to provide new insight into the coordination of excitatory and inhibitory synapses. It would thus be helpful for the authors to include a discussion about the strengths and limitations of the reporter and TKIT methods relative to other approaches used to live image synapses (e.g., intrabodies (FingR and nanobodies)).

      We have now significantly expanded the Discussion to include these important points.

      (7) In the discussion, can the authors elaborate on whether it is experimentally feasible to apply their TKIT labeling of gephyrin and Homer1c in the same neuron to assess the endogenous excitatory and inhibitory synapse dynamics from the same neuron?

      We have added discussion of this point and also proof-of-concept data supporting tagging of two postsynaptic targets within the same neuron (new Figure S5D).

      Reviewer #3 (Recommendations for the authors):

      (1) While the new tools described in the current manuscript can undoubtedly be used for the described purposes, the novelty of these tools is unclear to me. Viral vectors expressing fluorescently tagged versions of Homer1, synaptobrevin, and gephyrin are commercially available, e.g., via Addgene, and they are in routine use in many labs. CRISPR-mediated strategies for this purpose have also been previously reported (e.g., Willems et al. 2020, PLOS Biology; Fang et al. 2021, eLife). It is not clear to me how the tools reported here present a significant improvement over existing resources, other than that they use different fluorescent tags. If this aspect is a central part of the current manuscript, it should be expanded on in the discussion, including a direct comparison with available tools to highlight the novel aspects.

      We agree and have significantly expanded the Discussion to include these important points. Also, rather than argue that our tools are superior to pre-existing approaches, we adjust the text to argue that our tools and analytical approaches have been designed and optimized for the purposes we apply them to.

      (2) In addition to generating new tagged constructs, the authors also state that they have developed new imaging and analysis strategies to facilitate long-term assessment of synaptic dynamics. However, in many figures, they present only sample images, with little quantification to allow assessment of the wider relevance of the imaged synapses. For example, in Figures 2C and 4C, they present one example each of, e.g., a stable, nascent, transient, or eliminated synapse. However, they do not provide any quantification on how frequently any of these events occur, or whether they can be reliably quantified at all. These quantifications (i.e., percentage of each event type across a large population of synapses) would be necessary and should be added to demonstrate that this tool can be used for more than single example images.

      We have generated custom-made drift correction and particle tracking approaches for the revised manuscript. Based on the reviewer’s suggestion, we have quantified the relative frequencies of stable, nascent, transient, and eliminated synapses (Fig 2B-G, Fig3A-F, Fig 5A-F, Fig 7B-C). These metrics greatly enhance the biological interpretation of our results. We have also added a supplemental movie with an example image with corresponding categorized tracks for each puncta type (Movie S3)

      (3) The authors do present an automated visual representation of spatial track length across the neuron, e.g., in Figure 2E and 4E, although this is also not quantified. Moreover, the track lengths appear surprisingly short, despite the authors' claims that their analyses 'highlight the dynamic nature of excitatory synapses over these timescales'. It is not clear to me whether these short tracks are more than just jitter, either in the synapses themselves or in the images due to technical limitations. E.g., in panel 2E, I see very few examples in which the track is not simply centered around one point, but actually expands over a distance. Quantification of the distance between start and end points of the tracks would be important to support the claim that these synapses are dynamic in terms of spatial translocation (if that is what the authors meant). Or if the 'dynamic nature' of the synapses referred to temporal dynamics, it is unclear to me how this information can be gained from the represented tracks.

      We thank the reviewer for these excellent points. To accurately access spatial motion, we drift-corrected our images with a custom correction algorithm to eliminate stage or microscope drift as a source of contaminating motion (See Methods, Movie S2), in addition to collecting time-lapse imaging with Nikon perfect focus. We noticed heterogeneity in our cultures such that some areas contained very mobile neurites, while other remained stationary (Fig. S1). We binned movies into either moving or still neurites and assessed spatial metrics as suggested (Fig. S1A). Consistent with our binning, puncta on moving neurites showed larger net displacement (distance between start and end points), but puncta on still neurites also showed ~1 µm net displacement (Fig. S1D). We also quantified puncta speed and found that puncta on moving neurites generally moved faster (Fig. S1C). We appreciate the reviewer’s insight that track length were surprisingly short, and after employing our drift correction and revised tracking methods, we now see substantially longer track lengths (Fig 2E, Fig 3C & F, Fig S2B & C). We additionally see a large fraction of tracks that persist throughout the imaging session (Fig 2E, Fig S2B & C).

      (4) In Figure 3, the authors now quantify track length, but in this case in the unit 'minutes', from which I would interpret that this is now meant to assess the temporal dynamics rather than the spatial dynamics. The lack of a clear distinction between spatial dynamics and temporal dynamics is very confusing to me, since these are entirely independent measures. 'Track length' to me indicates spatial dynamics, and I would expect the units to be a measure of distance. 'Track duration', which the authors also use in some places, but inconsistently as far as I can tell, makes sense to me for the assessment of temporal dynamics, with the units being a measure of time. I would strongly recommend being very clear about this distinction, since the current representation of the data is very difficult to follow and interpret.

      In addition to new spatial metrics, we have clarified in the text when we are referring to spatial dynamics (distance) versus temporal dynamics (time). As suggested, we use duration when referring to time, and speed or distance when referring to spatial metrics.

      (5) The images from the newly generated CRISPR-based tags in Figures 5-7 are striking and very compelling - these will be very useful tools. However, here too, it seems that the interpretation of the data does not really match the results. All quantification indicates that there is very little change in synapse density or other assessed parameters over the time course of the imaging, and yet the authors emphasize the dynamic nature of visualized synapses. More compelling quantification would be needed to support this claim.

      We have quantified spatial and temporal metrics for live neuron culture imaging for all tools developed including CRISPR-based tags (Figure 7).

      (6) The discussion is extremely short and provides almost no integration of the results of the study into the framework of existing knowledge. Instead, it focuses almost exclusively on unanswered questions and future perspectives, which are also important, but not helpful in interpreting the findings from the current study. The latter aspects should be added to provide essential context for the current findings.

      We agree and have added additional discussion of our current findings to help contextualize their significance.

      We thank the Reviewers again for their positive feedback and insightful input, which has undoubtedly strengthened our study.

    1. eLife Assessment

      This important study shows that regions of the human auditory cortex that respond strongly to human voices are also sensitive to vocalizations from closely related primate species. The evidence is convincing and methodologically strong. The work offers significant insight into the evolutionary continuity of voice processing and would be of interest to researchers studying auditory processing and evolutionary neuroscience in general.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance u

      Comments on revised version.

      I thank the authors for having carefully considered and implemented my remarks on the first version.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding-that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations-suggests a potential neural sensitivity to calls form phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weakness:

      The authors only tested vocalizations from three non-human primate species other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      Comments on revised version.

      I have no further comments.

    4. Reviewer #3 (Public review):

      Summary:

      Using fMRI, the authors demonstrate that human temporal voice areas (TVA) respond not only to human vocalizations but also to those of other primates, particularly chimpanzee calls, which share acoustic features with human voices. These findings provide compelling evidence for cross-species vocal processing in the human auditory system and carry important theoretical implications for understanding the evolutionary underpinnings of speech perception.

      Strengths:

      The study offers a valuable comparative design, rigorous acoustic and phylogenetic modeling, and consistent evidence that bilateral anterior TVA regions respond more strongly to chimpanzee vocalizations than to other species' calls. The inclusion of both great apes and monkeys provides a rare cross-species perspective.

      Weaknesses:

      Minor limitations include the acoustic-phylogenetic confound (which the authors partially address with additional analyses), the lack of non-vocal controls to establish true selectivity.

      Overall, the methods, data, and analyses broadly support the claims, with only minor weaknesses that do not undermine the main conclusions. The findings are valuable for the subfield of auditory neuroscience and comparative cognition, with solid evidence supporting the primary claims.

      Comments on revised version.

      After revision, this work has shown great improvement in data analysis, figure organization, and writing. I have no further suggestions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance understanding of the neural bases of voice perception and the evolutionary roots of voice sensitivity in the human brain.

      Weaknesses:

      (1) Acoustic-phylogenetic confound: The design does not fully disentangle acoustic similarity from phylogenetic proximity, as species co-vary along both dimensions. A promising way to address this would be to include an additional model focusing on the acoustic features that specifically differentiate bonobo from chimpanzee calls, which share equal phylogenetic distance to humans.

      (2) Selectivity vs. sensitivity: Without non-vocal control sounds, the study cannot determine whether TVA responses reflect true selectivity for primate vocalizations or general auditory sensitivity.

      (3) Task demands: The use of an active categorization task may engage additional cognitive processes beyond auditory perception; a passive listening condition would help clarify the contribution of attention and task performance.

      (4) Figures and presentation: Some results are partially redundant; keeping only the most representative model figure in the main text and moving others to the Supplementary Material would improve clarity.

      We thank the reviewer for contributing to the improvement of the present study and for the extremely constructive criticism. Concerning the identified weaknesses of our work, we provide here some general answers while the detailed review (below) addresses point-by-point the reviews in high detail.

      (1) We totally agree that acoustics and phylogeny cannot be disentangled in our study, which is a limitation. We now provide the suggested analysis on the acoustic specificities of chimpanzee and bonobo calls.

      (2) This point on selectivity vs. specificity is indeed crucial, and we now provide a more careful viewpoint and phrasing on this aspect, since our study can only provide partial arguments for this important distinction.

      (3) Task demand following species categorization might rightfully yield to the engagement of distinct brain network compared to merely listening to the stimuli. We discuss this aspect and put forward the argument that, while we cannot control for this aspect, our attentional control study performed by an independent sample, N=28 provides clear evidence that no species triggered an attention bias. In other words, task demand might play a role, but at least in the study we know that attentional resources were not biased towards one species in particular since no effects were observed.

      (4) We agree that results were not articulated in a clear fashion and that figures were redundant. We addressed this aspect and regrouped the figures where appropriate while we include the rest in the supplementary material now.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding - that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations - suggests a potential neural sensitivity to calls from phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weaknesses:

      The interpretation of the findings in this paper regarding the evolutionary continuity of voice processing lacks sufficient evidence. A simple explanation is that the observed effects can be attributed to the similarity in low-level acoustic features, rather than effects specific to phylogenetically close species. The authors only tested vocalizations from three non-human primate species, other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      We want to thank the reviewer for the constructive criticism and for evaluating the manuscript.

      Concerning the principal weakness highlighted, we provide new analyses behavioral, acoustics, model-based fMRI that improve our understanding of the influence of both phylogeny and bioacoustics in our data. We argue that the explanation proposed by the reviewer cannot explain our results, as also observed in several other research from us and others. We discuss this aspect and emphasize that including stimuli from more species would greatly improve the understanding of phylogeny and bioacoustics in this context.

      Reviewer #3 (Public review):

      Summary:

      Ceravolo et al. employed functional magnetic resonance imaging (fMRI) to examine how the temporal voice areas (TVA) in the human brain respond to vocalizations from different nonhuman primate species. Their findings reveal that the human TVA is not only responsible for human vocalizations but also exhibits sensitivity to the vocalizations of other primates, particularly chimpanzee vocalizations sharing acoustic similarities with human voices, which offers compelling evidence for cross-species vocal processing in the human auditory system. Overall, the study presents intellectually stimulating hypotheses and demonstrates methodological originality. However, the current findings are not yet solid enough to fully support the proposed claims, and the presentation could be enhanced for clarity and impact.

      Strengths:

      The study presents intellectually stimulating hypotheses and demonstrates methodological originality.

      Weaknesses:

      (1) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy.

      We thank the reviewer for evaluating our manuscript and for the constructive criticism as well as the many suggestions. Concerning the weaknesses of the study, we provide here some quick answers while more detailed responses can be found below.

      (1) We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator.

      (2) We totally agree that figure redundancy was a problem and we now reduced confusion by combining congruent aspects while pushing other results to the supplementary material.

      Recommendations for the authors:

      Reviewing Editor Comments:

      With additional analyses and discussions, the work has the potential to offer important insight into the evolutionary continuity of voice processing.

      We thank the Reviewing Editor for this additional motivation and for offering us the possibility to revise our manuscript. We will now provide our point-by-point reviewing, referring to manuscript modifications by section and/or line number(s). All modifications are also highlighted in light grey in the text.

      Reviewer #1 (Recommendations for the authors):

      The manuscript is clearly written and addresses an important comparative question about the specificity of human TVA responses. The acoustic analyses are well designed, and the imaging work is careful and thorough. However, several conceptual and methodological issues need clarification or tempering of claims, particularly regarding (i) the distinction between sensitivity and selectivity, (ii) the confounding of acoustic and phylogenetic factors, and (iii) the interpretation of "chimpanzee-specific" TVA activity.

      (1) Introduction

      Line 48: cite more recent infant EEG evidence for early voice sensitivity (Calce, Curr Biol).

      The reference and explanation were added, lines 46-48.

      Line 53: mention recent data on voice processing in marmosets (Jafari, Cell Rep; Dureux, Curr Biol).

      We added the references and the mention of these interesting studies on common marmosets, lines 53-54.

      Line 59: Fecteau et al. (2004) already explored cross-species selectivity; please integrate and discuss.

      We now mention here the work from Fecteau and colleagues and its relevance, see lines 57-59.

      Line 70: clarify that in [27] (Bodin et al., 2021) human TVA responded similarly to human nonverbal vocalizations and macaque coos, likely due to acoustic similarity.

      We added this important aspect, thank you for this precision. See lines 71-72.

      Clarify why an active species-categorization task was chosen instead of passive listening, which is standard in TVA research. Were participants familiarized with stimuli beforehand?

      We added a sentence on this aspect, but basically to summarize it here: we wanted to be able to test human recognition of nonhuman primate species’ calls. From the start, we wanted to test the frontal mechanisms related to decision-based processes of humans when categorizing non-human primate calls hence the 2023 article we published. See lines 75-77 and we also added information on familiarization to the stimuli in the Methods, lines 679-682.

      The 16 acoustic features mentioned should be briefly defined earlier, as they are central.

      We feel like describing 16 acoustic parameters in the introduction would be heavy on the reader, so we instead added a reference to the supplementary table (Table S1) in which these are named and described. See line 80.

      Explain why only chimpanzees and bonobos were selected among the great apes, and discuss the value of including both, given their equal phylogenetic proximity but largely dissimilar acoustics.

      The stimuli were obtained by Thibaud Gruber and his team and through collaborations with Katie Slocombe and Zanna Clay. Unfortunately, at the time we could only use chimpanzee and bonobo calls for the great apes. Therefore, it was mainly a material constraint rather than a deliberate choice to exclude other great apes. We now discuss this aspect and present the absence of other great apes as a limitation (lines 587-591).

      Rephrase references to "recruitment" of TVA - this term implies general activation, while the key question concerns selectivity (stronger responses to voices vs. non-vocal controls).

      We rephrased throughout the manuscript, thank you for this suggestion.

      The hypothesis section should more clearly separate the acoustic and phylogenetic predictions, and clarify which earlier data motivate each.

      We now explicitly categorize the hypotheses according to either Bioacoustics or Phylogeny to clarify. We also added references motivating each hypothesis. See lines 114-120.

      (2) Methods

      Clarify whether stimuli were RMS-normalized or otherwise balanced for energy (line 128).

      Sound pressure level was kept constant but the stimuli were not normalized, specifically to avoid a negative impact on their naturality. We added a sentence (lines 131-132) including a reference on this aspect.

      The task design could benefit from reporting accuracy in addition to reaction times for the 4AFC species classification task.

      We agree this aspect was missing. We now report accuracy data (controlled for reaction times and acoustics of Model 3) for the species categorization task (lines 147-165; Fig.1B), and in the Methods (lines 769-786). The fitted regression values of this analysis are also used for a new fMRI model (Model 4), to uncover within-TVA correlates of the probability of correct species categorization (lines 309-325; Fig.4).

      Please note that previously, the behavioral data of the species categorization task were completely absent (N=23), and the reaction times data previously part of Fig.1 were for the species attentional bias task (independent sample of N=28). Since this aspect was not clear at all (same remark by all reviewers—apologies for that), we now include a clear separation in Fig.1, with newly added panels D & E part of a distinct figure area named: “Control task: Testing for Species attentional bias (N=28)”. Panel D illustrates the control task paradigm (each species as exogenous cue; “dot-probe” paradigm) while panel E shows the results (target sine wave tone or “bip” detection reaction times), showing that no species triggered more attentional capture than the others (Species effect non-significant).

      The acoustic parameters used in Models 2 and 3 should be explicitly listed in the Methods (even if already published elsewhere).

      In addition to their description in Table S1, we now include the 16 acoustic parameters used to calculate acoustic distance between the species in the Methods, see lines 828-844.

      Consider simplifying the presentation of the three models: a figure summarizing their relationships would help.

      We now include only one figure (Fig.2) for Model 3, and we pushed model 1&2 to the supplementary material. We also simplified Fig.3 for a clearer view of the overlaps between the 3 models within the TVA.

      The description of “systematic and thorough control of phylogeny” (line 119) is overstated, given that only three nonhuman species were included.

      We agree with the reviewer and we suppressed both “systematic” and “thorough” from the sentence.

      Provide rationale for not including a nonvocal control category (e.g., scrambled vocalizations or environmental sounds) to assess TVA selectivity.

      The main objective of the study was to uncover whether human participants could recognize the vocalizations from nonhuman primates—from both great apes and monkeys—as compared to the human voice. We therefore did not include nonvocal or noise stimuli. We added this point as a limitation in the Discussion (lines 593-596 and 609-611).

      Even though we did not include such stimuli for the reason mentioned above, the delineation of subtypes of nonvocal material within the TVA of our participants (Fig.2) are, in our opinion, clarifying the message: chimpanzee-selective activations are fully within ‘voice vs. animal’ and ‘voice vs. nature’ TVA subareas, while it is not the case in ‘voice vs. music’ and ‘voice vs. noise’ TVA subareas.

      Clarify if participants were trained or had a practice session to recognize the four species before scanning.

      The participants were indeed trained on 3 stimuli per species before entering the MRI scanner. These stimuli were discarded from the species categorization task. We added a sentence about this aspect, see lines 131-132.

      Specify what is meant by "no good or bad response" in the attentional control task (line 724).

      We suppressed this wording as it was highly confusing.

      (3) Results

      Behavioral accuracy should be reported to complement reaction times.

      We now added behavioral data for the species categorization task as well as the neural correlates of accurate species categorization. See our previous response above (‘‘‘).

      Figures 2-4 largely overlap; consider merging or simplifying to reduce redundancy.

      We agree and this point was raised by the other reviewers as well. Task-based results are now presented only for Model 3 as Fig.2, while Fig.3 (previously Fig.5) summarizes the overlap between the three models. Figures for Models 1 & 2, previously labelled Fig.3 and Fig.4, were moved to the supplementary material.

      Figure 2: Please indicate more clearly where "chimp-selective" areas are located (perhaps with zooms).

      We agree, we now modified Fig.2 with zoomed-in panels and a clearer outline of chimp-selective areas (solid blue outline). This outline is also referenced in the text (lines 236-237).

      Correction for multiple contrasts: With many pairwise tests, adjustments (Bonferroni or FDR) should be mentioned explicitly.

      We now specify ‘FDR correction at the voxel level’ at the beginning of the Results section (lines 195-198) as well as in each figure.

      Replace "specific to chimpanzee" with "selective for chimpanzee" to avoid implying exclusivity.

      We made the suggested replacement throughout the manuscript.

      Discuss whether the small macaque-related clusters might simply reflect acoustic overlap rather than true category selectivity.

      We added a section on this important aspect, including results that support the role of mid-STG/STS regions for more noise-like stimuli, including the use of macaque coos. See lines 450-461.

      (4) Discussion

      The discussion overstates claims of "chimpanzee-selectivity" in TVA. The evidence shows relative preference, not absolute selectivity.

      We now specify from the start of the Discussion that we are not interpreting the results as absolute selectivity but rather as more relative preference, see lines 371-373.

      The authors repeatedly conflate acoustic and phylogenetic factors; this should be explicitly acknowledged as a limitation.

      We agree, and we completed the limitations section already dedicated to this aspect by a more explicit account of the confound, see lines 609-611.

      Clarify what is meant by "recruitment" and "selectivity" (lines 411-419, 577). TVA activity often reflects enhanced responses to voices compared to non-vocal sounds, not exclusive activation.

      We clarified this wording in the Discussion (lines 377-378) and replaced another instance by “activated the […]” to make it clearer what we imply, namely enhanced activity triggered by chimpanzee calls within human TVA.

      The lack of non-vocal control conditions should be discussed as a major interpretive limitation.

      We added this point as a limitation in the Discussion (lines 593-596).

      The statement that "chimpanzee-selective activity" arose in humans who have never been exposed to chimp calls (line 450) invites evolutionary speculation but should be more cautiously phrased.

      We agree, and we rephrased by: “[…] with chimpanzee calls triggering responses in the anterior STG/TVA of our human participants […]”. See lines 432-433.

      The comparison to recent macaque data (Giamundo et al., 2024 PNAS) is crucial: these findings of human-voice-selective neurons in macaques directly parallel the present human-chimp result.

      We agree with the reviewer, and we are hopeful to read similar results for other apes/great apes in the future.

      Reviewer #2 (Recommendations for the authors):

      (1) The primate vocalizations used in this study were recorded in diverse social and emotional contexts, which may have contributed to the observed differences in TVA activation. Since the temporal voice areas are known to be sensitive to affective and socially relevant cues, these contextual differences could confound the interpretation of species-specific neural responses. Therefore, I suggest that the authors conduct a post-hoc analysis to quantify and compare the affective valence, arousal levels, and social contexts associated with each stimulus set.

      We agree that the TVA are sensitive to social—or socially relevant—cues, motivating the very thorough work of the expert reserve personnel on-site to accurately categorize the calls according to the very specific context they were produced in. If the reviewer meant presenting these stimuli to non-expert participants and asking them to categorize the context or valence, we think it would make no sense since the ratings would be completely below chance level and therefore uninformative. The newly added behavior—and model-based fmri—data include this crucial point, a factor that we named ‘Context’ in our analyses. In fact, for each species’ 18 stimuli, we control for agonistic and affiliative production context—split evenly, per species. Also, computing an additional posthoc analysis by splitting the stimuli according to Context would result in too few trials to get sensible and reliable fMRI results.

      That being said, our study targets this specific aspect by extracting the acoustic features that characterize our stimulus set the best, across context-species-valence-arousal, which is exactly what we want. Through the three types of modeling we used—from more simplistic to more elaborate the results converge only for one species: chimpanzee calls.

      We think the addition of behavioral data, model-based fMRI data, and the specific analysis on acoustic differences between chimpanzee and bonobo calls strengthens the message and the validity of our findings.

      (2) Although the author mentioned that the behavioral effects triggered by these vocalizations have been reported previously, the behavioral responses of the participants in the current study are also crucial for our understanding of the results. If the MRI data can be combined with the participants' behavioral responses for comprehensive analysis, the conclusions of this study will be more compelling.

      We agree with the reviewer, and we added the behavioral data—controlling for reaction times, production context and acoustics of interest—and we also included a model-based fMRI modeling of the probability of correct species categorization as Model 4, Fig.4. See, respectively: lines 147-165, Fig.1B; Methods, lines 769-786; Neuroimaging results, lines 309-325.

      (3) I am still not convinced that phylogenetic proximity drives the observed neural selectivity. While chimpanzee vocalizations do elicit stronger responses in anterior STG, the claim that this reflects evolutionary relatedness lacks evidence. If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree with the reviewer that generalizing our results in terms of phylogenetic proximity alone is not a viable option. Including many more primate species including other great apes would be necessary, and we mention this crucial aspect in the limitations section. We also insist in the Discussion on the interdependence between phylogeny and acoustics in our data, since: 1) we cannot fully disentangle these factors here, 2) we cannot attribute our results to either one or the other. See lines 387-390, 410-411, 473-477, 587-591.

      If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree, and nobody could disagree: if an auditory object is extremely similar to the human voice in terms of acoustics, it would therefore potentially activate the TVA. This is exactly our message: in the natural ‘auditory world’, the calls from chimpanzees seem to be among the very few animal auditory signals that are sufficiently close, acoustically, to the human voice and therefore trigger TVA activity. They also happen to be the calls from a species which is phylogenetically the closest to humans with minimal differences with other great apes. Our results are in that sense very aligned with work from the laboratory of Pascal Belin, namely on ‘voice patches’ in the primate brain located in the (anterior) TVA, cited in our manuscript.

      We therefore think our interpretation does not exclude that in the near future, similar results within the TVA could be observed for other auditory objects, and if animal, from a species potentially much more distant phylogenetically or from vocal signals of other great apes.

      We added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as scrambled or spectrum shifted per-species stimuli, which would have made the interpretation clearer identical acoustics but alteration/destruction of the species auditory object. See lines 593-596 and 609-611.

      Reviewer #3 (Recommendations for the authors):

      While the manuscript presents intriguing results, several concerns are raised for further consideration, detailed below.

      We thank the reviewer for evaluating the manuscript and for the constructive criticism and suggestions.

      Major concerns:

      (1) This study claims that bilateral anterior superior temporal gyrus (aSTG) in humans can be specifically activated by chimpanzee vocalizations rather than all other primate species after regressing out relevant acoustic parameters using three distinct analyses. I am wondering if a control stimulus (e.g., scrambled chimpanzee vocalizations) were presented, would the activation patterns in these same temporal voice areas (TVA) exhibit significant differences compared to the natural chimpanzee vocalizations?

      We completely agree with the reviewer, and this point was also raised by the other reviewers. We therefore added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as per-species scrambled or spectrum shifted stimuli, which would have made the interpretation clearer—identical acoustics but alteration/destruction of the species auditory object. See lines 609-611.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy. E.g:

      Figure 1C is the same as Figure S1. In addition, Figure 1C lacks a figure legend and descriptive label.

      The scatter plots in Figures 2D, 2H, 3D, 3H, and 4D, 4H are same as those in Figures S2, S3, and S4. However, some of these duplicate plots even have inconsistent axis labels.

      In several panels, the main figures appear to be summaries derived from the supplementary figures. The authors should organize these figures well to eliminate redundancy.

      Please double-check all the figures to make sure of accuracy.

      We agree that the figures were badly organized and were too crowded and redundant. We now suppressed the redundancy between Fig.1 and Fig.S1, and we reduced fMRI results to one figure for statistical Model 3 while the other models are in the supplementary data—we also justify this decision in the text by highlighting that model 3 is the most elaborate and sensitive one. Fig.3 (previously ‘Fig.5’) shows the overlaps between models and was simplified and clarified as well.

      (3) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task. It is possible that processing vocalizations from certain species requires more cognitive effort or induces higher decision uncertainty. Could the observed neural effects be confounded by the decision-making process itself?

      We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator. We now display these results in Fig.4 and we introduce the motivation factor for including a categorization task rather than more traditional passive listening (lines 75-77), as well as limitations, lines 595-596.

      (4) One interesting attempt of this study is to dissociate biologically salient information in animal vocalizations from their low-level acoustic properties. This presents a fundamental conceptual challenge: how to rigorously disentangle a vocalization's species-specific attributes from its inherent acoustic correlates. More precisely, what essential biological information persists in a species' vocal signal after statistically accounting for all quantifiable acoustic features? I recommend that the authors address it in the discussion.

      We thank the reviewer for this very important comment, and for suggesting we discuss it in the manuscript. We completely agree: we cannot fully orthogonalize species and acoustics, and this aspect relates also more broadly to cognitive and affective neuroscience studies involving vocal material. Namely: “What is an auditory object without acoustics?”

      We included a full paragraph on this aspect, see Discussion, lines 570-584.

      (5) If a brain region, such as TVA, is responsive to both acoustic parameters and biological meanings of animal vocalizations, the method used in this study might be inadequate by setting covariates to zero. It is possible that species information is embedded within a specific acoustic pattern. The current modeling approach may not capture such complex information and could potentially introduce bias when estimating the species effect. I recommend that the authors address this issue in the discussion.

      We thank the reviewer for this point once again, we addressed it in the Discussion, lines 581-584, and also in the section dedicated to study limitations, lines 609-613.

      (6) In the discussion, non-human primate vocalizations are "unreadable" to humans. If this is the case, what is the fundamental perceptual difference between these vocalizations and those from the other animal species? An alternative and highly plausible explanation for the findings is the differential familiarity of the participants with the various species, driven by media exposure (e.g., documentaries) or zoo visits and interactions. The authors need to provide a stronger justification for their control stimuli and directly address, either through discussion or additional analysis, how the factor of familiarity might explain their results better than the proposed "evolutionary distance" hypothesis.

      We now discuss this important aspect, see lines 560-569.

      We thought about doing additional analyses on this aspect but we concluded that we did not have any reliable indicators of familiarity for our participants, and additionally they were all recruited for being ‘unfamiliar’ with great apes or old-world monkeys’ vocalized communication.

      Also, frequent mismatches in the media between images of apes and the associated vocal signals (for instance, the depiction of a chimpanzee but with background audio of macaque coos) are not helping this cause.

      Minor:

      (1) No figure legend and result description for Figure 1.

      Figure 1 has a legend, maybe it was cut out during the uploading process, but it is present and verified now.

      (2) In the main text, three statistical models were referenced. Was the data used in each subsequent statistical model derived from the processed data of the preceding model? Please clearly explain this in the main text.

      We now specify this aspect in the Methods and the Results section to clarify that each model is independent from the others (lines 964-966 and 189-191, respectively).

      (3) In Figure 5, the two dashed lines representing Model 1 and Model 2 are confusing for readers.

      We modified the figure (now Fig.3) and simplified it by removing some outlines and clarifying the colors, therefore improving readability.

      (4) Lack of reaction times in the species categorization task.

      We clarified behavioral data, including the results for the species categorization task and for the control, exogenous cueing task, see modified Fig.1 and behavioral results section of the Results.

      (5) Figures 2, 3, 4, 5, Please keep the font size of the figure title consistent.

      Figure title font size were uniformized.

      (6) Line 201, Line 224, and so on, (EFG) → (E, F, G).

      We modified this aspect in every figure legend, including the supplementary material.

    1. eLife Assessment

      Argunşah et al. investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in shaping the responses to single vs multiple whiskers. Based on the observation of a higher density of SST+ interneurons in the septa, the authors investigate the hypothesis that Elfn1-dependent short-term plasticity shapes these responses. This important study is, however, supported by incomplete evidence; factors restricting the strength of evidence are the limited spatial resolution of the multi-unit activity, as well as the lack of a mechanistic explanation. This provocative and intellectually stimulating hypothesis provides a contribution to work on how different cell types shape cortical representation.

    2. Reviewer #1 (Public review):

      Summary:

      Argunşah et al. describe and investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in the whisker-related primary somatosensory cortex (S1). Upon repeated stimulation, the authors report that the response ratio between multi- and single-whisker stimulation increases in layer (L) 4 neurons of the septal domain, while remaining constant in barrel L4 neurons. The authors attribute this divergence to differences in short-term synaptic plasticity, particularly within somatostatin-expressing (SST⁺) interneurons. This interpretation is supported by 1) the increased density of SST+ neurons in L4 of the septa compared to barrel domain, 2) the stronger response of (L2/3) SST+ neurons to repeated multi- vs single-whisker stimulation and 3) the reduced functional difference in single- versus multi-whisker response ratios across barrel and septal domains in Elfn1 KO mice, which lack a synaptic protein that confers characteristic short-term plasticity, notably in SST+ neurons. Consistently, a decoder trained on WT data fails to generalize to Elfn1 KO responses. Finally, the authors report a relative enrichment of S2- and M1-projecting cell densities in L4 of the septal domain compared to the barrel domain, suggesting that septal and barrel circuits may differentially route information about single vs multi-whisker stimulation downstream of S1.

      Strengths:

      This paper describes and aims to study a circuit underlying differential response between barrel columns and septal domains of the primary somatosensory cortex. This work supports the view these two domains contribute distinctly to the processing single versus multi-whisker inputs and highlight the role of SST+ neuron and their short-term plasticity. Together, this study suggests that the barrel cortex multiplexes whisker-derived sensory information across its domains, enabling parallel processing within S1.

      Weaknesses:

      Although the divergence in responses to repeated single- versus multi-whisker stimulation between barrel and septal domains is consistent with a role for SST⁺ neuron short-term plasticity, the evidence presented does not conclusively demonstrate that this mechanism is the critical driver of the difference. The lack of targeted recordings and manipulations limits the strength of this conclusion: SST⁺ neuron activity is not measured in L4, nor is it assessed in a domain-specific manner. The Elfn1 knockout manipulation does not appear to selectively affect either stimulus condition, domain or interneuron subtype. Finally, all experiments were performed under anesthesia, which raises concerns about how well the reported dynamics generalize to awake cortical processing.

    3. Reviewer #3 (Public review):

      Summary:

      This study investigates the functional differences between barrel and septal columns in the mouse somatosensory cortex, focusing on how local inhibitory dynamics (particularly involving SST⁺ interneurons) may mediate temporal integration of multi-whisker (MW) stimuli in septa. Using a combination of in vivo multi-unit recordings, calcium imaging, and anatomical tracing, the authors propose a model in which Elfn1-dependent synaptic facilitation onto SST⁺ interneurons contribute to the distinct sensory responses to MW input in barrels and septa, enabling functional segregation between these domains.

      Strengths:

      The study presents a thought-provoking and useful conceptual model for understanding sensory processing in the somatosensory cortex. While barrel columns have been widely studied, septal regions remain relatively understudied in mice. If septa indeed act as selective integrators of distributed sensory input, this would suggest a novel computational role for cortical microcircuits beyond the classical view focused on barrels. Although still hypothetical, the proposed model in which SST⁺ interneurons contribute to domain-specific sensory responses between barrel and septal domains is intriguing and opens new avenues for investigating inhibitory circuit mechanisms.

      Weaknesses:

      The primary limitation of this study lies in the spatial and cellular specificity of the recording techniques. The physiological data rely predominantly on unsorted multi-unit activity (MUA) recorded with low-channel-count silicon probes. Because MUA aggregates signals from multiple neurons over a radius of approximately 50-100 µm (comparable to or larger than the width of septal domains in mice), it remains difficult to confidently attribute the recorded activity exclusively to septal versus barrel populations. The authors have now addressed this concern more carefully by reframing their interpretation in terms of "septal-enriched" populations and by providing additional threshold-based analyses suggesting that the principal effects are more robust in Layer 4. These additions substantially improve the manuscript and support a more cautious interpretation of the findings. Nevertheless, the proposed Elfn1/SST⁺ mechanism remains supported primarily by indirect evidence. Although the calcium imaging data provide useful support for stimulus-dependent SST⁺ recruitment, these experiments were restricted to L2/3 interneurons and therefore do not directly test the Layer 4 circuit mechanism proposed to underlie the electrophysiological observations. Direct in vivo cell-type-specific recordings and manipulations in Layer 4 would ultimately be required to establish the proposed mechanism more conclusively.

      Comments on revised version.

      I have read the revised manuscript and overall, I think the authors have addressed my major concerns appropriately. I appreciate the substantially moderated interpretation of the findings and the additional analyses clarifying the limitations of the MUA recordings.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Argunşah et al. describe and investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in the whisker-related primary somatosensory cortex (S1). Upon repeated stimulation, the authors report that the response ratio between multi- and single-whisker stimulation increases in layer (L) 4 neurons of the septal domain, while remaining constant in barrel L4 neurons. The authors attribute this divergence to differences in short-term synaptic plasticity, particularly within somatostatin-expressing (SST<sup>+</sup>) interneurons. This interpretation is supported by

      (1) The increased density of SST+ neurons in L4 of the septa compared to barrel domain,

      (2) The stronger response of (L2/3) SST+ neurons to repeated multi- vs single-whisker stimulation and

      (3) the reduced functional difference in single- versus multi-whisker response ratios across barrel and septal domains in Elfn1 KO mice, which lack a synaptic protein that confers characteristic short-term plasticity, notably in SST+ neurons.

      Consistently, a decoder trained on WT data fails to generalize to Elfn1 KO responses. Finally, the authors report a relative enrichment of S2- and M1-projecting cell densities in L4 of the septal domain compared to the barrel domain, suggesting that septal and barrel circuits may differentially route information about single vs multi-whisker stimulation downstream of S1.

      Strengths:

      This paper describes and aims to study a circuit underlying differential response between barrel columns and septal domains of the primary somatosensory cortex. This work supports the view these two domains contribute distinctly to the processing single versus multi-whisker inputs and highlight the role of SST+ neuron and their short-term plasticity. Together, this study suggests that the barrel cortex multiplexes whisker-derived sensory information across its domains, enabling parallel processing within S1.

      Weaknesses:

      Although the divergence in responses to repeated single- versus multi-whisker stimulation between barrel and septal domains is consistent with a role for SST<sup>+</sup> neuron short-term plasticity, the evidence presented does not conclusively demonstrate that this mechanism is the critical driver of the difference. The lack of targeted recordings and manipulations limits the strength of this conclusion: SST<sup>+</sup> neuron activity is not measured in L4, nor is it assessed in a domain-specific manner. The Elfn1 knockout manipulation does not appear to selectively affect either stimulus condition, domain or interneuron subtype. Finally, all experiments were performed under anesthesia, which raises concerns about how well the reported dynamics generalize to awake cortical processing.

      We thank the reviewer for their careful reading of the manuscript and their balanced assessment of both its strengths and limitations. We acknowledge the reviewer’s concerns regarding the lack of direct, layer- and cell-type–specific recordings and manipulations of SST<sup>+</sup> interneurons, as well as the use of anesthesia. As noted in the Discussion, these factors limit the extent to which causal mechanisms can be established and the degree to which the reported dynamics can be generalized to awake cortical processing. For this reason, we intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, developmental, physiological, and genetic evidence, rather than as definitive proof. We believe this conceptual framing appropriately reflects the scope of the current data while highlighting clear directions for future work.

      Reviewer #2 (Public review):

      Summary:

      Argunsah and colleagues demonstrate that SST expressing interneurons are concentrated in the mouse septa and differentially respond to repetitive multi-whisker inputs. Identifying how a specific neuronal phenotype impacts responses is an advance.

      Strengths:

      (1) Careful physiological and imaging studies.

      (2) Novel result showing the role of SST+ neurons in shaping responses.

      (3) Good use of a knockout animal to further the main hypothesis.

      (4) Clear analytical techniques.

      Comments on revisions:

      The authors have effectively responded to my initial critiques - I have no further concerns.

      We thank the reviewer for their positive evaluation of our work and for recognizing the novelty of the findings, the careful physiological and imaging approaches, the use of the Elfn1 knockout model, and the clarity of the analytical framework. We are pleased that the reviewer has no further concerns and appreciates the contribution of this study to understanding the role of SST<sup>+</sup> interneurons in shaping sensory processing in the barrel cortex.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the functional differences between barrel and septal columns in the mouse somatosensory cortex, focusing on how local inhibitory dynamics (particularly involving SST<sup>+</sup> interneurons) may mediate temporal integration of multi- whisker (MW) stimuli in septa. Using a combination of in vivo multi-unit recordings, calcium imaging, and anatomical tracing, the authors propose a model in which Elfn1-dependent synaptic facilitation onto SST<sup>+</sup> interneurons contributes to the distinct sensory responses to MW input in barrels and septa, enabling functional segregation between these domains.

      Strengths:

      The study presents a thought-provoking and useful conceptual model for understanding sensory processing in the somatosensory cortex. While barrel columns have been widely studied, septal regions remain relatively understudied in mice. If septa indeed act as selective integrators of distributed sensory input, this would suggest a novel computational role for cortical microcircuits beyond the classical view focused on barrels. Although still hypothetical, the proposed model in which SST<sup>+</sup> interneurons contribute to domain-specific sensory responses between barrel and septal domains is intriguing and opens new avenues for investigating inhibitory circuit mechanisms.

      Weaknesses:

      The primary limitation of this study lies in the spatial and cellular specificity of the recording techniques. The physiological data rely predominantly on unsorted multi-unit activity (MUA) recorded with lowchannel-count silicon probes. Because MUA aggregates signals from multiple neurons over a radius of approximately 50-100 µm (often wider than the typical septal width in mice), this approach makes it difficult to confidently isolate activity originating strictly from within septal domains. The manuscript would benefit from additional analyses to validate the spatial specificity of these recordings, such as systematically varying spike detection thresholds to test the robustness of domain attribution, as suggested by the reviewer. Furthermore, although the authors now appropriately frame their findings in the Elfn1 knockout mice as indirect evidence, it is worth emphasizing that the study lacks direct in vivo, cell-type-specific recordings and manipulations to more definitively test the proposed mechanism.

      We thank the reviewer for their thorough and constructive evaluation of the manuscript and for highlighting both the conceptual strengths of the study and its technical limitations. We agree that the spatial and cellular specificity of unsorted multi-unit recordings imposes inherent constraints on the interpretation of domain-specific activity, particularly given the narrow width of septal compartments in mice. As now clarified in the manuscript, we do not claim absolute cellular specificity of “septal” recordings but rather interpret them as septal-enriched populations. To directly address this concern, we performed additional threshold-based analysis demonstrating that the key domain-specific effects persist selectively in Layer 4 under stricter spike-detection criteria, supporting a local circuit origin of the critical findings. Further, the more stringent detection criteria (Suppl Fig 3A) collapse the divergence seen in Layer2/3 (Suppl Fig 4C), suggesting that this divergence arises in Layer 4, where SST+ interneuron distributions diverge between barrel and septa.

      We further agree that the Elfn1 knockout results provide indirect, rather than definitive, evidence for causal involvement of SST<sup>+</sup> interneurons and therefore intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, physiological, developmental, and genetic observations. We believe this explicitly moderated interpretation appropriately reflects the scope of the current data while establishing a clear conceptual framework and motivation for future studies employing cell-type-specific recordings and manipulations to directly test the proposed mechanism.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Major comments

      (1) Interpretation of "septal" recordings: The authors claim that the activity recorded from electrodes placed in the septa can be confidently attributed to septal neurons. In my previous review, I raised a major concern that such "septal" recordings likely include spikes from adjacent barrels, given the broad spatial resolution of MUA and the narrowness of the septa in the mouse S1. In fact, the intermediate properties observed in septal recordings from wild-type mice could be explained by a mixture of activity from principal and neighboring barrels-an interpretation that contrasts with the authors' conclusion. Upon reviewing the probe model used (A8x8-Edge-5mm-100-200-177), I noticed a discrepancy between the manufacturer's design and the schematic provided in the manuscript. The electrodes are located near the right edge of the probe rather than the center, suggesting that neurons in adjacent barrels could easily be sampled. In my previous review, I therefore suggested alternative approaches, such as calcium imaging, to more convincingly support the authors' claims. However, the revised manuscript does not include new experiments or additional analyses addressing this issue. Instead, the authors argue that using a high spike detection threshold (SD > 7.5) ensures that recorded activity originates from septal neurons, even though this value does not appear particularly conservative, as it was merely adopted from a previous study without justification in the present context. While I agree that a higher threshold may reduce contamination from distant sources, it does not guarantee that only septal neurons contribute to the signal. By nature, MUA reflects activity from multiple neurons within a radius of at least 50-100 µm. To more rigorously support the claim of spatial specificity, I strongly encourage the authors to reanalyze their existing dataset by systematically varying the spike detection threshold and quantifying how the properties and selectivity of detected units change. If neurons closer to the electrode indeed exhibit distinct domain-specific properties, they should become more prominent as the threshold increases. Such an analysis would strengthen the authors' interpretation and improve the manuscript's impact, even in the absence of new experimental data. Alternatively, the authors could revise their claims to acknowledge that the "septal" electrodes likely record from a population that includes septal neurons as well as neurons located at the periphery of principal and adjacent barrels.

      We agree with the reviewer that, by nature, MUA reflects the activity of multiple neurons within a spatial radius and that recordings obtained from electrodes positioned in the septa may include contributions from neurons located at the periphery of adjacent barrels. This concern is further compounded in superficial layers by probe geometry and orientation: given the narrow width of septa and the lateral spread of processes in upper cortical layers, recordings in L2/3 are inherently more susceptible to spatial mixing than those in layer 4, where columns are more compact and cytoarchitecturally distinct. To directly address these issues, we reanalyzed the same dataset using a more stringent spike detection threshold (SD > 9.5), compared to the originally reported SD > 7.5. Importantly, increasing the threshold selectively reduced or eliminated effects in L2/3, while the key domain-specific differences in L4 responses both the differential MW/SW dynamics in wild-type animals and their attenuation in Elfn1 knockout mice remained robust (the new Supp. Fig. 3. In the manuscript). This threshold-dependent dissociation is consistent with the interpretation that the critical effects reported in L4 arise from neurons spatially closer to the electrode and are less influenced by probe orientation or distant sources, rather than reflecting simple mixing of barrel signals. While this analysis does not claim absolute cellular exclusivity of septal neurons, it provides empirical support that the principal conclusions of the study are robust to stricter spatial sampling criteria and are particularly anchored in L4 circuitry. Accordingly, we now explicitly acknowledge in the manuscript that “septal” recordings likely represent septal-enriched populations rather than purely septal neurons, while emphasizing that the persistence of L4 effects under higher spike-detection thresholds strengthens the conclusion that local L4 inhibitory dynamics underlie the reported functional differences between barrel and septal domains.

      The greater sensitivity of L2/3 results to spike-detection threshold is also expected based on both anatomical considerations and probe geometry. Neurons in L2/3 possess broader horizontal dendritic and axonal arbors and participate in more laterally distributed integration across columns, making population signals in these layers intrinsically less spatially focal. As a result, conservative spike-detection criteria preferentially suppress L2/3 effects, particularly when recordings are obtained with probes optimized for deeper layers. Importantly, our two-photon calcium imaging data while similarly limited to L2/3 demonstrate that SST<sup>+</sup> interneurons show locally measurable and stimulus-specific responses at the single-cell level, providing independent support that L2/3 SST<sup>+</sup> activity is stimulus-modulated rather than artifactual. Taken together, these observations suggest that L2/3 results reflect more distributed and integrative network activity, whereas the L4 effects that persist across thresholds are more directly attributable to local circuit mechanisms. This layer-specific dissociation further supports our interpretation that the central findings of the study are driven by local inhibitory dynamics in L4, with L2/3 activity reflecting downstream integration rather than primary domain-specific computation.

      (2) Interpretation of the Elfn1 KO data: The authors' interpretation that Elfn1-dependent facilitation of SST<sup>+</sup> interneurons underlies the differential sensory responses between barrel and septal domains is conceptually appealing and supported by several converging, albeit indirect, lines of evidence. Specifically, the consistent correspondence among the differential activation of SST<sup>+</sup> neurons upon SWS and MWS, the late development of the barrel-septa differences in the responses to SWS and MWS, and the attenuation of this difference in Elfn1 knockout mice lends plausibility to the proposed model. However, it should be emphasized that the data remain indirect: the study does not include direct recordings of SST<sup>+</sup> neuronal activity from the knockout mice, nor cell-type- specific manipulations to demonstrate causal involvement. The mechanistic explanation therefore represents a hypothesis rather than definitive proof. That said, the authors clearly acknowledge these limitations in the Discussion and appropriately moderate their claims by presenting the SST-Elfn1 mechanism as a working model. Given this careful framing, the current manuscript can be regarded as a valuable conceptual contribution that advances our understanding of how inhibitory dynamics may shape temporal processing in the barrel cortex. Further experiments, as mentioned above, will be essential to test the causal role of this mechanism directly.

      We thank the reviewer for this thoughtful and balanced assessment. We fully agree that the Elfn1 knockout experiments provide indirect rather than definitive evidence for a causal role of SST<sup>+</sup> interneurons in mediating the domain-specific MW/SW response dynamics between barrels and septa For this reason, throughout the revised manuscript we explicitly frame the Elfn1–SST mechanism as a working model rather than a proven mechanism.

      Minor comments:

      The authors have adequately addressed my previous minor comments. In this round, I carefully reviewed the revised manuscript and identified several issues related to references. I would also like to add a brief comment regarding the Discussion section:

      (1) Stachniak et al., 2021 is included in the reference list but is not cited anywhere in the main text. Please either remove this entry or cite it appropriately in the manuscript.

      Removed.

      (2) Yamashita et al., 2018 is cited in the main text (Line 767), but it is not included in the reference list.

      Fixed.

      (3) Sylwestrak and Ghosh, 2012 is cited at Line 261 and Line 270, but likewise absent from the reference list.

      Fixed.

      (4) At Line 497, Chen et al., 2015 is cited, but, the appropriate and original reference would be Chen et al., 2013 (PMID: 23792559), which should either replace or precede the 2015 citation.

      Added.

      (5) At Line 221, El-Boustani et al., 2018 is cited. However, this study is based on the visual cortex, whereas the manuscript concerns the barrel cortex. A more relevant citation (e.g., Lefort et al., 2009 [PMID: 19186171]) would better support the discussion of cellular organization in the barrel cortex. Please consider updating the citation.

      Thank you for this suggestion. We agree with the reviewer and now we have changed El-Boustani with Lefort et al. 2009 as suggested by the reviewer.

      (6) Furthermore, Chakrabarti & Alloway (2006) performed tracer-based mapping of projections from barrel and septal columns in rat S1 and similarly suggested differential organization of M1- and S2projection neurons in the barrel and septal regions.

      Although the current study thoroughly analyzed the layer-specificity of the location of these projection neurons, the lack of explicit discussion of this relevant prior work is a notable omission.

      The authors should incorporate a comparison with these results to better contextualize their findings.

      The following text is added to the discussion: “Our retrograde labeling data supports and expands on previous work proposing similar models (Alloway, 2008; Chakrabarti and Alloway, 2006).”

    1. eLife Assessment

      This potentially valuable study aims to investigate neural correlates of spatial attention in whisker somatosensory cortex (S1) in mice, finding increased sensory-evoked spiking when the mice appear to be attending to the contralateral whiskers. Although some of the results appear to be robust despite relatively small effect sizes, overall the findings are incompletely supported, because attentional modulation is insufficiently distinguished from learning of stimulus-response contingencies, and because the analyses do not adequately consider orofacial movements that may contribute key confounds.

    2. Reviewer #1 (Public review):

      The paper uses a passive whisker detection task in mice to identify a behavioral phenomenon that can reasonably be interpreted as spatial attentional capture. The attentional effect occurs transiently after a successful whisker stimulus detection yields reward, and lasts for a few trials before subsiding. The attentional effect is to the right or left whiskers, depending on whether right or left whiskers are rewarded; no finer spatial resolution for attention was tested. By recording whisker-evoked spiking from single units in S1, the authors show that this form of spatial attention increases the gain of whisker-evoked neuronal responses in S1 for a large subset of S1 units. In contrast, neural responses are not modulated by overall task engagement. Together, these findings show a neural signature of spatial attention in S1 cortex. Because whisker or facial movements were not tracked, it is not clear whether this represents covert attention or whisker movement in response to previously rewarded stimuli, which would be a form of overt attention.

      Substantial attentional modulation of neural responses was observed for a subset of whisker-responsive S1 units, but the effect size was small on average for the total unit population. The top 25% of units showed a ~12% attentional response modulation (relative to firing rate range for each unit), but the median unit showed only a 1.3% response modulation. It would have been useful to analyze the magnitude or prevalence of attentional modulation across layers or in fast-spiking vs. regular spiking units, but this was not reported.

      Major

      (1) It is hard to interpret the underlying causes of the attentional modulation of neural activity without having measured whisker and facial movement. This is a particular issue in S1, where whisker movement against the stimulation grid can alter the mechanical efficiency of stimulus delivery. Such movements would represent overt attention, which would engage an entirely different neural mechanism than covert attention.

      (2) An interesting debate is whether the behavioral phenomenon is best described as attention or as dynamic learning of the stimulus-response association for that block. In Posner-type cued attention tasks, and also in many block-type attention tasks in rodents, animals receive reward for successfully detecting either cued or uncued stimuli, and thus attention (higher response probability or improved psychometric sensitivity for cued stimuli) is at least partially dissociated from the stimulus-reward contingency. That is not the case here. The fact that mice have difficulty learning the contingency reversal suggests that the phenomenon is better explained by attention than by learning the contingency; however, to prove this clearly, the existence of the attentional effect on neural activity in Block 1 vs. Block 2 would have to be shown.

      (3) Some of the graphical representations of the attentional modulation of neural activity are unclear. The single-unit example of attentional modulation is quite strong (Figure 3d). The mean response for the top 25% of units is also visually clear (Figure 3f). But the effect is not apparent at all in Figure 3e, which the figure legend says shows every unit. What is the yellow point and line in this figure? Why isn't the attentional effect visible in this panel? Perhaps I am misunderstanding Figure 3e, but it is not clear to me why it compares Pref>0.5 to Pref<0.5, when the intended analysis suggests it should be Pref>0 to Pref<0? Also in Figure 3, it is critical for the reader to know whether panels 3g-3h represent the top 25% of units or all units. Neither the results text nor the legend is clear on this.

      (4) There is a missed opportunity to quantify attentional modulation across cortical layers, since laminar probes and Neuropixels probes were used for the recordings. In addition, there is no separation of fast-spiking from regular-spiking units, and no quantitative metrics are provided to assess the quality of single units. This could reveal key aspects of cortical processing of attentional signals.

    3. Reviewer #2 (Public review):

      Summary:

      Dyce et al investigate the modulation of sensory responses in the somatosensory 'barrel' cortex during a novel whisker vibration detection task in head-fixed mice, aiming to find correlates of spatial attention in both the animals' behavior and their neuronal activity.

      Strengths:

      The authors produced an extensive and parameterized dataset of both behavioral responses and neuronal activity, with >3000 single units of which >1400 were responsive.

      Weaknesses:

      In my view, the main conclusions of the manuscript are not currently well supported by the data.

      The authors effectively define "spatial attention" as a state where an animal responds more to a stimulus that gives more rewards (out of two possible stimuli presented on different sides of the snout, i.e., segregated spatially). If one defines spatial attention purely in these terms, then their findings do show neuronal correlates of spatial attention. However, those neuronal correlates can be explained by known aspects of neuronal responses in the barrel cortex.

      This plays out in several different ways:

      From the behavioral point of view, greater attention may correlate with an increased hit rate to stimuli on the rewarded side, but in the absence of other supporting measurements, the relationship could well be the opposite: an animal could pay more (rather than less) attention to the stimulus delivered on the unrewarded side, to make sure it suppresses the incorrect response. It is impossible to tell, as the data don't provide an independent measurement of whether the animal is paying greater attention to, or is more aware of, one side than the other, nor do they provide an independent measurement of neuronal tuning on either side. There is no separate measurement of arousal either (e.g., via pupillometry or locomotion).

      The experimental design involved two blocks on each daily task session, with the second block reversing the side on which rewarded stimuli were delivered. Reinforcing one's doubts about the behavior and its interpretation, mice had much poorer performance on each day's second block, to the extent that perceptual sensitivity (d') was the same for both sides: d' did not increase after reward reversal for stimuli on the initially unrewarded side. This further emphasizes the lack of a separate demonstration of focused "spatial attention".

      Much of the data (both behavioral and neuronal) could be accounted for, e.g., by a strategy where the mouse keeps a token in working memory of what side seems to be driving rewards, while maintaining equally strong sensory drive on both sides, but with no attentional shift at all. The policy would be to respond more whenever the stimulated side matches the token in memory (thus also reinforcing the token, thus enhancing performance next time). This would be easily implemented with a disinhibitory reward-modulation signal such as the one multiple researchers have found carried by VIP neurons (e.g., Szadai et al DOI: 10.7554/eLife.78815).

      Similarly, the fact that "attended trials" (Pref > 0) produced greater responses than "unattended trials" appears to be explainable as follows. Here, "attended" trials are those where the contralateral stimulus is presented (and, if responded to, is rewarded), "unattended" trials are those where the stimulus is ipsilateral (and not rewarded). The animal responds more (at least in the first block) to stimuli delivered to the contralateral pad - i.e., rewarded as opposed to unrewarded ones. Beyond the knowledge mentioned above that cortex-wide VIP sensitivity to rewards can drive disinhibition in general, activity modulation dependent on rewards and outcomes (and stimulus value) has been established specifically in the barrel cortex (e.g., Lacefield et al DOI: 10.1016/j.celrep.2019.01.093, Bale et al DOI: 10.1016/j.cub.2020.10.059, Banerjee et al DOI: 10.1038/s41586-020-2704-z, Chereau et al 10.1038/s41467-020-17005-x). The reward- and value-evoked activity demonstrated in those papers would suffice to predict more activity at the contralateral electrode on "attended" trials, along the lines of the findings in Ramamurthy et al (DOI: 10.1038/s41467-025-60592-w) and consistent also with the enhanced "attentional modulation" on hit trials.

      Other aspects of the analysis and terminology lead to confusing outcomes. For example, in the analysis in Figure 3, Performance averaged in a set of trials around a given trial is defined as the mean rate of responses to stimulation on either side - regardless of whether those responses are correct (since the stimuli can be on either side, but only one side is correct and gets rewarded and putatively reinforced). Thus, this definition of "Performance" can increase with the rate of incorrect licks to the wrong side and is at odds with the normal use of the word. On trials where this Perf = 1 and the stimuli are balanced on either side, this corresponds to a true performance (and reward rate) of only 0.5 - what one would normally consider random discrimination between the sides. Thus, Perf = 1 trials may still give a low reward rate and, if responses scale with reward, a small effect of reward. Hence, based on known properties of reward dependence, greater correlation of neuronal activity with "Preference" than with "Performance" would be expected, rather than reflecting a new aspect of "spatial attention". A definition of performance more in line with established practice and measuring side-to-side discrimination (corresponding more closely to the authors' "Preference" parameter) would have shown this more clearly.

    4. Author response:

      (1) Introduction & Roadmap

      We are grateful to the Reviewers for engaging with outstanding questions relating to our findings’ connections to multiple subdisciplines of cognitive neuroscience. Noting that Reviewers 1 and 2 interpreted our findings differently, we welcome the opportunity to engage in what Reviewer 1 characterised as “an interesting debate”. To promote a shared understanding and discussion of our findings, we have organised our response to address more technical comments first.

      Our provisional response is organised as follows: Section 2 addresses selected technical comments relating to our Results. Section 3 addresses comments related to the design of our behavioural paradigm. Section 4 focuses on the broader interpretation of our findings. Section 5 concludes our provisional response with potential future directions and a summary of the significance of our findings.

      (2) Selected technical comments related to our Results

      We apologise to Reviewer 2 for the confusion in relation to the meaning of “attended” and “unattended” trials. What we said was “Positive Pref values indicate a higher response rate to the contralateral side than the ipsilateral side (relative to the electrode)” (Figure 3c caption), “we indexed all contralateral whisker vibrations according to their associated Perf and Pref” (Results text), and “we divided trials into (contralaterally) attended (Pref<sub>C/L</sub>: Pref>0) and unattended (Pref<sub>I/L</sub>: Pref<0) groups” (Results text). We can confirm that we defined an “unattended trial” (Pref<0) as a contralateral stimulus trial in the centre of an epoch (10-15 trials) within which the mouse responded (licked) more frequently to ipsilateral stimuli. Critically, we did not define an unattended trial as an ipsilateral stimulus trial. Furthermore, attention thus defined (i.e. Pref>0) can vary independently of the whisker stimulus associated with rewards. Indeed, while we initially did not include this result in our paper for the sake of brevity, even unrewarded “attended” trials (Pref>0) evoked significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials. We note that this is an analysis suggested by Reviewer 1, and we will include and discuss this result in our revised manuscript (e.g. in relation to literature suggested by Reviewer 2). For additional clarity, we use “Performance” (Perf) in relation to overall stimulus detection, consistent with the analysis of Lee et al. (2020), which found this measure was correlated with pupil diameter in a vibrissal target detection task.

      We thank Reviewer 1 for noticing that the axes on Figure 3e should be labelled “Pref>0” (Y axis) and “Pref<0” (X axis), as suggested by the figure caption. We will correct this in our revised submission. The yellow point on Fig 3e shows the unit from Fig 3d, while the yellow line in Fig 3e shows the magnitude of that unit’s (non-normalised) gain modulation. While this is alluded to in the Results text (“The example unit in Figure 3d is in the 93rd percentile of units for raw modulation depth (ΔHits(attended – unattended) = 3.3 spikes/second; yellow line in Fig.3e)”, this should be explained in the Figure caption, and it will be in our revised manuscript. We would also like to clarify that Figures 3g–3h display results for all units, not just the top 25%. We agree this is not sufficiently clear and we will rectify this in our revised manuscript. Addressing Reviewer 2, while we acknowledge that mice responded less to both stimuli in the second block, they also meaningfully adjusted their behaviour to the reversal in reward contingencies: their responses to the previously rewarded stimulus reduced significantly more than those to the previously unrewarded stimulus.

      (3) Design of the behavioural paradigm

      We made a deliberate design choice to maximise the ecological validity of our behavioural paradigm, and note that there are advantages to doing so. For example, our paradigm can be used to show that even unrewarded “attended” trials (Pref>0) evoke significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials (see Section 2, above). Indeed, it is precisely this finding that makes our paradigm uniquely suited to the investigation of value-driven attentional capture (Anderson et al., 2011): in this instance attention directed to stimuli that are no longer rewarded despite equal availability of rewarded stimuli. This finding also demonstrates that our paradigm dissociates attention from stimulus-reward contingency at least as well as other paradigms which have been successfully used to study spatial attention in mice. As noted in Section 2, we will discuss this result in relation to other relevant research (e.g. Ramamurthy et al., 2025) in our revised manuscript.

      Briefly, the direct manipulation of reward contingencies is one of two noteworthy methodological distinctions between our own paradigm and that of Ramamurthy and colleagues (2025). The task of Ramamurthy et al. (2025) associated all whisker stimuli with rewards and delivered stimuli to different whiskers on a single whisker pad. These methodological distinctions may have reduced the relevance of the spatial differences between stimuli to the mice undertaking the task. Indeed, it is not certain that a mouse would treat the unilateral variation in whisker stimulation Ramamurthy and colleagues delivered as primarily spatial or featural. The psychophysical and neural differences between spatial and featural attention in humans suggest dissociable underlying mechanisms, and the same may be true in mice. Thus, our own paradigm may more effectively isolate spatial attention from featural attention. Conversely, to the extent that the findings of Ramamurthy and colleagues do reflect spatial attention, our combined findings and paradigms help elucidate the associated mechanisms across spatial scales in mice.

      We acknowledge that spatial cueing is well-suited to isolating the effects of covert attention from other forms of attention. However, it should be noted that spatial cueing in rodents is subject to its own challenges, including limitations in trial numbers due to the required manipulation of stimulus intensity (Reynolds et al., 2000; Herrmann et al., 2010), cue validity and associated trial probabilities (Peterson & Gibson, 2011; Girardi et al., 2013). Such experiments are further complicated by the duration and efficacy of training (i.e. the number of mice that learn the task; Wang & Krauzlis, 2018; Hu & Dan, 2022). It is also worth noting that trial probability manipulations introduce the same limitation in trial numbers with block-type attention tasks (You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022).

      While there are clear differences between our own paradigm and those mentioned above, there are also important similarities. First, these tasks are all goal-directed, stimulus-driven, and reliant on learned task contingencies (e.g. Peterson & Gibson, 2011; Girardi et al., 2013). Furthermore, these paradigms are all operant conditioning protocols which leverage learned stimulus-reward contingencies to train attention-related behaviours in mice. A noteworthy similarity between our findings and those of authors using block-type attention tasks in particular (e.g. You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022) is the observation of apparent attentional biases in behavioural responses independent of the experimental manipulations (i.e. stimulus probability / reward contingency).

      (4) Comments relating to the broader interpretation and discussion of our findings

      Fundamentally, attention involves dedicating limited processing resources to some stimulus events at the expense of others. The design of our behavioural paradigm was informed by existing literature on spatial attention in humans, non-human primates, and mice. Our choice of behavioural and neuronal measures as proxies for attention in mice is consistent with this literature. It is technically possible “an animal could pay ‘more’ (rather than less) attention to the stimulus delivered on the unrewarded side, to make sure it suppresses the incorrect response”, but this seems unlikely given what is known about how attention is typically allocated in such tasks, based on the previously mentioned literature.

      With respect to the interpretation and discussion of our findings, Reviewer 1 describes them as “a behavioral phenomenon that can reasonably be interpreted as spatial attentional capture” but suggests they do not clearly distinguish whether this attentional capture is covert or overt. We respectfully disagree for three reasons. First, as discussed in our paper, whisker motion during detection tasks has consistently been associated with reduced detection performance (Ollerenshaw et al., 2012; Kyriakatos et al., 2017; Vandevelde et al., 2023), suggesting that a “receptive” strategy (Diamond & Arabzadeh, 2013) of whisker immobilisation is more applicable to the current data than a “generative” strategy of asymmetric whisker movement (O'Connor et al., 2010; Dominiak et al., 2019). Second, if our behavioural and neuronal findings were due to the mice moving their whiskers to maximise contact with the meshes, we would expect increased evoked neuronal responses to be associated with greater Perf, not just with greater Pref. This pattern was not observed. Of course, the mice might have employed different whisker movement strategies during epochs of high Pref and Perf, but this seems unlikely and is not a parsimonious explanation for our findings. Third, as noted in the Methods section of the paper, we deliberately positioned the meshes close to the base of the whiskers, limiting the impact of whisker movements on stimulus detectability and the incentive to make them.

      In contrast, Reviewer 2 questions the interpretation of our findings as evidence of spatial attention and suggests they might reflect working memory instead. Current research suggests attention and working memory are intimately related integrative brain functions. Indeed, some researchers have even proposed that working memory might be a form of internally directed attention (Awh & Jonides, 2001; Chun, 2011; Gazzaley & Nobre, 2012; Kiyonaga & Egner, 2013; or vice versa: Libedinsky & Fernandez, 2019). Consistent with the comments of Reviewer 2, more recent work seems to emphasise the coordination of attention and working memory (e.g. Joe & Kim, 2023; Zhu et al., 2026; for reviews see Huynh Cong & Kerzel, 2021; van Ede & Nobre, 2023), along with shared mechanisms (Kiyonaga et al., 2021; Panichello & Buschman, 2021), and nuanced dissociations (Liu et al., 2025). Attention is difficult to dissociate from working memory partly because there are multiple definitions (and/or types) of attention. We did not discuss the various definitions and/or forms of attention at length in our paper, but we will briefly discuss this in the revised manuscript.

      The “interesting debate” to which Reviewer 1 refers could also be described as vigorous, despite approximately three decades of research. This debate broadly relates to the degree to which attentional control is driven by exogenous (e.g. colour contrast) versus endogenous factors (e.g. the focus of spatial attention, see Fig.2 in Belopolsky et al., 2007; see also: Liesefeld & Mueller, 2020; Manini et al., 2021; Beffara et al., 2022), and the degree to which this is a function of experimental context. The review article by Luck et al. (2021) entitled “Progress toward resolving the attentional capture debate” provides a striking illustration of this debate, as do the twenty-two commentaries (and three commentary responses) associated with it. Admittedly, this debate largely revolves around human attention experiments, and human cognition may be more complex than mouse cognition. However, the complexity of human cognition may also be easier to study and appreciate because complex behavioural experiments can be explained to, understood, and performed by human participants with relative ease.

      (5) Comments relating to future directions and the significance of our findings

      The complexity of the attentional capture debate underscores the importance of developing accessible and scalable animal experiments which can be used to provide mechanistic insights. If the human attention literature is any indication, a diversity of rodent experimental paradigms will be necessary to thoroughly map the neuronal implementation of spatial attention. Returning to our paradigm, Reviewer 1 noted that valuable insights into the mechanisms of vibrissal spatial attention might be obtained from comparing the magnitude of attentional modulation we observed between putative regular and fast-spiking categories of units, and between units located in different cortical layers. We agree it is important to understand spatial attention with cell-type and circuit (including laminar) specificity. However, because we could not persuasively cluster our units based on waveform width, and because of the lack of histological data, segregating units on the basis of such variables is not feasible. Despite our assertion that our findings reflect the effects of covert attention (contra Reviewer 1), we agree that future experiments will be required to conclusively rule out overt attention. Noting the proximity of the meshes to the base of the whiskers in our paradigm, and the difficulty of tracking whiskers in this context, Botulinum toxin injections (as in Ramamurthy et al., 2025) might be a means of achieving this.

      The above notwithstanding, our findings provide multiple contributions to the literature on spatial attention (and perhaps working memory). We detected significant attentional gain modulation across a population of 1461 responsive units. While the gain modulation exhibited by the median unit was modest (albeit statistically significant), the top 25% of responsive units showed a ~12% response modulation (relative to firing rate range for each unit), and ~21% of responsive units were suppressed by the average vibrissal stimulus in the unattended state. Our experimental framework offers an accessible platform for future studies leveraging genetic and circuit-level interventions to dissect the cell-type specific mechanisms of spatial attention. Our work is timely, noting the recent focus of human research on the nexus of attention, selection history, and valence (e.g. Serences, 2008; Della Libera & Chelazzi, 2009; Della Libera et al., 2011; van den Berg et al., 2014; Kim & Anderson, 2019, 2023). Our work is also uniquely poised to stimulate new interdisciplinary research into the circuit mechanisms of value-driven attentional capture, with translational relevance to psychopathologies such as ADHD, addiction, and depression; where value-driven attentional capture is altered (for a review see Anderson, 2021).

      References

      Anderson, B. A. (2021). Relating value-driven attention to psychopathology. Curr Opin Psychol, 39, 48-54. https://doi.org/10.1016/j.copsyc.2020.07.010

      Anderson, B. A., Laurent, P. A., & Yantis, S. (2011). Value-driven attentional capture. Proceedings of the National Academy of Sciences of the United States of America, 108(25), 10367-10371. https://doi.org/10.1073/pnas.1104047108

      Awh, E., & Jonides, J. (2001). Overlapping mechanisms of attention and spatial working memory. Trends Cogn Sci, 5(3), 119-126. https://doi.org/10.1016/s1364-6613(00)01593-x

      Beffara, B., Hadj-Bouziane, F., Ben Hamed, S., Boehler, C. N., Chelazzi, L., Santandrea, E., & Macaluso, E. (2022). Dynamic causal interactions between occipital and parietal cortex explain how endogenous spatial attention and stimulus-driven salience jointly shape the distribution of processing priorities in 2D visual space. Neuroimage, 255. https://doi.org/10.1016/j.neuroimage.2022.119206

      Belopolsky, A. V., Zwaan, L., Theeuwes, J., & Kramer, A. F. (2007). The size of an attentional window modulates attentional capture by color singletons. Psychonomic Bulletin & Review, 14(5), 934-938. https://doi.org/10.3758/Bf03194124

      Chun, M. M. (2011). Visual working memory as visual attention sustained internally over time. Neuropsychologia, 49(6), 1407-1409. https://doi.org/10.1016/j.neuropsychologia.2011.01.029

      Della Libera, C., & Chelazzi, L. (2009). Learning to Attend and to Ignore Is a Matter of Gains and Losses. Psychological Science, 20(6), 778-784. https://doi.org/10.1111/j.1467-9280.2009.02360.x

      Della Libera, C., Perlato, A., & Chelazzi, L. (2011). Dissociable Effects of Reward on Attentional Learning: From Passive Associations to Active Monitoring. PLoS One, 6(4). https://doi.org/10.1371/journal.pone.0019460

      Diamond, M. E., & Arabzadeh, E. (2013). Whisker sensory system - from receptor to decision. Prog Neurobiol, 103, 28-40. https://doi.org/10.1016/j.pneurobio.2012.05.013

      Dominiak, S. E., Nashaat, M. A., Sehara, K., Oraby, H., Larkum, M. E., & Sachdev, R. N. S. (2019). Whisking Asymmetry Signals Motor Preparation and the Behavioral State of Mice. J Neurosci, 39(49), 9818-9830. https://doi.org/10.1523/JNEUROSCI.1809-19.2019

      Gazzaley, A., & Nobre, A. C. (2012). Top-down modulation: bridging selective attention and working memory. Trends Cogn Sci, 16(2), 129-135. https://doi.org/10.1016/j.tics.2011.11.014

      Girardi, G., Antonucci, G., & Nico, D. (2013). Cueing spatial attention through timing and probability. Cortex, 49(1), 211-221. https://doi.org/10.1016/j.cortex.2011.08.010

      Herrmann, K., Montaser-Kouhsari, L., Carrasco, M., & Heeger, D. J. (2010). When size matters: attention affects performance by contrast or response gain. Nat Neurosci, 13(12), 1554-1559. https://doi.org/10.1038/nn.2669

      Hu, F., & Dan, Y. (2022). An inferior-superior colliculus circuit controls auditory cue-directed visual spatial attention. Neuron, 110(1), 109-119 e103. https://doi.org/10.1016/j.neuron.2021.10.004

      Huynh Cong, S., & Kerzel, D. (2021). Allocation of resources in working memory: Theoretical and empirical implications for visual search. Psychon Bull Rev, 28(4), 1093-1111. https://doi.org/10.3758/s13423-021-01881-5

      Joe, J., & Kim, M. S. (2023). Spatial Attention in Visual Working Memory Strengthens Feature-Location Binding. Vision (Basel), 7(4). https://doi.org/10.3390/vision7040079

      Kanamori, T., & Mrsic-Flogel, T. D. (2022). Independent response modulation of visual cortical neurons by attentional and behavioral states. Neuron, 110(23), 3907-3918 e3906. https://doi.org/10.1016/j.neuron.2022.08.028

      Kim, H., & Anderson, B. A. (2019). Dissociable neural mechanisms underlie value-driven and selection-driven attentional capture. Brain Research, 1708, 109-115. https://doi.org/10.1016/j.brainres.2018.11.026

      Kim, H., & Anderson, B. A. (2023). Primary Rewards and Aversive Outcomes Have Comparable Effects on Attentional Bias. Behavioral Neuroscience, 137(2), 89-94. https://doi.org/10.1037/bne0000543

      Kiyonaga, A., & Egner, T. (2013). Working memory as internal attention: toward an integrative account of internal and external selection processes. Psychon Bull Rev, 20(2), 228-242. https://doi.org/10.3758/s13423-012-0359-y

      Kiyonaga, A., Powers, J. P., Chiu, Y. C., & Egner, T. (2021). Hemisphere-specific Parietal Contributions to the Interplay between Working Memory and Attention. J Cogn Neurosci, 33(8), 1428-1441. https://doi.org/10.1162/jocn_a_01740

      Kyriakatos, A., Sadashivaiah, V., Zhang, Y., Motta, A., Auffret, M., & Petersen, C. C. (2017). Voltage-sensitive dye imaging of mouse neocortex during a whisker detection task. Neurophotonics, 4(3), 031204. https://doi.org/10.1117/1.NPh.4.3.031204

      Lee, C. C. Y., Kheradpezhouh, E., Diamond, M. E., & Arabzadeh, E. (2020). State-Dependent Changes in Perception and Coding in the Mouse Somatosensory Cortex. Cell Rep, 32(13), 108197. https://doi.org/10.1016/j.celrep.2020.108197

      Libedinsky, C. D., & Fernandez, P. F. (2019). Graded Memory: A Cognitive Category to Replace Spatial Sustained Attention and Working Memory
 Yale J Biol Med, 92(1), 121-125. https://www.ncbi.nlm.nih.gov/pubmed/30923479

      Liesefeld, H. R., & Mueller, H. J. (2020). A theoretical attempt to revive the serial/parallel-search dichotomy. Attention Perception & Psychophysics, 82(1), 228-245. https://doi.org/10.3758/s13414-019-01819-z

      Liu, Y., Fu, Y., Tang, E., Wu, H., Han, J., Xie, M., Zhang, Y., Peng, B., Huang, J., Liu, H., Chen, H., & Qin, P. (2025). Neural dissociation of attention and working memory through inhibitory control. Nat Commun, 17(1), 22. https://doi.org/10.1038/s41467-025-66553-7

      Luck, S. J., Gaspelin, N., Folk, C. L., Remington, R. W., & Theeuwes, J. (2021). Progress toward resolving the attentional capture debate. Visual Cognition, 29(1), 1-21. https://doi.org/10.1080/13506285.2020.1848949

      Manini, G., Botta, F., Martin-Arevalo, E., Ferrari, V., & Lupianez, J. (2021). Attentional Capture From Inside vs. Outside the Attentional Focus. Frontiers in Psychology, 12. https://doi.org/10.3389/fpsyg.2021.758747

      O'Connor, D. H., Clack, N. G., Huber, D., Komiyama, T., Myers, E. W., & Svoboda, K. (2010). Vibrissa-based object localization in head-fixed mice. J Neurosci, 30(5), 1947-1967. https://doi.org/10.1523/JNEUROSCI.3762-09.2010

      Ollerenshaw, D. R., Bari, B. A., Millard, D. C., Orr, L. E., Wang, Q., & Stanley, G. B. (2012). Detection of tactile inputs in the rat vibrissa pathway. J Neurophysiol, 108(2), 479-490. https://doi.org/10.1152/jn.00004.2012

      Panichello, M. F., & Buschman, T. J. (2021). Shared mechanisms underlie the control of working memory and attention. Nature, 592(7855), 601-605. https://doi.org/10.1038/s41586-021-03390-w

      Peterson, S. A., & Gibson, T. N. (2011). Implicit attentional orienting in a target detection task with central cues. Conscious Cogn, 20(4), 1532-1547. https://doi.org/10.1016/j.concog.2011.07.004

      Ramamurthy, D. L., Rodriguez, L., Cen, C., Li, S., Chen, A., & Feldman, D. E. (2025). Reward history guides focal attention in whisker somatosensory cortex. Nat Commun, 16(1), 5580. https://doi.org/10.1038/s41467-025-60592-w

      Reynolds, J. H., Pasternak, T., & Desimone, R. (2000). Attention increases sensitivity of V4 neurons. Neuron, 26(3), 703-714. https://doi.org/10.1016/s0896-6273(00)81206-4

      Serences, J. T. (2008). Value-Based Modulations in Human Visual Cortex. Neuron, 60(6), 1169-1181. https://doi.org/10.1016/j.neuron.2008.10.051

      van den Berg, B., Krebs, R. M., Lorist, M. M., & Woldorff, M. G. (2014). Utilization of reward-prospect enhances preparatory attention and reduces stimulus conflict. Cognitive Affective & Behavioral Neuroscience, 14(2), 561-577. https://doi.org/10.3758/s13415-014-0281-z

      van Ede, F., & Nobre, A. C. (2023). Turning Attention Inside Out: How Working Memory Serves Behavior. Annu Rev Psychol, 74, 137-165. https://doi.org/10.1146/annurev-psych-021422-041757

      Vandevelde, J. R., Yang, J. W., Albrecht, S., Lam, H., Kaufmann, P., Luhmann, H. J., & Stuttgen, M. C. (2023). Layer- and cell-type-specific differences in neural activity in mouse barrel cortex during a whisker detection task. Cereb Cortex, 33(4), 1361-1382. https://doi.org/10.1093/cercor/bhac141

      Wang, L., & Krauzlis, R. J. (2018). Visual Selective Attention in Mice. Curr Biol, 28(5), 676-685 e674. https://doi.org/10.1016/j.cub.2018.01.038

      You, W. K., & Mysore, S. P. (2020). Endogenous and exogenous control of visuospatial selective attention in freely behaving mice. Nat Commun, 11(1), 1986. https://doi.org/10.1038/s41467-020-15909-2

      Zhu, P., Guan, C., Fu, Y., Shen, M., & Chen, H. (2026). Working memory encoding of attended information is adaptive to future relevance. J Exp Psychol Learn Mem Cogn. https://doi.org/10.1037/xlm0001582

    1. eLife Assessment

      This valuable study compares hippocampal-cortical functional connectivity to various other brain measures and examines their development across youth. It uses sophisticated analyses replicated in multiple datasets, but provides incomplete evidence to support the primary claim that hippocampal-cortical connectivity relates to cognitive maturation. The manuscript would benefit from a more nuanced consideration of the biological basis of some of the derived imaging measures and the limitations of the cross-sectional design. This work will be of interest to neuroimaging specialists and cognitive neuroscientists.

    2. Reviewer #1 (Public review):

      Summary:

      The authors studied the development of hippocampal connectivity gradients based on open datasets and performed correlation analyses with other MRI features as well as gene expression information from other datasets. Although the main findings are correlational and cross-sectional, the analyses are overall sophisticated and replicated in several datasets.

      Strengths:

      The hippocampus is a key region in understanding large-scale brain organization and cognition, and the authors applied advanced and suitable analytics to study its development. The paper is overall well-organized and well-written, and the findings are relevant for studying large-scale brain development.

      Weaknesses:

      While sophisticated, several of the analyses appear mainly correlational, cross-sectional, and rely on cross-dataset contextualization, which should also be stated as a limitation of the current work.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors aim to assess how the functional organisation of the hippocampus is related to the geometry and neurobiological differences of the hippocampus. In particular, the authors focus on the first three eigenvectors of hippocampal-cortical functional connectivity, based on non-linear dimensionality reduction on resting-state functional MRI data. Furthermore, the work aims to describe changes in these functional axes and their relation to other factors throughout youth and evaluate whether they are predictive of individual variations in cognition.

      Strengths:

      A major strength of this study is the attempt to replicate key findings across multiple developmental cohorts.

      Weaknesses:

      The major weaknesses of the manuscript center on gaps in technical transparency and several conceptual inaccuracies. The machine learning methodology used for cognitive prediction is scarce, leaving little means to evaluate whether the behavioral results suffer from data leakage or overfitting. The introduction sets up an oversimplified historical premise regarding the field's understanding and appreciation of hippocampal connectivity, and contains several incorrect references that throw doubt on the argumentation. Additionally, T1w/T2w signal intensity is incorrectly used as synonymous with myelin, despite gold-standard histological validation showing a non-significant correlation between T1w/T2w and myelin staining (Sandrone et al., 2013).

      Appraisal of Aims and Conclusions:

      The authors partially achieve their aims by illustrating certain age-related changes in hippocampal function; however, the correlative study design is not equipped to examine how these changes are "shaped" by geometry, myelination, or gene expression (especially the latter two). Furthermore, conclusions were often overstated based on small effect sizes.

      Context and Field Impact:

      This work adds to a growing body of literature focused on gradient-based representations of hippocampal topology. By applying these methods across a wide developmental age bracket, it provides a useful reference point for how the hippocampus and wider cortex interact during maturation. To improve utility to the neuroimaging and cognitive neuroscience communities, the nesting of subfields within the eigenvector topology should be addressed, too.

    1. eLife Assessment

      This valuable manuscript investigates how Drosophila larvae make foraging decisions in patchy environments with controlled resource density and valence; using movement tracking in bounded arenas, the authors show that larvae's patch residence time (PRT) differs depending on resource type, environmental context, and prior experience. A drift-diffusion model is used to describe patch-leaving behaviour, suggesting that an integration process may underlie stay-leave decisions during foraging. The strength of the evidence is mostly solid, but the interpretation and use of PRT needs further investigation, as PRT could be a direct effect of resource concentration on locomotion. Explicit reports of PRT statistical tests are needed for rigorous interpretation.

    2. Reviewer #1 (Public review):

      Summary:

      Mudunuri et al. investigate the foraging response of Drosophila larvae in response to patchy resources of distinct value (concentration of nutrient or valence). They show that larvae adjust their behavior according to both the quality and valence of available resources. Interestingly, previous exposure to resources of lower value increases the permanence time in resources of greater value. This suggests that larvae can value, remember and adapt their behaviour in response to previous foraging experience.

      They perform a simple integration model that recapitulates the larval behaviour.

      Strengths:

      This paper uses a very well-controlled foraging set-up where larvae are tested individually and for 3 hours, allowing for a good statistical analysis of their behaviour.

      They investigate for the first time the ability of Drosophila larvae to perceive, remember and compare the quality and valence of distinct resources. It is very exciting, as it will open up the field of foraging decision studies using the fruitfly larvae.

      Weaknesses:

      (1) Most of the analysis depends on the thresholding, but it is not clear what increasing the radius of analysis means in terms of foraging. There are two issues here:

      a) What is the behaviour of the larvae on the edges of the patch? It is obvious that the fructose or the NaCl will diffuse at the edge, so are they remaining in the proximity because they are actively feeding (exploiting) on this decaying concentration, or are they sensing the lower gradient and they are actually looking (chemosensing) for the higher concentration? The behaviour at the edge is really different (check sucrose in Wosniack et al. 2022), and there might be a way of avoiding the diffusion by actually adding a plastic ring and pouring the agar + resource in there. The effect of the ring, per se, would still have to be tested.

      b) How was the threshold selected? It is very likely that the concentration at the patch boundary will be very different for 1M and 0.1 M. Could the authors explain why they chose such a distance? What does majority of larvae mean? Is the "majority" the same for 0.1M and 1M? Is there a relationship between the threshold chosen and the diffusion of fructose and NaCl?

      (2) The word exploitation is used in the paper, but there are many instances where it is unclear whether that is the case. This should be clarified since there are no controls for exploitation.

      (3) In the experiments analysing the adaptation of foraging behaviour, it is not clear if the first and second patch means that only 2 patches were analysed per larva or the first and second in a sequence of patches visited. I think it is the second option (because of Figure S3D), but the authors should clarify this. Also, we do not know how many animals were tested. The number of data points in 4C (4G) compared to 4D (4H) seems very different.<br /> Regarding the results, which are very interesting, why aren't the larvae spending less time in the 0.1M sucrose patch after having fed on a 1M patch, while they spend more time in a 1M after a 0.1M? Could it be that the difference in residence time is correlated with their hunger rather than the comparison between conditions?

      (4) I am not an expert in this type of model, and I would appreciate it if the authors could explain how the values of the drift and leak have been fitted in Figure 5H. If possible, I would recommend adding a graph showing the parameter exploration of distinct possible combinations of values.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how Drosophila larvae make foraging decisions in patchy environments with controlled resource density and valence. Using movement tracking in bounded arenas, the authors show that larvae's patch residence time (PRT) differs depending on resource type, environmental context, and prior experience.

      The authors vary whether the environment is homogenous (all patches are equal) or heterogenous (mixed patches) and whether a higher density of the resource is appetitive (food) or aversive (salt). The most salient results are that in heterogeneous environments, larvae remain longer on higher-density patches of fructose, while they stay shorter in higher-density salt patches. The study further demonstrates that prior foraging experience influences subsequent patch residence time (PRT).

      A drift-diffusion model is used to describe patch-leaving behavior, suggesting that an integration process may underlie stay-leave decisions during foraging. Overall, the work provides a useful behavioral system for studying foraging behaviour and highlights the role of context and experience in shaping larval foraging strategies.

      Strengths:

      A major strength of the manuscript is the behavioral system. The assay is simple, well-controlled, and suitable for realistic spatial and temporal scale tracking of individual larvae. The use of non-volatile resources and embedded patches minimizes confounds from olfactory navigation and allows the authors to focus on local patch exploitation, return behavior, and experience-dependent decisions.

      The results regarding patch resident time (how long larvae stay in patches of different resource density) are convincing. In homogeneous environments, larvae spend more time on patches with a higher density of food (0.1M > 0.01M) and less time in patches with a lower density of salt (0.01M > 0.1M), indicating that their behaviour is sensitive to the valence of the resource. Further, larvae do not simply respond to current circumstances, since PRT in a given patch is sensitive to the quality of the preceding one encountered, showing some kind of memory.

      Weaknesses:

      (1) The theoretical background of the experiment, as exposed in the Introduction, is somewhat misleading. The experiment is based on patches of sufficient size for the individual larvae not to deplete them through their activity, so that the intake rate is constant while exploiting a given patch. In those circumstances, the theoretical rate-maximizing strategy would be to either reject a patch on encounter or stay in it indefinitely (until pupation). The threshold for rejection or acceptance will depend on travel time, but patch residence time would be either zero (or minimal identification time) or lifelong. In the introduction, it appears as if the system follows the classical Marginal Value Theorem assumptions as used in classical foraging theory. In that case, patch residence time is fundamentally sensitive to a decline in intake rate while in a patch. This raises questions about what factors drive patch-leaving in the present protocol. A better theoretical framework would focus on behavioural variables that can be expected to depend on the circumstances of the experiment, as discussed below.

      (2) Rather than make predictions about time in the patch, which as explained above do not reflect the present system, larval behaviour could be modelled and described as a function of observable properties such as: (a) speed of locomotion; (b) tendency to deviate from straight progress (area restricted searching); (c) probability of return after leaving a patch, possibly controlled through rea restricted searching; (d) a response to concentration gradient, since patch boundaries are probably gradual through diffusion. There is a useful literature in this regard in studies of parasitic wasps such as Venturia canescens (formerly Nemeritis canescens, see Waage 1979). Larva may respond directly to local resource concentration (see van Alphen, J. J., Bernstein, C., & Driessen, G., 2003), where higher concentration leads to increased feeding rate, reduced locomotion, and consequently results in longer time in each patch. This could still be a normative model, but based on realistic driving inputs. The dimensions of the system make it unlikely that larvae have the opportunity to adjust to travel time, or patch composition, on which classical foraging models are based. The original versions of the marginal value theorem were thought for cases where birds exploited pine cones, so that each bird had multiple encounters, and also on dung flies that mated in dung patches, which also dried out. A system with heritable optimised parameters could work for other natural systems where the parameters can be heritable, but not here.

      (3) The previous argument indicates that patch time, while it is a real quantitative consequence, is not ideal as the major dependent variable for this system. Given that the authors have the full trajectories, they could treat movement in discrete time bins and ask if the tendency to depart from linear progression (i.e. from moving straight ahead) is a function of the density of the resource. It would appear as if all the results, including return to patches (but not memory), could be explained by area-restricted searching (see Dorfman, A., Hills, T. T., & Scharf, I. (2022). A guide to area‐restricted search: a foundational foraging behaviour. Biological Reviews, 97(6), 2076-2089.). Slower movement (perhaps directly caused by eating) and more twisted progress could generate longer times in higher food densities.

      (4) The evidence for an effect of prior experience is interesting but could be strengthened. The authors state that PRT on the second patch depends on the concentration in the first patch. However, statistically significant modulation of prior experience was only found when the second food patch was richer, namely 1M fructose (Figure 4C). If the change in patch time is due to a form of learning and contrast, one might expect significantly shorter times in any second patch if the first one was richer, which is not the case. One difficulty is that the 'patchy' nature of the environment may not be evident to the larvae, because they are much smaller than the patches. From a larva's perspective, a patch is an environment, potentially suitable to remain in until pupation (which is what they ought to do in richer food patches).

      (5) The modelling section is promising but currently somewhat underdeveloped relative to the strength of the claims. The authors fit a drift-diffusion model to data and report that a drift-only model captures homogeneous environments, whereas adding a leak term improves the fit in heterogeneous environments. This provides a useful quantitative summary of behavior but the biological interpretation of the leak parameter is not clear. In addition, the valence condition was not modelled.

    4. Reviewer #3 (Public review):

      Summary:

      The work investigates how the foraging behaviour of Drosophila larvae depends on resource quality, valence, and heterogeneity in the foraging environment. A specific focus of the work was to study how foraging decisions depend on the prior experience of alternative resource patches in the same environment. Moreover, the work presents computational models (drift diffusion models) that recapitulate foraging decisions, and whose parameters appear to depend on resource quality and environment statistics, providing potential insights into the dynamics of the decision-making process.

      I am not familiar with previous literature on foraging decisions in Drosophila, but I was specifically consulted to comment on the computational modelling. Therefore, my comments will mostly focus on the modelling aspects.

      Strengths:

      In my understanding, the two strengths of the current study are that:<br /> (1) it uses non-volatile resources, providing better control of the available cues that could guide foraging decisions, and<br /> (2) it tracks foraging behaviour over an extended period of time (3h), generating a rich dataset of foraging behaviour in the same environment.

      Overall, the study appears to have been carefully conducted.

      Weaknesses:

      The computational modelling currently provides limited additional value beyond the empirical results. There are no prior hypotheses that are addressed by the computational models. Given the flexibility of DDMs, fitting foraging times is expected to be feasible. The question is whether the fits provide mechanistic insight. The main insight appears to be that describing foraging times in a homogeneous environment requires a single free parameter (drift rate), while the heterogenous environment requires a second parameter (leak). However, the effective complexity of the model is higher than the stated parameter count suggests, as each patch quality is fit with a different drift rate, which does not generalise across environments: in the heterogeneous environment, the drift rate differs substantially across fructose concentrations, whereas in the homogeneous environment, the same concentrations yield nearly identical drift rates. Counter their claims, the authors also do not systematically explore the effect of specific prior foraging experience on computational parameters, but only contrast model fits to environments with different statistics, in which prior experiences will be generally different. Overall, at the moment these modelling results have a rather descriptive character, and provide very little insight into the underlying computational principles that drive foraging decisions.

      A second weakness is that the study does not report the detailed results of the statistical tests, and it seems that the authors interpret several differences that are not marked as statistically significant in the figures. Furthermore, the model comparisons do not account for different degrees of freedom of the models, and the goodness of fit values alone are insufficient to conclude that one model is better than the other (rather than overfitting).

    1. eLife Assessment

      This useful study investigates noise-robust and energy-efficient circuit mechanisms for working memory by optimizing connectivity and reports that the resulting networks exhibit rotational dynamics and better match aspects of PFC population recording. However, the supporting evidence remains incomplete, given the restricted linear, task-specific training and analysis, and limited comparisons with other prominent models. The manuscript would be strengthened by extending the analysis to nonlinear dynamics, providing more rigorous comparisons with alternative models, and establishing a stronger link to prior theoretical and experimental work.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors address the question of working memory maintenance, starting from the experimental observation that recordings of neural activity during the delay period of working memory tasks are sometimes observed to be dynamic. They introduce a new combination of metrics (noise-robustness and energy efficiency) to quantify the performance of various network mechanisms of memory maintenance, in linear networks. They compared attractor networks, feed-forward networks, and networks trained with a loss that includes a robustness and an energy-efficiency component. They show, by plotting state-space trajectories, that networks optimized with this loss exhibit a form of rotational dynamics. They analyzed the data recorded during the delay of a working memory task in PFC, and observed state-space trajectories similar to those of the trained networks.

      The comparison with other network mechanisms is interesting in principle, but limited by the fact that only linear networks are considered. This led to counter-intuitive and misleading statements, like the fact that attractor networks are not robust to noise, or that feed-forward networks have energy consumption that is exponential in the number of neurons.

      Strengths:

      (1) The idea to use both robustness to noise and energy efficiency to assess the performance of networks on working memory tasks is interesting.

      (2) The manuscript is clearly written.

      (3) There is an interesting combination of methodologies: theory on simple models, network training, and data analysis.

      Weaknesses:

      (1) Linear networks only.

      The main feature of attractor networks is their robustness to noise, which is typically allowed by the non-linearity of neural responses. To fit their modeling framework, the authors focused only on continuous attractor neural networks (e.g., Seung 1996) and ignored point-attractor models such as the Hopfield model, which are typically used to model WM tasks, and which would presumably lead to very different results, e.g., in Figure 1D.

      The linearity assumption is also problematic for the comparison with feed-forward models. It seems that the authors obtained runaway firing rates, explaining Figure 1F middle, which are typically prevented in non-linear networks.

      The choice of parameters for the attractor network in Figure 1 is not explained. Why is t_slow = 10^4 chosen, and what does it correspond to? We expect in linear networks that activity goes back to zero or diverges as an exponential, but in principle, the time constant can be chosen to be of the same order as the time delay, with approximately linearly decreasing SNR.

      Regarding the comparison of the different mechanisms, it would have been nice to better define the notion of rotational dynamics, beyond only considering state-space analysis, which is limited to providing mechanistic interpretations.

      (2) Fixed duration of delay periods.

      I have understood that for a given network, the duration of the delay period is fixed, as opposed to a delay duration that would fluctuate from trial to trial. This would be an important assumption to relax as well, to better match common experimental paradigms, as well as to expose a fairer comparison with other network mechanisms. See Orhan and Ma (2023) for such a discussion.

      (3) Relationship with previous works

      Many other works addressed the question of dynamic firing rates during maintenance periods of WM tasks; they should be discussed and compared to the mechanism proposed here. This includes: Barak et al, Progress in Neurobio. 2013, Pereira-Obilinovic, Aljadeff, Brunel, PRX 2023, Hansel, Mato, 2013, or works pertaining to the activity-silent neural states (allowed by short-term plasticity), the framework in which the data of Panichello et al are interpreted in the original publication.

    3. Reviewer #2 (Public review):

      In this manuscript, Ritter et al. propose a model of working memory (WM) that combines feedforward and rotational dynamics. The model is discovered by optimizing a linear RNN using a loss function that encourages maximization of signal-to-noise ratio (SNR) and minimization of activation magnitude. The authors argue that the optimized model outperforms other WM models in terms of SNR and energetic efficiency, while also better replicating key features of neural responses recorded in monkey pre-frontal cortex (PFC) during a WM task. The authors also draw connections to state space models (SSM) used for other machine learning applications.

      My main issue with this manuscript is that it does not appear to convincingly demonstrate that rotational dynamics offer any advantage over purely feedforward dynamics. The authors adopt three criteria according to which they compare models:<br /> (1) SNR.<br /> (2) Energy efficiency.<br /> (3) Similarity to neural data.

      In terms of SNR, purely feedforward models seem to perform similarly to the optimized models (Figure 1). Figure 1 does seem to show that the optimized network produces responses of smaller magnitude when the number of units is large, but the authors do not explain why adding rotational dynamics would produce such a relationship. In fact, the responses that are plotted for the feedforward network in Figures 1B, 2C, and 5E look similar, if not smaller in magnitude than those of the optimized model. Lastly, while the authors claim in the body of the text that the optimized model replicates key features of monkey PFC responses better than the purely feedforward model, this is not apparent to me from the comparisons plotted in Figure 5E-J. The authors thus do not show strong evidence that the model they propose beats what they claim is an established baseline on any of the three criteria.

      Another weakness of the manuscript is that the comparison to attractor and feedforward models seems somewhat unfair. In Figure 1, the rotational model is optimized, while the parameters for the attractor and feedforward models seem to have been at least partially chosen by hand. Figure 5C again shows the three models side by side, but the fact that it compares the same network at different stages during training complicates the comparison. Instead, one should compare the rotational solution to the optimal attractor and feedforward models, respectively (obtained by constrained optimization). From looking at the flow-fields, it seems that a feedforward network with an optimized level of amplification may work just as well. On a mechanistic level, it is unclear what computational advantage rotations offer over feedforward dynamics in the WM context.

      The choice of baseline models to compare against might be questionable. The simple line attractor model by Seung et al. (1996) was initially designed to explain oculomotor integration. It is true that a line attractor has been suggested as a mechanism for working memory, e.g., in the seminal work by Machens et al (2005). However, it seems fair to say that most studies employing non-linear networks have focused on point attractors as mechanisms of working memory (e.g., Wong & Wang, 2006; Driscoll, Shenoy, Sussillo, 2024). A point attractor arguably does not suffer the SNR issues of a line attractor, because it does not lead to integration of the noise over time. However, non-trivial point attractors cannot be implemented in linear networks of the kind studied by the authors of the present study.

      The authors should expand their discussion to include other, potentially closely related work proposing rotation-like dynamics in artificial neural networks during working memory. In particular, the manuscript does not discuss Sharma, Proca, et al, ICML 2026, which describes a rotational solution to a similar WM task obtained by optimizing linear RNNs (Sharma et al., 2026, Fig. 6). Notably, Sharma et al. arrive at a similar rotational (and likely also non-normal) mechanism without using either noisy inputs or a constraint on energy efficiency. The authors of the present manuscript should discuss to what extent this finding contradicts their claim that "normative pressures on noise-robustness and energetic cost shape the complex dynamics of WM circuits." (present manuscript, Introduction). Given the obvious parallels between the two studies, a comparison between the present work and Sharma et al. (2026) would add necessary context to the Discussion.

      The authors should also clarify the significance of the "novel method for optimization of continuous-time RNNs driven by noisy inputs" (see Discussion) that the authors propose. This method is mentioned in the first line of the Discussion section but is barely discussed, let alone sufficiently explained, in the previous Sections. The only time a comparison to BPTT with a simple MSE loss is mentioned, it is stated that the two procedures produce the same results. The novel method appears to consist of a loss with two terms, the second of which is a well-known L2-penalty on unit activations (Sussillo et al., 2015). It is not clear that the method is either novel or necessary to obtain the reported results.

      Except for the fact that higher-dimensional networks also converge on rotational solutions, Figure 3 does not add much to the reader's understanding of the optimized model (except for panel F). I find the comparison to SSMs too superficial to provide real insight.

      Figure 4 claims to show that the optimized model recapitulates "a range of properties observed in prefrontal cortex and other brain areas during WM tasks" (p. 7) but does not show neural data for comparison.

    4. Reviewer #3 (Public review):

      Summary:

      The authors optimize continuous-time linear recurrent networks driven by noisy input, computing the gradient of decoding performance numerically and analytically. Optimizing for stimulus discriminability after a delay, with a penalty on firing rate, they find networks that adopt what they call high-dimensional rotational dynamics. They argue that these outperform attractor and feedforward models on noise robustness and energetic cost, and resemble state-of-the-art state-space models. They then fit a targeted dimensionality reduction model to prefrontal recordings from monkeys performing a spatial working memory task and argue that the population structure matches the rotational solution.

      Strengths:

      The evolution of the dynamics throughout learning is a nice observation, as are the analytical calculations, although I am not sure they are new since there is a fair share of work on the learning dynamics of linear networks.

      Weakness:

      I see many weaknesses. I will classify them into five groups.

      (1) Strawman comparison and no clear definition of what is rotational. The paper is centered on comparing a trained model with two models meant to represent "attractor dynamics" and non-normal dynamics. Both are picked as the weakest member of their class.

      I use quotation marks for "attractor dynamics" because I am not sure a linear system with an eigenvalue equal to zero is a representative model for the class. This is a particular linear instantiation of the line attractor from Seung 1996, but most attractor models are nonlinear and far more robust to noise, and they are robust through error correction that this linear model does not have. Even modern continuous attractors (Rivkind and Darshan) are very robust to noise through multiple mechanisms. So what the authors picked as an "attractor model" is a limited zero-eigenvalue case that, of course, will drift. "Attractor networks are highly susceptible to noise" is therefore true only of the toy they built, not of the class.

      Second, what they call a non-normal model is in fact a feedforward chain, the extreme of non-normality. There are degrees of non-normality in any matrix, and the homogeneous delay line is the corner that requires the largest firing rates. This is not representative. See Daie et al., which has a skip and recurrent structure, or Stroud, which is not a pure chain. So the feedforward chain was also picked as a strawman, chosen so that the energetic cost they then complain about is guaranteed.

      This brings me to the real problem in this section. "Rotational" is never defined. If it means complex eigenvalues, then it is a spectral property of any non-normal matrix, and "rotational versus feedforward" is not a dichotomy; it is two regions of the same continuous space of non-normal connectivity. Their own Figure 2C shows the network passing continuously through an attractor, then feedforward, then rotational during optimization. If these are points on a continuum, then "rotational dynamics is optimal" is just a statement about where the optimizer lands under this particular loss and input normalization, not the discovery of a new dynamical class. They need to define the term operationally and show the solution is qualitatively, not just quantitatively, different from non-normal feedforward. I do not think it survives that test.

      This brings me to the references.

      (2) The dynamical mechanisms of working memory have been studied for more than two decades, and I am surprised how much directly relevant work is missing. First, Druckmann and Chklovskii 2012, where a linear system produces stable encoding from oscillating modes. This is essentially their result more than a decade earlier, and it is not cited. They also miss Murray et al. on stable encoding and heterogeneous timescales in data. They oversimplify the attractor picture; for example, Pereira-Obilinovic et al. 2023 show you can have genuinely stable attractors. They do cite Daie et al., but they ignore its central claim, that non-normality is the underlying mechanism, which is more troubling than not citing it because it means they read it and did not engage. Overall, the references are idiosyncratic, missing relevant work, and not engaging the results of papers they cite.

      This brings me to the third point.

      (3) Novelty and the relationship to Stroud and Orhan. Those papers take a similar optimization approach and find that, depending on the task parameters, the optimal solution is non-normal, non-normal plus attractor, or attractor. My impression is that what this work calls rotational is just the dynamics of a strongly non-normal A, selected here by the firing-rate regularizer. They never clarify the connection with Stroud. Is the only difference the energy penalty?

      The way to settle this is quantitative, and they have the handle and do not use it: report the Henrici departure-from-normality of their optimized A and place the solution inside Stroud's regime structure.

      There is also a tension they leave implicit. In Stroud, the early loading direction is orthogonal to the late persistent readout, and that orthogonality is the source of dynamic coding. This paper's subspace alignment result (Figure 5G, H) shows exactly this early-to-late orthogonalization in both model and data, and then presents it as evidence for the rotational account and against Stroud's hybrid. You cannot reproduce a Strout's stim vs. decoder orthogonality and claim it against Strout's without doing more work.

      (4) I did not understand the SSM section, and I think it should be cut. Is this a result? Either "SSM" just means a linear dynamical system, in which case it is trivial since every linear network here, including the LMU is an SSM, or it means the network matches a fixed-connectivity model like the LMU, which it does not seem to either. So in what sense is it a result?

      (5) The data analysis is one section, and the analysis could be described as feeling somewhat like an afterthought on a very rich dataset. The coding structure they show for the rotational model also looks like the Stroud non-normal-plus-attractor model to me. They even state that the hybrid reproduces the cross-temporal subspace. What are the quantitative, cross-session metric that discriminates rotational from the non-normal-plus-attractor hybrid? Is it eyeballed trajectories?

    1. eLife Assessment

      This important study provides a detailed characterization of individual sarcomeres' contractility and of their synchrony in spontaneously beating cardiomyocytes derived from human induced pluripotent stem cells. The combination of high-resolution tracking, statistical analysis and mesoscopic modeling leads to compelling evidence that sarcomeres operate as dynamically unstable units, leading to stochastic heterogeneities in their contraction-elongation cycles depending on substrate stiffness. The work will be relevant to scientists interested in muscle biophysics, nonlinear dynamics and synchronization phenomena in biological systems.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this manuscript, the authors present comprehensive experimental observations and a theoretical framework to explain the heterogeneous behaviour of sarcomeres in cardiomyocytes. They show that a stochastic component exists in their contractile activity, which may act as a feedback mechanism regulating physiological function.

      Strengths:

      Experiments and data analysis are robust and valid. The rigorous statistical analysis and unbiased methods enable the authors to draw well-supported conclusions that go beyond the existing literature. Their outcomes inform about cellular activity at the individual level and the authors explain how the transient dynamics of single sarcomeres are governed by a force-velocity relationship and lead to the complex contractile patterns. The similarity of the results to the study cited in [24] demonstrates the validity of the in vitro setup for answering these questions and the feasibility of such in-vitro systems to extend our knowledge of out-of-equilibrium dynamics in cardiac cells.

      Very interesting the suggestion that the interplay between intrinsic fluctuations and the dynamic instability are part of a feedback mechanism for maintaining structural and functional homeostasis.

      The addition of the theoretical model and the new text of the manuscript improves the clarity of the study.

    3. Reviewer #2 (Public review):

      Summary:

      Sarcomeres, the contractile units of skeletal and cardiac muscle, contract in a concerted fashion to power myofibril and thus muscle fiber contraction.

      Muscle fiber contraction depends on the stiffness of the elastic substrate of the cell, yet it is not known how this dependence emerges from the collective dynamics of sarcomeres. Here, the authors analyze contraction time series of individual sarcomeres using live imaging of fluorescently labeled cardiomyocytes cultured on elastic substrates of different stiffness. They find that a reduced collective contractility of muscle fibers on unphysiologically stiff substrates is partially explained by a lack of synchronization in the contraction of individual sarcomeres.

      This lack of synchronization is at least partially stochastic, consistent with the notion of a tug-of-war between sarcomeres on stiff sarcomeres. A particular irregularity of sarcomere contraction cycles is 'popping', the extension of sarcomers beyond their rest length. The statistics of 'popping' suggest that this is a purely random process.

      Strengths:

      This study thus marks an important shift of perspective from whole-cell analysis towards an understanding the collective dynamics of coupled stochastic sarcomeres.

    4. Reviewer #3 (Public review):

      The manuscript of Haertter and coworkers studied the variation of the length of a single sarcomere and the response of microfibrils made by sarcomeres of cardiomyocytes on soft gel substrates of varying stiffness.

      The measurements at the level of a single sarcomere are an important new result of this manuscript. They are done by combining the labeling of the sarcomeres z line using genetic manipulation and a sophisticated tracking program using machine learning. This single sarcomere analysis shows strong heterogeneities of the sarcomeres that can show fast oscillations not synchronized with the average behavior of the cell and what the authors call popping events which are large amplitude oscillations. Another important result is the fact that cardiomyocyte contractility decreases with the substrate stiffness, although the properties of single sarcomeres do not seem to depend on substrate stiffness.

      The authors suggest that the cardiomyocyte cell behavior is dominated by sarcomere heterogeneity. They show that the heterogeneity between sarcomere is stochastic and that the contribution of static heterogeneity (such as composition differences between sarcomeres) is small.

      Strengths:

      All the results are, to my knowledge, new and original. The authors also made a theoretical model where each sarcomere is described by a Langevin equation based on a non-linear coupling between force and velocity of the sarcomeres. This model accounts well for the experimental results including the observation of what the authors call popping events.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors present comprehensive experimental observations and a theoretical framework to explain the heterogeneous behaviour of sarcomeres in cardiomyocytes. They show that a stochastic component exists in their contractile activity, which may act as a feedback mechanism regulating physiological function.

      Strengths:

      Experiments and data analysis are robust and valid. The rigorous statistical analysis and unbiased methods enable the authors to draw well-supported conclusions that go beyond the existing literature. Their outcomes inform about cellular activity at the individual level and the authors explain how the transient dynamics of single sarcomeres are governed by a force-velocity relationship and lead to the complex contractile patterns. The similarity of the results to the study cited in [24] demonstrates the validity of the in vitro setup for answering these questions and the feasibility of such in-vitro systems to extend our knowledge of out-of-equilibrium dynamics in cardiac cells.

      Very interesting the suggestion that the interplay between intrinsic fluctuations and the dynamic instability are part of a feedback mechanism for maintaining structural and functional homeostasis.

      The addition of the theoretical model and the new text of the manuscript improves the clarity of the study.

      Reviewer #2 (Public review):

      Summary:

      Sarcomeres, the contractile units of skeletal and cardiac muscle, contract in a concerted fashion to power myofibril and thus muscle fiber contraction.

      Muscle fiber contraction depends on the stiffness of the elastic substrate of the cell, yet it is not known how this dependence emerges from the collective dynamics of sarcomeres. Here, the authors analyze contraction time series of individual sarcomeres using live imaging of fluorescently labeled cardiomyocytes cultured on elastic substrates of different stiffness. They find that a reduced collective contractility of muscle fibers on unphysiologically stiff substrates is partially explained by a lack of synchronization in the contraction of individual sarcomeres.

      This lack of synchronization is at least partially stochastic, consistent with the notion of a tug-of-war between sarcomeres on stiff sarcomeres. A particular irregularity of sarcomere contraction cycles is 'popping', the extension of sarcomers beyond their rest length. The statistics of 'popping' suggest that this is a purely random process.

      Strengths:

      This study thus marks an important shift of perspective from whole-cell analysis towards an understanding the collective dynamics of coupled, stochastic sarcomeres.

      Reviewer #3 (Public review):

      The manuscript of Haertter and coworkers studied the variation of the length of a single sarcomere and the response of microfibrils made by sarcomeres of cardiomyocytes on soft gel substrates of varying stiffness.

      The measurements at the level of a single sarcomere are an important new result of this manuscript. They are done by combining the labeling of the sarcomeres z line using genetic manipulation and a sophisticated tracking program using machine learning. This single sarcomere analysis shows strong heterogeneities of the sarcomeres that can show fast oscillations not synchronized with the average behavior of the cell and what the authors call popping eveents which are large amplitude oscillations. Another important result is the fact that cardiomyocyte contractility decreases with the substrate stiffness, although the properties of single sarcomeres do not seem to depend on substrate stiffness.

      The authors suggest that the cardiomyocyte cell behavior is dominated by sarcomere heterogeneity. They show that the heterogeneity between sarcomere is stochastic and that the contribution of static heterogeneity (such as composition differences between sarcomeres) is small.

      Strengths:

      All the results are, to my knowledge, new and original. The authors also made a theoretical model where each sarcomere is described by a Langevin equation based on a non-linear coupling between force and velocity of the sarcomeres. This model accounts well for the experimental results including the observation of what the authors call popping events.

      We thank you and the reviewers for the positive evaluation of our revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Origin of the 3-Hz oscillation and required model extension. These oscillations are reproduced by our model, and their origin is already discussed in the manuscript (see lines 403–406).

      (2) Inclusion of all 5085 LOIs vs. the selected 2321. We have expanded the explanation of the LOI selection criteria in the manuscript and clarified that the main conclusions are not sensitive to this choice (lines 161-166)

      (3) Fig. 3G caption — popping rate. The caption has been updated to clarify the units and normalization. 

      (4) Fig. 4G — "Length x" vs. ΔL. Notation corrected for consistency.

      (5) Fig. 4G — gray data points. Confirmed: these represent the mean, and the caption has been updated accordingly.

      (6) Relation of k_l to the true substrate stiffness. We have added the following clarification: "The model evaluation compared the distributions of sarcomere length changes and velocities from simulations with representative experimental LOIs from substrates (5, 15, and 85 kPa, mapped to k_l = 0.5, 1.5 and 8.5 in our 1-D model; k_l is unitless, so only the ratios between values are meaningful — rescaling k_l leaves model output unchanged under correspondingly rescaled parameters) covering the full range of mechanical loads." (lines 365-369)

      (7) Could a simpler model fit the data? The cubic polynomial in Eq. (3) was deliberately chosen as a generalist ansatz rather than imposed: its coefficients were obtained by data-driven inference via Differential Evolution, and if lower-order terms within this family had sufficed, the higher-order coefficients would have been driven toward zero. The inferred nonmonotonic force–velocity relation has two extrema separated by an unstable negative-slope branch, which sets a lower bound on the polynomial order — a linear F–v is monotonic and a quadratic admits only a single extremum, so cubic is the minimum polynomial order capable of producing the observed shape. Furthermore, the qualitative phenomena we report — popping events, dynamic instability, and stochastic heterogeneity — cannot arise from any monotonic force–velocity relation, as discussed in the section on the non-monotonic instability. With 10 parameters covering complex contractile dynamics at the individual sarcomere and myofibril level across different substrate stiffnesses, the present model is parsimonious within the family of polynomial force–velocity ansätze; we have not exhaustively searched alternative non-polynomial functional families, but any such alternative would still need to reproduce the same non-monotonic shape that the data require.

      (8) Lines 497–507 in the Discussion. On reflection, we feel these lines provide useful context for the broader interpretation and would prefer to retain them.

      (9) Line 331 — motivation of Eq. (3). We have added citations to prior work motivating this form of the equation for the broader readership.

      (10) Line 427 — "scaled". Corrected.

      Reviewer #3 (Recommendations for the authors):

      We thank the reviewer for the recommendation of a theoretical appendix. The full model code, with the formulation and implementation documented in detail, is publicly available in our GitHub repository accompanying the paper, which we believe provides a complete reference for readers wishing to explore the model further. We therefore feel an additional appendix is not necessary within the scope of this revision.

    1. eLife Assessment

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells and trachea. They show that the KO mice exhibit severe hydrocephalus due to mislocated basal bodies and impaired ciliary beating. The findings are valuable with implications in the subfield of cell biology. The evidence is solid in that the methods, data and analyses largely support the claims with only a few remaining weaknesses.

    2. Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights, regarding the function of the CCP5 N-domain.

      Comments on revised version.

      The authors have appropriately revised the manuscript in response to most of my comments.

    3. Reviewer #2 (Public review):

      Summary:

      This study analyzed consequences of Agbl5 mutation on ependymal cells development and function. Authors first characterize their mutant mouse line reporting a reduced lifespan and severe hydrocephalus. Next, they report defect in ependymal cell cilia number and motility. They provide evidence for impaired basal bodies organisation, cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicate Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype are incomplete:

      Previous comment: Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting that the author checks whether the subapical network of microtubule is glutamylated or not during ependymal cells differentiation and how this network is affected in their mutants.

      Although authors now provide images of glutamylation in figure S8 their conclusion claiming that GT335 signal is increased in cilia of Agbl5M1/M1 mutant is not supported convincingly by those pictures. Quantification would be needed.

    4. Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele by extending the deletion to the N-terminus of CCP5 to investigate its function in mouse ependymal cells and trachea.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      The manuscript is well-written, and the experiments are convincing.

      Comments on revised version.

      The authors have taken all of my comments into account and have revised their manuscript to my satisfaction.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Editors for the positive assessment on our manuscript. We also thank the Reviewers for their positive remarks and constructive comments. Based on the Reviewers’ feedback, we have conducted additional experiments and provided supporting data to address Reviewers’ comments. Particularly, we provided quantitative measurement for rotational polarity of ependymal cells in Agbl5<sup>M1/M1</sup> mutants and assessed the microtubule polarization. We quantified the intensity of apical actin network in ependymal cells to strength the role of CCP5 in organizing actin network. Using scanning electron microscopy, we demonstrated the affected polarity of trachea multicilia in Agbl5<sup>M1/M1</sup>. We co-immunostained ependymal cilia with GT335 and acetylated tubulin to address the effects on their length in cilia in the mutant. We assessed the presence and length of primary cilia in ependymal cell progenitors to identify their potential contribution to the defective polarity in Agbl5<sup>M1/M1</sup> ependymal cells. We feel that these revisions have much strengthened this MS.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights.

      We thank the Reviewer’s positive comments. To address the molecular insights of the dysregulated planar cell polarity (PCP) in Agbl5<sup>M1/M1</sup> ependyma, we have conducted additional experiments to assess the microtubule polarization in ependymal cells (Figure 7O-P). We quantified the intensity of actin networks around BB patches to better understand how it is affected in the ependyma of the mutants and contributes to the dispersion of BBs (Figure 4M-N), (Please see Recommendations for the authors).

      We also assessed trachea multicilia in Agbl5<sup>M1/M1</sup> mutants using SEM and found that the polarity of trachea multicilia was affected as well (Figure S2).

      Reviewer #2 (Public review):

      Summary:

      This study analyzed the consequences of Agbl5 mutation on ependymal cell development and function. The authors first characterize their mutant mouse line reporting a reduced lifespand and severe hydrocephalus. Next, they report a defect in ependymal cell cilia number and motility. They provide evidence for impaired basal body organisation and cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicates Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype is incomplete:

      We thank the Reviewer’s constructive comments. We have performed additional quantitative analysis of the phenotypes in Agbl5<sup>M1/M1</sup> that we feel strengthen this study.

      Figure 3G - the sequence from the movie is not really informative. Providing beating frequencies as quantification of the data would be more informative.

      We have provided the beating frequency as well as the mean vector length of cilia beating directions (that reflects the coordination of cilia) in Figure 3H and 3I respectively in the revised manuscript.

      Figure 3 - the quantification of actin network would strengthen the message.

      We agree with the Reviewers. We have quantified the total intensity of actin around BBs and the actin intensity normalized to signals of the BB marker (CEP164). The data have been provided in Figure 4M and 4N respectively. The quantitative analysis showed that both the total intensity of apical actin network and the intensity of F-actin per BB are reduced in Agbl5<sup>M1/M1</sup> ependymal cells compared to that in wild-type mice, suggesting that CCP5 is involved in organizing actin network around BB. This analysis certainly improves the clarity of this message.

      Lines 219 -220 - the authors conclude «Taken together, in Agbl5M1/M1 ependymal cells, the expression of genes promoting multiciliogenesis were not impaired but certain proteins associated with differentiated ependymal cells are not properly expressed». However, they do not assess gene but protein expression (IF). In addition, their quantification shows differences in the number of FoxJ1 positive cells which indeed is an impaired expression.

      We will clarify this statement and emphasize the number of FoxJ1-positive cells.

      Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting for the authors to check whether the subapical network of microtubules is glutamylated or not during ependymal cell differentiation and how this network is affected in their mutants.

      We thank the Reviewer’s constructive comments. We conducted an immunostaining on whole-mount lateral walls of lateral ventricles for GT335 and Centrin1, the position of the latter being used to localize the subapical layer. While the GT335 signal in multicilia is increased in Agbl5<sup>M1/M1</sup> ependyma (Figure S8E), its signals underneath BBs are not much different between the mutant and wild-type (Please see Figure S8C, D, G, H).

      Showing the data mentioned in the discussion on Cep110 would be a nice addition to the paper.

      These data have been provided in Supplementary Figure S9.

      Line 354: "The latter serves as a component of tissue polarity that is required for asymmetric PCP protein localization in each cell (Boutin et al., 2014; Vladar et al., 2012)." The cited reference did not demonstrate that this microtubule network is required for asymmetric PCP localization.

      We thank the Reviewer for critical reading. The cited reference (Bountin et al., 2014) has been removed.

      Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      Weaknesses:

      The manuscript is well-written but lacks specific interpretations of the results presented. Further experiments are needed to be fully convincing.

      We thank the Reviewer’s comments. We have performed further analysis and conducted additional experiments to strengthen this study.

      (1) We have quantified the intensity of actin staining around BB patches and its intensity relative to the number of BBs to assess to which extent the actin networks in Agbl5<sup>M1/M1</sup> ependymal cells are affected (please refer to the above response to the comments of Reviewer 2#). The results were shown in Figure 4M-N.

      (2) We Co-stained tdTomato with an ependymal cell-specific markers to strengthen the expression of Agbl5 in ependymal cells (please see Figure 6C-E).

      (3) We have conducted co-immunostaining of GT335 and Ac-Tub and compared the length of their signals in ependymal multicilia between WT and Agbl5<sup>M1/M1</sup> mice (please see Figure 6O, P, R, S).

      (4) We quantified the area of ependymal cells in the wild-type and Agbl5<sup>M1/M1</sup> mice. Indeed, the area of ependymal cells is increased in the mutants. However, the primary cilia are present in the ependymal cell progenitors of Agbl5<sup>M1/M1</sup> mice and have similar length with that in the wild-type (Please see Figure 7M, N and our response to this point below).

      (5) We performed additional analysis to address the affected rotational polarity in the Agbl5<sup>M1/M1</sup> mutant mice (please see Figure 3I, Figure 7E).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors showed that the actin networks were severely affected, leading to impaired stability of basal bodies and that the intensity and length of acetylated tubulin signal in the multicilia were dramatically reduced in AGBL5M1/M1mutant mice (Figures 3 and 5). Data also suggested the dysregulation of planar cell polarity. Are expression and localization of other planar cell polarity proteins such as tyrosinated tubulin and Fzd6 affected in mutant mice?

      We thank the Reviewer’s recommendations. We have assessed the expression of tyrosinated tubulins and found they are similarly polarized in ependymal cells from wild-type and Agbl5<sup>M1/M1</sup> mice. The results are presented in Figure 7O, P in the revised MS. We also tried to assess the expression of Fzd6. However, with the antibody we tested, Fzd6 signals were not convincing. Therefore, we prefer to not showing the results and drawing a conclusion on it.

      (2) The phenotype of multiciliated cells in tracheas should also be examined in mutant mice. It is important to elucidate whether AGBL5 commonly functions in multiciliated cells of other organs.

      We thank the Reviewer’s suggestion. We have assessed the multicilia in the tracheas of P30 mice using scanning electron microscopy. Indeed, unlike the multicilia in wild-type mice that orientate to the same direction, those in the tracheas of Agbl5<sup>M1/M1</sup> mice often radiate to different directions in individual cells (Figure S2). Therefore, Agbl5 appears commonly involved in the alignment of multicilia.

      (3) According to Figure 1B, AGBL5 is highly expressed in the brain. Which cells in the brain express it besides ependymal cells?

      Based on the localization of tdTomato tracer engineered in Agbl5 mutant alleles (Figure 5B), Agbl5 is broadly expressed in the brain, including most if not all neurons, but its expression is much weaker in the subventricular zone (Please see Figure 5B). We clarified this in the revised MS.

      (4) From a mechanistic point of view, it is necessary to identify binding proteins with the N-domain of AGBL5 and perform functional analyses.

      We agree with the Reviewer. We feel that identification of the binding partners of CCP5 N-domain and functional analysis may be more suitable to go along with other mechanistic analysis on the function of CCP5 in ependymal cell polarities in our future study.

      Reviewer #2 (Recommendations for the authors):

      (1) Movie 3: The authors could comment on beating direction that seems impaired at the cell scale here, analysis of rotational polarity would be a plus.

      We thank the reviewer’s recommendation. We have analyzed the beating directions of cilia in individual cells and presented their consistency in each cell using mean vector length. These results indeed demonstrated defective rotational polarity in the cell level in Agbl5<sup>M1/M1</sup> mice (please refer to Figure 3I). We also analyzed the beating directions of ependymal multicilia in earlier stage in tissue level (Figure 7E). The mean vector length of cilia beating direction in Agbl5<sup>M1/M1</sup> mice is significantly reduced compared to that in wild-type, suggesting an aberrant rotational polarity in the tissue level in the mutant (Figure 7E).

      (2) Line 166 : ref to Werner et al., 2011 is not correct (no ependymal cells in that paper).

      We thank the reviewer’s critical reading. This reference has been removed.

      (3) Figure S4: B and D look similar picture to me same for C and F.

      We apologize for using the wrong images in this Figure. It has been corrected (Revised Figure S5).

      (4) Line 328: "Therefore, CCP5 apparently contributes to the establishment of both translational and tissue polarities in ependymal cells." Should be rephrased since translational polarity is also a tissue-level parameter which is the coordinated positioning of the ciliary patch. Cf Mirzadeh et al., 2010; Boutin et al., 2014.

      We thank the Reviewer’s comments. The sentence has been rephrased. This concept has been clarified where else needed in the revised manuscript. 

      (5) Line 348: "Planar cell polarity (PCP) pathway is essential for the establishment of rotational and tissue polarities in ependymal cells" Rotational polarity also has a tissular component (ie coordination of beating direction across tissue which is reflected by coordination of basal body polarities across tissue).

      We thank the Reviewer’s comments. We have clarified this point in the revised MS.

      (6) Incomplete bibliography citation (ie Walentek et al. without date).

      We thank the Reviewer’s critical reading. This bibliography citation has been fixed.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 3: The authors assert that the mutant's apical actin networks are significantly disrupted. However, the cell shown in Figure 3Q-R exhibits less compact centrioles than the controls, which could account for the reduction in phalloidin staining. Because centriole dispersion is variable in the mutant, quantifying actin staining in representative cells would be necessary to support such a statement.

      We thank the Reviewer’s comments. To address this concern, we have quantified the total intensity of actin network around BBs as well as the intensity of F-actin signals normalized to the level of immunosignals of BBs ((revised Figure 4M, N) please also refer to our response to Reviewer 1#). The results indicated the intensity of actin signal per BB is reduced in the mutant compared to that of wild-type mice. We feel that this analysis strengthened our statement.

      (2) Figures S3 and 4A-B show that the authors examine tdT expression to show that Agbl5 is expressed in ependymal cells but not in the SVZ. However, the tdT signal intensity is very low, and cells are very dense in this brain region. Double staining with specific markers of ependymal and/or SVZ cells would help convince readers that tdT is not expressed in SVZ cells.

      We agree with the Reviewer that the intensity of tdT signal is low, but broadly detectable in brain. Compared with its expression in ependymal cells, that in SVZ is much lower if any (Figure 4B’). To further confirm the identity of tdT-positive cells along the surface of ventricles, we have co-stained the brain sections of Agbl5<sup>WT/M1</sup> mice for tdT and S100b, a marker of mature ependymal cells (Figure 5C-E). The signal of tdt is colocalized with that of S100b and is much lower in cell layers next to S100b-positive cells.

      (3) Figure 4C-D and S4: The authors demonstrate that the number of FoxJ1+ cells per section increases at P7 (4C-E), while the number of S100β+ cells per mm decreases. Quantifications should be carried out in a similar manner to ensure comparability (number of positive cells per mm). Additionally, it remains unclear how to interpret these results, as S100β and FoxJ1 are two markers of differentiated cells, yet they exhibit opposite trends compared to controls. Is this a direct or indirect effect of Agbl5 mutation? The increase in the number of FoxJ1+ cells is particularly surprising given that the number of GT335 multicilia per mm remains unchanged (Figure 5).

      We agree with the Reviewer that quantifications should be carried out in a similar manner. In the revised MS, the quantification of Foxj1-positive cells is presented in number per mm (Figure 5I). To be noted, the expression of Foxj1 was assessed at P7 when ependymal cells are differentiating. while the expression of S100β was assessed at P17 when ependymal cells are supposed to be fully mature. Although S100b is used as a marker of mature ependymal cells, given its unclear function, we removed the results of S100b-positiving cell counting to avoid confusion in the revised manuscript.

      (4) Figure 5: In this figure, the authors analyze the labeling obtained with GT335, Acetylated Tubulin, and Arl13b antibodies. They show that the area of the cilium labeled by GT335 has increased, while the area labeled by the Acetylated Tubulin antibody has decreased in the knockout (KO) compared to the control. However, the length of the cilia observed through labeling with the Arl13b antibody remains unchanged. These observations are intriguing, but the low-magnification images in Figure 4 do not allow for the differences in ciliary axoneme labeling to be seen. Double GT335/AcTub labeling and higher magnifications are necessary for improved visualization of the differences in labeling along the axonemes.

      We thank the Reviewer comments. We have co-stained the cilia with GT335 and Ac-Tub antibodies, re-quantified cilia length labeled with respective antibodies and provided high magnification images. Please see the revised Figure 6O,P,R,S.

      (5) Figure 6: An analysis of ciliary beats using a high-speed camera shows no difference in ciliary beat frequency between the control and KO groups. At least, 3 animals should be analyzed. According to Figure 5, these findings indicate that the decrease in ciliary acetylation and the increase in ciliary glutamylation do not affect the beat frequency; instead, they disrupt the orientation of the beats. While these results are intriguing, they require further confirmation. Analyzing ciliary beats with a high-speed camera is informative, but at least three animals per genotype should be examined to ensure rigor. Furthermore, if the coordination of ciliary beats is impaired within the cells, this should be validated by double-labeling centrioles and basal feet to demonstrate that the orientation of cilia within the cells is abnormal.

      We thank the Reviewer’s comments. Sections shown in Figure 5 (currently Figure 6) are from P7 mice, while the ciliary beating analysis shown in Figure 6 (currently Figure 7) is from P15 mice. As the PTM changes in cilia were also observed in Agbl5<sup>M2/M2</sup>, we don’t think this is the cause that disrupts the orientation of the beats. The rotational polarity of Agbl5<sup>M1/M1</sup> ependymal cells is affected. Please refer to the analysis in Figure 3I and Figure 7E in the revised manuscript.

      (6) Figure 6F-G: β-Catenin labeling reveals cells of varying sizes in the KO. This phenotype is typical of ciliary mutants that lack primary cilia (Mirzadeh et al., 2010). Hence, it is essential to examine the mutation's impact on the presence, length, and positioning of the primary cilium in ependymal cell progenitors.

      We thank the Reviewer’s constructive comments. We assessed the area of ependymal cells labeled with β-Catenin. Indeed, the ependymal cells in the mutant showed larger area than that of wild-type. The ratio of the area of BB patch over that of cell surface is reduced (please see Figure 7O, P in the revised manuscript). However, primary cilia are present in ependymal cell progenitors in the mutant and exhibit comparable length with those in the wild-type (Figure S8). Due to some technique problems, we were unable to get convincing results from whole-mount ventricle walls for the primary cilium positioning at this time. We speculate that the localization of certain sensory proteins in primary cilia or the positioning of primary cilia might be affected in Agbl5<sup>M1/M1</sup> mice. We discussed this possibility and will certainly systemically assess this intriguing aspect in our future investigation.

      (7) Given the regular beating frequency in the KO at P15, how do the authors explain the complete absence of ciliary beating in the adult? How many animals were analyzed? One would expect ciliary beating to remain unaffected as it was at P15 unless the cilia structure was specifically altered at the adult stage. Is that the case?

      We thank the Reviewer’s critical questions. We do think that the ciliary structure of Agbl5<sup>M1/M1</sup> ependymal cells is likely altered during aging. Given that only Agbl5<sup>M1/M1</sup> but not Agbl5<sup>M2/M2</sup> mice develop hydrocephalus, we speculate the N-domain of CCP5 may contribute to the integrity of ependymal multicilia. We have added this in the Discussion section. For each genotype, 2 mice were analyzed.

      (8) Line 264 of the manuscript: replace intercellular with intracellular.

      It has been revised.

      (9) Indicate the number of animals analyzed in each experiment

      It has been included in figure legends.

    1. eLife Assessment

      This paper addresses a valuable research question on the modest heritability of the brain's response to movie watching, and how heritability varies under different parameters such as regional spatial hyperalignment and BOLD frequency bands. The topic of this paper is of interest to fMRI methodological experts, and potentially to a broader cognitive neuroscience audience, and those with an interest in understanding the heritable sources of individual differences in brain function. Although some of the conclusions could be strengthened by future cross validation studies in independent and larger family-based samples, and through complementary twin/family and SNP-based models, taken altogether, the analyses and results provide convincing evidence for the overall conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales, and more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor our topographic differences and found the relationship between heritability and neural time scales very interesting. The writing is clear and the results are compelling. In general, I don't have many complaints after a couple reads through the manuscript; most of my comments below are relatively minor suggestions and points of clarification.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified:

      On page 16, you compare heritability in functional connectivity (FC) and response time series and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and time-series differences cannot, by definition). This makes me wonder how this connectivity result would change if you used intersubject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript-but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. You could even color the data points according to networks like in Figure 3C. (You also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      On page 9, if I understand correctly, you regress the vector of ISC values across parcels out of the vector of heritability values across parcels and then plot the residual heritability values. Do you center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can you explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could you go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?)

      On page 4 (line 155), you say "we shuffled dyad labels"-is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure your approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned at line 189?

      I found panel A in Figure 4 to be a little bit misleading because your parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels you're performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248-259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Comments on revised version.

      The authors have adequately addressed my previous comments. This is a strong contribution: the methods are sophisticated, the statistical treatment is rigorous, and the results are quite interesting/compelling. I'm happy to endorse the revised manuscript as a finalized version.

      Just to confirm: The subjects watched all different movies across the two days, right? For a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

    3. Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      Comments on revised version.

      The whole manuscript has been improved a lot, and the concerns have been clarified.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained by local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales and is more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor out topographic differences, and I found the relationship between heritability and neural time scales very interesting. The writing is clear, and the results are compelling.

      We thank Reviewer 1 for their kind words and enthusiastic support of our manuscript.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified.

      (1) On page 16, the authors compare heritability in functional connectivity (FC) and response time series, and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and timeseries differences cannot, by definition). This makes me wonder how this connectivity result would change if the authors used inter-subject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript, but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      We agree that investigating the heritability of ISFC (or stimulus-driven functional connectivity) would make for a very interesting future direction. Ultimately, we chose to analyze FC (vs. ISFC) profiles to allow for direct comparison with the sizable existing literature on the heritability of FC (such as in our Movie vs. Rest FC analysis) and decided to refrain from analyzing ISFC data in order to keep the present manuscript focused. ISFC analysis of this dataset will be a focus of future work.

      (2) The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. The authors could even color the data points according to networks, like in Figure 3C. (They also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      We thank R1 for this helpful suggestion- we originally set the y-axis limits to r = 1 in order to facilitate comparison between ISC (Fig. 1B) and FC profile (Fig. 6B) similarity, but we agree that this renders the group differences harder to discern and have updated the plot accordingly (along with thicker lines to enhance readability). We prefer to keep the line plots in the main body as they allow for direct comparison of all three groups on the same plot, but we have included the scatter plot version in Fig. S2 for those who are interested.

      (3) On page 9, if I understand correctly, the authors regress the vector of ISC values across parcels out of the vector of heritability values across parcels, and then plot the residual heritability values. Do they center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can the authors explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could they go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?).

      We indeed included an intercept in this model using MATLAB’s fitlm function. This means that the model estimates the best-fitting line of the following form: heritability<sub>i</sub>=β0+β1ISC<sub>i</sub> +ε<sub>i</sub>. We agree that the interpretation of these ε<sub>i</sub> values and alternative approaches to controlling for ISC should be clarified. As such, we have added the following passages to the text:

      Methods: “Because the heritability of ISC is constrained by the degree of synchronization in a given area, we also sought to identify areas in which BOLD time courses were more/less heritable than would be expected based on ISC alone by fitting a linear model of the form heritability<sub>i</sub>=β0+β1ISC<sub>i</sub>+ε<sub>i</sub> and plotting the residuals. Regarding alternative approaches to controlling for ISC, although the heritability model introduced by Ge et al. allows for the inclusion of covariates defined at the subject level (e.g., age), it does not allow for covariates that are defined at the dyad level (e.g., pairwise ISC).”

      Results: “Here, negative values in the residual map indicate parcels where heritability is lower than expected based on ISC, while positive values indicate higher-than expected heritability.”

      (4) On page 4 (line 155), the authors say "we shuffled dyad labels"- is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure their approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned in line 189?

      Briefly, shuffling the kinship matrix involved permuting the rows and columns of the matrix in the same manner (also known as the quadratic assignment procedure), whereas shuffling the dyad labels involved random permutations of the three group labels (MZ, DZ, unrelated), which could not be done through matrix operations as the age- and gender matching precluded the use of a complete similarity matrix. However, given concerns raised by Reviewer 2, we have removed our significance claims from this (and similar) sections, which we discuss in more detail in response to Reviewer 2’s weakness A.

      (5) I found panel A in Figure 4 to be a little bit misleading because their parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels they are performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      We agree that our efforts to present a simplified depiction of hyperalignment may mislead less familiar readers and have amended Fig. 4A according to this suggestion. We have also added text to the methods section (below) to clarify that the outputs of hyperalignment are time series that reflect linear combinations of other voxels’ time series from that parcel.

      “This approach independently transforms each subject's data within discrete anatomical parcels into the common space, yielding functionally aligned vertex time series that are calculated as weighted linear combinations of the original time series from all other vertices within that same parcel for that subject.”

      (6) I believe the subjects watched all different movies across the two days, however, for a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

      We agree that this would be helpful and have added the following text to the relevant sections:

      “All clips were only viewed once by each subject, with the exception of the brief montage which was included at the end of each of the four runs for test-retest purposes.”

      “To characterize the heritability of brain responses to complex stimuli, we used 7T fMRI data from 178 HCP Young Adult subjects acquired across two days (using two largely non-overlapping sets of movie stimuli, see Methods)…”

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to estimate the heritability of brain activity evoked from a naturalistic fMRI paradigm. No new data were collected; the authors analyzed the publicly available and well-known data from the Human Connectome Project. The paper has 3 main pieces, as described in the Abstract:

      (1) Heritability of movie-evoked brain activity and connectivity patterns across the cortex.

      (2) Decomposition of this heritability into genetic similarity in "where" vs. "how" sensory information is processed.

      (3) Heritability of brain activity patterns, as partially explained by the heritability of neural timescales.

      Strengths:

      The authors investigate a very relevant topic that concerns how heritable patterns of brain activity among individuals subjected to the same kind of naturalistic stimulation are. Notably, the authors complement their analysis of movie-watching data with resting-state data.

      Weaknesses:

      The paper has numerous problems, most of which stem from the statistical analyses. I also note the lack of mapping between the subsections within the Methods section and the subsections within the Results section. We can only assess results after understanding and confirming the methods are valid; here, however, Methods and Results, as written, are not aligned, so we can't always be sure which results are coming from which analysis.

      (A) Intersubject correlation (ISC) (section that starts from line 143): "We used nonparametric permutation testing to quantify average differences in ISC for each parcel in the Schaefer 400 atlas for each day of data collection across three groups: MZ dyads, DZ dyads, and unrelated (UR) dyads, where all UR dyads were matched for gender and age in years." ... "some participants contributed to ISC values for multiple dyads (thus violating independence assumptions)"

      This is an indirect attempt to demonstrate heritability. And it's also incorrect since, as the authors themselves point out, some subjects contribute to more than one dyad.

      Permutation tests don't quantify "average differences", they provide a measure of evidence about whether differences observed are sufficient to reject a hypothesis of no difference.

      Matching subjects is also incorrect as it artificially alters the sample; covarying for age and sex, as done in standard analyses of heritability, would have been appropriate.

      It isn't clear why the authors went through the trouble of implementing their own nonparametric test if HCP recommends using PALM, which already contains the validated and documented methods for permutation tests developed precisely for HCP data.

      The results from this analysis, in their current form, are likely incorrect.

      We appreciate that permutation tests do not quantify average differences and intended to write “We used non-parametric permutation testing to quantify [the significance of] average differences…”. Our intention with this analysis was not to demonstrate heritability, but rather to quantify group differences in ISC in a manner that is interpretable for readers who are unfamiliar with h<sup>2</sup> (e.g., “identical twins’ BOLD time courses were 59% more similar than those from pairs of unrelated individuals”) and motivate the formal heritability analysis used later in the paper. Indeed, all of the heritability analyses in this paper leveraged a validated multidimensional heritability method first introduced by Ge et al. (2016) and used by many other investigators since then. Furthermore, we covaried for age and sex at the subject level in all our heritability analyses, and always tested the significance of these heritability values using a validated permutation procedure (the quadratic assignment procedure; Hubert & Schultz, 1976) that respects the non-independence of dyadic data.

      Regarding the shuffling procedure used for Figure 1, while PALM is the standard for univariate, subject-level GLMs in the HCP pipeline and can accommodate nested designs (i.e., subjects within families), it is not designed to handle the unique relational dependencies of dyadic ISC analysis (i.e., the same subject contributing to multiple dyads). Although the element-wise resampling approach was the most appropriate approach available, it is known to inflate the false positive rate (Chen et al., 2016; doi:10.1016/j.neuroimage.2016.05.023); given that this analysis was simply meant to motivate our later hypothesis testing heritability analyses, we have removed significance claims from this section of the manuscript. Still, we emphasize that this has no bearing on the validity of our conclusions which were supported by our formal heritability analyses; throughout our paper we have correctly used the appropriate methods to back the stated claims.

      (B) Functional connectivity (FC) (section that starts from line 159): Here the authors compute two 400x400 FC matrix for each subject, one for rest, one for movie-watching, then correlate the correlations within each dyad, then compared the average correlation of correlations for MZ, DZ, and UR. In addition to the same problems as the previous analysis, here it is not clear what is meant by "averaging correlations [...] within a network combination". What is a "network combination"? Further, to average correlations, they need to be r-to-z transformed first. As with the above, the results from this analysis in its current form are likely incorrect.

      We regret that R2 had difficulty understanding our analysis and have added the following text to the relevant Methods section to clarify our approach:

      “For example, there are 16 parcels in the Kong et al. Auditory network and 17 parcels in the Language network, so the FC profile for a given subject’s Auditory-Language network combination consists of the (16 * 17 =) 272 correlation coefficients between all unique pairs of one parcel from each network.”

      As we stated in the previous Methods paragraph, “All Pearson r values in this and all other analyses were Fisher z-transformed before averaging (and converted back to Pearson r for visualization)”. Thus, contrary to the reviewer’s assertion, these analyses were performed correctly. Once again, we emphasize that this analysis was not intended to demonstrate heritability, but rather to describe group differences in FC in familiar units.

      (C) ISC and FC profile heritability analyses (section that starts from line 175): Here, the authors use first a valid method remarkably similar to the old Haseman-Elston approach to compute heritability, complemented by a permutation test. That is fine. But then they proceed with two novel, ill-described, and likely invalid methods to (1) "compare the heritability of movie and rest FC profiles" and (2) to "determine the sample size necessary for stable multidimensional heritability results". For (1), they permute, seemingly under the alternative, rest and movie-watching timeseries, and (2), by dropping subjects and estimating changes in the distribution.

      The (1) might be correct, but there are items that are not clearly described, so the reader cannot be sure of what was done. What are the "153 unique network combinations"? Why do the authors separate by day here, whereas the previous analyses concatenated both days? Were the correlations r-to-z transformed before averaging?

      The (2) is also not well described, and in any case, power can be computed analytically; it isn't clear why the authors needed to resort to this ad hoc approach, the validity of which is unknown. If the issue is the possibility that the multidimensional phenotypic correlation matrix is rank-deficient, it suffices that there are more independent measurements per subject than the number of subjects.

      Regarding (1), we have clarified in section 2.6 that the 153 unique network combinations reflect each unique pair of 17 Kong networks. All of our analyses, including this one, were performed separately for each day of data collection, as we state throughout the paper and visualize in our figures (although we acknowledge that, on some occasions, we [conservatively] performed FDR-correction on a combined set of p-values, as discussed in our response to K). Given that the null hypothesis for this analysis is that rest FC and movie FC are equally heritable, we are not sure why permuting rest and movie FC matrices would be invalid. All Pearson r values were z-transformed before averaging, as we stated in our paper.

      Regarding (2), we included this analysis in response to editorial concerns that our heritability analyses were not sufficiently powered, and we chose this approach because it serves as a simple way to demonstrate the stability of our results at various sample sizes whose validity is self-evident. Furthermore, this sort of subsampling approach has been used many times before in our field (e.g., Marek et al., 2022) and others (e.g., Manyara et al., 2024) to demonstrate the sample-size dependence and stability of statistical effects. We have added text explaining this to the relevant Methods section (2.6).

      (D) Frequency-dependent ISC heritability analysis (from line 216): Here, the authors decompose the timeseries into frequency bands, then repeat earlier analyses, thus bringing here the same earlier problems and questions of non-exchangability in the permutations given the dyads pattern, r-z transforms, and sex/age covariates.

      We did not use dyadic permutation testing for any of the frequency-dependent ISC analyses; rather, we used the jackknife SEMs to compare heritability across frequency bands and have added an explicit description of this to section 2.7. We have addressed the r-z transform and covariate concerns in previous comments.

      (E) FC strength heritability analysis (from line 236): Here, the authors use the univariate FC to compute heritability using valid and well-established methods as implemented in SOLAR. There is no "linkage" being done here (thus, the statement in line 238 is incorrect in this application. SOLAR already produces SEs, so it's unclear why the authors went out of their way to obtain jackknife estimates. If the issue is non-normality, I note that the assumption of normality is present already at the stage in which parameters themselves are estimated, not just the standard errors; for non-normal data, a rank-based inversenormal transformation could have been used. Moreover, typically, r-to-z transformed values tend to be fairly normally distributed. So, while the heritabilities might be correct, the standard errors may not be (the authors don't demonstrate that their jackknife SE estimator is valid). The comparison of h2 between dyads raises the same questions about permutations, age/sex covariates, and r-z transforms as above.

      We used jackknife SEs for these analyses to maintain consistency with the multidimensional heritability package used here, which only outputs jackknife SEs. We note that this jackknife approach (and the corresponding multidimensional heritability analysis) was detailed in prior work (Anderson et al., 2021), and that the leave-one-family-out jackknife has a long history of being used to estimate SEs in heritability studies, especially when working with smaller samples (Knapp et al., 1989). We are also not sure what “the comparison of h2 between dyads” means- heritability cannot be compared “between” dyads; rather, it is defined across dyads.

      (F) Hyperalignment (from line 245): It isn't clear at this point in the manuscript in what way hyperalignment would help to decompose heritability in "where vs. how" (from the Abstract). That information and references are only described much later, from around line 459. The description itself provides no references, and one cannot even try to reproduce what is described here in the Methods section. Regardless, it isn't entirely clear why this analysis was done: by matching functional areas, all heritabilities are going to be reduced because there will be less variance between subjects. Perhaps studying the parameters that drive the alignment (akin to what is done in tensor-based and deformation-based morphometry) could have been more informative. Plus, the alignment process itself may introduce errors, which could also reduce heritability. This could be an alternative explanation for the reduced heritability after hyperalignment and should be discussed. An investigation of hyperaligment parameters, their heritability, and their co-heritability with the BOLD-phenotypes can inform on this.

      To help set up our hyperalignment analyses, we have added text to the introduction explaining how hyperalignment would help to decompose heritability. The description in the Methods section included a reference to Bazeille et al., 2021, in which the hyperalignment method used here is discussed in detail. Still, we have added citations to additional papers (also cited in the Bazeille et al. paper, and elsewhere in our paper) in case that might be helpful. We note that it is not the case that all heritabilities were reduced by hyperalignment- as can be seen in Figs. 4D, 8A, and S15, hyperalignment did increase heritability in some voxels and network combinations. This would be expected under the alternative (albeit unlikely) hypothesis that functional topographies are not heritable, such that topographic variation between related individuals would obscure similarities in their (heritable) topography-independent brain responses. Recognizing that this alternative is unlikely, we believe the main novelty of this analysis comes from the magnitude of the hyperalignment effect (up to 40% of brain-wide heritability) and its spatial pattern (e.g., larger heritability decreases in visual vs. auditory cortex, the opposite of our NT result).

      We agree that we would see lower post-hyperalignment heritability if the alignment process itself introduced errors/noise, but this would be deeply surprising as hyperalignment increases ISC by design (and errors/noise could only decrease ISC). To demonstrate this, we have added Figure S7 which shows that (as expected) ISC across all voxels and subject pairs increases after hyperalignment (and that this increase is larger when hyperalignment is performed in larger parcels). Given that hyperalignment increased ISC, and that it is blind to twin status, we are unsure how it could have introduced errors that would have confounded this result.

      (G) Relationships between parcel area and heritability (from line 270): As under F), how much the results are distorted likely depends on the accuracy of the alignment, and the error variance (vs heritable variance) introduced by this.

      We agree that alignment accuracy could potentially impact parcel-level differences in how much heritability changes following hyperalignment, and we included the frequency dependent h<sup>2</sup><sub>residuals</sub> (controlling for differences in ISC) in Fig. 3 for this reason, as more accurate hyperalignment should result in greater increases in ISC, raising the heritability ceiling. We note that we observe similar relationships between parcel rank and frequency dependent changes in these residualized maps, suggesting that our parcel-level differences are not simply the result of better alignment in more sensory parcels.

      (H) Neural timescale analyses (from line 280): Here, a valid phenotype (NT) is assessed with statistical methods with the same limitations as those previously (exchangability of dyads, age/sex covariates, and r-z transforms). NT values are combined across space and used as covariates in "some multivariate analyses". As a reader, I really wanted to see the results related to NT, something as simple as its heritability, but these aren't clearly shown, only differences between types of dyads.

      We have addressed the exchangeability, covariates, and r-z transform comments above (in A). As we explained for our FC strength analyses, we are underpowered to evaluate the heritability of unidimensional traits (like the heritability of NT magnitude), and the heritability of a closely-related measure (BOLD turnover magnitude) has already been established in a larger sample of HCP subjects (https://doi.org/10.1152/jn.00402.2022). Still, we agree that more results related to the heritability of NTs would be of interest to our readers. As such, we have added an analysis in section 3.4 quantifying the heritability of multivariate NT topographies and used SOLAR to quantify the heritability of NT magnitudes, with the disclaimer that this and similar analyses are underpowered (hence the large difference in day 1 and day 2 heritability effect sizes). We also removed significance claims for the dyadic NT similarity analysis.

      (I) Significance testing for autocorrelated brain maps and FC matrices (from line 310): Here, the authors suddenly bring up something entirely different: reliability of heritability maps, and then never return to the topic of reliability again. As a reader, I find this confusing. In any case, analyses with BrainSMASH with well-behaved, normally distributed data are ok. Whether their data is well behaved or whether they ensured that the data would be well behaved so that BrainSMASH is valid is not described. As to why Spearman correlations are needed here, Mantel tests, or whether the 1000 "surrogate" maps are valid realizations of the data under the null, remains undemonstrated.

      We brought up reliability in this section because we show the reliability of our results across the two days of data collection several times in the paper. R2 is correct to point out that BrainSMASH was validated using normally distributed brain maps, and although some of our brain maps contain normally distributed values, others are right skewed (due largely to the fact that many voxels/parcels exhibit low ISC while visual/auditory areas have very high ISC). In preparing our original manuscript, we visualized BrainSMASH’s variogram outputs for one of the most skewed inputs (vertex-wise BOLD time course heritability) and found that the autocorrelation structures of the empirical and null maps were well-matched. We did not include this in the original manuscript as it is not commonplace in the field to report the variograms, see Author response image 1. Furthermore, our use of Spearman (vs. Pearson) correlations renders these distributional differences less relevant, as the Spearman correlation transforms all inputs to a uniform distribution. To empirically check that these distributional differences do not bias our results, we retested the significance of all brain map associations using the spin test (10.1016/j.neuroimage.2018.05.070), an alternative method that does not assume normally distributed inputs, and obtained identical p-values for all analyses (P<.001 in all cases).

      Author response image 1.

      (J) Global signal was removed, and the authors do not acknowledge that this could be a limitation in their analyses, nor offer a side analysis in which the global signal is preserved.

      Although we agree that GSR is a contentious preprocessing step for certain analyses, it has explicitly been shown to increase ISC signal-to-noise without compromising FC fingerprints (Graff et al., 10.1016/j.dcn.2022.101087), and it is uncommon to perform ISC analyses with and without GSR. Still, we have added additional text to our Methods section explaining our rationale for using GSR and that this could affect our results. We also re-ran our main analysis (BOLD time course heritability) with and without GSR and found that GSR had little impact on our results; we have included this in our manuscript as Fig. S4.

      Specifically, we see that GSR resulted in a slight increase in heritability (average Day 1 h<sup>2</sup> with/without GSR = .064/.060; Day 2: .068/.061) and almost no effect on the spatial pattern of our results (With GSR/without GSR Spearman ρ = .99, P<sub>brainSMASH</sub> < .001 on both Day 1 and Day 2).

      (K) FDR is used to control the error rate, but in many cases, as it's applied to multiple sets of p-values, the amount of false discoveries is only controlled across all tests, but not within each set. The number of errors within any set remains unknown.

      We agree that the FDR usage in our original manuscript was inconsistent, in that for two analyses we FDR-corrected p-values from the two days of data collection together (instead of correcting p-values from each day separately and reporting voxels/parcels/etc. that were significant at q<.05 on both days, as in the rest of our analyses). We note that both approaches are more conservative than reporting significant results at q<.05 separately; regardless, to maintain consistency we have updated all analyses such that FDR correction is always performed separately for each day of data collection.

      (L) Generally, when studying the heritability of a trait, the trait must be defined first. Here, multiple traits are investigated, but are never rigorously defined. Worse, the trait being analyzed changes at every turn.

      Here, we analyze the heritability of movie-evoked BOLD time courses (Figures 1-5) as well as FC profiles (Figures 6-8). We defined FC profiles in our Introduction as an individual’s pattern of pairwise FC strengths (and further detailed how we quantified FC profiles in the relevant Methods section), and believe that “BOLD time course” is a well understood phrase in the field and does not need to be further defined. We also used hyperalignment to decompose the heritability of these traits into topography-dependent and independent portions, and (new to this version) also explicitly quantify the heritability of neural timescales, which we defined as the AUC of the ACF until the first negative ACF value in both the relevant Results and Methods sections.

      To make this clearer, we have modified the last paragraph of our Introduction to begin with:

      In the present work, we address these questions by analyzing 7T fMRI recordings of a twin sample acquired by the Human Connectome Project (Van Essen et al., 2013) to quantify the heritability of two distinct high-dimensional traits—stimulus-evoked BOLD time courses and functional connectivity profiles—across the cortex.

      Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      We believe that this question is getting at the difference between pairwise ISC (i.e., correlating one BOLD time course from one subject with that from another subject) and leave-one-subject-out ISC (i.e., correlating one BOLD time course from one subject with the corresponding average time course across all other subjects). We chose to use the pairwise ISC method because it allows us to capitalize on the information contained in the n<sup>2</sup> pairwise ISC matrix (whereas the other approach averages out meaningful information to yield a n<sup>1</sup> ISC matrix) and leverage a more sophisticated multidimensional heritability approach. Also, the leave-one-subject-out approach introduces additional issues re: handling family-level data (e.g., should we include a subject’s twin in the leave-one-subject-out average? If so, how should we handle subjects who don’t have a twin in the dataset, as averaging data from different numbers of subjects will lead to different ISC magnitudes? etc.).

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      We report p-values for heritability throughout our paper (e.g., stating that BOLD time courses are significantly heritable in 99% of parcels in Figure 2), and we believe that the reliability of our spatial maps across days of data collection (also quantified with p-values) further demonstrates the trustworthiness of our results. Finally, as we demonstrate in Figure S5, our sample size is more than sufficient to reliably detect small effects.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      In addition to driving shared neuronal responses (which are captured in BOLD signal oscillations <.1 Hz or so), movies also elicit shared cardiac, respiratory, and motion responses across participants at higher frequencies. Although we used a relatively conservative denoising approach here, we believe some of these non-neuronal signals are still present in our data; alternatively, it is also possible that these signals reflect “fast” BOLD responses at >.15 Hz (as discussed in 10.1016/j.neuroimage.2021.118658). In any case, the fact that information in this frequency band is considerably less heritable than information in slower frequency bands supports the idea that this band is noisier and suggests that our heritability results are driven by canonical neuronal activity-related BOLD signals.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      Although the decrease of 0.025 is small, we note that this constitutes around ~50% of BOLD time course heritability in some voxels (seen in comparison to Fig. 4C), and the spatial pattern of this result is quite consistent across days of data collection, indicating its reliability. Furthermore, the whole-brain distributions of results shown in Fig. 5B are clearly skewed towards negative values, indicating that controlling for NT partially reduces (or “explains”) BOLD time course heritability. Still, we agree that showing raw h<sup>2</sup> values in addition to the difference maps would be helpful for some readers and have added a corresponding supplementary figure (S12) which shows these.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      We did consider adding standard errors for these heritability estimates, but found that visualizing standard errors for each of the 153 unique network combinations in our heatmaps rendered the visualizations difficult to parse, and given that our hypotheses concerned global (e.g., hyperaligned vs. MSM-aligned) or network-level (e.g., sensory vs. associative) patterns, we focused on calculating standard errors/p-values for these analyses (although we note that dyad-level standard errors can be found in Fig. 6B, where they are clearly marginal compared to the group effects).

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      We agree that this result was relatively under-explored in our Discussion section and have added additional text (lines 851-855) to connect this result to recent work on arousal-dependent uniqueness of FC.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Do the authors have any ideas why we see this hotspot of heritability in pMTG/LOTC? It really jumps out in Figure 1A and Figure 2. The more posterior sensory MT+ area seems to drop when regressing out ISC in Figure 2B, but this pMTG area stays hot. Is there anything special about this kind of multimodal biological motion/action observation / social perception area (Pitcher & Ungerleider, 2021)? I don't think this is necessary to discuss in the manuscript, but I'm curious if the authors have any speculation.

      We are not certain as to why BOLD time courses in this parcel are particularly heritable- although this area is associated with biological motion, that particular function tends to be more right lateralized, and here we see nominally higher heritability in the left hemisphere. Per a Neurosynth review (and consistent with the left lateralization), we believe this may have more to do with speech processing, but a more definitive answer will require further investigation.

      (2) Page 3, line 127: "More information on these clips"-it might be worth saying a little bit more here just to make sure people understand that these are audiovisual clips, they include language, they're long enough to convey meaningful social and narrative information, etc.

      We agree and have added additional details on the clip composition to the relevant methods paragraph.

      (3) Figure 1 caption: can you add a sentence reminding readers what's going on with Day 1 and Day 2?

      We thank R1 for this suggestion and have added a sentence to this effect at this location.

      (4) Page 9, line 379: "although these more associative parcels do not encode a substantial amount of stimulus-specific information"-is this really true? I suspect these association areas still have decent ISCs, even if there are many processing stages downstream of the raw stimulus.

      Although these parcels are not the most synchronized by the stimulus, we agree that it is unfair (and vague) to say that they do not encode a substantial amount of stimulus-specific information. We have edited this sentence to make a more specific claim and highlight the relatively lower ISC in these parcels vs. more unimodal sensory areas.

      (5) Page 9, line 417: Can you unpack a bit more what you mean by "supra-BOLD frequency band"?

      Here, we refer to the fact that BOLD signals resulting from neuronal firing events have frequencies below ~.15 Hz (Josephs and Henson, 1999). We have added additional text and the Josephs and Henson citation to this line to further unpack this point.

      (6) Page 18, line 695: This discussion of how attention and gaze might partly shape response time series reminded me of recent work by Borovska & de Haas (2024)-might be worth citing.

      We are grateful to R1 for alerting us to this very relevant work and have included a reference to it in our discussion.

      (7) Page 19, line 755: I'm not sure I'd describe the hyperalignment results here as a "deleterious effects [on] heritability"-my reading was that hyperalignment allows you to say something more specific about heritability of function by allowing you to effectively factor out heritability effects that reduce to individual differences cortical topography; this seems like a good thing!

      We agree that “deleterious” was a poor word choice given its negative connotation, and have edited this sentence to read:

      “With this in mind, future studies investigating genetic correlations between brain function and behavioral variables may benefit from hyperalignment, as it can factor out individual-specific cortical topography and thus yield more precise estimates of functional heritability.”

      (8) I would love to see a ventral view in some of these plots! Not asking you to recreate the figures, but the ventral temporal cortex is an area of interest for many folks in the movie fMRI space (e.g., Haxby et al., 2011).

      We agree that ventral views would be of interest to some readers and have added the corresponding maps for our main results in supplementary figures S3 and S9.

      References:

      Borovska, P., & de Haas, B. (2024). Individual gaze shapes diverging neural representations. Proceedings of the National Academy of Sciences, 121(36), e2405602121. https://doi.org/10.1073/pnas.2405602121

      Haxby, J. V., Guntupalli, J. S., Connolly, A. C., Halchenko, Y. O., Conroy, B. R., Gobbini, M. I., Hanke, M., & Ramadge, P. J. (2011). A common, high-dimensional model of the representational space in human ventral temporal cortex. Neuron, 72(2), 404416. https://doi.org/10.1016/j.neuron.2011.08.026

      Pitcher, D., & Ungerleider, L. G. (2021). Evidence for a third visual pathway specialized for social perception. Trends in Cognitive Sciences, 25(2), 100-110. https://doi.org/10.1016/j.tics.2020.11.006

      Reviewer #2 (Recommendations for the authors):

      (1) To address the common core analytical problems listed under A), B), C), D), E), and basically throughout the methods:

      (a) Conduct permutations with exchangability restrictions to account for the pattern of dyad-relationships as e.g. implemented in PALM.

      (b) Control for age and sex covariates as covariates (e.g. as in SOLAR), rather than by matching.

      (c) Perform r-to-z transforms when conducting further analyses on correlations that assume normality.

      (d) For all analyses that assume normal distributions, e.g. in SOLAR and BrainSMASH, check that this is the case.

      We have explained how PALM is not suited for the study of effects that are defined at the dyad level (A), that we controlled for age and sex covariates in all our formal heritability analyses in our original submission (B), that we always performed r-to-z transforms when indicated in our original submission (C), and that our spatial permutation results don’t hinge on distributional differences (D).

      (2) Replace SEs derived from kacknife approach with those from SOLAR, or provide a comparison and motivation and/or demonstrate that SEs are correct.

      A more thorough explanation of the block jackknife procedure can be found in prior work introducing the multidimensional heritability method used here (Anderson et al., 2021).

      (3) Given problem (F & G):

      (a) Consider studying the parameters that drive the hyperalignment. They can be included as covariates in heritability analyses, and/or their heritability is of interest to understand the reasons for the heritability reduction post-hyperaligment.

      We agree that this would be interesting but the specific parameters that drive hyperalignment are beyond the scope of this study.

      (b) Include the alternative explanation of hyperalignment-induced noise in the discussion.

      We have added a figure showing that hyperalignment does not increase noise in ISC and explained here why “hyperalignment-induced noise” does not constitute a reasonable alternative explanation for our results.

      (4) Add heritability results for NT phenotypes.

      We have added heritability analyses for NT topography and (global) NT magnitude, as detailed above.

      (5) Motivate global signal removal, and acknowledge this process typically alters results substantially.

      We have added an explanation of our rationale for using GSR and shown in this response that it does not in fact substantially alter the results.

      (6) Rephrase and/or clarify the following:

      (a) "permutations quantify average differences" (under A).

      (b) "network combinations" and related analyses (under B & C).

      (c) why some analyses are separated per visit/day and others not (C).

      (d) methods and reasons for sample size estimation (C).

      We have rephrased or clarified all of the above.

      Reviewer #3 (Recommendations for the authors):

      (1) Participants should be recleared. I know HCP 7T data has 184 subjects. How can the authors have 176 twins and 690 unrelated subjects?

      As we reported in our Methods section, 178 subjects had complete movie-watching datasets, and 176 subjects had complete movie-watching and resting-state datasets. Of the 178 subjects with complete movie-watching data, we identified 690 age- and sex-matched dyads.

      (2) Figure 1. I don't find Figure S1A in Figure S1.

      We thank R3 for catching this error- we have amended this reference to read Fig. S1.

      (3) I could also suggest putting Figure 1 and Figure 2 together.

      We thank R3 for this suggestion- ultimately, we prefer to keep these figures separate to reinforce the difference between our dyadic similarity and formal heritability analyses.

    1. eLife Assessment

      This study presents important findings by identifying small molecules that can stabilize and refold missense-mutated VHL tumor suppressor protein, offering a potential therapeutic approach for clear cell renal cell carcinoma. The computational design approach is well-executed, but the evidence is incomplete due to insufficient demonstration that HIF2 downregulation occurs through on-target VHL rescue rather than off-target effects. Additional experiments with appropriate controls are needed to establish the specificity of the mechanism.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed some of comments raised in the previous round of review and have opted to proceed to a Version of Record without additional review.]

      Summary:

      This is an excellent and strong paper. The authors not only show the mechanisms of action of destabilizing mutations in VHL, but notably, they also go on to computationally design and experimentally test an inhibitor that restores wild-type pVHL function, offering starting points for a new class of kidney cancer drugs. The approach that the authors take here can be used to target destabilizing mutations in repressor proteins, common in diseases, including cancer.

      Strengths:

      This paper is the culmination of an extraordinary amount of work, over years, including method development and testing by a broad range of tools and experiments. It is thorough and comprehensive. It is also well-written and easy to follow.

    3. Reviewer #2 (Public review):

      Summary:

      Inactivating VHL mutations are common in clear cell renal cell carcinoma, and about half of those mutations unfold/destabilize the protein rather than directly interfering with critical protein-protein interactions. The authors identify a compound that can stabilize/refold mutant VHL and seemingly restore its ability to downregulate its major downstream targets.

      Strengths:

      The authors use a clever combination of virtual and cell-based screens, followed by suitable biophysical and cell-based validation assays, to arrive at a VHL refolder. This compound is suboptimal from an ADME point of view, but could be a starting point for further medicinal chemistry optimization. Success would have implications for other diseases linked to similar loss-of-function mutations.

      Weaknesses:

      In going from CP4 to CP4.29 the authors screened based on downregulation of HIF. This is logical but also introduces the danger of identifying chemicals that can downregulate HIF in an "off-target" manner i.e. non-specifically. It therefore essential to clearly show that CP4.29 downregulates steady-state levels of HIF and HIF target genes in cells with suitable (hydrophobic core) VHL mutants but not in isogenic cells lacking VHL.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We are most grateful to both reviewers for providing valuable feedback on our manuscript.

      Reviewer 1 had solely favorable comments, with no suggestions for revision.

      Reviewer 2 pointed out that experiment evaluating the effect of CP4 on pVHL half-life (originally included as Figure 3c) was difficult to evaluate because of CP4’s effect on pVHL abundance prior to cycloheximide treatment. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      Reviewer 2 also pointed out that experiment evaluating the effect of CP4.29 on HIF-2α half-life (originally included as Figure 4g) was not very compelling. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      We agree with Reviewer 2’s suggestion that additional experiments could further solidify that C4.29 downregulates HIF2 in a purely “on-target” manner, however we prefer to reserve such studies for the future.

      Reviewer 2 also made several valuable suggestions for the text itself (awkward wordings / citations / clearer figure legends). We appreciate this feedback and have updated the text accordingly.

    1. eLife Assessment

      This important study advances our understanding of the biomechanics of seed processing in birds by providing a comprehensive 3D kinematic analysis of coordinated bill and tongue movements across two species with contrasting biting forces. The evidence is convincing, combining high-speed XROMM with Bayesian statistical modeling in a rigorous and technically innovative framework that advances the understanding of avian feeding kinematics. Strengthening the statistical validation of qualitative claims, particularly for tongue-seed velocity relationships, and improving the accessibility of the probabilistic modeling framework would further solidify the conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      The authors quantified and compared the 3D kinematics of bill and tongue movements between two seed-eating bird species: one that specializes on soft seeds, and one that is more adapted to feeding on hard seeds. Their goal was to determine specifically what the role of the tongue was for processing (e.g., dehusking) seeds, and to understand how differences in biting strength between species affect other aspects of seed processing. The authors provided intricate (visual) details of seed processing movements, and showed how coordination between the tongue and cranial kinesis (i.e., mobility of the upper bill relative to the cranium) is both critically important for properly positioning seeds to enhance feeding efficiency. Many studies have detailed how seed-eating birds process seeds, but this study has elevated those to a new level of quantification and visualization for readers to fully experience firsthand. Furthermore, the authors established that the force-velocity trade-off that has been observed between bill functions (e.g., feeding and singing) is largely driven by the contractile properties of the muscles. The conclusions are well supported by the results, and the authors placed the results more broadly into the context of manual grasping, making the argument that these birds achieve high levels of dexterity with far fewer degrees of freedom, which could have potential biomimetic applications.

      Strengths:

      This study builds upon - and advances - our understanding of the feeding mechanics of seed-eating birds using cutting-edge 3-dimensional modeling and kinematics. Their quantitative analyses of upper and lower bill, tongue, and seed displacements are complemented by elegant visualizations of seed processing in each species. Their comprehensive Bayesian modeling statistical framework tackles the issue of small sample sizes (i.e., few subjects) with volumes of data for each (i.e., lots of sequential kinematic variables) that plague comparative biomechanics studies, principally because (a) it is difficult to gather these high resolution XROMM and muscle contractile data on more than just a few subjects, and (b) these data streams are inherently very large, as they are gathered at high frame and sampling rates. Furthermore, I believe their approach to statistically testing for differences between species sets a new standard for our field that could (perhaps should?) be implemented in other similar types of studies. Another strength is in how the results were packaged: each subsection indicated how the objectives were addressed, and there were concluding statements trailing each subsection that helped deliver the key takeaways.

      Weaknesses:

      A potential weakness is one that the authors themselves mentioned, regarding the body (and skull) size differences between species. Because gape size limits bite force, and given the force-velocity tradeoff in muscle function, there could be limitations on the rapid manipulation of relatively large seeds for similar reasons in the smaller finches. I see that the small finches appear to overcompensate in their beak rotations, but it's not clear how those compensatory movements might affect their seed processing kinematics with their preferred seed sizes. This does not nullify the authors' conclusions, but the results for the smaller finches might not be entirely representative of seed processing mechanics in smaller species.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates coordinated beak-tongue movements in seed manipulation, biting, and dehusking in songbirds. A comparative analysis of the seed-eating process in two songbird species with different biting forces, the domestic canary and Java sparrow, was conducted using high-speed XROMM with anatomical marker tracking and quantitative behavioral analysis. The authors have done a great job analyzing upper and lower beak rotation and translation, seed orientation and movement speed, and tongue kinematics.

      Strengths:

      The methodological approach of using high-speed (500 fps) X-ray reconstruction for 3D kinematic tracking in small animals is novel and powerful. It enables high temporal resolution tracking of orofacial movements and could potentially inspire future orofacial research in mammals, including mice and marmosets. Moreover, this study encompasses a wide range of anatomical components involved in seed manipulation behavior, including the upper and lower beak, the tongue, and jaw muscles. The behavioral quantification of these components is solid. The findings that both the upper and lower beaks contribute to seed processing, that the lower beak exhibits greater up-and-down and left-to-right flexibility than the upper beak during seed processing, and that the tongue plays an important role in transporting seeds into the mouth are all solid conclusions consistent with observations of bird feeding behavior. Nevertheless, it is valuable to confirm and quantitatively characterize these observations experimentally. The videos are excellent and very informative.

      Weaknesses:

      (1) The paper often resorts to qualitative descriptions (e.g., "a high positive correlation of tongue velocity and seed velocity", "Compared to positioning, the measured velocities of both seed and tongue were much lower") instead of providing exact quantitative measurements or statistical results. The authors stated that temporal autocorrelation biases standard statistical analyses (lines 205-210), but this rationale does not justify the absence of statistical validation. Suggestion: use appropriate methods for time-series data, such as a permutation test, to test the significance of correlations between variables and avoid false positives.

      (2) (Minor) The marker-tracking image shown in Figure 1B could benefit from the inclusion of a higher-contrast, zoomed-in frame of the head showing the metal markers without the red tracking points, alongside the same frame with the red tracking points overlaid, to provide readers with a clearer view of the X-ray image and the methodology and its precision.

      (3) (Minor: possibly soften the mechanistic claim). The proposed mechanism of lingual papillae on the tongue surface may aid food manipulation and food movement towards the posterior region of the mouth is interesting, yet the evidence describing their morphology is not strong enough to support the claim about their functional roles. Furthermore, the claim that papillae orientation affects food transport in lines 294-296 lacks supporting experimental evidence. In addition, the roles of extrinsic and intrinsic tongue muscles in controlling dexterous tongue shape changes and movements are not discussed.

    4. Author response:

      We would like to express our gratitude for the thorough evaluation of our manuscript by the editors and reviewers. We are grateful for the overall positive assessment. The suggestions for improvement are reasonable, and we are certain that addressing these points will improve the clarity, accessibility, and scientific integrity of the study. Thus, we plan to conduct a revision of the manuscript, addressing all the points raised. The most important planned adjustments are outlined below.

      (1) Improving the accessibility of the probabilistic modeling framework

      Reviewer 1 kindly stated that our Bayesian modeling framework for testing for species differences 'sets a new standard for our field.' As a new standard, however, the method should be explained in a more accessible way. Hence, we plan to provide additional explanations for the statistical workflow, e.g., by providing comprehensible visuals, to make the workflow easier to understand and easier to apply.

      (2) Statistical validation of qualitative claims

      We acknowledge that a statistical validation of qualitative claims regarding the relationship between seed and tongue movements and between upper and lower beak movements would considerably strengthen the validity of our findings. We thank Reviewer 2 for bringing permutation tests to our attention for quantifying the correlation between time series. Since permutation tests involving index-shuffling of one of the data sets are generally not valid for time-series data [1, 2], we'll consider a variant of a trial-swapping permutation test, such as a permute-match test [3]. Alternatively, the truncated time shift (TTS) test [2] might be an option, as also this method is valid for auto-correlated time series data. At this point, we can't tell yet which method we'll use for the revised manuscript. We need more time to assess the requirements of each method and evaluate which test is most appropriate to answer our specific research questions and best fits our kind of data.

      (3) Adjustments in the discussion

      Following the suggestion by Reviewer 1, we'll refine our discussion on the effects of skull size differences, putting more emphasis on the implications of potential effects for feeding kinematics in small species.

      Furthermore, as suggested by Reviewer 2, we'll soften our discussion on potential functions of lingual papillae in seed processing, as the current literature lacks experimental evidence for the claimed mechanistic roles.

      References

      (1) Yuan, A. E., & Shou, W. (2022). Data-driven causal analysis of observational biological time series. Elife, 11, e72518.

      (2) Yuan, A. E., & Shou, W. (2024). A rigorous and versatile statistical test for correlations between stationary time series. PLoS biology, 22(8), e3002758.

      (3) Yuan, A. E., & Shou, W. (2025). Permute-match tests: Detecting significant correlations between time series despite nonstationarity and limited replicates. eLife, 14.

    1. eLife Assessment

      This valuable study investigates the neural basis for recovery of complex wheel running behaviour following a unilateral spinal cord injury in mice. By combining behavioural analyses, whole-brain mapping, and tracing techniques, the authors provide incomplete evidence that new cortico-medullary connections can drive effective motor recovery. The paper could be strengthened with manipulations to establish causality, a more fine-grained analysis of the behaviour, and some reorganisation of how the data are presented and discussed.

    2. Reviewer #1 (Public review):

      Summary:

      The authors seek to understand and identify the neural plasticity that underlies recovery from precise unilateral hemi-pyramidotomy. The corticospinal tract is severed on one side in the pyramids below the exit of corticoreticular projections. Recovery from the injury is achieved with an intensive wheel running rehabilitation regime. The anatomical sites of plasticity, the importance of plasticity in different reticular areas<br /> to recovery, and the impact of the degree of plasticity observed on recovery as correlated predictors, are shown.

      Strengths:

      Refined anatomical analysis using mouse line and genetic and viral intersectional tracing identifies specific reticular targets of likely enhanced cortical control that correlate with recovery of locomotor skill.

      Weaknesses:

      (1) The study is correlational at this time. This does not undercut the value of the data and the identification of targets of plasticity achieved in the work.

      (2) Generalization of motor gains beyond locomotion was not tested. Reach-to-grasp tasks for feeding were not tested.

      (3) Some discussions and use of the terms fine motor and skilled motor are fuzzy, and the limitations of the study are not sufficiently clearly stated.

    3. Reviewer #2 (Public review):

      Summary:

      Bonanno and colleagues combine unilateral pyramidotomy, continuous voluntary complex-wheel running, whole-brain intersectional CSN tracing, and c-Fos mapping to ask whether rehabilitation reorganizes the supraspinal collaterals of the intact corticospinal tract neurons. The study is technically ambitious and competent, the uPyX + complex-wheel + intersectional-tracing + BrainJ combination is smart and interesting, the behavioral effect is convincing, and the blinding and exclusion criteria are explicit. The central anatomical finding - a CSN-specific, whole-brain projectome comparison with subregional LPGi/GiA/MdV granularity - is a legitimate contribution that builds on Asboth 2018. However, the strength of evidence does not support the strongest causal wording in the current abstract, significance statement, and parts of the discussion: the results remain correlational, the MdV-behavior correlation is modest, and its significance is sensitive to the unit of analysis. A major revision is recommended, primarily of framing and quantitative robustness, rather than because the central dataset is unconvincing.

      Strengths:

      (1) Technically ambitious and technically competent study addressing a relevant gap: brain-wide mapping of intact-CSN reorganization under continuous voluntary rehabilitation.

      (2) The combination of uPyX, complex-wheel running, intersectional tracing, and BrainJ whole-brain projection analysis is novel and well integrated.

      (3) Behavioral effect is convincing, blinding, and exclusion criteria are explicit.

      (4) The central anatomical finding (CSN-specific whole-brain projectome under rehab, with LPGi/GiA/MdV subregional resolution) is a legitimate contribution that builds on Asboth 2018. The closest recent works (Lemieux et al. 2024, Jeleva et al. 2026) study reticulospinal rather than CSN plasticity and are complementary rather than competing.

      Weaknesses:

      (1) Causal framing extends beyond what the current evidence supports.

      The abstract and significance statement present MdV as a potential mediator, or even a central locus, through which rehabilitation re-establishes descending control of the impaired limb. This is stronger than the evidence. What the paper shows is that CSN collateral projection density in MdV has a mild-to-medium correlation with behavioral recovery, and that this region is already known from prior work (Esposito 2014) to be relevant for skilled forelimb function. That is an interesting anatomical correlation, not a demonstration of mediation. No manipulation of MdV or of MdV-projecting CST terminals is performed; there is no silencing, no pathway-specific perturbation during rehabilitation, and no test showing that the identified sprouting is necessary for recovery. The limitations section acknowledges this, but the prominent claims do not.

      (2) The behavioral caveat on what is actually novel.

      The cleanest way to state what is genuinely new, clearer than the abstract itself, is this: when a CSN population loses part of its spinal target domain (via contralateral uPyX denervating the opposite cord), some CSNs from the opposite cortex appear to redirect growth into brainstem collaterals (LPGi, GiA, MdV). The compensation is plausibly sufficient to restore gross descending drive to the impaired forelimb, but most probably inadequate for the fractionated, cortico-motoneuronal fine-grain control that the direct CST normally provides. That distinction - recovery of drive and even skilled locomotor control vs. recovery of fine precision - is consistent with the ladder-rung improvements the paper reports (footfall counts are an integrated gross-placement metric) and with the skilled-reaching literature (Esposito 2014 and similar), which suggests precision grip and digit individuation would not be fully recovered by an MdV-centered detour. This note is also translationally important when we ask what humans consider fine motor control, which is mostly object manipulation. Relatedly, the ladder task is "skilled" in the operational sense that it requires cortical control, but the motor output measured (gross paw placement, overreach) is not fine motor function in the sense of digit individuation, grip force modulation, or pellet manipulation. "Skilled" here does not even mean *acquired* skill: classical skilled reaching in rodents involves explicit training to acquire a novel motor program, whereas here mice are only habituated. The brainstem-compensation hypothesis is more comfortable with restoring cortex-dependent gross placement than with restoring acquired fine-motor skills.

      (3) The anatomy sample is modest for the precision of the claims.

      Projection analysis rests on n = 9 pooled controls, n = 5 uPyX−Rehab, and n = 5 uPyX+Rehab. For a whole-brain subregion analysis, this is not a large dataset, even with the sensible restriction to the Wang et al. spinally-projecting set. The three medullary hits are plausible, but some of the most specific conclusions rely on a relatively small number of animals for its most specific claims. This matters especially for the MdV-behavior correlation.

      (4) Normalization enforces a zero-sum structure.

      Projection density is normalized to the total CST tract signal. This is a reasonable way to control for tracing variability, but it imposes a relative structure on the data: an apparent increase in one region may partly force an apparent decrease elsewhere. This may matter and has to be looked into by the authors, because the manuscript interprets decreased density in some other targets as meaningful redistribution.

      (5) The decision to merge PMn and MdV under a single "MdV" label needs more justification.

      Since the discussion relies on prior literature assigning skilled forelimb function to MdV proper, the reader needs to know whether the signal truly localizes there or whether it may partly reflect a neighboring region grouped under the same atlas label. Related to this, laterality would be very informative: since the proposed compensatory route is anatomically directional, showing whether the increased signal is preferentially located on the expected side of the medulla would strengthen the interpretation.

      (6) The c-Fos / Fig. 3 section goes beyond what the data directly support.

      The section "Complex-wheel running recruits intact corticospinal neurons" and the figure title "Rehabilitation functionally recruits intact CSNs" go beyond the actual observation, which is that a higher fraction of CSNs in M1 and M2 are c-Fos+ in runners than in non-runners. "Functionally" is not supported: c-Fos is a transcriptional marker of recent activity, not a functional readout; it does not show that the CSN's output is used to drive behavior. "Rehabilitation" is not supported either: the contrast is runners vs non-runners, applied uniformly across Sham and uPyX groups - healthy Sham+Rehab animals are on wheels for leisure, and the c-Fos effect is present in them too. The finding is difficult to interpret without thinking of the simpler framing ("moving mice have more motor cortex activity than resting mice"), with no control for generic arousal or ambulation. This section is the softest link in the causal chain running - CSN activity - medullary sprouting - recovery.

      (7) MdV-recovery correlation: unstated multiple-comparison correction and possible pseudoreplication.

      The correlation (R² ≈ 0.33, p ≈ 0.01) is the backbone of the paper's "causal" claim. Panels L/M/N test three correlations (LPGi, GiA, MdV vs forelimb footfall recovery); only MdV is reported as significant. The Figure 5 legend applies Tukey adjustment to the t-tests in A-C but makes no analogous statement for the correlations in L-N. A 3-test Bonferroni (α = 0.017) would not flip the MdV result, but disclosure is warranted, and the three tested regions were pre-selected from the significant group contrasts in A-C, which, to a statistician, would further shrink effective α. More importantly, the figure legend states that closed and open circles represent CFA- and RFA-traced values, respectively, which suggests the correlation treats the two tracer channels per mouse as independent datapoints - doubling the apparent n (≈ 20 from 10 uPyX mice), with the result of a higher significance than one would have at the mouse level.

      (8) Reporting issues.

      The reader would benefit from judging statistical choices such as those above directly from a data table rather than interpreting the authors' choices. The SciScore rightfully flags multiple missing components of transparent reporting: missing RRIDs, no code availability, limited data availability, and no power calculation, among others.

      Almost all these weaknesses can be addressed with a revision of the manuscript, especially in the framing of results.

      Conclusion:

      The core message - that rehabilitation is associated with a selective pattern of CSN collateral remodeling in the motor medulla, and that MdV projection density covaries with behavioral recovery - is defensible from the data and already a useful result. The current wording in parts of the abstract, significance statement, and discussion goes beyond this and implies a mechanistic conclusion (mediation, central locus, re-establishment of descending control) that the data do not yet establish. The manuscript would better match its evidence with "associated with", "correlates with", or "candidate locus" framing, unless a causal experiment is added.

    4. Reviewer #3 (Public review):

      Summary:

      In this study, Bonanno et al. show that after a lesion of the corticospinal tract (CST), rehabilitation running in a complex wheel drives improvement in skilled forelimb performance in mice. Mice with unilateral CST injury can perform gross motor tasks (locomotion) at the same level as the non-injured mice, but injured mice still have deficits in another task involving fine motor control. Thus, it is well-suited to test the efficacy of locomotion-based rehabilitation in fine motor control. Mice that voluntarily engaged in the rehabilitation protocol improved in the fine motor control task more than those mice that did not perform any rehabilitation. Highlighting the role of rehabilitation in the recovery of motor function after the lesion.

      The authors aimed to study rehabilitation-driven intact CST sprouting to supraspinal areas. They identified one area in the motor medulla where rehabilitation significantly changes the projection density from the intact cortical spinal neurons. Interestingly, this area has ipsilateral connections and thus could be a pathway to convey motor commands from the intact corticospinal tract to the denervated area. However, as the authors acknowledge in the discussion, they only found a correlation between the change in the synaptic projections from intact CST to the medulla and the recovery. Future work should study if indeed the area of the motor medulla identified here increases its ipsilateral projections to the denervated area, confirming the re-routing of motor commands from the intact cortico spinal tract to the denervated area. The paper is strong and, in general, claims are supported by the data.

      Strengths:

      In this study, Bonanno et al. show that after a unilateral corticospinal tract lesion (CST), locomotion rehabilitation can improve motor function and improvements generalized to tasks that require fine motor control. Moreover, it identifies a potential pathway that could be used for the intact corticospinal tract to convey motor commands to the denervated area. The pathway identified here could become a target for rehabilitation therapies.

      Weaknesses:

      As the authors acknowledge in the discussion of the study, the main limitation of this study is that the reorganization observed at the motor medulla is only correlational. Thus, it is possible that the adaptation to running with an injured limb of the intact CST to adapt to an injured limb rather than a re-routing of the intact CST inputs to the denervated area underlies the synaptic changes observed in the motor medulla.

      The statistical analysis could be better described.

      The generalization of skilled movement is limited to only locomotion tasks.

    1. eLife Assessment

      The worldwide decline in the health of coral reefs is well documented, and overgrowth by microbial consortia can be a contributing factor. Kelman and colleagues used metagenomic analysis to interrogate potential changes in phage-associated genes predicted to be involved in central carbon metabolism. The study addresses the hypothesis that metabolic genes associated with carbon metabolism that are encoded by viruses reflect the health of the corals. The study contributes a valuable perspective on the potential role of phages in coral health, although limitations of the data and analyses offer an exploratory examination rather than a definitive result. Overall, the evidence supporting the major findings is incomplete, in part because the conceptual model relies on qualitative assumptions rather than empirical data.

    2. Reviewer #1 (Public review):

      Summary:

      Microbialization (bacterial overgrowth) is a recognized component of degraded, eutrophied coral reefs where there is a shift from coral to algal dominance on the benthos. In addition, previous work has demonstrated that virus communities shift from a lytic strategy dominated (kill-the-winner) to a temperate (lysogenic) strategy dominated with reef microbialization. Kelman et al. sought to leverage previously published virus metagenomes produced from the water column of healthy and degraded coral reefs to assess virus community metabolic shifts. The authors also produce a conceptual model to demonstrate the potential impact of the observed metabolism shifts on reef fates.

      Strengths:

      The main strength of the manuscript is the findings from their metagenomic analyses and results. The virus metagenomes were produced using established approaches in the field and yield sufficient data per sample for their analyses. Interesting results regarding the shift in the types of genes from anaplerotic to cataplerotic provide the foundation for testable hypotheses to determine the magnitude of impact virus strategies have on reef health. The introduction is also well written and sets up the scene very well.

      Weaknesses:

      (1) The methods text currently omits important information related to the sampling design. It is not clear how many metagenomes are from healthy and degraded communities. This impacts the interpretability and robustness of the statistical results. Furthermore, it is unclear if analyses are based on assembled contigs or read-based alignments. Improving the clarity and organization of the Methods is essential for reproducibility.

      (2) Regarding the bioinformatics approach, normalization using the "percent known" approach within samples may not fully account for discovery bias related to sequencing depth. While Supplementary Table 1 shows variability in read counts, the lack of community-level metadata makes it difficult to determine if sequencing depth covaries with community type (healthy vs. degraded). The study would benefit from a rarefaction analysis or subsampling to ensure that gene frequency trends and Spearman correlations are biological signals rather than artifacts of sequencing effort.

      (3) The qualitative model in Figure 5 is positioned as evidence for the role of viruses in reef health, but it does not provide independent support for the authors' hypotheses. Since the model is parameterized using "arbitrary units" to reflect the authors' assumptions rather than being derived from the empirical metagenomic data, it serves as a helpful illustration of a hypothesis but not as a validation of the findings.

      (4) Results and discussion require revisions to improve readability and connectivity across sections. Ensuring a clear distinction between empirical data and model-based speculation would help the audience better appreciate the science.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Kelman and coauthors investigates how viral communities differ in the genes they encode in healthy and degraded coral reef ecosystems. Across 19 viral metagenomes from Central Pacific reefs, the authors assess the frequency of integration/excision genes as a proxy for viral community temperateness and ask whether genes associated with central carbon metabolism covary with signatures of temperateness. The main finding is that viral communities with more temperate-related genes encode more genes from the Entner-Doudoroff pathway and other reactions interpreted as anaplerotic, whereas more lytic-associated viral communities show greater representation of some pentose phosphate pathway, TCA, and redox-associated genes interpreted as cataplerotic. The authors propose a model based on these patterns in which lytic viral metabolism helps suppress bacterial overgrowth on healthy reefs, while temperate viral metabolism may promote microbialization on degraded reefs. The study addresses an interesting and potentially important concept - that viral auxiliary metabolic genes are important components of microbial communities and can affect ecosystem functioning. Linking viral metabolism to coral reef microbialization is a creative conceptual advance. The manuscript is clearly written, and the reported enrichment of anaplerotic genes in temperate-associated viromes is an interesting pattern that could motivate future work on how viral metabolic potential varies across reef states.

      Strengths:

      (1) The study connects viral lifestyle, central carbon metabolism, bacterial overgrowth, and reef degradation in a framework that could be useful for future studies of coral reef ecosystems and viral ecology. This is an interesting synthesis that links viral auxiliary metabolism to broader questions about microbialization and reef state.

      (2) The manuscript is generally clearly organized around a testable prediction: viral metabolic gene content should vary along a lytic-to-temperate viral community gradient. The reported enrichment of anaplerotic genes in viromes with a larger fraction of temperate viruses is a compelling result.

      (3) The authors highlight several virus-encoded metabolic genes that may not have been previously reported in viral datasets or genomes. If supported by further validation, these observations could expand the known repertoire of viral metabolic potential.

      (4) The modeling helps clarify the feedbacks the authors propose may connect viral lifestyle, bacterial metabolism, and coral reef degradation. It provides a foundation for generating hypotheses about how viral metabolic genes could influence reef microbial dynamics.

      Weaknesses:

      (1) The main limitation is that the evidence for several key claims remains indirect. The core analysis is based on correlations between metabolic gene frequencies and integration/excision-related genes. This does not demonstrate that the metabolic genes occur in temperate viral genomes, are physically linked to lysogeny genes, are expressed during infection, or alter host metabolism. Thus, the data support an association between VLP-associated metabolic annotations and a community-level temperateness proxy, but not a direct link between temperate phages and these metabolic functions.

      (2) It is important not to equate community-level gene frequencies with genome-level or infection-level metabolic programs. A virome may contain more anaplerotic genes overall, but that does not demonstrate that individual viruses reprogram their hosts in an anaplerotic manner nor that infection produces a net anaplerotic effect. Individual viruses may encode both anaplerotic and cataplerotic genes, and a smaller number of cataplerotic genes could have stronger metabolic consequences depending on expression, enzyme efficiency, pathway position, and host context. This is an important limitation that should be acknowledged and, if possible, addressed with contig- or genome-level analyses.

      (3) The ecological interpretation assumes that viral infection is strong enough to influence reef-scale bacterial population dynamics. However, the study does not directly measure infection frequency, lysis rates, viral production, burst size, lysogeny frequency, prophage induction, gene expression, or bacterial mortality. If viral mortality or lysogenic conversion were rare in these systems, the observed gene-frequency patterns could have limited ecosystem-level consequences. This makes claims about viral metabolism suppressing bacterial overgrowth, accelerating microbialization, or acting as a conservation lever more speculative than suggested.

      (4) There are statistical limitations related to the use of relative gene frequencies. Because genes are normalized as percentages of known genes, the data are compositional. Apparent increases in some categories may partly reflect decreases in others. Bootstrapped Spearman correlations are useful for assessing the robustness of these associations, but they do not address compositionality or multiple testing.

      (5) The anaplerotic/cataplerotic classification is central to the manuscript's conclusions and would benefit from more support. The framework is useful, but it depends on both annotation confidence and biochemical context. Sequence-similarity annotations alone may be vulnerable to misannotation, especially for central metabolic enzymes that share conserved domains across functionally distinct proteins. Stronger evidence that key genes contain key functional domains and/or are phylogenetically related to characterized enzymes would help support the proposed functions. In addition, many central carbon enzymes are reversible or context-dependent, so a clearer rationale for each classification would strengthen the interpretation.

      Overall, the manuscript presents a valuable hypothesis and highlights new ecological patterns in coral reef viral metagenomes, but falls short of the evidence needed for the strongest claims. The work would be strengthened by analyses that directly link metabolic genes to viral genomes or lysogeny markers, address compositional effects, validate key annotations, and more clearly distinguish observed gene-frequency associations from hypothesized effects on infection, host metabolism, and reef state.

    1. eLife Assessment

      This Review Article puts forth a normative theory for the grid cell representations found in the entorhinal cortex. It discusses a range of theoretical models and experimental findings, organizing them around a proposed framework in which grid cells are interpreted as biologically constrained, high-fidelity codes for path integration. This framing can be potentially interesting both for readers seeking a conceptual entry point into the grid cell literature and for those more generally interested in the promises and limitations of normative theories in neuroscience. Some logical gaps and points requiring conceptual or technical clarification were nonetheless identified. Moreover, the empirical support for the path-integration account is not yet as definitive as the manuscript's framing sometimes suggests. The review would thus be strengthened by clearer justification of key arguments and fuller discussion of biological complexities, model limitations, and competing interpretations. Some stylistic choices in how arguments and literature are sometimes rhetorically framed may lessen the review's appeal for key segments of its intended audience.

    2. Reviewer #1 (Public review):

      Summary:

      The review by Dorrell and Whittington synthesizes the progress made over the past few years with respect to a normative theory of grid cells. The core question addressed by normative frameworks of grid cells is what primary computational function grid cells serve. The review discusses evidence from mechanistic models and experimental data that point to path integration as the computational function of grid cells, consistent with results from normative models. The main goal of the review is to clarify the normative grid cell theory literature. However, the current version of the article reads at times more like a perspective or opinion article in support of the path integration hypothesis rather than a critical review of normative frameworks in the grid cell literature that contrasts the benefits and limitations, as well as pitfalls and caveats, with other modelling approaches.

      Some specific comments are as follows:

      (1) Abstract: "The first question quickly attracted an answer: grid cells subserve path integration ..." - I am not sure if this statement is correct. The first grid cell paper by Hafting and Fyhn in 2005 suggested that grid cells are part of a path integration-based map, and the paper emphasizes the map part. It remained unclear, and is still debated, whether grid cells are part of a system performing path integration or whether grid maps reflect the output/result of a path integration process. Other theories about the function of grid cells were brought forward as well. Although the main competing theory is discussed in this review, this review article at times appears more as a perspective or opinion article with a clear bias toward the path integration hypothesis rather than objectively discussing the evidence.

      (2) Grid cells may serve multiple functions. What would be the implications for our understanding of grid cells and for interpreting the results of normative models? In general, the review could discuss some pitfalls or caveats of normative models in more detail.

      (3) A normative framework can be helpful in two ways: (a) Given sufficient details on biological constraints, a normative model can help identify the computational function of grid cells. If a computational function is given and - under the given simulated biological constraints - grid cells were part of the solution, the results of the model would support the hypothesis that grid cells serve the computational function in question. (b) If a computational function were identified beyond any doubt (e.g., assume experimental data demonstrated that grid cells are necessary and sufficient for path integration), a normative model would help identify biological parameters necessary to produce grid cell firing. Unfortunately, the review falls short in making this clear distinction between (a) and (b) and in discussing important caveats regarding mixing up these two ways. E.g., the neural network model approaches by Sorscher et al. and others have been criticized because they try to achieve two things at the same time: find support for the computational function of grid cells and identify optimal parameters that result in grid cells. But doing both at the same time provides a strong bias in tweaking the parameters in exactly the way you need for the model to produce grid cells as a solution (other solutions may be possible given other parameters), preventing strong conclusions regarding the computational function of grid cells and preventing conclusions about what the parameter choices mean for biological connectivity motifs. These caveats in setting up normative models and interpreting them could be discussed in greater detail.

      (4) A common assumption underlying most grid cell models is that head direction is viewed as identical to movement direction. However, head direction can differ at times from movement direction, and entorhinal head direction cells code head direction rather than movement direction (Raudies et al., 2015; 10.1016/j.brainres.2014.10.053). This missing link in how movement direction signals reach and inform grid cells could be discussed.

      (5) "Knowing that one neuron in a module is active and that you make a movement north uniquely determines which neuron in that module should be active next" - I agree that this rule follows from the fact that grid cells within one module differ in phase but share spacing and orientation. However, I am surprised that the authors do not also make the argument here for the value of a normative model. Rebecca R.G. et al. (10.7554/eLife.96627) use exactly the rule cited above as a normative function. They demonstrate that this rule begets grid cells. Isn't this a prime example of how a normative approach can contribute to scientific inquiry? First, a hypothesis about a computational function is derived from experimental data. And in turn, using a normative framework, the experimental data are derived from the computational function (under appropriate biological results). The paper is discussed later together with Nicolai Waniek's work (10.1162/neco_a_01255). However, in my opinion, their work seems to be somewhat misrepresented in that later paragraph. E.g., velocity is still required as an input to determine which neuron should be active next, neurons do not need to be binary units, and space is not discretized beyond the fact that space is encoded by neurons with spatial firing fields.

    3. Reviewer #2 (Public review):

      Summary:

      This review by Dorrell and Whittington covers a number of aspects related to normative modeling of grid cells. They begin by discussing key experimental insights on grid cell phenomenology. Then, they discuss how grid cells can be used to perform path integration and how they size up as efficient codes of space. These two sections then lead the authors to discuss how combining path integration and efficient coding objectives leads to models of axis-aligned grid cells in a single module. Discussion on non-linear objectives leading to multi-modules is presented. The review ends with several outstanding questions and an optimistic outlook of how normative models (particularly, task-optimized RNNs) can be used as tools for advancing understanding in neuroscience.

      Strengths:

      (1) The review is timely and covers an area that has seen a lot of recent activity. This discussion around many of the different results (and kinds of models), I think, will be generally helpful for the field.

      (2) Although I think the story could be a little more coherently made (see below), in general I enjoyed the author's flow from efficient coding -> efficient coding + path integration -> efficient coding + path integration + non-linear objective. This framing supports the specific conclusion the authors arrive at.

      (3) I also really liked the message that the review made of how normative modeling, despite some of its challenges/limitations, can be used effectively in neuroscience. The discussion of cycling between "experimental" modeling (e.g., vanilla RNNs) and theoretically-grounded models was nice, and I think it helps demonstrate the value of this approach.

      (4) Showing how the metric loss could be seen as a bandpass filter (Figure 3C) was nice and a contribution of the review.

      (5) While the focus of P4 (conjunctive HD-grid cells) felt initially a little cast aside, the discussion around "brain and task-optimised RNNs with standard architectural choices use fundamentally different path-integration mechanism" was nice and I think helpful for steering the community to an interesting open problem.

      (6) Identifying how "non-linear functionality" can lead to multi-modules was nice and not something that I have seen as clearly presented before.

      Weaknesses:

      (1) The authors view the experimental evidence for grid cells being linked to path integration as "specific and strong" and that the " key computational feature that defines entorhinal cortex [is] path-integration". I think experimentalists (at least the ones I work with) would push back on that. First, it's hard to isolate path integration in rodent experiments. So while Gil et al. (2018) did about as good a job as you could do, there are still other interpretations of the results that are not purely path integration dependent. And second, as the authors point out later in the review, there is experimental work finding that grid cells are disrupted in large environments and 3D. Path integration certainly happens (to some extent) in these spaces, which begs the question of how it is achieved with weakened grid coding. Thus, I think reducing the claims about how strongly grid cells are experimentally linked to path integration is called for.

      (2) The authors introduce the idea of efficient coding of space and discuss how grid cells are not optimal. It is later clarified (Sec. 5.3) that multi-module codes can be efficient (even if not the most optimal). I was confused reading Section 3, because in Section 2 the multiple modules are discussed, but then in Section 3, they are dropped, and only a single module is being considered. Equation 2 was also a little confusing to me. Alpha is not defined, and I would have thought that it would be x^Tx' - g(x)^T g(x') and not x^Tx' g(x)^T g(x'). Given that there is no page limit here, I think a little more detail in Section 3 would be helpful.

      (3) In Section 3, the authors make use of P2 (translation invariance within a module) to rule out (or, at least, question) certain models/approaches. While this is certainly a standard assumption made in theoretical work, it is not very well supported by experimental findings. In particular, Diehl et al. (2017), Ismakov et al. (2017), and Dunn et al. (2017) all found that individual grid fields systematically vary in their peak firing rate. In addition, Redman et al. (2025) found that, within a given module, there was a small but robust diversity of grid orientations and spacings. These suggest that grid cells within a single module may actually be able to encode properties of local space and give some support to normative models that find efficient space coding with grid cells by finding non-axis-aligned grid fields. I think this is all important to mention because: a) it provides more biological nuance to the question about spatial coding; b) it provides more ways in which to test models. For instance, in Redman et al. (2025), the Sorscher et al. (2022) model was shown to produce variability in grid properties that loosely matched what was found in real data. For tests like this (e.g., how much does a model reproduce variability in grid firing field peak rates), I think it is going to be important for continuing to evaluate models.

      (4) The focus of the review, I know, is grid cells, but of course, grid cells are part of the MEC and the larger hippocampal network. I totally understand, at some level, you have to make a decision of what to model, but it seems that there are other functional classes of neurons (border cells, head direction cells) that all play an important role in path integration. And while the models the authors consider at the end of the review capture properties of grid cells really well, they do so at the cost of not modeling anything else. The authors mention this in the context of the models not capturing conjunctive grid-head direction cells, but I think the point is a deeper one, and more discussion of at what level it makes sense to consider grid cells only is important.

      (5) As I mentioned in the Strengths section, I did enjoy the flow of the paper on how path integration + efficiency is needed to get grid single modules and path integration + efficiency + non-linearity is needed to get multiple grid modules. This creates the story that adding more of these theory-driven constraints helps lead to more "accurate" models of grid cells. But one alternative view is that, if path integration + efficiency is enough to get a single grid module (but only a single grid module), then maybe the utility (or need) of multiple grid modules comes from something else. That is, instead of saying "we need more constraints to get multiple modules", it could be evidence for "we need to re-think whether multiple modules might need a different theory to explain". While I understand this is a big picture question that maybe isn't entirely fair to ask of the authors, I think: 1) the authors do a nice job of positioning their review as a kind of discussion on what normative modeling can provide to neuroscience, so having this discussion on when the failure of a model to capture ALL aspects of the biological features motivates further constraints as opposed to a new approach, would be useful; 2) this question connects with the title of the paper, i.e. "what is the question?"

    4. Reviewer #3 (Public review):

      Summary:

      The authors present an extensive review of the literature on normative grid cell theory, asking what kind of cost function might be minimized by the entorhinal grid cell code. The authors show which of the main features of grid cells emerge from combinations of terms in a cost function that optimizes for spatial fidelity, biological plausibility, and path integration. They conclude by outlining potential future directions for the field.

      Strengths:

      The structure of the review makes it particularly useful for researchers who are familiar with grid cells but not necessarily with normative models. Equations are kept to a minimum and are usually explained conceptually.

      Weaknesses:

      I identified one main weakness, related to the fact that the introduction to experimental results around grid cells and what they allow us to conclude is less nuanced than the rest of the review. However, since this is not the main focus of the manuscript, I consider this a secondary limitation.

      The review organizes the current literature on the subject within a coherent conceptual framework, helping to define possible paths forward for the field.

    5. Author response:

      We thank the reviewers for their time and attention which will significantly improve the paper. Further, we are grateful for their appreciation of our goals and work. In sum, the reviewers point to our overstated discussion of experimental evidence which we will tone down, some slightly confusing points of argumentation which we will clarify, and some discussion points on the role of normative theories that we will add text to address. We believe this will improve the paper significantly and hope you agree!

      Major Concern: Experimental Support for Path-Integration is not as strong as suggested

      The major point raised by all reviewers (reviewer 1 comment 1, reviewer 2 comment 1, reviewer 3’s only weakness) was that our presentation of the experimental perturbation evidence for path-integration is stronger than the reality. On reflection, we agree with this evaluation. We thank the reviewers for raising it; we will moderate our writing and include the sensible caveats raised. In sum, we still think that the convergence of evidence points to path-integration: first, disruptions to grid cells lead to path-integration problems, though these perturbations admittedly aren’t perfectly precise; second, normative theories of path-integration lead to grid cells and predict grid cell behaviour; third, mechanistic models of path-integration match grid cell behaviour and predict connectivity subsequently measured in entorhinal cortex. However, the evidence is not as all-encompassing as we suggested.

      That said, we’d like to further comment on one point. It is argued (reviewer 1, comment 1) that there are other theories of grid cell function, and that we discuss these theories. We discuss efficient-coding only models of grid cells and emphasise strongly why we reject them. We also briefly discuss oscillatory-interference models of path-integration and our reasons for not pursuing them further. As such, the reviewer is correct that our reading of literature strongly points us towards path-integration rather than other theories. We will slightly change the framing of the paper to make it clear that we are making a case. However, we are not aware of other theories the reviewer might be referring to. If the reviewer can point us to the other suggested theories that we do not address we would be happy to evaluate and include them.

      We now turn to the remaining comments, and how we plan to address them.

      Reviewer 1, Comment 2 – There could be multiple roles for grid cells

      The reviewer is indeed right that grid cells might perform multiple functions. This could just mean that the same computational motif (e.g. path-integration) is reused across different computations though that introduces no changes to the required normative theory. A stronger claim would be that grid cells perform both path-integration and some other function. This, according to a normative perspective, would most likely change how grid cells were optimally structured. We use the fact that large parts of the grid cell code can be captured with only path-integration as an argument against additional roles for grid cells. That said, there exist properties of grid cells not well-captured by path-integration which could well be smoking guns for additional roles of grid cells. The review already discusses both discrepancies between grid cells in three and two dimensions, and inhomogeneities in the grid in complex environments, and we will add two more (heading direction and peak-to-peak/angular variability, discussed below) that we are grateful to the reviewers for raising, and we discuss each of these in detail below.

      That said, whether these are necessarily arguments against purely path-integration or a reflection of interesting mappings of the core path-integration mechanism to the measurements we make remains to be seen. We would argue that both 3D grid cells (as explained below: there appear to be 2D slices in which grid cells behave as you’d expect) and spatial inhomogeneities (as explained in the paper: mappings of torus to world can introduce warping) can be explained without reference to additional computational roles of grid cells, which remain to us the most parsimonious explanation. We discuss next the slight update to path-integration only that the heading direction story suggest. But in sum, our view is that these discrepancies are likely not fatal for our path-integration-centric view of grid cells, but may well suggest some very interesting clarifications.

      Reviewer 1, Comment 4 – The system has two heading signals: true & internal, why?

      The reviewer is right to point to the puzzle over true vs. purely internal heading direction and which drives grid cells. We believe recent work from Abraham Vollan has effectively solved this puzzle: there appear to be two parallel circuits, one theta-modulated and following internal heading direction, another theta-unmodulated and aligning more with true heading direction. We will make sure to include discussion of this exciting work in our revised submission. This serves as a good example of an update we concede to the most austere version of the path-integration only view. Rather, it seems there are two parallel path-integrators working with different heading signals. The reasons for this remain unclear, but seem to be related to attention and planning (Vollan et al. 2026).

      Reviewer 2, Comment 3: Real Grid Cells have peak-to-peak variability & Angular variability

      The reviewer is right to point to the discrepancy in peak-to-peak firing rate and angles within a module that we did not adequately address. First, it is Sorscher’s RNN models, not nonnegative PCA that can generate a distribution of grid angles (Redman et al. 2025), which suggests that path-integration and such variability are compatible. We emphasise this point because the non-path-integration results from nonnegative PCA produce grid cells oriented at 30 degree offsets, something not measured even when you’re careful as in Redman et al. 2025. Thus, this becomes an interesting target for future work: perhaps using theories of path-integration up to an error threshold (rather than perfect) such angular diversity would be recovered. We will include this in our discussion. Further, we will include discussion of peak-to-peak variability that, as yet, has no obvious role.

      Reviewer 2, Comment 1: grid cells are inhomogeneous in 3D or complex environments, doesn’t that break the theory?

      Disrupted grid coding in extended or 3D environments indeed deserve more discussion, which we will add. In particular, we will add recent evidence that grid cells in 3D can be understood via the correct sequence of 2D projections(Qi & Yartsev, 2026). These two phenomena seem, to us, consistent with a path-integration only view of grid cells, as discussed above, and we hope to make this position clearer.

      Reviewer 2, Comment 5: Couldn’t there be other reasons for multiple modules?

      We have suggested a consistent normative framework in which multiple modules are explained through their role in non-linear coding. We think this elegant, and the most parsimonious current theory. We could, of course, be wrong. The discrepancies pointed to above might be good clues to follow to work out what else these modules might be doing, but currently these alternative explanations seem not to exist. We will text to clarify this.

      Reviewer 1, Comment 3: The review confuses computational and parameter parts of normative theory

      We disagree with the reviewer’s dichotomisation of normative theory. We view a normative theory as the complete procedure that produces the predictions. Almost all such theories have parameters and hence fitting a theory to data comprises both elements (a) [computational role] and (b) [specific parameters] identified by the reviewer. Occasionally theories have no parameters in the traditional sense, e.g. Rebecca et al.; instead they have heavy assumptions that play an equivalent role. It is true that, as the reviewer says, Sorscher et al.’s work was criticised for producing grid cells only for specific parameter values. We never found this as damning as Schaeffer et al. argued: simply it says that that theory is only correct within the given parameter range. Rather, arbitrating between models, parameters, or assumptions seems the same basic process: see what they predict and keep working with models while they remain useful ways to understand measured phenomena. If a model with very specific parameter values remains useful, that seems okay. In fact, we argued extensively why we think the nonnegative PCA model is not a useful model, but this was for completely different reasons. To us this story just reinforces the importance of hygiene in normative research: perform parameter sweeps and clarify how they constrain the claims you are making, carefully arbitrate what models can capture. Indeed, that is the whole goal of this review. We might be misunderstanding and, if so, we welcome correction.

      Reviewer 2, Comment 4: Normative Models of Cells Beyond Grid Cells

      The reviewer is right that extending these models to other cell types is an interesting area for further work, and that other cell types do seem to be involved in aspects of navigational computations both in RNNs and the brain. We will include a discussion to this effect in the revised manuscript. That said, we think the modularity of grid cells and their tight-linking to path-integration calculations should also be appreciated as a win!

      Reviewer 2, Comment 2: Multi-modularity is not cleanly explained

      We thank the reviewer for the comments, we agree. We will clarify the story regarding multiple modules, and will explain the equation further.

      Reviewer 1, Comment 5: the early introduction of phase-shifted Grid Cells seem the perfect place to normatively argue for Path-integration!

      We agree with the reviewer that this point can be made both normatively (‘oh look! If I try to do this optimally, I get translations!’) or, as we did early in the paper, mechanistically (‘oh look! With these cells I can do this!’). Indeed, a large part of the point of our paper is that path-integration is what is required to normatively derive phase-shifted grid modules, something discussed by Rebecca et al., our earlier work, and RNN studies, and appreciated for two decades. The earlier part of the paper does not discuss these papers as that section is aimed at giving intuition for the solution (mechanism). Later sections then heavily discuss the normative angle. We hope that division of labour makes sense.

      Finally, we will refine our summary of Rebecca et al. The reviewer is right that neurons don’t have to be discrete, we apologise for that error, but our understanding is that the only meaningful role of a neuron in Rebecca et al.’s work is the region in which is active, effectively making every neuron a binary unit, which seems dubious. We will clarify that by “predict velocity from each current and next encoding” we mean that the normative constraint they enforce is axiom 1: sequential activity of sets of neurons i then j can be uniquely interpreted as a trajectory, i.e. a step or velocity. Their work is elegant, and we will try to do more justice to it in the revision.

      To conclude, we thank the reviewers for their extensive comments, and look forward to releasing a version that addresses their concerns.

    1. eLife Assessment

      This study presents a valuable and well-documented computational pipeline for the scalable analysis and spike sorting of large extracellular electrophysiology datasets, with particular relevance for high-density recordings such as Neuropixels. The authors demonstrate the pipeline's utility for benchmarking spike sorter performance and evaluating the effects of data compression, supported by thorough testing, clear figures, and openly available code. The workflow is reproducible, portable, and practical, providing concrete guidance on computational cost and runtime. Overall, the evidence supporting the pipeline's performance and output quality is compelling, and this work will be of broad interest to the systems neuroscience community.

    2. Reviewer #1 (Public review):

      Summary:

      Extracellular electrophysiology datasets are growing in both number and size, and recordings with thousands of sites per animal are now commonplace. Analyzing these datasets to extract the activity of single neurons (spike sorting) is challenging: signal to noise is low, the analysis is computationally expensive, and small changes in analysis parameters and code can alter the output. The authors address the problem of volume by packaging the well-characterized SpikeInterface pipeline in a framework that can distribute individual sorting jobs across many workers in a compute cluster or cloud environment. Reproducibility is ensured by running containerized versions of the processing components.

      The authors apply the pipeline in two important examples. The first is a thorough study comparing the performance of two widely used spike-sorting algorithms (Kilosort 2.5 and Kilosort 4). They use hybrid datasets created by injecting measured spike waveforms (templates) into existing recordings, adjusting those waveforms according to the measured drift in the recording. These hybrid ground truth datasets preserve the complex noise and background of the original recording. Similar to the original Kilosort 4 paper, which uses a different method for creating ground truth datasets that include drift, the authors find Kilosort 4 significantly outperforms Kilosort 2.5. The second example measures the impact of compression of raw data on spike sorting with Kilosort 4, showing that accuracy, precision, and recall of the ground truth units is not significantly impacted even by lossy compression. As important as the individual results, these studies provide good models for measuring the impact of particular processing steps on the output of spike sorting.

      Strengths:

      The pipeline uses the Nextflow framework, which makes it adaptable to different job schedulers and environments. The high-level documentation is useful, and the GitHub code is well organized. The two example studies are thorough and well-designed and address important questions in the analysis of extracellular electrophysiology data.

      Weaknesses:

      There are no major weaknesses in the revised manuscript. While no data analysis pipeline can cover the needs of all experiments, the authors have added and significant flexibility in the pipeline. Even experimenters who might opt for a simpler pipeline will benefit from this work as a model.

    3. Reviewer #2 (Public review):

      Summary:

      This work presents a reproducible, scalable workflow for spike sorting that leverages parallelization to handle large neural recording datasets. The authors introduce both a processing pipeline and a benchmarking framework that can run across different computing environments (workstations, HPC clusters, cloud). Key findings include demonstrating that Kilosort4 outperforms Kilosort2.5 and that 7× lossy compression has minimal impact on spike sorting performance while substantially reducing storage costs.

      Strengths:

      (1)Extremely high-quality figures with clear captions that effectively communicate complex workflow information.

      (2) Very detailed, well-written methods section providing thorough documentation.

      (3) Strong focus on reproducibility, scalability, modularity, and portability using established technologies (Nextflow, SpikeInterface, Code Ocean)

      (4) Pipeline publicly available on GitHub with documentation.

      (5) Clear cost analysis showing ~$5/hour for AWS processing with transparent breakdown.

      (6) Good overview of previous spike sorting benchmarking attempts in the introduction

      (7) Practical value for the community by lowering barriers to processing large datasets.

      Weaknesses

      No significant weaknesses. The authors have responded to all my review critiques and suggestions.

    4. Reviewer #3 (Public review):

      Summary:

      The authors provide a highly valuable and thoroughly documented pipeline to accelerate the processing and spike sorting of high-density electrophysiology data, particularly from Neuropixels probes. The scale of data collection is increasing across the field, and processing times and data storage are a growing concern. This pipeline provides parallelization and benchmarking of performance after data compression that helps address these concerns. The authors also use their pipeline to benchmark different spike sorting algorithms, providing useful evidence that Kilosort4 performs the best of out the tested options. This work, and the ability to implement this pipeline with minimal effort to standardize and speed up data processing across the field, will be of great interest to many researchers in systems neuroscience.

      Strengths:

      The paper is very well written and clear. The accompanying GitHub and ReadTheDocs are well organized and thorough. Benchmarks are exceptionally well applied to support the authors' claims, and it is clear that the pipeline has been very thoroughly tested and optimized by users at the Allen Institute for Neural Dynamics. The pipeline incorporates existing software and platforms that have also been thoroughly tested (such as SpikeInterface), so the authors are not reinventing the wheel, but rather putting together the best of many worlds. In the latest revision, the authors add a nice analysis showing that compression mostly affects the lowest SNR units. This is a great contribution to the field and it is clear the authors have put a lot of thought into making the pipeline as accessible as possible.

      Weaknesses:

      None noted. The authors have addressed all previous questions and requests for clarification.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Weaknesses:

      The pipeline is very complete, but also complex. Workflows (optimal artifact removal, best curation for data from a particular brain area or species) will vary according to experiment. Therefore, a discussion of the adaptability of the pipeline in the “Limitations” section would be helpful for readers.

      We added a dedicated paragraph in the Discussion section under “Limitations” focusing explicitly on the adaptability and flexibility of the pipeline. Furthermore, we took this feedback as an opportunity to make the pipeline itself significantly more modular and customizable with the most recent release (v1.2.0: https://aind-ephys-pipeline.readthedocs.io/en/latest/releases/1.2.0.html).

      Reviewer #1 (Recommendations for the authors):

      (1) In the description of the Phase-shift correction (Line 166-167): The current text reads “As a result, different groups of channels are sampled asynchronously.” A better description would be: “Sample times for different groups of channels are offset in time by a known amount.”

      We replaced the phrase in the manuscript text with the suggested formulation.

      (2) Figure 5 and description of the benchmarking overview (Line 326-336): How were spike trains (times) selected for the injected ground truth units? What was the range of firing rates?

      All injected spike trains were generated as independent Poisson processes featuring a mean firing rate of 15 Hz. We have now incorporated this explicitly into the main text to clarify the ground-truth injection process.

      (3) Figure 6, panel b: Are the gray points in the raster the original spikes in the test recording? From the pattern, it looks like there are 8 recovered ground truth units. Were the other 2 undetected by either sorter?

      That is correct; the two remaining units were undetected by both sorters. To clear up any confusion, we updated the caption for Figure 6 to state: “Note that spikes undetected by any of the sorter are not shown in the plot.”

      (4) Figure 7, panel c: Are all units returned from KS included in these distributions? (i.e., regardless of the KS refractory metric calculated by the sorter) - it would be useful to add that detail to the caption. It would also be helpful for panel C to include a total unit count from the two sorters... Also, since there are multiple ways to calculate the refractory period contamination, it would be good to state the calculation used here.

      Because we rely directly on the hybrid ground-truth for accurate validation, we included all raw units returned by Kilosort for this specific analysis. We have explicitly added a note detailing this to the caption. Panel C does report the total raw unit count returned by the two sorters (N = 3046 for KS2.5; N = 3652 for KS4).

      Additionally, to clarify the evaluation procedure, we appended the following statement to the main text: “For all results, we perform spike train comparisons and compute performance metrics as defined in (Buccino et al. 2020), using all units returned by the spike sorter (without any sorterspecific curation).”

      (5) Comments about the pipeline:

      The paper clearly demonstrates the immense utility of the pipeline in the authors’ work. I did some testing to try to understand its adaptability to workflows at my institution.

      I tested the pipeline on our local cluster running LSF. I’ve worked on a similar pipeline using Nextflow to automate ephys analysis with the same sorters. Questions that came up for me that would be usefully addressed in the ’Limitations’ section:

      (i) Is the pipeline meant to be run only in total? In particular, is it possible to start with preprocesseddata? (aind-ephys-preprocessing/code/params.json does not appear to include any means to turn off filtering, for example). Is the pipeline meant to be run only in total? In particular, is it possible to start with preprocessed data? (aind-ephys-preprocessing/code/params.json does not appear to include any means to turn off filtering, for example).

      To accommodate users who wish to run only parts of the workflow or use external preprocessing setups, we have refactored the codebase to support a custom preprocessing pipeline option. This makes it possible to turn off standard filtering or inject custom workflows.

      (ii) For debugging purposes, is there a means to go from preprocessing or sorting to result collection,so that interim results can be interpreted even when some steps of the pipeline aren’t working?

      The pipeline is designed to be a spike sorting pipeline, so the spike sorting step cannot be skipped. However, we have rewritten the post-sorting architecture to make it highly lightweight and fault-tolerant. The postprocessing step now only requires the random spikes and templates computation and downstream steps have been update to accomodate this lightweight option. As an example, if no quality metrics are computed, the curation step will be skipped. The visualization and QC steps also required updates to be tolerant to missing extensions. This required coordinate updates across several components:

      Postprocessing: PR #12

      Curation: PR #13

      Visualization: PR #21

      Quality Control: PR #20

      (iii) If these options to skip processes and output data ’partway’ are available, it would be great toadd that to the documentation.

      We have fully updated our online documentation for v1.2.0 (release notes: https://aind-ephys-pipeline.readthedocs.io/en/latest/releases/1.2.0.html), introducing a brandnew “Customization” guide page that comprehensively explains how to construct and provide custom preprocessing and postprocessing strategies, as well as how to integrate a new spike sorter in the pipeline: https://aind-ephys-pipeline.readthedocs.io/en/latest/customization.html

      Reviewer #2 (Public review):

      Summary:

      This work presents a reproducible, scalable workflow for spike sorting that leverages parallelization to handle large neural recording datasets. The authors introduce both a processing pipeline and a benchmarking framework that can run across different computing environments (workstations, HPC clusters, cloud). Key findings include demonstrating that Kilosort4 outperforms Kilosort2.5 and that 7× lossy compression has minimal impact on spike sorting performance while substantially reducing storage costs.

      Strengths:

      (1) Extremely high-quality figures with clear captions that effectively communicate complex workflow information.

      (2) Very detailed, well-written methods section providing thorough documentation.

      (3) Strong focus on reproducibility, scalability, modularity, and portability using established technologies (Nextflow, SpikeInterface, Code Ocean).

      (4) Pipeline publicly available on GitHub with documentation.

      (5) Clear cost analysis showing ~$5/hour for AWS processing with transparent breakdown.

      (6) Good overview of previous spike sorting benchmarking attempts in the introduction.

      (7) Practical value for the community by lowering barriers to processing large datasets.

      Weaknesses:

      No significant weaknesses were identified, although it is noted that the limitations section of the discussion could be expanded.

      We thank the reviewer for their constructive feedback on our manuscript.

      Reviewer #2 (Recommendations for the authors):

      The authors could discuss why 2.25 bps is the “lowest supported” level and whether more aggressive compression could be achieved with custom approaches, potentially exploring where performance breakdown occurs.

      The 2.25 bits-per-sample (bps) limit is an inherent constraint of the WavPack lossy compression library itself. While more aggressive, domain-specific, or custom compression schemes could be explored, we focused on WavPack due to its native support in modern neurophysiology ecosystems and its excellent performance in our prior simulated benchmarks (Buccino et al. 2023). We agree that using this hybrid benchmarking framework to explore alternative compression configurations is a highly valuable avenue for future work. We have added the following text to the Discussion: “The benchmarking pipeline will continue to develop as an open evaluation framework, enabling transparent and reproducible comparisons of spike sorting and preprocessing methods across the community. As one example, the work on lossy compression could be extended with additional codecs and parameter settings, exploiting our ability to read out spike sorting degradation directly from the hybrid ground truth spike times.”

      (2) The limitations section would benefit from expansion to include: (i) discussion of how simulated data limitations may affect generalization of benchmarking results to real neural data, and (ii) clarification of the effort required to add new spike sorters, including configuration complexities for coordinating Nextflow processes beyond simple SpikeInterface integration.

      We have expanded the Discussion section to address both items:

      (i) We added a paragraph detailing the specific limitations of hybrid ground-truth datasets (e.g., how idealized template injection might miss extreme multi-unit overlapping dynamics or nonstationary noise properties found in real tissue).

      (ii) We added a structural overview section clarifying the workflow complexity, detailing exactly what steps are required to map a new spike sorter into a Nextflow execution processes beyond its baseline addition to Spike Interface.

      (3) The authors should clarify the terminology of “hypothetical experiment” in the introduction to improve reader comprehension.

      We have removed the word hypothetical from the introduction to ground the explanation more directly.

      (4) The cost analysis could be improved by making it clearer whether “runtime” refers to wall-clock vs. total parallel compute time.

      We mean wall-clock time. While total parallel compute time aggregated across cloud workers remains roughly identical to the overall sequential execution on a lone cloud instance, cluster parallelization slashes the wall-clock time drastically. We have updated the text to explicitly state that reported runtimes represent wall-clock time.

      (5) The authors could address the Nextflow Java dependency limitation by discussing containerized execution options (Docker/Singularity) as a solution, while noting relevant HPC system restrictions.

      We have updated the text to mention the official pre-built Nextflow container images as an elegant workaround for environments where local Java installations are blocked or restricted: “However, one option to bypass installation issues is to run the main pipeline script in container images packaged with Nextflow (https://hub.docker.com/r/nextflow/nextflow).”

      (6) Figure 8 analysis would be strengthened by explicitly noting that compression effects are more substantial for lower-accuracy units, suggesting better preservation of higher SNR units.

      We appreciate this insight. To evaluate this systematically, we generated a new supplementary figure (Figure S3) which shows sorting performance during lossy compression as a function of the Signal-to-Noise Ratio (SNR) of ground truth units. The plot demonstrates that for Neuropixels 2.0 recordings, the slight drop in sorting accuracy is indeed heavily concentrated among low-SNR units. We have integrated this observation into the Results section.

      Reviewer #3 (Public review):

      (1) Could the authors please expand on the statement on line 274, that processing their test dataset serially “on a single GPU-capable cloud workstation... would take approximately 75 hours and cost over 90 USD.” How were these values calculated? I was a bit surprised that this is a ¿4-fold slowdown from their pipeline, but only increases the cost by 1.35x... More context on why this is, and maybe some context on what a g4dn.4xlarge is compared to the other instances, might help.

      We have expanded the cost analysis section in the manuscript methods to explain these figures explicitly. The serial run relies on a single continuous, higher-tier GPU workstation instance (g4dn.4xlarge) running uninterrupted for 75 hours.

      Our distributed pipeline, by contrast, dynamically provisions CPU-only instances to process chunked preprocessing steps concurrently, then spins up short-lived GPU spot instances only when Kilosort executes. While this parallel execution compresses the overall wall-clock time by over 4-fold, the cost is only moderately reduced because the CPU-only instances with many parallel processing cores are only slightly less expensive than GPU instances.

      (2) One of the most commonly used preprocessing pipelines for Neuropixels data is the CatGT/ecephys pipeline from the developers of SpikeGLX at Janelia. It may be worth commenting very briefly... on how the preprocessing steps available in this pipeline compare to the steps available in CatGT. For example, is “destriping” similar to the “-gfix” option in catGT to remove high-amplitude artifacts?

      We have added a section drawing direct comparisons to CatGT preprocessing workflows. We explicitly clarify that our phase-shift correction performs the exact same function as CatGT’s Tshift. We also point out that while our current version lacks a direct equivalent to CatGT’s saturation removal feature (-gfix), this capability is scheduled for incorporation in our upcoming pipeline release.

      (3) Why are there duplicate units (line 194), and how often is this an issue? I understand that this is likely more of a spike sorter issue than an issue with this pipeline, but 1-2 sentences elaborating why might be helpful for readers.

      Duplicate units are primarily an artifact of template-matching sorting routines (such as Kilosort), which can occasionally split a single biological neuron into multiple overlapping spatial templates or over-extract templates in highly active channel regions. We have added two clarifying sentences explaining this phenomenon in the text: “Next, duplicated units, that can arise when using template-matching methods if different templates are consistently fit to the same spikes, are removed based on the fraction of overlapping spikes.”

      Customizability of cluster curation parameters It seems from the parameter files on GitHub that the cluster curation parameters are customizable - correct? If so, it may be worth explicitly saying so in the curation section of the text... A presence ratio of >0.8 could be particularly problematic for some recordings (e.g. state transitions, behavior specific cells).

      (4) Yes, they are completely customizable. We agree that a rigid presence ratio cutoff of 0.8 would erroneously discard highly valid units that are modulated by specific behavioral states, or are active only during sleep vs. wake cycles. We have explicitly added text in the Curation section clarifying that all quality metric thresholds can be modified by the user: “Units are tagged as passing a default_qc when they satisfy the following criteria based on quality metrics thresholds. Thresholds can be user defined, and these are the default”.

      (5) The axis labels in Figures 3d-e are too small to see, and Figure 3d would benefit from a brief description of what is shown.

      We have updated the figures with enlarged, high-visibility axis labels and expanded the caption of Figure 3d to clearly describe the visualization.

      Figure 4 labels (“neural” vs “passing QC”) (6) What is the difference between “neural” and “passing QC” in Figure 4?

      We have updated the figure caption for Figure 4 to include an explicit cross-reference to the Curation methodology section, which defines the strict quantitative boundary between raw neural classification and formal automated QC passage.

      (7) I understand the current paper is focused on spike data... but I am curious about the NP2.0 probes that save data in wideband. Does the lossy compression negatively affect the LFP data? Is software filtering applied for the spike band before or after compression?

      Compression is applied to the raw streams prior to any secondary downstream software processing. For Neuropixels 1.0, compression is executed strictly on the action potential (AP) stream. For Neuropixels 2.0, compression operates directly on the unified wide-band data stream.

      Software filtering to separate bands is conducted post-decompression, as captured in our baseline workflow definitions (e.g., WavPack compression → decompression → preprocessing → Kilosort4). To clarify this, we added the following text: “In all cases, compression was applied before any preprocessing took place. For Neuropixels 1.0, we compressed the AP stream only. For Neuropixels 2.0, we compressed the full wide-band data.”

      Because LFP signals possess inherently smooth continuous dynamics across both space and time, they are much more amenable to lossless or near-lossless compression. Thus, the minor losses introduced by lossy compression are overwhelmingly localized to high-frequency spike band features, leaving LFP components virtually unaffected.

    1. eLife Assessment

      The authors describe a valuable finding that the Streptococcus pyogenes secreted protease SpeB is expressed in response to protease activity that degrades the Vfr repressor. Proteases can be released from host neutrophils (possibly by NETosis), as well as a positive feedback mechanism by SpeB itself. The authors utilize a dual fluorescent reporter system to simultaneously read speB and capsule gene expression, providing solid evidence that demonstrates that proteases can regulate Vfr; however, the data indicating that this is physiologically relevant and that extracellular traps themselves have a functional role are incomplete. This work will be of interest to microbiologists studying the regulation of virulence factors at the host-pathogen interface.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript examines how Streptococcus pyogenes regulates expression of the virulence factor SpeB in response to both bacterial and host-derived cues. The authors propose that Vfr acts as a repressor of speB expression and that degradation of Vfr by SpeB or by neutrophil-derived proteases relieves this repression. This creates a model in which S. pyogenes can sense proteolytic activity during infection and use that information to tune virulence factor expression.

      Strengths:

      The main strength of the study is the bacterial regulatory mechanism. The dual reporter system provides a useful way to follow speB and hasABC expression, and the genetic analysis of known regulators helps validate the system. The media-swap experiments, recombinant Vfr experiments, and SpeB-mediated degradation of Vfr support the conclusion that Vfr represses speB and that proteolysis can relieve this repression. The finding that SpeB can degrade Vfr is particularly interesting because it suggests an autoregulatory mechanism that could reinforce SpeB expression once it has been initiated.

      Weaknesses:

      The host side of the model is less completely supported. The authors show that neutrophil lysates and protease-containing fractions can induce the speB reporter and degrade Vfr, which supports the idea that neutrophil-derived proteases can affect this circuit. However, the in vivo interpretation relies heavily on PAD4-deficient mice to implicate neutrophil extracellular traps. PAD4 deficiency is a useful perturbation, but it does not by itself distinguish loss of extracellular trap formation from changes in neutrophil recruitment, survival, degranulation, phagocytosis, oxidative killing, or other neutrophil death pathways. As a result, the current data support a role for neutrophil-associated proteolytic activity more strongly than they support a specific role for extracellular traps. This distinction is important for interpreting the central model. The bacterial circuit is well developed, but the host-derived cue remains somewhat underdefined. If the relevant signal is extracellular protease activity more broadly, then the model is still interesting, but the conclusion should be framed around neutrophil-derived proteolytic stress rather than extracellular traps specifically. If extracellular traps are the key in vivo source of protease exposure, then additional evidence would be needed to separate that mechanism from other neutrophil effector functions that remain intact in PAD4-deficient cells.

      Overall:

      This is a valuable study with solid evidence for a bacterial protease-sensing regulatory mechanism controlling SpeB expression. The work should be useful to investigators interested in bacterial virulence regulation, host-pathogen interactions, and how pathogens integrate immune-derived cues during infection. The impact of the study would be stronger if the host-derived signal were defined more precisely, but the bacterial Vfr-SpeB circuit provides a compelling framework for thinking about how S. pyogenes links proteolytic activity to virulence gene expression.

    3. Reviewer #2 (Public review):

      Summary:

      The study examines how Streptococcus pyogenes integrates bacterial and host-derived signals to regulate SpeB, proposing that Vfr acts as a protease-sensitive repressor whose degradation relieves repression of speB. The authors further suggest that neutrophil-derived serine proteases, including those associated with inflammatory conditions, may promote this transition, and thereby counterbalance LL-37/CovRS-associated suppression of speB. The conceptual framework is interesting and potentially important for understanding how host inflammation feeds into bacterial virulence regulation.

      Strengths:

      The work addresses a biologically significant question and does so using a broad and generally well-integrated experimental approach, including bacterial genetics, reporter assays, recombinant protein analyses, neutrophil-derived material, human blood infection, and mouse infection models. A particular strength is the effort to connect host inflammatory processes to bacterial regulatory behavior, which gives the study conceptual reach beyond a narrow mechanistic observation. The data support the view that Vfr is relevant to speB control and that neutrophil-associated protease activity may influence this pathway.

      Weaknesses:

      The main limitations are mechanistic. The physiological form, localization, and abundance of Vfr are not sufficiently defined to support the proposed model at full strength, and the evidence that Vfr functions as a SpeB-labile repressor under biologically relevant conditions remains incomplete. The relationship between Vfr and the broader RopB/SIP regulatory framework is also not yet firmly established. In addition, the reporter system is not yet benchmarked closely enough against endogenous SpeB protein output, and its growth-phase dependence is insufficiently resolved, which makes it difficult in some settings to distinguish promoter activity from mature protease production. The neutrophil protease component is likewise not defined beyond a general serine protease signal, and the potentially important LL-37/CovRS/Vfr connection is underdeveloped in the main text. Overall, the conceptual advance is promising, but several of the central mechanistic claims would benefit from more direct experimental support and more cautious framing.

    4. Reviewer #3 (Public review):

      Summary:

      SpeB is a cysteine protease secreted during infection by Streptococcus pyogenes (Spy). SpeB has been extensively investigated for its role in pathogenesis, which involves proteolytic processing of both Spy virulence factors and host proteins. Regulation of speB expression is complex and includes growth phase regulation, a quorum-sensing system, the transcription factor RopB, and the global regulatory system CovRS (CsrRS). Guerra et al now attempt to refine the current model of regulation of SpeB expression, focusing on the Spy protein Vfr, which has been suggested previously to act as a negative regulator of SpeB expression. In the current study, neutrophil lysates (representing proteases released during NETosis) are shown to degrade Vfr and to relieve repression of SpeB. At high cell density, SpeB itself also degrades Vfr, which may allow autoregulation of SpeB expression. These observations are unsurprising as the broad protease activities of both neutrophil proteases and SpeB are well known. Nonetheless, the data presented fill in additional details in our understanding of the complex regulation of an important Spy virulence factor.

      Strengths:

      (1) Construction of a GFP reporter strain provided a facile methodology for tracking speB promoter activity in a variety of experimental setups.

      (2) A Vfr deletion mutant was a useful tool to investigate the role of Vfr in SpeB regulation, and mutants in speB and ropB were important controls.

      (3) Experiments using neutrophil lysates in vitro, as well as in vivo studies of mice depleted of neutrophils with anti-Ly6G or in PAD4-/- mice (that cannot form NETs) support the hypothesis that neutrophil proteases derepress speB expression by degrading Vfr.

      Weaknesses:

      (1) The introduction and all the experiments in Figure 1 focus on CovRS, which turns out to be largely tangential to the overall story developed by the rest of the study. On the other hand, the complex and well-studied regulation of speB expression by RopB and the SIP quorum-sensing system is only minimally described. A better framing would be a more detailed introduction to the current model of speB/RopB/SIP/quorum sensing/growth phase regulation. CovRS could be introduced later as its relevance is really just to show that neutrophil lysates or NETs do more than simply providing LL-37, which signals through CsrS, as another regulator of speB expression.

      (2) Vfr, as the central focus of the paper, also deserves a more thorough introduction to provide context for the study. For example, reference 19 (Shelburne et al, 2011) showed reduced transcription of speB in a vfr mutant, an effect that could be complemented by expressing vfr or a 39-aa N-terminal fragment in trans. That study presented evidence that the N-terminal peptide binds to RopB, which may prevent RopB from upregulating SpeB expression. Do the authors concur with that model? As it stands, the discussion and model in Figure 1A imply a direct regulatory effect of Vfr on speB expression rather than an indirect one through regulation of RopB. If direct regulation of speB by Vfr is a consideration, it should be investigated more thoroughly, e.g., by promoter-binding assays, CHIP-seq, etc.

      (3) Use of single-cell flow cytometry generally confirmed results observed in batch culture. The authors also comment repeatedly on the heterogeneity of individual cell fluorescence representing both speB and has operon expression. However, the reason(s) for heterogeneity in gene expression are not explored, e.g., differences in individual cell growth rate in batch culture, variable loss of reporter plasmid during infection experiments, etc).

      (4) Lines 116-118 and Figure 3C: Incubation of recombinant Vfr with Spy Dvfr reduced SpeB expression, but the degree of suppression is modest compared to that seen in wild-type Spy. How does the concentration of rVfr added compare to that present in the culture fluid of wild-type Spy? (Also, the concentration of rVfr used is unclear: the figure says 3 µg/ml and the legend says 0.3 mg/ml, i.e., 300 µg/ml).

      (5) Lines 125-126: "...the Vfr structure contains several potential protease SpeB cleavage sites..." The role of Vfr in degrading SpeB could be clarified by identifying the predicted cleavage products, e.g., by mass spec, after co-incubation of the two recombinant proteins.

      (6) Lines 122-124: "Notably, speB expression in Spy Dvfr is unaffected by LL-37 or MgCl2, further validating its [Vfr's?] dominance over CovRS regulation." This statement is an oversimplification and is potentially misleading: LL-37 is degraded by SpeB (Nyberg et al, JBC 2004), which likely explains why the addition of LL-37 fails to signal through CovRS to repress SpeB in Spy Dvfr since SpeB is produced continuously in that strain. By contrast, SpeB is only produced during the stationary phase in the wild type, so LL-37 remains active throughout the exponential phase and represses SpeB expression. The response to the CovRS ligand MgCl2 is similar (or greater) in Spy Dvfr compared to wild type (Figure S2C).

      (7) Lines 153-154 and Figure 6E: Growing wild type Spy in the presence of neutrophil lysates with or without a protease inhibitor stimulated or repressed speB expression in a manner consistent with degradation (or not) of Vfr. It would be confirmatory and informative to do the same experiment with the Spy Dvfr strain.

      (8) Clarity of writing could be improved, particularly by eliminating pronouns of indefinite reference (it, its, this) in contexts in which the subject is ambiguous (examples at lines 62, 89, 111, 114, 115, 123, 183, 190, 193, 204, 205, 210, 217, 221, 222, 224).

    1. eLife Assessment

      The study presents valuable findings of a new E. coli cell-free protein synthesis (eCFPS) system that has been simplified by reducing the number of core components from 35 to 7; furthermore, the findings communicate a simplified 'fast lysate' preparation that eliminates the need for traditional runoff and dialysis steps. It is interesting that the system's robustness is exhibited by its applicability to nanoluc, a protein that expresses readily in many systems, to more challenging proteins like the functional self-assembling vimentin and the active restriction endonuclease Bsal. Despite the study representing an advancement towards simplifying protein expression workflows, the evidence is solid and supports the main claims however minor weakness exists i.e. the efficiency claims about the new system needs to be supported by accurate comparisons with typical cell free expression systems, in addition, investigations into the mechanistic basis of the observations would provide more evidence. Despite this shortcoming, the paper remains of interest to scientists in cell and molecular biology, microbiology, biotechnology and protein synthesis.

    2. Reviewer #1 (Public review):

      Summary:

      The authors presented a simplified E. coli cell-free protein synthesis (eCFPS) system reduces core reaction components from 35 to 7, improving protein expression levels. They also presented a "fast lysate" protocol that simplifies extract preparation, enhancing accessibility and robustness for diverse applications.

      Strengths:

      The authors present a valuable new protocol for eCFPS, which simplifies its application.

      Weaknesses:

      The authors provide data for optimization but offer insufficient explanation of the fundamental mechanisms underlying the phenomenon based on data.

      Comments on revised version.

      The authors have satisfactorily addressed the concerns raised by the reviewers. However, the mechanistic basis of the observed performance gain remains insufficiently substantiated. The attribution of this improvement to enhanced transcription is currently speculative. This point could be directly tested by quantifying mRNA levels, for example, using real-time PCR, in both the initial and optimized systems. Such analysis would significantly strengthen the mechanistic interpretation of the results.

    3. Reviewer #2 (Public review):

      Summary:

      The authors have made a convincing argument that the current system of in vitro translation using E. coli extracts can be significantly optimized to work with much lesser components, while maintaining activity. They have showcased their improved activity using not only physical but also functional readouts.

      Strengths:

      The experiments are designed in a very logical and easy to understand manner, which makes it easier not only to follow the paper, but also reproduce the results. Functional assays with the synthesized proteins are a good way to demonstrate functionality and applicability of the system. They also benchmark their system against a commercial kit to show superior performance of their system.

      Weaknesses:

      The production of the lysate requires special instrumentation, limiting accessibility.

      Comments on revised version:

      Thank you to the authors for addressing the concerns both textually and experimentally. This work has significant value.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to overcome the challenges associated with complex, conventional prokaryotic cell-free protein synthesis (CFPS) systems, which require up to thirty-five components, by developing a streamlined and efficient E. coli CFPS platform to encourage broader adoption. The main objective was to reduce the number of reaction components from thirty-five to seven, while also developing an accessible 'fast lysate' preparation protocol that eliminates time-consuming runoff and dialysis steps. The authors also sought to demonstrate the robustness and translational quality of this streamlined system by efficiently synthesising challenging functional proteins, including the cytotoxic restriction endonuclease BsaI and the self-assembling intermediate filament protein vimentin.

      Strengths:

      This study presents several key strengths of the optimised E. coli cell-free protein synthesis system in terms of its design, performance and accessibility.

      - The reaction mixture has been dramatically simplified, with the number of essential core components successfully reduced from up to thirty-five in conventional systems to just seven.

      - The "fast lysate" protocol is a significant advance in terms of procedure.

      - The system's ability to synthesise challenging, functional proteins is evidence of its robustness.

      Weaknesses:

      (1) Title: "A simplified and highly efficient cell-free protein synthesis system for prokaryotes".

      - This title is misleading since one would expect a simplified and highly efficient cell-free protein synthesis system to yield similar protein levels compared to current cell-free protein synthesis systems. What this study shows is that the composition of cell-free protein synthesis systems can be simplified while maintaining a certain level of protein synthesis. Here, optimisation does not involve maintaining protein synthesis yield while simplifying the cell-free protein synthesis system; rather, it involves developing a simplified cell-free protein synthesis system. As mentioned in my comments below, this study lacks a comparison of protein levels with a typical cell-free protein synthesis system.

      - What do the authors mean by "highly efficient"? Highly efficient compared to what experimental conditions? If one is interested by the yield of protein synthesis, is this simplified system highly efficient compared to current systems?

      (2) Figure 1, 3-5:

      - What do relative luciferase units represent? How are these units calculated?

      - In this system, the level of expression depends mainly on the level of NLuc transcripts and the efficiency of NLuc translation. How did the authors ensure that the chemical composition of the different eCFPS buffers only affected protein translation and not transcript levels? In other words, are luciferase units solely an indicator of protein synthesis efficiency, or do they also depend on transcription efficiency, which could vary depending on the experimental conditions?

      - How long were the eCFPS reactions allowed to proceed before performing the luciferase activity measurement? Depending on the reaction time, the absence or presence of certain compounds may or may not impact NLuc expression. For example, it can be assumed that tRNA does not significantly affect NLuc levels over a short period of time, and that endogenous tRNA in the lysate is present at sufficient concentrations. However, over a longer period of time, the addition of tRNA could be essential to achieve optimal NLuc levels.

      - The authors show that tRNA and amino acids are not strictly essential for the expression of NLuc, likely due to residual amounts within the cell lysate. However, are the protein levels achieved without added amino acids and tRNA sufficient for biochemical assays that require a certain amount of protein? It is important to note that the focus here is on optimising the simplicity of the buffer rather than the level of protein expression. In fact, the simplicity of the buffer is prioritised over the amount of protein produced. This should be made clear.

      - How would the NLuc level compare if all the components were optimised individually and present in an optimised buffer, compared to a buffer optimised for simplicity as described by the authors?

      (3) Line 71, Streamlining eCFPS: removal of dispensable components. This title is misleading because it creates the false impression that proteins can be produced in vitro without the addition of certain compounds. While this is true, the level of protein produced may not be sufficient for subsequent biochemical analyses. This should be made clear.

      (4) Figure 2: In the legend, change "(A) Protein expression levels of the eCFPS system measured at varying concentrations of KGlu and MgGlu2" to "(A) Protein expression levels of the eCFPS system using an Nanoluciferase (NLuc) reporter DNA measured at varying concentrations of KGlu and MgGlu2".

      (5) Lanes 302-303: "The thorough optimization of the seven core components was a critical step in achieving high protein expression levels". What are "high expression levels"? Compared to what?

      Comments on revised version.

      The authors have adequately addressed my previous concerns.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      The superiority of the optimized system might simply be due to insufficient T7 RNA polymerase in the initial lysate.

      We performed a T7 RNA polymerase titration (0–1600 ng/µL) in the initial system to test this hypothesis. Standard CFPS protocols typically utilize T7 RNA polymerase at ~90–100 ng/µL<sup>1</sup>. To fully characterize the concentration-dependent effect and determine the exact saturation threshold of T7 RNA polymerase in our system, we tested an extended range from 0 to 1600 ng/µL. As shown in the revised Figure S3B, the initial system's output reaches a plateau at ~800 ng/µL—a concentration nearly ten times higher than standard protocols. Increasing the concentration further (up to 1600 ng/µL) led to a decline in yield, likely due to inhibitory effects of excess enzyme or buffer components. Even under these T7-saturated conditions, our optimized system achieved ~45-fold higher NLuc output compared to the maximum possible output of the initial system. Notably, when the lysate concentration is increased to 70%, the productivity gap reaches nearly 80-fold, further demonstrating the extraordinary efficiency of our platform.

      As revised in the Discussion, this improvement confirms that the performance gain is not a result of a mere increase in T7 concentration. Instead, it represents a systemic synergy where our streamlined buffer and the optimized metabolic environment of the fast lysate together alleviate the transcriptional bottlenecks inherent in traditional platforms.

      Reviewer #2 (Public review):

      Performance or efficiency claims... needs to be supported by comparisons with typical cell free expression systems.

      We agree that robust benchmarking is essential for validating our claims of high efficiency. Our comparative evaluation was conducted across three levels:

      (1) Literature-based benchmarking: As detailed in Figures 3C, 4A-D, S3A-B, S4, and S5C, we extensively compared our system against the "initial" (35-component) and "PEPbased" platforms, which are established benchmarks widely utilized in CFPS literature. These diverse comparisons consistently demonstrate the superior performance and robustness of our optimized system across various conditions.

      (2) Commercial benchmarking: To provide independent verification, we performed a head-to-head comparison with a high-end commercial E. coli CFPS kit (PePExpress, Shanghai Epizyme, EC010L). As shown in the comparative data provided in this response (See author response image 1), our system exhibited remarkable rapid-expression capability, significantly outperforming the commercial kit in both speed and absolute yield. Our platform reached near-maximum yield within 2 hours, demonstrating a significant efficiency advantage over the commercial alternative.

      (3) Robustness and translational quality: The comparison was extended to challenging targets beyond standard reporters. As shown in Figures 4E-H, the successful synthesis of active BsaI restriction enzyme (a cytotoxic protein) and the functional assembly of vimentin (an aggregation-prone protein) demonstrate that our optimized system maintains superior translational quality and robustness compared to typical platforms that often struggle with such complex targets. By outperforming established academic benchmarks and a leading commercial platform in both yield and the ability to handle challenging proteins, our results provide compelling evidence that the simplified 7component system is highly efficient. In the revised Conclusion, we have explicitly contextualized "efficiency" as the integration of high protein productivity, reduced reaction complexity, and accelerated preparation speed.

      Author response image 1.

      Comparative evaluation of sfGFP yields between our _e_CFPS system (70% lysate) and a commercial kit (PePExpress) over an 8-hour time course.

      Summary of revisions: T7 titration data have been added to Supplementary Figure S3B in the revised manuscript. To provide the additional benchmarking evidence requested, commercial comparison data (PePExpress kit) are provided in Author response image 1, while the main manuscript remains focused on the mechanistic synergy and streamlined architecture of the system.

      We hope that these substantial new data and the corresponding revisions satisfy the reviewers' queries.

      References:

      (1) Kigawa, T. et al. Cell-free production and stable-isotope labeling of milligram quantities of proteins. FEBS Lett. 442, 15–19 (1999).

    1. eLife Assessment

      The study presents important findings revealing previously unresolved conformational dynamics of the heterodimeric type IV ABC transporter TmrAB using single-molecule FRET. The evidence presented is convincing, integrating careful experimental design with computational approaches to uncover states that are typically masked and difficult to detect. The work will be of interest to scientists studying the molecular mechanisms of primary active transport processes.

    2. Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

    3. Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATP-bound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single-molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states but also enabled the real-time monitoring of protein conformational changes, precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      We thank the reviewer for this accurate and thoughtful summary of our work and its broader significance. We agree that the combination of single-molecule FRET with orthogonal validation approaches enables mechanistic resolution of conformational states and transitions that are not accessible by ensemble measurements. In particular, this framework allows direct discrimination of ATP-free and ATP-bound conformations, real-time tracking of transport cycle progression, and identification of transient intermediates in the heterodimeric ABC transporter TmrAB. We further agree that these capabilities support a generalizable strategy for dissecting conformation dynamics in related ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results, and conclusions supported by the experimental data. The authors have determined the conformational dynamics of TmrAB across different ATP concentrations, including physiological ones, and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies.

      Weaknesses:

      The scientific study needs a bit of in-depth analysis with respect to consistency in K<sub>d</sub> and its implications on the mechanism.

      The apparent K<sub>d,ATP</sub> values were determined using two complementary approaches that report on different aspects of the system. Ensemble FRET measurements yielded values of 51 ± 38 µM (TmrAB<sup>NBD</sup>), 68 ± 25 µM (TmrAB<sup>PG</sup>), and 95 ± 26 µM (TmrAB<sup>PG_EQ</sup>), which are in good agreement with previously reported biochemical estimates (~100 µM for TmrAB<sup>EQ</sup>) (Stefan et al, 2020). The slightly elevated value observed for the E→Q variant may reflect modest perturbation of nucleotide handling in this slow-turnover background. Notably, the close agreement between labeled and unlabeled variants indicates that fluorophore attachment does not measurably affect ATP binding.

      In contrast, smFRET-derived K<sub>d,ATP</sub> values (13 ± 1 µM for TmrAB<sup>NBD</sup> and 2 ± 1 µM for TmrAB<sup>PG</sup>) are systematically lower. This difference likely arises from the difficulty of deconvoluting overlapping FRET populations at sub-K<sub>d,ATP</sub> concentrations, particularly for TmrAB<sup>PG</sup>, where state assignment is less well separated. Despite this quantitative offset, both approaches consistently indicate ATP saturation well below physiological concentrations and therefore support the same mechanistic conclusion that ATP binding drives conformational switching in TmrAB.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATP-bound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I have three major points and a few minor criticisms.

      We thank the reviewer for the thoughtful and constructive evaluation of our manuscript and for highlighting the strength of combining structural and single-molecule approaches. We have addressed all major and minor points in detail below and revised the manuscript where appropriate to clarify limitations, justify analysis choices, and improve transparency.

      Major points:

      (1) The main weakness is that the authors base their conclusions on a very limited set of FRET pairs. While TmrAB has been extensively studied in terms of its structure, the authors should at least acknowledge this limitation more clearly.

      We agree that our conclusions are based on a limited number of FRET reporter pairs, and we now explicitly state this limitation in the revised manuscript. The chosen labeling positions were selected to probe two functionally critical regions—the nucleotide-binding domains and the periplasmic gate—based on prior structural and spectroscopic evidence. While this represents sparse sampling of the full conformational space, it is consistent with typical smFRET studies of membrane transporters, where experimental constraints generally limit the number of simultaneously accessible labeling positions (Asher et al, 2021; Asher et al, 2022; Levring et al, 2023; Wang et al, 2020).

      Importantly, both independent reporter variants yield consistent ATP-dependent population shifts, supporting the robustness of the observed trends. We further clarify that additional labeling sites could, in principle, resolve finer structural sub-states; however, given the already limited population separation in the current variants, such extensions would likely provide diminishing returns in state resolvability under the present experimental conditions. This trade-off is now explicitly discussed.

      (2) Most smFRET distributions were fitted with one, two, or three Gaussians. However, in several cases, additional populations with noticeable amplitudes appear to be present (e.g., Figure 3c at 0.1 mM and 3 mM ATP; Figure 4a, apo; Figure 4c, 0.3 mM R9L). Could the authors clarify why these populations were not included in the analysis?

      We thank the reviewer for this careful observation. Low-amplitude sub-populations are occasionally detected in individual histograms; however, they were not included in the quantitative model because they do not meet criteria for reproducibility, amplitude robustness, or structural assignability. Specifically, these features vary between replicates, contribute minimally to total population, and cannot be mapped to structurally or biochemically defined states based on available cryo-EM (Hofmann et al, 2019), DEER/PELDOR (Barth et al, 2018; Barth et al, 2020), or accessible-volume simulations.

      Similar minor subpopulations have been reported in smFRET studies and often attributed to photophysical or labeling heterogeneity effects (Asher et al, 2022; Husada et al, 2018). To avoid over-parameterization, we therefore restricted analysis to reproducible, structurally supported states. This rationale is now clarified in the revised manuscript.

      (3) Figure 3c (3 mM ATP): Is it truly possible to distinguish the two states in this distribution?

      We agree that state separation in the TmrAB<sup>PG</sup> variant is limited (ΔE = 0.11), and we now explicitly acknowledge this constraint in the manuscript. To improve robustness under these conditions, we used a constrained fitting strategy in which the apo-state distribution was fixed from nucleotide-free measurement, reducing parameter degeneracy during fitting of ATP-bound datasets.

      While single-molecule trajectory-based approaches such as Hidden Markov Modeling would be ideal for resolving dynamic interconversion, this was not feasible due to the low fraction of dynamic traces at the available temporal resolution. We therefore rely on population-level analysis, which remains consistent across replicates and reporter variants.

      Notably, independent measurements from two reporter positions (TmrAB<sup>NBD</sup> and TmrAB<sup>PG</sup>) yield similar ATP-bound population fractions at saturating ATP concentrations (~77% vs. ~80%), supporting the robustness of the inferred state distribution despite partial overlap.

      We have revised the manuscript to more clearly articulate methodological limitations, strengthen the justification of our analytical approaches, and improve the clarity of data presentation. These revisions enhance the transparency and robustness of the study and address the reviewer’s concerns.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are a few comments that can help to improve the study.

      (1) Line 115: The authors have checked the purity and monodispersity of the protein sample using SDS-Gel and size exclusion chromatography; however, additional characterization using negative stain electron microscopy, which clearly shows the monodispersity, will be useful.

      We agree that negative stain EM can provide an additional assessment of sample homogeneity. Given the extensive prior structural characterization (Hofmann et al, 2019; Nocker et al, 2026; Nöll et al, 2017) and the SEC profiles presented here, we believe that additional negative stain EM would unlikely provide substantial new information regarding sample homogeneity. We have clarified this point in the manuscript by explicitly referencing the relevant cryo-EM studies.

      (2) Line 116: The authors have mentioned that the enzymatic activity of TmrAB was retained after purification. Although smFRET results showing conformational dynamics of TmrAB confirm its ATPase activity, a comment on the effect of labelling on ATPase activity will be useful.

      We appreciate this important point. Previous studies on spin-labeled TmrAB<sup>NBD</sup> demonstrated transport activity comparable to wild-type TmrAB, indicating that cysteine substitution and label conjugation do not substantially perturb this variant (Barth et al, 2018). In addition, AV simulations showed that fluorophores at the TmrAB<sup>NBD</sup> labeling positions do not interfere with ATP- or substrate-binding sites, supporting the conclusion that FRET labeling does not affect ATP binding, hydrolysis, or transport. For TmrAB<sup>PG</sup>, however, equivalent transport data were not available, and AV simulations suggested interference of fluorophores with periplasmic gate dynamics. We therefore directly compared the transport activity of LD555/LD655-labeled TmrAB<sup>PG</sup> and unlabeled wild-type TmrAB using a single-liposome transport assay with the fluorescein-labeled peptide C4F (RRYC<sup>F</sup>KSTEL) (<sup>F</sup>, fluorescein; Fig. 1– Fig. S3a). Both variants showed indistinguishable transport activity, demonstrating that fluorophore conjugation at the periplasmic gate preserves transport function.

      (3) Line 117 and Figure S1c. Please add the reference for consistency of ATPase activity with previous studies on TmrAB.

      We have added a reference to previous biochemical studies reporting comparable ATPase activity and kinetic parameters for TmrAB to support the consistency of our measurements.

      (4) Line 119: It mentions that "Cysteine-maleimide labeling of detergent-solubilized TmrAB achieved site-specific labeling efficiencies exceeding 90%". The legend of Figure S1d mentions about labeling efficiency in the range of 40-50%. A clarification will be helpful for the reader. Also, calculations can be extended to the ratio of LD555 and LD655 labels on the molecule, which can be considered in analyzing results.

      We apologize for the lack of clarity. The reported >90% labeling efficiency refers to the site-specific cysteine labeling efficiency per accessible site, as determined by dye incorporation. In contrast, the 40–50% values shown in Fig.1–Fig. S1d reflect the per-site efficiency for donor-lonely and acceptor-only populations respectively, which together account for the >90% overall labeling efficiency. We have revised the main text and figure legend to clearly distinguish between per-cysteine labeling efficiency and the fraction of correctly double-labeled molecules. We also clarify that only complexes with appropriate donor– acceptor stoichiometry were included in the smFRET analysis.

      (5) Figure 1: Line 627: This line mentions "For all simulations, TmrA is shown in blue with LD655 (orange) and TmrB in yellow with LD555 (green)." Is it (which label on which subunit) known for the experimental setup?

      We thank the reviewer for pointing out this potential source of confusion. In the experimental system, fluorophore attachment occurs stochastically. Therefore, the assignment of donor and acceptor dyes to specific subunits is random. The representation shown in Figure 1 reflects one possible configuration for visualization purposes only. We have clarified this explicitly in the figure legend to avoid misinterpretation.

      (6) Figure S1-2a. Tau value can be better represented in a graph for visual readers instead of in the form of a table, and a dotted line with the threshold (~1 ns) will give a better representation of no change. Values can be included in the graph as well.

      We appreciate this helpful suggestion. We have revised Figure S1-2a to include a graphical representation of fluorescence life times, including a reference line around ~1 ns to facilitate visual comparison. Numerical values are retained alongside the plot for completeness.

      (7) Figure 2a: Each component of the assembly has been pointed with an arrow, which can mix two components and confuse readers. It would be good to make a legend column on the left or right and depict or indicate each component of the assembly clearly.

      We have changed the labeling in Figure 2a to improve clarity by separating the components and introducing a clearer legend layout, ensuring that each element of the assembly is unambiguously labeled.

      (8) The physiological concentration of ATP can range up to 5-10 mM. A comment on choosing the ATP concentration specifically to be 3 mM would be useful for the readers.

      We appreciate this suggestion. While intracellular ATP concentrations can reach up to 5–10 mM, values around 3 mM are commonly used as physiologically relevant conditions in in vitro biochemical and biophysical studies. We selected 3 mM ATP as a representative near physiological concentration that ensures saturation of ATP-dependent conformational transitions while remaining comparable to previous studies on TmrAB (Hofmann et al, 2019; Nocker et al, 2026; Nöll et al, 2017; Stefan et al, 2020). We have clarified this rationale in the manuscript.

      (9) Figure 2c is not cited in the text.

      We thank the reviewer for noting this oversight. Figure 2c is now explicitly cited in the main text.

      (10) Results in Figure 2 and 3 have been analyzed using 2 and 3 Gaussian distributions, respectively. It would be good to explain the rationale for it.

      We appreciate that this important point was brought to our attention. The number of Gaussian components was determined based on the minimal model required to describe reproducible and structurally supported populations. For ATP titration experiments (Figure 2 and Figure 3), two populations (apo and ATP-bound) were sufficient and consistent across replicates. In contrast, three populations were required under trapping conditions (Figure 4), where an additional state (OFF<sup>open</sup>) becomes kinetically stabilized and clearly resolved. We have clarified this rationale in the manuscript.

      (11) Figure 3b: data points do not seem to be saturated with respect to ATP concentration. It needs more points beyond 3 mM. Different K<sub>d</sub> at different sites in the structure could represent differential local dynamics over the structure.

      Previous structural studies demonstrated that 1 mM ATP is sufficient to saturate both nucleotide-binding sites under trapping conditions (Hofmann et al, 2019), indicating that the concentration range used here is adequate. Consistent with this, both ensemble and smFRET measurements approach saturation by 3 mM ATP, a near-physiological condition commonly used in biochemical studies. While additional data points above 3 mM could further define the plateau, they are unlikely to alter the mechanistic conclusion. We have clarified this point in the manuscript.

      (12) Figure 3 and Figure 1 - S1 have two different Kd values with respect to ATP concentration; both of these graphs measure conformational changes using smFRET. A comment specifying these Kd values based on single molecule verses ensemble measurement from will be helpful for readers.

      We appreciate this important point and have clarified it in the manuscript and the response to Reviewer #1 above. The K<sub>d,ATP</sub> values in Fig. 1–Fig. S1 are derived from ensemble FRET measurements, whereas those in Fig. 3 are obtained from smFRET population analysis. This difference likely arises from the difficulty of deconvoluting overlapping FRET populations at sub-K<sub>d,ATP</sub> concentrations, particularly for TmrAB<sup>PG</sup>, where state assignment is less well separated. Despite this quantitative offset, both approaches consistently indicate ATP saturation well below physiological concentrations and therefore support the same mechanistic conclusion that ATP binding drives conformational switching in TmrAB. We now explicitly distinguish these methods and their interpretation in the manuscript.

      (13) Figure 4: Slow-turnover TmrAB mutant has been employed in cysteine mutant on the PG opening side, but not towards the NBD side. Either experimental data or a comment on not pursuing it would be helpful for the reader. Similarly, experiments in the presence of peptide and in the absence of ATP, which can help to understand the role of substrate in conformational dynamics in the absence of ATP, are not pursued in this study. Along similar lines, experiments with wild type, in the presence of MgADP +/- substrate, are not shown in this study.

      We thank the reviewer for these insightful suggestions. The slow-turnover variant was specifically applied to the periplasmic gate reporter (TmrAB<sup>PG</sup>) because this construct provides direct sensitivity to outward-facing conformations, which are central to resolving the OF<sup>open</sup> state. In contrast, the NBD reporter primarily monitors nucleotide-binding domain (NBD) dimerization and is less suitable for distinguishing periplasmic conformational differences.

      Experiments in the absence of ATP but in the presence of peptide, as well as MgADP ± substrate, would indeed be valuable for further dissecting substrate effects. However, these conditions are beyond the scope of the current study, which focuses on ATP-driven conformational dynamics and the identification of kinetically hidden intermediates. We have added a statement in the Discussion to acknowledge these possibilities as directions for future work.

      (14) Figure 4, peptide concentration has been varied in the right panel. The result can also be presented as the % of OFopen and OFoccluded state with increasing concentration of peptide.

      We thank the reviewer for this suggestion. While such a plot would indeed be informative and could improve our understanding of substrate binding and substrate-induced trans-inhibition, the current dataset does not contain sufficient data points to construct a reliable concentration-dependent curve, particularly given that peptide saturation was not reached in our experiments. The characterization of substrate binding is further complicated by the presence of two distinct substrate-binding sites one in the outward-facing and one in the inward-facing state with likely completely different K<sub>d</sub> values and would require a more complex binding model. We have therefore decided against including this plot in the current manuscript. We do acknowledge, however, that future smFRET studies with improved temporal resolution are particularly well suited to investigating substrate binding to TmrAB and its effects on conformational equilibrium, and we have noted this in the Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) In all figures, can you please label the transporter schematics with the conformational states they represent?

      We thank the reviewer for this suggestion. All transporter schematics in the main and supplementary figures have been updated to include clear labels indicating the corresponding conformational states, thereby improving clarity and consistency.

      (2) As a suggestion, it may improve clarity to include the labelling positions (residue numbers) directly in Figure 1a and b, even though they are provided in the legend.

      We appreciate this suggestion. Residue numbers corresponding to labeling positions have now been added directly to Figure 1a and b to improve readability and facilitate interpretation.

      (3) Lines 183-188: This is a key point. It would be helpful to include a reference line for the expected state (0.63). Interestingly, this value coincides with the shoulder observed in Fig. 3c (0.1 mM ATP). Is there an explanation for this (see also point 2)?

      We thank the reviewer for highlighting this point. We considered adding a reference line at 0.63 to the plot; however, we decided against it. While a subpopulation does appear at ~0.63 —consistent with the expected FRET efficiency of the OF<sup>open</sup> conformation—it is only present in a single condition (0.1 mM ATP) and is not observed across other ATP concentrations for this TmrAB variant. It more likely reflects a minor non-reproducible subpopulation or photophysical artefact, in line with our response to Point 2 of the public review (Reviewer #2).

      (4) The final section of the Results section seems like an afterthought, especially since the heading suggests a broader scope.

      We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      References

      Asher WB, Geggier P, Holsey MD, Gilmore GT, Pa; AK, Meszaros J, Terry DS, Mathiasen S, Kaliszewski MJ, McCauley MD, Govindaraju A, Zhou Z, Harikumar KG, Jaqaman K, Miller LJ, Smith AW, Blanchard SC, Javitch JA (2021) Single-molecule FRET imaging of GPCR dimers in living cells. Nat Methods 18: 397–405. doi:10.1038/s41592-021-01081-y

      Asher WB, Terry DS, Gregorio GGA, Kahsai AW, Borgia A, Xie B, Modak A, Zhu Y, Jang W, Govindaraju A, Huang LY, Inoue A, Lambert NA, Gurevich VV, Shi L, Lefkowitz RJ, Blanchard SC, Javitch JA (2022) GPCR-mediated beta-arrestin activation deconvoluted with single-molecule precision. Cell 185: 1661– 1675 e1616. doi:10.1016/j.cell.2022.03.042

      Barth K, Hank S, Spindler PE, Prisner TF, Tampé R, Joseph B (2018) Conformational coupling and transinhibition in the human antigen transporter ortholog TmrAB resolved with dipolar EPR spectroscopy. J Am Chem Soc 140: 4527–4533. doi:10.1021/jacs.7b12409

      Barth K, Rudolph M, Diederichs T, Prisner TF, Tampé R, Joseph B (2020) Thermodynamic basis for conformational coupling in an ATP-binding cassette exporter. J Phys Chem LeJ 11: 7946–7953. doi:10.1021/acs.jpclett.0c01876

      Hofmann S, Januliene D, Mehdipour AR, Thomas C, Stefan E, Brüchert S, Kuhn BT, Geertsma ER, Hummer G, Tampé R, Moeller A (2019) Conformation space of a heterodimeric ABC exporter under turnover conditions. Nature 571: 580–583. doi:10.1038/s41586-019-1391-0

      Husada F, Bountra K, Tassis K, de Boer M, Romano M, Rebuffat S, Beis K, Cordes T (2018) Conformational dynamics of the ABC transporter McjD seen by single-molecule FRET. EMBO J 37: e100056. doi:10.15252/embj.2018100056

      Levring J, Terry DS, Kilic Z, Fitzgerald G, Blanchard SC, Chen J (2023) CFTR function, pathology and pharmacology at single-molecule resolution. Nature 616: 606–614. doi:10.1038/s41586-023-05854-7

      Nocker C, Pečak M, Nocker T, Fahim A, Sušac L, Tampé R (2026) Single-molecule dynamics reveal ATP binding alone powers substrate translocation by an ABC transporter. Nat Commun 17 doi:10.1038/s41467-026-70021-1

      Nöll A, Thomas C, Herbring V, Zollmann T, Barth K, Mehdipour AR, Tomasiak TM, Bruchert S, Joseph B, Abele R, Olieric V, Wang M, Diederichs K, Hummer G, Stroud RM, Pos KM, Tampé R (2017) Crystal structure and mechanistic basis of a functional homolog of the antigen transporter TAP. Proc Natl Acad Sci U S A 114: E438–E447. doi:10.1073/pnas.1620009114

      Stefan E, Hofmann S, Tampé R (2020) A single power stroke by ATP binding drives substrate translocation in a heterodimeric ABC transporter. eLife 9: e55943. doi:10.7554/eLife.55943

      Wang L, Johnson ZL, Wasserman MR, Levring J, Chen J, Liu S (2020) Characterization of the kinetic cycle of an ABC transporter by single-molecule and cryo-EM analyses. eLife 9: e56451. doi:10.7554/eLife.56451

    1. eLife Assessment

      The authors present important evidence for a WIPI2-Retriever complex (termed CROP2) that couples cargo selection to carrier fission at endosomes. CROP2 appears to function analogously to the previously described CROP1 complex, formed by WIPI1 and Retromer, with which it shares structural similarities. They provide compelling evidence that CROP1 and CROP2 regulate the trafficking of distinct subsets of cargoes; however, the cellular evidence for the existence of these distinct complexes is mostly inferred from immunoprecipitation analysis and would benefit from further validation.

    2. Reviewer #1 (Public review):

      WIPI1 is a PROPPIN family protein that has been implicated in Retromer-mediated membrane fission events. Although the cargos that it has been tested to be important for are diverse, one of the cargos that is unaffected is Beta1-Integrin. This leads the authors to assess another PROPPIN family protein - WIPI2, which is a homolog of WIPI1. KD using siRNA is effective and had no consequences on LAMP1, EGFR trafficking or GLUT1 trafficking. Integrin-B1, however, had a large and significant defect in its recycling from the endosome, with a clear endosomal colocalisation. Complementation experiments with WT WIPI2 recovered the phenotype, but various mutant WIPI2 complements resulted in elongated tubules, and there was also a dominant negative effect of the mutant. Integrin is a classic retriever cargo, so the authors rationalise that WIPI2 may be playing a role with retriever that WIPI1 plays with retromer. To assess this, they perform a set of immunoprecipitations. SNX17, the retriever-associated sorting nexin, co-IPs with WIPI2 in a VPS26C-dependent manner. VPS26C but not VPS26 co-IPs with WIPI2, and the reciprocal with WIPI1. These interactions were not present for the FSSS mutation of WIPI2. WIPI2 localises to Rab11 endosomes mainly, as does retriever. Mutations of WIPI2 not only affected WIPI2 localisation, but also VPS35L mutations, indicating that there is a functional relationship between the two.

      Comments on revised version.

      The reviewers have responded appropriately to all the points. I have no remaining concerns.

    3. Reviewer #3 (Public review):

      Summary:

      The manuscript of Mayer and colleagues analyzes the function of WIPI proteins in mammalian cells. The authors identified previously CROP as a complex consisting of WIPI1 and the retromer complex, primarily in yeast cells. In mammalian cells, both WIPI1 and WIPI2 exist, whereas retromer has a homologous complex termed retriever. The now find that WIPI2 can form a complex with retriever subunits. They name this complex CROP2. Their data further indicate that CROP2 and CROP1 have distinct substrate specificities as knock down of CROP2 subunits affect beta1 integrin sorting, whereas knock down of CROP1 affects EGFR and GLUT1. The further identify a similar sequence (FSSS) in both WIPI1 and WIPI2, which is required for their specific binding to retromer and retriever.

      Strengths:

      CROP1 and CROP2 seem to use similar features for their formation, and have different substrates, which is convincingly shown.

      Weaknesses:

      The analysis lacks information that this is a complex as claimed. It can be deduced from the immunoprecipitation analysis.

      Comments on revised version.

      The authors answered my questions and adjusted the text accordingly. Figure 10 was not part of the submitted version. It should be checked by the editor.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      WIPI1 is a PROPPIN family protein that has been implicated in Retromer-mediated membrane fission events. Although the cargos that it has been tested to be important for are diverse, one of the cargos that is unaffected is Beta1-Integrin. This leads the authors to assess another PROPPIN family protein - WIPI2, which is a homolog of WIPI1. KD using siRNA is effective and had no consequences on LAMP1, EGFR trafficking or GLUT1 trafficking. Integrin-B1, however, had a large and significant defect in its recycling from the endosome, with a clear endosomal colocalisation. Complementation experiments with WT WIPI2 recovered the phenotype, but various mutant WIPI2 complements resulted in elongated tubules, and there was also a dominant negative effect of the mutant. Integrin is a classic retreiver cargo, so the authors rationalise that WIPI2 may be playing a role with retreiver that WIPI1 plays with retromer. To assess this, they perform a set of immunoprecipitations. SNX17, the retreiver-associated sorting nexin, co-IPs with WIPI2 in a VPS26C-dependent manner. VPS26C but not VPS26 co-IPs with WIPI2, and the reciprocal with WIPI1. These interactions were not present for the FSSS mutation of WIPI2. WIPI2 localises to Rab11 endosomes mainly, as does retriever. Mutations of WIPI2 not only affected WIPI2 localisation, but also VPS35L mutations, indicating that there is a functional relationship between the two.

      On the whole, I find the manuscript compelling. The manuscript is very clearly written, the results are convincing and well performed. The flow of experiments is logical, and although not comprehensive in the subsequent mechanistic understanding, the fundamental findings are important and convincing. My comments below are, on the whole, minor and are intended to support the communication of the findings to the field.

      We are happy that the reviewer has received our work quite positively.

      (1) The IP interaction data were convincing; however, for me and some others, an interaction is only convincing when performed in vitro, and understood at a structural level. I do not suggest the authors do that in this case; however, I think, at a minimum, some sensible moderation of claims would be useful here.

      Indeed, quantitative in vitro data on the affinities would be a nice addition. However, we have significant trouble to recombinantly express and purify well-behaved WIPI2 in sufficient quantities for such studies. We keep working in this direction but are not there yet.

      We have now inserted a phrase into the discussion section highlighting this limitation: "Our immunoprecipitation assays cannot distinguish and more detailed structural and interaction studies with pure compounds will be necessary to elucidate the nature of this interaction". We nevertheless think that the the isoform specificity of the IPs, the effect of the point mutations in WIPI2 on these interactions, and the functional effects in vivo lend signficant support to the notion of a complex even if there is no proof of direct binding of WIPI2 to Retriever.

      (2) I found the final localisation data and its interpretation confusing. My interpretation of that data would not be that the retreiver is relocalised, but rather that there is less of both recruited to the membrane and the remaining localisation distribution is shifted. In addition, I am not quite sure of the model here - is the idea that WIPI2 recruits retreiver, if that is the case, I find it hard to resolve with its role as a mediator of fission. Clarity would be appreciated here.

      We are not quite sure what "final" localisation data the reviewer refers to, but we guess it is Fig. 9. This figure primarily provides in vivo evidence supporting the connection between Retriever and WIPI2. It does this by showing that the S67 substitution shifts both proteins. In WIPI2 wildtype cells, WIPI2 and VPS35L strongly colocalize in Rab11 compartments. S67 substitutions in WIPI2 abolish this localisation; WIPI2 shifts mainly to Rab5 compartments, where VPS35L shows only a moderate increase, and to Rab7 compartments, where VPS35L shows no increase at all.

      We do not understand the reviewer's interpretation that less Retriever would be recruited to the membranes in the S67 variants. VPS35L remains completely associated with punctate, presumably membrane-bounded structures also in the mutants, providing no evidence for a detachment from the membrane. The same is observed in a WIPI2 knockdown. Therefore, we did not claim that WIPI2 is the main factor recruiting Retriever to the membrane, for which our experiments yield no hints. This does not exclude that the interaction of WIPI2 could strengthen membrane recruitment, or that two pools of Retriever exist, one interacting with Snx17 and another interacting with WIPI2, and that both link to each other in a coat. We did not dwell on this in the discussion because our experiments cannot distinguish these possibilities and were not conceived to analyse membrane recruitment of Retriever.

      (3) I am concerned that the repeats being compared for statistical analysis are not biological repeats but technical repeats (cells in the same experiment). I should think the idea of the statistical comparison is to show experimental reproducibility and variability across biological repeats. Therefore, I would expect an appropriate number of biological repeats (3 or more minimum), to be the data compared in the statistical analysis and graphs. I think it is appropriate to average the technical repeats from each biological repeat. I find these to be useful resources https://doi.org/10.1083/jcb.202401074, https://doi.org/10.1083/jcb.200611141

      The repeats being compared are biological repeats from independent experiments. This is described in Methods, where the reviewer may not have seen it. In order to make the independent experiments more evident in the figures, we have now colour coded the individual cell measurements from the three independent experiments. This allows to visualize both the individual data points, the average from each experiment and the variability across the independent experiments.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from De Leo and Mayer presents evidence that the PROPPIN protein, WIPI2, associates with the Retriever complex, and is required for the proper transport of the SNX17-Retriever cargo, beta1-integrin. This finding fits with prior papers from the Mayer lab, which showed that a related PROPPIN, WIPI1, is required for the transport of some SNX27-Retromer cargo, including GLUT1. The retromer and retriever complexes are architecturally similar. Importantly, they act at the same endosomes, and each transports cargo from endosomes to the plasma membrane. Thus, the possibility that each also requires a structurally related PROPPIN is of interest. However, the manuscript is incomplete, and the main claims are only partially supported.

      Strengths:

      The topic that PROPPIN proteins are important for the function of the Retromer and Retriever complexes expands our view of the trafficking complex.

      Weaknesses:

      Many important controls are missing. Several points that are made in the manuscript are only supported through a single approach.

      We made a serious effort and implemented many suggestions of this reviewer, but orthogonal approaches are not always available or accessible.

      Reviewer #3 (Public review):

      Summary:

      The manuscript of Mayer and colleagues analyzes the function of WIPI proteins in mammalian cells. The authors previously identified CROP as a complex consisting of WIPI1 and the retromer complex, primarily in yeast cells. In mammalian cells, both WIPI1 and WIPI2 exist, whereas retromer has a homologous complex termed retriever. They now find that WIPI2 can form a complex with retriever subunits. They named this complex CROP2. Their data further indicate that CROP2 and CROP1 have distinct substrate specificities as knockdown of CROP2 subunits affects beta1 integrin sorting, whereas knockdown of CROP1 affects EGFR and GLUT1. They further identify a similar sequence (FSSS) in both WIPI1 and WIPI2, which is required for their specific binding to retromer and retriever.

      Strengths:

      CROP1 and CROP2 seem to use similar features for their formation, and have different substrates, which is convincingly shown.

      Weaknesses:

      The analysis lacks information that this is a complex as claimed. It can be deduced from the interaction analysis, but was not shown.

      It is of course desirable to obtain a detailed structural and in vitro characterisation of this interaction, which we have not provided because we currently do not have sufficient amounts of well-behaved source material for this. We nevertheless think that the interaction we show, which is strictly isoform-specific and dependent on single amino acid substitutions in a motif that in CROP1 is necessary for the interaction its recombinant subunits, supports that CROP2 is a similar a complex. We don't show a direct interaction but also don't claim in the manuscript that the interaction between WIPI2 and Retriever is direct and independent of additional factors.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, the reviewers generally value the contribution to the field, but they feel that some claims require additional experimental support.

      (1) I have summarized the major points below.

      (a) Both reviewers 1 and 2 agree that the quality of localization data presented in Figure 9 and S5-S7, and the interpretation of the data, could be improved. See comment 2 from reviewer 1 and comments 23, 24 and 25 from reviewer 2. They not only suggest ways to improve the presentation of the data, but additionally suggest improving the staining of the Rab11 marker and additionally explain the lack of co-localization between VPS35 and Rab5, which has been reported in the literature.

      This impression was due to the fact that some figures showed projections of image stacks, which was not indicated clearly in the figure legend. We have changed this and now show single image planes throughout all figures.

      (b) Both reviewers 1 and 3 note that the evidence supporting a functional WIPI2-Retriever complex in vivo is currently weak. We agree that additional biochemical data demonstrating the presence of the CROP1 and CROP2 complexes in vivo would strengthen the central message of the paper and elevate it to a more fundamental discovery.

      We understood that the reviewers did not ask for further in vivo evidence but would welcome structural characterisation of the complex and quantitative binding data in vitro with purified proteins. Structural characterisation is out of scope of our study and in vitro binding studies have remained hampered by the fact that WIPI2 is hard to express and purify and not well behaved in vitro.

      (c) All reviewers agree that the authors should carefully repeat their statistical analysis to account for the number of biological replicates. Reviewer 1 suggests publications that the authors could refer to.

      The reviewers have probably overlooked the respective description in the methods section, where it had been stated that we analysed biological replicates from independent experiments. In graphs showing measurements from individual cells we now make this evident through colour coded dots, in which each colour represents data points stemming from an independent experiment. This makes it evident that the variance from experiment to experiment is low. The means (n = 3) were generally compared using a two-tailed unpaired t-test.

      (d) Reviewer 2 additionally has various minor points that would greatly improve the readability and presentation of the work, and we recommend addressing (comments 1, 2, 3, 4, 12, 15, 17, 20, 27, 28, 29). All reviewers, in general, provide great minor suggestions. It would be great if the CROP1 and 2 complexes could be clearly introduced in each figure. We also agree that the WIPI2 CT labelling is confused and should be changed to "control" or similar.

      Many of the points raised by this reviewer were actually quite minor or questions of personal preference, not major problems as stated in the review. Nevertheless, we found a number of useful suggestions in this review and have addressed these points as detailed in the response to reviewer 2.

      (2) In addition to the major shared concerns laid out in the points above, reviewer 2 has some further minor suggestions:

      (a) Comment 6. Could the author explain the discrepancies between the example blot shown in Figure 1D and the quantification (1E).

      The two have actually been quite consistent. The reviewer might have mistaken the marker lane as the 0 min reference value to arrive at this impression. We have now removed the marker lane to avoid this.

      (b) Comment 9 - could the authors clarify how surface labelling experiments were carried out?

      This had been clearly described in the methods section, where this reviewer has probably not seen it.

      (c) Comment 11 - The reviewer suggests normalizing the surface levels of markers to the cell area and not per cell. This is a reasonable suggestion.

      The analysis had already been performed as proposed. This had been clearly described in the methods section, which the reviewer may not have looked at.

      (d) Comment 19 "In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures." The authors could try a Rab4 or Rab11 overexpression plasmid to show whether these are elongated recycling tubules.

      This has now been added.

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The figures are not colourblind friendly, and should be changed to be so. Additionally, single colour images should be grayscale.

      That was a good learning opportunity. We adapted the colour schemes of the images to make them more colourblind friendly, now using magenta, green, and white for the overlaps. In doing so we have relied on published recommendations, but we have not found a colourblind colleague to check the efficacy of this change.

      (2) WIPI2^CT labels are confusing, as people may think they are a mutant. I suggest changing to "control" or similar.

      These have been changed.

      (3) "The effect was comparable to that of a knockdown of SNX17 (Figure 3 A, B)." On page 6. Based on this sentence, I was expecting to see a comparison to SNX17 KD, but it was not there as far as I can tell.

      This statement referred to a publication by P.Cullen and collaborators. We have changed the wording and inserted the (missing) reference to make this clear.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is modest. In addition, many of the claims should be better supported by the addition of orthogonal data. Moreover, the quality of some of the data presented needs to be improved. Overall, the manuscript requires better descriptions of the methods. In many figures, it was not clear how the experiments were performed.

      The experimental descriptions that the reviewer refers to had been provided in the Methods section, where this reviewer may have overlooked them.

      The paper should also be better organized. Some less important findings are in the main figures, whereas some critical results are in the supplemental figures. In addition, there were multiple issues with the readability of the paper, and the authors should consider using a professional editor to make the paper easier to read.

      We had given the paper to colleagues who found it clear, and also Reviewer 1 has underlined its clarity. Nevertheless, we have re-phrased the manuscript in some parts to optimise it.

      One of the main claims in the paper is that the FSSS motif of WIPI2, as well as a conserved amphipathic helix, is critical for WIPI2 function in the CROP2 complex. It is notable that these are the same regions that are also critical for the role of WIPI2 in autophagy (Gubas et al., 2024 PMID: 39152217). The authors should include this information in the manuscript and cite the paper.

      Indeed. We mention this now in the introduction of the revised version.

      Additional Major Issues:

      While some of the issues raised below are actually minor and/or matters of personal preference, several comments led us to improve and correct the figures and we thank this reviewer for the constructive suggestions.

      (1) In Figure 1, it appears from the representative images that WIPI2 KD cells have higher levels of EGFR (Figure 1A and 1B). Is this correct?

      To some degree. This increase is not systematic. A moderate increase has been observed only in 2 experiments out of 4. Therefore, we did not investigate this.

      (2) Also in Figure 1, the colocalization is difficult to see. The authors should add the separate channels in addition to the merged images. Since the point is supposed to be that there is no impact on EGFR, all of this data could go into the supplement.

      We had considered this already for the original version but dismissed the idea. The overlap is quantified in Fig. 1C, which provides the relevant values from four experiments. Fig. 1A/B provide only sample pictures, which also permit to see overlap (yellow) 0 and 5 min after the induction of degradation, which vanishes at later timepoints. Separating the channels would quadruple the space that this figure occupies, which would not be practical and not change the point to be made.

      (3) The scale bars for each panel differ from each other. To better assess the data, the exact same magnification should be shown for each panel.

      Corrected

      (4) Figure 1C is confusing. The authors should explain which lines correspond to EEA1 and LAMP1.

      Corrected

      (5) In Figure 1D, the authors show different blots for control and WIPI2 KD. Could the authors compare WIPI2 and EGFR in the same blot? Without a comparison on the same blot, it is impossible to know whether the starting levels of EGFR are the same. Moreover, the quantitation in Figure 1E sets the value for each cell line to 100%. Instead, the starting levels in each cell line should be compared. The authors should use the amount of EGFR at zero time in the control cells to define 100%, and then indicate the relative initial EGFR levels in the WIPI2KD cells.

      A new blot is shown now and the quantification has been performed as proposed.

      (6) The quantification in Figure 1E does not match the representative blot shown in Figure 1D. According to the graph, the rate of degradation of EGFR is similar in both cell lines. But the representative blot shows that there are large differences.

      We do not understand this comment. The representative blot shows similar kinetics for both. Perhaps the reviewer got confused by the fact that a marker lane was still present on the left blot and not labelled as such. The new version of the figure corrects this.

      (7) The blot showing the WIP2 knockdown in Figure 1D has a lot of background. However, the blot of the WIPI2 knockdown in Figure S1 looks very good. The authors should make sure that they load enough sample and use a good antibody for the experiments in Figure 1.

      The new blot that we added in response to comment 5 corrects this.

      (8) In Figure 2 and Figure 3A, the cells are too confluent. This is an issue because the cells might not be metabolically active. In addition, the signal is saturated. The authors should make sure that all of the data is collected on cells that are not too confluent.

      The confluency of the culture cannot be judged from single frames, which were selected to show several cells. We had controlled confluency and underlined in the Methods section that “For microscopy, the cells were plated on 18-mm-diameter glass coverslips on 24-well plates and grown for 2 or 3 days according to the protocol of DNA or siRNA transfection by reaching a confluency of 70-80%”. The reviewer may not have seen this.

      (9) One main issue with these figures, especially the non-permeablized cells, is that it is impossible to assess how much of the signal is on the cell surface. The authors should provide the methods that they used to prevent inadvertent permeabilization of the cells. Were these experiments performed at 4 degrees? The authors should include a control of an antibody to a protein that is not found on the cell surface.

      There is an internal control in that the non-permeabilised WIPI2KD cells, which have been treated with the same antibody, show no much less staining than the control cells (Fig. 3A). In WIPI2KD cells, integrin becomes accessible for antibody staining only upon detergent permeabilization. This demonstrates that our procedure does not lead to significant inadvertent permeabilization of the cells.

      (10) The authors should perform surface biotinylation assays as an orthogonal approach to determine GLUT1 levels and beta1-integrin levels at the cell surface, respectively.

      There is a strong, qualitative difference in the surface labelling of beta1-integrin that is not observed for GLUT1. Given that, it is not obvious to us what additional argument would be provided by surface biotinylation or subfractionation experiments.

      (11) In quantifying surface levels of GLUT1 or beta1-integrin by microscopy, the authors should normalize to the cell area, rather than per cell.

      The reviewer has probably not seen that the Methods section states that the cell area has been used for normalisation.

      (12) In Figure 3, the nuclear DAPI stain in the KD cells is much less bright than in the control cells. The authors should make sure to choose representative images.

      The nuclear DAPI signal has been visible in all cells. Depending on the position of the nucleus, is shape and dimension in the z-direction, individual nuclei can show different degrees of staining. The images shown are representative. We have adjusted the settings now to make the nuclei in the WIPI2KD cells easier to spot.

      (13) For the immunofluorescence studies, the authors should be using single z planes rather than maximum projection.

      Images have been exchanged by single planes.

      (14) For the experiments in Figure 3, the authors should check the total levels of EEA1 and LAMP1 by western blot to test whether WIPI2 KD affects the levels of these proteins. If these organelle marker proteins are impacted, this could impact the colocalization measurements shown in Figures 3C and D.

      We have measured the total fluorescence intensity of EEA1 and LAMP1 in the images. It shows no significant difference between control and WIPI2 knockdown cells (new Fig. 3F, H).

      (15) In Figure 4A, the helical representation is rotated in the WIPI2-Sloop; the orientation of the residues that are not mutated should stay the same.

      Yes. Done.

      (16) In Figure 4B and 4C, cells that were not transfected with WIPI2 WT or WIPI2 Sloop should be shown.

      Since the transfection efficiency is limited, the fields contain both non-transfected (lacking green fluorescence) and transfected cells (showing green fluorescence). We have now marked transfected cells with an asterisk.

      (17) The cells in the lower panel of 4B have an unusual morphology and are much more round. The authors should choose cells that are representative of each experimental condition.

      We now provide another field.

      (18) In Figure 4C, it looks like the magnification of the top panels is different from the bottom panels. The same magnification for all the panels should be shown (and the size of the scale bars should be the same.

      Corrected

      (19) In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures.

      We have done this (new Fig. S4). Rab4 is on tubules. Rab5 on the structures from which the tubules emanate.

      (20) In Figure 5A, the top scale bar is missing.

      Corrected.

      (21) In Figure 5B, the confluency is too high.

      See our response above. A single field does not permit to judge this. Confluency was controlled for all cultures. The cultures were not confluent.

      (22) The IP studies shown in Figures 6, 7 and 8, should be accompanied by colocalization studies.

      Colocalization measurments have now been integrated into the manuscript (Figs. S5, S6). They are consistent with the IP data.

      (23) Figure 9 was very confusing and should be broken up into multiple figures. Data showing that localization did not change in any of the cell lines can be put in figures that are distinct from figures that show that localization changed in the various mutants. Figures that show no change can go in the supplement.

      Since every panel of Fig. 9 shows a statistically significant difference we left the figure unchanged.

      (23) Representative figures should be shown in the same figure as the corresponding graph. In addition, the order of the colocalization data shown in the graphs and figures should match the order described in the text.

      We consider the graphs of Fig. 9 as the relevant information. Representative images are just illustration. Integrating them with the graphs would make it necessary to split everything up into multiple figures, making it harder to compare the different combinations. Therefore, we left the figures unchanged.

      (24) In Figure S7, the Rab11 signal looks continuous, which makes the colocalization analysis meaningless. The authors should determine how to take images that can be evaluated. On a more minor note, the zoomed panels should be labeled as well.

      This is a result of having shown a projections of multiple planes. The images have now been replaced by single plane images. Zoomed panels have been labelled and the scale bar added.

      (25) The low colocalization of VPS35L with Rab5 is surprising, as SNX17 has been previously shown to co-localize with early endosomes positive for EEA1. This result may have occurred due to overexpression because the authors chose to utilize plasmids that express a tagged protein. There are antibodies to each of the endogenous proteins, and this is what should be used for this set of experiments.

      This comment made us control the analysis performed for these images, which by mistake had been performed on z-projections rather than on single planes. This distorted the values. The re-analysed data shows a higher colocalisation with Rab5, but it remains inferior to colocalisation with Rab11.

      (26) The authors should determine whether β1-integrin colocalizes with WIPI2 in endosomal compartments.

      This was done. WIPI2 colocalizes with beta-integrin on EEA1-and SNX17-positive strcutures but not positive for LAMP1 (Fig. 3E/F).

      Minor points

      (27) In one of the panels in Figure 1A, "30 min" is duplicated.

      Removed

      (28) In Figures 5C and 5D, the y-axis should indicate that this is surface β1integrin.

      Changed and added “surface”

      (29) In Figure 9 there is a typo in panel A. It is VPS35L and not VPS35.

      Corrected

      Reviewer #3 (Recommendations for the authors):

      This is an overall convincing study, which shows that the two complexes, CROP1 and CROP2 function at different membranes and serve different substrates. While I agree with their localization analysis, I have one key issue. The authors claim that each of the two forms a complex and base this on their specific pull-down and western blot analyses.

      I find it important that they show that both indeed form stable complexes in vivo, using pull-down and mass spectrometry approaches. They have all the necessary tools in hand and could use WIPI1 and WIPI2 to demonstrate the existence of the two complexes. The FSSS mutants of each are good controls for such an analysis.

      The manuscript actually presents the demanded in vivo experiments. Figs. 6 to 8 show pull-downs of WIPI1 and WIPI2 from cells, including also the FSSS mutant. While we haven't analysed this interaction by mass spectrometry, the Western blot analysis confirms the analysis. Cooperation of these proteins is further supported by the in vivo phenotypes, where the S67A substitution in WIPI2 produces a similar phenotype on integrin beta1 localisation as inactivation of Retriever.

      A second aspect is the general presentation. The paper would be a lot more accessible if the subunits of each complex (CROP1 and CROP2) were also introduced in the figures of each part. For readers, a final model is helpful to put the data into context and show where each complex operates in the cell.

      We have introduced a scheme of the respective complexes, including the names of the compunds, in Figs. 6 and 7 to avoid confusion.

      Finally, it is not clear how the statistics compare to repeats in their data. This should be clarified.

      This had been described in methods. Statistics has always been done on biological replicates stemming from independent experiments. We have added a cartoon (Fig. 10) depicting the trafficking pathways affected by CROP1 and CROP2.

    1. eLife Assessment

      This paper reports the findings of a neuroimaging experiment that tested the hypothesis that the cortex, specifically early visual areas, reinstates certain content from past episodic events. This is a useful study that highlights the role of early sensory cortices in supporting rapid, one-shot learning of location information for long-term memory. The strength of the evidence is solid, with the methods, data, and analyses broadly supporting the claims.

    2. Reviewer #1 (Public review):

      Summary:

      This paper reports the findings of a neuroimaging experiment that tested the hypothesis that the cortex, specifically early visual areas, reinstates the content from single events during our lives. The researchers tested this hypothesis by presenting to-be-remembered pictures of objects at spatial locations on the computer screen and then testing subjects with both recall and recognition. They show that during memory testing, the spatial location of the object can be decoded from the pattern of cortical BOLD responses measured with fMRI. They go on to show that the spatial tuning is higher during recognition than recall, that the tuning is correlated with memory retrieval accuracy, and that the retrieved precision is predicted by the encoded precision, particularly in the higher-level visual areas. Thus, the paper finds evidence of cortical reinstatement of details from a single event in a human life.

      Strengths:

      This is a strong manuscript that I have had the luxury of commenting on during a round of review at another prestigious journal. As a result, the authors have already made changes to address previous comments about highlighting the complementary learning systems approach more to motivate the alternative prediction that the cortex should only show evidence of reinstatement after repeated presentations. In addition, the authors have fleshed out the discussion of working memory in this task. They also revised their review of the literature to include citations suggesting spatial locations are normal parts of our episodic representations, likely obligatory in nature, as my group and others have argued in completely unrelated work. I applaud the authors for being responsive to a previous round of review and using the comments to address relatively minor issues with the paper, even though they moved on to a different journal. Thus, I found the paper even stronger than at first approach, and at first blush, the results were intriguing and the paper well written.

      Weaknesses:

      There is a logical perspective in the narrative that seems to unnecessarily weaken the paper. The paper shows evidence consistent with the conclusion that mnemonic representations are contained in early visual cortex, but then argues that those representations are not actually stored therein. For example, the first half of the last sentence of the conclusions (see page 19 of the manuscript). I understand the perspective that subcortical mechanisms must be involved in the act of retrieval, given the neuropsychology and other evidence. But if storage is elsewhere with the same fidelity so as to code this information, then how would such a memory system work? The MTL neurons would need to have the real, precise representation of all the orientations encoded at all the retinotopic locations, a mirror to V1 in terms of precision, because that's the actual memory representation being retrieved, so its fidelity will be limited by what is stored in the file, so to speak. Then, at retrieval, the paper proposes that the brain just reactivates the encoding context in V1 to help with the response output and ensure the precision of the behavioral responses. This must mean that the hippocampus/MTL has cells and networks with tuning functions that match the precision in all the cortical sensory systems that they are integrating context across, given the episodic memory models like Polyn and colleagues (2009, Psych Rev). So, there are little MTL maps that are completely redundant with V1, M1, A1, S1, etc.? Why such redundancy?

      Why not propose that what the subcortical systems do is to encode a unique pattern for that episode, that is separated from others, that just links (or provides pointers to, in computer science jargon) the contextual details stored in the cortical networks themselves? In this way, we can explain why neglected patients also neglect their memories of the town square. This has always been my interpretation of the results of the Polyn et al. (2006, Science) paper and the models tested with those whole-brain results. That is, you see widespread cortical context reinstatement during (one-shot) free recall events that included visual selective cortex for faces when faces were being recalled, but included a broad network, probably V1, and activating sounds in A1, body posture in M1, etc., though the latter three examples did not discriminate between categories of memoranda, in their experiments. Given that you show that activity in V1 during retrieval looks like it is being used, you should propose that the early cortex really participates in memory storage functions. V1 neurons are wired up to neurons of other selectivities in a competitive network with plastic synaptic connections. How would experience be prevented from changing activity in the cortex? Yes, cortical changes slow after the critical periods, as studied in the classic eye suturing experiments to study ocular dominance, but changes in cortical representations do not stop with maturity, with the pinwheel centers looking like they are context sensitive, thus, changing rapidly to events across time (Okamoto, Ikezoe, et al., 2011, Sci Reports). The brain would need a no-plasticity mechanism, and instead, it looks like the cortex can completely rewire even in adulthood (Buonomano & Merzenich, 1998, Annu Rev Neuro).

      I believe that the paper needs to describe the strong/radical interpretation of the current findings; that they are consistent with the view that the entire brain may be a memory structure, with encoding linking representations across sensory cortices. But also activating semantic and lexical systems, emotional networks encoding those aspects of context which we know can sometimes strongly drive effects, a nice prediction that could be made in the discussion/conclusions. Here you are looking at how precise the visual reinstatement is in V1 during retrieval following one exposure. One parsimonious mechanism to explain this effect is that the brain stores details of events using the neurons that do the high-fidelity perception of the event. Given that our goal is to stimulate thinking among fellow scientists so that this paper can be a citation classic, I think the paper should be revised so that it paints a complete picture of the theoretical possibilities of its findings.

    3. Reviewer #2 (Public review):

      Summary:

      The study aims to show that the early visual cortex is not merely a sensory-perceptual region that encodes stimuli while they are physically present, but also supports the formation and retrieval of long-term episodic memories. Instead, the authors demonstrate that spatially tuned reactivation of early visual cortex after a single encoding event supports memory-guided behavior, such as recalling an object's original location.

      Strengths:

      The study provides solid evidence that location information for single, trial-unique objects is reinstated in early visual cortex during both recognition and recall, even without explicit spatial demands, and the remembered vs. forgotten analyses link spatial tuning to behavior. The one-shot design and absence of explicit spatial instructions are important strengths that bring the paradigm closer to everyday, incidental episodic experiences and go beyond highly trained cue-target associations.

      Weaknesses:

      (1) Conceptually, the main findings would appear less surprising without a sharper theoretical contrast. Given basic retinotopic coding, it is natural that object identity and location are jointly encoded when an object is presented at a particular position, so spatially tuned reinstatement in V1-V3 can be interpreted as a reconfirmation of known properties unless more clearly contrasted with theories that emphasize more abstract, position-invariant cortical representations following hippocampal-cortical recoding. As currently framed, the introduction does not fully articulate what existing accounts might predict, or what pattern of results would have challenged those accounts, which somewhat weakens the perceived theoretical payoff.

      (2) It also remains somewhat unclear why early visual cortex (V1-V3), specifically, is the critical locus for the spatial information of interest, as opposed to higher-level visual or parietal regions that could also provide a spatial scaffold; clearer rationale and, if possible, control analyses in additional regions would help here.

      (3) Since gaze behavior is central to any spatial account, it would be helpful to report basic eye-tracking analyses comparing remembered versus forgotten trials, especially at encoding, to rule out systematic differences in fixation patterns that could contribute to the spatial tuning results.

    4. Reviewer #3 (Public review):

      Summary and Overall Evaluation:

      This is an elegant paper addressing an important question: whether spatial location is automatically activated during the recall of object memories. Building on prior work that relied on trained or repeated stimuli, the present study uses unique objects with one-time encoding across four spatial locations - a meaningful advance in ecological validity. The experimental design is clean, the data analysis is well-executed, and the reported effects, while small, are intriguing and open up interesting questions about the role of spatial structure in visual memory. Overall, this is a solid contribution, and my comments below are intended to help the authors strengthen the paper further.

      Major Comments

      (1) Incidental encoding.<br /> Was the memory task fully incidental - that is, were participants unaware that a subsequent memory test would follow encoding? This seems important for interpreting the automaticity claim that is central to the paper's contribution, and should be clarified explicitly.

      (2) Spatial extent of the analysis - higher visual regions and negative pRFs.<br /> The analysis appears restricted to regions V1-V3. Have the authors examined higher visual areas as well? This seems like an important omission given that object memory likely engages regions well beyond the early visual cortex. Relatedly, recent work by Adam Steel and colleagues suggests that spatially tuned negative pRFs may play an important role in memory. Have the authors considered examining these? Expanding the analysis in these directions could substantially enrich the findings.

      (3) Mechanism - retinotopic or spatiotopic?<br /> The paper makes a compelling case that spatial structure supports memory, but the nature of that spatial structure deserves more discussion. Are the effects retinotopic or spatiotopic in nature? The current design may not be able to fully dissociate these possibilities, but this distinction is theoretically important, and the authors should engage with it directly. Even a careful discussion of what the current data can and cannot tell us on this point would be valuable.

      (4) Relationship between encoding failure and retrieval failure.<br /> For trials where memory performance is worse, and the encoding models fail, is there a systematic relationship between how the pRFs fail at object retrieval versus spatial retrieval? In other words, are the pRFs wrongly tuned in the same way at both stages? This analysis could provide meaningful insight into whether object and location retrieval draw on shared spatial representations.

      (5) Object shape and spatial mapping.<br /> Real-world objects vary considerably in surface structure and shape, which may affect how cleanly they map onto a specific spatial location. Was this considered in the analysis? What was taken as the correct or peak location for each object, and how was this defined when objects extended across space? Apologies if this was addressed in the methods and I missed it.

      (6) Time course of pRF activation.<br /> Is there a way to examine the time course of pRF activation within a trial? Do the spatially tuned responses arise immediately upon retrieval, or do they build up over time? Even a preliminary analysis of this would be of considerable theoretical interest, as it would speak to whether spatial reinstatement is an early automatic process or a later, more deliberate one.

      (7) Effect size and functional significance.<br /> The authors acknowledge that the reported effects are very small, which I appreciate. However, this does raise genuine questions about functional significance that I think deserve a more direct response. One approach that would help contextualize the spatial effects would be to compare their magnitude to that of another feature - object identity, for example - to give readers a sense of the relative importance of spatial versus non-spatial information in memory representations. I recognize this may not be straightforward with the current design, but even a brief discussion of how one might benchmark the spatial effects would be helpful.

      (8) The attention account.<br /> I found the discussion of attention less than fully convincing. The authors appear to argue against an attentional interpretation of the spatial effects, but it is not clear why participants wouldn't attend to the encoded location during retrieval - particularly in a design with relatively few retrieval cues, where spatial location may be one of the most useful available. The attention account thus seems difficult to rule out on the basis of the current data, and the discussion should engage more seriously with this alternative rather than setting it aside.

      (9) Later-remembered versus later-forgotten objects - BOLD signal.<br /> Were later-remembered objects associated with stronger overall BOLD responses during encoding compared to later-forgotten objects, or was the effect specific to the pRF modelling? Clarifying this would help readers understand whether the spatial effects are part of a broader pattern of stronger encoding or something more specific to the spatial reinstatement mechanism.

    1. eLife Assessment

      This fundamental study provides convincing evidence that distinct molecular mechanisms underlie AAV-associated retinal toxicity in retinal pigment epithelial cells and photoreceptors, advancing our understanding of gene therapy-related retinal injury. The authors employ a rigorous and comprehensive experimental approach, including multiple knockout mouse models, transcriptomic analyses, and genetic loss-of-function studies, which substantially strengthen the mechanistic conclusions. Some concerns remain regarding vector characterization, the absence of procedural injection controls, and the limited interpretation of adult versus neonatal studies; nevertheless, the study makes a substantial contribution to the field and provides a strong foundation for future translational investigations.

    2. Reviewer #1 (Public review):

      This study examines the mechanisms underlying retinal toxicity associated with certain AAV gene therapy vectors, particularly in the retinal pigment epithelium (RPE) and photoreceptors following expression of transgenes such as GFP. The findings suggest that AAV-related retinal toxicity is driven less by transgene identity itself and more by distinct pathogenic mechanisms, including stress-induced injury in RPE cells and interferon-mediated damage in photoreceptors. The comments are as follows:

      (1) The AAV vectors were manufactured in-house, and the production method is described in sufficient detail. However, were any characterization assays performed beyond qPCR-based titer determination, such as vector genome titer, capsid titer, empty/full capsid ratio, sterility, bioburden, endotoxin, mycoplasma, residual host cell DNA, residual plasmid DNA, or residual host cell protein testing? These analyses, particularly those assessing residual impurities and microbial contamination, are critical, as such contaminants may provoke inflammatory responses following subretinal injection. This, in turn, could confound the interpretation of the results, including the identification of the molecular pathways contributing to toxicity as well as the specific role of GFP-associated toxicity. Please provide any characterization information for the AAV vectors.

      (2) The study uses contralateral or uninjected eyes as controls, but this choice may not adequately account for changes induced by the subretinal injection procedure itself. Because the earliest assessment of RPE toxicity was performed at 2 weeks post-injection, any injury, inflammation, retinal detachment-related stress, or wound-healing responses triggered by the surgical procedure could have contributed to the observed phenotype. As a result, comparisons to uninjected eyes alone make it difficult to distinguish vector or transgene-specific toxicity from procedure related effects. Inclusion of a more appropriate procedural control, such as sham-injected eyes or eyes injected with vehicle/buffer alone, would have strengthened the study by enabling clearer discrimination between injection-related retinal responses and toxicity attributable to the AAV construct or transgene expression.

      (3) The authors used phalloidin staining on RPE-choroid flatmounts to evaluate RPE toxicity, which provides useful information on RPE morphology and structural disruption. However, it would be highly informative to also assess the presence and distribution of subretinal microglia/macrophages, for example, by Iba1 immunostaining, in the same preparations. Specifically, determining whether Iba1-positive cells accumulate in or around areas of RPE dystrophy would help clarify the contribution of local inflammatory responses to the observed pathology. Such analysis could strengthen the interpretation of the toxicity phenotype by revealing whether RPE degeneration is accompanied by focal immune cell recruitment and whether these cells spatially associate with regions of tissue damage. This would also provide additional insight into whether inflammation is likely to be a downstream consequence of RPE injury or a more direct contributor to disease progression, especially in light of publications by Danial Saban's group regarding the characterization of microglia phenotypes using RNA-seq analysis.

      (4) The Discussion should also address the anatomical and procedural differences between neonatal and adult mouse eyes, particularly with respect to retinal thickness and the potential impact of subretinal injection-related injury. Because the RPE toxic effects appeared less severe in adult mice, it would be valuable for the authors to consider whether this difference reflects true age-dependent biological susceptibility or, at least in part, differences in the mechanical consequences of the injection procedure. Neonatal retinas are thinner and structurally less mature than adult retinas, which may render them more vulnerable to injection-associated stress, retinal detachment, or secondary tissue injury following subretinal delivery. In contrast, the greater retinal thickness and maturity of the adult eye may provide some degree of resilience to procedural trauma, thereby reducing the apparent severity of RPE damage. Expanding the Discussion to consider these factors would strengthen the interpretation of the age-related differences observed in toxicity and help distinguish vector- or transgene-driven effects from potential confounding effects introduced by the delivery method itself.

      Overall, this manuscript presents a detailed and comprehensive analysis of transgene-induced retinal toxicity and makes effective use of multiple mouse models to dissect the contribution of relevant molecular pathways. The study is particularly strengthened by its systematic approach, combining histologic, transcriptomic, and genetic loss-of-function strategies to distinguish the mechanisms underlying toxicity in the RPE versus photoreceptors. By evaluating several knockout mouse lines, the authors can move beyond descriptive observations and begin to assign causality to specific stress and immune signaling pathways, thereby providing important mechanistic insight into AAV-associated retinal injury. These findings are timely and relevant to the broader field of ocular gene therapy, as they highlight the complexity of vector- and transgene-related toxicity and underscore the need for careful pathway-level evaluation during preclinical development.

    3. Reviewer #2 (Public review):

      Summary:

      Adeno-associated viruses (AAVs) are popular gene therapy vectors, but AAVs can cause toxicity. This is particularly evident following expression of some transgenes, e.g., GFP, in the retinal pigment epithelium (RPE), which leads to loss of RPE cells and photoreceptors. Here, we sought to unravel the toxicity mechanism(s). Several transgenes, self and non-self, were tested for toxicity, with no clear correlation for this variable. RPE RNA-sequencing revealed upregulation of translational processes, cell stress, cytokine release, antiviral responses, and leukocyte infiltration pathways. Toxicity-inducing pathways were explored for causality by injecting toxic AAVs into mice deficient for intrinsic, innate, or adaptive immune pathways. The CHOP KO partially alleviated toxicity for RPE but not photoreceptors, whereas the type I interferon receptor KO partially alleviated toxicity for photoreceptors but not RPE. In situ hybridization of interferon pathway transcripts (IFNB1, IFNAR1) revealed that the RPE and retina can produce and potentially respond to interferon. These data suggest that transgene-induced cell stress responses in the RPE lead to RPE cell death, while interferon signaling contributes to the death of photoreceptors.

      Strengths:

      This manuscript used numerous KO mouse models to evaluate the interferon pathway, inflammatory cytokine pathways, the complement pathway, toll-like receptor signaling, cytosolic DNA sensing, double-stranded RNA sensing strain, intrinsic cellular stress pathways, as well as strains deficient for B cells and T cells or B cells, T cells, and natural killer cells. This is a robust piece of work with rigorous controls, groups, and timepoints tested. The RNA-sequencing data provided helpful guidance on the pathways that should be assessed when analyzing AAV toxicity to the retina.

      Weaknesses:

      The main weakness of the study is that it focuses on subretinal administration to neonatal mice, and the canonical TLR9-MyD88 was not found to have an impact on the AAV toxicity measured. More information could have been provided to understand the discrepancy.