10,000 Matching Annotations
  1. Aug 2026
    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the referees for noting the substantive revisions and for the praise of the work. While we each have somewhat different weightings of likelihood, we feel the appraisals are fair and reasonable.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors addressed an important biological question, namely the role of glutamine metabolism in humoral responses, and they obtained solid conclusions. The strength of this study is that the authors used state-of-the-art transgenic mouse models together with in vitro analysis, thereby providing significant insights into the question posed. The following would strengthen the manuscript: i) adding more in-depth functionality/physiological relevance in the discussion part, and ii) regarding the experiments, the inclusion of more appropriate controls and a clearer and more accurate description of the methods.

      We are grateful for the decision of the Editors to select this submission for in-depth peer review and to the Reviewing Editor and referees for the thoughtful and constructive comments.

      We mostly agree with the specific comments and evaluation of strengths of what the work adds as well as with indications of limitations and caveats that apply to the breadth of conclusions. We have edited the text to be more clear and provide more details about certain aspects of the Methods and Legends. In addition, although we try to avoid Discussion sections that are unduly long or have flights of fancy, we will add to the Discussion as well as edit it for directness about potential relevance, basic explorations of mechanisms, and functionality.

      The revised manuscript also contains new data, some of it dealing with comments of the referees, other additions representing work done while the manuscript was under review. While we would be inclined to do more, the sad practical problem is one of limits placed by both the absence of any grant funds and the institution's terminations (RIFs) of the two experimenters in the lab.

      While we believe the original data interpretable as presented originally, up to a point it nonetheless is good to enhance scope or have even better data and add refinements about some of the technical issues. Ultimately, the question becomes "when is enough enough?"

      In the detailed point-by-point response below, we outline changes prompted by the reviewers. We also comment on a few points more expansively that would be suitable for the paper itself, and offer some skepticism or disagreement, (longer and more detailed explanations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Cho et al. present a comprehensive and multidimensional analysis of glutamine metabolism in the regulation of B cell differentiation and function during immune responses. They further demonstrate how glutamine metabolism interacts with glucose uptake and utilization to modulate key intracellular processes. The manuscript is clearly written, and the experimental approaches are informative and well-executed. The authors provide a detailed mechanistic understanding through the use of both in vivo and in vitro models. The conclusions are well supported by the data, and the findings are novel and impactful. I have only a few, mostly minor, concerns related to data presentation and the rationale for certain experimental choices.

      Detailed Comments:

      (1) In Figure 1b, it is unclear whether total B cells or follicular B cells were used in the assay. Additionally, the in vitro class-switch recombination and plasma cell differentiation experiments were conducted without BCR stimulation, which makes the system appear overly artificial and limits physiological relevance. Although the effects of glutamine concentration on the measured parameters are evident, the results cannot be confidently interpreted as true plasma cell generation or IgG1 class switching under these conditions. The authors should moderate these claims or provide stronger justification for the chosen differentiation strategy. Incorporating a parallel assay with anti-BCR stimulation would improve the rigor and interpretability of these findings.

      We edited the manuscript to be clear that total splenic B cells were used in this set-up figure and the rest of the paper. In addition, we performed new experiments to improve this "set-up figure (Fig. 1)" and moved the older data using alternative experimental conditions to a supplemental figure, Figure 1 - supplement 1. We also used new conditions that included styles of stimulating proliferation and differentiation - to foster an increased sense of generality. The findings in no way change the supported conclusions of the work. Specifically, we used mitogenic stimulation with anti-IgM <sup>+</sup> anti-CD40, all with BAFF, IL-4, and IL-5 in addition to the anti-CD40 stimulation of the original manuscript, bearing in mind excellent work from Aiba et al, Immunity 2006; 24: 259-268, and similar papers. In addition, we added a panel with representative flow cytometric profiles. These new data are presented in Figure 4 - supplement 1 (panels ae).

      To be transparent and add to a more open public discussion (using the virtues of this forum), the senior author and colleagues would caution about whether any in vitro conditions exist that warrant complete confidence. That is the reason for proceeding to immunization experiments in vivo. That is not said to cast doubt on our own in vitro data - there are some experiments (such as those of Fig. 1a-c and associated Fig 1 - supplement 1) that only can be done in vitro or are better done that way (e.g., because of rapid uptake of early apoptotic B cells in vivo).

      For instance: Well-respected papers use the CD40LB and NB21.2D9 systems to activate B cells and generate plasma cells. Those appear to be BCR-independent and yet continue in common use. [We found that these cellular systems (CD40LB; NB21.2D9) cannot be used in experiments with a.a. deprivation or the inhibitors due to effects on the engineered stroma-like cells.] In considering BCR engagement, Reth has published salient points about signaling and concentrations of the Ab, the upshot being that this means of activating mitogenesis and plasma cell differentiation (when the B cells are costimulated via CD40 or TLR (4 or 7/8) is also artificial. Moreover, although Aiba et al, Immunity 2006; 24: 259-268 is a laudable exception, one rarely finds papers using BAFF despite the strong evidence it is an essential part of the equation of B cell regulation in vivo and a cytokine that modulates BCR signaling - in the cultures.

      (2) In Figure 1c, the DMK alone condition is not presented. This hinders readers' ability to properly asses the glutaminolysis dependency of the cells for the measured readouts. Also, CD138<sup>+</sup> in developing PCs goes hand in hand with decreased B220 expression. A representative FACS plot showing the gating strategy for the in vitro PCs should be added as a supplementary figure. Similarly, division number (going all the way to #7) may be tricky to gate and interpret. A representative FACS plot showing the separation of B cells according to their division numbers and a subsequent gating of CD138 or IgG1 in these gates would be ideal for demonstrating the authors' ability to distinguish these populations effectively.

      In the revised manuscript, we have added new experimental data (Figure 1).

      We agree that exact placement of divisions and deconvolution by FlowJow is more fraught than might be thought from presentations in many or most papers. We include the data shown to the right as representative FACS plot(s) with old and new data that illustrate the gating on CTV fluorescence. With the representative examples pasted in here and presented in Fig 1 - supplement 1f, g of the revised manuscript, we will aver that using divisions 0-6, and ≥7 was and is entirely reasonable.

      Ditto for DMK with normal glutamine. However, in the spirit of eLife transparency lacking in many other journals, this comparison is more fraught than the referee comment would make things seem. The concentration tolerated by cells is highly dependent on the medium and glutamine concentration, and perhaps on rates of glutaminolysis (due to its generation of ammonia). In practice, DMK becomes more toxic to B cells unless glutamine is low or glutaminolysis is restricted. Thus, the concentration of DMK that is tolerated and used in Fig. 1b, c can become toxic to the B cells when using the higher levels of glutamine in typical culture media (2 mM or more) - at which point the "normal conditions <sup>+</sup> DMK" "control" involves the surviving cells in conditions with far greater cell death and less population expansion than the "low glutamine <sup>+</sup> DMK". condition.

      (3) A brief explanation should be provided for the exclusive use of IgG1 as the readout in classswitching assays, given that naïve B cells are capable of switching to multiple isotypes. Clarifying why IgG1 was preferentially selected would aid in the interpretation of the results.

      On lines ~112-3 and ~182-5, we edited the text in light of the referee's suggestion that we focus the presentation of serologic data on IgG1 in the immunization experiments. We also rearranged figures and panels to be more explicit and harmonize. That said, and [Brief explanation - IgG1 provides the strongest signal and hence better signal/noise both in vitro and with the alum-based immunizations that are avatars for the adjuvant used in the majority of protein-based vaccines for humans. Perhaps for this reason, the majority of papers on molecular mechanisms seem only to analyze IgG1. Nonetheless, since molecular regulation can differ according to isotype, and the more pro-inflammatory mouse IgG2c is more pertinent to some forms of anti-pathogen immunity and some auto-immune disease models, we believe it valuable to retain these data in supplements to the related Figures.]

      (4) The immunization experiments presented in Figures 1 and 2 are well designed, and the data are comprehensively presented. However, to prevent potential misinterpretation, it should be clarified that the observed differences between NP and OVA immunizations cannot be attributed solely to the chemical nature of the antigens - hapten versus protein. A more significant distinction lies in the route of administration (intraperitoneal vs. intranasal) and the resulting anatomical compartment of the immune response (systemic vs. lung-restricted). This context should be explicitly stated to avoid overinterpretation of the comparative findings.

      We appreciate the positive assessment, and agree with the referee that it is possible the conditions of immune challenge or re-exposure may contribute to the observed differences. We edited the text of the revised manuscript accordingly [lines ~152-153; ~159-160]. Certainly, the difference in how the anti-ova response is elicited compared to the anti-NP response in the same mice or with a bit different an immunization regimen might be another factor - or the major factor - explaining why glutaminolysis was important after ovalbumin inhalations (used because emergence of anti-ova Ab / ASCs is suppressed by the NP hapten after NP-ova immunization) but not needed for the anti-NP response unless Slc2a1 or Mpc2 also was inactivated. Thank you prompting addition of this important caveat!

      Nevertheless, it seems fair to note that in Figures 1 and 2, the ASCs and Ab are being analyzed for NP and ova in the same mice, albeit with the NP-specific components not being driven by the inhalations of ovalbumin. With that in mind, when one compares the IgG1 anti-NP ASC and Ab to those for IgG1 anti-ovalbumin (ASC in bone marrow; Ab), the ovalbumin-specific response was reduced whereas the anti-NP response was not. [lines ~171-172]

      (5) NP immunization is known to be an inducer of an IgG1-dominant Th2-type immune response in mice. IgG2c is not a major player unless a nanoparticle delivery system is used. However, the authors arbitrarily included IgG2c in their assays in Figures 2 and 3. This may be confusing for the readers. The authors should either justify the IgG2c-mediated analyses or remove them from the main figures. (It can be added as supplemental information with proper justification).

      We rearranged the Figure panels to move IgM and IgG2c data to Supplemental Figures (Figure 3 - supplements 1, 2, 4, 5 in the eLife system).

      For purposes of public discourse, we note first that in contrast to the premise about weak IgG2c responses, the data [previously, Figure 3(c, g); now in the supplements] show substantial levels of NP-specific IgG2c. The referee is quite right that the class switching and in vitro ASC generation were done with IL-4 / IgG1-promoting conditions.

      To assist readers, the revised manuscript takes note of the important role of IgG2c (mouse - IgG1 in humans) in controlling or clearing various pathogens as well as in autoimmunity [lines ~182-5]. Moreover, we continue to think that these measurements add substantial value both from the standpoint of providing a better sense of generality to the loss-of-function effects, and in considering potential ways of translating the findings to B cell-dependent autoimmune conditions such as systemic lupus erythematosus.

      [As a scientific aside, we speculate that a greater or lesser IgG2c anti-NP response may arise due to different preparations of NP-carrier obtained from the vendor (Biosearch) having different amounts of TLR (e.g., TLR4) ligand. In any case, the points of presenting the IgG2c (and IgM) data were to push against the limiting boundaries of convention (which risks perpetuating a narrow view of potential outcomes) and make the breadth of results more apparent to readers.

      (6) Similarly, in affinity maturation analyses, including IgM is somewhat uncommon. I do not see any point in showing high affinity (NP2/NP20) IgMs (Figure 3d), since that data probably does not mean much.

      As noted in the reply immediately preceding this one, we appreciate this suggestion from the reviewer and moved the IgM and IgG2c to supplemental status.

      Nonetheless, in collegial discourse we disagree a bit with the referee in light of our data as well as of work that (to our minds) leads one to question why inclusion of affinity maturation of IgM is so uncommon - as the referee accurately notes. Of course a defect in the capacity to class-switch is highly deleterious in patients but that is not the same as concluding that recall IgM or its affinity is of little consequence.

      In some of the pioneering work back in the 1980's, Bothwell showed that NP- carrier immunization generated hybridomas producing IgM Ab with extensive SHM (~11% of the 18 lineages; ~ 1/3 of the IgM hybridomas) [PMID: 8487778], IgM B cells appear to move into GC, and there is at least a reasonable published basis for the view that there are GC-derived IgM (unswitched) memory B cells (MBC) that would be more likely, upon recall activation, to differentiate into ASCs. [As an example, albeit with the Jenkins lab anti-rPE response, Taylor, Pape, and Jenkins generated quantitative estimates of the numbers of Ag-specific IgM<sup>+</sup> vs switched MBC that were GC-derived (or not). [PMID: 22370719]. While they emphasized that ~90% of IgM<sup>+</sup> MBC appeared to be GC-independent, their data also indicated that ~1/2 of all GC-derived MBC were IgM<sup>+</sup> rather than switched (their Fig. 8, B vs C; also 8E, which includes alum-PE). And while we immensely respect the referee, we are perhaps less confident that IgM or high-affinity Ag-specific IgM doesn't mean that much, if only because of evidence that localized Ab compete for Ag and may thus influence selective processes [PMCID: PMC2747358; PMID: 15953185; PMID: 23420879; PMID: 27270306].

      (7) Following on my comment for the PC generation in Figure 1 (see above), in Figure 4, a strategy that relies solely on CD40L stimulation is performed. This is highly artificial for the PC generation and needs to be justified, or more physiologically relevant PC generation strategies involving anti-BCR, CD40L, and various cytokines should be shown.

      In line with our response to point (1), we tested BCR-stimulated B cells (anti-CD40 plus anti-IgM with BAFF, IL-4, and IL-5, parallel to the analyses with anti-CD40 but no BCR engagement). These results align with and reinforce the utility of the data with anti-CD40 as the sole mitogen.

      (8) The effects of CB839 and UK5099 on cell viability are not shown. Including viability data under these treatment conditions would be a valuable addition to the supplementary materials, as it would help readers more accurately interpret the functional outcomes observed in the study.

      We added presentation of data that provide cues as to relative viability / cxmsurvival under the experimental conditions used.

      [FSC X SSC as well as 7AAD or Ghost dye panels; we also generated new data that in[ further experiments scoring annexin V staining (see Fig 4 - supplement 1d, e, and Fig 5 - supplement 1e, f)].

      (9) It is not clear how the RNA seq analysis in Figure 4h was generated. The experimental strategy and the setup need to be better explained.

      Including text added at lines ~291-293 and ~582-585, the revised manuscript provides more information in the Results, Methods and Legend for Fig 4j-l. We agree entirely with the concern and apologize that in this and a few other instances we inadvertently sacrificed sufficiency of detail on the altar of attempting brevity.

      [As a synopsis: In three temporally and biologically independent experiments, cultures were harvested 3.5 days after splenic B cells were purified and cultured as in the experiments of Fig. 4a-e. Total cellular RNA was prepared from the twelve samples (three replicates for each of four conditions - DMSO vehicle control, CB839, UK5099, and CB839 <sup>+</sup> UK5099), then analyzed by RNA-seq. RNA-seq data were initially processed using the pipeline described in the Methods. For panels g & h of Fig 4, DESeq2 was used to quantify and compare read counts in the three CB839 <sup>+</sup> UK5099 samples relative to the three independent vehicle controls and identify all genes for which variances yielded P<0.05. In Fig 4g, all such genes for which the difference was 'statistically significant' (i.e., P<0.05) were entered into the indicated Immgen tool and thereby mapped to the B lineage subsets shown in the figure panels (i.e., g, h). In (g), these are displayed using one format, whereas (h) uses the 'heatmap' tool in MyGeneSet.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigate the functional requirements for glutamine and glutaminolysis in antibody responses. The authors first demonstrate that the concentrations of glutamine in lymph nodes are substantially lower than in plasma, and that at these levels, glutamine is limiting for plasma cell differentiation in vitro. The authors go on to use genetic mouse models in which B cells are deficient in glutaminase 1 (Gls), the glucose transporter Slc2a1, and/or mitochondrial pyruvate carrier 2 (Mpc2) to test the importance of these pathways in vivo.

      Interestingly, deficiency of Gls alone showed clear antibody defects when ovalbumin was used as the immunogen, but not the hapten NP. For the latter response, defects in antibody titers and affinity were observed only when both Gls and either Mpc2 or Slc2a1 were deleted. These latter findings form the basis of the synthetic auxotrophy conclusion. The authors go on to test these conclusions further using in vitro differentiations, Seahorse assays, pharmacological inhibitors, and targeted quantification of specific metabolites and amino acids. Finally, the authors document reduced STAT3 and STAT1 phosphorylation in response to IL-21 and interferon (both type 1 and 2), respectively, when both glutaminolysis and mitochondrial pyruvate metabolism are prevented.

      Strengths:

      (1) The main strength of the manuscript is the overall breadth of experiments performed. Orthogonal experiments are performed using genetic models, pharmacological inhibitors, in vitro assays, and in vivo experiments to support the claims. Multiple antigens are used as test immunogens--this is particularly important given the differing results.

      (2) B cell metabolism is an area of interest but understudied relative to other cell types in the immune system.

      (3) The importance of metabolic flexibility and caution when interpreting negative results is made clear from this study.

      Weaknesses:

      (1) All of the in vivo studies were done in the context of boosters at 3 weeks and recall responses 1 week later. This makes specific results difficult to interpret. Primary responses, including germinal centers, are still ongoing at 3 weeks after the initial immunization. Thus, untangling what proportion of the defects are due to problems in the primary vs. memory response is difficult.

      We performed new experiments and added the data on differences prior to a boost [see below; new Fig 3d, e; etc].

      (2) Along these lines, the defects shown in Figure 3h-i may not be due to the authors' interpretation that Gls and Mpc2 are required for efficient plasma cell differentiation from memory B cells. This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. The more likely interpretation is that ongoing primary germinal centers are negatively impacted by Gls and Mpc2 deficiency, and this, in turn, leads to reduced affinities of serum antibodies.

      We have edited the wording of the conclusion to add a possibility we consider unlikely and downplay a conclusion that MBCs bearing switched BCRs are affected once reactivated. [see lines ~221-230] We also have added citations pertaining to the topic, including work from the Victora lab which seems to put the point succinctly: "Recall GCs in mice consist almost entirely of naïve B cells, whereas recall antibodies derive overwhelmingly from memory B cells." [emphasis added] [PMID: 38838672; new ref #83]. While unclear as to the reasoning - as one looks at the data - and skeptical as to the accuracy of the referee's point (2), it suggests that the matter is open to reasonable doubt. In line with the point and the edits, we also have added citation of a bioRxiv preprint from the Victora lab, which touches on the concept of what one could call boost-induced reinvigoration of a pre-existing GC [new ref #82].

      Beyond the textual changes, we performed a new series of experiments to investigate partially, and present the results in Fig 3d, e as well as Fig 3 - supplement 1d, e. Unfortunately, time before lab closure was an enemy both for the period between primary and recall immunizations in performance and multiple replication of work to extend that presented in Figure 3, panels g & h, and the related Supplemental Data (Fig 3 - supplements 4d, 5a-g). Unfortunately, it was not possible to do a longer-term memory experiment with recall immunization out at 8 weeks.

      The intriguing concerns and questions of points 1 & 2 provide a springboard for consideration of generalizations and simplifications. Germinal center durability is not at all monolithic, and instead is quite variable**. It is true that in the literature (especially with the substantially different approach of transferring BCR-transgenic / knock-in versions of an NP-biased BCR) there may be meaningful pools of IgG1 and IgG2c GC B cells. The premise (cognitive bias, perhaps?) in our interpretation is that in our previous work we measured few if any GC B cells - NP-APC-binding or otherwise - above the background (non-immunized controls) three weeks after immunization with NP-ovalbumin in alum. While recognizing that the immunogen can matter, we note for the readers and referee that Fig. 1 of the Taylor, Pape, & Jenkins paper considered above [PMID: 22370719] reported 10-fold more Ag-specific MBCs than GC B cells at day 29 post-immunization (the point at which the boost/recall challenge was performed in our Figure 3g, h. [That work did not use NP-carrier in alum to immunize, or measure the anti-NP response.]

      Viewing Fig. 3i from that perspective, the surmise of the comment is that a major contribution to the differences in both all-affinity and high-affinity anti-NP IgG1 (whose production requires differentiation into plasma cells) derived from the immunization at 4 wk stimulating persistent GC B cells as opposed to memory B cells.

      The issue and question also relate to rates of output of plasma cells or rises in the serum concentrations of class-switched Ab. To this point, our prior experiences agree with the long-published data of the Kurosaki lab in Figure 3c of the Aiba et al paper noted above (Immunity, 2006) (and other such time courses). Readers can note that the IgG1 anti-NP response (alum adjuvant, as in our work) hits its plateau at 2 wk, and did not increase further from 2 to 3 wk. The most likely interpretation is that GC are on the decline and Ab production has reached its plateau by the time of the 2nd immunization in Fig. 3h.

      Assuming we understand the comment and line of reasoning correctly, we also lean towards disagreeing with the statement " This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. Our evidence shows that both low-affinity as well as high-affinity anti-NP Ab (IgG1) were reduced due to combined gene-inactivation after the peak primary response (Fig. 3h; also, see the new data in Fig 3 and Fig 3 - supplement 1). Recent papers show that affinity maturation is attributable to greater proliferation of plasmablasts with high-affinity BCR. Accordingly, the findings with loss of GLS and MPC function are quite consistent with the interpretation that much of the response after the second immunization draws on MBC differentiation into plasmablasts and then plasma cells, where the proliferative advantage of high-affinity cells is blunted by the impaired metabolism. Notwithstanding these issues, the revised manuscript includes the alternative, if less likely, interpretation proposed by the review [lines ~221-230].

      **In some contexts, of course, especially certain viral infections or vaccination with lipid nanoparticles carrying modified mRNA, germinal centres are far more persistent; also, in humans even the seasonal flu vaccine

      (3) The gating strategies for germinal centers and memory B cells in Supplemental Figure 2 are problematic, especially given that these data are used to claim only modest and/or statistically insignificant differences in these populations when Gls and Mpc2 are ablated. Neither strategy shows distinct flow cytometric populations, and it does not seem that the quantification focuses on antigen-specific cells.

      The revised manuscript improves these aspects of the presentation, using old and new data. See Fig 3 - supplement 3a, c; Fig 3 - supplement 4a. We note for readers that many other papers in the best journals show plots in which the separation of, say, GC-Tfh from overall Tfh is based on cut-off within what essentially is a continuous spectrum of emission as adjusted or compensated by the cytometer (spectral or conventional).

      The revised manuscript presents results from new experiments that deal with the subset of GC B cells whose BCRs bind NP-APC with enough affinity to retain a positive signal after washing. These new data are presented in Fig 3 - supplement 3c & 3e. In practice, the new findings suggest that the metabolic requirement applied more to the NP-binding B cells than the overall GC B cell population.

      (4) Along these lines, the conclusions in Figure 6a-d may need to be tempered if the analysis was done on polyclonal, rather than antigen-specific cells. Alum induces a heavily type 2-biased response and is not known to induce much of an interferon signature. The authors' observations might be explained by the inclusion of other ongoing GCs unrelated to the immunization.

      We apologize for ambiguity or insufficient clarity and, as noted above, have edited the text to be more clear that the in vitro experiments do not represent GC B cells and that the RNA-seq data were from experiments that did not involve alum and were not an Ag (SRBC)-specific subset.

      New text in the Results, an expanded Legend, and tweaking the Methods make it more readily clear that the RNA-seq data (and hence the GSEA) involved immunizations with SRBC (not the alum / NP system. That said, we note that the hapten-carrier experiments in which the immunogen was adjuvantized with alum actually generated a robust IgG2c (type 1-driven) response along with the type 2-enhanced IgG1 response, in line with what has been reported by others with alum-adjuvanted vaccination.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors investigate how glutaminolysis (GLS) and mitochondrial pyruvate import (MPC2) jointly shape B cell fate and the humoral immune response. Using inducible knockout systems and metabolic inhibitors, they uncover a "synthetic auxotrophy": When GLS activity/glutaminolysis is lost together with either GLUT1-mediated glucose uptake or MPC2, B cells fail to upregulate mitochondrial respiration, IL 21/STAT3 and IFN/STAT1 signaling is impaired, and the plasma cell output and antigen-specific antibody titers drop significantly. This work thus demonstrates the promotion of plasma cell differentiation and cytokine signaling through parallel activation of two metabolic pathways. The dataset is technically comprehensive and conceptually novel, but some aspects leave the in vivo and translational significance uncertain.

      Strengths:

      (1) Conceptual novelty: the study goes beyond single-enzyme deletions to reveal conditional metabolic vulnerabilities and fate-deciding mechanisms in B cells.

      (2) Mechanistic depth: the study uncovers a novel "metabolic bottleneck" that impairs mitochondrial respiration and elevates ROS, and directly ties these changes to cytokinereceptor signaling. This is both mechanistically compelling and potentially clinically relevant.

      (3) Breadth of models and methods: inducible genetics, pharmacology, metabolomics, seahorse assay, ELISpot/ELISA, RNA-seq, two immunization models.

      (4) Potential clinical angle: the synergy of CB839 with UK5099 and/or hydroxychloroquine hints at a druggable pathway targeting autoantibody-driven diseases.

      We agree and thank the referee for the positive comments and this succinct summary of what we view as contributions of the paper.

      Weaknesses:

      (1) Physiological relevance of "synthetic auxotrophy"

      The manuscript demonstrates that GLS loss is only crippling when glucose influx or mitochondrial pyruvate import is concurrently reduced, which the authors name "synthetic auxotrophy". I think it would help readers to clarify the terminology more and add a concise definition of "synthetic auxotrophy" versus "synthetic lethality" early in the manuscript and justify its relevance for B cells.

      We edited the Abstract, Introduction, and Discussion to try to do better on this score. Conscious of how expansive the prose and data are even in the original submission, we appear to have taken some shortcuts that we will try to rectify or at least mitigate. Thank you for highlighting this need to improve on key concepts !!

      Specifically, the revised text expands a bit on the notion that synthetic auxotrophy represents effects on differentiation that go beyond additional mechanisms of reducing division efficiency and a modest impact on selective death. [see the 10th - 11th lines in Abstract and lines ~84-85, Introduction] Even though decreased population expansion is observed and new evidence supports a model in which the altered metabolism contributes to enhanced death in vivo, at equal division numbers the frequency of CD138<sup>+</sup> progeny is lower once glutaminolysis and mitochondrial pyruvate are reduced by either genetic or pharmacological means.

      This comment of the review raises interesting semantic questions about what represents "physiological relevance". The fundamental point is to explore a basic science question - what, if any, are limits to metabolic flexibility? In principle, shouldn't B cells be able to use fatty acid metabolism to generate enough ATP and provide the backbones for biosynthesis during growth? Put a different way, the point is that a basic curiosity to understand why decreasing glucose influx did not have an even more profound effect than what was observed, combined with curiosity as to why glutaminolysis was dispensable in relatively standard vaccine-like models of immunize/boost, provided a springboard to identification of new vulnerabilities. The manuscript shows one physiological limitation (and hence vulnerability). Be that as it may, the revised text of the Discussion section more clearly addresses this issue (lines ~531-549 at the end of the Discussion).

      While the overall findings, especially the subset specificity and the clinical implications, are generally interesting, the "synthetic auxotrophy" condition feels a little engineered.

      CAR-T cells are 'a little engineered' (or more than a little) and yet they do seem to have had an impact on understanding the centrality of B cells in various autoimmune conditions as well as in the direction of cancer therapy research. So it is a matter of balancing this perspective of the referee against the strengths they highlight in points 1, 2, and 4. In editing the revision, we try to expand and be more explicit about this in the Discussion of the revised manuscript.

      In brief, even were the money not all gone, we would not believe that expanding the heft of this already rather large manuscript and set of data would be appropriate. As matters stand, a basic new insight about metabolic flexibility and its limits leads to evidence of a way to reduce generation of Ab and a novel impairment of STAT transcription factor induction by several cytokine receptors. The vulnerability that could be tested in later work on B cell-dependent autoimmunity includes the capacity to test a compound that already has been to or through FDA phase II in patients together with an FDA-approved standard-of-care agent.

      Therefore, the findings strongly raise the question of the likelihood of such a "double hit" in vivo and whether there are conditions, disease states, or drug regimens that would realistically generate such a "bottleneck".

      Hence, the authors should document or at least discuss whether GC or inflamed niches naturally show simultaneous downregulation/lack of glutamine and/or pyruvate. The authors should also aim to provide evidence that infections (e.g., influenza), hypoxia, treatments (e.g., rapamycin), or inflammatory diseases like lupus co-limit these pathways.

      Again, we appreciate some 'licensing' to be more expansive and explicit, and will try to balance editing in such points against undue tedium or tendentiously speculative length in the Discussion. In particular, we will note that a clear, simple implication of the work is to highlight an imperative to test CB839 in lupus patients already on hydroxychloroquine as standard-of-care, and to suggest development of UK5099 (already tested many times in mouse models of cancer) to complement glutaminase inhibition.

      As backdrop, we note that the failure to advance imaging mass spectrometry to the capacity to quantify relative or absolute (via nano-DESI) concentrations of nutrients in localized interstitia is a critical gap in the entire field. Techniques that sample the interstitial fluid of tumour masses or in our case LN as a work-around have yielded evidence that there can be meaningful limitations of glucose and glutamine, but it needs to be acknowledged that such findings may be very model-specific and, as can be the case with cutting-edge science, are not without controversy. That said, yes, we had found that hypoxia reduced glutamine uptake but given the norms of focused, tidy packages only reported on leucine in an earlier paper [PMID27501247; PMCID5161594].

      Beyond all that, another impetus to and inspiration for these experiments stems from quite data that we generated in a model of short-term protein-restricted diet (loosely akin to kwashiorkor in humans), based on an excellent publication showing that such a regimen quickly led to lower circulating glutamine and mTORC1 activity (**). In brief, we found that a low-protein diet did, in our experiments, preferentially lower glutamine but - importantly - led to reduced Ab responses (which would match what we have modeled here). The findings were not a well-enough connected evidentiary component to include in the "story" but I'll append slides with the relevant data to this Response to Reviews for the referee's perusal (and anyone else who reads this online discourse).

      It would hence also be beneficial to test the CB839 + UK5099/HCQ combinations in a short, proof-of-concept treatment in vivo, e.g., shortly before and after the booster immunization or in an autoimmune model. Likewise, it may also be insightful to discuss potential effects of existing treatments (especially CB839, HCQ) on human memory B cell or PC pools.

      We certainly agree that the suggestions offered in this comment are important next steps and the right approach to test if the findings reported here translate toward the treatment of autoimmune diseases that involve B cells, interferons, and pathophysiology mediated by auto-Ab. As practical points, performance and replication of such studies would take more time than the year allotted for return of a revised manuscript to eLife and in any case neither funds nor a lab remain to do these important studies.

      Concrete evidence for our concurrence was embodied in a grant application to NIH that was essential for keeping a lab and doing any such studies. [We note, as a suggestion to others, that an essential component of such studies would be to test the effects of these compounds on B cells from patients and mice with autoimmunity]. Perhaps unfortunately for SLE patients, the review panelists did not agree about the importance of such studies. However, it can be hoped that the patent-holder of CB839 (and perhaps other companies developing glutaminase inhibitors) will see this peer-reviewed preprint and the public dialogue, and recognize how positive results might open a valuable contribution to mitigation of diseases such as SLE.

      (2) Cell survival versus differentiation phenotype

      Claims that the phenotypes (e.g., reduced PC numbers) are "independent of death" and are not merely the result of artificial cell stress would benefit from Annexin-V/active-caspase 3 analyses of GC B cells and plasmablasts. Please also show viability curves for inhibitor-treated cells.

      This comment leads us to see that the wording on this point may have been overly terse in the interests of brevity, and thereby open to some odd misunderstanding. The CD138<sup>+</sup> events are scored among VIABLE CELLS, so a decrease in the %CD138<sup>+</sup> at similar division number represents an effect independent from (or beyond) survival and division-counting. Accordingly, we expanded the text of the Abstract and elsewhere in the manuscript, to be more clear. In addition, we added data from new experiments addressing death in vitro and among GC-phenotype B cells in vivo. To clarify in this public context, it is not that an increase in death (along with the reported decrease in cell cycling) can be or is excluded. The point is that beyond any such increase, and taking into account division number (since there is evidence that PC differentiation and output numbers involve a 'division-counting' mechanism), the frequencies of CD138<sup>+</sup> cells and of ASCs among the viable cells are lower, as is the level of Prdm1-encoded mRNA even before the big increase in CD138<sup>+</sup> cells in the population.

      (3) Subset specificity of the metabolic phenotype

      Could the metabolic differences, mitochondrial ROS, and membrane-potential changes shown for activated pan-B cells (Figure 5) also be demonstrated ex vivo for KO mouse-derived GC B cells and plasma cells? This would also be insightful to investigate following NP-immunization (e.g., NP+ GC B cells 10 days after NP-OVA immunization).

      We performed a series of new experiments to have enough biologically independent replications for meaningful and statistical analyses. The new results, added in as Fig 5 - supplement 1, showed that the combined pathway interruption by loss-of-function increased ROS, mtROS, and death (annexin V / 7AAD) upon analyzing GCphenotype B cells immediately upon harvest. The findings align well with the data in Fig 5 (cultured B cells).

      (4) Memory B cell gating strategy

      I am not fully convinced that the memory-B-cell gate in Supplementary Figure 2d is appropriate. The legend implies the population is defined simply as CD19+GL7-CD38+ (or CD19+CD38++?), with no further restriction to NP-binding cells. Such a gate could also capture naïve or recently activated B cells. From the descriptions in the figure and the figure legend, it is hard to verify that the events plotted truly represent memory B cells. Please clarify the full gating hierarchy and, ideally, restrict the MBC gate to NP+CD19+GL7-CD38+ B cells (or add additional markers such as CD80 and CD273). Generally, the manuscript would benefit from a more transparent presentation of gating strategies.

      In considering the referee's viewpoint, we further expanded the supplemental data displays to include more of the gating and analytic schemes, which we believe should mitigate one concern noted here. In addition, we now include flow data from the non-immunized control mice that had been analyzed concurrently in the experiments.

      Third and finally, we performed new experiments and analyses in which the focus was the frequencies of memory-phenotype (IgD<sup>neg</sup> GL7<sup>neg</sup> CD38<sup>+</sup> / CD38<sup>hi</sup> aka CD38<sup>+</sup><sup>+</sup>) NPbinding B cells after immunization. While this time, as opposed to previously, the NP-APC staining met our standard for interpretability, the gist of the findings was that the two independent repeat experiments yielded a split decision and a degree of variability. With time being up due to the funds running out, we have elected to delete the issue and the data panel in question.

      That said, it bears noting that in the previous figure panel, the labeling indicated that the gating included the important criterion that cells be IgD<sup>neg</sup>, which excludes the vast majority of naive B cells but measures memory-phenotype B cells independent from consideration of whether or not they were NP-binding.

      [In principle marginal zone (MZ) B cells might fall within this gate. However, the MZ B population is unlikely to explain the differences shown.

      (5) Deletion efficiency - [The] mRNA data show residual GLS/MPC2 transcripts (Supplementary Figure 8). Please quantify deletion efficiency in GC B cells and plasmablasts.

      Even were there resources to do this, the degree of reduction in target mRNA (Gls; Mpc2) renders this question superfluous. To the best of our understanding, the proteins (for which there might be some phenotypic lag) are translated from RNA. Might there be a small subpopulation of B cells (or their PC progeny) with only one, or even neither, allele converted from fl to D? Yes, but they would be a minor subset in light of the magnitude of mRNA reduction, in contrast to our published observations with Slc2a1. As to plasmablasts and plasma cells, the pre-existing populations make such an analysis misleading, while the scarcity of such cells recoverable with antigen capture techniques is so low as to make both RNA and genomic DNA analyses questionable. We also refer readers to the supplemental figure that presents the results of experiments testing the issue one might infer from the question about extents of deletion in PC (i.e., how much counter-selection might have occurred by the PC stage).

    1. eLife Assessment

      This interesting study adapts machine learning tools to analyze movements of a chromatin locus in living cells in response to serum starvation. The machine learning approach developed is useful, the experiments are well controlled, and the data are solid. The study will benefit from future work testing predictions by perturbation experiments. This work will be of interest to those studying chromosome biology and gene expression patterns.

    2. Reviewer #1 (Public review):

      Summary:

      Redchuk et al. explore the dynamic properties of chromatin upon serum starvation using machine learning approaches. They use CRISPR-tagging to visualize a region on chromosome 1 in human cells and show that in their system, chromosome 1, but not the previously reported chromosomes 10, 13, and X, undergo a change in radial position upon serum starvation. Live cell imaging showed a position change towards the periphery after serum starvation. They then apply a machine learning algorithm for the analysis of the imaging data, which reveals changes in nuclear area during serum starvation and longer displacements of the chromosome 1 locus near the nuclear periphery. Differential behavior of homologues is also reported.

      Strengths:

      (1) The study of chromatin dynamics is an interesting and important area of research.

      (2) The use of machine learning approaches to analyze live cell imaging data is timely.

      (3) With serum starvation, the authors use a simple, well-controllable model system.

      Weaknesses:

      (1) This study provides limited new insight into chromatin dynamics.

      (2) It was not immediately evident what the use of machine learning approaches added to this study. It appears that the main conclusions could have been reached by conventional analysis.

      Comments on revised version:

      The authors have added some technical information, but have not made any major efforts to clarify some of the major points or to strengthen the paper. The degree of advance remains limited and several conclusions are not convincingly supported by the presented data.

    3. Reviewer #2 (Public review):

      Summary:

      The study demonstrates that CRISPR-Sirius provides a powerful approach to investigating chromosome dynamics in living cells during environmental stress. By focusing on serum starvation, the authors show that this process induces global nuclear changes, including a reduction in nuclear area and increased morphological dynamism, while at the same time driving specific reorganization of chromosome 1. Chromosome 1 relocates toward the nuclear periphery and displays distinctive patterns of motion, maintaining overall motility but punctuated by occasional long-distance displacements, particularly near the nuclear envelope. Importantly, the analysis reveals that homologous copies of chromosome 1 do not behave uniformly: peripheral loci become more mobile and responsive to starvation, whereas central homologs remain comparatively stable, often associated with nucleolar subcompartments. By integrating live imaging with machine learning and explainable AI analysis, the study highlights the complexity of nuclear organization and provides valuable insights into how chromosome-specific and locus-specific responses to stress are orchestrated within the three-dimensional nuclear landscape.

      Strengths:

      The study uses live-cell imaging to investigate the dynamics of loci during starvation. Live-cell tracking and data interpretation are carried out using machine learning and AI models, which is a major strength.

      Weaknesses:

      The manuscript is at times difficult to follow, partly because the methodological descriptions are highly specialized, especially for non-expert biologists. In addition, the observations are not tested for a mechanistic basis. Experiments that could provide deeper insights are missing, for example, why chromosome 1 moves, why the peripheral homologue dislocates, or why a "long jump" is observed at the periphery even though the speed of the loci does not change. It is also unclear whether a displacement of 0.5 μm is functionally meaningful.

      Comments on revised version:

      The authors have added some technical information and provided a better discussion of the data, but beyond that, they have not strengthened the conclusions. The observations are not supported by any perturbation assays.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This interesting study adapts machine learning tools to analyze movements of a chromatin locus in living cells in response to serum starvation. The machine learning approach developed is useful, the experiments are well controlled, and the data are solid. The study would be greatly strengthened by testing key predictions made using perturbation experiments. This work will be of interest to those studying chromosome biology and gene expression patterns.

      We thank eLife for this nice assessment. We indeed believe that the presented machine learning approach will be useful for many types of research questions, and this was the main aim of this manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Redchuk et al. explore the dynamic properties of chromatin upon serum starvation using machine learning approaches. They use CRISPR-tagging to visualize a region on chromosome 1 in human cells and show that in their system, chromosome 1, but not the previously reported chromosomes 10, 13, and X, undergo a change in radial position upon serum starvation. Live cell imaging showed a position change towards the periphery after serum starvation. They then apply a machine learning algorithm for the analysis of the imaging data, which reveals changes in nuclear area during serum starvation and longer displacements of the chromosome 1 locus near the nuclear periphery. Differential behavior of homologues is also reported.

      Strengths:

      (1) The study of chromatin dynamics is an interesting and important area of research.

      (2) The use of machine learning approaches to analyze live cell imaging data is timely.

      (3) With serum starvation, the authors use a simple, well-controllable model system.

      Weaknesses:

      (1) This study only provides limited new insight into chromatin dynamics.

      We respectfully disagree with this conclusion. To the best of our knowledge, our study is the first to provide any insights into chromatin dynamics upon serum starvation. Previous studies are solely based on studies in fixed cells, and the dynamics have remained unexplored. Moreover, for example the notion that homologous chromosomes show differential dynamic behavior is novel and will likely have implications and relevance to many chromatin-based processes beyond the example studied here.

      (2) It was not immediately evident what the use of machine learning approaches added to this study. It appears that the main conclusions could have been reached by conventional analysis.

      First, we would like to point out that the other reviewer found our machine learning analysis pipeline a major strength of our manuscript. Indeed, analyzing single features and assessing their impact on the studied phenomenon could have been achieved relatively easily by conventional analysis. However, this analysis would have ignored the interactions (some of which were not intuitively obvious) between different features and thereby limited the knowledge gain from the experiment.

      Unbiased analysis of the interactions between the different features would have been already very difficult and time-consuming with conventional approaches. We believe that our analysis pipeline, especially with the Shapley values, addresses the key issue of combinatorial explosion prominent to multiparametric data, such as imaging data, and helps the researcher to navigate complex datasets.

      (3) There are several specific technical points:

      (a) It was not clear what the CRISRP-Sirius probes actually labelled. The chromosome 1 sgRNA sequence is provided, but I could not find information as to which region(s) of the chromosome are actually labelled (size, location, etc.).

      We have added a schematic as Supplementary Figure 1A to show the region of the chromosome that is labelled. In addition, the target sequence, together with the relevant references can be found in the Materials and methods (page 16). Please see also below Reviewer #1 (Recommendations for the authors) point 4a.

      (b) The authors visualize a relatively small region of chromosome 1 but make conclusions regarding the entire chromosome. Additional probes on the same chromosome should be used.

      Related to this point, the discussion of why the authors are unable to reproduce the prior findings of relocation of chromosomes 10, 13, and X is not satisfying. It would be worth comparing the FISH-based painting of entire chromosomes, which generated the results suggesting relocation of these chromosomes, with the point-labelling method used here.

      We agree that our approach to labeling chromosome 1 is very different than the FISH-based probes utilized before. However, we also feel that we discuss this aspect, and the difference between our and previous results, which may also stem from the used cell model, in quite a detail in the first paragraph of the results (page 4). Also, we are very careful throughout the manuscript to indicate that here we study the dynamics of a specific chromosome loci, not the entire chromosome, and have further amended the text to emphasize this. In the future, it would be very interesting to study the dynamics of also other loci of chromosome 1. As indicated also below in response to reviewer 2, we have failed to identify further gRNAs that would reliably and reproducibly label further chromosome 1 loci, suggesting that we would need to change the labeling system entirely. Unfortunately, this is not in the scope of this manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 1.

      (c) The study lacks controls. Since in their hands chromosomes 10, 13, and X do not change position, they should be used as a negative control in all experiments demonstrating a shift in the location of chromosome 1.

      We disagree that our study lacks controls, since we use telomeres as controls throughout the manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 2,3.

      (d) I did not find information about the spatial or temporal resolution of the imaging modality. This is important to assess whether the observed changes in position, relative to time, are meaningful.

      To estimate the spatial resolution, we have added new data using fixed cells (Supplementary figure 1E; corresponding text in results on page 5); temporal resolution is indicated in Materials and methods (page 17). Please see also below Reviewer #1 (Recommendations for the authors) point 4d.

      (e) The authors analyze surprisingly early timepoints (up to 40 minutes) of serum starvation. Would these results look different if longer serum starvation timepoints of several hours were analyzed?

      We chose to analyze early time points of serum starvation based on the previous literature reporting the chromosome relocation within the first 15 minutes of starvation. Indeed, the results might look very different later during serum starvation, since we already observe differences between 0-20 min vs 20-40 min into starvation (see for example Figure 5A-D). Analyzing further time points is not in the scope of this manuscript.

      (f) The authors can do a better job of explaining what the biological meaning of the various parameters (DistR, TDist, etc.) they measure is.

      We have amended Table 1 to describe the measured features more clearly. Please see also below Reviewer #1 (Recommendations for the authors) point 4e.

      (g) I did not understand the reasoning for the authors' conclusion of differential behavior of homologues. Please explain this better, or idealy use more direct labeling methods that identify the individual homologues.

      The differential behavior of homologues is best demonstrated in Figure 6H, which shows that in serum-containing media, the peripheral homolog has equal probability of being faster or slower compared to its homolog. However, the distribution changes upon starvation, with the peripheral loci being more frequently the faster homolog. We completely agree that further studies are needed to understand this phenomenon better, but changing the labeling method is not in the scope of this manuscript.

      (h) In many figures, statistical analysis of the data is missing, including, but not limited to, Figures 1B, C, G, Figures 4, 5, 6.

      We have added a Supplementary table to include inferential statistics. See also below Reviewer #1 (Recommendations for the authors) point 4b.

      (i) No information is provided throughout the manuscript as to how many cells were analyzed in each experiment. This should be indicated in every figure legend.

      The number of analyzed loci or nucleus is indicated in every figure. See also below Reviewer #1 (Recommendations for the authors) point 4c.

      Reviewer #2 (Public review):

      Summary:

      The study demonstrates that CRISPR-Sirius provides a powerful approach to investigating chromosome dynamics in living cells during environmental stress. By focusing on serum starvation, the authors show that this process induces global nuclear changes, including a reduction in nuclear area and increased morphological dynamism, while at the same time driving specific reorganization of chromosome 1. Chromosome 1 relocates toward the nuclear periphery and displays distinctive patterns of motion, maintaining overall motility but punctuated by occasional long-distance displacements, particularly near the nuclear envelope. Importantly, the analysis reveals that homologous copies of chromosome 1 do not behave uniformly: peripheral loci become more mobile and responsive to starvation, whereas central homologs remain comparatively stable, often associated with nucleolar subcompartments. By integrating live imaging with machine learning and explainable AI analysis, the study highlights the complexity of nuclear organization and provides valuable insights into how chromosome-specific and locus-specific responses to stress are orchestrated within the three-dimensional nuclear landscape.

      Strengths:

      The study uses live-cell imaging to investigate the dynamics of loci during starvation. Livecell tracking and data interpretation are carried out using machine learning and AI models, which is a major strength.

      Weaknesses:

      The manuscript is at times difficult to follow, partly because the methodological descriptions are highly specialized, especially for non-expert biologists. In addition, the observations are not tested for a mechanistic basis. Experiments that could provide deeper insights are missing, for example, why chromosome 1 moves, why the peripheral homologue dislocates, or why a "long jump" is observed at the periphery even though the speed of the loci does not change. It is also unclear whether a displacement of 0.5 μm is functionally meaningful.

      We appreciate the comment about the readability of our manuscript, and have seriously evaluated this point. We also completely agree that it would be interesting and important to understand the mechanistic and functional basis of the observed changes in chromatin dynamics take place upon serum starvation. However, we feel that it is not in the scope of the present manuscript. See also below Reviewer #2 (Recommendations for the authors) points 3,7-11.

      Recommendations for the authors:

      Reviewing Editor Comments:

      I would like to first offer my congratulations on a very interesting study; second, I would like to encourage you to test a few key predictions using a perturbation experiment. Two reviewers with deep expertise in this area were supportive of the work, and both noted that such an addition would greatly increase the impact and visibility of this work in the field. I welcome a revision that addresses this seminal point. Thank you for sending your work to eLife!

      We thank eLife for the positive assessment. We have aimed to address all of the reviewers comments and suggestions. However, we feel that some of the suggestions are not in the scope of this particular manuscript, since they would require setting up a different chromatin labeling system.

      Reviewer #1 (Recommendations for the authors):

      The following experiments would strengthen the study:

      (1) Please label additional regions on chromosome 1 so as not to rely on a single point to represent the behavior of the entire chromosome.

      This is an excellent suggestion, but unfortunately, despite our extensive efforts, we have failed to identify further gRNAs that would reliably label chromosome loci with the CRISPR-Sirius system. Changing the labeling system is not in the scope of the presented manuscript.

      (2) Please use chromosomes 10, 13, or X as a negative control since these chromosomes do not change position in the authors' hands.

      (3) Please compare the behavior of the homologues to that of either random loci or control loci on 10, 13, or X to assess whether the differential behavior observed for chromosome 10 is a specific effect.

      Related to points 2 and 3, we opted to use telomeres as controls in this study. Throughout the manuscript, the behavior of chromosome 1 loci is compared to telomeres, demonstrating the specific effect of serum starvation on chr 1. For example, Figure 5A and 5B show that when analyzing mean locus displacement, chr1 and telomeres show the opposite behavior.

      (4) In addition:

      (a) Please provide detailed information on the sequence and location of the probes used.

      We have added a schematic showing the location of the probes as Supplementary Figure S1A. In addition, the sequences are indicated in Materials and methods (page 16).

      (b) Please provide a statistical analysis in all graphs.

      To make statistical analysis more comprehensive, we have added supplementary table 1, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. See also below Reviewer #2 (Recommendations for the authors) point 5.

      (c) Please provide throughout the manuscript in each figure legend information as to how many cells were analyzed in each experiment.

      The number of analyzed loci (or nucleus) is indicated in each graph.

      (d) Please provide information on the spatial and temporal resolution of the imaging modality.

      The imaging settings are indicated in Materials and methods, including the temporal resolution of 0.25 frames per second (page 17). To estimate spatial resolution, and especially its relationship with the observed repositioning of the chromosome loci, we performed experiments in fixed cells, using an optically identical set-up as utilized for live imaging. Unfortunately, the microscope utilized for live imaging was taken out of use by the core facility after submission of the original draft of this manuscript, but we used a microscope with essentially a similar set-up. The data from fixed cells is now presented as Supplementary figure 1E and discussed in results on page 5. This analysis indicates that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error.

      (e) Please better explain what the various measured parameters mean in biological terms.

      We have amended Table 1 to provide better explanation of the measured parameters.

      (f) Please add a scale bar to Figure 6I.’

      Scale bar has been added to figure 6I.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) SHAP analysis identified nuclear area (MA) and its change (sA) as the most predictive features of starvation state, while motility features (MD, MaxD, TD) showed strong interactions with nuclear morphology. Discrete features, such as displacement outliers and homolog subclassification by speed/proximity, influenced classification, particularly in MLP models. Could the authors clarify why morphological and motility features act in combinatorial and context-dependent ways? A biological interpretation of this interdependence would strengthen the study.

      Unfortunately, we do not have a good biological interpretation for this. The fact that some interactions are context-dependent indicates that there could be subpopulations of cells/analyzed loci. For example, we found that the predictive value of nuclear area was high in a subgroup of low motility loci (Figure 4F and Supplementary figure 4D-F). We do not believe that adding more speculation would strengthen the study.

      (2) The manuscript shows that chromosome 1 moves toward the periphery within the first 20-40 minutes of serum withdrawal. However, it remains unclear whether the locus eventually "touches" the periphery and whether it subsequently stabilizes or retracts. It would be valuable to compute the time point of minimal nuclear distance and examine whether this is transient or sustained.

      With the experimental set-up utilized here, we imaged the loci for only two minutes at random time point within the first 40 minutes of the starvation. Hence extracting the time point of minimal nuclear distance is not meaningful from this dataset. As we discuss in the manuscript, following the dynamics of the same locus for longer periods of this would be very interesting in the future. However, this is not in the scope of the present manuscript.

      (3) The manuscript is at times difficult to follow, partly because methodological descriptions are highly detailed in the main text. Consider moving more of the methodological content into Supplementary Methods and emphasizing the main results and interpretations in the main text for clarity.

      We have carefully evaluated this point. Most methodological descriptions in the manuscript relate to the machine learning models and their explanation with SHAP. As we feel that this combination is an essential part of the manuscript, and likely the aspect that can have widest impact beyond chromatin dynamics studies, we feel that the background and our reasoning related to the chosen methods are important.

      (4) The distinction between the first 20 minutes and the latter 40-minute window is intriguing. Could these different time scales be paralleled with early versus delayed gene expression responses to serum starvation? A discussion of this temporal connection would add biological depth.

      This is an intriguing idea, and we have added a short note on this in the discussion (page 13). However, as we do not know how the U2OS cells utilized here respond transcriptionally to serum starvation, we are hesitant to speculate too much.

      (5) If the observed interpretations are robust, could this be demonstrated more explicitly through statistical principles or reproducibility tests across independent datasets?

      To provide further evidence of the robustness of our findings, we have 1) added new data to estimate the spatial resolution (Supplementary figure 1E) and 2) expand the statistical analysis as supplementary table 1. Regarding the spatial resolution (see also the response to reviewer 1), our experiments on fixed cells demonstrate that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. To make statistical analysis more comprehensive, we added a supplementary table, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. MW test was used as a non-parametric statistic, with null hypothesis assuming the samples are coming from the same distribution. Null hypothesis was rejected at the p-value below 0.05. In most cases (bold font in table) MW test results were in accordance with the descriptive statistics confirming the conclusions in the study. In case of exceptions (MD, TD for telomeres and TDist), the results were reported, for example, as an “appearing trend” to reflect descriptive statistics while highlighting certainty levels.

      (6) Figure labeling is difficult to follow. Please include abbreviation explanations directly in the figure panels or legends for clarity.

      Abbreviations have been added to figure legends. Adding them to figures themselves would have made the figures too busy.

      (7) The manuscript reports higher displacement at the nuclear periphery. Can the authors explain why displacement amplitudes increase near the periphery and how this relates to nuclear architecture?

      We speculate in the manuscript (results, page 11; discussion, page 14) that actually the lower displacement observed with the central locus may, at least partially, result from anchoring this locus to the nucleolus (Fig 6I). Nevertheless, alternative explanations, such as differences in transcriptional and/or chromatin states may exist (see also the response to point 11), and this is now mentioned in the discussion (page 14).

      (8) How is the movement of chromosome 1 directed specifically toward the periphery, rather than being random fluctuations? This point requires clarification.

      This is an important question, but unfortunately our data does not provide an answer to this, and suggesting any mechanism would be pure speculation. Nevertheless, our results agree with previous studies utilizing fixed cells that also demonstrated movement of chromosome 1 towards nuclear periphery (Mehta et al., 2010), arguing against random fluctuation.

      (9) Only chromosome 1, and not the other tested chromosomes, undergoes this relocalization. Could the authors elaborate on why some chromosomes but not others display this behavior?

      Previous studies (Mehta et al., 2010) utilizing chromosome paints in fixed cells actually show the relocalization of several chromosomes upon serum starvation. The fact that we observed the relocalization of only chr1 loci is likely due to the labeling method and/or the cell model utilized in this study. This is quite explicitly discussed in the first paragraph of results (page 4).

      (10) The magnitude of these movements appears relatively small (0.5 micron). Can the authors discuss whether such small but reproducible displacements are likely to be biologically meaningful in terms of nuclear function or gene regulation?

      At the moment, our experimental set up allows us to analyze the dynamics of only a small portion of chr1, which indeed shows an average 0.5 micron displacement towards the nuclear periphery. Based on the chromosome painting data from fixed cells, the displacement at the level of whole chromosome is significantly larger. As mentioned in the discussion (page 13), the functional implications of radial repositioning of chromosomes upon serum starvation is not known. Therefore further discussion on the relevance of the magnitude reported here would be pure speculation.

      (11) Peripheral homologs of chromosome 1 became faster and more dynamic under starvation. Why might these loci be more prone to movement? Could this be linked to differences in transcriptional activity or chromatin state between central and peripheral homologs?

      At the moment we favour the idea that the central homolog is constrained by its anchorage to the nucleolus (Figure 6I). However, transcriptional activity and/or chromatin state may also play a role, and this possibility is now mentioned in the discussion on page 14.

      Minor points:

      (1) Figure legends use inconsistent capitalization and panel labels. These should be standardized across all figures for better readability.

      We apologize for these inconsistencies, and have aimed to standardize all labeling.

      References

      Mehta, I.S., Amira, M., Harvey, A.J., and Bridger, J.M. (2010). Rapid chromosome territory relocation by nuclear motor activity in response to serum removal in primary human fibroblasts. Genome Biol 11, R5.

    1. eLife Assessment

      This important study reports insights into how the caspase Dcp-1, best known for cell death, can also promote tissue growth in Drosophila, extending the authors' earlier work by identifying regulatory factors that shape this non-lethal activity. The compelling findings identify a physical and functional interaction between Dcp-1 and Bruce, as well as new Dcp-1-interacting proteins that function in autophagy: Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a. This work helps broaden the understanding of the non-lethal roles of Dcp-1.

    2. Reviewer #1 (Public review):

      The authors clearly demonstrate that overexpressed Dcp-1, but not Drice, is activated without canonical apoptosome components.

      Using TurboID-based proximity labeling they revealed distinct proximal proteomes, among which Sirtuin 1, an Atg8a deacetylase, which promotes autophagy, was specifically required for Dcp-1 activation. Additionally, the show that autophagy-related genes, including Bcl-2 family members Debcl and Buffy, are required for Dcp-1 activation. Using structure-based prediction using AlphaFold3 they identified that Bruce, an autophagy-regulated inhibitor of apoptosis, as a Dcp-1-specific regulator acting outside the apoptosome-mediated pathway. Finally, they show that Bruce suppresses wing tissue growth. These findings indicate that non-lethal Dcp-1 activity is governed by the autophagy- Bruce axis, enabling distinct non-lethal functions independent of cell death.

      Comments on revised version.

      No further comments.

    3. Reviewer #2 (Public review):

      Summary:

      The Drosophila executioner caspase Dcp-1 has established roles in cell death, autophagy, and imaginal disc growth. This study reports previously unrecognized factors that work together with Dcp-1. Specifically, the authors performed a turboID-based proximal ligation experiment to identify factors associated Dcp-1 and Drice. Dcp-1-specific interactors were further examined for their genetic interaction. The authors report autophagy-related genes, including Debcl and Buffy, to be required for Dcp-1 activation. In addition, the authors present evidence of an interaction between Bruce and Dcp-1. Bruce expression blocks the Dcp-1 overexpression phenotype. Inhibition of effector caspases or overexpression of Bruce commonly reduced wing growth, suggesting a relationship between the two proteins.

      Strengths:

      The study identifies new Dcp-1-interacting proteins and provides a functional link between Dcp-1 and Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a. During the revision, the authors have also added convincing new data supporting the interaction between Dcp-1 and Bruce. They further make a strong case regarding the quality of the turboID-proteomics data. Overall, this is a strong manuscript supporting an interesting discovery.

    4. Reviewer #3 (Public review):

      Summary:

      The present paper by Shinoda et al. from the Miura group builds upon findings reported in an earlier study by the same team (Shinoda et al., PNAS, 2019), which identified a non-apoptotic role for the Drosophila executioner caspase Dcp-1 in promoting wing tissue growth. That earlier work attributed this function primarily to Dcp-1 and to Decay, a caspase structurally related to executioner caspases, but not to DrICE, the principal apoptotic executioner caspase. The authors further proposed that this non-apoptotic caspase activity operates independently of the initiator caspase Dronc.

      In the current study, the authors both corroborate aspects of their previous findings and extend the investigation to mechanisms regulating Dcp-1 in this context. They identify roles for the giant IAP Bruce, two BCL-2 family members, and autophagy-related components in modulating non-apoptotic Dcp-1 activity. Moreover, they show that Bruce binds to a BIR-like peptide exposed upon Dcp-1 cleavage, but not to DrICE. The study further suggests that low levels of Dcp-1 activity promote wing tissue growth, whereas excessive activity induces cell death, as evidenced by impaired wing development following Dcp-1 overexpression. Overall, the manuscript provides several intriguing insights into the non-apoptotic regulation of the comparatively weak apoptotic executioner caspase Dcp-1 and complements the group's earlier work. However, several concerns remain regarding certain interpretations of the data and the experimental rigour of some of the results.

      Strengths:

      A major strength of the work is its systematic genetic and biochemical approaches, which combine tissue-specific manipulation with protein interaction mapping to explore how Dcp-1 is regulated. The identification of several regulatory factors, including an inhibitor of cell death protein and components linked to autophagy, provides a coherent framework for understanding how Dcp-1 activity might be tuned.

      Weaknesses:

      The evidence supporting some key claims remains incomplete. In particular, the type of cell death form induced when Dcp-1 is overexpressed is not clearly established, and additional tests would be needed to distinguish between the different cell death types.

      Likely impact:

      The study contributes to a growing body of work showing that proteins traditionally associated with cell death can have broader roles in tissue development. This conceptual advance is likely to be of interest to researchers studying growth control and tissue maintenance.

      Specific points:

      (1) Nature of the wing ablation phenotype<br /> A central concern is whether the wing ablation phenotype observed upon Dcp-1 overexpression truly reflects apoptotic cell death. The authors show in Fig. 1c that nuclei in cells overexpressing Dcp-1, but not DrICE, zymogens are highly condensed, which is suggestive of apoptosis. However, it is equally plausible that this phenotype reflects a form of non-apoptotic, Dcp-1-dependent cell death (e.g. autophagy-dependent cell death). This distinction could be readily addressed using TUNEL labelling and direct caspase activity assays. The latter would be particularly informative, as it remains unclear whether zymogen Dcp-1 is capable of cleaving standard effector caspase reporters in vivo. Does the anti-cleaved Dcp-1 antibody detect Dcp-1 activation following overexpression of the Dcp-1 zymogen?

      (2) Role of Decay<br /> In their earlier study, the authors identified Decay as another caspase influencing wing growth, albeit more modestly than Dcp-1. It is therefore unclear why this line of investigation was not pursued further in the current work. This omission is notable, as Decay is not implicated in apoptosis and, to date, no substantial physiological function has been assigned to this caspase in any system. At minimum, this point should be discussed explicitly.

      (3) Fig. 2: Proximity labelling analysis<br /> The authors use TurboID-mediated proximity labelling to reveal distinct Dcp-1- and DrICE-associated proteomes across tissues, with a particular focus on the wing disc. They further demonstrate that RNAi-mediated knockdown of the Dcp-1-associated proteins Sirt1 and Fkbp59 suppresses the wing ablation phenotype induced by Dcp-1 overexpression, suggesting that these factors are required for Dcp-1 activity. However, it should be clarified whether Bruce was identified as a Dcp-1 interactor in the proximity labelling dataset, given its proposed central regulatory role. In addition, further discussion of Fkbp59, its known functions and how it might mechanistically influence Dcp-1 activity, would be valuable.

      (4) Fig. 3: Autophagy-related factors<br /> Given that Sirt1 is known to promote autophagy, the authors next examine autophagy-related proteins and identify roles for Atg2, Atg8a, Debcl, and Buffy in Dcp-1 activation. Notably, these proteins do not promote cell death in the Hid-induced canonical apoptotic pathway. However, it is important to determine whether knockdown of Debcl, Buffy, Atg2, or Atg8a alone affects wing development in the absence of Dcp-1 overexpression, to exclude the possibility that these perturbations independently impair wing formation.

      (5) Evidence for canonical autophagy<br /> The involvement of autophagy would be more convincingly demonstrated by testing additional core autophagy genes, such as Atg7, Atg5, and Atg12, as well as performing a combined knockdown of Atg8a and Atg8b. Moreover, direct assessment of autophagy at the cellular level using established genetic reporters would substantially strengthen the conclusions.

      (6) Figs. 4-5: Functional consequences<br /> It would be informative to determine whether Synr, Debcl, or Buffy influence wing size on their own and whether their overexpression enhances wing growth.

      (7) Terminology and interpretation of cell death<br /> Taken together, the results suggest that Dcp-1 zymogen overexpression induces a form of non-apoptotic cell death, potentially autophagy-dependent or related. The reviewer does not understand the authors' insistence on referring to this process as apoptosis. The authors should be more cautious in their terminology: there is no canonical versus non-canonical apoptosis, there is simply apoptosis. Without stronger evidence, these effects should not be described as apoptotic cell death.

      Comments on revised version.

      In the revised manuscript, the authors addressed each of my concerns in good faith and, in my opinion, responded to them thoroughly and satisfactorily. I have no further concerns.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We are grateful to all the reviewers for dedicating time to review our manuscript and for providing insightful comments and suggestions. We have revised our manuscript in line with the reviewers' feedback. The major revisions include characterization of Dcp-1 overexpression-induced cell death, demonstration of the involvement of autophagy in Dcp-1 activation, characterization of the interaction between full-length Bruce and cleaved Dcp-1. We have introduced new figures (Figure 1 – figure supplement 1, Figure 2 – figure supplement 2, Figure 3 – figure supplement 1, Figure 5 – figure supplement 1), new panels (Figures 1D, Figure 3C, Figure 4I, J) and a new table (Table S2). The previous Figure 5 – figure supplement 1 has been relocated to Figure 4 – figure supplement 2.

      With all concerns and suggestions from the reviewers addressed, our conclusion—that Bruce suppresses autophagy-regulated caspase activity and wing tissue growth in Drosophila— is now more robustly supported. We are confident that our revised manuscript makes a significant contribution to the fields of cell death, autophagy, and developmental biology, as it provides a new conceptual framework for understanding non-lethal caspase regulation. We remain hopeful that the reviewers will find it suitable for publication in eLife.

      Reviewer #1 (Public review):

      Summary:

      The authors clearly demonstrate that overexpressed Dcp-1, but not Drice, is activated without canonical apoptosome components. Using TurboID-based proximity labeling, they revealed distinct proximal proteomes, among which Sirtuin 1, an Atg8a deacetylase, which promotes autophagy, was specifically required for Dcp-1 activation. Additionally, the show that autophagy-related genes, including Bcl-2 family members Debcl and Buffy, are required for Dcp1 activation. Using structure-based prediction using AlphaFold3, they identified that Bruce, an autophagy-regulated inhibitor of apoptosis, acts as a Dcp-1-specific regulator acting outside the apoptosome-mediated pathway. Finally, they show that Bruce suppresses wing tissue growth. These findings indicate that non-lethal Dcp-1 activity is governed by the autophagy-Bruce axis, enabling distinct non-lethal functions independent of cell death.

      Strengths:

      This is an excellent paper with very good structure, excellent quality data and analysis.

      Weaknesses:

      This reviewer did not identify any weaknesses or recommendations for revision.

      We sincerely thank the reviewer for their highly positive evaluation of our work. We are pleased that the reviewer found the overall structure, data quality, and analyses to be strong, and that they clearly recognized the key findings of our study. No changes to the manuscript were required in response to this review.

      Reviewer #2 (Public review):

      Summary:

      The Drosophila executioner caspase Dcp-1 has established roles in cell death, autophagy, and imaginal disc growth. This study reports previously unrecognized factors that work together with Dcp-1. Specifically, the authors performed a turboID-based proximal ligation experiment to identify factors associated Dcp-1 and Drice. Dcp-1-specific interactors were further examined for their genetic interaction. The authors report autophagy-related genes, including Debcl and Buffy, to be required for Dcp-1 activation. In addition, the authors present evidence of an interaction between Bruce and Dcp-1. Bruce-expression blocks the Dcp-1 overexpression phenotype. Inhibition of effector caspases or overexpression of Bruce commonly reduced wing growth, suggesting a relationship between the two proteins.

      Strengths:

      On the positive side, the study identifies new Dcp-1-interacting proteins and provides a functional link between Dcp-1 and Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a.

      Weaknesses:

      The data supporting the Dcp-1/Bruce interaction are not strong, even though the title of this manuscript highlights Bruce. For example, the authors' turboID data does not support Dcp1/Bruce interaction. The case for the interaction is based on a single experiment that overexpresses a truncated Bruce transgene in S2 cells.

      We sincerely thank the reviewer for their constructive and detailed evaluation of our manuscript. We appreciate the positive assessment that our study identifies new Dcp-1-associated factors and provides functional links between Dcp-1 and Sirt1, Fkbp59, and multiple autophagy-related genes, including Debcl, Buffy, Atg2, and Atg8a. We also thank the reviewer for clearly pointing out concerns regarding the strength and interpretation of the evidence connecting Bruce and Dcp-1. In the revised manuscript, we have addressed these concerns in two major ways. First, we provided additional experimental evidence explaining why TurboID-mediated labeling did not identify Bruce. Specifically, we showed that the majority of TurboID-tagged Dcp-1 expressed in wing imaginal discs remains in its full-length form, which is unlikely to engage Bruce. Second, and more importantly, we now demonstrated that endogenously expressed full-length Bruce interacts with cleaved Dcp-1 in wing imaginal discs. These new data provide strong support for a physiologically relevant interaction between Bruce and cleaved Dcp-1. Detailed descriptions of these experiments and results are provided in the point-by-point responses in the “recommendations for the authors” section. Together, these newly added data substantially strengthen the evidence for the Dcp-1/Bruce interaction and support the focus of the original manuscript title.

      Reviewer #2 (Recommendations for the authors):

      (1) The title of the manuscript highlights Dcp-1/Bruce interaction, even though the evidence there is not strong. The evidence for Dcp-1/Sirt1 and Dcp-1/Fkbp59 is stronger. How about changing the title to highlight these other Dcp-1 interactions?

      We thank the reviewer for the thoughtful suggestion. We agree that several Dcp-1-associated factors identified in our study, particularly Sirt1 and Fkbp59, are supported by functional evidence. Specifically, our data show that Sirt1 and Fkbp59 are required for Dcp-1 overexpression-mediated activation. However, Bruce differs from these factors in both the scope and the nature of its effects on Dcp-1. Bruce is not only shown to specifically suppress Dcp-1 activity, but also to suppress wing tissue growth, indicating a broader physiological role in modulating non-lethal Dcp-1 function. Importantly, we further demonstrate that Bruce can specifically physically interact with cleaved Dcp-1. In addition, in this revised manuscript, we show that using the endogenously mStayGold::V5-tag knock-in-tagged Bruce allele, cleaved Dcp-1, induced by overexpression of Dcp-1::VENUS in wing imaginal discs, can be co-immunoprecipitated with full-length Bruce (new Figure 4I, J). These results support a physical interaction between full-length Bruce and activated Dcp-1 in vivo, consistent with a direct inhibitory role. Based on these findings, we decided to retain Bruce in the manuscript title, as it is the only factor for which both physiological and functional interactions with Dcp-1 are supported by multiple independent lines of evidence.

      (2) The case for Dcp-1/Bruce interaction is not strong because the Dcp-1 turboID fails to identify Bruce. In fact, the Dcp-1 turboID approach may not have been effective, as it failed to detect many established interactions, including Diap1 (Wang et al. 1999 PMID 10481910; Tenev et al., 2006 PMID 15580265). The authors may want to comment on this.

      We thank the reviewer for raising this important point. We agree that Bruce, as well as DIAP-1, was not identified in our TurboID-MS labeling dataset (Figure 2C, Table S1). Previous studies have shown that DIAP1 interacts with Dcp-1 and Drice only after exposure of the IAP-binding motif (IBM) at the neo-N-terminus of the large executioner caspase subunit following cleavage (Tenev et al., 2005). Similarly, our co-immunoprecipitation analyses show that Bruce interacts specifically with cleaved Dcp-1, but not with full-length Dcp-1. In the revised manuscript, we confirmed by western blot that the majority of endogenously expressed Dcp-1 in wing imaginal discs is present in the full-length pro-form (new Figure 2 – figure supplement 2A). Thus, the failure to identify Bruce and DIAP1 by TurboID-MS using full-length Dcp-1 as bait is expected, as this approach primarily labels interactors of the inactive, full-length form of Dcp-1. To evaluate whether our proximity labeling approach was nevertheless effective, we compared our TurboIDMS dataset with a previously published immune-affinity purification (IAP)-MS dataset generated using catalytically inactive, C-terminally V5-tagged Dcp-1 overexpressed in Drosophila 1(2)mbn cells (Choutka et al., 2017). Although the experimental conditions differ in several respects, we observed a substantial overlap between the TurboID-MS-mediated and IAP-MS-mediated interaction lists (new Figure 2 – figure supplement 2B, new Table S2). Importantly, SesB, one of the best-characterized Dcp-1 interactors located in mitochondria (DeVorkin et al., 2014), was also identified in our mass spectrometry dataset (new Figure 2 – figure supplement 2B, new Table S2). Based on these analyses, we now more explicitly describe the experimental context and limitations of the TurboID approach, clarifying that it preferentially labels interactors of full-length Dcp-1 in the revised manuscript. We also incorporate comparisons with prior studies to further support the validity of our mass spectrometry experiments in the revised manuscript

      (3) The best experimental evidence for Bruce/Dcp-1 interaction can be found in Figure 4H. But here, they see a weak interaction only when a truncated Bruce construct is overexpressed in S2 cells. Whether Dcp-1 interacts with Bruce in a physiological setting remains unsupported.

      We thank the reviewer for the important comment. We agree that, in the original manuscript, the biochemical evidence for the Bruce/Dcp-1 interaction relied primarily on experiments using an overexpressed truncated Bruce construct in S2 cells and therefore did not sufficiently establish whether this interaction occurs in vivo, especially in wing imaginal discs. To address this concern, we performed additional experiments to examine the Bruce/Dcp-1 interaction. In the background of the mStayGold::V5-tag knocked-in Bruce allele, we overexpressed Dcp-1::VENUS using WPGal4 driver to induce Dcp-1 activation and tested whether full-length Bruce under endogenous expression interacts with cleaved Dcp-1 in wing imaginal discs. Following immunoprecipitation with anti-V5 antibody-conjugated magnetic agarose, we found that cleaved Dcp-1 signal was enriched by co-immunoprecipitation (new Figure 4I, J). These new data demonstrate that Bruce associates with cleaved Dcp-1 in vivo and thus support the physiological relevance of the Bruce/Dcp-1 interaction. We have clarified this point in the revised manuscript and included the corresponding data.

      (4) The genetic interaction between Bruce and Dcp-1 is interesting, but the interpretation becomes complicated because Bruce inhibits Reaper, and at the same time, Dcp-1 genetically interacts with Reaper, Hid, and Grim (Figures 1E, F, G). Thus, it remains unclear if the genetic interaction between Bruce/Dcp-1 is due to a direct interaction between Bruce/Dcp-1 or alternatively, because Bruce inhibits Reaper and Grim.

      We thank the reviewer for the comment. The primary function of Reaper, Hid, and Grim (RHG proteins), collectively referred to as IAP antagonists, is to directly interact with inhibitor of apoptosis proteins (IAPs), most notably DIAP-1 (Kornbluth and White, 2005; Ryoo and Baehrecke, 2010), leading to the inhibition of DIAP-1 function. RHG proteins have not been shown to directly inhibit caspases. Because inhibition of RHG proteins results in the stabilization of DIAP-1, it is likely that the effects observed upon RHG gene knockdown are mediated through DIAP-1. Consistent with this idea, overexpression of DIAP-1, while less potent than Bruce, can also suppress Dcp-1 activation (Figure 5B, C). However, we also acknowledge that Bruce suppresses Reaper- and Grim-dependent, but not Hid-dependent, cell death (Vernooy et al., 2002). In addition, Bruce directly targets Reaper through non-lysine ubiquitination, promoting its degradation (Domingues and Ryoo, 2012). Thus, it is possible that Bruce overexpression suppresses Reaper and thereby strengthens DIAP-1 function, which could indirectly contribute to the inhibition of Dcp-1 activation. Nevertheless, because the effect of Bruce overexpression is stronger than that of DIAP-1 overexpression (Figure 5B, C), and together with our physical interaction data of Bruce with cleaved Dcp-1, we propose that Bruce most likely inhibits Dcp-1 directly to attenuate its activation.

      (5) In general, the manuscript could benefit from highlighting the strong data on Sirt1 and Fkbp59, while clearly acknowledging the limitations of the Bruce/Dcp-1 interaction.

      We thank the reviewer for the comment. As described above, in the revised manuscript we now demonstrate that endogenously expressed full-length Bruce physically interacts with cleaved Dcp-1 in wing imaginal discs (Figure 4I, J). These new data provide strong support for a physiologically relevant interaction between Bruce and cleaved Dcp-1. Based on this evidence, we decided to highlight Bruce in the manuscript, as it is the only factor for which both physiological and functional interactions with Dcp-1 are supported by multiple independent lines of evidence.

      Reviewer #3 (Public review):

      Summary:

      The present paper by Shinoda et al. from the Miura group builds upon findings reported in an earlier study by the same team (Shinoda et al., PNAS, 2019), which identified a nonapoptotic role for the Drosophila executioner caspase Dcp-1 in promoting wing tissue growth. That earlier work attributed this function primarily to Dcp-1 and to Decay, a caspase structurally related to executioner caspases, but not to DrICE, the principal apoptotic executioner caspase. The authors further proposed that this non-apoptotic caspase activity operates independently of the initiator caspase Dronc.

      In the current study, the authors both corroborate aspects of their previous findings and extend the investigation to mechanisms regulating Dcp-1 in this context. They identify roles for the giant IAP Bruce, two BCL-2 family members, and autophagy-related components in modulating nonapoptotic Dcp-1 activity. Moreover, they show that Bruce binds to a BIR-like peptide exposed upon Dcp-1 cleavage, but not to DrICE. The study further suggests that low levels of Dcp-1 activity promote wing tissue growth, whereas excessive activity induces cell death, as evidenced by impaired wing development following Dcp-1 overexpression. Overall, the manuscript provides several intriguing insights into the non-apoptotic regulation of the comparatively weak apoptotic executioner caspase Dcp-1 and complements the group's earlier work. However, several concerns remain regarding certain interpretations of the data and the experimental rigour of some of the results.

      Strengths:

      A major strength of the work is its systematic genetic and biochemical approaches, which combine tissue-specific manipulation with protein interaction mapping to explore how Dcp-1 is regulated. The identification of several regulatory factors, including an inhibitor of cell death protein and components linked to autophagy, provides a coherent framework for understanding how Dcp-1 activity might be tuned.

      Weaknesses:

      The evidence supporting some key claims remains incomplete. In particular, the type of cell death form induced when Dcp-1 is overexpressed is not clearly established, and additional tests would be needed to distinguish between the different cell death types.

      Likely impact:

      The study contributes to a growing body of work showing that proteins traditionally associated with cell death can have broader roles in tissue development. This conceptual advance is likely to be of interest to researchers studying growth control and tissue maintenance.

      We sincerely thank the reviewer for their thoughtful and constructive evaluation of our study. In response to these concerns, we have performed additional experiments to clarify the nature of the cell death induced by Dcp-1 overexpression. Based on the detection of cleaved Dcp-1, the detection of executioner caspase activity, and TUNEL assay, we now conclude that excessive Dcp-1 expression induces typical executioner caspase activity-dependent apoptotic cell death. Detailed explanations and experimental results are provided in the point-by-point responses below. Overall, we believe that these additions strengthen the manuscript by clarifying the dual roles of Dcp-1 in promoting tissue growth at low activity levels while triggering apoptosis when excessively activated.

      Specific points:

      (1) Nature of the wing ablation phenotype

      A central concern is whether the wing ablation phenotype observed upon Dcp-1 overexpression truly reflects apoptotic cell death. The authors show in Figure 1c that nuclei in cells overexpressing Dcp-1, but not DrICE, zymogens are highly condensed, which is suggestive of apoptosis. However, it is equally plausible that this phenotype reflects a form of non-apoptotic, Dcp-1-dependent cell death (e.g. autophagy-dependent cell death). This distinction could be readily addressed using TUNEL labelling and direct caspase activity assays. The latter would be particularly informative, as it remains unclear whether zymogen Dcp-1 is capable of cleaving standard effector caspase reporters in vivo. Does the anti-cleaved Dcp-1 antibody detect Dcp-1 activation following overexpression of the Dcp-1 zymogen?

      We thank the reviewer for this important point regarding the nature of cell death. We agree that nuclear condensation alone is not sufficient to conclude apoptotic cell death, and we therefore performed additional experiments. First, we performed TUNEL staining and detected robust TUNEL-positive signals in wing imaginal discs upon Dcp-1 overexpression (new Figure 1D), supporting apoptotic DNA fragmentation. Second, to directly test whether Dcp-1 overexpression leads to executioner caspase activity in vivo, we used two independent executioner caspase activity probes, GC3Ai (Schott et al., 2017; Zhang et al., 2013) and CD8::PARP::VENUS (Williams et al., 2006). Both probes showed clear executioner caspase activity-positive signals in wing imaginal discs upon Dcp-1 overexpression (new Figure 1 – figure supplement 1C–F), demonstrating that Dcp-1 overexpression leads to executioner caspase activity capable of cleaving standard substrates in vivo. In addition, staining with an anti-cleaved Dcp-1 antibody was positive upon Dcp-1 zymogen overexpression (new Figure 1 – figure supplement 1B), indicating that the overexpressed Dcp-1 zymogen is converted into its active form. Consistent with this result, western blot analysis revealed that Dcp-1 zymogen overexpression results in the appearance of a cleaved Dcp-1 (new Figure 1 – figure supplement 1A). Importantly, consistent with our original observation that the wing ablation phenotype is suppressed by expression of the caspase inhibitor p35, we further showed that p35 overexpression completely abolished the appearance of cleaved Dcp-1 in western blot (new Figure 1 – figure supplement 1A), suggesting that Dcp-1 activation is mediated by self-cleavage. Taken together, these new results demonstrate that Dcp-1 zymogen overexpression induces typical executioner caspase activity-dependent apoptotic cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (2) Role of Decay

      In their earlier study, the authors identified Decay as another caspase influencing wing growth, albeit more modestly than Dcp-1. It is therefore unclear why this line of investigation was not pursued further in the current work. This omission is notable, as Decay is not implicated in apoptosis and, to date, no substantial physiological function has been assigned to this caspase in any system. At a minimum, this point should be discussed explicitly.

      We thank the reviewer for the comment regarding the role of Decay. In our previous study (Shinoda et al., 2019), we demonstrated that both Dcp-1 and Decay promote wing tissue growth in a non-lethal manner. In the present study, however, we focused our analysis on Dcp-1. This decision was based on both technical and biological considerations. From a technical perspective, we had established TurboID knock-in lines and UAS overexpression lines for Dcp-1, Drice, and Dronc, whereas corresponding genetic tools are not available for Decay. From a biological standpoint, Dcp-1 exerts a stronger effect on wing growth than Decay, as shown in our previous work, and exhibits a dual functional spectrum: Dcp-1 promotes tissue growth at low activity levels, whereas excessive activation induces overt cell death. By contrast, Decay has not been implicated in cell death in wing imaginal discs (Kondo et al., 2006). Given that a central aim of the present study was to dissect how executioner caspase activity is differentially regulated to support nonlethal functions versus apoptotic cell death, we therefore focused on the two executioner caspases that are known to participate in apoptosis, Dcp-1 and Drice. We agree with the reviewer that Decay remains an intriguing caspase with largely unexplored physiological roles, and further investigation into its regulation and function will be an important direction for future studies. Importantly, Decay has been shown to mediate Hid-induced cell death in the DIAP1- and apoptosome-independent manner in differentiating photoreceptors and accessory cells of the eye (Leulier et al., 2006). In addition, although not required for cell death, Decay accounts for most of the caspase activity during metamorphic midgut programmed cell death, which is executed by autophagy (Denton et al., 2009). Thus, similar to Dcp-1, Decay might be an executioner caspase that can be regulated independently of the canonical apoptosome-mediated pathway, potentially involving autophagy-Bruce axis, and thereby contributing to the regulation of tissue growth. We have now discussed this point in the revised manuscript.

      (3) Figure 2: Proximity labelling analysis

      The authors use TurboID-mediated proximity labelling to reveal distinct Dcp-1- and DrICEassociated proteomes across tissues, with a particular focus on the wing disc. They further demonstrate that RNAi-mediated knockdown of the Dcp-1-associated proteins Sirt1 and Fkbp59 suppresses the wing ablation phenotype induced by Dcp-1 overexpression, suggesting that these factors are required for Dcp-1 activity. However, it should be clarified whether Bruce was identified as a Dcp-1 interactor in the proximity labelling dataset, given its proposed central regulatory role. In addition, further discussion of Fkbp59, its known functions and how it might mechanistically influence Dcp-1 activity would be valuable.

      We thank the reviewer for the comment regarding the TurboID-based proximity labeling analysis and the interpretation of the identified Dcp-1-associated factors. With respect to Bruce, we clarify that Bruce was not identified as a Dcp-1 interactor in the TurboID proximity labeling dataset. Our co-immunoprecipitation analyses in S2 cells indicate that Bruce interacts specifically with cleaved Dcp-1, but not with the full-length, inactive form. In the revised manuscript, we confirmed by western blot that the majority of endogenously expressed Dcp-1 in wing imaginal discs exists in the full-length pro-form (new Figure 2 – figure supplement 2A). Therefore, the failure to detect Bruce in the TurboID experiment using full-length Dcp-1 as bait is expected, as this approach primarily labels proteins proximal to the inactive form of Dcp-1. To examine the Bruce/Dcp-1 interaction under more physiological conditions, we performed additional in vivo experiments. Using the mStayGold::V5-tag knock-in allele of Bruce, we overexpressed Dcp1::VENUS using WP-Gal4 driver to induce Dcp-1 activation and assessed whether endogenously expressed full-length Bruce associates with Dcp-1 in wing imaginal discs. Following immunoprecipitation with anti-V5 antibody-conjugated magnetic agarose, we found that cleaved Dcp-1 signal was enriched by co-immunoprecipitation (new Figure 4I, J). These new data demonstrate that Bruce associates selectively with the cleaved, active form of Dcp-1 in vivo, thereby supporting the physiological relevance of the Bruce/Dcp-1 interaction. We have clarified this point in the revised manuscript and included the corresponding data.

      FK506-binding proteins (FKBPs) are a conserved group of proteins known to bind FK506, an immunosuppressive drug. FKBPs contain FK domains, which correspond to peptidyl cis-trans isomerase (PPIase) domains. Drosophila Fkbp59 is an orthologue of the mammalian FKBP4 and FKBP5, both of which possess a C-terminal tetratricopeptide repeat (TPR) domain that functions independently of the PPIase domain by mediating protein-protein interactions. The mammalian orthologues of Drosophila Fkbp59 function as Hsp90 co-chaperones (GharteyKwansah et al., 2018). Importantly, loss of Fkbp59 results in pupal lethality (Iki et al., 2020), which precludes further mechanistic analysis on Dcp-1 activation using adult wing phenotypes. To date, the involvement of Fkbp59 in caspase regulation has not been reported. Given that Fkbp59 functions as a co-chaperone, it may facilitate Dcp-1 activation by promoting proper folding, stability, or subcellular positioning of Dcp-1 or its regulatory factors. Importantly, Dcp1 proximal proteins are enriched in chaperone-related factors, including CCT2, CCT8, Droj2, CG16817, Fkbp59, Sgt1, and nudC; seven out of sixteen identified proximal proteins are chaperone-related. These observations suggest that Dcp-1 activity may be regulated by chaperone proteins or that Dcp-1 activity may be spatially restricted to regions enriched in chaperone machinery. Further analysis of the relationship between Dcp-1 activity and chaperone-related proteins will be important to elucidate the mechanisms and functions underlying non-lethal Dcp1 activation.

      (4) Figure 3: Autophagy-related factors

      Given that Sirt1 is known to promote autophagy, the authors next examine autophagy-related proteins and identify roles for Atg2, Atg8a, Debcl, and Buffy in Dcp-1 activation. Notably, these proteins do not promote cell death in the Hid-induced canonical apoptotic pathway. However, it is important to determine whether knockdown of Debcl, Buffy, Atg2, or Atg8a alone affects wing development in the absence of Dcp-1 overexpression, to exclude the possibility that these perturbations independently impair wing formation.

      We thank the reviewer for the comment. To address whether knockdown of Debcl, Buffy, Atg2, or Atg8a independently affects wing development, we performed RNAi-mediated knockdown of each gene using the WP-Gal4 driver in the absence of Dcp-1 overexpression. Under these conditions, knockdown of Debcl, Buffy, Atg2, or Atg8a did not cause any detectable defects in wing morphology (new Figure 3 – figure supplement 1A), indicating that these autophagy-related factors specifically function to suppress Dcp-1-mediated cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (5) Evidence for canonical autophagy

      The involvement of autophagy would be more convincingly demonstrated by testing additional core autophagy genes, such as Atg7, Atg5, and Atg12, as well as performing a combined knockdown of Atg8a and Atg8b. Moreover, direct assessment of autophagy at the cellular level using established genetic reporters would substantially strengthen the conclusions.

      We thank the reviewer for the constructive comment regarding the involvement of canonical autophagy. To further strengthen the evidence that autophagy is required for Dcp-1 activation, we examined additional core autophagy-related genes that function at distinct steps of the autophagy process, in addition to the previously tested Atg2, which mediates autophagosomal membrane expansion, and Atg8a, a core component directly associated with autophagosomal membranes. Specifically, we performed knockdown of genes including FIP200/Atg17, which is required for the initiation of autophagosome formation; Atg9, which is required for autophagosomal membrane nucleation; Atg5, which is required for autophagosomal membrane expansion through Atg12-Atg5-Atg16 ubiquitin-like conjugation system; and Stx17, which is required for autophagosome-lysosome fusion (Umargamwala et al., 2024). Because Atg8b is known to be specifically expressed in the male germline and is dispensable for autophagy, at least in fat body cells (Jipa et al., 2021), we did not further examine Atg8b in wing imaginal discs. Using WPGal4 driver, knockdown of each of these genes significantly suppressed Dcp-1-induced wing ablation phenotype (new Figure 3 – figure supplement 1C), supporting a requirement for canonical autophagy components across multiple stages of autophagosome biogenesis in Dcp-1 activation. Importantly, knockdown of these autophagy-related genes alone did not affect wing morphology in the absence of Dcp-1 overexpression (new Figure 3 – figure supplement 1B), as observed previously for Atg2 and Atg8a, suggesting the suppressive effects are specific to Dcp-1 overexpression-dependent cell death. Together, these results indicate that inhibition of autophagy at any of several key steps can suppress Dcp-1-dependent cell death, demonstrating that intact canonical autophagy is required for Dcp-1 activation. In addition, to directly assess autophagy at the cellular level, we monitored autophagosome formation using mCherry::Atg8a reporter. Upon overexpression of Dcp-1::VENUS in the wing pouch region, we observed a clear accumulation of Atg8a-positive puncta in wing imaginal discs (new Figure 3C), demonstrating that Dcp-1 overexpression induces autophagy in vivo. Together, these results provide both genetic and cellular evidence that canonical autophagy is activated upon Dcp-1 overexpression and is required for Dcp-1-dependent cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (6) Figures 4-5: Functional consequences

      It would be informative to determine whether Synr, Debcl, or Buffy influence wing size on their own and whether their overexpression enhances wing growth.

      We thank the reviewer for the suggestion regarding the functional consequences of Synr, Debcl, and Buffy on wing size. As requested, we knocked down Debcl or Buffy using WP-Gal4 driver and found that this led to reduced wing size (new Figure 5 – figure supplement 1A), indicating that endogenous Debcl and Buffy promote wing growth potentially through regulating endogenous Dcp-1 activity. We have included the corresponding data in the revised manuscript. Because Synr RNAi did not show any detectable effect on the Dcp-1 overexpression-induced phenotype (Figure 3A, B), we did not further examine the effect of Synr knockdown on wing development alone. Overexpression of Synr was not examined in this study. However, Synr overexpression has previously been reported to induce cell death in wing imaginal discs, resulting in malformed adult wings (Ikegawa et al., 2023), suggesting that increased Synr expression is likely to have deleterious rather than growth-promoting effects. Because Debcl and Buffy are both required for Synr-induced cell death, overexpression of Debcl or Buffy may lead to similar phenotypes. Therefore, we did not test Debcl or Buffy overexpression in the wing imaginal discs.

      (7) Terminology and interpretation of cell death

      Taken together, the results suggest that Dcp-1 zymogen overexpression induces a form of nonapoptotic cell death, potentially autophagy-dependent or related. The reviewer does not understand the authors' insistence on referring to this process as apoptosis. The authors should be more cautious in their terminology: there is no canonical versus non-canonical apoptosis; there is simply apoptosis. Without stronger evidence, these effects should not be described as apoptotic cell death.

      We thank the reviewer for the important comment on terminology and interpretation of the cell death phenotype. As explained in our response to comment #1, we have performed additional experiments to clarify the nature of the cell death induced by Dcp-1 overexpression. Based on the detection of cleaved Dcp-1, the detection of executioner caspase activity, and TUNEL assay, we now conclude that excessive Dcp-1 expression induces typical executioner caspase activity-dependent apoptotic cell death. At the same time, as explained in our response to comment #5, we provide both genetic and cellular evidence that canonical autophagy is activated upon Dcp-1 overexpression and promotes Dcp-1 activation. We recognized that the phrase “Dcp1 activity-regulating alternative apoptosis signaling pathway” used in Figure 5L could be misleading, as it may imply the existence of an “alternative apoptosis”. To avoid this confusion, we have revised the figure legend to read “autophagy-facilitated alternative caspase activation pathway.”

      Reviewer #3 (Recommendations for the authors):

      Figure 1c should be annotated more clearly so that it is evident that the images shown are grouped by genotype.

      We thank the reviewer for the helpful suggestion. We have added lines to Figure 1C to improve clarity by indicating that the images are grouped by genotype.

      References

      Choutka C, DeVorkin L, Go NE, Hou Y-CC, Moradian A, Morin GB, Gorski SM. 2017. Hsp83 loss suppresses proteasomal activity resulting in an upregulation of caspase-dependent compensatory autophagy. Autophagy 13:1573–1589.

      Denton D, Shravage B, Simin R, Mills K, Berry DL, Baehrecke EH, Kumar S. 2009. Autophagy, not apoptosis, is essential for midgut cell death in Drosophila. Curr Biol 19:1741–1746.

      DeVorkin L, Go NE, Hou Y-CC, Moradian A, Morin GB, Gorski SM. 2014. The Drosophila effector caspase Dcp-1 regulates mitochondrial dynamics and autophagic flux via SesB. J Cell Biol 205:477–492.

      Domingues C, Ryoo HD. 2012. Drosophila BRUCE inhibits apoptosis through non-lysine ubiquitination of the IAP-antagonist REAPER. Cell Death Differ 19:470–477.

      Ghartey-Kwansah G, Li Z, Feng R, Wang L, Zhou X, Chen FZ, Xu MM, Jones O, Mu Y, Chen S, Bryant J, Isaacs WB, Ma J, Xu X. 2018. Comparative analysis of FKBP family protein: evaluation, structure, and function in mammals and Drosophila melanogaster. BMC Dev Biol 18:7.

      Ikegawa Y, Combet C, Groussin M, Navratil V, Safar-Remali S, Shiota T, Aouacheria A, Yoo SK. 2023. Evidence for existence of an apoptosis-inducing BH3-only protein, sayonara, in Drosophila. EMBO J 42:e110454.

      Iki T, Takami M, Kai T. 2020. Modulation of Ago2 loading by Cyclophilin 40 endows a unique repertoire of functional miRNAs during sperm maturation in Drosophila. Cell Rep 33:108380.

      Jipa A, Vedelek V, Merényi Z, Ürmösi A, Takáts S, Kovács AL, Horváth GV, Sinka R, Juhász G. 2021. Analysis of Drosophila Atg8 proteins reveals multiple lipidation-independent roles. Autophagy 17:2565–2575.

      Kondo S, Senoo-Matsuda N, Hiromi Y, Miura M. 2006. DRONC coordinates cell death and compensatory proliferation. Mol Cell Biol 26:7258–7268.

      Kornbluth S, White K. 2005. Apoptosis in Drosophila: neither fish nor fowl (nor man, nor worm). J Cell Sci 118:1779–1787.

      Leulier F, Ribeiro PS, Palmer E, Tenev T, Takahashi K, Robertson D, Zachariou A, Pichaud F, Ueda R, Meier P. 2006. Systematic in vivo RNAi analysis of putative components of the Drosophila cell death machinery. Cell Death Differ 13:1663–1674.

      Ryoo HD, Baehrecke EH. 2010. Distinct death mechanisms in Drosophila development. Curr Opin Cell Biol 22:889–895.

      Schott S, Ambrosini A, Barbaste A, Benassayag C, Gracia M, Proag A, Rayer M, Monier B, Suzanne M. 2017. A fluorescent toolkit for spatiotemporal tracking of apoptotic cells in living Drosophila tissues. Development 144:3840–3846.

      Shinoda N, Hanawa N, Chihara T, Koto A, Miura M. 2019. Dronc-independent basal executioner caspase activity sustains Drosophila imaginal tissue growth. Proc Natl Acad Sci U S A 116:20539–20544.

      Tenev T, Zachariou A, Wilson R, Ditzel M, Meier P. 2005. IAPs are functionally non-equivalent and regulate effector caspases through distinct mechanisms. Nat Cell Biol 7:70–77.

      Umargamwala R, Manning J, Dorstyn L, Denton D, Kumar S. 2024. Understanding developmental cell death using Drosophila as a model system. Cells 13:347.

      Vernooy SY, Chow V, Su J, Verbrugghe K, Yang J, Cole S, Olson MR, Hay BA. 2002. Drosophila Bruce can potently suppress Rpr- and Grim-dependent but not Hid-dependent cell death. Curr Biol 12:1164–1168.

      Williams DW, Kondo S, Krzyzanowska A, Hiromi Y, Truman JW. 2006. Local caspase activity directs engulfment of dendrites during pruning. Nat Neurosci 9:1234–1236.

      Zhang J, Wang X, Cui W, Wang W, Zhang H, Liu L, Zhang Z, Li Z, Ying G, Zhang N, Li B. 2013. Visualization of caspase-3-like activity in cells using a genetically encoded fluorescent biosensor activated by protein cleavage. Nat Commun 4:2157.

    1. eLife Assessment

      The authors describe a new member of the KCNE auxiliary subunits of potassium channels from a lamprey. This new subunit represents an early evolutionary member which confers new properties when expressed along with KCNQ channels. In the revised version of the manuscript, the authors present convincing evidence from several experimental approaches. The contents of this manuscript are important and should be relevant to understanding both the mechanism of modulation of KCNQ channels by KCNE subunits and the evolutionary history of these subunits, which this manuscript now extends to the divergence of early vertebrates.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

    3. Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

      Thank you for reviewing our manuscript and for your constructive comments. Our point-to-point responses are shown below.

      Weaknesses:

      (1) No error bars visible for lamprey Q1 isoforms (open symbols) in Figure 2G. No statistical comparison was provided to indicate whether lamprey Q1 isoform V1/2s are significantly different (nor in Supplementary Table 1).

      (2) There is the same issue in Figures 3 and 4. No appropriate statistical comparison is made between V1/2s for different truncations of PmKCNE0 (Figure 3), or between KCNQ1 species isoforms with and without PmE0.

      We thank you for these helpful comments. Based on your suggestions, we revised the presentation of error bars in Fig. 2G and in other panels showing G–V or F–V relationships (Figs. 2J, 3F, 3N, 4C, 4F, 4I, 4L, and 5E; Supplementary Fig. 5D) to make the SEM bars clearer. We also added statistical comparisons of V<sub>1/2</sub> values among the three lamprey KCNQ1 orthologs in Fig. 2G and among truncation-series constructs in Figs. 3F and 3N using one-way ANOVA followed by Tukey–Kramer multiple-comparison tests. For Fig. 4, we added statistical comparisons between KCNQ1 species isoforms expressed with or without PmKCNE0 (Figs. 4C, 4F, and 4I), and between PmKCNQ1 expressed alone or with human KCNE1 or KCNE3 (Fig. 4L), using unpaired two-tailed Welch’s t-tests. These statistical comparisons are included in Supplementary Table 1.

      Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

      We thank you for the positive assessment of our work and for the constructive suggestions. Our point-by-point responses are provided below.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What is the physiological role of this KCNE0 and lamprey KCNQ1 in the lamprey species? While the authors mention that the physiological roles of KCNE0 are the next focus, it is preferable to discuss some of the potential functional significance of this newly characterised KCNE.

      We thank you for this helpful suggestion. We agree that discussing the potential physiological roles of KCNQ1–KCNE0 complexes in lamprey strengthens the manuscript. We have therefore expanded the Discussion to raise the possibility that, given the broad tissue distribution of kcne0 transcripts and the ability of KCNE0 to render lamprey KCNQ1 constitutively active, KCNQ1–KCNE0 complexes may contribute to general ion homeostasis, potentially analogous to the epithelial K<sup>+</sup> recycling function of mammalian KCNQ1–KCNE3, rather than to the highly specialized KCNQ1–KCNE1 function in the mammalian heart and inner ear (page 14, lines 263–267).

      (2) Human KCNQ1 has 676 amino acids, but LcKCNQ1 contains just 507 amino acids. The species-specific regulatory effects of KCNE0 may not only be attributed to KCNE0 itself but might also be influenced by the species of KCNQ1. Some discussion on this possibility will be helpful.

      We thank you for raising this important point. We agree that the species-specific regulatory effects observed in our cross-species pairing experiments are unlikely to be determined by KCNE0 alone and may also be influenced by species-specific features of the KCNQ1 α-subunit. To address this point, we expanded the Discussion to note that the KCNQ1 proteins used in this study vary in amino-acid length, largely reflecting differences in the cytoplasmic C-terminal region, which may affect KCNQ1–KCNE compatibility and thereby influence channel gating and coupling to KCNE subunits (pages 12–13, lines 227–238). We also updated Supplementary Fig. 4 to include LrKCNQ1 and LcKCNQ1, and clarified that the LcKCNQ1 construct used in this study encodes 644 amino acids.

      (3) In Supplementary Figure 3, bands corresponding to LcKCNQ1 (507 amino acids, Supplementary Figure 1) were not seen.

      We thank you for pointing out this potentially confusing point. Supplementary Fig. 3 shows RT-PCR products amplified from tissue cDNA, not full-length amplification of the LcKCNQ1 ORF. As stated in the Methods section (pages 19–20, lines 374–399), the primers used for RT-PCR in Fig. 1F and Supplementary Fig. 3 were different from those used for cloning the full-length LcKCNQ1 cDNA shown in Supplementary Fig. 1. Therefore, a band corresponding to the full-length LcKCNQ1 coding sequence was not expected in Supplementary Fig. 3.

      To clarify this point, we revised the figure legends and indicated the RT-PCR primer-binding sites with orange arrows in Supplementary Figs. 1 and 2.

    1. eLife Assessment

      This useful study addresses a timely question about semantic prioritisation in visual working memory, using behavioural manipulations and drift-diffusion modelling. However, the strength of evidence is incomplete for the broader claims about working-memory representations because the main interpretation relies on indirect inferences from non-decision time, which cannot uniquely identify memory access or retrieval.

    2. Reviewer #1 (Public review):

      Summary:

      This paper investigates whether semantic prioritization in visual working memory reflects pre-decisional access, evidence accumulation, or both, using drift diffusion modeling across a reanalysis of prior data and two new experiments. The core finding - that semantic information receives a robust pre-decisional access advantage that is amplified by attentional disruption rather than temporal delay alone - is novel and contributes meaningfully to ongoing debates about the format and accessibility of working memory representations.

      Strengths:

      The experimental approach is well-motivated, and the use of drift-diffusion modeling to decompose decision components adds analytical value beyond standard RT and accuracy measures. The two new experiments are pre-registered and address important questions. The broader theoretical conclusion - that working memory limits are shaped not only by storage capacity but by which representational formats remain accessible under attentional uncertainty - is an important and timely contribution to the field.

      Weaknesses:

      The central interpretive claims rely heavily on differences in non-decision time, a parameter that aggregates many processes unrelated to memory retrieval, making it rather difficult to uniquely attribute the observed effects to access or retrieval mechanisms specifically. Additionally, the characterization of the two memory conditions as genuinely perceptual versus semantic warrants further justification, as both may primarily require categorical rather than format-specific knowledge.

    3. Reviewer #2 (Public review):

      This manuscript aims to characterize how semantic information is prioritized relative to perceptual details in visual working memory. The central claim is that semantic judgements benefit from faster pre‑decisional access (shorter non‑decision time), and that advantages in evidence accumulation emerge under higher cognitive demands (e.g., when items are outside the focus of attention or must be maintained under interference). Based on this, the paper argues that unattended working‑memory contents are reformatted into more abstract, long‑term‑memory‑like semantic representations that remain more readily accessible than fine‑grained perceptual features.

      Strengths:

      (1) The question is timely and relevant to current research about the format of visual working memory.

      (2) Behaviorally, the semantic advantage is carefully documented in many conditions across datasets.

      (3) The use of hierarchical drift-diffusion modelling is helpful to decompose the semantic advantage into cognitive processes such as non‑decision time and drift‑rate components.

      Weaknesses:

      (1) The strong claims about visual working‑memory representation and "long‑term‑memory‑like" formats rest on an indirect inference from decision‑model parameters to representational content, and this link is not convincingly established. Non‑decision time, as implemented here, bundles many things, such as probe processing, cue processing, retrieval/access, and motor preparation, so reduced non‑decision time for semantic probes could reflect easier question reading, simpler response mapping, or more efficient decision preparation rather than a genuine advantage in accessing semantic memory representations. Although the manuscript acknowledges that non‑decision time includes multiple processes, it nonetheless treats this parameter as primary evidence for a retrieval‑stage semantic advantage, which overstates what the data can uniquely support.

      (2) The modelling approach is relatively constrained and does not fully address the underdetermination inherent in mapping latent drift-diffusion parameters onto specific psychological mechanisms. The preferred model that allows multiple parameters (non‑decision time, drift rate, threshold) to vary provides only modest improvements in predictive accuracy over simpler models, and several key drift‑rate effects are present only in particular load or lag conditions. As a result, the theoretical interpretation that semantic prioritization primarily reflects faster access and secondarily more efficient accumulation under high demand appears rather post hoc, and alternative accounts focused on generic task efficiency or strategy differences remain plausible.

      (3) The operationalization of "semantic" is narrow and largely categorical, focusing on animacy (animal/object) and a perceptual format dimension (photo/drawing), rather than richer semantic or associative relations among items. This makes it difficult to generalize the conclusions to broader claims about semantic structure and its integration into working‑memory representations. Important recent work on how semantic and associative relationships facilitate the formation, maintenance, and retrieval of visual working memory is not adequately integrated into the theoretical framing. Consequently, the discussion tends to generalize from a specific probe structure to a broader semantic prioritization theory without engaging fully with the existing literature on semantic facilitation and neural decoding of working‑memory content.

      (4) The paper contrasts its behavioral/model‑based results with prior neural decoding findings, but the comparison is not fair. Neural decoding provides complementary evidence about the content and format of working memory representations, whereas drift-diffusion parameters reflect downstream decision dynamics given a probe. Because the current work does not include any direct representational or neural measure, its conclusions about representational "reformatting" and long‑term‑memory‑like access remain speculative and, in places, feel like a stretch.

      (5) Overall, while the data show a semantic advantage in decision‑stage measures and the modelling provides an informative decomposition of this advantage, the manuscript does not fully achieve its stated aim of characterizing the representational format of visual working memory or demonstrating a mechanistic shift toward long‑term‑memory‑like semantic representations. The work primarily informs decision‑process analyses of the conditions under which semantic judgements are faster and more robust, rather than the nature of visual working‑memory representations themselves.

    4. Author response:

      We thank the reviewers for their constructive and careful assessment of our manuscript. We are encouraged that both reviewers recognised the value of the empirical contribution: the semantic advantage is robust across experiments, the two new experiments are pre-registered, and the drift-diffusion modelling provides an informative decomposition of behavioural performance. At the same time, both reviewers raise an important and convergent point: the manuscript currently places too much interpretive weight on non-decision time and sometimes moves too quickly from decision-model parameters to claims about the representational format of working memory.

      We agree that this aspect of the manuscript should be revised. In the next version, we will substantially soften claims about adaptive reformatting and long-term-memory-like formats. We will instead frame the central contribution more precisely: semantic-category judgements show a reliable advantage at stages preceding evidence accumulation and this advantage is modulated by attentional prioritisation and interference during the maintenance interval. Our data constrain the dynamics with which different kinds of information become available for WM-guided decisions, but they do not, on their own, provide a direct measure of representational format. This hypothesis should be tested in future experiments.

      At the same time, we think the data provide stronger constraints on alternative explanations than the current manuscript makes clear. The reviewers correctly note that non-decision time is not a pure retrieval parameter, as we also note in the discussion. It can include probe encoding, response preparation, motor execution, and other processes. We will therefore avoid more explicitly equating NDT directly with retrieval latency. However, many of the alternatives raised by the reviewers, such as easier question reading or simpler response mapping for semantic probes, predict a relatively fixed semantic–perceptual offset. In our experiments, the probes and response mappings are held constant across attentional conditions, while the semantic NDT advantage changes as a function of whether the relevant item can be prioritised in advance or must be selected/reactivated at test. We will restructure the Results and Discussion to make these condition × feature interactions central to the argument.

      We will also clarify the logic of Experiment 1. We agree with Reviewer 1 that a valid retro-cue likely triggers retrieval or reactivation of the cued item. Our original phrasing, which described the valid-cue condition as reducing retrieval demands, was imprecise. The critical manipulation is better described as shifting item prioritisation/retrieval earlier in the trial. Under valid cueing, the relevant item can be prioritised before the probe appears, whereas under neutral cueing, item selection and access must occur after probe onset. We will rewrite this section accordingly.

      We will also clarify our operationalisation of semantic and perceptual categories. The present contrast is specifically between semantic category information (animate versus inanimate) and perceptual-format information (photograph versus drawing). We agree that the perceptual judgement is still categorical and does not measure fine-grained perceptual fidelity. We will therefore avoid broad claims about semantic structure or perceptual detail in general. However, as pointed in the manuscript, we believe the contrast remains meaningful: the two dimensions are orthogonal within the same stimuli, and previous work using the same feature space showed the opposite ordering during perception (Linde-Domingo et al., 2019), where perceptual-format information was available before semantic-category information. We will move this argument earlier in the manuscript and present it as converging evidence for dissociable access dynamics, while acknowledging that it does not by itself prove representational format.

      In response to the modelling concerns, we will expand the model-validation section. Specifically, we plan to add posterior predictive checks for the reported models, report model comparisons more transparently, clarify when more complex models do or do not provide practically meaningful improvements, and include sensitivity analyses using alternative parameterisations where identifiable.

      We will also make several methodological clarifications. First, because the reanalysis of Kerrén et al. (2022) forms a substantial part of the manuscript, we will add a fuller description of the original task in the main text, including how the probed item was indicated at test. Second, we will rewrite the unclear sentence describing pseudo-random stimulus selection in Experiment 1 and add a control analysis testing whether performance differs when the probed item belongs to the majority versus minority category within the trial. Third, we will clarify the stimulus repetition scheme and discuss possible long-term-memory contributions. Importantly, because semantic and perceptual probes are applied to the same items from the same trials, any repetition history or proactive-interference contribution is shared across the two probe types, although we agree that this should be discussed explicitly.

      Finally, we will revise the broader theoretical framing. We will remove or substantially qualify claims linking the present data directly to episodic memory and imagery. We will also integrate the recent literature suggested by Reviewer 2 on semantic structure, associative relations, long-term-memory contributions to working memory, and boundary conditions for semantic labelling effects. This will allow us to position the study as one piece of a broader literature on how semantic information influences WM performance, rather than as direct evidence for a general representational reformatting mechanism.

      In summary, the revised manuscript will make a narrower but stronger claim: semantic-category information shows a robust pre-accumulation advantage during WM-guided decisions, and this advantage is shaped by attentional prioritisation and interference during maintenance. We will present this as evidence about WM access dynamics and decision components, not as direct evidence that WM representations are transformed into long-term-memory-like formats.

    1. eLife Assessment

      The manuscript by von Velsen et al. offers valuable structural insights into the mitogen-activated protein kinase (MAPK) pathway by providing cryo-EM structures of stabilized MEK1-ERK2 kinase-substrate complexes in inactive, active, and nucleotide-free states, complemented by HDX-MS, SAXS, ITC, crystallography, and molecular dynamics. The work provides solid evidence for the overall architecture of the complex and identifies interaction sites that help explain MAPK pathway specificity. However, some mechanistic conclusions are not yet fully supported, particularly the designation of one state as an active phosphoryl-transfer configuration, the claim that substrate binding releases the MEK1 catalytic machinery, the proposed link to processive phosphorylation, and the extrapolation to disease-associated mutations.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes three conformers derived from a complex between ERK2-T185V, a variant of MEK1-DD with the KIM sequence replaced by the KIM from the p38 activator, GRA24, ADP, and AlF4-. The goal was to try to capture the complex in its active state. The results show contacts between the kinases between their N-lobe and their C-lobes that resemble MKK6-p38 complexes previously reported by the authors. Two MEK1-ERK2 conformers (States 1,3) are deemed inactive based on the lack of access of ERK-Y187 to the MEK1 active site, and the absence of ADP bound to MEK1 in State 3, while one conformer (State 2) is deemed active, but not fully active due to disorder in MEK1 activation loop (A-loop) and an essential salt bridge between strand beta3 and helix aC. HDX-MS and SAXS solution measurements and all-atom MD simulations are used to model the mutant complex and variants with WT ERK2. The study concludes that substrate recognition involves low-energy contacts with MEK, allowing substantial protein flexibility within the complex in a manner that may accommodate processive phosphorylation of ERK2.

      Strengths:

      The strengths of the work are that the findings provide important structural insights for MEK-ERK signaling and protein phosphorylation in general. These are valuable given that atomic resolution structures of kinase-substrate complexes are still limited in number. The authors succeeded in showing key contacts between subunits and conformational variations within the complex.

      Weaknesses:

      Weaknesses were that some of the conclusions about activity state, dynamics, and effects of ligand binding were less convincing. For example, that State 2 truly represents an active configuration seemed ambiguous, given the absence of Mg2+ and AlF4- in the cryoEM structure and disorder in the activation loop and the K97-E114 salt bridge. Conclusions by SAXS that ADP-AlF4 binding increases active site compaction while increasing local flexibility were not rigorously supported by HDX data, given that the latter were performed without ligand. Sections of the narrative and figures throughout were often confusing, and many assertions were made without clear explanation. Data shown in the supplementary materials were not always described in the Results, even those important for the conclusions. Figure legends and text lacked clear descriptions of specific complexes analyzed. Substantial changes are recommended to improve the readability and clarity of the work.

    3. Reviewer #2 (Public review):

      Summary:

      The authors used Cryo-EM to obtain a complex between MEK1 and ERK2. They used the same method as previously used by the same authors to form a stable complex between MKK6 and p38, an extra-strong KIM replacing the wild-type KIM in MEK1. Three conformers were resolved, with the highest resolution of 3.0 Å. The multiple conformers indicate more flexibility in MEK2 than in ERK2. These data suggest that nucleotide exchange is possible while maintaining MEK1-ERK2 interactions. SAXS and HDX data reinforce the idea of flexibility. They point to interactions between the two N-terminal domains between histidines at the N-terminus of helix C and between the G helices that are maintained in each of the 3 conformers, and sequence and structure suggest these histidines may be a source of specificity in MEK1-ERK2 versus MKK6-p38 interactions. A 2.2 Å structure of a complex between ERK1 (88% identical to ERK2) and the docking peptide used was also presented. Molecular dynamics simulations suggest that the MEK1-ERK2 complex can assume a fully active configuration of MEK1.

      Strengths:

      This is the first structure of a MEK1-ERK2 complex. The structural data are valuable additions to our understanding of MAP2K-MAPK interactions. The discussion points offered in the results section are palatable. These include the origins of specificity and the idea of flexibility in the MAP2K in support of a processive mechanism for the dual phosphorylation activity of MAP2Ks.

      Weaknesses:

      (1) This reviewer considers that the abstract is overstated. Specifically, this paper does not reveal the molecular details of phosphoryl transfer, nor does it demonstrate that substrate binding releases the catalytic machinery.

      (2) The discussion is in some places not supported by evidence and in others has superfluous text. Examples follow:

      - "Once the αG-helix is docked, and the C-lobe histidine triad is in place, the N-lobe interactions must then be fulfilled." The data in this paper does not suggest an order of events.<br /> - "If the substrate MAPK is incorrect, the N-lobe interaction will not be stabilised, preventing alignment of the MAPK A-loop with the MAP2K active site." This statement could be described as obvious.

      (3) Much of the discussion is embedded in the results, such that it is difficult to separate new facts offered by the paper from speculation.

    4. Author response:

      eLife Assessment

      The manuscript by von Velsen et al. offers valuable structural insights into the mitogen-activated protein kinase (MAPK) pathway by providing cryo-EM structures of stabilized MEK1-ERK2 kinase-substrate complexes in inactive, active, and nucleotide-free states, complemented by HDX-MS, SAXS, ITC, crystallography, and molecular dynamics. The work provides solid evidence for the overall architecture of the complex and identifies interaction sites that help explain MAPK pathway specificity. However, some mechanistic conclusions are not yet fully supported, particularly the designation of one state as an active phosphoryl-transfer configuration, the claim that substrate binding releases the MEK1 catalytic machinery, the proposed link to processive phosphorylation, and the extrapolation to disease-associated mutations.

      We would like to counter the final statement. We were very careful in our description of state 2, while we describe it as ‘active’ we clearly explain that the resolution of the reconstruction is not sufficient to define all the classical indicators of a kinase active state; however, the map is consistent with the active conformation, the complex is active in vitro, and the MEK1 variant used is the well known DD mutant that is constitutively active. While the A-loop is not observed, this is in agreement with many crystal structures of other DD mutants. We therefore decided to define this state as ‘active’ as the A-loop of ERK2 approaches the active site, the alpha-C helix has moved in and the A-loop of MEK1 no longer occludes the active site - to clarify the state we refer to the classically active confirmation as ‘fully active’. Our supporting data also show that the complex is highly dynamic during turnover, meaning we have captured MEK1 in a number of conformations on the landscape of an active state – we feel that rather than a limitation, this is an important observation in MAP2K studies. Finally, the determination of an 80 kDa complex by cryoEM to resolutions well below 4 Å is a huge technical achievement allowing the first snapshots of the MEK1-ERK2 complex to be visualised.

      Regarding the A-helix release – our observation is that the helix becomes less folded on binding of substrate. There are many studies, which we cite, that show that destabilising this helix leads to release of the catalytic machinery, see Mansour et al, 1996, Biochemistry, 35, 15529-15536 and Jindal et al 2017 J. Biol. Chem. 292, 18814-18820 for initial studies. Our observation shows that this is linked to substrate binding – a very relevant new insight that demonstrates the importance of this helix, in addition to many previous studies, but links unfolding to substrate recognition for the first time.

      For the mechanism of processive phosphorylation – it has been well established that both processive and distributive mechanisms exist. While the way that a distributive mechanism could work is obvious (complete dissociation of the two proteins), it has not been clear how a MAP2K can remain bound to its substrate and exchange nucleotides. While caution should be employed in interpreting our state 3 structure, it clearly shows what nucleotide exchange when bound to substrate can look like and that this low nucleotide affinity state is linked to disorder in the P-loop, the A-helix and substrate binding via the KIM. We would love to perform experiments that could demonstrate this but cannot at present think of an appropriate method – the reviewers did not suggest a route either.

      Finally, for the cancer-causing mutations – there are many studies demonstrating that the mutations lead to a destabilisation of the A-helix. Our study links this to substrate recognition. While this is inference, it seems justified to describe a link between substrate recognition, A-helix unfolding and disease mutations given the large body of literature describing these events.

      We are currently performing a series of in-cell activity assays that should strengthen our claims regarding the A-helix and other observations in the structure - the histidine interactions in particular.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes three conformers derived from a complex between ERK2-T185V, a variant of MEK1-DD with the KIM sequence replaced by the KIM from the p38 activator, GRA24, ADP, and AlF4-. The goal was to try to capture the complex in its active state. The results show contacts between the kinases between their N-lobe and their C-lobes that resemble MKK6-p38 complexes previously reported by the authors. Two MEK1-ERK2 conformers (States 1,3) are deemed inactive based on the lack of access of ERK-Y187 to the MEK1 active site, and the absence of ADP bound to MEK1 in State 3, while one conformer (State 2) is deemed active, but not fully active due to disorder in MEK1 activation loop (A-loop) and an essential salt bridge between strand beta3 and helix aC. HDX-MS and SAXS solution measurements and all-atom MD simulations are used to model the mutant complex and variants with WT ERK2. The study concludes that substrate recognition involves low-energy contacts with MEK, allowing substantial protein flexibility within the complex in a manner that may accommodate processive phosphorylation of ERK2.

      Strengths:

      The strengths of the work are that the findings provide important structural insights for MEK-ERK signaling and protein phosphorylation in general. These are valuable given that atomic resolution structures of kinase-substrate complexes are still limited in number. The authors succeeded in showing key contacts between subunits and conformational variations within the complex.

      Weaknesses:

      Weaknesses were that some of the conclusions about activity state, dynamics, and effects of ligand binding were less convincing. For example, that State 2 truly represents an active configuration seemed ambiguous, given the absence of Mg2+ and AlF4- in the cryoEM structure and disorder in the activation loop and the K97-E114 salt bridge. Conclusions by SAXS that ADP-AlF4 binding increases active site compaction while increasing local flexibility were not rigorously supported by HDX data, given that the latter were performed without ligand. Sections of the narrative and figures throughout were often confusing, and many assertions were made without clear explanation. Data shown in the supplementary materials were not always described in the Results, even those important for the conclusions. Figure legends and text lacked clear descriptions of specific complexes analyzed. Substantial changes are recommended to improve the readability and clarity of the work.

      We thank reviewer #1 for in-depth comments and analysis of our manuscript. However, there is a misunderstanding regarding HDX-MS and SAXS data. First, the HDX-MS data were performed on the ADP.AlF<sub>4</sub><sup>-</sup> inhibited complex - this was not made sufficiently clear in the text, and we will amend this. Secondly, we are not trying to support local flexibility observed in the SAXS data with the HDX data. The HDX data support the interactions observed in the cryoEM structure and demonstrate flexibility in the proline-rich loop, the ERK A-loop and unfolding of the MEK1 A-helix. The SAXS data demonstrate that when the transition state complex is formed, the complex is more compact but flexibility within the complex increases - as observed in the dimensionless Kratky plot, supporting our observations in the cryoEM maps. Therefore, the HDX data and SAXS data are separate observations. We thank reviewer #1 for all the comments and will rewrite the manuscript in order to increase clarity as suggested.

      Reviewer #2 (Public review):

      Summary:

      The authors used Cryo-EM to obtain a complex between MEK1 and ERK2. They used the same method as previously used by the same authors to form a stable complex between MKK6 and p38, an extra-strong KIM replacing the wild-type KIM in MEK1. Three conformers were resolved, with the highest resolution of 3.0 Å. The multiple conformers indicate more flexibility in MEK2 than in ERK2. These data suggest that nucleotide exchange is possible while maintaining MEK1-ERK2 interactions. SAXS and HDX data reinforce the idea of flexibility. They point to interactions between the two N-terminal domains between histidines at the N-terminus of helix C and between the G helices that are maintained in each of the 3 conformers, and sequence and structure suggest these histidines may be a source of specificity in MEK1-ERK2 versus MKK6-p38 interactions. A 2.2 Å structure of a complex between ERK1 (88% identical to ERK2) and the docking peptide used was also presented. Molecular dynamics simulations suggest that the MEK1-ERK2 complex can assume a fully active configuration of MEK1.

      Strengths:

      This is the first structure of a MEK1-ERK2 complex. The structural data are valuable additions to our understanding of MAP2K-MAPK interactions. The discussion points offered in the results section are palatable. These include the origins of specificity and the idea of flexibility in the MAP2K in support of a processive mechanism for the dual phosphorylation activity of MAP2Ks.

      Weaknesses:

      (1) This reviewer considers that the abstract is overstated. Specifically, this paper does not reveal the molecular details of phosphoryl transfer, nor does it demonstrate that substrate binding releases the catalytic machinery.

      (2) The discussion is in some places not supported by evidence and in others has superfluous text. Examples follow:

      - "Once the αG-helix is docked, and the C-lobe histidine triad is in place, the N-lobe interactions must then be fulfilled." The data in this paper does not suggest an order of events.

      - "If the substrate MAPK is incorrect, the N-lobe interaction will not be stabilised, preventing alignment of the MAPK A-loop with the MAP2K active site." This statement could be described as obvious.

      (3) Much of the discussion is embedded in the results, such that it is difficult to separate new facts offered by the paper from speculation.

      We thank reviewer #2 for comments and thorough analysis of our manuscript. We agree that perhaps the abstract should be toned down in terms of claims of an active conformation even though we feel that the combination of the first structure of the MEK1-ERK2 complex combined with MD simulation studies clearly demonstrate how phosphoryl transfer will occur. We are now also performing in-cell assays to support our theory of A-helix regulation. For point 2 we based the order of events on data from Juyoux et al 2023 Science, 381, 1217-1225, where in long-timescale MD simulations and experimentally validated adaptive Markov state model simulations the KIM interaction was the last to dissociate after the alpha-G helix interaction. In our MD simulations of the MEK1-ERK2 complex, the interactions formed by the N-lobe were weaker than those formed by the alpha-G. Indeed, dissociation of the N-lobe was observed in various independent simulations, whereas dissociation of the alpha-G was observed only once.

      Assuming that, as in the MKK6-p38a complex, the association proceeds along the reverse of the dominant dissociation pathway, the simulations suggest that the KIM interaction forms first, followed by the alpha-G and finally the N-lobes. While alternative association pathways are possible, this interpretation is consistent with the MD and in line with the main association and dissociation pathway observed for the MKK6-p38a complex. This is additionally supported by the observation that there is no catalytic activity if the KIM is removed, demonstrating this as the first essential recruitment event. We will expand this section to include our arguments.

      For the second example, we feel this is rather unfair. The statement that if the His-His interaction is absent, catalysis will be prevented is only obvious if one knows about the His-His interaction - this is the first structure showing pathway-specific interactions in the variable loop regions of a MAPK. If it is obvious, why has no one described these residues as important before?

      We have taken on board the comments on the style of the manuscript and will make significant changes as suggested by both reviewers.

    1. eLife Assessment

      The authors use a novel patch-leaving task to reveal a reward-reset strategy when mice choose to leave a depleting resource, and find that accumulated step-like activity in the dorsomedial striatum is correlated with the timing of these decisions. These important behavioral and neural findings are supported by substantial and convincing data. Additional control analyses would strengthen the evidence that these signals are specifically related to timing and patch-leaving decisions, rather than alternative task-related processes.

    2. Reviewer #1 (Public review):<br /> <br /> Summary:

      In this study, Shuler and colleagues record neurons from the DMS in mice performing a patch foraging task. In this task, mice had the choice between harvesting rewards from 2 ports - one the time-investment port where the rate of reward declined over time and the other a context port where the rate of reward was either high or low. Mice performed the task appropriately, switching between ports as the rate of reward declined in the time-investment port and switching more rapidly when the context port delivered high versus low rewards. The behavior of the mice was also strongly driven by time since the most recent reward receipt, in conflict with normative accounts of patch foraging. Individual DMS neurons showed bistable firing patterns, transitioning to high rates of activity at various times from reward. Overall, the population tiled the temporal space, and the accumulation of the number of neurons in the high firing state was predictive of patch exit. The rate of accumulation varied with things that also affected behavior.

      Strengths:

      Overall, the aims of the study were clear and important, the experiment directly addresses them, and the results are clear and provide compelling support for the authors' conclusions.

      Weaknesses:

      I have only a few comments and questions to consider, none of which are criticisms of what was done, really.

      (1) Probably my chief question, alluded to in the discussion, is what the evidence is that DMS plays a causal role in generating these correlates and the resulting behavior, in light of the lack of causal evidence here. What are other options? Could such information depend on upstream areas such as OFC or mPFC, with DMS just a pass-through? And while I would not ask for causal data, is there a specific prediction? That is, if the area were inactivated, would mice stay longer or shorter? Not do the task? If I wanted to do a causal test of the authors' idea regarding the contribution of DMS to this behavior, what would be predicted, and what result would invalidate the hypothesis? Speculating on this a bit, beyond just saying DMS is involved, would be useful.

      (2) Not much is said about the suboptimal strategy. Would DMS continue to play the same role if the mice showed no effect of recent reward and instead performed appropriately? Or is some other area doing that job? Or is this not important? I thought it was interesting that the mice basically did not treat the game quite like they were supposed to. Is it important to go back and look at what is happening in DMS under normative conditions to really know how this area contributes to proper foraging?

      (3) Do these neurons also track time in the context port? Or do they only exhibit this behavior in the port where rewards are depleting? This seems like an interesting question. Do they show the same profile in different ports, if so?

    3. Reviewer #2 (Public review):

      Summary:

      Here, Sutlief et al. use a novel patch-foraging task to investigate the role of dorsomedial striatal (DMS) neurons in determining when animals disengage from a resource. They show that mice, contrary to canonical optimal-foraging predictions, adopt a strategy in which reward receipt resets timing behavior, with decisions further shaped by both cumulative time spent in a patch and the overall quality of the environment. The authors further demonstrate that a subset of DMS neurons exhibits step-like activity patterns during task performance. Importantly, the accumulation of these state transitions across the neuronal population predicts the timing of patch-leaving decisions on a trial-by-trial basis, providing a potential neural mechanism underlying decisions about when to abandon a currently exploited resource.

      Strengths:

      This study addresses an important question using a well-designed, interesting behavioral task. The finding that mice employ a reward-triggered exit-timing policy is particularly interesting, as it is pertinent to the many patch foraging-style tasks that have been developed for use in mice, where rewards are delivered as discrete events. The identification of step-like activity in DMS neurons is mostly compelling, and the authors' trial-by-trial analysis linking this activity to behavior provides some support for its relevance to patch-leaving decisions.

      Weaknesses:

      A key interpretational issue is whether the DMS signal reflects timing specifically, rather than movement initiation or other task-related factors. The authors argue that once a sufficient number of neurons transition, the animal exits the time-investment port. However, it remains unclear whether this population threshold reflects a timing computation that determines when to leave in the more abstract sense, or a signal more directly related to movement onset (that may also be initiated after some proportion of the population has changed its activity). An important control would be to examine neural activity while animals are engaged at the context port. In this epoch, animals presumably do not need to time their departure in the same way, but they still eventually initiate movement. If the DMS signal reflects timing rather than movement, one would not expect the same accumulation-to-threshold pattern of step-like transitions at the context port.

      It would also be helpful for the authors to clarify the behavioral definition of the leaving decision. Can mice return to the time-investment port after exiting it if they do not subsequently enter the context port? How exactly is "exit" defined: as withdrawal from the time-investment port, entry into the context port, or some other behavioral event? Is there variability in the latency between time-investment port exit and context-port entry, and if so, is this latency related to DMS step-like activity? These details are important for interpreting whether the neural activity is aligned with a timing decision, movement initiation, or the execution of a transition between task states.

      The classification approach for identifying step-like activity seems generally reasonable, and the low false-positive rate against homogeneous Poisson controls is reassuring. However, one potential issue is that the identification of trial-by-trial state transitions is not independent of the session-level characterization of each neuron. The algorithm first fits a sigmoid to the pooled session data and then uses the resulting high- and low-firing-rate states to constrain interval-level fits. This may bias the analysis toward finding step-like transitions in neurons whose activity is only approximately step-like at the session level, effectively reducing the space of alternative solutions available to the interval-level fits. As implemented, the approach therefore functions more as a detector of consistency with a session-defined step model than as an unbiased test of whether individual intervals are better described by discrete state transitions versus alternative dynamics such as ramps or gradual drifts. This concern could be addressed by comparing the constrained sigmoid model against alternatives, such as constant-rate or ramping models, on held-out intervals, or by deriving state parameters from an independent subset of trials and testing classification on the remaining trials.

      The inclusion threshold for the accumulation analysis is difficult to evaluate. Sessions were included if they contained at least seven simultaneously recorded step-like units, but this number is hard to interpret without knowing the total number of recorded units per session and the fraction classified as step-like. Seven units may be sufficient for fitting a population accumulation trajectory, but because the cutoff is based on an absolute number rather than a proportion of the recorded population, it is unclear whether included sessions reflect robust population-level step-like dynamics or a relatively small selected subset of DMS activity. Reporting the number and fraction of step-like units per session, as well as the sensitivity of the accumulation results across different inclusion thresholds, would help clarify this point.

    4. Reviewer #3 (Public review):

      Sutlief and colleagues report behavioral and neural results from mice performing a patch foraging task. Behaviorally, they argue that time since last reward is a major determinant of when mice decide to leave a patch. In the brain, they find neurons in the dorsomedial striatum that show step-like changes in their firing rate at a range of times following reward. Population analyses show that the cumulative fraction of neurons that have undergone such a step-like change in firing rate can be used to predict patch-leaving times with impressive accuracy.

      Overall, this is an interesting set of results that has been analyzed in a principled way. The manuscript is well written, the results are explained clearly, and the evidence supporting the authors' conclusions is strong. The manuscript is therefore a potentially valuable contribution to the growing literature assessing how the brain solves stopping problems like the patch foraging scenario. I have suggestions for the authors to consider that might further increase the rigor of their results, and a few suggestions for improving the clarity of the work for readers.

      (1) I don't quite understand how the behavioral task works. Are mice rewarded for making discrete nose poke responses in the investment and context ports? Or are they required to nose poke and hold? Is reward given with some probability per response (which decreases with time in the patch), or is the reward probability a function of elapsed time in the patch, time since last response, or dwell time in the port? Also, exactly what equation defines how reward probability changes over time for the high- and low-value contexts? I couldn't find these details anywhere in the manuscript, and they would be helpful for better understanding the behavior and the later neural results.

      (2) How was the optimal strategy determined? Several features of the author's task violate the assumption of the marginal value theorem, so computing the optimal residence time is not a straightforward application of the classic model. There's a diagram in Figure 1h that depicts an MVT-like graphical solution, but the conventions of the plot are not familiar to me, and there's no description of how it works in the results or methods. More detail here would be much appreciated. In a similar vein, the authors report that mice generally exceeded optimal residence times in patches, but no statistical comparison is provided to back up that statement. There should be some formal test of this if it is to be included in the results.

      (3) The authors argue that time since last reward is the predominant determinant of patch leaving time. However, as the authors note, time since last reward is correlated with other task variables (patch reward rate, time in patch, etc.). I don't trust that SVM coefficients can be interpreted as straightforward measures of a variable's importance for classification performance in the case of correlated predictors. A better approach would be to assess how well the model performs as subsets of variables are added or removed from the model.

      (4) For the SVM analysis, I'm not quite understanding how or why the authors are using 5 s after mice left the patch as additional "Leave" examples. For instance, is time since entry computed for the investment patch, or the context patch that mice enter after they leave the investment patch? Similarly, is the time since the last reward relative to the investment patch, or the reward the mouse is likely to receive at the context patch? Moreover, I'm not sure it's safe to assume that because the mouse left at time t, time t+1 necessarily reflects conditions on which the mouse would definitely leave again. If we're thinking about the stay/leave decision as something that is being repeated sequentially on a fast time scale to determine how long mice stay in the patch, it doesn't follow that observing a mouse leave means that any patch conditions after that would necessarily result in the same decision. If that were the case, it would mean that seeing a mouse leave a patch after 2 s would preclude ever observing a residence time longer than 2 s, which is clearly not compatible with the authors' data. Ultimately, it's only possible to observe one decision to leave per trial; including data points beyond that as additional leave examples seems overly speculative to me.

      (5) The authors validate their approach for quantifying step-like changes in firing rate using simulations of constant-rate Poisson spiking and observe a low false positive rate. This is encouraging, but it doesn't seem like the only way in which their method could go awry, or even the most concerning way. I would be much more interested in seeing the false positive rate for continuous, ramp-like changes in firing rate, which would be much more likely to trip up the authors' approach and are also the major relevant alternative hypothesis to step-like changes in firing rate. Random walks in firing rate might also be worth testing.

      (6) The finding that cumulative "transitioned" neurons is predictive of patch leaving is interesting. However, I can't help but wonder how truly informative this variable is for predicting patch leaving. It seems as though neurons can only transition firing rates one time. That means that as time in the patch increases, the fraction of transitioned neurons naturally increases. Similarly, all visits must eventually end with the mouse leaving the patch, so the hazard rate of leaving increases with time in the patch. Given that, can the authors be certain that the cumulative transitioned neurons are really what's predicting patch leaving time, or would any generically increasing function perform roughly the same? An interesting test would be to mismatch the neural predictor and behavior at the level of trials. If this mechanism is really specific, rather than something that captures the general structure of an increasing hazard rate of leaving, then prediction of leaving time should work substantially better when the neural predictor is correctly matched to behavior on the trial for which it was recorded.

    5. Author response:

      Reviewer #1

      (1) Causality and the role of DMS; a specific, falsifiable prediction. We agree that the paper should not leave the causal question implicit, and we will expand the Discussion to state a concrete prediction rather than a general claim of involvement. Briefly, if the accumulation signal we describe carries the animal's intended departure time, then suppressing DMS during patch occupancy should not simply shift exit times in one direction but should degrade their structure: exit-time variability should increase, and exit timing should lose its systematic dependence on reward-rate context and on the time of the most recent reward. A plausible alternative outcome is disengagement from the task altogether, which would be uninformative and would need to be controlled for. The result that would falsify our hypothesis is the one we will state explicitly: exit timing that remains as predictable, and as sensitive to context and reward history, under DMS suppression, as without it. We will also discuss the alternative the reviewer raises, that these signals are inherited from cortical inputs such as OFC or mPFC with DMS acting as a relay, and note that our data cannot presently distinguish this from a locally generated signal.

      (2) Behavior under a normative strategy. This is an interesting question and we will address it in the Discussion. Our expectation, which we will frame as a prediction rather than a result, is that an animal timing from patch entry rather than resetting at each reward would show accumulation that begins at entry and proceeds to a context-dependent threshold at the reward-rate-optimal time, rather than the reward-triggered resets we observe. In this view, the reset structure of the neural signal is a reflection of the behavioral policy rather than a property of the region. We will make clear that this is a testable prediction that our current dataset does not address.

      (3) Do these neurons also track time at the context port? We intend to answer with new analysis, and it converges with Reviewer #2's suggested analysis (below), so we treat the two together there.

      Reviewer #2

      (a) Timing versus movement initiation: activity at the context port. We take this to be a central interpretational concern. We will examine whether the step-like DMS activity extends to the context port, testing the interval between the final context-port reward and departure for the same step-like transitions and accumulation we observe in the time-investment port. We will apply the same comparison to the context-port inter-reward intervals, which addresses Reviewer #1's third point about whether these neurons also track time at the context port.

      We want to flag one feature of the task that bears on how the outcome should be read. The context port is not a timing-free epoch. Its four rewards are delivered at predictable, regularly spaced intervals, and the interval between the final reward and the animal's departure is self-timed. Departure from the context port is therefore also a self-timed action, and observing accumulation there would not by itself indicate that the signal reflects movement initiation rather than timing. What the comparison can inform is whether the accumulation is specific to a decision about when to disengage from a depleting resource, or is a more general feature of self-timed departures. This is a meaningful distinction either way, and one we will report and interpret whichever direction the result falls.

      (b) Operational definition of leaving; the exit-to-entry latency. We agree these details are necessary for interpretation and their absence is our omission. Exit is the final withdrawal from the time-investment port preceding the next context-port visit, and we will make that clear in the revised methods. Mice can and occasionally do re-enter the time-investment port without an intervening context-port visit (especially early in training). Such re-entries are not counted as exits. We will also examine whether the latency between time-investment-port exit and context-port entry relates to the accumulation slope on the corresponding interval, to test whether the neural signal relates to the decision or to the execution of the transition.

      (c) Independence of interval-level fits from the session-level model. This is a fair characterization of the procedure, and we accept the distinction the reviewer draws between a detector of consistency with a session-defined step model and an unbiased test of discrete versus continuous dynamics. We will address it with a held-out validation: estimating each unit's state parameters and transition time from one half of its intervals and testing whether the transition times recovered from the withheld half agree. The discrete-versus-continuous comparison is addressed directly by the ramp simulations under Reviewer #3's point (5) below.

      (d) The inclusion threshold for the accumulation analysis. We will add a supplementary figure reporting the total number of recorded units per session and the fraction classified as step-like, so that the seven-unit criterion can be evaluated against the recorded population rather than in the abstract. Yield varied substantially across sessions, from a handful of units to roughly one hundred, and we will show this distribution directly. We will also report the accumulation results across a range of inclusion thresholds spanning approximately five to eight simultaneously recorded step-like units, so that readers can assess sensitivity to the choice.

      Reviewer #3

      (1) Specification of the task. We agree the task description was insufficient, and we will correct this at the front of the Results and in the Methods. The time-investment port operates on a poke-and-hold basis: the mouse maintains its head in the port and rewards are delivered stochastically over time for as long as it remains, with no requirement to withdraw and re-poke. Reward delivery follows an exponentially decaying rate in time since port entry, with a time constant of eight seconds, integrating to an expected eight rewards of one microliter each (8 µL total) for indefinite occupancy; we will give the explicit function. The reward probability function in the time-investment port is identical across blocks. The high- and low-reward-rate contexts are properties of the context port alone (four rewards over five seconds versus four rewards over ten seconds), and we will make this contrast unambiguous, since it is the manipulation on which the design rests.

      (2) Derivation of the optimum, Figure 1h, and a formal test of overstaying. We appreciate this comment. The optimal residence time in our task is not obtained by the classical Charnov tangent construction. It is computed by explicit maximization of the overall reward rate over the full cycle, following the framework in Sutlief et al. (2025) and shown graphically in Figure 1h. We will expand the legend of Figure 1h so its conventions are stated explicitly, give the reward-rate-maximizing derivation as an explicit equation in the Methods, and reframe the surrounding text around reward-rate maximization as the normative principle, with MVT identified as the special case it is. We will also add the formal statistical comparison of observed residence times against the computed optimum, which the reviewer correctly notes was asserted rather than tested.

      (3) Interpretation of SVM coefficients with correlated predictors. We accept this criticism. We will not rest the ordering of predictors on coefficient magnitudes alone. We will add a variable inclusion-and-ablation analysis, reporting cross-validated classification performance as each predictor is added to and removed from the model, so that the contribution of time since last reward is assessed by its effect on performance rather than by its normalized weight. We will additionally add a complementary analysis of the leave hazard that estimates the contribution of each variable without requiring the classification framework.

      (4) The five-second post-exit window. The reviewer is right that we did not explain this choice, and right that it rests on an assumption. Our reasoning was that a single exit moment per trial leaves the decision boundary badly under-constrained, and that treating the moments immediately following an exit as conditions under which the animal would also have left is licensed by the fact that within-patch reward rate declines monotonically with time, so conditions in the counterfactual continued visit would have been strictly less favorable than those already rejected. We accept that this is an assumption rather than an observation and will state it as such. We will also report the analysis across a range of window durations so that the independence of the result on this choice is visible. To the reviewer's specific questions: both time since entry and time since last reward are computed with respect to the time-investment port throughout, and we will state this explicitly.

      (5) False positive rate against ramps and random walks. We agree this is the more informative validation, and that continuous ramping is the most relevant alternative to ours. We will generate simulated units with continuous ramp-like rate changes, matched to the firing rates and interval structure of our recorded units, and pass them through the identical classification pipeline to obtain false positive rates comparable to the Poisson analysis already reported. We will retain the flat-rate Poisson simulation and present the ramp results as additional panels of the same supplement. We will also explore random-walk dynamics. Together with the held-out validation of transition times described under Reviewer #2(c), this converts the step characterization from a single-null validation into a comparison against the relevant continuous alternatives.

      (6) Specificity of the accumulation signal versus a generic increasing function. This is a valuable challenge and we will address it directly. We will implement the trial-mismatch control the reviewer proposes, randomly reassigning accumulation trajectories to reward-to-exit intervals within session and showing the extent to which predictive performance degrades relative to the correctly matched case.

      We would also note two features of the existing results that speak to this concern, and which we will bring forward in the revision because we did not make them salient enough. First, a signal that merely tracked elapsed time would be expected to shift its starting level as well as its rate across trials with different exit times; instead, the accumulation slope is strongly related to exit time (mean r = -0.551) while the intercept is not (mean r = 0.013), indicating a variable rate from a stable origin. Second, and more to the point, the accumulation arrives at a common level at the moment of exit whether the animal leaves early or late. The rate of accumulation shifts with the animal’s policy on that trial such that the threshold is met at the intended time. 

      What makes this predictor non-trivial is its trial-by-trial correspondence to behavior, not simply that it increases over time. We will make this argument explicitly alongside the shuffle control.

      Summary

      To summarize the planned additions: (i) analysis of step-like activity and its accumulation at the context port, with the interpretive caveat noted above; (ii) validation of step detection against ramping alternatives, together with held-out estimation of transition times; (iii) a trial-mismatch control for the specificity of the accumulation predictor; (iv) an inclusion-and-ablation analysis of the behavioral predictors and a complementary hazard model of the leave decision; (v) reporting of unit yield, step-like fraction, and sensitivity of the accumulation results to the inclusion threshold; (vi) analysis of the exit-to-context-entry latency in relation to the neural signal; (vii) a formal statistical test of overstaying relative to the computed optimum; and (viii) substantial clarification of the task specification, the operational definition of exit, and the derivation of the optimal residence time, including an expanded Figure 1h legend.

      Our aim in the revision is to meet the specificity concern raised in the assessment as directly as the existing data allow, and we hope the revised manuscript will warrant reconsideration of the strength-of-evidence characterization.

      We are grateful to the reviewers for the care evident in their reports, and to you both for handling the manuscript.

      References

      Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology, 9(2), 129–136.

      Sutlief, E., Walters, C., Marton, T., & Hussain Shuler, M. G. (2025). The value of initiating a pursuit in temporal decision-making. eLife. https://doi.org/10.7554/eLife.99957.2.

    1. eLife Assessment

      This important study uses longitudinal EEG to chart how neural tracking of syllables and word-level statistical structure develops over the first two years of life in infants at high and low likelihood for autism and links these measures to verbal outcomes at 18-20 months. The strength of evidence is convincing: the prospective longitudinal design, careful data-quality handling, and partial least squares analyses are appropriate and well executed, though some interpretations of the group differences in syllable tracking, along with the possible contributions of multilingual exposure and sleep state during recording, warrant caution. The work will be of interest to developmental cognitive neuroscientists studying language acquisition and early neural markers of neurodevelopmental conditions.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript reports a prospective longitudinal study examining whether infants with high likelihood (HL) for autism differ from low-likelihood (LL) infants in two levels of word learning: brain-to-speech cortical entrainment and implicit word segmentation. The authors report reduced syllable tracking and post-learning word recognition in the HL group relative to the LL group. Importantly, both the syllable-tracking entrainment measure and the word recognition ERP measure are positively associated with verbal outcomes at 18-20 months, as indexed by the Mullen Verbal Developmental Quotient. Overall, I found this to be a thoughtfully designed and carefully executed study that tackles a difficult and important set of questions. With some clarifications and modest additional analyses or discussion on the points below, the manuscript has strong potential to make a substantial contribution to the literature on early language development and autism.

      Strengths:

      This is an important study that addresses a central question in developmental cognitive neuroscience: what mechanisms underlie variability in language learning, and what are the early neural correlates of these individual differences? While language development has a relatively well-defined sensitive period in typical development, the mechanisms of variability-particularly in the context of neurodevelopmental conditions-remain poorly understood, in part because longitudinal work in very young infants and toddlers is rare. The present study makes a valuable contribution by directly targeting this gap and by grounding the work in a strong theoretical tradition on statistical learning as a foundational mechanism for early language acquisition.

      I especially appreciate the authors' meticulous approach to data quality and their clear, transparent description of the methods. The choice of partial least squares correlation (PLS-c) is well motivated, given the multidimensional nature of the data and collinearity among variables and the manuscript does a commendable job explaining this technique to readers who may be less familiar with it.

      The results reveal interesting developmental changes in syllable tracking and word segmentation from birth to 2 years in both HL and LL infants. Simply mapping these trajectories in both groups is highly valuable. Moreover, the associations between neural indices of brain-to-speech entrainment and word segmentation with later verbal outcomes in the LL group support a critical role for speech perception and statistical learning in early language development, with clear implications for understanding autism. Overall, this is a rich dataset with substantial potential to inform theory.

      Comment on revised version.

      The revised manuscript has provided additional analyses that lead to critical clarification of the main findings, including the longitudinal nature of the relationship between neural tracking of speech and language, the role of sleep, and the potential modulation effect of stream structure on syllable-level neural tracking. The overall results highlight the robustness of the findings as well as the specific relevance of the structured speech tracking to verbal outcomes of infants with high likelihood (HL) of autism.

    3. Reviewer #2 (Public review):

      Summary:

      This article looks at differences in how the brain entrains to, or tracks, the rhythmic presentation of syllables and words in speech in infants at increased likelihood versus low likelihood for autism. The authors first sought to characterize how brain responses are modulated by learning the statistical probability of a given syllable following the one before it over the first two years of life. They then sought to identify at which stages of word learning infants at increased likelihood for autism showed difficulties, and whether those difficulties worsened over time. Finally, they sought to indicate whether infants' statistical learning and word learning abilities could predict later verbal skills. The authors found similar developmental trajectories of neural entrainment to syllables in infants at high and low likelihood for autism, but infants at high likelihood for autism had overall weaker syllable-level entrainment. Infants at high versus low likelihood for autism showed different developmental trajectories for word entrainment. Lower syllable entrainment in high-likelihood infants corresponded with poorer verbal outcomes, but word entrainment was not associated with verbal outcomes. Event-related potential responses to words and part words were positively associated with verbal outcomes, however, but only in low-likelihood infants.

      Strengths:

      Overall, the article provides rigorous statistical analysis of longitudinal EEG data to provide strong support for the claims that neural entrainment to syllable and word features of speech may be a useful marker for language development difficulties, particularly in infants at increased likelihood for neurodevelopmental disorders. The EEG data collection and preprocessing procedures are well within standards within the field. Readers should take care to note that authors indexed neural entrainment to speech using phase-locking values instead of spectral power.

      Comments on revised version.

      While the statistical analyses are rigorous, there are a few potential confounds to the results. The authors now do a nice job addressing these limitations to the work. For example, sleep status may modulate some of the biomarkers relevant for language learning. Exposure to additional languages may influence performance on the verbal assessment, though the authors do clarify that participants came from majority French-speaking households. As a result, readers should be encouraged to interpret that neural entrainment to speech features is likely a useful mechanism to explain differences in language development, while taking this interpretation with some caution.

    4. Author response:

      The following is the authors’ response to the current reviews.

      Reviewer #1 (Recommendations for the authors):

      (1) Interpretation of Syllable-Tracking in the RND Condition:

      The finding of greater syllable-tracking in the LL group compared to the HL group in the RND condition warrants cautious interpretation. Currently, there is no direct statistical evidence demonstrating greater PLV at 4 Hz in the Structured versus Random conditions for either group; readers must infer this solely from numeric differences in Figure S5 B and D. Therefore, while the interpretation on Page 14 (Lines 443-446) "successful segmentation may enhance syllable tracking via top-down predictions of the next syllable" is an interesting speculation, it feels somewhat far-reaching. Additionally, the authors should discuss whether this upregulated syllable tracking in the structured condition (which is specific to the HL group) represents an adaptive or maladaptive response.

      The reviewer correctly highlights the lack of direct  comparison between conditions (RND versus STR). We tempered our claims in the cited paragraph and insisted on the speculative nature of this part of the discussion. We also clarified that, to us, it may represent an adaptive compensatory strategy:

      Page 14, line 441: “Interestingly, our supplementary analyses (Supplementary Material Figure S4-5) suggest that syllable entrainment may be differentially affected in HL versus LL infants, depending on the statistical structure of the input stream (RND versus STR). However, as our experiment was not explicitly designed to test stream effects, these results should be interpreted with caution. Future studies could explore how successful segmentation may enhance syllable tracking via top-down predictions of the next syllable in both LL and HL infants. If confirmed, such a mechanism may improve alignment to syllable onsets, potentially constituting a compensatory process allowed by preserved segmentation abilities.”

      (2) Preservation of Statistical Learning in HL Infants:

      The text added on Pages 17-18 (Lines 562-566) regarding a "heightened dependence on bottom-up mechanisms (in autism)" does not appear to be supported by the data or by theories of implicit statistical learning. Because greater syllable-level entrainment was observed in the LL group than the HL group across both the random and structured conditions, the data actually point toward impaired bottom-up processes. Furthermore, implicit statistical learning typically involves an interplay of both bottom-up and top-down mechanisms; the implicit nature of a task does not guarantee a strictly bottom-up process. Consequently, this interpretation is not entirely convincing.

      We agree with the reviewer that the concepts of “top-down” and “bottom-up” were not fully appropriate to support our point in the cited paragraph. We should have used the concepts of implicit versus explicit learning instead, in line with previous literature suggesting increased reliance on preserved implicit learning in autism to compensate for altered explicit processes. The paragraph was slightly modified.

      Page 18, line 564: “According to these studies, autistic impairments in explicit attentional processes, such as social orienting - which are critical for bootstrapping language acquisition (70) - may result in a heightened dependence on implicit mechanisms, including statistical learning. As previously discussed, preserved word segmentation abilities may further compensate for alterations in lower-level implicit processes, such as syllable tracking.”

      Reviewer #2 (Recommendations for the authors):

      Potential typo on line 199 - I think an apostrophe is needed here.<br /> Potential typo on line 255 - do you mean Central electrodes?

      We addressed the typos spotted by reviewer.

      Line 199: variables’

      Line 255: Centro-frontal electrodes


      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports a prospective longitudinal study examining whether infants with high likelihood (HL) for autism differ from low-likelihood (LL) infants in two levels of word learning: brain-to-speech cortical entrainment and implicit word segmentation. The authors report reduced syllable tracking and post-learning word recognition in the HL group relative to the LL group. Importantly, both the syllable-tracking entrainment measure and the word recognition ERP measure are positively associated with verbal outcomes at 18-20 months, as indexed by the Mullen Verbal Developmental Quotient. Overall, I found this to be a thoughtfully designed and carefully executed study that tackles a difficult and important set of questions. With some clarifications and modest additional analyses or discussion on the points below, the manuscript has strong potential to make a substantial contribution to the literature on early language development and autism.

      Strengths:

      This is an important study that addresses a central question in developmental cognitive neuroscience: what mechanisms underlie variability in language learning, and what are the early neural correlates of these individual differences? While language development has a relatively well-defined sensitive period in typical development, the mechanisms of variability - particularly in the context of neurodevelopmental conditions - remain poorly understood, in part because longitudinal work in very young infants and toddlers is rare. The present study makes a valuable contribution by directly targeting this gap and by grounding the work in a strong theoretical tradition on statistical learning as a foundational mechanism for early language acquisition.

      I especially appreciate the authors' meticulous approach to data quality and their clear, transparent description of the methods. The choice of partial least squares correlation (PLS-c) is well motivated, given the multidimensional nature of the data and collinearity among variables, and the manuscript does a commendable job explaining this technique to readers who may be less familiar with it.

      The results reveal interesting developmental changes in syllable tracking and word segmentation from birth to 2 years in both HL and LL infants. Simply mapping these trajectories in both groups is highly valuable. Moreover, the associations between neural indices of brain-to-speech entrainment and word segmentation with later verbal outcomes in the LL group support a critical role for speech perception and statistical learning in early language development, with clear implications for understanding autism. Overall, this is a rich dataset with substantial potential to inform theory.

      Weaknesses:

      (1) Clarifying longitudinal vs. concurrent associations

      Because the current analytical approach incorporates all time points, including the final visit, it is challenging to determine to what extent the brain-language associations are driven by longitudinal relationships vs. concurrent correlations at the last time point. This does not undermine the main findings, but clarifying this issue could significantly enhance the impact of the individual-differences results. If feasible, the authors might consider (a) showing that a model excluding the final visit still predicts verbal outcomes at the last visit in a similar way, or (b) more explicitly acknowledging in the discussion that the observed associations may be partly or largely driven by concurrent correlations. Either approach would help readers interpret the strength and nature of the longitudinal claims.

      We thank the reviewer for this insightful comment. We agree that distinguishing between longitudinal predictive power and concurrent correlations at the final visit is crucial for clarifying the nature of these brain-language associations. Following the reviewer’s suggestion (a), we re-ran the two critical Partial Least Squares Correlation (PLS-c) analyses by excluding all EEG and behavioral data from the final 18–21 month visit (n = 54 recordings kept) to test whether earlier trajectories still predict the final verbal outcome.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2C–D): The PLS-c restricted to the 3- to 15-month visits still identified a single significant component (p=.001, r=.56, 64.0% explained covariance, Figure 2 -figure supplement 3). Bootstrap ratios (BSR) were: contrast (low vs. high autism likelihood) 4.1; mean age −1.3; contrast*mean-age −2.1; delta-age 10.5; contrast*delta-age 2.0; age<sup>2</sup> −6.0; contrast*age<sup>2</sup> −5.8; and notably verbal outcome 7.1; contrast*verbal-outcome −6.3.

      The latent component and its spatial electrode configuration remain highly consistent with the original analysis (Figure 2C–D). This confirms that excluding the final visit preserves the model’s predictive validity: lower syllable entrainment correlates with poorer verbal outcomes at 18–21 months, particularly in the high-likelihood group.

      (2) Late evoked response to novel words (original analysis on Figure 6): The PLS-c analysis on the ERP late time window (1500–3000 ms), excluding the final visit, also revealed one significant component (p=.002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (part-word versus word) 14.5; mean age -5.2; contrast*mean-age 6.2; delta-age -0.8; contrast*delta-age 3.5; age<sup>2</sup> -0.8; contrast*age<sup>2</sup> -12.1; verbal-outcome 20.1; contrast*verbal-outcome -7.2. The latent component closely mirrored the original analysis (Figure 6), with frontal electrodes contributing negatively and posterior electrodes positively. Minor divergences in age-related parameter contributions were observed, likely due to the absence of 18-21 month timepoints, which previously contributed to the convex/concave shapes of the group age trajectories in figure 6B (left panel).

      Crucially, both models (with and without the final visit) positively predicted verbal outcomes (Figure 6B, right panel, and Figure 6 -figure supplement 1B). However, excluding the final visit reversed the direction of the group*verbal-outcome interaction (from 3 to -7.2): This indicates that after ruling out cross-sectional correlations at 18–21 months, the early predictive value of the late ERP to word novelty is more prominently observed in high-likelihood infants, suggesting that the original result was influenced by concurrent cross-sectional correlations at the final visit. This aligns with the syllable entrainment findings (Figure 2 -figure supplement 3), as both 4 Hz neural tracking and late ERP responses to novelty predominantly predict verbal outcomes in infants at high likelihood for autism.

      We reported these supplementary analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 3 and Figure 6 -figure supplement 1. In general, most of figures that were present in Supplementary materials were moved as figure supplements to enhance readability.

      Page 8 lines 234-240 (pages and lines refer to the reviewed uploaded manuscript): “To rule out the possibility that the association between syllable entrainment and verbal outcome was driven by concurrent measures taken at 18–21 months, we re-ran the PLS-c analysis excluding EEG data from the final visit (n = 54 recordings kept). The resulting latent component remain significant (p = .001) and showed contributions from behavioral and EEG variables that were highly similar to those observed in the previous analysis, with a verbal outcome BSR of 7.1 and a group’verbal-outcome interaction BSR of −6.3 (Figure 2 -figure supplement 3).”

      Page 12 lines 387-394: “As we did for neural entrainment to syllables, we conducted a new analysis on late ERP to word novelty, excluding EEG data from the final visit. This PLS-c yielded one significant latent component (p = .002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1) with globally similar EEG parameter contributions and age trajectory modelling. Verbal outcome still significantly contributed to the latent component (BSR=20.1), with a negative verbal outcome*group interaction (BSR=-7.2). These results suggest that, after ruling out cross-sectional correlations at 18–21 months, the late ERP to word novelty predominantly predicts verbal outcomes in high-likelihood infants for autism.”

      Page 17 lines 547-548: “As with syllable entrainment, the late ERP to novel words primarily predicted verbal outcomes in high-likelihood (HL) infants.”

      Page 18 lines 588-590: “Likewise, the absence of a late ERP orientation response in HL participants may represent an early neural signature of altered attention to novelty that can be used both as a non-invasive predictor of language development and as a potential target for early intervention.’

      (2) Incorporating sleep status into longitudinal models

      Sleep status changes systematically across developmental stages in this cohort. Given that some of the papers cited to justify the paradigm also note limitations in speech entrainment and word segmentation during sleep or in patients with impaired consciousness, it would be helpful to account for sleep more directly. Including sleep status as a factor or covariate in the longitudinal models, or at least elaborating more fully on its potential role and limitations, would further strengthen the conclusions and reassure readers that these effects are not primarily driven by differences in sleep-wake state.

      The reviewer is highlighting here a limitation of our study design that comprised sleeping status that varied from one timepoint to another among participants. To rule out any confounding effect of wake status (coded as a binary variable: sleeping or awake during recording) on analyses comparing groups, a linear mixed-effect model with repeated measures was fitted finding no significant difference between high- and low-likelihood participants (p=.769, reported at page 20, lines 646-647). However, as rightly suggested by the reviewer, this doesn’t prevent from a sleep bias on age trajectories, especially given that sleeping status significantly decreases with age in our sample.

      Including sleep status as a covariate in our analyses, as suggested by the reviewer, would be difficult to implement in our PLS-c methods, since a categorical behavioral parameter that varies within participants is not possible in the models provided by myPLS toolbox.

      As an alternative option, we re-ran all analyses that explored the condition effect on the whole sample within the sleeping participants only (n=25 recordings) to confirm that the same age-trajectories of EEG parameters were highlighted. However, negative results should be interpreted with caution since the sample is small for such a multivariate approach, resulting in modest statistical power.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2A–B): The PLS-c identified one significant component (p <.001, r = .78, 85.1% explained covariance, Figure 2 -figure supplement 2 and Figure 3 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (4hz vs. adjacent frequencies) 30.3; mean age -2.9; contrast*mean-age -2.4; delta-age 3.8; contrast*delta-age 3.4; age<sup>2</sup> -1.1; contrast* −2.5. The spatial distribution of contributing electrodes globally matched that shown in Figure 2A. The high contrast BSR (30.3) confirms robust syllable entrainment in sleeping infants. Critically, the contrast*age<sup>2</sup> parameter contributed negatively to the latent component (BSR = −2.5), confirming that the convex age trajectory of syllabic entrainment (Figure 2B) is also present in the sleeping subsample.

      (2) Word entrainment (1.3 Hz) (original analysis on Figure 3A–B): The PLS-c identified one significant component (p <.001, r = .63, 37.3% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (1.3hz vs. adjacent frequencies) 24.9; mean age -7.5; contrast*mean-age -3.2; delta-age 2.9; contrast*delta-age 0.0; age<sup>2</sup> 1.5; contrast* 1.1. The spatial distribution of significant electrodes partially overlaps with the ones in the original analysis, primarily showing fronto-central positive contribution to the latent component. The high contrast BSR confirms a robust word entrainment in sleeping participants, in line with previous studies (e.g., Flò et al, Sci Rep, 2022). However, the lack of a significant contrast* age<sup>2</sup> suggests that the U-shape age trajectory illustrated on Figure 3 might be modulated by wakefulness or due to a lack of power in the present analysis. A non-significant trend towards a U-shape pattern with a 12-month nadir is visible in sleeping participants, but additional data from sleeping 18-21 months sleeping infants would be required to confirm or refute this trend.

      (3) Early evoked response to novel words (original analysis on Figure 4): The PLS-c analysis on the ERP early time window (0–1000 ms) in sleeping participants revealed no significant component. The absence of early response to word novelty in sleeping participant might account for the lack of response observed in the whole sample, illustrated on Figure 4. To test this hypothesis, we conducted the same PLS-c in awake participants (n=58 recordings), which also yielded no significant latent component. This suggests that the lack of a measurable early response to word novelty observed in the whole sample is consistent across both sleeping and awake infants, and not driven by any of the two subsamples.

      (4) Late evoked response to novel words (original analysis on Figure 5): The PLS-c analysis on the ERP late time window in sleeping participants revealed no significant component. This suggests that sleeping participants might present a reduced or even absent late response to novel words. Given this identified effect of sleep on late ERP response, we reran the PLS-c on the late ERP window using group as contrast (original analysis on figure 6), excluding the sleeping participants to avoid any confounds. This PLS-c revealed one significant component (p = .006, r = .68, 30.5% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (low versus high likelihood) 7.6; mean age -5.9; contrast*mean-age -3.8; delta-age 2.5; contrast*delta-age -3.5; age<sup>2</sup> -2.5; contrast*age<sup>2</sup> -4.9; verbal-outcome 8.5; contrast*verbal-outcome -1.0. Behavioral parameters contribute to this latent component with similar magnitude and polarity as in the original analysis. Electrode contributions are also highly consistent, with frontal negative and posterior positive contributions. This confirms that sleeping participants, despite their potentially reduced late response, did not significantly bias the results presented in Figure 6.

      Author response image 1.

      Late evoked response potential (ERP) to word novelty in awake participants. A. Design and brain saliences derived from the significant latent component. Brain topographies of bootstrap ratios (BSR) are displayed at 250ms intervals. Black dots indicate BSR > 2.3. B. Participants’ brain scores for part-word and word conditions, as a function of age (left panel) and verbal DQ (right panel). For details on brain scores, see Figure 6 -figure supplement 1. Linear fitting is used for illustrative purposes only. HL: high likelihood for autism; LL: low likelihood for autism.

      We reported these analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 2A and Figure 3 -figure supplement 1.

      Page 7, lines 218-224: “Because some infants were asleep during the recording session, particularly at younger ages, we performed a supplementary control analysis restricted to this sleeping subsample (n = 25 recordings, Figure 2 -figure supplement 2). This PLS-c also identified a significant latent component (p < .001, r = .78, 85.1% explained covariance), with a significant contrast effect (BSR = 30.3) and a significant negative contrast*age<sup>2</sup> interaction (BSR = −2.5). These findings confirm that the convex age trajectory observed in the main analysis remains present and observable even in sleeping infants.”

      Page 8, lines 252-259: “We further investigated word entrainment in sleeping participants (n=25), which yielded one significant latent component (p<.001, r=.63, 37.3% explained covariance, Figure 3 -figure supplement 1). Centro-frontal electrode contributed to this component, with a high contrast BSR (24.9), confirming a similar word entrainment pattern in the sleeping subsample. The contrast*age<sup>2</sup> was also positive but not significant (1.1), suggesting a trend toward a U-shape age trajectory with a 12-month nadir in sleeping infants. Additional 18-21 month recording would be required to confirm this trend.”

      Page 11, lines 347-349: “The same PLS-c, conducted separately in sleeping (n=25) and awake subsamples (n = 58), yielded no significant latent component, indicating a consistent absence of early response to word novelty in both sleeping and awake infants.”

      Page 11-12 lines 368-370: “The same PLS-c in the sleeping subsample yielded no significant latent component, suggesting that sleep may reduce or even abolish the late response to word novelty.”

      Page 12 lines 382-385: “Given that no late response was detected in sleeping participants, we re-ran the PLS-c analysis using group as a contrast in the awake subsample (n=58). This yielded one significant latent component (p=.006, r=.68, 30.5% explained covariance), with behavioral and electrode contributions highly overlapping with those in Figure 6.”

      Page 18 lines 590-592: “This potential biomarker might nevertheless be modulated by participants’ sleep status, warranting careful consideration of vigilance state in future studies.”

      (3) Use of PLS-c and potential group × condition interactions

      I am relatively new to PLS-c. One question that arose is whether PLS-c could be extended to handle a two-way interaction between group and condition contrasts (STR vs. RND). If so, some of the more complex supplementary models testing developmental trajectories within each group (Page 8, Lines 258-265) might be more directly captured within a single, unified framework. Even a brief comment in the methods or discussion about the feasibility (or limitations) of modeling such interactions within PLS-c would be informative for readers and could streamline the analytic narrative.

      The reviewer raises a valid concern regarding the capacity of PLS-c to accommodate multi-way interactions among categorical and continuous variables. While PLS-c has no inherent theoretical constraints on the number of predictor terms (they can even exceed the sample size in number), practical limitations arise from model stability and interpretability when the ratio of predictors to sample size becomes excessive. As noted by Geladi and Kowalski (1986), exceeding ~10% of the sample size with predictors increases noise sensitivity and overfitting.

      In our study, the PLS-c analyses already reach this ~10% limit, with a maximum of nine predictors for a sample size of n=83. Attempting to integrate both group and condition as contrasts — along with necessary age parameters to account for developmental trajectories — would result in 12 predictors (or 15 if verbal outcome is included). Specifically, the model would require behavioral terms for Group, Condition, Group*Condition, Mean-age, Group*Mean-age, Condition*Mean-age, Delta-age, Group*Delta-age, Condition*Delta-age, Age<sup>2</sup>, Group*Age<sup>2</sup>, Condition*Age<sup>2</sup>, Verbal-outcome, Group*Verbal-outcome, and Condition*Verbal-outcome.

      Although a unified multivariate model capturing the complex dynamics at play in our sample is theoretically appealing, the substantial risk of overfitting precludes its feasibility. Therefore, we opted to use only one categorical predictor per PLS-c analysis to maintain model parsimony and reliability. However, a larger sample could overcome this limitation, allowing a stable and unified model of longitudinal EEG data that simultaneously captures age trajectories, group, clinical outcome, and condition.

      Reference:

      Geladi, P., & Kowalski, B. (1986). Partial least-squares regression: A tutorial. Analytica Chimica Acta, 185, 1–17. https://doi.org/10.1016/S0003-2670(00)82582-3

      We added the following comment in the method section:

      Page 24, lines 773-777: “We limited the number of behavioral variables to nine to mitigate noise sensitivity and overfitting risks associated with exceeding the 10% sample size threshold (Geladi & Kowalski, 1986). This limitation precluded the implementation of a single PLS-c model incorporating group, condition (STR vs. RND), age, and their interactions.”

      (4) STR-only analyses and the role of RND

      Page 8, Lines 241-245: This analysis is conducted only within the STR condition. The lack of group difference observed here appears consistent with the lack of group difference in word-level entrainment (Page 9, Lines 292-294), suggesting that HL and LL groups may not differ in statistical learning per se, but rather in syllabic-level entrainment. As a useful sanity check and potential extension, it might be informative to explore whether syllable-level entrainment in the RND condition differs between groups to a similar extent as in Figure 2C-D. In other work (e.g., adults vs. children; Moreau et al., 2022), group differences can be more pronounced for syllable-level than for word-level entrainment. Figure S6 seems to hint that a similar pattern may exist here. If feasible, including or briefly reporting such an analysis could help clarify the asymmetry between the two learning measures and further support the interpretation of syllabic-level differences.

      The reviewer points to the interesting pattern highlighted in supplementary figure S6, suggesting that group differences in syllabic entrainment might be modulated by the structure of the stream (STR versus RND). Such modulatory effect of stream structure on entrainment to syllables has been suggested by many studies, like Moreau et al (2022), as pointed by the reviewer, and seems at play in our sample, as illustrated on supplementary figure S5 (decline in the 4hz PLV that exceeds the size of confidence intervals, ~90 s after STR onset).

      Following the reviewer’s suggestion, we ran a PLS-c testing group effect on 4hz PLVs in each stream:

      (1) in the RND stream: the analysis yields one significant component (p<.001, r=.49, 52.7% explained covariance, Author response image 2A-B). Bootstrap ratios (BSR) are: contrast (low versus high likelihood) 6.9; mean age -1.0; contrast*mean-age 0.1; delta-age 11.8; contrast*delta-age 0.1; age<sup>2</sup> -4.9; contrast*age<sup>2</sup> 4.6; verbal-outcome 9.7; contrast*verbal-outcome -0.6. Interestingly, the model still highlights a strong link between syllable tracking and group, suggesting that RND also discriminate between HL and LL. However, RND syllable tracking doesn’t appear to be linked to group x verbal-outcome as we observed in Figure 2C-D.

      (2) In the STR stream, we obtained one significant latent component (p=.002, r=.51, 57.8% explained covariance, Author response image 2C-D). Bootstrap ratios (BSR) are: contrast 2.2; mean age -1.6; contrast*mean-age -1.1; delta-age 6.1; contrast*delta-age -0.7; age<sup>2</sup> -4.4; contrast*age<sup>2</sup> 0.5; verbal-outcome 8.2; contrast*verbal-outcome -7.3. Here, the strong association between syllable tracking and group x verbal-outcome is similar to the model presented in Figure 2C-D.

      Taken together, these results suggest that the apparent STR/RND dissociation illustrated in Figure S6 might primarily reflect a Group*Verbal-outcome divergence, with syllable tracking in the STR stream being related to verbal outcome mainly in high likelihood for autism.

      Author response image 2.

      Syllable entrainment within RND (A-B) and STR (C-D).

      These results were reported in the revised manuscript in the Result section (Time course of the entrainment along experiment subheader), implying a slight reframing of the result presentation of supplementary analysis S6. Author response image 2 was added in supplementary material as Figure S5.

      Page 10, lines 307-319: “The group, age and verbal outcome parameters were mainly correlated (BSR>2.3) with the neural entrainment occurring~90 seconds after the onset of the STR stream, coinciding with the time participants began tracking word boundaries (Supplementary material, S3). This result suggests that the group differences in syllable entrainment, as shown in Figure 2C-D, as their associations with verbal outcome, are modulated by the structure of the stream (STR versus RND). We ran one additional PLS-c for each stream separately, using group as contrast. In both streams, the PLS-c yielded a significant LC (p<.001 for RND and p=.002 for STR), with a positive group effect (BSR>2.3) in both LC (Supplementary material, S5). Most strikingly, the group*verbal outcome parameter reached significance exclusively within the STR latent component (BSR:-7.3). These results suggest that while syllable tracking is generally decreased in HL infants across both streams, its association with verbal outcome is prominently driven by the stream containing words (STR).”

      Page 14, lines 443-446: “This temporal overlap suggests that successful segmentation may enhance syllable tracking via top-down predictions of the next syllable, improving alignment to syllable onsets in LL infants as well as in HL with better verbal outcome.’

      (5) Multi-speaker input and voice perception (Page 15, Lines 475-483)

      The multi-speaker nature of the speech input is an interesting and ecologically relevant feature of the design, but it does add interpretive complexity. The literature on voice perception in autism is still mixed: for example, Boucher et al. (2000) reported no differences in voice recognition and discrimination between children with autism and language-matched non-autistic peers, whereas behavioral work in autistic adults suggests atypical voice perception (e.g., Schelinski et al., 2016; Lin et al., 2015). I found the current interpretation in this paragraph somewhat difficult to follow, partly because the data do not directly test how HL and LL infants integrate or suppress voice information. I think the authors could strengthen this section by slightly softening and clarifying the claims.

      We acknowledge the reviewer’s concern regarding the potential ambiguity in the cited paragraph. To address this, we have revised the text to explicitly clarify the aims of our study and its design. Furthermore, we now emphasize the speculative and post-hoc nature of the hypotheses and interpretations presented, thereby ensuring transparency regarding the limitations of our findings.

      Page 16 lines 520-530), as follows: “HL infants, on the other hand, did not show this transient disruption. In this group, word entrainment remained stable over time. To account for this unexpected finding, we followed up on the post-hoc hypothesis proposed above: a reduced sensitivity to social and vocal cues observed in HL infants may have spared segmentation abilities by limiting the interference introduced by speaker variability. If this post-hoc hypothesis holds true, LL and HL infants would differ not in their intrinsic ability to learn statistical regularities per se, but rather in how they integrate or suppress competing cues (such as speaker changes) during the segmentation process. It is important to note, however, that the present study was not designed to isolate and evaluate the specific impact of speaker changes on word segmentation. Consequently, this interpretation remains speculative, and additional research is required to further address this question.”

      (6) Asymmetry between EEG learning measures

      Page 16, Lines 502-507 touches on the asymmetry between the two EEG learning measures but leaves some questions for the reader. The presence of word recognition ERPs in the LL group suggests that a failure to suppress voice information during learning did not prevent successful word learning. At the same time, there is an interesting complementary pattern in the HL group, who show LL-like word-level entrainment but does not exhibit robust word recognition. Explicitly discussing this asymmetry - why HL infants might show relatively preserved word-level entrainment yet reduced word recognition ERPs, whereas LL infants show both - would enrich the theoretical contribution of the manuscript.

      We concur with the reviewer’s observation that our findings imply a theoretically significant double dissociation between HL and LL groups, specifically concerning the asymmetries between word-level neural entrainment and word recognition mechanisms. We believe this point was partly addressed in the subsequent paragraph, where we stated that “in contrast” to LL, HL infants “showed no clear ERP difference between novel and familiar triplets”, while “both groups showed similar word neural entrainment during learning”. We further explored potential explanations for this apparent dissociation, such as a possible deficit in novelty orientation that may be specific to HL infants and unrelated to statistical learning itself. We cited Liu et al (2023) as a reference showing the dissociation between mechanisms underlying implicit versus explicit traces of statistical learning. We acknowledge that we can discuss more in depth the potential preservation of statistical learning in HL infants. We have incorporated the following discussion in the reviewed manuscript, supported by relevant references:

      Pages 17-18, lines 562-566: “Interestingly, this dissociation between spared implicit versus impaired explicit statistical learning in autism has been previously discussed in the literature (Zwart et al, 2018, Kissine, 2021). According to these studies, autistic impairments in top-down attentional processes, such as social orienting — which are critical for bootstrapping language acquisition (Kuhl, 2007) — may result in a heightened dependence on bottom-up mechanisms, including implicit statistical learning.”

      References:

      Zwart, F.S., Vissers, C.T.W.M., Kessels, R.P.C. and Maes, J.H.R. (2018), Implicit learning seems to come naturally for children with autism, but not for children with specific language impairment: Evidence from behavioral and ERP data. Autism Research, 11: 1050-1061. https://doi.org/10.1002/aur.1954

      Kissine, M. (2021). Autism, constructionism, and nativism. Language 97(3), e139-e160. https://dx.doi.org/10.1353/lan.2021.0055.

      Kuhl, P.K. (2007), Is speech learning ‘gated’ by the social brain?. Developmental Science, 10: 110-120. https://doi.org/10.1111/j.1467-7687.2007.00572.x

      References:

      (1) Moreau, C. N., Joanisse, M. F., Mulgrew, J., & Batterink, L. J. (2022). No statistical learning advantage in children over adults: Evidence from behaviour and neural entrainment. Developmental Cognitive Neuroscience, 57, 101154. https://doi.org/10.1016/j.dcn.2022.101154

      (2) Boucher, J., Lewis, V., & Collis, G. M. (2000). Voice processing abilities in children with autism, children with specific language impairments, and young typically developing children. Journal of Child Psychology and Psychiatry, 41(7), 847-857. https://doi.org/10.1111/1469-7610.00672

      (3) Schelinski, S., Borowiak, K., & von Kriegstein, K. (2016). Temporal voice areas exist in autism spectrum disorder but are dysfunctional for voice identity recognition. Social Cognitive and Affective Neuroscience, 11(11), 1812-1822. https://doi.org/10.1093/scan/nsw089

      (4) Lin, I.-F., Yamada, T., Komine, Y., Kato, N., Kato, M., & Kashino, M. (2015). Vocal identity recognition in autism spectrum disorder. PLOS ONE, 10(6), e0129451.https://doi.org/10.1371/journal.pone.0129451

      Reviewer #2 (Public review):

      Summary:

      This article looks at differences in how the brain entrains to, or tracks, the rhythmic presentation of syllables and words in speech in infants at increased likelihood versus low likelihood for autism. The authors first sought to characterize how brain responses are modulated by learning the statistical probability of a given syllable following the one before it over the first two years of life. They then sought to identify at which stages of word learning infants with increased likelihood of autism showed difficulties, and whether those difficulties worsened over time. Finally, they sought to indicate whether infants' statistical learning and word learning abilities could predict later verbal skills. The authors found similar developmental trajectories of neural entrainment to syllables in infants at high and low likelihood for autism, but infants at high likelihood for autism had overall weaker syllable-level entrainment. Infants at high versus low likelihood for autism showed different developmental trajectories for word entrainment. Lower syllable entrainment in high-likelihood infants corresponded with poorer verbal outcomes, but word entrainment was not associated with verbal outcomes. Event-related potential responses to words and part words were positively associated with verbal outcomes, however, but only in low-likelihood infants.

      Strengths:

      Overall, the article provides rigorous statistical analysis of longitudinal EEG data to provide strong support for the claims that neural entrainment to syllable and word features of speech may be a useful marker for language development difficulties, particularly in infants at increased likelihood for neurodevelopmental disorders. The EEG data collection and preprocessing procedures are well within standards in the field. Readers should take care to note that authors indexed neural entrainment to speech using phase-locking values instead of spectral power.

      Weaknesses:

      While the statistical analyses are rigorous, a few of the components of the models are not clearly defined, and some corrections and thresholds for significance warrant further justification. Further, a few stimuli and participant details that could influence results are not specified. It is not clear whether all participants came from majority French-speaking families; differences in the amount of French language exposure (compared to other languages that may be spoken by a participant's family) could influence results. The standardized volume of the stimuli is also not included. As a result, readers should be encouraged to interpret that neural entrainment to speech features is likely a useful mechanism to explain differences in language development, while taking this interpretation with some caution.

      We thank the reviewer for these remarks.

      Regarding the amount of French exposure: while all participants were raised in primarily French-speaking environments (i.e., French as the dominant language at home and daycare), the parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. We did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism. The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      Regarding the volume of stimuli, they were played at 50cm distance with an intensity of 75dB. Both considerations have been included in the new version of the manuscript. In general, we moved most of the figures present in Supplementary material to figure supplements to improve readability.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Comments:

      Figure 6: The figure caption is not complete (there is no description for the right half of panel B).

      We thank the reviewer for this observation, figure 6 caption has been completed.

      Reviewer #2 (Recommendations for the authors):

      Broadly speaking, I would recommend reducing the number of abbreviations in this article, and I would recommend that the authors take care as to where these abbreviations are being introduced. Many of the abbreviated terms are defined in the Materials and Methods section, which is presented after the abbreviations are used in the main results.

      We acknowledge that our extensive use of abbreviations compromises the readability of the manuscript. Consequently, we have removed the following abbreviations:

      - SL (replaced by statistical learning)

      - LC (replaced by latent component)

      - ASD (replaced by autism)

      - TP (replaced by transition probability)

      - MEG (replaced by magneto-encephalogram)

      - MSEL (replaced by Mullen Scale of Early Learnings)

      The remaining abbreviations are:

      HL (high likelihood for autism), LL (low likelihood for autism), EEG (electroencephalogram), PLS-c (partial least square correlation), ERP (event-related potential), RND (random), STR (structured), BSR (bootstrap ratio), AIC (Akaike Information Criterion), PLV (phase locking value), DQ (developmental quotient), APSI (Autism Parent Screen for Infants).

      Moreover, we carefully reviewed how abbreviations were introduced and identified that PLS-c, STR and RND were not defined prior to the Method section. This oversight has been corrected in the reviewed manuscript.

      I would also recommend that the authors be careful with the structuring of the Introduction, particularly with their research questions and hypotheses. The article initially makes clear that the research questions are focused on the developmental trajectory of statistical learning, the levels of word learning that may differentiate high-likelihood versus low-likelihood infants, and the stability of those differences, and associations between statistical learning and various levels of word learning with verbal outcomes. The use of acoustic variability across syllables, while a valuable methodological tool, is somewhat presented as an additional research question, but not clearly stated or tested as such.

      We acknowledge that the introduction (particularly the paragraph from lines 173 to 184) may have implied that speaker variability across syllables was one of our primary research aims. We clarify here that speaker variability was introduced as a mean to increase task difficulty, particularly for high-likelihood (HL) participants, with the aim of amplifying the effect sizes in our analyses.

      To address this, we have removed the theoretical discussion on speaker variability in autism and typical development (lines 173–184) and explicitly stated that speaker variability was not a research question in this study. Crucially, our experimental design did not include a control condition without speaker variability, and thus we could not test its specific effects on statistical learning across age trajectories and groups.

      Page 6, lines 173-176 (pages and lines refer to the reviewed uploaded manuscript): “It is worth noting, however, that our study was not designed to isolate or quantify the specific impact of speaker variability on statistical learning, as the experimental design did not include a baseline control condition omitting this acoustic variation.”

      The authors do a nice job in the Materials & Methods explaining PLS-c and defining the latent components and bootstrapped ratios that will be shared in the Results. An additional brief iteration defining these statistical elements is needed at the beginning of the Results section.

      We thank the reviewer for their appreciation of our Method section. We agree that an additional iteration in the result section would improve readability. We added the following paragraph at the very beginning of the Result section, briefly defining PLS-c and its main statistical output (latent components and bootstrap ratios):

      Pages 6-7, lines 193-202: “Briefly, PLS-c is a data-driven multivariate modelling approach designed to identify significant patterns of electrode clusters (from a brain data matrix containing electrophysiological measures, here PLV) and their associations with “behavioral” variables (from a behavioral design matrix, here age-related parameters). Patterns of brain x behavior associations are called latent components, and their statistical significance is evaluated using permutation testing (n=1000, Bonferroni correction for number of components tested, alpha=.006). Brain and behavioral variables respective contributions to any significant latent component are tested with bootstrapping (500 random samples and replacement), with bootstrap ratios (BSR) greater than 2.3 indicating a stable contribution (for details, see the Materials and Methods section).”

      (1) Page 18 Line 576. The authors need to clarify whether participants were required to be in primarily French-speaking environments and whether there was a minimum amount of French language exposure that participants were required to have if they were exposed to additional languages besides French in their everyday life.

      The reviewer raises a valid concern regarding participants’ language exposure. In this study, all participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare. The parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. However, we did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism.

      The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      To include these considerations, Limitations and Material and methods sections were modified as follows:

      Page 19, lines 604-607: “Second, although all participants were primarily exposed to French, we did not quantify additional language exposure, precluding any analysis of its potential moderator effects on statistical learning in our groups and age-trajectories. However, prior work has reported no effect of bilingualism on auditory triplet segmentation in children (Yim & Rudoy, 2013).”

      Page 20, lines 630-631: “All participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare.”

      References:

      Weiss DJ, Schwob N, Lebkuecher AL. Bilingualism and statistical learning: Lessons from studies using artificial languages. Bilingualism: Language and Cognition. 2020;23(1):92-97. doi:10.1017/S1366728919000579

      Yim D, Rudoy J. Implicit statistical learning and language skills in bilingual children. J Speech Lang Hear Res. 2013 Feb;56(1):310-22. doi: 10.1044/1092-4388(2012/11-0243). Epub 2012 Aug 15. PMID: 22896046.

      (2) Page 18 Line 588. Further, the authors should clarify whether the 7 infants in the HL group, due to early parental concerns were defined by the 18-21-month APSI scores or by parental report prior to study enrollment.

      These 7 infants were recruited based on early parental concerns prior to intake. The APSI score at 18-21 months is only reported to provide an illustration of the amount of early autistic signs that were present in these 7 infants, and to provide an estimation of their probability to develop autism later on based on Sacrey et al., 2018 longitudinal study on the APSI predictive value. We agree with the reviewer that our phrasing suggests that the APSI was used as an inclusion criterion. We rephrased the page 20 lines 642-646 as follows:

      “The 7 other HL infants presented with early parental concerns for autism, based on parental report prior to enrollment. Their Autism Parent Screen for Infants (APSI) total score at their 18-21 months visit was 15.6±6.4, [8-22] range – a score greater than 8 reflecting a 63% positive predictive value for autism in HL populations.”

      (3) Page 20 Line 641. The authors should specify the volume of the stimuli.

      The volume of stimuli was reported in the main text (page 22, lines 695-696) as follows:

      “Stimuli were played on a Bose® Companion 2 Series III at a 50cm distance with an intensity of 75dB.”

      (4) I'd prefer Figure 1 to be reorganized slightly - at present, the placement of the arrows explaining the analysis steps is not intuitive.

      We addressed the reviewer’s comments (4) and (5) together as they both refer to Figure 1B.

      (5) Page 23 lines 718-719. I think it would be helpful to explicitly define each of the interaction variables included in the behavioral design matrix. Further, this matrix should be labeled consistently in both Figure 1B and in the main text.

      We refined figure 1B and its corresponding main text (in Methods section) for clarity. The arrows are now simpler and more parsimonious, labels (e.g., participant i, visit n, behavior design matrix and its parameters) are now standardized between the figure and the main text, and the interaction terms at lines 718-719 are explicitly defined.

      (6) Page 23 lines 726-731: It would be helpful to know whether applying a Bonferroni correction in addition to completing permutation testing is standard when evaluating latent components derived from PLS-c. The authors should also cite justification for a bootstrap ratio cutoff of 2.3 for defining stability.

      In PLS-c analyses, multiple comparisons correction across latent components and bootstrap ratio (BSR) thresholding at 2.3 are commonly adopted practices.

      - Correction for multiple comparisons in PLS-c: PLS-c performs singular decomposition of the data into latent components equal in number to the variables included in the behavior design matrix (7-9 in our study, depending on the inclusion of Verbal outcome as an input variable). Each latent component’s statistical significance is assessed through permutation testing, generating a null distribution for its singular value (Krishnan et al., 2011). Given the multiple tests (one permutation test per latent component), Type I error inflation must be addressed. Recent PLS-c studies commonly applied Bonferroni correction (default procedure in the myPLS toolbox, used by Zoeller et al., 2017, and Delavari et al, 2021), though FDR correction has also been used (Lombardo et al, 2018).

      - Stability threshold for bootstrap and replacement: Within each latent component, saliences’ stability (brain/behavior parameter contributions to each latent component) are evaluated using bootstrapping (Krishnan et al., 2011). The bootstrap ratio (BSR) of each parameter, calculated as the saliency divided by its bootstrap-derived standard error, functions analogously to a z-score under normality assumptions. The BSR can then be used to assess the stability of the saliency (i.e., how stable is its contribution to the latent component). BSR thresholds in the literature typically range from 1.96 to 3.0. Krishnan et al (2011) state that when BSR are “larger than 2 the corresponding saliences are considered significantly stable”. Delavari et al (2021) and our study used a 2.3 thresholding, corresponding to a 99.0% bootstrap confidence interval not crossing the zero line – roughly equivalent to a two-tailed p<.001. Lombardo et al (2018) used a looser threshold of 1.96, corresponding to a 95% confidence interval not crossing the zero line (~two-tailed p<.05), while Zöller et al (2017) used a more stringent 3.0 thresholding (~p<.001, or 99.9% confidence interval not crossing the zero line).

      Thus, our application of Bonferroni correction for multiple comparisons and our 2.3 BSR threshold aligns with established conventions.

      We added following lines in the manuscript:

      Page 25 lines 784-785: “Bonferroni correction was applied to account for multiple comparisons across the 9 tested latent components in the PLS-c, yielding an adjusted alpha of .006 (Zoeller et al, 2017; Delavari et al, 2021).”

      Page 25 lines 789-792: “BSR are analogous to Z-scores and can be used to assess the stability of the saliency. We considered BSR > 2.3 as stable, corresponding to a 99.0% bootstrap confidence interval not crossing zero – roughly equivalent to a two-tailed p<.001 (Delavari et al., 2021; Krishnan et al., 2011).”

      References:

      Delavari F, Sandini C, Zöller D, Mancini V, Bortolin K, Schneider M, Van De Ville D, Eliez S. Dysmaturation Observed as Altered Hippocampal Functional Connectivity at Rest Is Associated With the Emergence of Positive Psychotic Symptoms in Patients With 22q11 Deletion Syndrome. Biol Psychiatry. 2021 Jul 1;90(1):58-68. doi: 10.1016/j.biopsych.2020.12.033. Epub 2021 Jan 18. PMID: 33771350.

      Lombardo, M.V., Pramparo, T., Gazestani, V. et al. Large-scale associations between the leukocyte transcriptome and BOLD responses to speech differ in autism early language outcome subtypes. Nat Neurosci 21, 1680–1688 (2018). https://doi.org/10.1038/s41593-018-0281-3

      Daniela Zöller, Marie Schaer, Elisa Scariati, Maria Carmela Padula, Stephan Eliez, Dimitri Van De Ville. Disentangling resting-state BOLD variability and PCC functional connectivity in 22q11.2 deletion syndrome. NeuroImage, Volume 149, 2017, Pages 85-97, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2017.01.064

      Anjali Krishnan, Lynne J. Williams, Anthony Randal McIntosh, Hervé Abdi, Partial Least Squares (PLS) methods for neuroimaging: A tutorial and review, NeuroImage, Volume 56, Issue 2, 2011, Pages 455-475, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2010.07.034

      (7) I have a few minor grammar/formatting recommendations for the authors as well:

      (a) Should the Geneva Autism Cohort be capitalized? At present, it is not.

      We agree with the reviewer’s suggestion, and we capitalized the Geneva Autism Cohort in the main text (page 18, line 571)

      (b) Page 24, line 750. Do the authors mean that the data was re-referenced to average?

      The preprocessed data is not average-referenced (see section Data pre-processing). Therefore, both for neural entrainment computation and ERPs, the data were average-referenced.

      (c) It would be nice to have a figure of the actual ERP for each condition and age group.

      We agree that PLS-c can be difficult to interpret without the raw actual ERPs on which it was modelled. We direct the reviewer to supplementary figure S6 at page 59, which displays the raw ERPs for each condition (part-word, word, and their subtraction) per age group. Supplementary figures S7-8 at pages 60-61 further illustrate topographical ERPs for each group (high and low likelihood for autism). We deemed these figures too extensive for the main text. Instead, the most relevant ERP topographies are presented in Figures 4-6 to facilitate PLS-c interpretation.

    1. eLife Assessment

      In this valuable study, the authors identified a rare population of Nestin-expressing cells within the external granule layer of the early postnatal mouse cerebellum. They demonstrated that these cells are distinct from Sox2+ progenitors and can give rise to medulloblastoma. Collectively, the findings provide convincing evidence that this unique Nestin+ population is susceptible to oncogenic transformation and may underlie the preferential emergence of Sonic hedgehog-driven medulloblastomas.

    2. Reviewer #1 (Public review):

      Summary:

      GCPs, which drive postnatal cerebellar growth and can give rise to SHH-MB, are not uniform. The authors show that GCPs include a rare Nestin-expressing subpopulation with distinct molecular features. This subpopulation is spatially restricted, enriched for stem cell-like properties, and shows a high competency for tumor formation comparable to larger GCP pools, with tumors preferentially arising in the posterior-lateral cerebellum. Overall, the findings indicate that SHH-MB might originate preferentially from this small, tumor-competent Nestin-expressing GCP subset.

      Strengths:

      (1) The authors use a breadth of approaches from histology, mouse genetics, and single-cell RNA sequencing.

      (2) Throughout, this paper uses very elegant genetic approaches, such as the double Nes-FlpoER; Atoh1-FSF-Cre; LSL-Smo-M2, to generate tumors only from Atoh1+; Nes+ double-positive cells. This intersectional genetic experiment makes for a very clear answer.

      (3) The findings reported in this manuscript are valuable since they reveal a novel GCP subpopulation defined by spatial and molecular identity. Some of their experiments suggest that these cells could represent the main cell-of-origin of SHH MB. The experiments are carefully performed, and the evidence is convincing.

      Weaknesses or elements that could be improved:

      (1) A transgenic Nestin-CFP mouse is used in this study. However, it is not clear whether CFP accurately reflects the Nestin protein. Figure 1: After the promoter is turned off, these cells might remain positive for CFP for longer than they are positive for Nestin, due to CFP protein stability. Is the Nestin protein present in these cells? Nestin double immunofluorescence with CFP and Sox2 and Barhl1 could be performed to address this. Related to this comment, it is also important to note that this is a rat promoter transgene. So the transgene might not reflect exactly the endogenous Nestin expression.

      (2) Could the posterior restriction of Nestin-CFP be due to the timing (P1) at which the authors looked? In other words, if they look earlier, would the authors see Nestin-CFP cells more anterior?

      (3) Since only one medulloblastoma mouse model (Smo-M2) is used to conclude that "the Nes-expressing GCP population in the normal cerebellum is transcriptionally closer to SHH MB tumor cells than the remainder of the GCPs", the findings might not apply to other SHH-MB models. This should be mentioned.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors studied transgenic reporter mice to profile Nestin expression in the postnatal mouse cerebellum. They discovered a small population of Nestin+; Atoh1+ granule neuron precursors (GNPs) in the external granule cell layer (EGL). Using immunostaining, qPCR, and RNA-sequencing, the authors showed that these Nestin+ cells are not identical to Sox2+ cells (e.g., the majority of Nestin+ cells are Sox2-). Using various mouse genetic strategies, including an elegant intersectional strategy that specifically targets Nestin+; Sox2+ cells, the authors showed that Nestin+ cells are capable of initiating Sonic hedgehog (SHH) medulloblastoma when they express the SmoM2 allele that drives constitutively active SHH signaling. Lastly, the authors profiled the transcriptomes of these cells and showed that they display enriched stem cell genes and are closer to the transcriptomes of GNP-like cells in medulloblastoma compared to Nestin- GNPs in the developing cerebellum.

      Strengths:

      (1) The comprehensive mouse genetics experiments, in combination with immunostaining, lineage tracing, and RNA-seq studies, provided compelling evidence that rare Nestin+ cells are present in the EGL, predominantly at the posterior lateral cerebellum in early postnatal mice.

      (2) The intersectional genetics experiment unequivocally show that Nestin+; Atoh1+ cells can be oncogenically transformed by SmoM2, leading to SHH medulloblastoma.

      (3) The more stem cell-like transcriptomic features of the Nestin+ GNPs compared to Nestin- GNPs provide support for the heterogeneity of this transient progenitor cell population, with implications for development, congenital diseases, and tumors from the cerebellum.

      Weaknesses:

      Main comments:

      My main concern relates to whether these Nestin+; Atoh1+ cells are restrictively localized in the EGL. Both the title "A Rare Nestin-Expressing Granule Cell Precursor Subpopulation Underlies SHH Medulloblastoma Formation" and what the authors described throughout the manuscript propose that Nestin+; Atoh1+ cells in the EGL are the cell-of-origin of SHH medulloblastoma. To definitively conclude this, the authors need to comprehensively analyze all regions of the developing cerebellum.

      Most importantly, are Nestin+; Atoh1+ cells present in the rhombic lip? Are there any rhombic lip cells genetically labeled in their intersectional mouse mutants (e.g., the Atoh1Frt-Cre/+; Nes-FlpoER; R26LSL-SsmoM2/+ mice)?

      If Nestin+; Atoh1+ cells are present at non-EGL regions in the developing cerebellum, the authors would have to reconsider many of their conclusions and also the title of this paper.

      Additional comments:

      (1) To investigate Nestin expression, the authors used Nes-CFP transgenic mice expressing CFP from promoter/enhancer sequences from the rat Nes gene (Encinas et al., 2006). Given that Nestin expression is of central importance for this study, it is important to validate that these reporter mice faithfully report Nestin protein expression (e.g., by co-labeling CFP with Nestin antibody and systemically comparing signals throughout the cerebellum, ideally in several developmental stages).

      (2) The authors mostly presented immunostaining data of the cerebellum from P1 mice. It is important to systematically profile the appearance and disappearance of these Nestin+, Atoh1+ cells in mouse cerebellum across developmental stages (e.g., embryonic, early, and late postnatal stages).

      (3) How different is the proliferative ability of the Nestin+ versus Nestin- GNPs at various developmental stages? Also, the difference between EdU+; Barhl1+; Nestin+ and EdU+; Barhl1+; Nestin- cells is quite small despite statistical difference (Figure 1N). Do the authors think this very small EdU incorporation difference can translate into a biological difference (in developmental and/or disease context)?

      (4) Lines 145-147: "Compared to double-negative cells, Atoh1 and Nes were significantly higher in the double-positive fraction, supporting the identity of the cells as a previously unrecognized rare population of GCPs at P1 that expresses both the GCP marker Atoh1 and ventricular zone marker Nes." Nestin is not a ventricular zone marker. This should be rephrased.

      (5) In Figure 3, the authors showed mouse survival data and concluded that Nes-driven and Atho1-driven SHH medulloblastoma models show similar tumor penetrance. This is not an entirely accurate description of their data. The Nes-SmoM2 mice displayed significantly longer survival compared to the Atoh1-SmoM2 mice (Figure 3B). This conclusion needs to be revised.

      (6) In Figure 5, the authors showed that genes enriched in cluster 10 included Sox2, Nes, Wls, and Wnt1, while Neurod1 and Rbfox3 were preferentially expressed in the other GCP clusters. They conclude that cluster 10 represents a less differentiated, more stem-like GCP state, potentially positioned upstream in the lineage hierarchy. While these few markers are useful, it is more informative to formally support this conclusion by comparing the stem cell transcriptomic signature (using a larger gene list) between cluster 10 and other GCPs.

      (7) In Figure 5, the authors performed gene ontology analysis and showed that cluster 10 is enriched for biological processes linked to WNT signaling and proposed that this molecular profile supports their identity as a transient, developmentally plastic population within the GCP lineage related to the rhombic lip. The authors are recommended to use an orthogonal approach (i.e., immunostaining to compare nuclear localization of beta-Catenin) to validate their transcriptome-based finding.

    1. eLife Assessment

      This manuscript presents important new findings showing that the transcription factor and regulator of lipid and glucose metabolism PPARγ is methylated by the enzyme SETD6. The data convincingly demonstrate that methylation of PPARγ by SETD6 regulates its function in controlling transcription of lipid storage and metabolism genes and thereby modulates lipid accumulation in liver cells. This work uncovers a new role for post-translational modulation by lysine methylation in controlling transcription factor activity and lipid metabolism in a relevant physiological context.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript from the Levy lab, the authors investigate whether SETD6 regulates hepatic lipid accumulation through direct methylation of PPARγ. They show that SETD6 binds and mono-methylates PPARγ at K170 and provide evidence that this modification enhances PPARγ occupancy at target promoters, promotes expression of lipid metabolism genes, as well as facilitates lipid droplet accumulation in HepG2 cells. The authors also find a positive feedback loop or circuit in which PPARγ activates SETD6 transcription in a methylation-dependent manner, thereby reinforcing this lipogenic program. Overall, the work presents a novel SETD6-PPARγ regulatory axis linking lysine methylation to transcriptional control of lipid storage genes, with possible relevance to NAFLD-associated biology.

      In all, I find this to be an important paper that describes and advances a new regulatory pathway that has significance to human health and disease. It would also be of interest to a broad audience. That said, there are also some concerns that the authors should address, as outlined below.

      Major concerns (pertains to rigor - highest priority)

      (1) Overall, the work presented is of high quality and the data nicely support the conclusions; however, a few panels should be strengthened that have missing controls or information:<br /> a. The co-IP panel in Fig. 1B lacks a lane where HA SETD6 is expressed without PPARγ. This control is needed to verify that the SEDT6-HA signal depends on PPARγ.<br /> b. In Fig. 1C, the authors should show that the co-IP works in both directions (include IP for PPARγ/blot for SETD6). I am a bit confused also over the labeling with IP on the left and on top of the panel next to the beads label. More importantly, the data would be stronger if the authors take advantage of a deletion line to validate the co-IP is specific to the presence of both.<br /> c. The same IP labeling issue exists for Fig 3B (label is on the same and on top).<br /> d. Antibody information (e.g., where the pan-methyl Ab comes from and at what dilutions they are used at) is missing.

      Nice to have experiments (medium priority - strongly consider)

      (2) A missing gap is how K170me1 contributes to DNA binding and gene transcription. One possibility is that methylation enhances the DNA binding activity of PPARγ. Given the authors have all of the reagents, it would be possible to perform a gel shift assay (or other approach) with and without SETD6-mediaetd methylation. Is DNA binding affected/enhanced?

      (3) Along these lines, I wonder if there is another possibility: could SETD6-mediated methylation of PPARγ drive SETD6-PPARγ interaction? In other words, in the K170R, is SETD6 still even associated with PPARγ, and this interaction is required for promoter recruitment? Alternatively, would a catalytic dead version of SETD6 fail to associate with PPARγ? Currently, no experiments test the impact of an unmethylatable version of PPARγ or catalytic dead version of SETD6 on SETD6-PPARγ interaction or SETD6 recruitment to promoters.

      Minor concerns (text and figure display)

      (4) The text has multiple typos and grammatical errors.

      Comments on revised version.

      Great job on addressing the comments. It is a nice study.

    3. Reviewer #2 (Public review):

      Summary:

      In this work, the authors investigated the regulation of the transcription factor PPARγ by the post-translational modification lysine methylation The data demonstrate that the lysine methyltransferase SETD6 targets PPARγ for methylation using biochemical and cell-based assays. Methylation of PPARγ occurs in its DNA binding domain, and the authors demonstrate that loss of methylation limits PPARγ chromatin binding, particularly to lipid storage and metabolism genes promoters. As a physiological output, the authors demonstrate that deletion of SETD6 and loss of PPARγ methylation also disrupt lipid droplet accumulation in hepatocytes. In addition, the authors uncover a positive feedback loop in which SETD6 methylation of PPARγ also regulates its binding to the SETD6 promoter and expression of the gene.

      Strengths:

      One of the key strengths of this manuscript is the novelty of the findings in terms of identifying a new mode of regulation of PPARγ that modulates its chromatin association in cells and thereby regulating lipid metabolism genes. The authors nicely combine biochemical studies of SETD6 activity with cell-based assays investigating PPARγ and SETD6 function in regulating lipid storage. Data supporting this conclusion is largely convincing and frequently, multiple assays are used to provide sufficient support to the conclusions. This work therefore expands regulatory modes of PPARγ and identifies a new target for SETD6, an enzyme that targets a number of other transcription factors. Furthermore, the regulatory loop that controls SETD6 expression via PPARγ methylation is likely important for understanding SETD6 function in different cell types that have high levels of lipid accumulation or regulation. The gene expression and lipid accumulation assays are useful for testing the physiological outcome of loss of SETD6 activity or PPARγ methylation directly. In the revised manuscript, the authors have added useful structural modeling to better define potential roles of methylation of PPARγ in regulating its function, particularly relative to DNA binding, and to better define the physical interaction between PPARγ and SETD6.

      Weaknesses:

      The revised manuscript substantially improved on the presentation of the data and broadened the discussion to provide more context to both the role of SETD6 and to elaborate on potential mechanisms by which methylation impacts PPARγ function and under what physiological conditions this interaction and regulation is important. This improves and strengthens the manuscript and its impact overall.

      Comments on revised version.

      The authors addressed all of my major concerns following this round of review and I do not have additional recommendations. The presentation of the manuscript including text and figures is improved compared to the previous version. I have updated my public review to reflect these changes.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript from the Levy lab, the authors investigate whether SETD6 regulates hepatic lipid accumulation through direct methylation of PPARγ. They show that SETD6 binds and monomethylates PPARγ at K170, and provide evidence that this modification enhances PPARγ occupancy at target promoters, promotes expression of lipid metabolism genes, as well as facilitates lipid droplet accumulation in HepG2 cells. The authors also find a positive feedback loop or circuit in which PPARγ activates SETD6 transcription in a methylation-dependent manner, thereby reinforcing this lipogenic program. Overall, the work presents a novel SETD6PPARγ regulatory axis linking lysine methylation to transcriptional control of lipid storage genes, with possible relevance to NAFLD-associated biology.

      In all, I find this to be an important paper that describes and advances a new regulatory pathway that has significance to human health and disease. It would also be of interest to a broad audience. That said, there are also some concerns that the authors should address, as outlined below.

      We are grateful to the reviewer for the positive feedback and appreciation of our work.

      Major concerns (pertains to rigor - highest priority)

      (1) Overall, the work presented is of high quality, and the data nicely support the conclusions; however, a few panels should be strengthened that have missing controls or information:

      (a) The co-IP panel in Figure 1B lacks a lane where HA SETD6 is expressed without PPARγ. This control is needed to verify that the SEDT6-HA signal depends on PPARγ.

      We thank the reviewer for this valuable suggestion. The overexpression co-immunoprecipitation experiment referred to by the reviewer has been moved to Supplementary Figure S1 in the revised manuscript. In this experiment, immunoprecipitation was performed using an anti-FLAG antibody to pull down FLAG-tagged PPARγ. In the absence of FLAG-PPARγ, the anti-FLAG immunoprecipitation does not recover a bait protein, and therefore HA-SETD6 is not expected to be specifically immunoprecipitated. Thus, an HA-SETD6-only condition would primarily serve as a negative control for the anti-FLAG pull-down rather than provide additional information regarding the specificity of the interaction.

      Importantly, in the revised manuscript we have substantially strengthened the evidence supporting the SETD6–PPARγ interaction by adding two independent complementary experiments. First, we included a reciprocal endogenous co-immunoprecipitation (new Figure 2B), demonstrating that endogenous PPARγ co-immunoprecipitates with endogenous SETD6. Second, we added an independent proximity ligation assay (PLA) (new Figure 2D), which further confirms the interaction between SETD6 and PPARγ in cells. Together with the in vitro binding assay presented in Figure 2A, these orthogonal approaches provide compelling evidence for the specificity of the SETD6–PPARγ interaction. Therefore, we believe that the requested HA-SETD6-only control would not provide additional mechanistic insight beyond the comprehensive validation now included in the revised manuscript.

      (b) In Figure 1C, the authors should show that the co-IP works in both directions (include IP for PPARγ/blot for SETD6). I am a bit confused also over the labeling with IP on the left and on top of the panel next to the beads label. More importantly, the data would be stronger if the authors took advantage of a deletion line to validate that the co-IP is specific to the presence of both.

      We thank the reviewer for this helpful suggestion. We have revised the manuscript to strengthen the evidence supporting the endogenous interaction between SETD6 and PPARγ. Specifically, we now include a reciprocal endogenous co-immunoprecipitation (new Figure 2B), demonstrating that endogenous SETD6 co-immunoprecipitates with endogenous PPARγ and, conversely, that endogenous PPARγ co-immunoprecipitates with endogenous SETD6. These reciprocal experiments independently validate the specificity of the interaction.

      In addition, we have revised the figure layout and labeling to more clearly distinguish the immunoprecipitating antibody from the bead control, thereby addressing the reviewer's concern regarding the presentation of the co-immunoprecipitation data.

      Although we did not perform the co-immunoprecipitation in a depletion/knockout background, we believe that the combination of reciprocal endogenous co-immunoprecipitation (New Figure 2B), the independent proximity ligation assay (New Figure 2D), and the direct in vitro binding assay (Figure 2A) provides multiple orthogonal lines of evidence supporting a specific interaction between SETD6 and PPARγ.

      (c) The same IP labeling issue exists for Figure 3B (label is on the same and on top).

      We have revised the labeling in Figure 3B to clearly distinguish the immunoprecipitating antibody from the bead control, thereby improving the clarity of the figure.

      (d) Antibody information (e.g., where the pan-methyl Ab comes from and at what dilutions they are used at) is missing.

      We thank the reviewer for pointing this out. We have now added the missing information regarding the pan-methyl antibody to the Materials and Methods section, including the supplier, catalogue number, and experimental conditions used. Specifically, the pan-methyl antibody used in this study was purchased from Abcam (ab23366) and was used at a 1:500 dilution for western blot analysis and 2 μg per reaction for immunoprecipitation experiments.

      Nice to have experiments (medium priority - strongly consider)

      (2) A missing gap is how K170me1 contributes to DNA binding and gene transcription. One possibility is that methylation enhances the DNA-binding activity of PPARγ. Given that the authors have all of the reagents, it would be possible to perform a gel shift assay (or other approach) with and without SETD6-mediated methylation. Is DNA binding affected/enhanced?

      We thank the reviewer for raising this important point. To investigate whether K170 methylation could directly affect PPARγ binding to DNA, we performed structural modeling based on the available co-crystal structure of PPARγ bound to DNA (PDB: 3DZU). As shown in the new Figure 6F, K170 is positioned near the DNA-binding region; however, the modeled K170me1 side chain is predicted to face away from the DNA interface and does not appear to sterically interfere with the PPARγ–DNA interaction. In addition, modeling of multiple K170me1 rotamers did not suggest any major disruption of the DNA-bound conformation.

      These observations suggest that K170 methylation is unlikely to directly alter the intrinsic DNA-binding affinity of PPARγ. In contrast, our ChIP-qPCR experiments demonstrate that K170 methylation positively regulates PPARγ occupancy at target promoters in cells. Together, these findings support a model in which K170 methylation promotes PPARγ chromatin association and transcriptional activity through mechanisms other than direct modulation of DNA binding, such as altered cofactor recruitment or protein–protein interactions.

      We agree with the reviewer that future biochemical approaches, including EMSA/gel shift assays or quantitative DNA-binding measurements, will be valuable to directly determine whether K170 methylation affects the intrinsic DNA-binding affinity of PPARγ. We have incorporated this new structural analysis and the corresponding discussion into the revised manuscript.

      (3) Along these lines, I wonder if there is another possibility: could SETD6-mediated methylation of PPARγ drive SETD6-PPARγ interaction? In other words, in the K170R, is SETD6 still even associated with PPARγ, and this interaction is required for promoter recruitment? Alternatively, would a catalytic dead version of SETD6 fail to associate with PPARγ? Currently, no experiments test the impact of an unmethylatable version of PPARγ or a catalytic dead version of SETD6 on SETD6-PPARγ interaction or SETD6 recruitment to promoters.

      We thank the reviewer for this insightful suggestion. To address whether SETD6 catalytic activity is required for its association with PPARγ, we performed an additional PLA experiment comparing SETD6 WT and the catalytic mutant SETD6 Y285A. As shown in the revised New Figure 2D, both SETD6 WT and SETD6 Y285A showed comparable proximity to PPARγ in cells, indicating that SETD6 catalytic activity is not required for the physical association between SETD6 and PPARγ.

      These findings support a model in which SETD6 first recognizes and binds PPARγ independently of its catalytic activity. Subsequent methylation of PPARγ at K170 is therefore likely to regulate the downstream functional consequences of this interaction, including enhanced chromatin occupancy and transcriptional activation, rather than the initial SETD6–PPARγ association itself.

      In addition, we generated using AlphaFold a structural model of the SETD6–PPARγ complex as a supportive visualization (New figure S2). Given the limited confidence of the prediction, we interpret this model cautiously and include it in the Supplementary Information rather than the main figures. We agree with the reviewer that future studies examining SETD6 recruitment to PPARγ target promoters and the effect of the PPARγ K170R mutant on SETD6–PPARγ association will further refine the molecular mechanism.

      Minor concerns (text and figure display)

      (4) The text has multiple typos and grammatical errors, and there are some issues with the figure display.

      We thank the reviewer for this comment. We carefully revised the manuscript to correct typographical and grammatical errors throughout the text and also addressed the figure display issues noted by the reviewer.

      Reviewer #2 (Public review):

      Summary:

      In this work, the authors investigated the regulation of the transcription factor PPARγ by the post-translational modification lysine methylation. The data demonstrate that the lysine methyltransferase SETD6 targets PPARγ for methylation using biochemical and cell-based assays. Methylation of PPARγ occurs in its DNA binding domain, and the authors demonstrate that loss of methylation limits PPARγ chromatin binding, particularly to lipid storage and metabolism gene promoters. As a physiological output, the authors demonstrate that deletion of SETD6 and loss of PPARγ methylation also disrupt lipid droplet accumulation in hepatocytes. In addition, the authors uncover a positive feedback loop in which SETD6 methylation of PPARγ also regulates its binding to the SETD6 promoter and expression of the gene.

      Strengths:

      One of the key strengths of this manuscript is the novelty of the findings in terms of identifying a new mode of regulation of PPARγ that modulates its chromatin association in cells and thereby regulates lipid metabolism genes. The authors nicely combine biochemical studies of SETD6 activity with cell-based assays investigating PPARγ and SETD6 function in regulating lipid storage. Data supporting this conclusion is largely convincing, and frequently, multiple assays are used to provide sufficient support to the conclusions. This work therefore expands regulatory modes of PPARγ and identifies a new target for SETD6, an enzyme that targets a number of other transcription factors. Furthermore, the regulatory loop that controls SETD6 expression via PPARγ methylation is likely important for understanding SETD6 function in different cell types that have high levels of lipid accumulation or regulation. The gene expression and lipid accumulation assays are useful for testing the physiological outcome of loss of SETD6 activity or PPARγ methylation directly.

      We thank the reviewer for his/her positive feedback on the manuscript.

      Weaknesses:

      The data presented in the manuscript are largely convincing in support of the authors' conclusions; however, there are some errors in the presentation of the figures and some issues in the text that would benefit from editing. Furthermore, there are some important questions not fully addressed in the results or discussion. 

      It would be great if the authors could speculate more on the diverse roles of SETD6 in methylated transcription factors and/or provide more context regarding the conditions that are likely to support methylation of PPARγ by SETD6. 

      We thank the reviewer for this important suggestion. In the revised Discussion, we expanded the manuscript to better place our findings within the broader context of SETD6-mediated regulation of transcription factors. Previous studies from our group and others demonstrated that SETD6 methylates multiple chromatin-associated transcriptional regulators, including RelA, E2F1, TWIST1, and BRD4, thereby modulating transcriptional selectivity, chromatin occupancy, and cofactor recruitment. We now discuss the possibility that SETD6 functions as a context-dependent signalling integrator that selectively regulates transcription factor activity through lysine methylation under distinct physiological conditions.

      In addition, we expanded the Discussion regarding potential cellular contexts that may favour PPARγ methylation by SETD6. Because PPARγ activity is strongly induced during lipid overload and fatty acid exposure, conditions associated with steatosis and metabolic stress may enhance the functional importance of SETD6-dependent methylation. We also discuss the possibility that chromatin accessibility, ligand-dependent activation of PPARγ, and metabolic signaling pathways may collectively influence the formation and stability of the SETD6–PPARγ complex.

      Also, while a potential cross-talk between methylation and phosphorylation is described in the discussion, it would be great to provide more structural insight into how this might regulate DNA binding of PPARγ and/or discuss whether there are other possibilities given the location of the target lysine in the DNA binding domain.

      We thank the reviewer for this valuable suggestion. To provide additional structural insight into the potential interplay between methylation and phosphorylation within the PPARγ DNA-binding domain, we performed structural modeling based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU). As shown in the new Figure S6, the modeled K170me1 side chain is predicted to face away from the DNA interface and does not introduce steric clashes with DNA, suggesting that K170 methylation is unlikely to directly alter the DNA-binding interface through steric effects. In contrast, phosphorylation of the neighboring residue T166 is predicted to introduce multiple intramolecular steric clashes within the DNA-binding domain. These structural changes could influence the local conformation or dynamics of the DNA-binding domain and thereby indirectly modulate PPARγ DNA binding or transcriptional activity.

      In addition, as suggested by the reviewer, we expanded the Discussion to consider alternative mechanisms by which K170 methylation may regulate PPARγ function. While our ChIP-qPCR experiments demonstrate that K170 methylation positively regulates PPARγ chromatin occupancy at target promoters, the structural modeling suggests that this effect is unlikely to arise from direct steric modulation of the DNA interface. Instead, K170 methylation may influence chromatin occupancy by regulating protein–protein interactions, cofactor recruitment, local conformational dynamics, or other chromatin-associated mechanisms. We have incorporated these new structural analyses and the expanded discussion into the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The KEGG panels are too low resolution to read.

      We thank the reviewer for this comment. We improved the resolution of the KEGG pathway enrichment panels and, as suggested by the reviewer (see below), we separated the upregulated and downregulated gene sets to improve clarity and readability. These changes are now reflected in the revised new Figures 5C, 5D, 6B and 6C.

      (2) Figure 3 panel D is hard to visualize; a better version or repeat experiment is needed.

      We thank the reviewer for this comment. To address this concern, we replaced the original Figure 3D with a new independent experiment that more clearly demonstrates the methylation of endogenous PPARγ by SETD6. We believe that the new data provide substantially stronger evidence and improve the clarity of the revised manuscript.

      (3) There is a typo in Figure 2A (line present in "S" of Signal).

      We thank the reviewer for pointing out this typo. The error in Figure 2A has now been corrected in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Overall, the experiments and analyses presented are sufficient to support the conclusions and interpretations of the work. However, there are some issues of presentation and writing that are worth addressing. These comments are listed below:

      (1) My only substantial recommendation is to improve the discussion to provide more context to the findings in terms of both the larger role of SETD6 in methylating transcription factors (some of whom also regulate its expression) and the potential modes through which PPARγ DNA binding activity could be regulated by methylation.

      We thank the reviewer for this insightful suggestion. In the revised manuscript, we substantially expanded the Discussion to better place our findings within the broader context of SETD6mediated regulation of transcription factors. We now discuss previous studies demonstrating that SETD6 methylates multiple chromatin-associated transcriptional regulators, including RelA, E2F1, TWIST1, and BRD4, thereby modulating chromatin occupancy, cofactor recruitment, and transcriptional selectivity. We further propose that SETD6 functions as a context-dependent signaling regulator that integrates distinct cellular pathways through the selective lysine methylation of transcription factors.

      In addition, we expanded the Discussion regarding the potential mechanisms by which PPARγ K170 methylation regulates transcriptional activity. We incorporated new structural modeling based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU), which suggests that K170 methylation is unlikely to directly alter the DNA-binding interface through steric effects, whereas phosphorylation of the neighboring residue T166 may induce intramolecular steric clashes within the DNA-binding domain. We also expanded the Discussion to consider alternative mechanisms by which K170 methylation may regulate PPARγ function, including modulation of chromatin occupancy, local conformational dynamics, protein–protein interactions, and recruitment of transcriptional cofactors or chromatin-associated proteins. Finally, we discuss that future biochemical studies will be important to determine whether K170 methylation also influences the intrinsic DNA-binding affinity of PPARγ.

      Can the authors incorporate any other published structural data to speculate on the role of methylation or describe more about how it is expected that the methylation-phosphorylation crosstalk modulates DNA binding? This type of discussion would better highlight the potential importance of this new modification on PPARγ.

      We thank the reviewer for this important suggestion. In the revised manuscript, we expanded the Discussion and incorporated additional structural analyses (new Figure 6F and new figure S6) based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU). K170 is positioned within the DNA-binding domain, between the two zinc-finger motifs that mediate DNA recognition and stabilization on PPRE-containing DNA. Our structural modeling predicts that the K170me1 side chain is oriented away from the DNA interface and does not introduce steric clashes with DNA, suggesting that methylation is unlikely to directly alter the DNA-binding interface through steric effects. Nevertheless, lysine methylation can influence protein surface properties, protein– protein interactions, and recognition by regulatory binding partners, raising the possibility that K170 methylation modulates PPARγ chromatin occupancy or promoter selectivity through indirect mechanisms.

      In contrast, structural modeling predicts that phosphorylation of the neighboring residue T166 introduces multiple intramolecular steric clashes within the DNA-binding domain. These clashes could alter the local conformation or dynamics of the DNA-binding domain and thereby indirectly influence PPARγ–DNA interactions and transcriptional activity. Together, these observations suggest that K170 methylation and T166 phosphorylation may represent a regulatory crosstalk that fine-tunes PPARγ function through distinct structural mechanisms.

      We further expanded the Discussion to consider additional, non-mutually exclusive mechanisms by which K170 methylation may regulate PPARγ function, including modulation of chromatin occupancy, protein–protein interactions, cofactor recruitment, and stabilization of transcriptional complexes at target genes. While these possibilities require further mechanistic investigation, we agree with the reviewer that these structural considerations highlight the potential regulatory importance of this newly identified PPARγ modification.

      (2) In the gene expression experiments presented in Figure 5, it would be useful if the GO terms were described as enriched in either the up- or down-regulated gene sets. From the way it is presented, it is not clear if specific categories of genes are found enriched in those upregulated or downregulated upon KO of SETD6. This is also true for the experiments presented in Figure 6 regarding the PPARγ mutant.

      We thank the reviewer for this important suggestion. In the revised manuscript, we separated the pathway enrichment analyses into upregulated and downregulated gene sets for both the SETD6 knockout RNA-sequencing experiments (Figure 5) and the PPARγ WT versus K170R mutant analysis (Figure 6). This revision provides improved clarity regarding which biological pathways are positively or negatively associated with SETD6 depletion or disruption of PPARγ K170 methylation. The updated KEGG enrichment analyses are now presented in the revised new Figures 5C, 5D, 6B and 6C.

      In addition, if there is a significant overlap in genes misregulated in both mutants, this could be shown in the figure.

      We thank the reviewer for this suggestion. We compared the differentially expressed genes identified in the SETD6 KO cells and the PPARγ K170R mutant cells to evaluate the extent of overlap between the two datasets. However, we did not observe a substantial or statistically significant overlap in misregulated genes under the thresholds used in our analysis. Therefore, we decided not to include this comparison in the revised figure. Nevertheless, both datasets consistently showed enrichment for pathways associated with lipid metabolism and PPAR signaling, supporting a functional connection between SETD6 and PPARγ-mediated transcriptional regulation.

      (3) There are some typos and grammatical errors throughout the work, and it should be carefully edited. One example is the following heading: PPARγ K170 methylation by SETD6 regulates mediates lipid droplets formation.

      We thank the reviewer for this comment. The manuscript was carefully revised to correct typographical and grammatical errors throughout the text. In particular, the heading mentioned by the reviewer was corrected in the revised manuscript.

      Minor errors in the figures:

      (1) Formatting issue in the labeling for the x-axis of Figure 1E.

      We thank the reviewer for pointing out this formatting issue. The labeling of the x-axis in Figure 1E has been corrected in the revised manuscript.

      (2) Figure 2B - HA-SETD6 should be labeled as minus for the first lane.

      We thank the reviewer for pointing this out. We corrected the labeling in Figure 2B (now Figure S1), and the first lane is now properly indicated as negative for HA-SETD6.

      (3) Figure 2D - Labeling needs improvement. Is this FLAG-PPARγ? "NC" was not defined in the legend. If negative control, what type?

      We thank the reviewer for this comment. We improved the labeling and figure legend of Figure 2D for clarity. Specifically, we now clearly indicate that the experiment was performed using Flag-PPARγ, and we defined “NC” in the legend as the negative control condition. In addition, during the revision process we noticed that the original PLA experiment was performed in HeLa cells rather than HepG2 cells, as previously indicated. This has now been corrected throughout the revised manuscript.

      (4) Figure 3 - The CRSIPR control should be described somewhere in the legend or the methods.

      We thank the reviewer for this comment. We clarified the description of the CRISPR control cells in the Materials and Methods section. Specifically, we now explicitly state that the CRISPR control (CT) cells were generated using the empty lentiCRISPR vector without SETD6-targeting sgRNAs.

      (5) Figure 5B and 6A - The legend is not labeled nor defined in the text- fold-change, log2 fold change, z score?

      We thank the reviewer for pointing this out. We revised the figure legends for Figures 5B and 6A to explicitly define the heatmap scale and normalization method. Specifically, we now indicate that the heatmaps represent normalized gene expression values displayed as Z-scores.

    1. eLife Assessment

      This study presents valuable data suggesting that ATP-induced modulation of alveolar macrophage (AM) functions is associated with NLRP3 inflammasome activation and enhanced phagocytic capacity. While the in vivo and in vitro data reveal an interesting phenotype, the evidence provided is incomplete and does not fully support the paper's conclusions. Additional investigations would be of value in complementing the data and strengthening the interpretation of the results. This study should be of interest to immunologists and the mucosal immunity community.

    2. Reviewer #1 (Public review):

      Summary:

      Alveolar macrophages (AMs) are key sentinel cells in the lungs, representing the first line of defense against infections. There is growing interest within the scientific community in the metabolic and epigenetic reprogramming of innate immune cells following an initial stress, which alters their response upon exposure to a heterologous challenge. In this study, the authors show that exposure to extracellular ATP can shape AM functions by activating the P2X7 receptor. This activation triggers the relocation of the potassium channel TWIK2 to the cell surface, placing macrophages in a heightened state of responsiveness. This leads to the activation of the NLRP3 inflammasome and, upon bacterial internalization, to the translocation of TWIK2 to the phagosomal membrane, enhancing bacterial killing through pH modulation. Through these findings, the authors propose a mechanism by which ATP acts as a danger signal to boost the antimicrobial capacity of AMs.

      Strengths:

      This is a fundamental study in a field of great interest to the scientific community. A growing body of evidence has highlighted the importance of metabolic and epigenetic reprogramming in innate immune cells, which can have long-term effects on their responses to various inflammatory contexts. Exploring the role of ATP in this process represents an important and timely question in basic research. The study combines both in vitro and in vivo investigations and proposes a mechanistic hypothesis to explain the observed phenotype.

      Weaknesses:

      Although these findings are convincing and intrinsically interesting, they do not support the conclusion that ATP induces trained immunity. By definition, trained immunity refers to long-lasting metabolic and epigenetic reprogramming initiated by a primary stimulus. Importantly, some of these changes persist after the cells have returned to a basal activation state, thereby generating an altered response upon secondary stimulation (https://doi.org/10.1038/s41590-020-00845-6). In the present study, the data demonstrate a sustained increase in inflammasome activation and enhanced microbicidal activity for up to seven days following ATP exposure. While this sustained activation is noteworthy as well as metabolic shift, it does not demonstrate the existence of trained immunity. The terms priming or sustained activation would therefore be more appropriate than trained immunity.

      Similarly, the observation of increased chromatin accessibility at inflammasome-related genes is expected given the robust activation of this pathway. The presence of open chromatin at these loci does not, by itself, constitute evidence for long-term trained immunity. The authors should therefore be cautious with their terminology and avoid overinterpreting their findings.

      The authors have revised the manuscript to address the comments raised during the first rounds of review. However, several figures, figure legends, and methodological sections still require additional adjustments and clarification.

      The Methods section remains incomplete and requires substantial revision. For instance, the methodology used to quantify immune cell populations presented in Figure 2 is still not described. It is not stated how immune cells were isolated and identified (e.g. flow cytometry from lung tissue). No information is provided regarding tissue digestion, cell isolation procedures, or gating strategy (presumably by flow cytometry). These details are essential and should be included, together with the corresponding gating strategy and absolute cell numbers.

      There are inconsistencies throughout the manuscript. For example, the authors report n = 3 in the figure legend 2 and 3 independent experiments, whereas 3 or 4 points are represented in the graphs. This discrepancy is unclear and should be clarified.

      Overall, while the study addresses an interesting biological question, the manuscript would benefit from substantial revision prior to publication. In particular, clarifications and improvements regarding the methodology, data presentation, and interpretation are required to strengthen the rigor and reproducibility of the conclusions. Several of the conclusions extend beyond what is directly supported by the data. In particular, the interpretation that these findings demonstrate trained immunity should be revised, and additional methodological clarifications and corrections are required.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Please include an uptake control (early time point) or time‑course to distinguish phagocytosis from intracellular killing.

      We agree that distinguishing uptake from intracellular killing is important for interpreting bactericidal assays. Due to a recent transition, we are unable to conduct additional early‑time‑point assays. To address this transparently, we have revised the manuscript to clarify that our measurements represent overall bacterial load reduction, reflecting the combined effects of uptake and killing.

      The normalization as ‘fold killing’ is non‑standard; please report absolute CFU (log scale).

      We have retrieved the raw data and re‑expressed all bactericidal activity measurements as absolute CFU. All relevant figures, legends, and text have been updated accordingly.

      Authors report quantification of cytokine concentrations, yet no information is provided regarding how these measurements were performed.

      Cytokine concentrations were quantified by ELISA. We have now added details regarding assay kits, sample preparation, and detection parameters to the Methods section.

      While the choice of IL‑1β and IL‑6 is straightforward, the focus on IL‑18 requires explicit justification.

      IL‑18 is a macrophage‑associated pro‑inflammatory cytokine with established links to inflammasome activation and trained immunity pathways, providing a clear justification for its inclusion.

      The methodology used to quantify immune cell populations presented in Figure 2 is not described.

      We have added a detailed description of the flow cytometry methodology, including staining and gating strategies.

      Immune cell quantification would be expected in the context of the challenge experiment as well.

      While we agree that such data would be valuable, additional mouse experiments cannot be performed because the animals used in the challenge model are not currently available during the transition period. If feasible, we are exploring ex vivo flow cytometry data from TWIK2 mutant versus wild‑type macrophages.

      AMs are not considered recruited immune cells; this should be corrected.

      We have corrected this in the figure legend and throughout the manuscript.

      The authors report n = 5 for the survival curves in the figure legend, whereas n = 7 is stated in the Methods section.

      We have corrected the sample size to ensure consistency between the figure legend and Methods.

      ATAC‑seq peaks are referred to as ‘genes’ and ‘differentially expressed genes’.

      We have corrected the terminology in the manuscript. ATAC‑seq identifies differentially accessible chromatin regions, which are then annotated to the nearest downstream gene.

      In Figure 7, trained WT and Nlrp3-/- mice display similar levels of bacterial clearance. How should this result be interpreted?

      A portion of this phenotype is via the normalization of phagocytosis to ‘fold killing’. Presentation of the raw CFU data shows a trend towards reduced bacterial clearance in Nlrp3-/- mice.

      Reviewer #2 (Public review):

      Sample numbers for experiments 1, 2, and 6 are not provided.

      We have added explicit n values for all experiments and verified their accuracy against the original records.

      The Discussion would benefit from a clear summary of study caveats.

      We agree and have added a dedicated paragraph outlining key Caveats as described.

      Specific identities of DEGs are not provided; only pathway enrichment is shown.

      We have now included a supplementary table listing the differentially expressed genes identified in our analysis.

      Controls for subcellular fractionation and dye microscopy should be included.

      Controls for subcellular fractionation have added in figure 3B and controls for dye microscopy have now been added in supplementary figure 1.

      The text states that protease inhibitors diminish ATP‑induced training effects, but the figure does not show significance.

      We have re‑examined the data and updated the figure to include statistical testing where appropriate.

    1. eLife Assessment

      This manuscript reports valuable results on the role of MDC1 and Treacle in DSB repair in rDNA repeats. It has been previously established that MDC1 is replaced by Treacle as the main adaptor in the nucleolar DNA damage response. This work provides convincing evidence that MDC1 is required for the recruitment of RAD51 and BRCA1 to DSBs in rDNA. The work involves multiple MDC1 knockout models and establishes that RFN8-RNF168 act downstream of MDC1in parallel with the RAP80-ABRAXAS pathway in the recruitment of the HR machinery to nucleolar DSBs.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The revised version of the manuscript addresses the previous concerns. Importantly, a major role of the RAP80-ABRAXAS pathway is now demonstrated in the recruitment of BRCA1, PALB2, and RAD51 to nucleolar DSBs.]

      This study elucidates the molecular linkage between the mobilization of damaged rDNA from the nucleolus to its periphery and the subsequent repair process by HDR. The authors demonstrate that the nucleolar adaptor protein Treacle mediates rDNA mobilization, and the MDC1-RNF8-RNF168 pathway coordinates the recruitment of the BRCA1-PALB2-BRCA2 complex and RAD51 loading. This stepwise regulation appears to prevent aberrant recombination events between rDNA repeats. This work provides compelling evidence for the recruitment of the Treacle-TOPBP1-NBS1 complex to rDNA DSBs and demonstrates the critical role of MDC1 in the rDNA damage response. There are some issues with the over-interpretation of results as described subsequently. Some aspects could be strengthened, for example, a potential role of the RAP80-Abraxas axis, the origin of the repair synthesis (HDR vs. NHEJ), and a direct comparison of the RNF8 and RNF168 recruitment in the absence or presence of MDC1.

    3. Reviewer #2 (Public review):

      Summary:

      DNA double-strand breaks (DSB) in repeated DNA pose a challenge for repair by homologous recombination (HR) due to the potential of generating chromosomal aberrations, especially involving repeats on different chromosomes. This conceptual caveat led to a long-held notion that HR is not active in repeated DNA, which was disproven in groundbreaking work by Chiolo showing in Drosophila that DSBs in pericentromeric repeats are mobilized to the nuclear periphery for repair by HR. A similar mechanism operates in mouse cells, as shown by the Gautier laboratory, but the mobilization goes to the nucleolar periphery, called nucleolar caps. In this manuscript, the authors reexamine the role of MDC1 in the mobilization of DSBs in rDNA in human cells. Previous work has shown that MDC1 is replaced by Treacle, the gene associated with Treacher Collins syndrome 1, in its role as the main adaptor of the DNA damage response, and these results are confirmed here. The novelty of this contribution lies in the discovery that MDC1 is required downstream in the recruitment of BRCA1 and RAD51 to nucleolar DSBs that were mobilized to the nucleolar cap. Using multiple MCD knockout models and DSBs induced by the nuclease PpoI, which cleaves at nuclear sites as well as in the 28S rDNA, convincingly documents this role of MDC1 and shows that it acts upstream of the RNF8-RNF168 ubiquitylation axis. Using a proxy assay of co-localization of EdU incorporation at DSBs (gammaH2AX), evidence is provided that MDC1 is required for HR in rDNA. MDC1 was not required for RAD51 recruitment to IR-induced foci, but it is unclear whether this is related to the different DSB chemistry (enzymatic versus IR) or to the localization of the DSB (rDNA versus unique sequence genome).

      Strengths:

      (1) The manuscript is well-written, and the experimental evidence is nicely presented.

      (2) Multiple MDC1 knockout models are used to validate the results.

      (3) Convincing back-complementation data clarify the relationship between MDC1 and RNF8.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      This study elucidates the molecular linkage between the mobilization of damaged rDNA from the nucleolus to its periphery and the subsequent repair process by HDR. The authors demonstrate that the nucleolar adaptor protein Treacle mediates rDNA mobilization, and the MDC1-RNF8-RNF168 pathway coordinates the recruitment of the BRCA1-PALB2-BRCA2 complex and RAD51 loading. This stepwise regulation appears to prevent aberrant recombination events between rDNA repeats. This work provides compelling evidence for the recruitment of the Treacle-TOPBP1-NBS1 complex to rDNA DSBs and demonstrates the critical role of MDC1 in the rDNA damage response. There are some issues with the over-interpretation of results as described subsequently. Some aspects could be strengthened, for example, a potential role of the RAP80-Abraxas axis, the origin of the repair synthesis (HDR vs. NHEJ), and a direct comparison of the RNF8 and RNF168 recruitment in the absence or presence of MDC1.

      We thank the reviewer for the positive assessment of our work and for the constructive suggestions. We agree that certain aspects of the manuscript required clarification, in particular the potential contribution of the RAP80–ABRAXAS pathway, the interpretation of the repair synthesis assay, and the role of RNF168 recruitment. We have addressed these points experimentally where feasible and have revised the manuscript accordingly to avoid overinterpretation.

      Reviewer #1 (Recommendations for the authors):

      Major comments

      (1) In Figures 4C, 4D, and S4B-D, BRCA1 and RAD51, recruitment to nucleolar caps is partially reduced upon RNF168 depletion. Despite this, the authors broadly conclude that recruitment mainly depends on the MDC1-RNF8-RNF168 pathway. Since the RAP80-Abraxas pathway may also contribute, as briefly mentioned in the Discussion, siRNA knockdown of Abraxas would help clarify the relative roles of these two pathways.

      We thank the reviewer for this important suggestion. To directly address the potential contribution of the RAP80–ABRAXAS pathway, we depleted RAP80 by siRNA and analysed BRCA1 and RAD51 recruitment to nucleolar caps following I-PpoI-induced rDNA damage.

      Strikingly, RAP80 depletion strongly impaired the formation of both BRCA1 and RAD51 nucleolar caps in two independent cell lines (U2OS and RPE1) (new Figure 5). These findings demonstrate that the RAP80–ABRAXAS pathway plays a critical role in BRCA1 recruitment at nucleolar caps.

      Together with our observation that RNF168 depletion only partially reduces BRCA1 and RAD51 recruitment, these results indicate that both RNF168-dependent and RAP80–ABRAXAS-dependent pathways contribute to HDR factor recruitment downstream of RNF8-mediated chromatin ubiquitylation.

      We have revised the model Figure (Figure 9) and the Results and Discussion sections accordingly to reflect this dual-pathway model and to avoid overemphasising the contribution of RNF168.

      (2) In Figure 7C, the EdU-γH2AX PLA assay detects DNA synthesis at rDNA breaks, but it remains unclear whether this signal reflects HDR- or NHEJ-mediated repair. Since MDC1 functions upstream of the DSB repair pathway choice, the observed reduction in PLA signal upon MDC1 depletion does not necessarily reflect impaired HDR alone. Synchronizing cells in G2 or using cell cycle markers would help clarify the repair context and strengthen the interpretation.

      We thank the reviewer for raising this important point. We agree that the EdU–gH2AX PLA assay does not exclusively report on HDR-mediated DNA synthesis and may also capture other forms of repair-associated DNA synthesis.

      In the revised manuscript, we have therefore tempered our interpretation and now describe this assay more cautiously as a readout of DNA synthesis at sites of rDNA damage, rather than as a direct measure of HDR activity.

      Importantly, our conclusion that MDC1 promotes HDR factor recruitment at nucleolar caps is based primarily on the reduced accumulation of BRCA1, PALB2, and RAD51, which are well-established markers of HDR. The PLA assay is now presented as supportive evidence for ongoing DNA synthesis at these sites rather than as a definitive indicator of HDR.

      We agree that further experiments, such as cell cycle synchronization or the use of phase-specific markers, would help to more precisely define the repair context, and we have included this point in the Discussion.

      (3) The authors propose that MDC1 is essential for RNF8-RNF168 recruitment, specifically at nucleolar rDNA breaks. A side-by-side comparison of RNF8 or RNF168 localization in the presence and absence of MDC1, with IR-treated conditions, would provide important validation of this model. Including representative images in Figure S2C would further support the claim.

      We agree with the reviewer that a direct analysis of RNF8 and RNF168 recruitment in the presence and absence of MDC1 would provide valuable mechanistic insight. We therefore attempted to address this experimentally.

      However, despite testing multiple antibodies, we were unable to obtain specific and reproducible signals for RNF8 and RNF168 at nucleolar caps, precluding a reliable analysis of its recruitment under these conditions.

      Given this technical constraint, we have revised the manuscript to avoid overinterpretation regarding direct RNF168 recruitment and instead focus on functional readouts of downstream ubiquitylation-dependent signalling, such as BRCA1 and RAD51 accumulation.

      We note that the requirement for MDC1 in BRCA1 and RAD51 recruitment at nucleolar caps is consistent with a role of MDC1 upstream of RNF8-dependent chromatin ubiquitylation, in line with its established function at IR-induced DSBs.

      Minor comments:

      (1) The legend for Figure 8 should more clearly explain the proposed mechanism and include concise titles or descriptions for each sub-panel.

      We agree with the reviewer that the model should be described in the Figure legend. We have thus updated the model to accommodate the new data and wrote a legend that concisely explains the proposed model. We do not think that titles for each sub-panel are required. Instead, we separately referred to the sub-panels in the legend.

      (2) Typos:

      (a) Page 8: PRE1 MDC1, as "RPE1 MDC1;

      (b) S3 Figure legend: Dhermacon";

      (c) Page 29: "80.103"-please clarify or correct.

      We thank the reviewer for pointing out these errors. These have been corrected in the revised manuscript

      Reviewer #2 (Public review):

      Summary:

      DNA double-strand breaks (DSB) in repeated DNA pose a challenge for repair by homologous recombination (HR) due to the potential of generating chromosomal aberrations, especially involving repeats on different chromosomes. This conceptual caveat led to a long-held notion that HR is not active in repeated DNA, which was disproven in groundbreaking work by Chiolo showing in Drosophila that DSBs in pericentromeric repeats are mobilized to the nuclear periphery for repair by HR. A similar mechanism operates in mouse cells, as shown by the Gautier laboratory, but the mobilization goes to the nucleolar periphery, called nucleolar caps. In this manuscript, the authors reexamine the role of MDC1 in the mobilization of DSBs in rDNA in human cells. Previous work has shown that MDC1 is replaced by Treacle, the gene associated with Treacher Collins syndrome 1, in its role as the main adaptor of the DNA damage response, and these results are confirmed here. The novelty of this contribution lies in the discovery that MDC1 is required downstream in the recruitment of BRCA1 and RAD51 to nucleolar DSBs that were mobilized to the nucleolar cap. Using multiple MCD knockout models and DSBs induced by the nuclease PpoI, which cleaves at nuclear sites as well as in the 28S rDNA, convincingly documents this role of MDC1 and shows that it acts upstream of the RNF8-RNF168 ubiquitylation axis. Using a proxy assay of co-localization of EdU incorporation at DSBs (gammaH2AX), evidence is provided that MDC1 is required for HR in rDNA. MDC1 was not required for RAD51 recruitment to IR-induced foci, but it is unclear whether this is related to the different DSB chemistry (enzymatic versus IR) or to the localization of the DSB (rDNA versus unique sequence genome).

      Strengths:

      (1) The manuscript is well-written, and the experimental evidence is nicely presented.

      (2) Multiple MDC1 knockout models are used to validate the results.

      (3) Convincing back-complementation data clarify the relationship between MDC1 and RNF8.

      Weaknesses:

      (1) The recruitment of BRCA2 was not directly demonstrated. This caveat could be recognized, as IF for BRCA2 is challenging.

      (2) PpoI also induces DSBs in the non-rDNA genome. These DSBs would be an ideal control to establish nucleolar specificity of the events described and clarify whether the difference between IR and PpoI is the chemical structure of the DSB or the location of the DSB.

      We thank the reviewer for the positive and insightful evaluation of our work. We appreciate the recognition of the conceptual advance and the robustness of our experimental approaches. We have carefully considered the reviewer’s suggestions and have revised the manuscript to clarify interpretation where appropriate, particularly regarding BRCA2 recruitment and the specificity of I-PpoI-induced DNA damage. Where possible, we have also added new analyses to strengthen the conclusions.

      Reviewer #2 (Recommendations for the authors):

      (1) The claim that the BRCA1-PALB2-BRCA2 is recruited (abstract, end of results section, discussion page 15) should be qualified as BRCA2 recruitment was not directly demonstrated.

      We thank the reviewer for this important point. We agree that BRCA2 recruitment was not directly demonstrated in our study, as reliable immunofluorescence detection of BRCA2 remains technically challenging.

      We have therefore revised the manuscript throughout (Abstract, Results, and Discussion) to avoid overstatement and now refer more precisely to the recruitment of BRCA1, PALB2, and RAD51, rather than implying direct recruitment of a BRCA1–PALB2–BRCA2 complex.

      We note that BRCA2 function is supported indirectly by the observed RAD51 loading, which depends on BRCA2 activity. However, we have clarified this point to ensure that our conclusions remain fully supported by the presented data.

      (2) The temporal sequence established in Figure 1, 1hr BRCA1 and 2 hrs PALB2, argues against recruitment of a stable BRCA1-PALB2-(BRCA2) complex. This should be acknowledged.

      We thank the reviewer for this insightful observation. We agree that the temporal separation between BRCA1 accumulation (1 h) and PALB2/RAD51 recruitment (2 h) argues against the recruitment of a pre-assembled, stable BRCA1–PALB2–BRCA2 complex.

      We have revised the manuscript to reflect this interpretation and now describe the recruitment of HDR factors as a sequential process rather than as the assembly of a pre-formed complex. This is consistent with current models in which BRCA1 promotes subsequent PALB2 and BRCA2 recruitment, ultimately leading to RAD51 loading.

      (3) The model predicts that MDC1-KO cells are proficient for transcriptional repression after nucleolar DSB induction. Has this been tested?

      We did not specifically test this in the current work, but previous results published by our group revealed that siRNA-mediated depletion of MDC1 in human cells had a minimal effect on rDNA transcriptional inhibition after DNA damage (Larsen et al., 2024).

      (4) The nuclear PpoI DSBs could be analyzed as a specificity control, and clarify whether the difference between IRIF and PpoI DSBs relates to the DSB chemistry or location.

      We thank the reviewer for this important point. We agree that I-PpoI induces DNA breaks both within rDNA repeats and at additional genomic loci.

      To address this, we have now analysed the formation of gH2AX-positive nucleolar caps and non-nucleolar gH2AX foci over time following I-PpoI expression (new Figure 1–figure supplement 2). We find that nucleolar caps form rapidly and are prominent at early time points, whereas gH2AX foci accumulate more gradually.

      These results indicate that nucleolar caps and non-nucleolar DNA damage responses can be distinguished both spatially and temporally, and support the use of nucleolar caps as a specific readout for rDNA damage in our study.

      In addition, we note that RAD51 recruitment to IR-induced foci is not affected by MDC1 loss, suggesting that the requirement for MDC1 in RAD51 loading is specific to nucleolar rDNA breaks rather than reflecting differences in DSB chemistry alone. We have clarified this point in the Discussion.

      Additional points:

      (5) Page 4 top: Shieldin.

      Corrected.

      (6) The general reader will be interested to learn about the connection of the Treacle function with Treacher Collins syndrome. Maybe a paragraph could be added to discuss this?

      We thank the reviewer for this suggestion. We agree that the relationship between Treacle and Treacher Collins syndrome may be of interest to a broad readership. Since the developmental pathology of Treacher Collins syndrome is currently thought to arise primarily from impaired ribosome biogenesis and nucleolar dysfunction rather than defective nucleolar DNA damage signalling, we felt that an extensive discussion would be beyond the scope of the present study. We have, however, added a brief statement introducing Treacle as the product of the TCOF1 gene mutated in Treacher Collins syndrome and noting that whether its DNA damage response function contributes to disease pathology remains an open question.

      (7) Figure 7: A short explanation could be added as to why hypoxia conditions were chosen for the p53-deficient cell lines.

      We thank the reviewer for pointing this out. We have added a brief explanation in the figure legend to clarify that hypoxia conditions were used to stabilise replication stress and enhance detection of DNA repair intermediates in p53-deficient cells.

      (8) A short statement on whether the repair of nuclear DBS is affected by Treacle could be added.

      We thank the reviewer for this interesting point. While our study focuses on nucleolar DNA damage, we did not observe evidence that Treacle is required for the repair of non-nucleolar DSBs. We have added a brief statement in the Discussion to clarify that Treacle appears to function specifically in the nucleolar DNA damage response.

    1. eLife Assessment

      This study offers valuable insights into the role of post-translational modifiers, specifically SUMO2ylation at K81 in p66Shc, and its impact on endothelial function through reactive oxygen species. A series of compelling experiments demonstrated that lysine 81 of p66Shc is the site of SUMO2 conjugation, which is crucial for mitochondrial localization and essential for S36 phosphorylation, leading to specific pathological effects. The combination of cell overexpression and animal studies provides solid data supporting this mechanistic link.

    2. Reviewer #2 (Public review):

      Summary:

      The manuscript titled "p66Shc Mediates SUMO2-induced Endothelial Dysfunction" by Kumar et al. builds upon established literature demonstrating that both p66Shc and SUMOylation are essential players in nitric oxide (NO)-mediated endothelial vascular homeostasis and development (PMID: 10580504, 28760777, and 35187108).

      In this study, the authors uncover a novel mechanism showing how the SUMO2ylation of p66Shc drives reactive oxygen species (ROS) production in endothelial cells. Specifically, they identify Lysine 81 (K81) as the critical residue on p66Shc conjugated to SUMO2, proving it is essential for the protein's mitochondrial localization.

      The authors convincingly demonstrate that:

      p66Shc is actively SUMO2ylated at the K81 site in cellular models.

      Phosphorylation at Serine 36 (S36) is significantly reduced upon the loss of this critical SUMOylation site.

      Conclusion:

      Overall, this study provides strong evidence for a novel regulatory axis in endothelial cells. It successfully opens the door for further dissection of the complex mechanistic crosstalk between three key post-translational modifications on p66Shc: S36 phosphorylation, K81 SUMO2ylation, and acetylation.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a role of sumoylation at K81 in p66Shc which affects endothelial dysfunction. This explores a new mechanism for understanding the role of PTMs in cellular processes.

      Strengths:

      The experiments are well planned and the results are well represented.

      Vascular tonality experiments were carried out nicely, given the amount of time and effort one needs to put in to get clean results from these experiments.

      Weaknesses:

      (1) The production of ROS has been measured in a very superficial way.

      The term "ROS" confers a plethora of chemical species which exerts different physiological effects on different cells and situations.

      Mitochondria through one of the source, but not the only source of ROS production. Only measuring ROS with mitosox do not reflect the cellular condition of ROS in a specific condition. I would suggest authors consider doing IF of oxidative stress specific markers, carbonyl group and also, maybe, Amplex red for determining average oxidative stress and ros production in the cells.

      As suggested, we employed an additional ROS-sensitive probe, H<sub>2</sub>DCFDA, which revealed an overall increase in intracellular ROS levels upon SUMO2 overexpression; this effect was reversed by knockdown of p66Shc. In addition, we performed the Amplex Red assay on conditioned media to assess extracellular ROS release. SUMO2 overexpression did not alter ROS levels detected in the media, whereas a paradoxical increase was observed following p66Shc knockdown.

      Author response image 1.

      Amplex Red assay performed in HUVECs with and without knockdown of p66Shc expressing SUMO2 (Ad-SUMO2) or a control virus (Ad-LacZ).

      Amplex Red predominantly detects hydrogen peroxide (H<sub>2</sub>O<sub>2</sub>), which may originate from NADPH oxidases or be generated through the dismutation of superoxide. In contrast to superoxide, H₂O₂ is relatively stable, membrane-permeable, and well recognized as a second messenger in endothelial signaling, where it modulates kinase and phosphatase activity and supports physiological vascular functions. Thus, the elevated H<sub>2</sub>O<sub>2</sub> detected following p66Shc knockdown may represent a signaling-competent redox state rather than a pathological increase in oxidative stress. Moreover, the SUMO2-p66Shc-mediated increase in mitochondrial ROS may be efficiently buffered by cellular antioxidant defense mechanisms, thereby limiting detectable extracellular ROS release.

      (2) 8-OHG signal seems very confusing in Figure 7E. 8-ohg is supposed to be mainly in the nucleus and to some extent in mitochondria. The signal is very diffused in the images. I would suggest a higher magnification and better resolution images for 8-ohg. Also, the VWF signal is pretty weak whereas it should be strong given the staining is in aorta. Authors should redo the experiments.

      We have provided a better image for figure 7E. We repeated the staining with another antibody for vWF which showed stronger signal for vWF.

      (3) PCA analysis is quite not clear. Why is there a convergence among the plots? Authors should explain. Also, I would suggest that the authors do the analysis done in Figure 8B again with R based packages. IPA, though being user-friendly, mostly does not yield meaningful results and the statistics carried out is not accurate. Authors should redo the analysis in R or Python whichever is suitable for them.

      We thank the reviewer for their valuable feedback and insightful suggestions.

      Regarding the PCA analysis, the observed convergence among the data points reflects the underlying biological similarity between the samples within each group. Given the relatively low abundance and limited number of quantified peptides in our dataset, the variance captured by PCA is modest, leading to partial overlap between groups. This convergence is therefore likely due to shared biological characteristics and inherent sample variability, rather than technical issues.

      In response to the reviewer’s suggestion on pathway analysis, we have re-performed the analysis using R-based approaches. Specifically, we utilized established pipelines PROGENy-based signaling pathway activity inference. These methods provide statistically robust and reproducible results. The updated analyses and corresponding figures have been included in the revised manuscript, replacing the previous IPA-based results (Fig. 8). We believe these additions strengthen the interpretation of signaling pathway alterations in our dataset.

      (4) The MS analysis part seems pretty vague in methods. Please rewrite.

      We have revised the methodology for MS.

      Reviewer #2 (Public review):

      Summary:

      The article builds on the earlier work that both p66Shc and SUMOylation are essential nitric oxide (NO) based development of endothelial vasculature (PMID: 10580504; 28760777 and 35187108). The current manuscript brings forward a finding of how SUMO2ylation of p66Shc mediated ROS production which is essential for endothelial cells. They further identify that lysine 81 of p66Shc is the residue which is conjugated to SUMO2 and is crucial for mitochondrial localization. They further show that K81 SUMO2ylation is essential for S36 phosphorylation.

      Strengths:

      Convincingly shows that p66Shc is SUMO2ylated on lysine 81 in cells and also shows that the phosphorylation (serine 36) reduces upon loss of this critical SUMOylation site.

      Weaknesses:

      All the experiments performed here are in overexpression background therefore, it would be crucial to show that p66Shc is SUMO2ylated at physiological levels.

      As detecting endogenous SUMO2-p66Shc is technically challenging considering the almost 92% homology between SUMO2 and SUMO3 and the absence of a specific antibody to detect p66Shc, we generated a custom-made antibody which can detect SUMO2-p66Shc (YenZym, CA). Using this antibody, we performed immunoprecipitation which showed endogenous SUMO2-p66Shc at a molecular weight higher than p66Shc suggesting the SUMO2 modification of p66Shc at physiological level.

      Reviewer #3 (Public review):

      Summary:

      The authors set out to determine how SUMO2 impairs endothelial function through direct modification of the protein p66Shc. p66Shc is known to promote reactive oxygen species production, and here the authors demonstrate that SUMO2 modifies p66Shc at lysine-81, resulting in increased phosphorylation, mitochondrial translocation. These are prosed to mediate the detrimental effects of SUMO2 in a mouse model of hyperlipidemia.

      Strengths:

      A major strength of this work is the multi-pronged approach combining biochemical assays, proteomic analyses, and a genetically modified mouse model expressing a SUMOylation resistant mutant of p66Shc. These experiments comprehensively illustrate that lysine-81 SUMOylation of p66Shc is necessary for the observed endothelial dysfunction in hyperlipidemic conditions.

      Weaknesses:

      One notable weakness is that the link between the observed cellular changes and the ultimate in vivo phenotype remains only partially explored. While the authors successfully show that p66ShcK81R knockin mice are protected from endothelial dysfunction in a hyperlipidemic context, additional experiments characterizing the broader tissue-specific roles, or examining further endothelial assays in vivo, would strengthen the mechanistic conclusions. It would also be beneficial to see more direct evaluations of p66Shc subcellular localization in the protective knockin mice to complement the proteomic findings.

      We agree with the reviewer’s suggestion. However, due to limited resources, we could not pursue additional studies.

      Despite these gaps, the data broadly support the authors' main conclusions. The authors lay out a plausible mechanistic pathway for how hyperlipidemia and increased global SUMOylation can converge on the oxidative stress pathway to provoke vascular dysfunction.

      The likely impact of this work on the field is noteworthy. Beyond clarifying how a single post-translational modification event can influence the pathophysiology of endothelial cells, the study provides a model for investigating broader roles of SUMO2 in other cardiovascular conditions and highlights the importance of identifying additional SUMOylation sites and their downstream impact.

      In conclusion, by demonstrating the direct SUMOylation of p66Shc at lysine-81 and linking that modification to endothelial dysfunction in a hyperlipidemic mouse model, this paper offers valuable insights into how broadly acting post-translational modifiers can evoke specific pathological effects.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please rearrange the figures. It is very hard to follow.

      We rearranged some of the figures for a better presentation.

      Reviewer #2 (Recommendations for the authors):

      **Please note that I am not an expert of mouse work therefore I will not be able to comment on the mouse work (Figure 7).

      The following suggested changes major concerns will strengthen the work, and the minor concerns will increase the paper's accessibility to eLife's broad scientific community.

      Major concerns.

      (1) All the work done here is based on overexpression studies, therefore it will be very useful to show that p66Shc gets SUMO2ylated under physiological conditions using SUMO-Trap beads or using tandem SUMO-Interacting Motifs (as shown in Silva-Ferrada et al., 2013: https://doi.org/10.1038/srep01690).

      We have demonstrated endogenous SUMO2 conjugation of p66Shc under physiological conditions by immunoprecipitation using a custom-generated antibody specific for SUMO2-p66Shc. This approach allows detection of SUMO2ylated p66Shc without reliance on overexpression systems, thereby confirming that SUMO2 modification of p66Shc occurs endogenously.

      (2) The authors claim that K81 is the only site of SUMO2ylation and based on the HUVEC experiments (Fig 3C, D) it appears that p66Shc K81R mutant still gets SUMO2ylated indicating that there are other residues which can get SUMO2ylated. Do you see the other sites getting SUMO2ylated in your mass spec data?

      Our mass spectrometry analysis did not identify SUMO2 modification on lysine residues other than K81. However, we acknowledge that SUMOylation detected in vitro on recombinant protein may differ from SUMOylation occurring in a cellular context, where protein conformation, interacting partners, and local enzyme availability can influence modification patterns. These differences may account for the appearance of multiple SUMOylated p66Shc species in cell lysates despite the absence of additional SUMOylation sites in the mass spectrometry dataset. Our emphasis on K81 is based on its localization within the CH2 domain of p66Shc, a region unique to the p66 isoform. Previous studies have demonstrated that post-translational modifications within the CH2 domain, such as phosphorylation, are critical for activation of the oxidative and pro-apoptotic functions of p66Shc. Accordingly, we focused our mechanistic analyses on SUMO2ylation at K81. Nevertheless, we do not exclude the possibility that additional lysine residues within the PTB or CH1 domains of p66Shc may also undergo SUMOylation in cells. Importantly, our functional data support the conclusion that SUMO2ylation at K81 is a key regulatory modification driving the oxidative activity of p66Shc.

      (3) In Figure 4 the authors claim that SUMO2ylation at K81 is a prerequisite for S36 phosphorylation, however the phosphorylation changes are marginal (Figure 4 A, E) and why do they not see the higher molecular weight bands in the western blots.

      We appreciate the critique and agree with the reviewer that our data do not provide direct evidence that K81 SUMO2ylation and S36 phosphorylation occur simultaneously on the same p66Shc molecule. Rather, our findings suggest that SUMO2ylation at K81 may facilitate or enhance S36 phosphorylation, but is not an absolute requirement for this modification. Accordingly, we have revised the wording in the main text to avoid implying a strict prerequisite relationship and to more accurately reflect the magnitude of the observed changes in S36 phosphorylation.

      (4) The authors further claim that the SUMO2ylation at K81 effects S36 phosphorylation which has consequences for mitochondrial localization. However, K81R still translocate to mitochondria which is indicative of SUMO2ylation is important but not necessary (Fig 5A, C).

      We agree with the reviewer’s observation and have revised the main text accordingly, as described in our prior response, to clarify that K81 SUMO2ylation facilitates, but is not essential for, S36 phosphorylation and mitochondrial translocation.

      (5) To know if the K81 SUMO2ylation had any effect on phosphorylation it would be interesting to see how the phosphomimic (S36D) mutant along with/without K18R will behave in mitochondrial localization experiments.

      We agree that examining the double mutant (S36D/K81R) would provide valuable mechanistic insight into the relationship between K81 SUMO2ylation and S36 phosphorylation in regulating mitochondrial localization. Although we are currently unable to perform these experiments due to constraints in manpower and resources, we have acknowledged this as a limitation in the Discussion and plan to pursue these studies in future work to further validate the mechanism.

      Minor Concerns

      (a) In Figure 1A the authors claim that SUMO2ylation of p66Shc promotes ROS production however they do not include any well-known ROS induced proteins in the western blots such as KEAP1 or NRF2 or any other marker.

      We agree that NRF2 is a well-established transcriptional regulator of antioxidant responses. However, the primary aim of Figure 1A was to directly examine the effect of SUMO2ylation on p66Shc-mediated ROS production. While NRF2 and other ROS-responsive proteins reflect downstream adaptive responses, they do not provide a direct measure of ROS generation. Therefore, we focused on measuring SUMO2-induced changes in cellular ROS levels. To further strengthen this conclusion, we have performed additional complementary ROS assays, which are now included in the revised manuscript.

      (b) In Figure 2D, the concentration of Anacardic acid used is low and therefore the SUMO2ylation is not completely inhibited. It would be advisable if the authors go to concentration of complete inhibition.

      We acknowledge that some residual SUMO2ylation is visible in the immunoblots following anacardic acid treatment. However, the assay demonstrates a clear, dose-dependent reduction in SUMO2ylation levels, which is sufficient to interpret the effect of SUMO2 inhibition on p66Shc function. Increasing the concentration further could introduce off-target effects, and the current assay provides a reliable and physiologically relevant readout.

      (c) The figure panels need uniform fonts and also the size/resolution of the western blots needs to be improved

      We thank the reviewer for this suggestion. All figures have been updated to ensure uniform fonts, and the western blot images have been improved for size and resolution in the revised manuscript.

      (d) Supplemental figures are very low resolution.

      High-resolution images have been provided.

      (e) Please introduce abbreviations before using them (like LDLr and ND).

      The abbreviations have been explained.

      (f) Figure 6A seems like data that belongs in the SI rather than the main text.

      We kept it in main figure as the knock-in mouse was generated for this study.

      (a) Figure 7B needs a WT HFD control.

      It is a valid suggestion. However, it is known in the literature (we have prior experience as well) that wild-type mice do not exhibit a drastic increase in serum lipid level with high-fat diet feeding.

      (b) The differences in Figure 7E are not profound and immediately obvious. Is there a better way to show the data, potentially with quantification?

      We agree, the difference is not that huge, we have provided the qualification as Fig. 7F.

      (h) Figure panels 6E-H could use some labels to differentiate the data from WT vs p66ShcK81R mutant mice within the figure panel.

      We mentioned that on the top and have inserted a vertical line to separate the groups.

      (i) Lines 165-168, 182-185, 214-218 and 282-286: These sentences can be rewritten for more clarity.

      We have modified these sentence for clarity.

      (j) Line 512 should read, "Two-way ANOVA, ***P<0.001."

      Thank you, we have corrected the mistake.

    1. eLife Assessment

      This valuable study introduces an innovative experimental design to address a crucial and timely issue in microbial ecology: the potential bias in soil microbial community analyses caused by extracellular DNA degradation. The work contributes to the field by providing convincing evidence that variable extracellular DNA degradation rates impact microbial ecology studies. This research will appeal to microbial ecologists and researchers interested in using molecular techniques to evaluate microbial community structure.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates the degradation dynamics of extracellular DNA in soils and its impact on estimates of microbial abundance and diversity. By combining a broad geographic sampling design with a primer-labeling strategy, qPCR quantification, amplicon sequencing, and PMA treatment, the authors aim to disentangle total versus intracellular DNA signals and explore sequence-specific degradation patterns. The topic is relevant, particularly given the increasing awareness of relic DNA as a confounding factor in microbial ecology. The experimental design is ambitious and potentially impactful. However, several conceptual inconsistencies, methodological ambiguities, and statistical limitations currently weaken the robustness of the conclusions. These issues need to be addressed.

      Strengths:

      The manuscript addresses a timely and important question in microbial ecology, particularly given the growing recognition that relic DNA can bias interpretations of community composition derived from amplicon sequencing. The study is ambitious in scope, incorporating a broad geographic sampling design across multiple soil types, which enhances the generalizability of the findings. The use of a controlled microcosm experiment combined with a primer-labeling strategy to track extracellular DNA dynamics is conceptually innovative and provides a structured framework to investigate degradation processes.

      In addition, the integration of multiple approaches, including qPCR for absolute quantification, high-throughput sequencing for community profiling, and PMA treatment to differentiate extracellular from intracellular DNA, represents a comprehensive attempt to disentangle complex sources of bias in soil microbiome analyses. The effort to link degradation dynamics with environmental variables and to explore sequence-level patterns further demonstrates the authors' intent to move beyond descriptive analyses toward a mechanistic understanding.

      Weaknesses:

      Several conceptual and methodological issues currently limit confidence in the study's conclusions. Key terms such as "sequence-specific degradation" are not clearly defined or supported by a mechanistic or structural hypothesis, making it difficult to interpret the biological meaning of the results. In addition, the bioinformatic workflow presents inconsistencies, particularly the use of ASVs followed by clustering at 97% similarity, which undermines the resolution required to support sequence-level inferences. Statistical analyses are also insufficiently described, including unclear definitions of "T values," a lack of detail on pairing structure, and no indication of multiple testing correction.

      Furthermore, important methodological details are missing or unclear, including primer design (e.g., GAPDH tag vs ACTF), Illumina library preparation (e.g., adapter and indexing strategy), and validation of PMA treatment efficiency. The interpretation of PMA-treated samples as representing "living communities" is likely overstated, given the known limitations of the method in soil systems. Finally, typographical errors, inconsistent terminology, and unclear phrasing throughout the manuscript reduce readability and further complicate interpretation.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript describes the results of an interesting study examining the rate of degradation of extracellular DNA in soil ecosystems using a clever experimental approach. 16S ribosomal RNA genes were amplified from soil samples, and then purified PCR amplicons, containing a 5' linker sequence on the forward primer, were introduced to soils and monitored over time using real-time quantitative PCR and NGS amplicon sequencing. The study was able to measure rates of overall extracellular DNA degradation, but also sequence-specific degradation rates. I like the idea and execution of the study, and the results are interesting. The manuscript needs some help to improve the overall readability. Please see general and editorial comments below.

      Strengths:

      Innovative experimental design that is well deployed across a large number of soil types, revealing interesting variability in extracellular DNA degradation.

      Weaknesses:

      (1) The manuscript needs another review to improve the readability of the document.

      (2) The authors have used 16S genes to look at sequence-specific degradation. But 16S rRNA genes are actually pretty well conserved, and there isn't as much genetic variation across this gene among organisms as there is for other genes. It might be more relevant to look at metagenomic DNA degradation from high AT, high GC organisms, etc. This would be more generalizable than 16S genes.

      (3) Consideration of differential cell lysis during soil DNA extraction needs to be considered as well.

      (4) It is not clear why the authors didn't put GAPDH linkers on the reverse primer as well. This would have given an easier amplicon to amplify (no degeneracies at all).

    4. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely appreciate you and the reviewers for investing time and effort in evaluating our manuscript. After carefully reading the comments and suggestions, we found they are insightful, constructive, and critical for improving the quality of our work. Based on these valuable recommendations, we have substantially revised the manuscript as summarized below.

      Abstract: Inappropriate or ambiguous statements have been revised to improve clarity.

      Introduction: (1) The study purpose and hypotheses have been re-organized in a clearer and more concise way. (2) A mechanistic rationale for sequence-specific degradation has been provided and the use of PMA treatment has been explained. (3) The terminologies related to extracellular DNA and 16S rRNA gene amplicons have been clarified.

      Materials and Methods: 1) More detailed description of the microcosm experiment has been added. 2) The design and rationale of GAPDH F-tagged primers and the use of fusion primers for Illumina library preparation have been clarified. 3) More details about PCR amplification, DNA purification, and pooling strategies have been added. 4) We have corrected and standardized primer naming throughout the manuscript; 5) More details about bioinformatic workflow have been added. 6) We have defined statistical parameters and multiple testing corrections. 7) All abbreviations have been defined and standardized.

      Results: 1) The terminology for PMA-treated DNA has been revised and it has been clarified interpretation as “PMA-treated prokaryotic community” rather than “living community”. 2) The figures and legends have been updated for clarity, and the explicit explanation of “ASV I” and “ASV II” in pairwise comparisons have been added. 3) the figures (e.g., Figs. 2–5, S2–S8) have been reorganized to better reflect results; 4) Inappropriate statements or misleading interpretations have been removed.

      Discussion: A detailed section on technical limitations have been added. The limiatons added mainly include: 1) PCR amplification bias and recommendations for spike-in standards or multi-primer approaches; 2) differential DNA extraction efficiency due to variable cell lysis; and 3) limitations of using 16S rRNA amplicons as proxies for natural extracellular DNA and the limitations of PMA treatment efficiency in soil matrices;

      eLife Assessment

      This valuable study introduces an innovative experimental design to address a crucial and timely issue in microbial ecology: the potential bias in soil microbial community analyses caused by extracellular DNA degradation. While the evidence showing variable degradation rates of extracellular DNA is convincing, additional conceptual, methodological, and statistical clarifications could reinforce the claims and the study's contribution to the field. This research will appeal to microbial ecologists and researchers interested in using molecular techniques to evaluate microbial community structure.

      We sincerely appreciate the editors for the careful assessment of our work and for recognizing the value of addressing extracellular DNA degradation in soil microbial community analyses. We also greatly appreciate the reviewers’ constructive feedbacks concerning the need for additional conceptual, methodological, and statistical clarifications. We agree that further refinement in these areas will strengthen our claims and enhance the study’s contribution to the field. Based on these insightful suggestions, we have carefully revised the manuscript to provide clearer conceptual framework, more detailed methodological descriptions, and more rigorous statistical analyses. We believe these revisions have substantially improved the clarity and robustness of our work. More details about the revisions have been provided in the following responses.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates the degradation dynamics of extracellular DNA in soils and its impact on estimates of microbial abundance and diversity. By combining a broad geographic sampling design with a primer-labeling strategy, qPCR quantification, amplicon sequencing, and PMA treatment, the authors aim to disentangle total versus intracellular DNA signals and explore sequence-specific degradation patterns. The topic is relevant, particularly given the increasing awareness of relic DNA as a confounding factor in microbial ecology. The experimental design is ambitious and potentially impactful. However, several conceptual inconsistencies, methodological ambiguities, and statistical limitations currently weaken the robustness of the conclusions. These issues need to be addressed.

      We sincerely appreciate the reviewer for the constructive assessment of our work. We also appreciate the reviewer’s critical insights regarding the conceptual inconsistencies, methodological ambiguities, and statistical limitations that currently weaken the robustness of the conclusions. We agree with the reviewer that addressing these issues is essential to strengthen our work. Based on these valuable comments, we have carefully revised the manuscript to clarify the conceptual framework. Additionally, we have provided more detailed methodological descriptions, and enhance the statistical rigor of our analyses. We believe these revisions have substantially improved the clarity, consistency, and overall robustness of our conclusions.

      Strengths:

      The manuscript addresses a timely and important question in microbial ecology, particularly given the growing recognition that relic DNA can bias interpretations of community composition derived from amplicon sequencing. The study is ambitious in scope, incorporating a broad geographic sampling design across multiple soil types, which enhances the generalizability of the findings. The use of a controlled microcosm experiment combined with a primer-labeling strategy to track extracellular DNA dynamics is conceptually innovative and provides a structured framework to investigate degradation processes.

      In addition, the integration of multiple approaches, including qPCR for absolute quantification, high-throughput sequencing for community profiling, and PMA treatment to differentiate extracellular from intracellular DNA, represents a comprehensive attempt to disentangle complex sources of bias in soil microbiome analyses. The effort to link degradation dynamics with environmental variables and to explore sequence-level patterns further demonstrates the authors' intent to move beyond descriptive analyses toward a mechanistic understanding.

      We sincerely thank the reviewer for the positive and encouraging comments of our work.

      Weaknesses:

      Several conceptual and methodological issues currently limit confidence in the study's conclusions. Key terms such as "sequence-specific degradation" are not clearly defined or supported by a mechanistic or structural hypothesis, making it difficult to interpret the biological meaning of the results. In addition, the bioinformatic workflow presents inconsistencies, particularly the use of ASVs followed by clustering at 97% similarity, which undermines the resolution required to support sequence-level inferences. Statistical analyses are also insufficiently described, including unclear definitions of "T values," a lack of detail on pairing structure, and no indication of multiple testing correction.

      Furthermore, important methodological details are missing or unclear, including primer design (e.g., GAPDH tag vs ACTF), Illumina library preparation (e.g., adapter and indexing strategy), and validation of PMA treatment efficiency. The interpretation of PMA-treated samples as representing "living communities" is likely overstated, given the known limitations of the method in soil systems. Finally, typographical errors, inconsistent terminology, and unclear phrasing throughout the manuscript reduce readability and further complicate interpretation.

      We sincerely appreciate the reviewer’s thorough and critical evaluation of the manuscript’s weaknesses. The issues raised regarding conceptual clarity, bioinformatic consistency, statistical rigor, methodological transparency, and the interpretation of PMA treatment have been fully acknowledged. We also recognized that typographical errors, inconsistent terminologies, and unclear phrasing largely reduced readability. In response to these valuable comments, the manuscript has been carefully revised as follows. (1) The clearer definition of the term “sequence-specific degradation” has been provided. (2) The bioinformatic workflow was streamlined to ensure consistency. (3) The descriptions of statistical analyses were substantially expanded, including explicit definitions of “t values,” detailed clarification of the pairing structure, and the application of appropriate multiple testing corrections. (4) Missing details regarding primer design, Illumina library preparation, and PMA treatment validation have been added to the Method section. (5) Interpretations of PMA-treated samples have been revised to more accurately reflect methodological limitations in soil systems. (6) The manuscript has been thoroughly proofread to correct typographical errors, standardize terminology, and enhance overall clarity. These revisions are believed to substantially address the concerns raised and significantly strengthen the manuscript. A point-by-point response to the specific comments is provided below.

      Reviewer #2 (Public review):

      Summary:

      This manuscript describes the results of an interesting study examining the rate of degradation of extracellular DNA in soil ecosystems using a clever experimental approach. 16S ribosomal RNA genes were amplified from soil samples, and then purified PCR amplicons, containing a 5' linker sequence on the forward primer, were introduced to soils and monitored over time using real-time quantitative PCR and NGS amplicon sequencing. The study was able to measure rates of overall extracellular DNA degradation, but also sequence-specific degradation rates. I like the idea and execution of the study, and the results are interesting. The manuscript needs some help to improve the overall readability. Please see general and editorial comments below.

      We sincerely thank the reviewer for the positive and encouraging assessment of our study. We have carefully revised the manuscript to enhance clarity, streamline the presentation, and refine the language throughout. We believe these improvements have made the manuscript more readable and easier to follow. We are also grateful for the general and editorial comments provided, which have been addressed as outlined below.

      Strengths:

      Innovative experimental design that is well deployed across a large number of soil types, revealing interesting variability in extracellular DNA degradation.

      We sincerely thank the reviewer for the positive and encouraging assessment of our work.

      Weaknesses:

      (1) The manuscript needs another review to improve the readability of the document.

      We thank the reviewer for this helpful suggestion. We fully agree that improving readability is essential for effectively communicating our findings. Based on the comment, we have carefully revised the manuscript to enhance clarity and readability. We have streamlined sentence structures, standardized terminology, corrected typographical errors, and improved the logical organization of the text. We believe these revisions have substantially improved the overall readability of the manuscript.

      (2) The authors have used 16S genes to look at sequence-specific degradation. But 16S rRNA genes are actually pretty well conserved, and there isn't as much genetic variation across this gene among organisms as there is for other genes. It might be more relevant to look at metagenomic DNA degradation from high AT, high GC organisms, etc. This would be more generalizable than 16S genes.

      We thank the reviewer for this insightful comment. We agree with the reviewer that 16S rRNA genes are relatively conserved compared to functional genes or whole metagenomic DNA, and that studying degradation of more variable sequences (e.g., high‑AT, high‑GC regions, or metagenomic DNA) would provide greater generalizability. However, we would like to clarify the rationale for using 16S rRNA gene amplicons in the present study. First, the 16S rRNA gene remains the most widely used phylogenetic marker in soil microbial ecology (Knight et al., 2018). Demonstrating sequence‑specific degradation with this well‑established marker directly informs a large body of existing research that relies on 16S RNA gene‑based community analyses. Second, despite its conserved nature, the targeted fragment in this study is belong to the highly varied region (V4) of 16S rRNA gene. Accordingly, we indeed observed significant sequence‑specific variation in degradation rates among different 16S rRNA gene amplicon sequence variants (ASVs) (Fig. 2c, 3a). This indicates that even within a conserved marker gene, sequence‑dependent degradation biases exist and can affect diversity estimates. Third, our study was designed as a proof‑of‑concept to establish a methodological framework for quantifying both overall and sequence‑specific degradation rates. Using a single, well‑characterized marker allowed us to develop and validate the primer‑labeling and qPCR/sequencing workflow without the additional complexity of metagenomic DNA (e.g., variable fragment lengths and complex mineral associations). In the revised manuscript, we have added the following sentence to the Discussion section to address the concerns from the reviewer.

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015)”

      (3) Consideration of differential cell lysis during soil DNA extraction needs to be considered as well.

      We thank the reviewer for raising this important technical consideration. We agree that differential cell lysis during soil DNA extraction is a well‑recognized source of bias in microbial community analysis. Different microbial taxa (e.g., Gram‑positive vs. Gram‑negative bacteria, spores, or fungi) vary in their cell wall structure and susceptibility to lysis, which can lead to under‑representation of certain groups and over‑representation of others in the extracted DNA. This bias affects both total DNA extracts and PMA‑treated fractions, potentially influencing our estimates of the relative contributions of intact‑cell derived DNA versus extracellular DNA. However, currently, eliminating these biases are still challenging, and thus we have added the following sentence to the Discussion to address this concern.

      L305-311

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015).”

      (4) It is not clear why the authors didn't put GAPDH linkers on the reverse primer as well. This would have given an easier amplicon to amplify (no degeneracies at all).

      The decision to place the GAPDH linker only on the forward primer (515F) was intentional to balance the need for tracking exogenous extracellular DNA with amplification efficiency, sequencing quality, and cost-effectiveness. Adding linkers to both primers would increase the total amplicon length, potentially reducing amplification efficiency, especially in complex soil samples with degraded or low-quality DNA. More importantly, the reverse primer used in our study is a degenerate primer designed to target the 16S rRNA gene across diverse bacterial taxa, and extending it with an additional GAPDH linker could introduce further complexity, decrease amplification efficiency, and increase primer-dimer formation. Additionally, single-end labeling allows the usage of standard 16S rRNA reverse primers with existing barcodes, whereas dual-end labeling would require synthesis of new barcode-labeled primers, increasing both cost and time. Our preliminary experiments confirmed that single-end labeling produced reproducible amplification curves (~85% efficiency) and high-quality sequencing reads, which were sufficient for quantifying degradation rates. We have added a clarification in the Methods section to explain this rationale.

      L365-371

      “The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Inconsistency between ASV inference and 97% sequence recruitment

      The bioinformatic pipeline presents a major conceptual inconsistency. ASVs are inferred using UNOISE3, but reads are subsequently mapped to ASVs at a 97% similarity threshold, effectively reintroducing OTU-level clustering. Given that the manuscript's central claim is sequence-specific degradation, this step undermines the single-nucleotide resolution that ASVs provide and may obscure biologically meaningful differences. The authors should either reanalyze the data using a consistent ASV framework (exact matching), or explicitly treat the analysis as OTU-like and moderate claims of sequence specificity.

      We appreciate the reviewer's critical evaluation of our bioinformatic pipeline. We also apologized for our unclear statements in our original manuscript. We understand the concern that mapping reads to ASVs at 97% similarity might appear to reintroduce OTU-level clustering. Actually, we used the default pipeline provided by the authors of USEARCH with otutab command to generate the ASV table.

      Following the logic and recommendations of the USEARCH/UNOISE3 developer (Robert Edgar), this approach is a standard procedure for robust noise management rather than a conceptual inconsistency. First, in our pipeline, ASVs (ZOTUs) are first inferred using the UNOISE3 algorithm, which effectively identifies “true” biological sequences at single-nucleotide resolution. Secondly, according to the USEARCH manual, while the ASVs themselves represent exact biological sequences, the raw reads inevitably contain stochastic sequencing errors. Using “exact matching” for recruitment would discard a significant portion of the data that originates from a specific ASV but carries minor random errors. Mapping at 97% identity is the recommended method to recruit these noisy reads back to their correct biological origin (the ASV centroid). Meanwhile, during the recruitment process, reads are not randomly assigned to any ASV within the 97% identity radius. Instead, the algorithm follows a “highest similarity first” principle. For instance, if a specific read exhibits 98% similarity to ASV1 and 97% similarity to ASV2, it is strictly assigned to ASV1. A read is only discarded if its highest similarity to any ASV falls below the 97% threshold. Unlike traditional OTU clustering (where sequences are clustered together based on similarity from the start), our approach maintains the ASV as a fixed biological reference. The quantification of degradation rates is performed on these high-resolution centroids. Thus, our claims regarding sequence-specific degradation remain valid, as the underlying biological variation is defined by the ASVs. To avoid any possible confusion, we have rewritten the relevant paragraph in the Methods section (subsection 4.6) as follows.

      L446-457

      “ASVs were generated using the UNOISE3 non‑clustering denoising algorithm (Edgar, 2016), which infers 100% exact sequence variants by distinguishing biological sequences from PCR/sequencing errors. ASVs with total sequence counts fewer than 9 across all samples were removed to reduce noise. To quantify the abundance of each ASV, an ASV table was generated by mapping the quality‑filtered raw reads back to the ASV set using the otutab command. A 97% similarity threshold was applied for this recruitment to accommodate stochastic sequencing noise while maintaining biological resolution. Crucially, the mapping followed a best-hit priority rule, where each read was assigned to the ASV with the highest per cent identity within the 97% radius. This approach ensures that reads derived from the same biological template are accurately counted toward their respective ASV, preventing the underestimation of abundances that would occur with exact matching while strictly preserving the single-nucleotide resolution of the ASV framework.”

      (2) Undefined "ASV I" and "ASV II" groups

      The manuscript refers to "ASV I" and "ASV II" groups in pairwise comparisons of degradation rates (e.g., Fig. 3), but these groups are not defined anywhere in the text.

      It is unclear whether these represent: predefined biological categories, arbitrary pairwise ASV comparisons, or groupings based on taxonomy, abundance, or degradation rate.

      In addition, the pairing structure underlying these comparisons is not described. While a paired t-test is mentioned, it is unclear how ASVs were paired (e.g., within sites, across samples, or across time points).

      The current terminology ("groups") is potentially misleading and suggests biological structure where none may exist. The authors should explicitly define these terms, clarify the pairing scheme, and revise terminology if these are simply pairwise comparisons.

      We thank the reviewer for this keen observation. We completely agree that the terms "ASV I" and "ASV II" were poorly defined and potentially misleading.

      We would like to clarify that "ASV I" and "ASV II" were not intended to represent predefined biological categories (such as groupings based on taxonomy, abundance, or degradation rates). Instead, they were merely used as a labeling convention to indicate the directionality of pairwise comparisons within the heatmap matrix. Specifically, "ASV I" referred to the ASVs represented in the rows, while "ASV II" referred to those in the columns. In the original Fig. 3, blue indicated that the degradation rate of the row ASV was significantly lower than that of the column ASV, and red indicated the opposite. To avoid any confusion, we have removed the “ASV I/II” terminology throughout the manuscript and figures, replacing them with “Row ASVs” and “Column ASVs”. To address this issue, we have revised the Figure 3 legend to include a more explicit explanation:

      L820-824

      “In the heatmap, each cell represents a pairwise comparison between two ASVs. Blue indicates that the degradation rate of the ASVs listed in the row (row ASVs) is significantly lower than that of the ASVs listed in the column (column ASVs); red indicates that the row ASVs has a significantly higher degradation rate than the column ASV. A positive t value indicates that the row ASVs degrades significantly faster than the column ASVs; a negative t value indicates the opposite.”

      We thank the reviewer for raising the important issue regarding the definition of “t values” in our statistical analysis. We apologize for the lack of clarity in the original manuscript. To clarify, the T values presented in Figure 3a represent the test statistics (t-values) from paired t-tests comparing the degradation rate constants of two ASVs across the 30 study sites. The T-value was obtained from a paired t-test between two ASVs across the same samples. The t-value indicates the magnitude and direction of the difference between the two ASVs’ degradation rates relative to the variability across sites. A positive t-value (colored red in the heatmap) indicates the row ASVs degrades significantly faster than the column ASVs; a negative t-value (colored blue) indicates the opposite.

      L544-548

      “As for the analysis, we performed paired t‑tests across all the study sites. Thus, the degradation rates were essentially compared within each site, with both values originating from a same soil sample under identical incubation conditions. A positive t value indicates that the first ASV has a significantly higher degradation rate than the second one, and a negative t value indicates the opposite. The p values were adjusted for multiple comparisons using the FDR method.”

      (3) Lack of definition and justification of "T values"

      The manuscript reports "T values" for comparisons between ASVs but does not clearly define how these values are calculated. Although a paired t-test is mentioned, it remains unclear how the pairing was constructed, whether assumptions (normality, independence) were evaluated, and whether corrections for multiple comparisons were applied. Given the large number of ASVs, failure to control for multiple testing could inflate false positives. More broadly, the use of a simple paired t-test may not be appropriate given the hierarchical and compositional structure of the data.

      We sincerely thank the reviewer for pointing out the need to clarify the definition and justification of the t-values presented in our manuscript. Each t-value represents the test statistic from a paired t-test comparing the degradation rates of two ASVs across the same set of samples. The paired t-test assumes that the differences between paired observations are approximately normally distributed and that the pairs are independent across columns. We have evaluated the normality of differences using standard diagnostic plots and verified that the assumption is reasonably satisfied given the sample size. We performed a correction for multiple comparisons using the False Discovery Rate (FDR) procedure to control for potential false positives. We have revised the Methods section to clearly define t-values.

      (4) Conceptual validity of "sequence-specific degradation"

      The manuscript repeatedly refers to "sequence-specific degradation" of extracellular DNA; however, this concept is not clearly defined nor supported by a biological or structural hypothesis. It is unclear what "sequence-specific" refers to (e.g., nucleotide composition, GC content, secondary structure, taxonomic identity), whether differences are expected in conserved versus variable regions of the 16S rRNA gene, or what mechanistic basis would explain differential degradation among sequences. Given that the analysis is based on short 16S V4 amplicons, and no structural or biochemical framework is provided, it is difficult to interpret whether the observed differences truly reflect intrinsic sequence-dependent degradation or are instead driven by methodological or statistical artifacts (e.g., abundance effects, amplification bias).

      I believe the authors should explicitly define what is meant by "sequence-specific degradation," provide a biologically grounded hypothesis (e.g., structural accessibility, GC content, stem-loop stability), and align their interpretation with the resolution and limitations of the data.

      We thank the reviewer for this critical conceptual comment. We apologize that “sequence‑specific degradation” was not clearly defined and lacked a biological or structural hypothesis. To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation. This adjustment ensures that the hypotheses are directly grounded in the theoretical framework (e.g., GC content, thermodynamic stability, and secondary structures) presented in the paragraph.

      We now define “sequence‑specific degradation” as statistically significant differences in first‑order degradation rate constants among distinct ASVs, mainly arising from intrinsic DNA properties (base composition, secondary structure, and restriction sites) or differential mineral adsorption.

      L99-104

      “Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      We also expanded the mechanistic discussion to include GC content and secondary structure.

      L230-235

      “We also examined whether GC content could explain the observed sequence‑specific patterns, but no significant correlation was found (Fig. S4), suggesting that simple base composition is not the primary driver in this study. However, this does not exclude the possibility that higher‑order structural features (e.g., hairpin loops) or sequence‑specific nuclease recognition motifs contribute to differential degradation (Wang et al., 2007). This should be tested in future studies using synthetic DNA constructs with controlled structural elements.”

      We acknowledge that inferring sequence‑specific degradation from combined relative abundance and qPCR data is subject to potential methodological artifacts, including compositional effects, PCR amplification bias, and abundance‑dependent detection limits. However, we have taken several stringent steps to minimize these concerns. Specifically, we restricted our analysis to ASVs that were present in more than 90% of the study sites and for which the degradation curve fits yielded R<sup>2</sup> > 0.5, ensuring that only robustly detected and reliably modeled sequences were retained. Because our analysis tracks the ratio of each ASV across a time series, any sequence-specific PCR amplification bias remains constant for that particular sequence. By focusing on the rate of change rather than absolute read counts, such systematic biases are mathematically canceled out during the calculation of degradation kinetics.

      (5) Conceptual ambiguity in "GAPDH F-labeled 16S rRNA genes"

      The manuscript repeatedly refers to "GAPDH F-labeled 16S rRNA genes," which is confusing and may be misinterpreted as targeting GAPDH rather than 16S. It should be clearly stated that GAPDH refers to glyceraldehyde-3-phosphate dehydrogenase, and a GAPDH-derived sequence is used as a synthetic tag appended to a 16S primer. Additionally, the divergence of this tag from microbial sequences should be justified to ensure specificity. There is also an inconsistency in primer naming (e.g., "GAPDH F" vs "ACTF" in the figures), which should be corrected.

      We sincerely thank the reviewer for this important comment. We agree that the phrase “GAPDH F‑labeled 16S rRNA genes” could be confusing, as it may be misinterpreted as targeting the GAPDH gene rather than the 16S rRNA gene. We have revised the manuscript to avoid this ambiguity and to provide clear justification for the use of the GAPDH tag. GAPDH (glyceraldehyde‑3‑phosphate dehydrogenase) is a human housekeeping gene. Its forward primer sequence (GAPDH F: 5′‑CAT TGG CAA TGA GCG GTT C‑3′) was used as a synthetic tag appended to the 16S primer because (i) no homologous sequences exist in soil DNA (confirmed by PCR), and (ii) its melting temperature is compatible with the reverse primer. This tag allows specific tracking of exogenous DNA without interference from native soil sequences.

      Throughout the manuscript, ambiguous phrases such as “GAPDH F‑labeled 16S rRNA genes” have been replaced with more precise terms, “GAPDH F‑tagged 16S rRNA gene amplicon fragments” clarifying that the tag is an appendage and not the amplification target.

      We have checked the entire manuscript and confirm that “ACTF” does not appear anywhere. To avoid confusion, the primer is now consistently referred to as “GAPDH F” in all figures, legends, and text.

      L360-371

      “GAPDH is a primer for a human housekeeping gene and it has no homologous sequences in soils. Subsequently, GAPDH was selected as the label primer based on two criteria. First, this primer was selected to avoid interference from the original soil sequences (Huang et al., 2014; Yang et al., 2021; Arvizu-Hernandez et al., 2025), and no detectable PCR amplification was observed for the primer set GAPDH F-806R across all the soil DNA samples included in this study. Second, the melting temperature (Tm) value of GAPDH F approximately matched that of 806R. The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      (6) Limitations of using PCR amplicons as proxies for extracellular DNA

      The study uses PCR-generated amplicons to simulate extracellular DNA. While useful for controlled comparisons, these fragments may not reflect the physicochemical diversity of natural extracellular DNA (e.g., adsorption to minerals, fragment size variability, protection within aggregates). This limitation should be explicitly acknowledged, and conclusions should be framed accordingly.

      We appreciate the reviewer’s constructive feedback. We fully acknowledge that using PCR-generated amplicons to simulate extracellular DNA (eDNA) has inherent limitations in capturing the full physicochemical diversity of naturally occurring eDNA in soils. Specifically, we agree that PCR fragments may not replicate features such as highly variable fragment size distributions, associations with complex cellular components (e.g., vesicles or protein complexes), or long-term physical sequestration within soil micro-aggregates. Despite of these limitations, the use of uniform primer-tagged PCR amplicons was a deliberate choice to enable precise tracking of exogenous DNA degradation kinetics while eliminating background interference from endogenous soil eDNA. This design is a prerequisite for the high-resolution kinetic modeling of sequence-specific decay. Furthermore, in our bioinformatic pipeline, the 97% mapping threshold was specifically applied to minimize the influence of stochastic sequencing and PCR errors on abundance quantification. In the revised manuscript, these potential limitations have been addressed.

      L295-299

      “First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020).”

      (7) Interpretation of sequence-specific degradation

      Sequence-specific degradation rates are inferred from combining relative abundance data with total qPCR estimates. This approach is sensitive to compositional effects, amplification biases, and abundance-dependent detection limits. It remains unclear whether observed differences reflect true sequence-specific degradation or methodological artifacts. This limitation should be discussed more explicitly.

      We thank the reviewer for highlighting this critical methodological point. In our study, sequence-specific degradation rates were estimated by combining ASV-relative abundances with total qPCR-derived 16S rRNA gene copy numbers. We acknowledge that this approach may be influenced by compositional effects, PCR amplification biases, and abundance-dependent detection limits. However, the degradation rate constant (k) in our study, represents the rate of change for a specific sequence over time. Since PCR amplification biases are generally sequence-specific and consistent across samples processed under identical conditions, these systematic errors are mathematically canceled out when calculating the relative change (slope) for the same ASV across a time series. Second, all qPCR measurements were performed with three technical triplicates with standard curves to ensure quantitative reliability. Third, relative abundances were converted to absolute abundances using total qPCR estimates, allowing cross-taxa comparisons that reduce compositional bias. This approach is widely recognized in microbial ecology as a robust method. To address this concern, we have added some explanations in the revised manuscript.

      L84-86

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequence variants (ASVs) under identical soil and incubation conditions.”

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      L311-314

      “While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017)..”

      (8) Overinterpretation of PMA-treated samples as "living communities"

      The manuscript interprets PMA-treated DNA as representing intracellular or "living" microbial communities. While PMA is useful, this interpretation should be treated with caution in soils. PMA efficiency can be affected by soil matrix complexity, DNA adsorption to particles, incomplete light penetration, and permeability of compromised cells. Importantly, no validation of PMA efficiency is presented.

      We thank the reviewer for this important caution. We agree that interpreting PMA‑treated DNA as representing “living” or “intracellular” communities is an overstatement in soil systems. In the revised manuscript, we no longer describe PMA-treated DNA as a direct proxy for the “living community,” but instead refer to it as the “PMA-treated prokaryotic community”.

      Although we did not directly validate PMA efficiency in this study, we used a standardized PMA protocol that has been widely applied in microbial ecology, and our goal was to obtain a comparative estimate of the influence of extracellular DNA on community analysis across soils under a consistent methodological framework. Based on previous studies (Carini et al., 2016; Du et al., 2025), which found that in similar soil types, PMA treatment can significantly reduce the interference of extracellular DNA and alter the community structure, this indirectly proves the effectiveness of this technique.

      Nevertheless, we agree that future studies should include explicit validation controls, such as live/dead cell mixtures, heat-killed controls, or soil-specific PMA efficiency tests, to better quantify method performance across diverse soil matrices. We have added a dedicated paragraph in the "Methodological Considerations and Limitations" section to discuss how soil-specific properties (e.g., turbidity, adsorption capacity) might lead to incomplete exclusion of extracellular DNA, thereby advising a more cautious interpretation of the "viable" community data.

      L496-500

      “To inhibit amplification of eDNA, soils were incubated with propidium monoazide (PMA), as described previously (Carini et al., 2016). Upon photoactivation, eDNA can form covalent bonds through cross-linking, leading to the inhibition of its PCR amplification. In contrast, microbes with intact cell membranes exclude PMA, and their DNA is not cross-linked with PMA, and remains amenable to PCR amplification.”

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      Minor Comments

      (1) Line 79: Provide examples of how extracellular DNA contributes to nutrient cycling (e.g., P, N sources) and signal transduction (e.g., horizontal gene transfer).

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have added specific examples to clarify how extracellular DNA contributes to nutrient cycling and signal transduction. Specifically, we now note that extracellular DNA can serve as a source of phosphorus and nitrogen following enzymatic degradation, thereby contributing to soil nutrient turnover. We also clarify that extracellular DNA plays an important role in horizontal gene transfer, acting as a genetic reservoir that can be taken up by competent microorganisms and thereby facilitating the spread of functional traits such as antibiotic resistance. These examples have been added to improve the clarity and biological context of this statement.

      L48-52

      “EDNA serves as a critical vector for horizontal gene transfer (HGT), facilitating the uptake of genetic material by competent microorganisms and promoting the spread of functional traits such as antibiotic resistance (Liu et al., 2024). In addition, eDNA participates in soil biogeochemical cycling because its enzymatic degradation releases bioavailable nutrients, particularly phosphorus and nitrogen, which can be reused by soil microorganisms (Ye et al., 2022).”

      (2) Line 79: Replace "for an extended period of time" with a more precise or referenced timescale.

      We agree with the reviewer. We have replaced the vague phrase with a precise timescale. Extracellular DNA can persist in soils for months to years.

      (3) Line 94: Clarify what is meant by "high-level structure" (e.g., secondary structure, environmental association).

      We thank the reviewer for pointing out this ambiguity. In the original manuscript, the phrase “high-level structure” was not sufficiently precise. In the revised version, we have clarified that this refers primarily to higher-order structural properties of DNA molecules, such as secondary structure, local conformational features, and sequence-dependent interactions with minerals or organic matter in soil. These characteristics may influence the accessibility of extracellular DNA to nucleases and thus affect degradation rates. We have revised the text accordingly to improve clarity and precision.

      L84-104

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequences (ASVs) under identical soil and incubation conditions. The potential variations in sequence-specific eDNA degradation rates can be attributed to several factors. First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C: N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021). Third, the persistence of soil DNA is often associated with its adsorption and protection by minerals and humus in soils (Cai et al., 2006b; Vuillemin et al., 2017; McKinney and Dungan, 2020). Thus, sequence-dependent differences in the physicochemical behavior of DNA molecules, including their affinity for soil minerals and organic matter, may also contribute to variation in degradation rates among sequences (Levy-Booth et al., 2007; Morrissey et al., 2015). Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (4) Line 111: The hypothesis is not clearly linked to the rationale. If sequence-specific degradation is expected, clarify whether it relates to conserved vs variable regions or structural features (e.g., stems vs loops).

      We thank the reviewer for this helpful comment. We agree that the original manuscript did not clearly link the hypothesis regarding sequence-specific degradation to its mechanistic rationale. In the revised manuscript, we have clarified that the expectation of sequence-specific degradation is not simply based on conserved vs variable regions of the 16S rRNA gene, but rather on the potential for sequence differences to influence intrinsic physicochemical properties, including base composition, local conformational features, potential secondary structures, and motif-dependent nuclease susceptibility. These factors may alter DNA accessibility to extracellular nucleases, providing a mechanistic basis for sequence-specific degradation. This clarification is now reflected in the Introduction and linked to the formal hypothesis statement.

      To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses (H1–H3) immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation.

      L84-104

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequences (ASVs) under identical soil and incubation conditions. The potential variations in sequence-specific eDNA degradation rates can be attributed to several factors. First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C:N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021). Third, the persistence of soil DNA is often associated with its adsorption and protection by minerals and humus in soils (Cai et al., 2006b; Vuillemin et al., 2017; McKinney and Dungan, 2020). Thus, sequence-dependent differences in the physicochemical behavior of DNA molecules, including their affinity for soil minerals and organic matter, may also contribute to variation in degradation rates among sequences (Levy-Booth et al., 2007; Morrissey et al., 2015). Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (5) Lines 310-311: Clearly indicate which portion of the primers corresponds to the modified (GAPDH-derived) sequence. Provide full annotated primer sequences.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we now clearly indicate which portion of the forward primer corresponds to the GAPDH-derived synthetic tag and which portion corresponds to the 16S rRNA gene primer sequence. We have also provided the full annotated primer sequences in the Methods section to avoid ambiguity.

      Specifically, the modified forward primer is now described as:

      GAPDH-F-515F: 5′-CAT TGG CAA TGA GCG GTT C-GTG CCA GCM GCC GCG GTA A-3′,

      where CAT TGG CAA TGA GCG GTT C is the GAPDH-derived synthetic tag and GTG CCA GCM GCC GCG GTA A is the 16S rRNA gene forward primer sequence (515F).

      The reverse primer is:

      806R: 5′-GGA CTA CHV GGG TWT CTA AT-3′.

      L354-359

      “Briefly, exogenous eDNA was prepared by PCR amplification using a modified forward primer consisting of a GAPDH F tag fused to the 16S rRNA gene primer 515F, together with the reverse primer 806R. The full primer sequences were as follows: GAPDH-F-515F: 5'-CAT TGG CAA TGA GCG GTT C-GTG CCA GCM GCC GCG GTA A-3', in which CAT TGG CAA TGA GCG GTT C represents the GAPDH F tag and GTG CCA GCM GCC GCG GTA A represents the 16S rRNA gene forward primer sequence (515F); and 806R: 5'-GGA CTA CHV GGG TWT CTA AT-3'.”

      (6) Lines 310-311: Explicitly define GAPDH and justify its use as a synthetic tag.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we now explicitly define GAPDH as glyceraldehyde-3-phosphate dehydrogenase, a human housekeeping gene. Specifically, the GAPDH-derived sequence was selected for two reasons. First, it is highly divergent from known soil microbial 16S rRNA gene sequences and did not produce detectable amplification when tested with soil DNA using the GAPDH tagged 806R primer pair, indicating that it would not interfere with endogenous soil DNA signals. Second, its melting temperature was compatible with that of the reverse primer, which allowed stable amplification of the tagged 16S amplicons under our PCR conditions.

      L360-371

      “GAPDH is a primer for a human housekeeping gene and it has no homologous sequences in soils. Subsequently, GAPDH was selected as the label primer based on two criteria. First, this primer was selected to avoid interference from the original soil sequences (Huang et al., 2014; Yang et al., 2021; Arvizu-Hernandez et al., 2025), and no detectable PCR amplification was observed for the primer set GAPDH F-806R across all the soil DNA samples included in this study. Second, the melting temperature (Tm) value of GAPDH F approximately matched that of 806R. The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing”

      (7) Line 346: Start a new paragraph to clearly separate this as a distinct experiment.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have started a new paragraph.

      (8) Line 346: Specify the number of samples analyzed for consistency.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have now explicitly specified the number of samples.

      A total of 120 samples were analyzed in this moisture gradient experiment: 2 ecosystems (Kaiyuan and Dashanbao) × 5 moisture levels (10%, 25%, 50%, 75%, and 100% of water holding capacity) × 6 incubation time points (0, 1, 3, 6, 12, and 24 days) × 2 replicates. Just two technical replicates were performed for this validation experiment, as the primary aim was to assess the trend of moisture effects rather than statistical inference across replicates.

      L408-410

      “This complementary experiment included two sites, five moisture levels, six incubation time points, and two replicates per treatment combination, resulting in a total of 120 soil samples.”

      (9) Lines 351-352: Replace "harvested" with "collected."

      We have replaced “harvested” with “collected” as suggested.

      (10) Line 367: Clarify how Illumina adapters and indices were added (e.g., two-step PCR, fusion primers).

      We thank the reviewer for this helpful suggestion. We have revised the Methods section to clarify how Illumina adapters and indices were incorporated. We used a pooled amplicon library preparation strategy. Individual samples were first amplified with primers containing sample-specific barcode sequences. The barcoded amplicons from multiple samples were then pooled and used for library preparation with the ALFA-SEQ DNA Library Prep Kit. Universal Illumina-compatible adapters were first ligated to the pooled amplicons. After bead-based purification, an indexing PCR was performed using the index primer mix, which introduced the complete P5/P7 sequences and a library-level Illumina index into the library molecules. Thus, sample demultiplexing was based on the sample-specific barcodes introduced during amplicon PCR, whereas the Illumina index was used to identify the pooled sequencing library. We have clarified this procedure in the revised manuscript.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).

      (11) Provide more detail on chimera removal, filtering thresholds, and normalization choices.

      We thank the reviewer for this helpful suggestion. The raw paired-end reads were first merged, and primer sequences were removed using the search_pcr2 script in USEARCH. Reads with more than two primer mismatches were discarded. Quality filtering was then performed using fastq_filter, and sequences with quality scores below 20 were removed. Redundant reads were collapsed using fastx_uniques. Amplicon sequence variants (ASVs) were generated using the UNOISE3 denoising algorithm, which also performs built-in chimaera filtering during ASV inference. In addition, ASVs with total sequence counts fewer than 9 were excluded to reduce the influence of low-frequency noise.

      L443-460

      “Briefly, paired-end reads were merged using USEARCH, and primer sequences (GAPDH-F-515F and 806R) were removed using the search_pcr2 script. Reads with more than two primer mismatches were discarded. Quality filtering was performed using the fastq_filter script, and sequences with quality scores below 20 were removed. Redundant sequences were dereplicated using the fastx_uniques script. ASVs were generated using the UNOISE3 non‑clustering denoising algorithm (Edgar, 2016), which infers 100% exact sequence variants by distinguishing biological sequences from PCR/sequencing errors. ASVs with total sequence counts fewer than 9 across all samples were removed to reduce noise. To quantify the abundance of each ASV, an ASV table was generated by mapping the quality‑filtered raw reads back to the ASV set using the otutab command. A 97% similarity threshold was applied for this recruitment to accommodate stochastic sequencing noise while maintaining biological resolution. Crucially, the mapping followed a best-hit priority rule, where each read was assigned to the ASV with the highest per cent identity within the 97% radius. This approach ensures that reads derived from the same biological template are accurately counted toward their respective ASV, preventing the underestimation of abundances that would occur with exact matching while strictly preserving the single-nucleotide resolution of the ASV framework. Taxonomic annotation of the ASVs was performed in QIIME2 with the Silva v138 database. A total of 89322 prokaryotic ASVs were obtained. To standardize sequencing depth across samples, the read number of each sample was rarefied to 53251 using the rarefy function in the vegan package in R.”

      (12) Line 412: Rephrase to refer to 16S amplicon addition rather than 16S rRNA genes (along the whole text), as only the V4 region is analyzed.

      We thank the reviewer for this helpful suggestion. we have rephrased references to “16S rRNA genes” to “16S rRNA gene amplicon fragments”

      (13) Ensure consistent primer naming throughout (e.g., GAPDH F vs ACTF).

      We have checked the entire manuscript and confirm that only “GAPDH F” is used as the label primer.

      (14) Finally, the manuscript would benefit from careful language editing. Several typographical errors, grammatical inconsistencies, and unclear phrases are present throughout. Examples include:

      Misspellings such as "diffrence" (e.g., figure legends) and inconsistent capitalization. Inconsistent terminology (e.g., "genes," "amplicons," and "fragments" used interchangeably without clarification). Redundant or awkward phrasing (e.g., repeated use of "extracellular 16S rRNA genes"). Occasional subject-verb agreement issues and missing articles.

      We sincerely apologize for the language issues. The manuscript has now undergone a thorough language editing process by a native English‑speaking colleague.

      Recommendation

      Major revision: The manuscript addresses an important problem and presents a promising approach. However, key issues related to conceptual clarity, bioinformatic consistency, statistical rigor, and interpretation of PMA-based results must be resolved. With substantial revision and clarification, the study has the potential to make a meaningful contribution to the field.

      We sincerely thank the reviewer for the thorough, constructive, and critical evaluation of our manuscript. We greatly appreciate the recognition that our study addresses an important problem and presents a promising approach. We also acknowledge the key issues raised regarding conceptual clarity, bioinformatic consistency, statistical rigor, and interpretation of PMA‑based results. We have taken these comments very seriously and have substantially revised the manuscript accordingly, more details about the revisions are described in the following point-by-point responses.

      Reviewer #2 (Recommendations for the authors):

      Editorial comments:

      (1) Title: I recommend removing "across China" from the title. In many ways, the study has nothing to do specifically with China, and you limit the broad applicability of the study. The same work could have been done with soils from Africa, for example. Also, it might be ok to remove 16S rRNA as well. The 16S rRNA genes are a proxy for rates of extracellular DNA degradation, but the study isn't exactly about 16S either.

      We thank the reviewer for this thoughtful suggestion regarding the title. We have revised the title to “The overall and sequence-specific degradation of soil extracellular DNA fragments: rates and influential factors.”

      (3) L44-45: "...such as real-time PCR, high-throughput amplicon sequencing, and metagenomic analysis...".

      We thank the reviewer for this suggestion. We have revised the order according to the suggestions of the reviewer.

      L44-45

      “The investigation of soil microbial abundance and diversity heavily relies on DNA-based technologies, such as real-time PCR, high-throughput amplicon sequencing, and metagenomic analysis.”

      (4) L48: remove "they".

      We agree with the reviewer and have removed the extraneous “they”.

      (5) L51: "noise factor"; "...persistence can lead to...".

      We have revised the sentence as suggested.

      (6) L53: remove theoretical.

      We have removed “theoretical”.

      (7) L58: remove "the".

      We have removed "the".

      (8) L86: Is restriction digestion of DNA a likely extracellular process in soil?

      We thank the reviewer for this thoughtful comment. We agree that the original wording may have overstated the likelihood of classical restriction digestion as a dominant extracellular process in soils. Our intention was not to suggest that intracellular restriction enzyme systems operate directly in the soil matrix in the same manner as they do within living cells. Rather, we aimed to indicate more generally that sequence-dependent nuclease susceptibility could contribute to differential degradation among extracellular DNA fragments.

      L87-95

      “First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C: N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021).”

      (9) L94-99: The authors might also consider the different nitrogen content of different bases; this might also affect sequence-specific selection of DNA for degradation.

      We thank the reviewer for this insightful suggestion. We agree that differences in the elemental composition of DNA bases, including nitrogen content, may provide an additional mechanistic explanation for sequence-dependent degradation. In the revised manuscript, we have incorporated this point into the Introduction.

      L92-95

      “Differences in base composition also alter the elemental stoichiometry (e.g., C:N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021).”

      (10) L110-112: These are not really written in hypothesis form. Also, what about a hypothesis about degradation rates and soil type/temperature/moisture?

      We thank the reviewer for this constructive critique. We have rewritten the hypotheses. To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation. This adjustment ensures that the hypotheses are directly grounded in the theoretical framework.

      L99-104

      “Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (11) L116: "GAPDH F-labeled 16S rRNA gene amplicon fragments....".

      We thank the reviewer for this helpful suggestion. we have rephrased references to “16S rRNA genes” to “16S rRNA gene amplicon fragments”

      (12) L117: "rapidly".

      We agree with the reviewer and have revised.

      (13) L118-120: "After a 48-day incubation period, 0.2 to 3.1% of the initial spike GADPH F-labeled 16S rRNA gene amplicon fragments ...".

      We agree with the reviewer and have revised.

      (14) L125: Spell out SEM in first usage.

      We thank the reviewer for this suggestion. In the revised manuscript, we have spelled out SEM as Structural equation modeling.

      (15) L128: I don't like the idea of putting this Figure in supplemental materials.

      We thank the reviewer for this suggestion. We have moved Figure S2 (moisture gradient microcosm experiment) to the main text as Figure 1f.

      (16) L154: The term "intracellular prokaryotic abundance" is not the right term. This makes one think of an intracellular parasite. I think you want something like: "Approximately 40% of sequences in total soil DNA extraction NGS amplicon libraries were derived from intact cells, while the remaining represented extracellular DNA. Conversely, greater than 80% of observed richness was derived from intact cells." (Please check that I stated this correctly.) I would also suggest some statistics or ranges here.

      We thank the reviewer for this important terminological clarification. We agree that the term “intracellular prokaryotic abundance” is misleading, as it could imply intracellular parasites. In the revised manuscript, we have replaced this with a clearer description and We have also added the across‑site ranges to provide statistical context.

      L163-166

      “The PMA treatment revealed that intact cells accounted for approximately 40% (range: 9–73%) of the total 16S rRNA gene copies. In contrast, over 80% (range: 27–97%) of the observed ASV richness was associated with sequences originating from intact cells (Fig. 4a and b).”

      (17) L168: "...a significant NEGATIVE correlation was observed...".

      We agree with the reviewer and have revised.

      (18) L169: "However, no significant relationship was observed...".

      We agree with the reviewer and have revised.

      (19) L194-195: What about pH and temperature?

      We thank the reviewer for this comment. We agree that pH and temperature are important environmental factors that can influence microbial DNA degradation and community composition. However, our results (Fig. 1c) indicate that soil moisture is the most dominant factor affecting extracellular DNA degradation. Therefore, in the revised manuscript, we have focused the explanation primarily on soil moisture, while acknowledging that pH and temperature may also be important influencing factors.

      L208-211

      “Third, environmental factors, including soil moisture, pH, and temperature, can predominantly govern enzymatic reaction rates (He et al., 2024; Shah et al., 2024). Indeed, strong positive correlations were observed between moisture content and eDNA degradation rates in both the survey and microcosm experiments (Fig. 1d-f).”

      (20) L199: "findings".

      We have revised as suggested.

      (21) L227-229: This sounds more like results.

      We thank the reviewer for this comment. We agree that the original first sentence in L227–229 reads more like results. Our intention was to introduce the discussion by linking extracellular DNA to potential impacts on prokaryotic community analysis, rather than to present specific findings at this point. We have reorganized this section as follows.

      L246-248

      “Accordingly, we further explored how DNA may influence prokaryotic community analyses using PMA treatment, and significant disparities were observed between the profiles of the total and PMA-treated soil prokaryotic communities (Fig. 4).”

      (22) L230: Need to also consider differential cell lysis during DNA extraction.

      We thank the reviewer for this important comment. We agree that differential cell lysis during DNA extraction could influence the observed community profiles, as microbial taxa differ in cell wall composition and resistance to mechanical or chemical lysis. In the revised manuscript, we explicitly acknowledge this limitation in the relevant section. We also clarify that a standardized DNA extraction protocol (DNeasy PowerSoil kit) was used to efficiently lyse a broad range of microbial taxa, but some taxon-specific lysis bias may remain. Future studies could combine multiple lysis methods or spike-in controls to quantify and correct for potential extraction bias.

      L262-265

      “However, as DNA extraction efficiency may differ between intact cells and eDNA, the actual differences between total and living prokaryotic abundance could be smaller than those observed in this study. Similarly, the overestimated prokaryotic richness may arise from historically accumulated microbial taxonomic information stored in eDNA pools (Deshpande and Fahrenfeld, 2023; Wang et al., 2024).”

      L309-311

      “Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015).”

      (23) L232: Need to also consider that PMA treatment is not perfect and can be affected by substrate, the ability of light to access DNA for crosslinking, etc.

      We thank the reviewer for this important reminder. We agree that PMA treatment is not perfect and that its efficiency can be affected by soil matrix properties (e.g., organic matter, clay minerals) and the ability of light to penetrate the sample for DNA crosslinking. In the revised manuscript, we have explicitly acknowledged these limitations in the discussion.

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (24) L240: Can extracellular DNA have an ecological role?

      We thank the reviewer for this thoughtful question. Yes, extracellular DNA (eDNA) does have important ecological roles beyond being a potential bias in molecular analyses. In the revised manuscript, we have added statements to highlight that extracellular DNA can serve as a nutrient source (e.g., nitrogen and phosphorus) for microbes and may also contribute to horizontal gene transfer. This emphasizes that extracellular DNA may actively influence microbial community structure and function, in addition to its role in potentially inflating observed abundance and richness.

      L267-275

      “We observed a significant correlation between eDNA degradation rates and the overall structure of the prokaryotic community, but this relationship was absent in PMA-treated communities (Fig. 5b). This discrepancy highlights the divergent ecological roles of extracellular and intracellular DNA. Analyses of the total community integrate intracellular DNA from metabolically active cells with eDNA which primarily originates from historical microbial residues (Lennon et al., 2018). EDNA incorporates signals that likely reflect the legacy effects of past environmental conditions (Wang et al., 2021). In contrast, the PMA-treated community reflects transient microbial activity driven by current selective pressures. Additionally, eDNA can serve as a nutrient source and facilitate horizontal gene transfer, which may further shape its interactions with contemporary microbial communities (Levy-Booth et al., 2007).”

      (25) L256: Why would microorganisms selectively degrade one DNA sequence vs another? This seems to be likely to be stochastic in terms of which sequences are taken up by microorganisms. However, different DNA sequences might hydrolyze differently or be otherwise damaged, and that could lead to differential degradation of a viable amplicon. It might be interesting to incorporate long pieces of DNA with different internal primer sites and use quantitative PCR to determine how sequences are degrading.

      We thank the reviewer for this important mechanistic insight. We agree that the observed correlation between degradation rate and sequence abundance does not necessarily imply active microbial preference. It could equally reflect stochastic encounter rates or intrinsic chemical differences (e.g., AT‑rich regions hydrolyzing faster). We have revised the corresponding paragraph in the Discussion.

      L279-292

      “This finding suggests that abundant eDNA degrades at a faster rate compared to rare eDNA. As mentioned earlier, this could be explained by several mechanisms. First, as soil eDNA is subject to enzymatic degradation and microbial recycling, abundant DNA sequences may be more likely to be encountered and degraded by extracellular nucleases simply due to their higher copy numbers (Levy-Booth et al., 2007; Nagler et al., 2018). Similarly, if microbes preferentially take up DNA as a nutrient source, they may degrade abundant sequences more frequently as a stochastic consequence of higher encounter rates (Finkel and Kolter, 2001). However, we also found that the relationships between the sequence-specific degradation rates and the effect sizes of extracellular 16S rRNA gene amplicon fragments varied across the study sites (Fig. S1g). The sequence-specific effect sizes of extracellular 16S rRNA gene amplicon fragments are mainly determined by both their production and degradation rates (Pietramellara et al., 2009; Sirois and Buckley, 2019). These inconsistent correlations emphasize the critical role played by the production rates of extracellular 16S rRNA genes in influencing the analysis of prokaryotic communities. Therefore, future studies should systematically determine both the production and degradation rates of eDNA.”

      (26) L282-283: This belongs in the discussion.

      We agree with the reviewer and have revised accordingly.

      (27) L289: "as well as measurements of total organic carbon".

      We agree with the reviewer and have revised accordingly.

      (28) L338: Any water content for these soils?

      We thank the reviewer for this comment. The water contents of soils from all study sites are reported in Supplementary Table 2.

      (29) L349-350: You mean that you measured the total soil extracted DNA and then added 1% as labeled 16S?

      Yes, for each soil sample, we extracted total soil DNA and quantified its concentration (ng DNA per gram of soil). We then added exogenous GAPDH‑tagged 16S amplicon fragments at an amount equal to 1% of this total DNA concentration. This concentration was chosen to mimic a realistic pulse of extracellular DNA input without overwhelming the endogenous DNA pool. We apologize for any confusion caused by the imprecise wording in the original manuscript.

      L392-398

      “The microcosm experiment was conducted using 30 g of soil for each sample. After pre-incubation at 20℃ for one week, each soil was thoroughly mixed with the GAPDH F‑tagged 16S rRNA gene amplicon fragments and incubated further at 20℃ (Fig. S8). The amount of exogenous GAPDH F‑tagged 16S rRNA gene amplicon fragments added to each soil sample was equivalent to 1% of the total DNA concentration naturally present in that soil, as determined fluorometrically prior to the experiment. This concentration was chosen to approximate natural eDNA fluxes resulting from microbial lysis, ensuring experimental relevance to in situ conditions (Table S2).”

      (30) L354: Remember that soil recovery from intact cells is going to be lower than for extracellular DNA. So, you are probably overestimating the contribution of extracellular DNA to the total DNA in the system.

      We thank the reviewer for this comment. We agree that DNA recovery from intact cells is generally lower than from extracellular DNA due to differential cell lysis efficiencies. Consequently, the contribution of extracellular DNA to total soil DNA may be somewhat overestimated in our study. We have clarified this limitation in the revised manuscript.

      L262-265

      “However, as DNA extraction efficiency may differ between intact cells and eDNA, the actual differences between total and living prokaryotic abundance could be smaller than those observed in this study. Similarly, the overestimated prokaryotic richness may arise from historically accumulated microbial taxonomic information stored in eDNA pools.”

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (31) L362: Amplification efficiency is pretty low. I think you would have been better served with GAPDH on both ends, and that would have given you a much higher efficiency qPCR.

      We thank the reviewer for this comment. The actual qPCR amplification efficiency in our assay was approximately 85%, which, although slightly below the ideal range, was still acceptable and produced reproducible amplification curves and reliable quantification for degradation-rate calculations.

      We acknowledge that the amplification efficiency in our qPCR experiments using a GAPDH F-labeled 16S primer on one end was suboptimal. The current design used a single GAPDH tag at the forward primer to avoid potential amplification bias or primer-dimer formation that could arise from extending the degenerate reverse primer. In addition, dual-end labeling would have required synthesis of new barcode-labeled tagged primers, increasing both cost and experimental complexity. Thanks again for the constructive comments, which provided us with the direction for future experiment optimization.

      L365-371

      “The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      (32) L367: Not enough detail on how barcoded libraries were made. UDIs?

      We thank the reviewer for this helpful comment. We have now clarified the library preparation and indexing strategy in the revised Methods section. This amplicon diversity sequencing used a pooled-library strategy. Individual samples were first distinguished by sample-specific inline barcodes introduced during the amplicon PCR step. After amplification, barcoded PCR products from multiple samples were pooled and subjected to library construction using the ALFA-SEQ DNA Library Prep Kit. Universal Illumina-compatible adapters were ligated to the pooled amplicons, followed by an indexing PCR that introduced the complete P5/P7 sequences and a library-level Illumina index. Thus, the Illumina index was used to identify the pooled sequencing library, whereas sample demultiplexing was performed according to the sample-specific inline barcodes. We have revised the Methods section to make this procedure explicit.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).”

      (33) L368: Why was the # of cycles increased?

      Thank you for your question. In the original manuscript (L368), we stated that the number of PCR cycles was increased to 35. This was mainly because the exogenously added GAPDH F‑labeled 16S rRNA genes had a relatively low initial abundance in the soil and gradually degraded during the microcosm incubation, with their copy numbers becoming particularly low at the last time points (see Fig. 1a). To ensure sufficient PCR product for high‑throughput sequencing from samples at all time points (especially those with low abundance at later stages), we appropriately increased the cycle number to 35.

      L429-431

      “The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing.”

      (34) L372: Were sequencing adapters ligated onto the pool?

      We thank the reviewer for this question. Yes, in this amplicon diversity sequencing workflow, sequencing adapters were ligated onto the pooled amplicon products. Briefly, individual samples were first amplified with sample-specific barcode sequences, allowing each sample to be distinguished after sequencing. The barcoded PCR products from multiple samples were then pooled for library construction. Universal Illumina-compatible adapters were ligated to this pooled amplicon library using the ALFA-SEQ DNA Library Prep Kit. After adapter ligation and purification, an indexing PCR was performed to introduce the complete P5/P7 sequences and a library-level Illumina index. We have clarified this pooled-library construction workflow in the revised Methods section.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).”

      (35) L378: "Amplicon sequence variants".

      We agree with the reviewer and have revised accordingly.

      (36) L380: Why were ASVs with fewer than 9 reads removed?

      We thank the reviewer for this question. The threshold of removing ASVs with fewer than 9 total reads across all samples was applied to reduce noise from sequencing errors and PCR artifacts. Our justification is supported by both the default parameters of the UNOISE3 algorithm and common practice in amplicon sequencing analysis.

      The USEARCH manual specifies that the -minsize parameter in the unoise3 command defaults to 8. This means that unique sequences occurring fewer than 8 times are discarded by the algorithm during ASV inference, as they are unlikely to represent true biological variants. Our threshold of 9 is slightly more conservative than the default (9 > 8), ensuring that only ASVs with a minimal level of abundance are retained. This choice is directly aligned with the algorithm’s intrinsic noise‑filtering logic.

      (37) L402: Please don't forget to discuss that PCR bias can contribute to uncertainty in the abundance of each taxon.

      Thank you for this important reminder. We agree that PCR bias (e.g., primer‑template mismatches, GC content differences, and variable amplification efficiency) can contribute to uncertainty in the abundance estimates of each taxon. Following your suggestion, we have now added a paragraph in the Discussion section to address this issue. We state that sequence‑specific degradation rates and PCR bias may jointly affect the accuracy of taxon abundance estimates, and future studies should incorporate internal standards or multiplex PCR strategies to correct for such biases. Thank you for your careful review.

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      (38) L414: Suggest: "To inhibit amplification of extracellular DNA, soils were incubated with propidium monoazide (PMA), as described previously (REF). Briefly, soil (X grams) was mixed with PMA in a total volume of Y (ml).

      We thank the reviewer for this suggestion. We have revised the Methods section to provide a clearer description of PMA treatment, specifying the soil amount (0.50 g) and the total volume (0.5 mL).

      L496-497

      “To inhibit amplification of eDNA, soils were incubated with PMA, as described previously (Carini et al., 2016).”

      L505-506

      “In this study, 0.50 g of soil was mixed with PMA in a total volume of 0.5 mL (40 µM PMA in phosphate‑buffered saline, PBS), while the control soil samples were mixed with PBS without PMA.”

      (39) L416: In contrast, microbes with intact cell membranes exclude PMA, and their DNA is not cross-linked with PMA, and remains amenable to PCR amplification.

      We agree and have revised.

      (40) L418-420: wording/sentence is strange and needs work.

      Thank you for pointing this out. We have reviewed the sentence at L418‑420 and agree that the wording is awkward. Moreover, the content only listed the advantages of the PMA method without acknowledging its limitations, making the statement less balanced. Therefore, in the revised manuscript, we have deleted this sentence. The limitations of the PMA method have been addressed in the Discussion section.

      L502-503

      “Currently, PMA treatment is a widely used to suppress PCR amplification of eDNA (Xue et al., 2023; Canini et al., 2024).”

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (41) L421-422: PMA treatment is a widely used method for inhibiting the enzymatic processing of extracellular DNA (Xue, Canini).

      We agree and have revised.

      (42) L425: include volume of PBA.

      We thank the reviewer for this comment. We have revised the Methods section to include the volume of PMA used

      L505-506

      “In this study, 0.50 g of soil was mixed with PMA in a total volume of 0.5 mL (40 µM PMA in phosphate‑buffered saline, PBS).”

      (43) L429-430: Don't use the word precipitates- use "pellets".

      We agree and have revised.

      (44) L433: "The abundance of 16S rRNA genes was determined using quantitative PCR employing a LightCycler...".

      We agree and have revised.

      (45) L445-: Section 4.9 - needs citations for PERMANOVA, NMDS, SEM, etc.

      Thank you for your suggestion. We have added the necessary citations for PERMANOVA, NMDS, SEM, and other methods in Section 4.9.

      L531-539

      Prokaryotic community structure differences among the study sites and incubation time points were examined through non-metric multidimensional scaling analysis (NMDS), permutation multivariate analysis of variance (PERMANOVA), and Permutational Analysis of Multivariate Dispersion (PERMDISP) (Kruskal, 1964; Anderson, 2001). Random forest modeling was conducted to assess the importance of environmental and soil variables in predicting the overall degradation rates of extracellular 16S rRNA gene amplicon fragments. Structural equation modeling (SEM) was employed to further evaluate the direct and indirect effects of soil moisture, soil pH, MAP, and prokaryotic abundance on the overall degradation rates of extracellular 16S rRNA gene amplicon fragments (Grace, 2006).

      (46) L698: A few comments. It would be nice to know how many different 16S sequences were tracked for differential degradation and shown in the figure.

      We thank the reviewer for this helpful comment. We would like to clarify that Fig. 1A does not track the degradation of individual 16S rRNA gene amplicon sequences, but instead shows the overall degradation dynamics of the total added exogenous DNA pool. The data points are derived from total 16S gene copy numbers measured via qPCR at each incubation time point. Consequently, this quantification inherently includes all sequences present within the added pool. The multiple lines visualized in the figure represent the collective degradation trajectories of the entire DNA pool across different study sites

      To address sequence-level changes, we further analyzed the richness and composition of the GAPDH F-tagged 16S rRNA gene amplicon fragments, which are presented in Fig. 2A and related analyses.

      (47) L699: Better to use "16S rRNA gene amplicon fragment abundance" as the term.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have replaced the original wording with “16S rRNA gene amplicon fragment abundance” where appropriate.

      (48) Y-axis for Figures 1A and 2A should be GAPDH-labeled, not ACTB-labeled.

      We apologize for this mistake. We have corrected this error in the revised manuscript.

      (49) For Figure 1b: Why not use box plots and ANOVA for different soil types?

      Thank you for your valuable suggestion. In the original Figure 1b, we used a bar plot to display the degradation rate constants across the 30 study sites. This choice was intended to emphasize the continuous variation among sites and their gradient relationships with environmental factors (e.g., soil moisture, MAP), which were then used in random forest and structural equation modeling. The bar plot better illustrates the spatial continuum of degradation rates rather than treating ecosystem types as discrete categories.

      Nevertheless, we fully agree that a boxplot grouped by ecosystem type (grassland, forest, cropland, desert) would help readers quickly grasp the overall differences among land‑use types. In the revised manuscript, we have added a boxplot grouped by ecosystem type and performed one‑way ANOVA followed by Tukey HSD post‑hoc tests (Fig. 1.). The results show that degradation rate constants differ significantly among ecosystem types (P < 0.05).

      L124-128

      “The degradation rate constants of the spiked extracellular 16S rRNA gene amplicon fragments displayed considerable variability among the study sites, ranging from 0.05 to 0.16 day<sup>-1</sup> (Fig. 1b). Furthermore, we found that degradation rate constants differed significantly among ecosystem types (Fig. 1c, P < 0.05). Specifically, cropland and forest soils exhibited significantly higher degradation rates than grassland soils (P < 0.05).”

      (50) For Figure 2: Where are PERMANOVA and PERMDISP values for the figure?

      We thank the reviewer for this comment. In the revised manuscript, we have added the PERMANOVA and PERMDISP values corresponding to Figure 2 in the figure legend and Results section (Fig. 2).

      (51) I found Figure 2b to be hard to see. The 48-day circles are almost invisible. Difficult to know what the authors are trying to show here, since there is so much variability associated with soil type.

      We thank the reviewer for this comment. Figure 2b is intended to illustrate the temporal changes in microbial community structure during the incubation. The different colored circles represent samples at different time points (1, 3, 6, 12, 24, and 48 days), showing how communities shift over time. We apologize that in the original Figure 2b, the 48‑day samples were nearly invisible and that the high variability among soil types obscured the intended message. In the revised manuscript, we have added a black border around every data point, which greatly enhances the visibility of the 48‑day samples (and all time points). We now use distinct shapes to represent different ecosystem types (grassland, forest, cropland, desert) in the NMDS ordination, and added PERMANOVA results in both the Results section and the figure legend (Fig. R2b).

      (52) Figure 4A: Y-axis need a label like "16S rRNA gene abundance".

      We agree and have revised.

      (53) Figure 4B: I'd like to see a Shannon index too, not just richness.

      Thank you for your suggestion. We agree that the Shannon index, which integrates both richness and evenness, provides a valuable complement to richness alone. In the revised manuscript, we added an analysis of the Shannon index to compare α‑diversity between total DNA (PMA‑untreated) and intact cell DNA (PMA‑treated) samples (Fig.4).

      (54) Figure 4D: Would be good to have lines linking the intact cell vs total abundance. Also, what about a box plot of Bray-Curtis (or similar) dissimilarity between intact cell and total microbial analysis across the dataset?

      Thank you for your suggestions. Regarding the addition of connecting lines in Figure 4D, after careful consideration we decided not to add them for the following reason: the total and PMA-treated communities from the same site are already coded with the same color (different colors for different sites), which effectively indicates the pairing. Adding lines would greatly reduce readability due to dense overlapping lines, especially given the number of sites. Therefore, we kept the original color‑based pairing design.

      To address your second suggestion, we have added a bar plot showing the distribution of Bray‑Curtis dissimilarities between total (PMA‑untreated) and intact cell (PMA‑treated) communities across all study samples (Fig. R3d).

      (55) Figure 5B: What do correlations with p > 0.05 show? I would remove these from the image.

      We thank the reviewer for this suggestion. We agree that correlations with p > 0.05 do not represent statistically significant relationships and may cause confusion. In the revised manuscript, we have removed these non-significant correlations from Figure 5B.

      (56) Figure 6: "Incubations of 0, 3, 6, 12, 24, and 48 days".

      We agree with the reviewer and have revised as suggested.

      References

      Albertsen, M., Karst, S.M., Ziegler, A.S., Kirkegaard, R.H., Nielsen, P.H., 2015. Back to basics–the influence of DNA extraction and primer choice on phylogenetic analysis of activated sludge communities. PLoS One 10, e0132783.

      Anderson, M.J., 2001. A new method for non‐parametric multivariate analysis of variance. Austral Ecology 26, 32-46.

      Arvizu-Hernandez, E., Ocadiz-Delgado, R., Gariglio, P., 2025. E7HPV16 Oncogene and 17beta-Estradiol Stress Promote Oncogenic microRNA Expression Patterns, Cell Proliferation and Cervical Intraepithelial Neoplasia 1. Cell Biochemistry and Function 43, e70065.

      Buitrago, D., Labrador, M., Arcon, J.P., Lema, R., Flores, O., Esteve-Codina, A., Blanc, J., Villegas, N., Bellido, D., Gut, M., 2021. Impact of DNA methylation on 3D genome structure. Nature Communications 12, 3243.

      Cai, P., Huang, Q., Zhang, X., Chen, H., 2006a. Adsorption of DNA on clay minerals and various colloidal particles from an Alfisol. Soil Biology and Biochemistry 38, 471-476.

      Cai, P., Huang, Q.Y., Zhang, X.W., 2006b. Interactions of DNA with clay minerals and soil colloidal particles and protection against degradation by DNase. Environmental Science & Technology 40, 2971-2976.

      Carini, P., Marsden, P.J., Leff, J.W., Morgan, E.E., Strickland, M.S., Fierer, N., 2016. Relic DNA is abundant in soil and obscures estimates of soil microbial diversity. Nature Microbiology 2, 1-6.

      Deshpande, A.S., Fahrenfeld, N.L., 2023. Influence of DNA from non-viable sources on the riverine water and biofilm microbiome, resistome, mobilome, and resistance gene host assignments. Journal of Hazardous materials 446, 130743.

      Du, Y., Wang, Z., Liu, K., Chai, G., Chi, Y., Li, T., Duan, Y., Xia, T., Liu, D., Che, R., 2025. The performance of different methods in characterizing soil live prokaryotic diversity and abundance is highly variable. iMetaOmics, e70011.

      Emerson, J.B., Adams, R.I., Román, C.M.B., Brooks, B., Coil, D.A., Dahlhausen, K., Ganz, H.H., Hartmann, E.M., Hsu, T., Justice, N.B., 2017. Schrödinger’s microbes: tools for distinguishing the living from the dead in microbial ecosystems. Microbiome 5, 86.

      Finkel, S.E., Kolter, R., 2001. DNA as a nutrient: novel role for bacterial competence gene homologs. Journal of Bacteriology 183, 6288-6293.

      Frostegård, Å., Courtois, S., Ramisse, V., Clerc, S., Bernillon, D., Le Gall, F., Jeannin, P., Nesme, X., Simonet, P., 1999. Quantification of bias related to the extraction of DNA directly from soils. Applied and Environmental Microbiology 65, 5409-5420.

      Grace, J.B., 2006. Structural equation modeling and natural systems. Cambridge University Press.

      He, P., Li, L.-J., Dai, S.-S., Guo, X.-L., Nie, M., Yang, X., Kuzyakov, Y., 2024. Straw addition and low soil moisture decreased temperature sensitivity and activation energy of soil organic matter. Geoderma 442, 116802.

      Heise, J., Nega, M., Alawi, M., Wagner, D., 2016. Propidium monoazide treatment to distinguish between live and dead methanogens in pure cultures and environmental samples. Journal of Microbiological Methods 121, 11-23.

      Huang, C., Xie, D.C., Cui, J.J., Li, Q., Gao, Y., Xie, K.P., 2014. FOXM1c Promotes Pancreatic Cancer Epithelial-to-Mesenchymal Transition and Metastasis via Upregulation of Expression of the Urokinase Plasminogen Activator System. Clinical Cancer Research 20, 1477-1488.

      Knight, R., Vrbanac, A., Taylor, B.C., Aksenov, A., Callewaert, C., Debelius, J., Gonzalez, A., Kosciolek, T., McCall, L.-I., McDonald, D., 2018. Best practices for analysing microbiomes. Nature Reviews Microbiology 16, 410-422.

      Kruskal, J.B., 1964. Nonmetric multidimensional scaling: a numerical method. Psychometrika 29, 115-129.

      Lennon, J.T., Muscarella, M.E., Placella, S.A., Lehmkuhl, B.K., 2018. How, when, and where relic DNA affects microbial diversity. mbio 9, e00637-00618.

      Levy-Booth, D.J., Campbell, R.G., Gulden, R.H., Hart, M.M., Powell, J.R., Klironomos, J.N., Pauls, K.P., Swanton, C.J., Trevors, J.T., Dunfield, K.E., 2007. Cycling of extracellular DNA in the soil environment. Soil Biology and Biochemistry 39, 2977-2991.

      Liu, Q.H., Yuan, L., Li, Z.H., Leung, K.M.Y., Sheng, G.P., 2024. Natural organic matter enhances natural transformation of extracellular antibiotic resistance genes in sunlit water. Environmental Science & Technology 58, 17990-17998.

      Marrone, A., Ballantyne, J., 2008. Sequence Specificity of BAL 31 Nuclease for ssDNA Revealed by Synthetic Oligomer Substrates Containing Homopolymeric Guanine Tracts. PLoS One 3, e3595.

      McKinney, C.W., Dungan, R.S., 2020. Influence of environmental conditions on extracellular and intracellular antibiotic resistance genes in manure-amended soil: A microcosm study. Soil Science Society of America Journal 84, 747-759.

      Morrissey, E.M., McHugh, T.A., Preteska, L., Hayer, M., Dijkstra, P., Hungate, B.A., Schwartz, E., 2015. Dynamics of extracellular DNA decomposition and bacterial community composition in soil. Soil Biology and Biochemistry 86, 42-49.

      Nagler, M., Insam, H., Pietramellara, G., Ascher-Jenull, J., 2018. Extracellular DNA in natural environments: features, relevance and applications. Applied Microbiology and Biotechnology 102, 6343-6356.

      Nocker, A., Sossa-Fernandez, P., Burr, M.D., Camper, A.K., 2007. Use of propidium monoazide for live/dead distinction in microbial ecology. Applied and Environmental Microbiology 73, 5111-5117.

      Pietramellara, G., Ascher, J., Borgogni, F., Ceccherini, M., Guerri, G., Nannipieri, P., 2009. Extracellular DNA in soil and sediment: fate and ecological relevance. Biology and Fertility of Soils 45, 219-235.

      Shah, A., Huang, J., Han, T., Khan, M.N., Tadesse, K.A., Daba, N.A., Khan, S., Ullah, S., Sardar, M.F., Fahad, S., 2024. Impact of soil moisture regimes on greenhouse gas emissions, soil microbial biomass, and enzymatic activity in long-term fertilized paddy soil. Environmental Sciences Europe 36, 120.

      Sirois, S.H., Buckley, D.H., 2019. Factors governing extracellular DNA degradation dynamics in soil. Environmental Microbiology Reports 11, 173-184.

      Vuillemin, A., Horn, F., Alawi, M., Henny, C., Wagner, D., Crowe, S.A., Kallmeyer, J., 2017. Preservation and significance of extracellular DNA in ferruginous sediments from Lake Towuti, Indonesia. Frontiers in Microbiology 8, 1440.

      Wang, X., Ganzert, L., Bartholomaus, A., Amen, R., Yang, S., Guzman, C.M., Matus, F., Albornoz, M.F., Aburto, F., Oses-Pedraza, R., Friedl, T., Wagner, D., 2024. The effects of climate and soil depth on living and dead bacterial communities along a longitudinal gradient in Chile. The Science of the total environment 945, 173846.

      Wang, Y.-T., Yang, W.-J., Li, C.-L., Doudeva, L.G., Yuan, H.S., 2007. Structural basis for sequence-dependent DNA cleavage by nonspecific endonucleases. Nucleic Acids Research 35, 584-594.

      Wang, Y., Yan, Y., Thompson, K.N., Bae, S., Accorsi, E.K., Zhang, Y., Shen, J., Vlamakis, H., Hartmann, E.M., Huttenhower, C., 2021. Whole microbial community viability is not quantitatively reflected by propidium monoazide sequencing approach. Microbiome 9, 1-13.

      Wolpe, J., Guertin, M., 2022. Regional and Single Nucleotide Correction of Sequence Bias in Chromatin Accessibility Data. The FASEB Journal 36.

      Yang, A., Liu, X., Liu, P., Feng, Y.Z., Liu, H.B., Gao, S., Huo, L.M., Han, X.Y., Wang, J.R., Kong, W., 2021. LncRNA UCA1 promotes development of gastric cancer via the miR-145/MYO6 axis. Cellular & Molecular Biology Letters 26, 33.

      Ye, M., Zhang, Z., Sun, M., Shi, Y., 2022. Dynamics, gene transfer, and ecological function of intracellular and extracellular DNA in environmental microbiome. iMeta 1, e34.

    1. eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

    2. Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. In particular, two types of learned associations are studied. One association involves associations between stimulus features and control states (SC), and the other involves stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates.

      The authors conclude that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing that SC and SR correlates are (1) not entirely overlapping in cross-decoding, (2) simultaneously observed on average over trials, (3) independently correlate with RT, and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Response:

      In their last response to the reviewers, the authors write:

      "... constructing a theoretically unbiased decoder requires perfectly counter-balanced training data (i.e., for every training trial of class A that is X trials away from the test data, there must be a training trial of all other classes that is exactly X trials away from the test data). As we were unable to achieve such a perfect design, we chose not to run an additional experiment."

      This isn't really an issue about whether this design is "perfectly" orthogonal. It's an issue regarding a clear confound among the decoded classes for SC/SR decoders. To be clear: of the 8 classes in the SC decoder, 4 are overwhelmingly presented in the first half (PHASE 2) of the session, whereas the other 4 are overwhelmingly presented in the second half (PHASE 3). The same is true for the SR decoder. So, session-half correlated noise could readily contribute to distinguishing among these classes. And counterbalancing this across subjects won't help because decoders lose sign.

      To me, the conducted control analyses don't really make strong contact with this issue. The split-half cross-validation is a nice idea but, as the authors acknowledge, it's also subject to slower cross-session noise, as is the original analysis. This sort of noise is not exactly exotic in EEG. Caps/hair/electrodes shift, gel dries and impedance changes, posture / muscle tension / skin conductance changes, fatigue may wax and wane (e.g., linked to increasing alpha), etc. And the newest analysis didn't really seem to engage with this issue either, as it only assessed minimum distances between classes, on the order of 5 +- 2 SD trials. This seems to assume that the dominant potential sources of noise will be scale-free, such that the strength of the relation at short time scales would generalize to longer ones. I'm not sure why that's expected here.

      Here are some suggestions for alternative control analyses that I think would be more targeted to this issue:

      (1) Explicitly train a decoder to separate the three levels of PHASE from each other. Successful decoding would provide positive evidence for the presence of structured noise at this timescale.

      (2) Specify an RDM for the PHASE variable and regress this component separately from each time-point/trial of the SC and SR decoders. This is a post-hoc band-aid, but it is in the spirit of correcting for a known confound.

      (3) In the spirit of the authors' distance analysis, but without assuming that the noise is scale-free: perform a time-series RSA like that in Alink et al. (2015; https://doi.org/10.1101/032391), Fig. 1 and 3. This would allow one, e.g., to estimate the structure & timescales of the noise processes across the session.

      Other readers may, like me, be puzzled by the selection of this particular experimental design to test this question of SC and SR coding, given the temporal confound among SC/SR classes, and given that there would seem to be many possible designs that are less confounded. For example, why not use a design where ISPC was swapped/shuffled several more times within each subject, so that PHASE is more orthogonal to long-timescale noise? Isn't ISPC learning fast enough to support learning phases shorter than 700 trials? Such readers would likely appreciate a frank discussion of this dilemma, and a motivation for the choice of the present design, within the manuscript.

      Pre-stimulus coding:

      To explain the apparent pre-stimulus coding of several task variables, the newest version of the manuscript proposes that subjects were proactively coding these variables via predictive mechanisms. This is an interesting account of item-specific control. It is also surprising, given that item-specific control mechanisms are typically conceptualized as reactive or stimulus-driven phenomena. But I think support for a proactive control account was incomplete. The mechanistic logic was not presented, and no hypotheses under this account were developed or tested. So I would suggest pinning down some hypotheses here and actually putting this account to the test.

      Outliers & t-values: thank you for checking this!

      Random slopes were omitted due to convergence failure, but this can inflate false positive inferences (e.g., Barr et al. 2013), and doesn't really motivate a minimal model. I'd suggest trying a slightly reduced model (e.g., drop correlations via `slope || subject`) using buildMer automated selection, or switching to brms.

    3. Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a solid design and the analyses are appropriate for the research questions.

      Strengths:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      Weaknesses:

      I still have some doubts on the effectiveness of the experimental manipulation on Phase 2: although the ISPC effect is still present, it is much weaker in comparison, suggesting the participants did not learn the contingency statistics in Phase 2 as well as they did in the other phases, due to either the lingering effect of the previous phase or an inherent bias towards one color pairs. Perhaps by separately plotting the earlier and later blocks of Phase 2 any difference can be revealed if it exists. This behavioral difference could result in unequal levels of SC/SR representation across phases, which may raise problems when data were combined for analyses that assume the neural effects are equivalent.

    4. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

      We thank the editors and the reviewers for their positive assessment and constructive feedback of our work, which led us to think more deeply about the conceptual and methodological aspects of this project and further strengthen the manuscript. Based on the comments, we included more control analyses and revised the manuscript accordingly. Please see below our responses to each comment raised in the reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. Two types of learned associations are characterized, one being associations between stimulus features and control states (SC), and the other being stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates, which are used to test properties of their dynamics.

      The conclusion is reached that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing SC and SR correlates are: (1) not entirely overlapping in cross-decoding; (2) simultaneously observed on average over trials in overlapping time bins; (3) independently correlate with RT; and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Thank you for acknowledging the strength in EEG decoding and design. We have addressed all your concerns raised below point by point.

      Weaknesses:

      I still have my concern from the first round that the decoders are overfit to temporally structured noise. As I wrote before, the SC and SR classes are highly confounded with phase (chunk of session). I do not see how the control analyses conducted in the revision adequately deal with this issue.

      In the figures, there are several hints that these decoders are biased. Unfortunately, the figures are also constructed in such a way that hides or diminishes the salience of the clues of bias. This bias and lack of transparency discourage trust in the methods and results.

      I have two main suggestions:

      (1) Run a new experiment with a design that properly supports this question.

      I don't make this suggestion lightly, and I understand that it may not be feasible to implement given constraints; but I feel that this suggestion is warranted. The desired inferences rely on successful identification of SC and SR representations. Solidly identifying SC and SR representations necessitates an experimental design wherein these variables are sufficiently orthogonalized, within-subject, from temporally structured noise. The experimental design reported in this paper unfortunately does not meet this bar, in my opinion (and the opinion of a colleague I solicited).

      An adequate design would have enough phases to properly support "cross-phase" cross-validation. Deconfounding temporal noise is a basic requirement for decoding analyses of EEG and fMRI data (see e.g., leave-one-run-out CV that is effectively necessary in fMRI; in my experience, EEG is not much different, when the decoded classes are blocked in time, as here). In a journal with a typical acceptance-based review process, this would be grounds for rejection.

      Please note that this issue of decoder bias would seem to weaken the rest of the downstream analyses that are based on the decoded values. For instance, if the decoders are biased, in the within-trial correlation analysis, how can we be sure that co-fluctuations along certain dimensions within their projected values are driven by signal or noise? A similar issue clouds the LMM decoding-RT correlations.

      We appreciate the reviewer’s concern with the potential confound of temporally structured noise (TSN) in the EEG data. As we understand it, TSN refers to a process that the noise structure drifts over time. It follows that noise structure should be more similar for temporally closer trials and that the TSN’s bias on decoding accuracy is stronger for test trials that are closer to the training data. In the previous round of revision, we conducted a control analysis that reduced the influence of TSN by maximizing the temporal distance between training and test data (the distance between the centers of the training and test data of the same SC/SR manipulation is about 400 trials given the experimental design) and showed comparable decoding accuracy with the main results. As the reviewer finds this analysis unconvincing, we reason that the reviewer believes that the TSN has a long-term effect, such that it remains relatively stable over time and can be picked up by trials temporally distant from the training data. With this assumption and the assumption that this effect may not be linear, constructing a theoretically unbiased decoder requires perfectly counter-balanced training data (i.e., for every training trial of class A that is X trials away from the test data, there must be a training trial of all other classes that is exactly X trials away from the test data). As we were unable to achieve such a perfect design, we chose not to run an additional experiment. Instead, we focused on testing whether and how much TSN systematically biased the reported decoding accuracy.

      Please note that the existence of TSN in the EEG data is not sufficient to rule that the decoding results are biased. As TSN is stronger for trials closer to each other, the idea that auto-correlation biases decoding results would predict a distance effect, such that if a test trial is closer to a training trial of the same trial type, the higher similarity in TSN between the training and test data would more strongly inflate the decoding accuracy of the test trial, resulting in a negative correlation between distance between a test trial and its closest training trial of the same type and the test trial’s decoding accuracy. To test this predicted negative correlation, in each fold and each repetition of the cross-validation reported in the SC-SC and SR-SR decoders in Fig. 4, we calculated the distance (mean=5.84 trials, SD=2.05, 5th percentile =2.87, 95th percentile=9.45, one trial = 2.4-2.6s) between each test trial and its closest training trial of the same trial type. This distance was used as the predictor to predict decoding accuracy in a linear regression. Note that even if the relation between distance and decoding accuracy is non-linear, the linear relation will be negative because the relation is monotonic (similarity in noise structure decreases monotonically with temporal distance between trials). Similarly, because the effect is monotonic, if a long-range effect exists, it should also exist in short-range and be picked up by the distance range in this analysis. The regression coefficient is averaged across cross-validation folds and repetitions for each subject to match how the decoding accuracy was reported in the main text. Finally, the averaged regression coefficient was tested against 0 using a one-sample t-test. This analysis was conducted at each time point (from -250ms to 1500ms) separately. As shown in the figure below, no time point exhibited the negative correlation as predicted by the auto-correlation account. An alternative explanation is that this result indicates that TSN remains stable over time. If this is the case, TSN will be shared by all trials and will be unable to bias decoding results. Together with the control analysis introduced previously, this new control analysis supports the notion that the decoding results are not inflated by TSN in the EEG data. We included all the control analyses in the revised manuscript (page 13-14). Please note that this analysis is specific for the present dataset and we strongly agree with the reviewer that TSN is a key confounding factor in EEG analysis in general and should be carefully addressed.

      Lastly, we understand the concern with the early onset of above-chance decoding accuracy. Here, we provide an explanation: because of the blocked design (i.e., participants performed hundreds of trials with the same SC/SR associations), it is possible the participants learned the associations and used them to guide proactive cognitive control. As proactive cognitive control is anticipatory and sustained (Braver, 2012; Khan et al., 2025), it may be able to be decoded early on a trial, or even before trial onset. In the revised manuscript, we discussed this account along with the TSN issue as a limitation of the current project and directions for future research (page 24).

      (2) Increase transparency in the reporting of results throughout main text.

      Please do not truncate stimulus-aligned timecourses at time=0. Displaying the baseline period is very useful to identify bias, that is, to verify that stimulus-dependent conditions cannot be decoded pre-stimulus. Bias is most expected to be revealed in the baseline interval when the data are NOT baseline-corrected, which is why I previously asked to see the results omitting baseline correction. (But also note that if the decoders are biased, baseline-correcting would not remove this bias; instead, it would spread it across the rest of the epoch, while the baseline interval would, on average, be centered at zero.)

      Please use a more standard p-value correction threshold, rather than Bonferroni-corrected p<0.001. This threshold is unusually conservative for this type of study. And yet, despite this conservativeness, stimulus-evoked information can be decoded from nearly every time bin, including at t=0. This does not encourage trust in the accuracy of these p-values. Instead, I suggest using permutation-based cluster correction, with corrected p<0.05. This is much more standard and would therefore allow for better comparison to many other studies.

      I don't think these things should be done as control analyses, tucked away in the supplemental materials, but instead should be done as a part of the figures in the main text -- including decoding, RSA, cross-trial correlations, and RT correlations.

      Thank you for your suggestions. we have added the baseline period from 200 to 0 ms prior to the stimulus onset in all the stimulus-locked analyses and tested the significance with cluster-based permutation test (cluster-forming threshold p < 0.001, cluster-level p < 0.05, (Collins & Frank, 2018)) in all the analyses including decoding, RSA, cross-trial correlations and RT correlations. The results showed similar patterns, and they are all reported in the main text (please see all the figures and page 30-32 in the main text).

      Other issues:

      Regarding the analysis of the within-trial correlation of RSA betas, and "Cai 2019" bias:<br /> The correction that authors perform in the revision -- estimating the correlation within the baseline time interval and subtracting this estimate from subsequent timepoints -- assumes that the "Cai 2019" bias is stationary. This is a fairly strong assumption, however, as this bias depends not only on the design matrix, but also on the structure of the noise (see the Cai paper), which can be non-stationary. No data were provided in support of stationarity. It seems safer and potentially more realistic to assume non-stationarity.

      This analysis was included in the supplemental material. However, given that the correlation analysis presented in the Results is subject to the "Cai 2019" bias, it would seem to be more appropriate to replace that analysis, rather than supplement it.

      Regardless, this seems to be a moot issue, given that the underlying decoders seem to be overfit to temporally structured noise (see point above regarding weakening of downstream analyses based on decoder bias).

      Thank you for this important point. We now replaced the previous control analysis with a new one that does not assume stationary noise structure (page 19 in the revised manuscript). In Cai et al (2019), the source of confound is the covariance between observations. Specifically, as the observations in fMRI data are the BOLD signal at different time points, TSN can introduce covariance between nearby observations, which further biases the observed correlation between experimental conditions/trial types. In our case, the observations are decoding accuracy for different trial types. Thus, bias in the correlation may come from covariance between trial types. In this study, potential covariance between trial types includes the constrain that the decoding accuracy of all trial types adds up to 1 for a given trial (although we transformed the accuracy into logits prior to RSA, so the constrain may not hold), and the blocked design (as discussed above). Thus, to establish a baseline level of correlation between SC and SR representation strength, we took a similar shuffling approach as in Cai et al (2019) and randomly shuffled the trial types within each block. The reason to shuffle within each block is to preserve the covariance structure in the blocked design. We then repeated the same analysis using the shuffled data. The results of 10 shuffled analysis were averaged to form a baseline. Please note that (1) this control analysis was performed separately at each time point, hence removing the assumption of stationary noise structure, (2) this analysis also included as noise any covariance introduced by the proactive cognitive control guided by the learned SC and SR associations (see response to comment 1), thus it is more stringent than intended and (3) this control analysis started from decoding and was intended to provide a baseline for all downstream analysis. As shown in figures 2A, 3A, 7C and 8C, the reviewer was correct that the bias was not stationary, as the baseline of correlation coefficient varies over time. Additionally, the SC-SR representation strength correlation remained significantly above baseline between ~100 and ~ 450 ms following stimulus onset and between -180 and + 50 ms relative to response, suggesting that the noise structure (even when including potential proactive cognitive control) cannot fully explain the observed the SC-SR representation strength correlation. Considering the fact that this control analysis treated proactive control as a source of confound, this result does not necessarily contradict the absence of distance effect reported above.

      Outliers and t-values:

      More outliers with beta coefficients could be because the original SD estimates from the t-values are influenced more by extreme values. When you use a threshold on the median absolute deviation instead of mean +/-SD, do you still get more outliers with beta coefficients vs t-values?

      Thank you for your suggestion. We calculated the proportion of outliers with a threshold of median absolute deviation (defined as values beyond median ± 5 median absolute deviation) for each subject. The outliers remained less frequent for t-values than for beta coefficients (t-values: mean = 1.08%, SD = 0.12%; beta-values: mean = 4.45%, SD = 0.28%). Based on these results and to maintain consistent with previous studies employing the methods (Cellier et al., 2022; Kikumoto & Mayr, 2020; Kikumoto et al., 2022a; Kikumoto et al., 2022b; Rangel et al., 2023), we still decided to stay with t-values.

      Random slopes:

      Were random slopes (by subject) for all within-subject variables included in the LMMs? If not, please include them, and report this in the Methods.

      Thank you for your suggestion. The model failed to converge with random slopes of all variables. Thus, we chose not to add random slopes in the LMM. But we have added the random effects structure in the methods (see page 34).

      Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a sold design and the analyses are appropriate for the research questions.

      Strength:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      We thank you for acknowledging our work on design and brain-behavior analysis. We have addressed all your concerns raised below point by point.

      Weakness:

      (1a) The distinction between Phase 2 and Phase 1&3 behavioral results, specifically the opposite effect of MC/MI in congruent trials raises some concerns with regard to the effectiveness of the ISPC manipulation. Why do RTs and error rates under MC congruent condition in Phase 2 seem to be worse than MI congruent?

      Thank you for raising these issues. In Phase 1, one color set (red and blue) was assigned to the MC condition, whereas another color set (yellow and green) was assigned to the MI condition. In Phase 2, these assignments were flipped, and they were flipped back again in Phase 3. Thus, the MC condition consisted of red and blue in Phases 1 and 3 but yellow and green in Phase 2, whereas the MI condition consisted of yellow and green in Phases 1 and 3 but red and blue in Phase 2 (Fig. 1b in the manuscript). This manipulation leads to seemingly opposite patterns between Phases 1 & 3 and Phase 2.

      However, when considering specific colors, the pattern is consistent across phases. In Phase 2, RTs and error rates for yellow and green (MC congruent) were worse than those for red and blue (MI congruent), which mirrors the pattern observed in Phases 1 and 3, where RTs and error rates for yellow and green (MI congruent) were worse than those for red and blue (MC congruent)

      We interpreted the results in Phase 2 as reflecting a typical ISPC effect, which is defined as a smaller conflict effect in the MI condition (MI incongruent – MI congruent) compared with the MC condition (MC incongruent – MC congruent). To our knowledge, the ISPC paradigm does not impose a specific prediction regarding the relative difference between MC-congruent and MI-congruent conditions.

      (1b) Could there be other factors at play here, e.g. order effect?

      We agree that order effect could play a role, such that memory from Phase 1 may influence the pattern in phase 2. For example, in phase 1, yellow and green were assigned to the MI condition, and participants therefore have associated these colors with a high control state (SC) and incongruent responses (SR). These prior associations could interfere with the newly learned mappings in Phase 2, where yellow and green were reassigned to the MC congruent condition (i.e., low control state and congruent responses). As a result, memory from Phase 1 may have weakened the expected MC in phase 2. A similar effect could also apply to the MI condition. Consequently, the same condition does not show parallel performance between phase 1 and phase 2, which may lead to different patterns in the difference between MC congruent and MI congruent conditions in phase 2.

      (1c) How does this potentially affect the neural analyses where trials from different phases were combined?

      Thank you for the question. As we mentioned above, the order effect could slow down the newly learned associations. However, we still found the ISPC effect in each phase, suggesting that all kinds of both SC and SR associations were formed and could be applied to the decoding and the following analyses cross phases. Relatedly, there might be confounded with temporal structured noise (TSN) when the neural analyses on decoding were combined the trials from different phases. However, we have performed the control decoding analyses and distance effect tests and confirmed that our decoding results were not driven by TSN (Please see comment #1 of R1).

      (1d) the manuscript does not mention whether there is counterbalancing for the color groups across participants, so far as I can tell.

      Thank you for the reminder. We have balanced the color groups by randomly dividing the participants into two groups and assigning different color sets to each group. The related interpretations have been included in task overview of the revised manuscript (page 6), which reads:

      “The color groups were counterbalanced across participants by red and blue as the color set of MC in one group while as the color set of MI in another group in the phase 1.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I commend the authors for addressing and clarifying my previous questions. One new comment regarding the newly added Figure 9: the response-locked behavioral correlation is much weaker compared to the stimulus-locked one, even never reaching the significance level. I think this difference should be discussed instead of simply glossing over it.

      Thank you for your suggestion. We discussed the difference in the discussion of revised manuscript (page 25), which reads:

      “Note that we found the negative prediction of the strength of SC and SR to RTs did not reach statistical significance with response-locked analysis as stimulus-locked analysis. It is possible that SC and SR representations have occurred before the stage of response processing, which is usually aligned with stimulus onset (Jiang et al., 2020a; Kang & Yu-Chin, 2024; Khan et al., 2025)”

      References

      Braver, T. S. (2012). The variable nature of cognitive control: a dual mechanisms framework. Trends Cogn Sci, 16(2), 106-113. doi:10.1016/j.tics.2011.12.010

      Cellier, D., Petersen, I. T., & Hwang, K. (2022). Dynamics of Hierarchical Task Representations. J Neurosci, 42(38), 7276-7284. doi:10.1523/JNEUROSCI.0233-22.2022

      Collins, A. G., & Frank, M. J. (2018). Within- and across-trial dynamics of human EEG reveal cooperative interplay between reinforcement learning and working memory. Proceedings of the National Academy of Sciences, 115(10), 2502-2507. doi:10.1073/pnas.1720963115

      Khan, A. U., Hoy, C. W., Anderson, K. L., Piai, V., King-Stephens, D., Laxer, K. D., . . . Bentley, J. N. (2025). Neural dynamics of proactive and reactive cognitive control in medial and lateral prefrontal cortex. iScience, 28(9), 113375. doi:10.1016/j.isci.2025.113375

      Kikumoto, A., & Mayr, U. (2020). Conjunctive representations that integrate stimuli, responses, and rules are critical for action selection. Proc Natl Acad Sci 117(19), 10603-10608. doi:10.1073/pnas.1922166117

      Kikumoto, A., Mayr, U., & Badre, D. (2022a). The role of conjunctive representations in prioritizing and selecting planned actions. Elife, 11. doi:10.7554/eLife.80153

      Kikumoto, A., Sameshima, T., & Mayr, U. (2022b). The Role of Conjunctive Representations in Stopping Actions. Psychol Sci, 33(2), 325-338. doi:10.1177/09567976211034505

      Rangel, B. O., Hazeltine, E., & Wessel, J. R. (2023). Lingering Neural Representations of Past Task Features Adversely Affect Future Behavior. J Neurosci, 43(2), 282-292. doi:10.1523/JNEUROSCI.0464-22.2022

    1. eLife Assessment

      This study offers a large-scale resource that maps the transcriptional landscape following combinatorial transcription factor overexpression. The authors propose a modular framework for understanding gene regulatory networks and identify transcription factor cocktails for cellular reprogramming. While the dataset and proposed principles are valuable, the strength of the evidence is incomplete, in part due to the complexity of the computational analyses and limited orthogonal experimental validation. The study will be of interest to the fields of gene regulation and developmental biology.

    2. Reviewer #1 (Public review):

      Summary:

      Duan, Li, Kulkarni et al. apply a multiplexed single-cell overexpression screen (Reprogram-Seq) to combinatorially perturb 105 transcription factors across 7 target cell types in mouse embryonic fibroblasts, generating a resource of ~200,000 single-cell transcriptomes spanning over 1,300 TF combinations. They develop a framework for classifying pairwise TF-TF interactions, identify a modular, shared architecture of gene regulatory programs across diverse TF combinations, and use these tools to nominate and partially validate new reprogramming cocktails.

      Strengths:

      The scale of the combinatorial screen is substantial, and the resulting dataset is a genuine resource for the field. The TF-TF interaction typing framework is a useful conceptual extension of prior genetic-interaction approaches to an overexpression/reprogramming context, and the modularity finding that diverse TF combinations converge on shared gene programs is a compelling organizing principle. The authors are, for the most part, careful and appropriately hedged in their claims; the overclaiming we flag below is the exception, not the rule. We also note that the core Reprogram-Seq assay itself builds directly on the authors' own prior work; the novelty here rests on scale, the interaction framework, and the modularity analysis.

      Weaknesses:

      Most of the concerns raised below relate to how existing data are quantified, cited, and reconciled with the text, rather than to the underlying experiments themselves. Several quantitative and comparative claims in the Results are not fully supported by the figures cited, and some conclusions are in tension with the authors' own data. Key methodological details relevant to interpreting the central TF-TF interaction framework, including TF expression dosage and within-combination transcriptional variability, are not reported or controlled for, which limits confidence in the resulting interaction classifications. The relationship between TF number and reprogramming efficiency is not clearly distinguished from a simple combinatorial coverage effect and does not consistently generalize across batches. Experimental validation of predicted cocktails is limited to a single target cell type. The manuscript would also benefit from addressing whether TF overexpression in fibroblasts can fully capture a factor's endogenous regulatory network, given that pioneer activity and chromatin accessibility are not addressed.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents a large-scale combinatorial transcription factor overexpression screen in mouse embryonic fibroblasts coupled with single-cell RNA-seq to systematically map the relationship between TF combinations, gene regulatory networks, and transcriptional reprogramming. Using approximately 100 transcription factors, the authors identify TF combinations that shift cells toward diverse transcriptional states, organize TF combinations into "perturbation clusters" with shared transcriptional outputs, infer modular gene regulatory networks, model pairwise TF interactions, and use pseudotime analyses to nominate candidate reprogramming TF combinations. The resulting dataset represents a potentially valuable resource for studying combinatorial TF activity and transcriptional reprogramming.

      Strengths:

      The primary strength of the study is its experimental scale and the breadth of the generated dataset. The Reprogram-Seq platform enables systematic interrogation of thousands of TF combinations that would be difficult to test individually, and the authors develop several computational frameworks to organize these data and generate biological hypotheses. The epicardial reprogramming analyses, including independent qPCR and protein localization experiments, provide proof-of-principle that the platform can recover biologically relevant TF combinations.

      Weaknesses:

      Many of the manuscript's central conclusions rely on a complex computational pipeline that is not sufficiently justified or independently validated. Identification of transcriptionally reprogrammed cells depends on co-embedding with reference atlases, yet the robustness of this analysis and the interpretation of cells occupying primary-cell clusters are not explored in depth. Similarly, the conclusions regarding modular gene regulatory networks depend critically on the perturbation clusters defined by MDE embedding and HDBSCAN clustering. These perturbation clusters form the basis for nearly all downstream analyses, including differential expression, gene module identification, gene specificity, and TF modularity, yet little evidence is provided that the clusters are robust to alternative embedding strategies, clustering parameters, or resampling approaches.

      The manuscript also provides relatively limited orthogonal validation of its computational predictions. Although the epicardial analyses are validated experimentally, comparable validation is not performed for most other predicted cell fates, TF interaction classes, or perturbation modules. Consequently, many conclusions regarding the generality of modular TF activity, TF cooperativity, and the predicted reprogramming cocktails remain supported primarily by computational inference.

      In addition, several aspects of the analytical workflow-including the quality filtering of TF combinations, interpretation of unclustered perturbations, selection of genes for downstream visualization, and robustness of pseudotime analyses across lineages-would benefit from greater methodological transparency.

      Overall, this work provides a valuable dataset and introduces analytical approaches that will likely be useful to the community. However, in its current form, I believe the strongest biological conclusions are insufficiently validated. Additional analyses demonstrating the robustness of the computational framework, together with broader orthogonal validation of representative predictions, would substantially strengthen confidence in the proposed principles governing combinatorial transcription factor activity and transcriptional reprogramming.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Duan et al perform a combinatorial TF overexpression screen combined with single-cell RNA-seq (Reprogram-Seq) to extract general principles of how combinatorial TF interactions drive distinct gene regulatory networks in reprogrammed cells. Using a library of 105 TFs, they induce different cell fates, many of which resemble in vivo cell identities. They observe that combinations of TFs have better reprogramming results than inductions driven by a single TF. By looking at gene expression enrichment/depletion in different reprogrammed cell clusters, they infer functional GRNs induced by specific TF combinations and identify GRNs specific for certain cell types. They also identify TFs that could improve known TF cocktails for the induction of certain cell fates. They observe that TFs with cooperative interactions regarding the regulation of gene expression may lead to better reprogramming results, and finally, they build a bottom-up approach that utilizes the single-cell transcriptomes to predict TFs driving certain reprogrammed fates.

      Strengths:

      Reprogram-Seq is not new, but the strength of the study lies in the fact that the authors assess the induction outcome from a large number of different TF combinations. The authors are thus able to make broad observations, such as the modularity of GRNs and TF cooperativity, as well as propose new TFs and TF interactions to be tested for the induction of certain cell fates. The manuscript is well written, and the conclusions are, in general, supported well by the authors' analyses and data.

      Weaknesses:

      The study would benefit from some further analysis and discussion to better tighten the conclusions:

      While both expression enrichment and depletion were used to define perturbation clusters, the authors then focused on analyzing functional gene groups only for the enriched genes. Are there any functional relations between the repressed genes within a perturbation cluster? Do the authors observe the same modularity (in terms of regulation by TFs) for repressed genes as they do for induced genes?

      Can the authors give a description of how neomorphic TF interactions work? How would gene expression be affected in single vs double perturbation in those cases?

      What does it mean functionally when a TF pair shows more than one type of interaction (as shown in Supplementary Table 6), and how do such interactions correlate with successful transcriptional reprogramming?

      In the last Results section, the authors are able to use the transcriptome to predict the TF that was used for the induction. Could the authors discuss some plausible applications of this TF prediction method? For example, could they use it on in vivo single-cell RNA-seq of a certain cell type to predict candidate TFs for the induction of that cell type?

      It would help the reader if, at the end of each Results section, the authors add a concluding paragraph, highlighting the most important conclusions and findings (same as they have done in the section titled "Combinatorial TF over-expression reprograms MEFs to diverse states").

    1. eLife Assessment

      This useful paper describes a software tool, "DrosoMating", which allows automated, high-throughput quantification of 6 common metrics of courtship and mating behaviors in Drosophila melanogaster. The validity of the tool is convincingly shown to perform as well as expert human assessments, across a wide range of conditions. While it is not suitable yet for posed based classification of individual actions, and its utility with respect to new deep learning-based tools such as DeepLabCut, SLEAP have yet to be assessed, the efficient coarse-grained detection of behavioral states in low resolution videos offered by DrosoMating represents a very helpful addition to analytical tools available for assessing courtship behaviour in Drosophila.

    2. Reviewer #1 (Public review):

      Summary:

      The study of Drosophila mating behaviors has offered a powerful entry point for understanding how complex innate behaviors are instantiated in the brain. The effectiveness of this behavioral model stems from how readily quantifiable many components of the courtship ritual are, facilitating the fine-scale correlations between the behaviors and the circuits that underpin their implementation. Detailed quantification, however, can be both time consuming and error prone, particularly when scored manually. Song et al. have sought to address this challenge by developing DrosoMating, software that facilitates the automated and high-throughput quantification of 6 common metrics of courtship and mating behaviors. Compared to a human observer, DrosoMating matches courtship scoring with high fidelity. Further, the authors demonstrate that the software effectively detects previously described variations in courtship resulting from genetic background or social conditioning. Finally, they validate its utility in assaying the consequences of neural manipulations by silencing Kenyon cells involved in memory formation in the context of courtship conditioning.

      Strengths:

      (1) The authors demonstrate that for three key courtship/mating metrics, DrosoMating performs virtually indistinguishably from a human observer, with differences consistently within 10 seconds and no statistically significant differences detected. This demonstrates the software's usefulness as a tool for reducing bias and scoring time for analyses involving these metrics.

      (2) The authors validate the tool across multiple genetic backgrounds and experimental manipulations to confirm its ability to detect known influences on male mating behavior.

      (3) The authors present a simple, modular chamber design that is integrated with DrosoMating and allows for high throughput experimentation, capable of simultaneously analyzing up to 144 fly pairs across all chambers.

      Weaknesses:

      (1) DrosoMating appears to be an effective tool for the quantification of key courtship and mating metrics, but similar tools for automated analysis already exist. The authors present a compelling use case for DrosoMating, where it has particular advantages over tools like FlyTracker and Ctrax for high-throughput analysis. This comparative analysis, however, leaves out modern pose-estimation methods (SLEAP, DeepLabCut), better able to tolerate low contrast and occlusion. It therefore remains unclear what specific advantages it might offer over current machine learning approaches.

      (2) The courtship behaviors of Drosophila males represent a series of complex behaviors that unfold dynamically in response to female signals. While metrics like courtship latency, courtship index, and mating duration are useful, they compress the complexity of actions that occur throughout the mating ritual. The authors suggest DrosoMating's modular architecture facilitates integration with behavioral classifiers like JAABA. Such integration could substantially expand the utility of this tool for the broader Drosophila neuroscience community, but in its current form its applications are confined to summary timing metrics.

      (3) Validation is limited to multiple D. melanogaster strains. Cross-species studies of mating behavior diversity are increasingly common, so demonstrating the tool's accuracy across species would strengthen claims about its broader applicability.

    3. Reviewer #2 (Public review):

      This manuscript introduces DrosoMating, an integrated hardware-software pipeline designed to automate quantification of Drosophila courtship and mating behavior. The authors aim to provide a low-cost, scalable alternative to existing behavioral tracking systems, focusing on extracting key temporal metrics including courtship index, copulation latency, and mating duration from high-throughput video recordings.

      A major strength of the work is the clear emphasis on experimental scalability and practical usability. The system is designed for multi-chamber recording and performs robustly under low-quality imaging conditions where conventional pose-tracking pipelines often fail. The revised manuscript substantially improves its scientific positioning through the inclusion of systematic benchmarking against widely used tools (Ctrax and FlyTracker), demonstrating that both fail to complete end-to-end analysis under these recording conditions: Ctrax due to segmentation instability and trajectory fragmentation, FlyTracker due to runtime errors during feature computation. A comparative table (Table 1) summarizes key features across tools, and the authors appropriately qualify that these limitations are specific to the low-quality video conditions tested and should not be interpreted as general shortcomings of those tools. The addition of individual-level behavioral ethograms (Figure S3) further strengthens the evidence by allowing direct assessment of the system's temporal resolution at the single-fly level.

      The authors also appropriately address statistical concerns raised in review. The re-analysis using ANOVA frameworks - one-way ANOVA with Tukey's HSD for multi-strain comparisons, two-way ANOVA with Sidak's correction for genotype × training interactions - improves the rigor of the behavioral comparisons and supports the revised conclusions regarding strain and learning effects. The addition of locomotor control analyses (Figure S4) further clarifies interpretation of mutant phenotypes by partially disentangling motor from courtship-specific effects, with the revised text appropriately acknowledging contributions from both general hypoactivity and sensory impairments.

      A key limitation remains the conceptual scope of the system. DrosoMating is optimized for state-based temporal segmentation rather than fine-grained behavioral annotation or posture-level decomposition. While the authors now clearly acknowledge this and position the tool appropriately, it inherently restricts its use cases compared to modern pose-estimation and classifier-based frameworks. Additionally, the benchmarking comparison is necessarily asymmetric: DrosoMating is tested on low-quality videos where it excels by design, while Ctrax and FlyTracker are evaluated under conditions outside their intended operating range. A comparison under more favorable conditions for the tracking-based tools, or an evaluation of whether modest improvements in video quality would bring conventional pipelines within functional range, would further contextualize the practical boundary between approaches. The comparison also does not include modern deep-learning-based tools (e.g., DeepLabCut, SLEAP), which may be more robust to low-contrast conditions than classical segmentation-based pipelines. These points do not diminish the demonstrated utility of DrosoMating for its intended niche but would help users make more informed decisions about tool selection.

      Overall, the revised manuscript presents a well-validated and clearly positioned contribution. It defines the niche in which DrosoMating provides substantial practical value while appropriately delimiting its limitations relative to more general behavioral analysis frameworks.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study of Drosophila mating behaviors has offered a powerful entry point for understanding how complex innate behaviors are instantiated in the brain. The effectiveness of this behavioral model stems from how readily quantifiable many components of the courtship ritual are, facilitating the fine-scale correlations between the behaviors and the circuits that underpin their implementation. Detailed quantification, however, can be both time-consuming and error-prone, particularly when scored manually. Song et al. have sought to address this challenge by developing DrosoMating, software that facilitates the automated and high-throughput quantification of 6 common metrics of courtship and mating behaviors. Compared to a human observer, DrosoMating matches courtship scoring with high fidelity. Further, the authors demonstrate that the software effectively detects previously described variations in courtship resulting from genetic background or social conditioning. Finally, they validate its utility in assaying the consequences of neural manipulations by silencing Kenyon cells involved in memory formation in the context of courtship conditioning.

      Strengths:

      (1) The authors demonstrate that for three key courtship/mating metrics, DrosoMating performs virtually indistinguishably from a human observer, with differences consistently within 10 seconds and no statistically significant differences detected. This demonstrates the software's usefulness as a tool for reducing bias and scoring time for analyses involving these metrics.

      (2) The authors validate the tool across multiple genetic backgrounds and experimental manipulations to confirm its ability to detect known influences on male mating behavior.

      (3) The authors present a simple, modular chamber design that is integrated with DrosoMating and allows for high-throughput experimentation, capable of simultaneously analyzing up to 144 fly pairs across all chambers.

      Weaknesses:

      (1) DrosoMating appears to be an effective tool for the high-throughput quantification of key courtship and mating metrics, but a number of similar tools for automated analysis already exist. FlyTracker (CalTech), for instance, is a widely used software that offers a similar machine vision approach to quantifying a variety of courtship metrics. It would be valuable to understand how DrosoMating compares to such approaches and what specific advantages it might offer in terms of accuracy, ease of use, and sensitivity to experimental conditions.

      (2) The courtship behaviors of Drosophila males represent a series of complex behaviors that unfold dynamically in response to female signals (Coen et al., 2014; Ning et al., 2022; Roemschied et al., 2023). While metrics like courtship latency, courtship index, and copulation duration are useful summary statistics, they compress the complexity of actions that occur throughout the mating ritual. The manuscript would be strengthened by a discussion of the potential for DrosoMating to capture more of the moment-to-moment behaviors that constitute courtship. Even without modifying the software, it would be useful to see how the data can be used in combination with machine learning classifiers like JAABA to better segment the behavioral composition of courtship and mating across genotypes and experimental manipulations. Such integration could substantially expand the utility of this tool for the broader Drosophila neuroscience community.

      (3) While testing the software's capacity to function across strains is useful, it does not address the "universality" of this method. Cross-species studies of mating behavior diversity are becoming increasingly common, and it would be beneficial to know if this tool can maintain its accuracy in Drosophila species with a greater range of morphological and behavioral variation. Demonstrating the software's performance across species would strengthen claims about its broader applicability.

      Reviewer #2 (Public review):

      This paper introduces "DrosoMating," an integrated hardware and software solution for automating the analysis of male Drosophila courtship. The authors aim to provide a low-cost, accessible alternative to expensive ethological rigs by utilizing a custom acrylic chamber and smartphone-based recording. The system focuses on quantifying key temporal metrics-Courtship Index (CI), Copulation Latency (CL), and Mating Duration (MD)-and is applied to behavioral paradigms involving memory mutants (orb2, rut).

      The development of open-source behavioral tools is a significant contribution to neuroethology, and the authors successfully demonstrate a system that simplifies the setup for large-scale screens. A major strength of the work is the specific focus on automating Copulation Latency and Mating Duration, metrics that are often labor-intensive to score manually.

      However, there are several limitations in the current analysis and validation that affect the strength of the conclusions:

      First, the statistical rigor requires substantial improvement. The analysis of multi-group experiments (e.g., comparing four distinct strains or factorial designs with genotype and training) currently relies on multiple independent Student's t-tests. This approach is statistically invalid for these experimental designs as it inflates the family-wise Type I error rate. To support the claims of strain-specific differences or learning deficits, the data must be analyzed using Analysis of Variance (ANOVA) to properly account for multiple comparisons and to explicitly test for interaction effects between genotype and training conditions.

      Second, the biological validation using $w^{1118}$ and $y^1$ mutants entails a potential confound. The authors attribute the low Courtship Index in these strains to courtship-specific deficits. However, both strains are known to exhibit general locomotor sluggishness (due to visual or pigmentation/behavioral defects). Since "following" behavior is likely a component of the Courtship Index, a reduction in this metric could reflect a general motor deficit rather than a specific lack of reproductive motivation. Without controlling for general locomotion, the interpretation of these behavioral phenotypes remains ambiguous.

      Third, the benchmarking of the system is currently limited to comparisons against manual scoring. Given that the field has largely adopted sophisticated open-source tracking tools (e.g., Ctrax, FlyTracker, JAABA), the utility of DrosoMating would be better contextualized by comparing its performance - in terms of accuracy, speed, or identity maintenance - against these existing automated standards, rather than solely against human observation.

      Finally, the visual presentation of the data hinders the assessment of the system's temporal precision. While the system is designed to capture time-resolved metrics, the results are presented primarily as aggregate bar plots. The absence of behavioral ethograms or raster plots makes it difficult to verify the software's ability to accurately detect specific transitions, such as the exact onset of copulation.

      We sincerely thank the reviewers for their constructive and detailed feedback, which has substantially improved the clarity and rigor of our work. Below is a summary of the major revisions.

      (1) Comparison with existing tools. We added a new main figure (Figure 5) and Table 1 systematically benchmarking DrosoMating against Ctrax and FlyTracker on identical low-quality, high-throughput mating videos. Both conventional tools failed under our recording conditions: Ctrax showed severe segmentation instability and fragmented trajectories, while FlyTracker frequently crashed during feature computation. These results demonstrate that DrosoMating's state-detection approach bypasses the pose-tracking limitations that impair established pipelines. We also revised the Discussion to clarify that DrosoMating is specialized for robust extraction of mating timing metrics from low-quality videos, not a replacement for general-purpose tracking tools.

      (2) Statistical analysis. Following Reviewer #2's recommendation, we re-analyzed Figure 3 using one-way ANOVA with Tukey's HSD for strain comparisons, and Figure 4 using two-way ANOVA with Sidak's post-hoc test for learning assays, including the critical Genotype by Training interaction term.

      (3) New supplementary data. We added Figure S4 showing basal locomotor velocity of single-housed males across all four strains to decouple motor defects from courtship deficits in w1118 and y1 mutants. We also added Figure S3 with behavioral ethograms for individual flies to visually demonstrate the system's temporal resolution and detection accuracy.

      (4) Additional improvements. We standardized statistical reporting across all figure legends, explicitly defined the segmentation threshold parameter s in the Methods, consistently used "Mating Duration (MD)" throughout, standardized MB247-GAL4 labeling, clarified that occluded frames are retained for mating-state detection via merged-contour analysis, and corrected minor errors including duplicate references and unnecessary quotation marks.

      We believe these revisions substantially strengthen the manuscript and clearly position DrosoMating within the existing ecosystem of behavioral analysis tools.

      We remain grateful for the valuable feedback from both reviewers and the editorial team, and we hope the revised version meets the standards for publication in eLife.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It's difficult to assess the utility of this tool in relation to the variety of alternative methods available for the automated scoring of courtship behavior, some of which offer more granular behavioral data than what DrosoMating has been presented to produce. A direct comparison to other approaches (e.g., FlyTracker, DeepLabCut, SLEAP) would be highly valuable for understanding what use cases DrosoMating is best suited for. Ideally, this analysis would include (i) a comparison of the accuracy of scoring courtship metrics and constituent behaviors, (ii) an assessment of the sensitivity of the various software to different experimental conditions, such as lighting and alternative chamber designs, and (iii) testing across Drosophila species, which could help highlight the particular strengths of DrosoMating over more established methods. At a minimum, a table comparing key features, requirements, and capabilities would help position DrosoMating within the existing ecosystem of tools.

      We sincerely appreciate the reviewer’s insightful comments on benchmarking DrosoMating against existing tools for automated courtship behavior analysis. We fully agree that direct comparisons are essential to clarify the unique advantages and ideal use cases of DrosoMating. We have extensively revised the manuscript by adding comparative experiments, a new figure(Figure 5) and comparative table (Table1), and expanded descriptions, as detailed below.

      (1) Direct comparison with conventional tracking tools

      We systematically tested Ctrax and FlyTracker on the same low-quality, high-throughput mating videos used for DrosoMating. The corresponding results are presented in Figure 5.

      Ctrax showed severe segmentation instability, including over-segmentation, under-segmentation, and fragmented trajectories, and failed to stably detect two flies.

      FlyTracker was able to generate background models and showed partially improved segmentation after threshold adjustment, but frequently crashed during feature computation and could not complete the end-to-end analysis pipeline.

      These results confirm that tracking-based tools are vulnerable to low contrast, chamber artifacts, and prolonged male–female overlap, whereas DrosoMating bypasses pose tracking and directly detects mating states.

      The revised manuscript text is as follows (line 265-342):

      "DrosoMating is more compatible with low-quality mating videos than conventional tracking-based pipelines.

      To evaluate whether conventional fly-tracking pipelines could be used as alternative tools for extracting mating-duration metrics, we tested Ctrax and FlyTracker on representative low-quality single-chamber mating videos recorded under our standard high-throughput conditions (Fig. 5). Ctrax is designed to estimate the position and orientation of multiple walking flies while maintaining individual identities over time (Branson et al., 2009) (https://ctrax.sourceforge.net/), whereas FlyTracker aims to track fly pose, including position, orientation, body size, wing and leg positions, and to generate trajectory- and feature-based outputs for downstream behavior analysis (Eyjolfsdottir et al., 2014). Because these programs were not readily compatible with our full high-throughput behavioral recording setup, we first cropped the original videos and tested single-chamber videos containing one male and one female fly.

      Using Ctrax, we observed that target detection was highly variable even within single-chamber videos. In representative frames, Ctrax could occasionally identify the two flies correctly (Fig. 5A). However, imperfect segmentation was frequently observed. In some frames, parts of the fly body were detected as additional targets, resulting in over-segmentation and an apparent increase in the number of detected flies (Fig. 5B, C). Conversely, when the male and female were close to each other or physically overlapped, the two animals were sometimes detected as a single target (Fig. 5D). These examples indicate that Ctrax detection was sensitive to the low contrast and overlapping fly bodies present in our mating videos. We then adjusted the Ctrax detection threshold to improve segmentation quality. Although threshold optimization improved detection in some frames, abnormal detections remained evident across randomly sampled frames (Fig. 5E). For example, some frames still showed incorrect target numbers, and a severe segmentation failure was observed in the lower-right example, where the detected objects did not correspond to two clearly separable flies. Thus, even after parameter optimization, Ctrax did not consistently maintain the expected two-target detection state in single-chamber mating videos.

      The instability of Ctrax detection was also reflected in the trajectory output. In the first 500 frames, the generated trajectories were fragmented into multiple colored track segments rather than two continuous trajectories corresponding to the male and female (Fig. 5F). In addition, some trajectories extended outside the chamber boundary, indicating tracking errors and identity instability. Consistently, frame-by-frame quantification of detected target number showed that the detected object count did not remain stable at the expected value of two flies per chamber (Fig. 5G). Because Ctrax failed to maintain stable two-fly detection and continuous trajectories under these video conditions, its output could not be reliably used for downstream extraction of copulation latency or mating duration.

      We next tested FlyTracker on the same type of cropped single-chamber mating videos. FlyTracker was developed to track multiple flies by estimating body position, orientation, size, wing and leg positions, and by maintaining fly identities across video frames; it also outputs per-frame features such as velocity, facing angle, and wing-angle-related measurements for downstream behavioral analysis (Eyjolfsdottir et al., 2014). In our videos, FlyTracker was able to generate a background model, indicating that the program could recognize the overall imaging field and chamber background (Fig. 5H). However, during calibration and segmentation, the default threshold setting produced inconsistent detection results (Fig. 5I). Although some frames were segmented relatively well under the default threshold (Fig. 5J), these successful examples were not representative of the overall tracking process, and the program frequently terminated with runtime errors during tracking or downstream feature computation. To improve detection stability, we lowered the segmentation threshold. Under this adjusted setting, FlyTracker produced more complete fly masks in representative frames (Fig. 5K), and the diagnostic output showed that two flies were detected in many sampled frames (Fig. 5L). Nevertheless, detection remained unstable in some sampled frames, including frames in which zero or one fly was detected despite the expected two flies per chamber (Fig. 5L). The full pipeline ultimately failed during feature computation, producing a runtime error before complete tracking and feature outputs could be generated (Fig. 5M). Therefore, even after threshold adjustment, FlyTracker could not provide a stable end-to-end workflow for extracting mating-duration metrics from these low-quality mating videos.

      This failure mode is relevant because FlyTracker depends on stable segmentation, identity maintenance, and per-frame feature extraction. In our assay videos, low contrast, chamber-edge artifacts, and prolonged male–female overlap during copulation interfered with these requirements. As a result, FlyTracker could occasionally identify the flies in individual frames, but it did not reliably complete the full analysis pipeline required for downstream behavioral quantification. This limitation is especially important for workflows such as JAABA, which use manually labeled examples to train behavior classifiers but still depend on upstream tracking-derived features. Thus, for our low-quality high-throughput mating recordings, FlyTracker-based analysis was substantially less practical than DrosoMating, which directly outputs mating-related timing metrics without requiring continuous high-fidelity two-fly pose tracking."

      (2) Evaluation of accuracy, robustness, and cross-species potential

      We addressed the three key points requested by the reviewer:

      (i) Accuracy: DrosoMating achieved 98–99% agreement with manual scoring. Under our experimental conditions, neither Ctrax nor FlyTracker completed an end-to-end workflow capable of reliably extracting copulation latency and mating duration. Ctrax produced fragmented trajectories with unstable target counts, and FlyTracker terminated with runtime errors during feature computation.

      (ii) Robustness: DrosoMating is highly robust to low-contrast lighting and common behavioral chamber setups, whereas conventional tools require high-quality videos and fail under fly occlusion.

      (iii) Cross-species testing: We did not perform cross-species validation in this revision. DrosoMating relies on mating state detection rather than species-specific morphology, suggesting potential transferability, although this remains to be experimentally validated. We have therefore revised the text to avoid claiming demonstrated cross-species performance and now describe this as a potential future application.

      The revised manuscript text is as follows (line 572-576):

      "Cross-species testing was not performed in this study. Although DrosoMating's state-detection approach is morphology-agnostic and therefore theoretically applicable across Drosophila species, we have revised the text to avoid claiming demonstrated cross-species performance. Formal validation across diverse species remains a promising future direction."

      We added Table1 to compare key features across DrosoMating, Ctrax, and FlyTracker, including low-quality video compatibility, multi-chamber support, robustness to fly overlap, direct output of mating timing metrics, accuracy.

      (3) Clarification of ideal use cases

      We revised the Discussion to emphasize that DrosoMating is not intended to replace general-purpose pose or tracking tools (e.g., FlyTracker, Ctrax) that provide fine-grained behavioral data.

      Instead, it offers a specialized, robust, and high-throughput workflow for scenarios requiring efficient extraction of mating timing metrics from low-quality videos, where conventional pipelines often fail.

      These revisions greatly improve the clarity and positioning of DrosoMating. We thank the reviewer for this valuable suggestion.

      The revised manuscript text is as follows (line 578-620):

      “Comparative Evaluation with Existing Courtship Analysis Tools

      A major advantage of DrosoMating is its compatibility with low-quality, high-throughput mating videos. Conventional tracking-based tools such as Ctrax and FlyTracker are powerful for trajectory- and pose-based behavioral analysis, but they generally require stable object segmentation, identity maintenance, and reliable feature extraction across frames. Ctrax was designed to estimate the positions and orientations of multiple walking flies while maintaining their identities, whereas FlyTracker aims to track detailed fly pose and generate per-frame behavioral features. Under our recording conditions, these requirements were difficult to satisfy because the videos were low contrast and the male and female frequently overlapped during copulation.

      In our tests, both tools showed limited compatibility with these videos. Even after cropping to single-chamber videos and adjusting detection parameters, Ctrax produced unstable target numbers, fragmented trajectories, and tracking errors. FlyTracker could generate a background model and occasionally segment flies successfully, but detection remained unstable and the full pipeline failed during feature computation. These issues are particularly relevant for copulation latency and mating duration analysis: during copulation, the male and female remain physically coupled for a long period, which makes identity-based tracking difficult. For these timing metrics, it is more important to robustly detect the onset and offset of the mating state than to reconstruct detailed individual trajectories. DrosoMating was purpose-built to address these specific challenges. It operates reliably on lower-quality video streams, requires no complex pre-processing or manual ROI definition, and is optimized for high-throughput multi-chamber analysis. While it does not offer the same level of pose or kinematic detail as other tools, it provides a unique solution for laboratories seeking a simple, fast, and robust pipeline to quantify core reproductive timing metrics—copulation latency (CL), courtship index (CI), and mating duration (MD)—without the overhead of more complex systems (Table 1).

      Although we directly tested only Ctrax and FlyTracker, this limitation may also affect workflows that depend on upstream tracking-derived features. For example, JAABA uses tracking-derived features to train behavior classifiers, and DANCE, a recent Drosophila aggression and courtship pipeline, uses JAABA-based classifiers and lists FlyTracker and JAABA as required software. Thus, our conclusion is not that these tools are generally unsuitable for Drosophila behavior analysis, but that DrosoMating provides a more practical workflow for low-quality, high-throughput videos focused specifically on mating timing.

      While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays.”

      We have also added this clarification to the Table 1 legend to avoid overgeneralizing the limitations of Ctrax and FlyTracker (line 786-790):

      "DrosoMating was designed to extract mating-related timing metrics from high-throughput videos without requiring continuous two-fly identity tracking. Ctrax and FlyTracker were tested on cropped single-chamber videos from the same recording setup. Their limitations described here refer specifically to these low-quality mating videos and should not be interpreted as general limitations of the tools."

      (2) The coarse nature of the metrics captured by DrosoMating may limit the usefulness of this tool for many researchers. Consider integrating DrosoMating with one or more behavioral classifiers (e.g., JAABA) and validating its performance to increase its utility across a wider range of uses. Demonstrating the feasibility of this would substantially increase the tool's appeal to those who need access to more discrete behavioral elements.

      We thank the reviewer for this constructive suggestion. We agree that DrosoMating is currently optimized for rapid, high-throughput quantification of core mating timing metrics (CL, CI, and MD) rather than discrete behavioral classification (e.g., wing extension, circling, or aggressive postures). This reflects a deliberate design trade-off: by prioritizing robust state detection over detailed pose tracking, DrosoMating achieves reliable performance on low-quality videos where conventional identity-based pipelines fail.

      We appreciate the reviewer’s vision for expanding the tool’s utility. While DrosoMating does not presently generate the per-frame kinematic features (position, orientation, wing angles, etc.) required as input for JAABA classifiers, its modular architecture and underlying video-processing framework provide a foundation for future integration. In the revised Discussion, we have clarified that extending DrosoMating to export trajectory-derived features compatible with downstream classifiers such as JAABA represents a promising future direction—one that would broaden its applicability to laboratories requiring granular behavioral elements, without necessitating a complete overhaul of the existing pipeline.

      We hope this clarification addresses the reviewer’s concern and accurately reflects both the current capabilities and future potential of the tool.

      The revised manuscript text is as follows (line 614-620):

      "While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays."

      (3) Please clarify how DrosoMating handles fly identification during mating? I would think mounted flies would be occluded, and if these frames are excluded from analysis, it would be expected to skew mating duration scores. It would be helpful if the authors could discuss these details.

      Occluded frames are excluded from CI calculation because individual courtship actions cannot be reliably assigned during prolonged overlap. However, these frames are not discarded from mating-duration analysis. Instead, prolonged merged contours are used as evidence for copulation-state detection, and the MD timer continues until physical separation:

      (1) When two flies are separate, the system detects two contours; when they overlap during mounting, it detects one merged contour. Our code processes both cases, so occluded frames are retained, not skipped.

      (2) To distinguish a single fly from two overlapping flies, we analyze the shape of the merged contour. Two overlapping flies produce a characteristically different aspect ratio (elongated shape) compared to one fly. This geometric cue is fed into the state classifier to label the frame as "mating."

      (3) Because these frames are classified as mating rather than excluded, the mating duration timer runs continuously through the occlusion period. Thus, mating duration is not artificially shortened.

      The high agreement between DrosoMating and manual scoring for mating duration (within 10 s, no significant difference) confirms that this approach does not introduce bias.

      We revised the manuscript as follow (line 191-193):

      "Occluded frames are excluded from CI calculation but retained for mating-state detection via merged-contour analysis, ensuring continuous MD measurement through copulation."

      (4) Please indicate the statistical tests used in Figures 3, 4, and S1. What methods were used to address multiple hypothesis testing?

      We thank the reviewer for this important comment. For comparisons between two groups, two-sided Student’s t-tests were used. For comparisons among three or more groups, one-way ANOVA followed by post-hoc tests were applied. For two-factor experimental designs, two-way ANOVA was used. We have clearly stated these statistical tests in the Statistical Analysis section and have added this information to the figure legends for Figures 3, 4, and S1 in the revised manuscript.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (5) Line 43: "High resolution video tracking enables...".

      We thank the reviewer for noting this imprecise expression. We have revised the statement to accurately describe that our method uses image-based video analysis to identify courtship and copulation events and quantify their temporal parameters. The description has been corrected in the revised manuscript line 44-46:

      "Our image-based video analysis enables precise identification of courtship and copulation events, as well as quantification of their timing and duration under controlled conditions. "

      (6) Line 51: remove quotation mark.

      We thank the reviewer for the careful correction. The unnecessary quotation mark at Line 51 has been removed in the revised manuscript.

      (7) Line 107: Chen et al, 2024 is listed twice in the references.

      We thank the reviewer for the careful correction. The duplicate reference of Chen et al., 2024 has been removed from the reference list in the revised manuscript.

      (8) Line 242: remove quotation mark.

      We thank the reviewer for the careful correction. The unnecessary quotation mark at Line 242 has been removed in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Please re-analyze the data in Figures 3 and 4 using ANOVA followed by appropriate post-hoc tests (e.g., Tukey's HSD). Specifically, use a One-way ANOVA for strain comparisons in Figure 3 and a Two-way ANOVA for the learning assays in Figure 4. The interaction term (Genotype $\times$ Training) is critical for demonstrating specific learning deficits. Update the "Statistical Analysis" section and figure legends accordingly.

      We appreciate the reviewer’s recommendation. We have re-analyzed the data in Figures 3 and 4 using the suggested ANOVA approaches: one-way ANOVA with Tukey’s HSD for strain comparisons (Figure 3), and two-way ANOVA (including the Genotype × Training interaction term) with Sidak's post-hoc test for learning assays (Figure 4). The updated statistical methods are now described in the Statistical Analysis section and figure legends.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (2) To decouple motor defects from courtship deficits in $w^{1118}$ and $y^1$ mutants, please use your tracking data to calculate and present a "General Locomotion" metric (e.g., average velocity or total distance traveled in the absence of a female).

      We thank the reviewer for this suggestion. We have now added Figure S4 showing basal locomotor velocity of single-housed males across all four strains. As expected, w^1118 and y^1 mutants move more slowly than wild-type controls.

      These data reveal that both motor and sensory factors contribute to the observed courtship phenotypes. Reduced basal locomotion likely limits the males' ability to approach and follow females. However, this generalized hypoactivity is compounded by strain-specific sensory deficits: w<sup>1118</sup> males suffer visual impairment that compromises female detection (Krstic et al., 2013), while y<sup>1</sup> males display altered cuticular hydrocarbons that disrupt pheromonal communication (Drapeau et al., 2006). These sensory defects impair courtship initiation and female recognition independent of locomotor capacity. Thus, the reduced CI, CL, and MD in these mutants reflect the combined effects of slower movement and courtship-specific sensory impairments, rather than motor defects alone.

      We have revised the manuscript to incorporate the velocity data and clarify this interpretation (line 219-226):

      "Notably, reduced basal locomotor activity in w<sup>1118</sup> and y<sup>1</sup> mutants has been well documented in previous studies, independent of courtship behavior (Drapeau et al., 2006; Krstic et al., 2013). Consistent with these reports, our tracking data show that single-housed w<sup>1118</sup> and y<sup>1</sup> males exhibit lower average velocity than Canton-S and Oregon-R controls (Fig. S4A). These general locomotor differences are insufficient to fully explain the observed courtship and mating timing phenotypes, indicating that additional courtship‑related processes contribute to the observed behavioral differences."

      (3) Please expand the discussion or provide a small comparative dataset contrasting DrosoMating with established tools like JAABA. Explain the specific advantages of your pipeline (e.g., cost, simplicity, focus on CL/MD) to justify its adoption over these alternatives.

      We sincerely appreciate the reviewer’s insightful comments on benchmarking DrosoMating against existing tools for automated courtship behavior analysis. We fully agree that direct comparisons are essential to clarify the unique advantages and ideal use cases of DrosoMating. We have extensively revised the manuscript by adding comparative experiments, a new figure (Figure 5) and comparative table (Table 1), and expanded descriptions, as detailed below.

      (1) Direct comparison with conventional tracking tools

      We systematically tested Ctrax and FlyTracker on the same low-quality, high-throughput mating videos used for DrosoMating. The corresponding results are presented in Figure 5.

      Ctrax showed severe segmentation instability, including over-segmentation, under-segmentation, and fragmented trajectories, and failed to stably detect two flies.

      FlyTracker was able to generate background models and showed partially improved segmentation after threshold adjustment, but frequently crashed during feature computation and could not complete the end-to-end analysis pipeline.

      These results confirm that tracking-based tools are vulnerable to low contrast, chamber artifacts, and prolonged male–female overlap, whereas DrosoMating bypasses pose tracking and directly detects mating states.

      The revised manuscript text is as follows (line 265-342):

      " DrosoMating is more compatible with low-quality mating videos than conventional tracking-based pipelines

      To evaluate whether conventional fly-tracking pipelines could be used as alternative tools for extracting mating-duration metrics, we tested Ctrax and FlyTracker on representative low-quality single-chamber mating videos recorded under our standard high-throughput conditions (Fig. 5). Ctrax is designed to estimate the position and orientation of multiple walking flies while maintaining individual identities over time (Branson et al., 2009) (https://ctrax.sourceforge.net/), whereas FlyTracker aims to track fly pose, including position, orientation, body size, wing and leg positions, and to generate trajectory- and feature-based outputs for downstream behavior analysis (Eyjolfsdottir et al., 2014). Because these programs were not readily compatible with our full high-throughput behavioral recording setup, we first cropped the original videos and tested single-chamber videos containing one male and one female fly.

      Using Ctrax, we observed that target detection was highly variable even within single-chamber videos. In representative frames, Ctrax could occasionally identify the two flies correctly (Fig. 5A). However, imperfect segmentation was frequently observed. In some frames, parts of the fly body were detected as additional targets, resulting in over-segmentation and an apparent increase in the number of detected flies (Fig. 5B, C). Conversely, when the male and female were close to each other or physically overlapped, the two animals were sometimes detected as a single target (Fig. 5D). These examples indicate that Ctrax detection was sensitive to the low contrast and overlapping fly bodies present in our mating videos. We then adjusted the Ctrax detection threshold to improve segmentation quality. Although threshold optimization improved detection in some frames, abnormal detections remained evident across randomly sampled frames (Fig. 5E). For example, some frames still showed incorrect target numbers, and a severe segmentation failure was observed in the lower-right example, where the detected objects did not correspond to two clearly separable flies. Thus, even after parameter optimization, Ctrax did not consistently maintain the expected two-target detection state in single-chamber mating videos.

      The instability of Ctrax detection was also reflected in the trajectory output. In the first 500 frames, the generated trajectories were fragmented into multiple colored track segments rather than two continuous trajectories corresponding to the male and female (Fig. 5F). In addition, some trajectories extended outside the chamber boundary, indicating tracking errors and identity instability. Consistently, frame-by-frame quantification of detected target number showed that the detected object count did not remain stable at the expected value of two flies per chamber (Fig. 5G). Because Ctrax failed to maintain stable two-fly detection and continuous trajectories under these video conditions, its output could not be reliably used for downstream extraction of copulation latency or mating duration.

      We next tested FlyTracker on the same type of cropped single-chamber mating videos. FlyTracker was developed to track multiple flies by estimating body position, orientation, size, wing and leg positions, and by maintaining fly identities across video frames; it also outputs per-frame features such as velocity, facing angle, and wing-angle-related measurements for downstream behavioral analysis (Eyjolfsdottir et al., 2014). In our videos, FlyTracker was able to generate a background model, indicating that the program could recognize the overall imaging field and chamber background (Fig. 5H). However, during calibration and segmentation, the default threshold setting produced inconsistent detection results (Fig. 5I). Although some frames were segmented relatively well under the default threshold (Fig. 5J), these successful examples were not representative of the overall tracking process, and the program frequently terminated with runtime errors during tracking or downstream feature computation. To improve detection stability, we lowered the segmentation threshold. Under this adjusted setting, FlyTracker produced more complete fly masks in representative frames (Fig. 5K), and the diagnostic output showed that two flies were detected in many sampled frames (Fig. 5L). Nevertheless, detection remained unstable in some sampled frames, including frames in which zero or one fly was detected despite the expected two flies per chamber (Fig. 5L). The full pipeline ultimately failed during feature computation, producing a runtime error before complete tracking and feature outputs could be generated (Fig. 5M). Therefore, even after threshold adjustment, FlyTracker could not provide a stable end-to-end workflow for extracting mating-duration metrics from these low-quality mating videos.

      This failure mode is relevant because FlyTracker depends on stable segmentation, identity maintenance, and per-frame feature extraction. In our assay videos, low contrast, chamber-edge artifacts, and prolonged male–female overlap during copulation interfered with these requirements. As a result, FlyTracker could occasionally identify the flies in individual frames, but it did not reliably complete the full analysis pipeline required for downstream behavioral quantification. This limitation is especially important for workflows such as JAABA, which use manually labeled examples to train behavior classifiers but still depend on upstream tracking-derived features. Thus, for our low-quality high-throughput mating recordings, FlyTracker-based analysis was substantially less practical than DrosoMating, which directly outputs mating-related timing metrics without requiring continuous high-fidelity two-fly pose tracking."

      (2) Clarification of ideal use cases

      We revised the Discussion to emphasize that DrosoMating is not intended to replace general-purpose pose or tracking tools (e.g., FlyTracker, Ctrax) that provide fine-grained behavioral data.

      Instead, it offers a specialized, robust, and high-throughput workflow for scenarios requiring efficient extraction of mating timing metrics from low-quality videos, where conventional pipelines often fail.

      These revisions greatly improve the clarity and positioning of DrosoMating. We thank the reviewer for this valuable suggestion.

      The revised manuscript text is as follows (line 578-620):

      “Comparative Evaluation with Existing Courtship Analysis Tools

      A major advantage of DrosoMating is its compatibility with low-quality, high-throughput mating videos. Conventional tracking-based tools such as Ctrax and FlyTracker are powerful for trajectory- and pose-based behavioral analysis, but they generally require stable object segmentation, identity maintenance, and reliable feature extraction across frames. Ctrax was designed to estimate the positions and orientations of multiple walking flies while maintaining their identities, whereas FlyTracker aims to track detailed fly pose and generate per-frame behavioural features. Under our recording conditions, these requirements were difficult to satisfy because the videos were low contrast and the male and female frequently overlapped during copulation.

      In our tests, both tools showed limited compatibility with these videos. Even after cropping to single-chamber videos and adjusting detection parameters, Ctrax produced unstable target numbers, fragmented trajectories, and tracking errors. FlyTracker could generate a background model and occasionally segment flies successfully, but detection remained unstable and the full pipeline failed during feature computation. These issues are particularly relevant for copulation latency and mating duration analysis: during copulation, the male and female remain physically coupled for a long period, which makes identity-based tracking difficult. For these timing metrics, it is more important to robustly detect the onset and offset of the mating state than to reconstruct detailed individual trajectories. DrosoMating was purpose-built to address these specific challenges. It operates reliably on lower-quality video streams, requires no complex pre-processing or manual ROI definition, and is optimized for high-throughput multi-chamber analysis. While it does not offer the same level of pose or kinematic detail as other tools, it provides a unique solution for laboratories seeking a simple, fast, and robust pipeline to quantify core reproductive timing metrics—copulation latency (CL), courtship index (CI), and mating duration (MD)—without the overhead of more complex systems (Table 1).

      Although we directly tested only Ctrax and FlyTracker, this limitation may also affect workflows that depend on upstream tracking-derived features. For example, JAABA uses tracking-derived features to train behavior classifiers, and DANCE, a recent Drosophila aggression and courtship pipeline, uses JAABA-based classifiers and lists FlyTracker and JAABA as required software. Thus, our conclusion is not that these tools are generally unsuitable for Drosophila behavior analysis, but that DrosoMating provides a more practical workflow for low-quality, high-throughput videos focused specifically on mating timing.

      While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays.”

      We have added this clarification to the Table 1 legend to avoid overgeneralizing the limitations of Ctrax and FlyTracker (line 786-790):

      "DrosoMating was designed to extract mating-related timing metrics from high-throughput videos without requiring continuous two-fly identity tracking. Ctrax and FlyTracker were tested on cropped single-chamber videos from the same recording setup. Their limitations described here refer specifically to these low-quality mating videos and should not be interpreted as general limitations of the tools."

      (4) Complement the aggregate bar plots with behavioral ethograms or raster plots for representative individual flies. Color-code these plots for specific states (resting, following, courting, copulating) to visually demonstrate the system's temporal resolution and detection accuracy.

      We thank the reviewer for this valuable suggestion. To address this comment, we have added a new supplementary figure (Figure S3) that presents behavioral ethograms for individual flies, directly complementing the aggregate bar plots in the main text. The figure displays the full temporal progression of courtship and copulation behaviors for all wells that exhibited successful mating. As requested, behaviors are color-coded (orange: courting, red: copulating) to clearly delineate different states. These ethograms visually demonstrate the system’s ability to resolve behavioral transitions with high temporal precision, confirming the accuracy of our automated detection of courtship initiation, duration, and copulation events. This addition provides critical individual-level validation that supports the aggregate statistical results presented in the main text.

      (5) Standardize the use of estimation statistics (DBMs). If used in Figure 3, they should also be applied to Figure 4, with appropriate statistical comparisons between groups.

      We thank the reviewer for this suggestion. We have standardized the statistical reporting format across all figure legends, with each legend explicitly stating the test used, the post hoc method (where applicable), and significance thresholds.

      Our approach is as follows:

      Fig. 2 and Fig.3D-I (two-group comparison, manual vs. automated scoring): DBM + Student's t-test.

      Fig. 3A-C and Fig. S1B-C, E-F (multi-group comparison, 4 strains or 2 rearing conditions): One-way ANOVA + Tukey's HSD.

      Fig. 4B-D and F (two-factor design, genotype × training): Two-way ANOVA + Sidak's post hoc.

      We have also added a sentence to the Methods clarifying that estimation statistics (DBM) are used for single two-group contrasts, while ANOVA-based approaches are used for multi-group or multi-factor designs. The statistical methods are now consistently documented in both the Methods section and the corresponding figure legends. We hope this clarification addresses the reviewer's concern.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (6) Fix inconsistent labeling (e.g., "MB-247" vs. "MB247") and redundant axis labels.

      Thank you for pointing this out. We have now standardized all labels to "MB247-GAL4" throughout the text and figures.

      We have also removed redundant axis labels from multi-panel figures. And redundant DBM and metric labels have been streamlined in Figures 2 and 4. All statistical tests and metric definitions are now fully described in the corresponding figure legends rather than being repeated on each sub-panel.

      (7) Significantly increase the size of data panels to make individual data points and error bars legible.

      Thank you for this suggestion. We have standardized the figure formatting and simplified the panels by removing redundant axis labels and consolidating descriptive details into the figure legends. We believe the current panel sizes, combined with these clarifications, provide sufficient legibility for both data points and error bars in the final high-resolution PDF.

      Minor Corrections:

      (1) Abstract: Rephrase "lack of certain timing-related behavioral repertoires" (Line 108) to "lack of precise quantification for temporal parameters of post-copulatory behavior."

      Thank you for the suggested rephrasing. We have updated the sentence accordingly.

      (2) Define the physical parameter "s" (e.g., is it a pixel threshold?) to ensure reproducibility.

      Thank you for this important suggestion. We have now explicitly defined the parameter s in the Methods section. Briefly, s is the grayscale intensity threshold (0–255, 8-bit) used for binary segmentation of flies from the background. It is automatically calculated as the maximum grayscale value of three user-selected reference flies plus an offset of 28, and can be manually adjusted to accommodate varying illumination conditions.

      The revised manuscript text is as follows (line 499-503):

      “Note on the s value: The parameter s represents the grayscale intensity threshold (range: 0–255 for 8-bit images) used for binary segmentation of flies from the background. It is automatically calculated as the maximum grayscale value at the three reference fly positions plus an offset of 28, and can be manually adjusted to accommodate varying illumination conditions.”

      (3) Figure 1 Legend: Clarify the "clockwise selection" of pillars and their relation to well numbering.

      Thank you for this suggestion. We have revised the Figure 1 legend to clarify that:

      (1) The four pillars are selected in clockwise order starting from the top-left corner to define the chamber corners for perspective transformation (homography), which corrects for camera angle and standardizes the field of view.

      (2) Well numbering is independent of pillar selection order. After automated perspective correction, wells are numbered sequentially in a left-to-right, top-to-bottom order (1–36) based on the standardized chamber layout.

      The updated Figure 1 legend now reads:

      "Columns (pillars) should be selected in clockwise order starting from the top-left corner to define the chamber boundaries for perspective transformation (Fig. 1A, lower). Well numbering (1–36) follows a left-to-right, top-to-bottom sequence after perspective correction and is independent of pillar selection order."

      (4) Add the citation for Eastwood and Burnet (1977) regarding Courtship Latency.

      Thank you for pointing this out. We have added the citation Eastwood and Burnet (1977) to the Introduction where Courtship Latency is first defined

      (5) Select one term ("Mating Duration" or "Copulation Duration") and use it consistently throughout the text and figures.

      Thank you for this suggestion. We have now standardized the terminology throughout the manuscript and figures. "Mating Duration" (MD) is used consistently in all instances where "Copulation Duration" previously appeared. The abbreviation MD has been retained for consistency with existing figure labels

      References

      Branson, K., Robie, A.A., Bender, J., Perona, P., & Dickinson, M.H. (2009). High-throughput ethomics in large groups of Drosophila. Nature Methods, 6(6), 451–458.

      Claridge-Chang, A., & Assam, P.N. (2016). Estimation statistics should replace significance testing. Nature Methods, 13(2), 108–109.

      Drapeau, M.D., Cyran, S.A., Viering, M.M., Geyer, P.K., & Long, A.D. (2006). A cis-regulatory sequence within the yellow locus of Drosophila melanogaster required for normal male mating success. Genetics, 172(2), 1009–1030.

      Eastwood, L., & Burnet, B. (1977). Courtship latency in male Drosophila melanogaster. Behavior Genetics, 7(3), 359–372.

      Eyjolfsdottir, E., Branson, S., Burgos-Artizzu, X.P., Hoopfer, E.D., Schor, J., Anderson, D.J., & Perona, P. (2014). Detecting social actions of fruit flies. Lecture Notes in Computer Science, 8692, 772–787.

      Gil-Martí, B., Barredo, C.G., Pina-Flores, S., Poza-Rodriguez, A., Treves, G., Rodriguez-Navas, C., Camacho, L., Pérez-Serna, A., Jimenez, I., Brazales, L., Fernandez, J., & Martin, F.A. (2023). A simplified courtship conditioning protocol to test learning and memory in Drosophila. STAR Protocols, 4(1), 101572.

      Greenspan, R.J., & Ferveur, J.F. (2000). Courtship in Drosophila. Annual Review of Genetics, 34, 205–232.

      Hall, J.C. (1994). The mating of a fly. Science, 264(5163), 1702–1714.

      Kabra, M., Robie, A.A., Rivera-Alba, M., Branson, S., & Branson, K. (2013). JAABA: interactive machine learning for automatic annotation of animal behavior. Nature Methods, 10(1), 64–67.

      Kitamoto, T. (2001). Conditional modification of behavior in Drosophila by targeted expression of a temperature-sensitive shibire allele in defined neurons. Journal of Neurobiology, 47(2), 81–92.

      Krstic, D., Boll, W., & Noll, M. (2013). Influence of the White locus on the courtship behavior of Drosophila males. PLoS ONE, 8(9), e77904.

      Levin, L.R., Han, P.L., Hwang, P.M., Feinstein, P.G., Davis, R.L., & Reed, R.R. (1992). The Drosophila learning and memory gene rutabaga encodes a Ca2+/calmodulin-responsive adenylyl cyclase. Cell, 68(3), 479–489.

      Pavlou, H.J., & Goodwin, S.F. (2013). Courtship behavior in Drosophila melanogaster: towards a ‘courtship connectome’. Current Opinion in Neurobiology, 23(1), 76–83.

      Yapici, N., Kim, Y.J., Ribeiro, C., & Dickson, B.J. (2008). A receptor that mediates the post-mating switch in Drosophila reproductive behaviour. Nature, 451(7176), 33–37.

    1. eLife Assessment

      This study provides an important advance in understanding how disordered proteins interact with cell membranes by identifying the sequence rules that enable aromatic residues to penetrate deeply into the membrane interior. The integration of complementary computational approaches, including molecular simulations, large-scale sequence analysis, and the development of an online prediction server, makes the work potentially impactful for the membrane protein and intrinsically disordered protein communities. The evidence supporting the main conclusions is generally convincing.

    2. Reviewer #1 (Public review):

      Summary:

      This work investigates the membrane insertion of aromatic-centered sequences in IDPs. Using a combination of all-atom MD simulations, the PPM method, and development of the sequence-based predictor AroMIP, the authors aim to establish a quantitative membrane insertion role for aromatic-centered motifs. The study demonstrates that flanking aliphatic and basic residues promote membrane insertion, whereas acidic and polar residues suppress insertion, and further reveals a difference between F/W-centered motifs and Y-centered motifs. The resulting AroMIP model achieves high predictive accuracy on human IDPs and is implemented as a publicly accessible web server.

      Strengths:

      This work addresses an important biological problem, as aromatic-driven membrane insertion remains poorly characterized despite mediating diverse functions like membrane remodeling and signaling. A key strength is the combination of complementary approaches, e.g., MD simulations provide mechanistic insight into insertion pathways, while PPM enables exhaustive sequence space exploration. The large-scale analysis clearly establishes L and R as promoters and E, N, and G as suppressors. The work also provides valuable mechanistic insight into how aromatic, aliphatic, and basic residues cooperate to stabilize membrane insertion states. Another important strength is the development of AroMIP as a practical prediction tool with a user-friendly online server that appears computationally efficient and broadly accessible to the community. The work is also well connected to prior experimental and computational literature, and the authors carefully position their findings within existing knowledge of membrane-associated IDPs.

      Comments on revised version:

      I think the authors have addressed all my concerns. I do not have further comments or requests for additional revisions. Thank you for all the hard work!

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides a valuable advance in understanding how disordered proteins interact with cell membranes by identifying the sequence rules that enable aromatic residues to penetrate deeply into the membrane interior. The integration of complementary computational approaches, including molecular simulations, large-scale sequence analysis, and the development of an online prediction server, makes the work potentially impactful for the membrane protein and intrinsically disordered protein communities. The evidence supporting the main conclusions is generally convincing, although its transferability across diverse membrane compositions and its validity as a prediction tool for real protein-membrane systems remain to be further established.

      We thank the editors for recognizing our study as a valuable advance. This work lays a solid foundation for future developments to account for diverse membrane compositions and further refinements after additional experimental tests.

      Public review:

      Reviewer #1 (Public review):

      A primary limitation is the heavy reliance on computational modeling. Training for AroMIP is generated using PPM rather than direct experimental measurements, and so the model may primarily reproduce PPM behavior rather than true membrane insertion thermodynamics. Moreover, all simulations use a single lipid composition (POPC:POPS:<sub>2</sub> 70:25:5), but biological membranes vary substantially in cholesterol, cardiolipin, and acidic lipid content. Whether AroMIP's predictions transfer to diverse lipid environments remains untested. The 5% <sub>2</sub> concentration used in the simulations is higher than that of a normal mammalian cell and may therefore overemphasize electrostatic contributions. Applicability beyond short 9-residue motifs is unclear, as longer-range interactions or secondary structure in full-length IDRs could modulate insertion in ways the current model does not capture. This could be considered for future development.

      The reviewer’s point on our reliance on PPM for training, a single lipid composition, and potential effects beyond a 9-residue motif is well taken. Regarding PPM, we chose it as the optimal compromise for high-throughput data. However, we complemented the high-throughput PPM data with experimental data on an initial set of 10 peptides. Moreover, we validate AroMIP on an additional 12 IDRs (intrinsically disordered regions; Table S2). On membrane composition, we now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph).

      On potential effects beyond a 9-residue motif, we now add justification and note neglected factors for future developments (paragraph running from p. 19-20), as suggested by the reviewer.

      Reviewer #2 (Public review):

      (1) Aromatic residues have been shown to partition preferentially to the headgroup region of the lipid bilayer. Most of the papers on this problem were published in the mid 1990s to early 2000s. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev.

      Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602-608; Killian & von Heijne, TIBS 2000, 25, 429434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, none of these articles is cited.

      We have now citations to the Landolt-Marticorena paper and the von Heijne reviews [refs 25-27]. The Doyle paper is not particularly relevant. As for the Fleming paper, we cited a 2016 JACS paper (original ref 27; now ref 30) that specifically dealt with aromatic residues.

      (2) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth"."

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is incorrect even in a qualitative sense.

      We have now revised these figures.

      Reviewer #3 (Public review):

      (1) Membrane composition and lipid shape characteristics: The authors chose to use a model membrane bilayer of a distinct lipid composition, POPC: POPS: PI4,5P2 (70:25:5 molar ratio), for their all-atom simulations of the various model peptides. While this may be pertinent for some of these peptides, it is not for many, such as sequence 2 derived from Drp1, which preferentially binds target conical lipids such as cardiolipin (CL) and phosphatidic acid (PA). The rationale behind using PI4,5P2, which can induce positive membrane curvature when sequestered, versus CL and PA, which both induce negative membrane curvature, is not explained.

      We now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). In this Discussion paragraph, we also speculate that conical lipids, by promoting membrane defects, may facilitate membrane insertion.

      (2) Parallel vs. perpendicular peptide orientation of sequence 2 in peripheral Drp1-lipid interactions: On page 11, the authors state that their simulation results of sequence 2 derived from Drp1 "contrasts with a transmembrane orientation proposed by Mahajan et al." However, upon review, a transmembrane orientation for this region has never been proposed anywhere. Drp1 is a peripheral membrane protein that reversibly binds CL- and PA-containing membranes via its intrinsically disordered variable domain containing an aromatic-centered WRG motif. Indeed, the model presented in Figure 9 of Mahajan et al. displays a peripheral and parallel orientation of the transiently helical WRG-containing motif rather than a transmembrane (i.e., across the bilayer) orientation. While the authors can distinguish between a parallel vs. perpendicular orientation of this sequence relative to the plane of the membrane bilayer surface from their simulations, suggesting that previous studies indicated a transmembrane orientation for Drp1 is disingenuous and misleading. The term "transmembrane" should be removed or replaced, as it presents a wrong image.

      We have now deleted the sentence mentioning “transmembrane orientation”.

      (2) Mutational analysis of W vs. F in membrane insertion of W-centered insertion motifs and vice versa: The PPM-based workflow suggests that F-centered sequences have the highest membrane insertion properties as opposed to W-centered ones. A W552F mutation in the WRGML sequence of Drp1 was, however, found to impair function. How do the authors rationalize this? A crossmutational analysis of W vs. F in W-centered motifs and F-centered motifs is warranted. 

      AroMIP predicts a membrane insertion propensity of 0.782 for the WRGML sequence and a moderately higher propensity, 0.837, with a W552F mutation. This increase contradicts the experimental observation of a 3.6-fold increase in membrane binding affinity by Mahajan et al. We now speculate that the specific lipid, cardiolipin, as the reason for the discrepancy (p. 19, 3rd paragraph). This discrepancy provides a concrete example for the need to account for membrane composition in future developments.

      Recommendations for the authors:

      Reviewing Editor Comments:

      (1) The membrane composition used in this study is highly specific. The manuscript would benefit from a clearer justification of the lipid composition and discussing the transferability of the approaches to other relevant membrane systems.

      We now acknowledge the limitation of our work to a single composition (p. 19, 3rd paragraph). This composition was chosen because it is widely used in both computational and experimental studies (e.g., PMID: 21144818; 21344950; 29845130; 29995324; 34813727; 37406927). However, as we now point out, future developments should account for the effects of lipid composition.

      (2) This work heavily relies on computational modeling. Thus, it remains unclear to what extent the model captures the thermodynamics of membrane insertion, rather than reproducing the behavior of the PPM framework. Further comparisons with available experimental results will make this work more impactful.

      We note that the manuscript already has substantial experimental support. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) Overall, the manuscript is lengthy. Shortening will improve the readability, clarity, and accessibility to a broad audience.

      We have shortened some text, as explained below.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript uses a single membrane composition… The high PIP<sub>2</sub> content (5%) in the simulations may overemphasize electrostatic contributions from basic residues. Please discuss how different membrane compositions (e.g., lower <sub>2</sub>, presence of cholesterol, cardiolipin in mitochondrial membranes) might alter the q parameters and whether AroMIP predictions would change qualitatively or quantitatively.

      We now discuss how different membrane compositions may alter q parameters (p. 19, 3rd paragraph). We believe that these alterations will change our predictions quantitatively but not qualitatively, given that our validation is against experimental results acquired on a variety of membrane compositions.

      (2) The limitations of the current framework should be discussed more explicitly. For example, the applicability of AroMIP beyond isolated 9-residue motifs remains unclear. In full-length IDPs, membrane insertion may be modulated by longer-range sequence interactions, transient secondary structure formation, multivalent interactions, or post-translational modifications.

      We now discuss the limitation of the 9-residue motif, and note these neglected factors for future developments (paragraph running from p. 19-20).

      (3) The AroMIP web server is user-friendly… The utility of the server could be enhanced by enabling visualization of full-length disordered proteins. For example, if users input a >300 residue IDP, the server could output residue-wise or sliding-window membrane insertion propensity profiles.

      We have revised the web server. We now display insertion propensity profiles as a plot and have added a link for users to download.

      (4) The manuscript is quite long and in several sections overly descriptive, e.g., in the first MD Results section and portions of the "Additional test cases" section.

      We have shortened the first subsection in Results and placed the expanded presentation in Supporting Information. However, we have kept the “Additional test cases” subsection, because Reviewer 2 appears to have overlooked the experimental validation of our method in these additional test cases.

      (5) The authors used the Berendsen barostat… known not to reproduce correct volume fluctuations and is generally considered less rigorous for equilibrium simulations, although this does not affect the main conclusions of the work. For future studies, the authors may consider using more modern barostats.

      Thank you for the suggestion! We will definitely be using the more modern barostats in future studies.

      Reviewer #2 (Recommendations for the authors):

      (1) The idea of the article seems very interesting. The problem of membrane association mediated by aromatic residues is definitely worth studying. Aromatic residues, especially Tryptophan (W), but also, albeit to a lesser extent, Phenylalanine (F), and Tyrosine (Y) are well known to partition preferentially to the headgroup region of the lipid bilayer. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev. Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602608; Killian & von Heijne, TIBS 2000, 25, 429-434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, however, none of these articles is cited. Have the authors read them?

      We now cite the relevant references as explained above.

      (2) The authors propose to decipher the sequence code for insertion of sequences containing aromatic residues in the membrane employing three types of calculation methods with decreasing order of detail and complexity, but increasing order of efficiency. First, all-atom MD simulations; second, the PPM method (protein positioning in membranes) from Lomize et al (2006), Protein Sci 15, 1318; and third, AroMIP, a mathematical model developed by the authors. Incidentally, I don't see anywhere in the text what AroMIP stands for. Is it Aromatic Membrane Insertion Prediction, or something like that? Please define. In any case, the proposed endeavor is commendable.

      We now spell out the acronym at its first occurrence in the Abstract and Introduction (Aromatic Membrane Insertion Predictor).

      (3) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth"." Thus, there must be a comparison with experiment, as elaborated in the next point.

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM.

      (4) I understand that the authors are computational or theoretical physical chemists and would not expect them to perform the experiments. However, there are plenty of data in the literature that can be used to corroborate the calculations. First and foremost, though, the authors must calculate the binding constants for their set of peptides and then compare them with experiment, or use a different set of peptides for which the experimental results are available, and test how PPM and AroMIP perform on those peptides. There are two possible approaches to calculate the binding constants. First, calculate the Gibbs energy of binding from simulation or calculation, and, from that, calculate the binding constant via the Boltzmann factor. Second, and better, is to calculate the probability of binding in the simulations or calculations from the fraction of time that the peptide spends bound to the membrane or in water. According to the ergodic principle, the ratio of the two times is the binding constant. I understand that most amphipathic peptide sequences whose binding constants to membranes have been determined by experiment are long. But some are not. For example, Mastoparan X is a 14-residue antimicrobial peptide, and its dissociation constant from POPC vesicles is known to be about 300 μM.

      We would like to make it clear that the aim of our study is to predict membrane insertion propensities of aromatic-centred motifs, not the membrane binding affinity of peptides. These two properties are related but require different ways of validating predictions. For membrane insertion, validation requires that not only the motifs are bound to membranes but also the aromatic side chains are placed in the acyl chain region. Our validation of the 12 IDRs listed in Table S2 targeted these requirements.

      That said, we note that free energy of binding and probability of binding calculations, mentioned by the reviewer, have been reported previously, including the free-energy cost of Ala substitutions of aromatic residues located at various depths reported by Waheed et al. (ref 35) and residue-specific insertion depths reported by Wang et al. (ref 13). Both of these studies highlighted the propensities of aromatic side chains in inserting into the acyl chain region.

      (5) Furthermore, the entire Wimley-White interfacial hydrophobicity scale was determined using pentapeptides, measuring the equilibrium binding constant to POPC membranes experimentally and then using those data to build a residue-based Gibbs energy of binding (Wimley & White, Experimentally determined hydrophobicity scale for proteins at membrane interfaces. Nature Struct. Biol. 1996, 3, 842-848; White & Wimley Membrane protein folding and stability: Physical principles. Annu. Rev. Biophys. Biomol. Struct. 1999, 28, 319-365.) This allows for the calculation of the binding affinity for any sequence. The calculations have been compared to experiment and shown to have a high accuracy. A note of caution: Be careful, though, because Wimley and White used a mole fraction concentration scale, which makes the Gibbs energy of binding more favorable than the value calculated using the more common molar concentration scale by -2.4 kcal/mol. Calculating the binding constant for the peptides used by the authors from the WW interfacial scale and comparing the results with the authors' results would be a good place to start.

      This is an excellent suggestion! We now compare our insertion scores with the binding free energies calculated from the WW interfacial scale (new Figure S10; p. 15, 2nd paragraph). Interestingly, we found moderately higher correlations with binding free energies calculated from the octanol scale, which we suggest is more in line with our insertion scores since octanol mimics the hydrophobic region as suggested by White and Wimley in their 1999 Annu Rev paper.

      (6) When we speak of insertion in a bilayer, we normally mean insertion in the nonpolar core, not in the interfacial region, which is what the authors mean in this paper. This is extremely misleading and should be changed, namely in the title, but also throughout the paper.

      Actually, by insertion we precisely mean into the nonpolar core, NOT the interfacial region. Throughout the Introduction, when we used the word “insert”, we added “into the acyl chain region” (p. 3, line 6 from bottom; p. 5, lines 6-7 from top and line 7 from bottom; p. 6, line 4 from top). To avoid any confusion, we now also explicitly add “into the membrane hydrophobic core” when the word “insertion” first occurs in the Abstract and in the opening paragraph of Introduction (replacing the previous “deep insertion”).

      (7) What is q? It appears in equation (1) on page 14, and is referred to several times afterwards, but it is never defined.

      We now elaborate, in the text above and below equation (1), on the meaning of the q parameters: they represent the contributions of flanking residues to the insertion score of a central aromatic residue.

      (8) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is wrong even in a qualitative sense.

      We have revised these figures.

      (9) In Figure 1, the location of the bilayer midplane should be indicated, for example, with a line. Currently, there is a red line on the figures, but that is not the bilayer midplane. A reader may easily - and is likely to - misinterpret the figure.

      As explained in the Figure 1B, C caption, the red line is drawn at Z = -3.1 Å, which is the mean position of glycerol C2 carbon atoms (setting Z = 0 for the phosphate plane) and thus the start of the acyl chain region.

      Reviewer #3 (Recommendations for the authors):

      (1) Please refrain from using the term "transmembrane" when describing Drp1-membrane insertion, as Drp1 is a soluble, peripheral membrane-binding protein that reversibly associates with membranes.

      We have removed the sentence mentioning “transmembrane orientation”.

    1. eLife Assessment

      This is an important study on the role of Slap in restricting Src activity and proliferation of colonic cells in vivo. The authors present solid evidence that an EPHB2-SRC signaling axis stimulates the proliferation of colon precursors and is controlled by the Src-binding protein SLAP, whose loss promotes tumorigenesis.

    2. Reviewer #1 (Public review):

      Naim et al., use genetically engineered mouse models and tissue culture cell lines to investigate the role of the SLAP adaptor protein in colonic epithelium and colon tumour formation. The SLAP adaptor protein is known to be a negative regulator of tyrosine kinase signaling in hematopoietic cells but its role outside the immune system is less well defined. Here the authors use genetically engineered SLAP deficient mice, tissue specific SLAP KO, and colonic organoids to demonstrate that SLAP is expressed in cells of the colonic epithelium where it acts as a cell autonomous regulator of proliferation and differentiation. In addition, they provide biochemical evidence that loss of SLAP expression in cultured colonic organoids results in increased Src family kinase activity and global tyrosine phosphorylation, consistent with its known role as a suppressor of tyrosine kinase activity in immune cells. Consistently, treatment with a SRC kinase inhibitor inhibited growth of SLAP deficient organoids. These data provide solid evidence of a cell autonomous role of SLAP in the colonic epithelium.

      Using a chemically induced model of colitis-associated cancer the authors demonstrate that inactivation of SLAP shows a trend toward increased tumor formation as well as significantly increased Src family kinase activity within tumors. Tumor spheres from SLAP deficient animals showed enhanced growth that was suppressed by treatment with a Src family kinase inhibitor. Of note, the latter effect was specific to SLAP deficient tumor spheres. These observations are convincing and support the authors conclusion that SLAP has a tumor suppressor role in CRC through inhibition of SFK signaling.

      Mechanistically, elevated expression of the EPHB2 receptor tyrosine kinase was detected in immunoblots and by IHC of SLAP KO colonic crypts. In addition, in SLAP deficient crypts, levels of phosphorylated EPHB2 are increased and associated with activated SRC family kinases. Using an EPHB2 inhibitor, the role of EPHB2 in the growth of SLAP deficient colonic organoids, and downstream SRC phosphorylation was demonstrated. The authors also show that low expression of SLAP in human CRC cell line organoids sensitizes to the growth inhibitory effects EPH inhibition which can be reversed by SLAP over expression but not expression of a SH2/SH3 mutant form of SLAP.

      Overall, this work provides evidence of SLAP adaptor function in restricting EPH tyrosine kinase signaling the colonic epithelium and suggests that loss of SLAP expression promotes tumorigenesis in this context.

    3. Reviewer #2 (Public review):

      Summary:

      Protein tyrosine kinases are submitted to diverse regulatory mechanisms controling their activity in normal situation. The authors previously identified SLAP (Src-like adaptor protein), a negative regulator of receptor tyrosine kinase (RTK) signaling, as a key suppressor of the cytoplasmic tyrosine kinase SRC in the normal colon and demonstrated that SLAP is downregulated in a majority of colorectal cancers (CRCs).

      In this study, the authors further explored slap functions in mouse models using constitutive and inducible epithelial-specific Slap deletion (villin-CreERT2 model). They found that loss of slap augments colonic epithelial cell proliferation and that induction of tumorigenesis by the AOM/DSS protocol mimicking CRC leads to more aggressive tumors in the absence of slap. This effect is apparently cell-autonomous as growth of normal and tumoral colonic organoids is SLAP-dependent in in vitro settings. Finally, the authors define that, in colon, SLAP represses EphB2, an RTK lying upstream of SRC, and show that inhibitors of EphB2 can partially limit tumorigenic development in vitro.

      Strengths:

      The manuscript is clearly and concisely written, making it easy to follow. Data obtained in the mouse models are very convincing.

      Weaknesses:

      Direct evidence that EphB2 is activated/phosphorylated in the absence of SLAP is lacking as conclusions are only based on results obtained with inhibitors. Some other issues have to be addressed before acceptance, in particular the relevance of the findings in CRC patients.

      Comments on revised version.

      The authors have satisfactorily addressed my concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Naim et al. use genetically engineered mouse models and tissue culture cell lines to investigate the role of the SLAP adaptor protein in colonic epithelium and colon tumour formation. The SLAP adaptor protein is known to be a negative regulator of tyrosine kinase signaling in hematopoietic cells, but its role outside the immune system is less well defined. Here, the authors use genetically engineered SLAP-deficient mice, tissue-specific SLAP KO, and colonic organoids to demonstrate that SLAP is expressed in cells of the colonic epithelium, where it acts as a cell-autonomous regulator of proliferation and differentiation. In addition, they provide biochemical evidence that loss of SLAP expression in cultured colonic organoids results in increased Src family kinase activity and global tyrosine phosphorylation, consistent with its known role as a suppressor of tyrosine kinase activity in immune cells. Consistently, treatment with an SRC kinase inhibitor inhibited the growth of SLAP-deficient organoids. These data provide solid evidence of a cell-autonomous role of SLAP in the colonic epithelium.

      This work would be improved by further description and interpretation of the SLAP expression pattern shown in the constitutive and tissue-specific KO to further support the conclusions made. In Supplementary Figure 1, magnification of the colon epithelium areas with SLAP expression shown by b-gal and anti-SLAP staining, highlighting regions of interest, would better support the conclusions regarding SLAP expression in specific regions of the colon epithelium. In Supplementary Figure 1B, the authors should indicate that the SLAP staining referred to is epithelial and in resident immune cells, as is mentioned in the text. Also, magnification of the boxed area of LRG5 staining in Figure 1 would improve this figure.

      We thank the reviewer for their positive and constructive evaluation of our work.

      We have revised Fig 1 and S1 to better highlight SLAP expression patterns. Specifically, we have included higher-magnification images of the colonic epithelial regions, with clearly indicated regions of interest (new Figure S1). We have also clarified in the legend that SLAP staining is observed in both epithelial and resident immune cells, as described in the text. Additionally, we have provided a magnified view of the boxed area showing LGR5 staining in Figure 1 to improve clarity.

      Using a chemically induced model of colitis-associated cancer, the authors demonstrate that inactivation of SLAP shows a trend toward increased tumor formation (though this did not reach significance) as well as increased Src family kinase activity within tumors. Tumor spheres from SLAP-deficient animals showed enhanced growth that was suppressed by treatment with a Src family kinase inhibitor. Of note, the latter effect was specific to SLAP-deficient tumor spheres. These observations are convincing and support the authors' conclusion that SLAP has a tumor suppressor role in CRC through inhibition of SFK signaling.

      Mechanistically, elevated expression of the RTK, EphB2, was detected in immunoblots of SLAP KO colonic crypts, while overexpression of SLAP in CRC cell lines downregulated EphB2 protein levels. Using an EPHB2 inhibitor, the role of EPHB2 in the growth of SLAP-deficient colonic organoids was demonstrated. While these data generally support the authors' conclusion that SLAP limits colonic organoid growth by downregulating RTKS such as EphB2 and downstream Src family kinase activity, they do not show which cell types/regions in the colonic epithelium have increased EPHB2 protein and how this relates to SLAP and phospho-SRC expression, as shown in Figure 1 and Figure S1 immunocytochemistry. The expression of EphB2 and its role in colonic tumorsphere growth were not investigated.

      Overall, this work provides evidence of SLAP adaptor function in restricting tyrosine kinase signaling in the colonic epithelium, and suggests that loss of SLAP expression could promote tumorigenesis in this context.

      We thank the reviewer for their positive assessment of our tumour studies and for recognizing the evidence supporting a tumor suppressor role for SLAP through inhibition of SFK signaling.

      To address the reviewer’s mechanistic concerns, we performed additional experiments that are now included in the revised manuscript. We confirmed that loss of Slap is associated with increased EPHB2 expression in colonic crypts by IHC (new Figure 4B) and directly tested the role of EPHB2 in the Slap-deficient phenotype: EPH inhibition reduced both pTyr levels and SRC activation in Slap-deficient organoids (new Figure S2), demonstrating that SFK hyperactivation depends on upstream EPHB2 signaling. Consistent with this mechanism, we also observed increased EPHB2 tyrosine phosphorylation and active SRC (pSRC) association in isolated colonic epithelial cells following Slap deletion (new Figure 4A).

      To extend these findings to the tumour context, we examined the effect of EPH inhibition in human CRC tumoroids (new Figure 5). Pharmacological EPHB2 inhibition reduced tumoroid growth in CRC cells expressing low levels of SLAP, whereas this effect was largely lost upon SLAP overexpression. An SH2- or SH3-inactivating point SLAP mutant failed to suppress tumoroid growth and restored sensitivity to EPHB2 inhibition, further supporting EPHB2 as a critical target of SLAP-mediated tumour suppression. Together, these new data identify EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and strengthen our conclusion that deregulated EPHB2-SRC signaling drives the hyperproliferative phenotype associated with SLAP loss.

      Reviewer #2 (Public review):

      Summary:

      Protein tyrosine kinases are subject to diverse regulatory mechanisms controlling their activity in normal situations. The authors previously identified SLAP (Src-like adaptor protein), a negative regulator of receptor tyrosine kinase (RTK) signaling, as a key suppressor of the cytoplasmic tyrosine kinase SRC in the normal colon and demonstrated that SLAP is downregulated in a majority of colorectal cancers (CRCs).

      In this study, the authors further explored SLAP functions in mouse models using constitutive and inducible epithelial-specific Slap deletion (villin-CreERT2 model). They found that loss of SLAP augments colonic epithelial cell proliferation and that induction of tumorigenesis by the AOM/DSS protocol mimicking CRC leads to more aggressive tumors in the absence of SLAP. This effect is apparently cell-autonomous as growth of normal and tumoral colonic organoids is SLAP-dependent in in vitro settings. Finally, the authors define that, in colon, SLAP represses EphB2, an RTK lying upstream of SRC, and show that inhibitors of EphB2 can partially limit tumorigenic development in vitro.

      Strengths:

      The manuscript is clearly and concisely written, making it easy to follow. The data obtained in the mouse models are very convincing.

      Weaknesses:

      Direct evidence that EphB2 is activated/phosphorylated in the absence of SLAP is lacking, as conclusions are only based on results obtained with inhibitors. Some other issues have to be addressed before acceptance, in particular, the relevance of the findings in CRC patients.

      We thank the reviewer for their positive and constructive evaluation of our work.

      We agree that direct evidence linking SLAP loss to activation of the EPHB2-SRC pathway would strengthen the study. To address this point, we performed additional experiments that are now included in the revised manuscript. In addition to demonstrating increased EPHB2 expression upon Slap deletion, we found that loss of Slap enhances the EPHB2 tyrosine phosphorylation (an index of EPHB2 activity) and association between EPHB2 and pSRC in isolated colonic epithelial cells, supporting increased signaling through this pathway (new Figure 4A). Furthermore, pharmacological inhibition of EPHB2 reduced both SRC activation and the hyperproliferative phenotype observed in Slap-deficient organoids (new Figure S2). Together, these findings provide functional and biochemical evidence that deregulated EPHB2 signaling contributes to SRC activation in the absence of SLAP. We also examined the effect of EPHB2 inhibition in human CRC tumoroids (new Figure 5).

      Pharmacological EPHB2 inhibition reduced tumoroid growth in CRC cells expressing low levels of SLAP, whereas this effect was largely lost upon SLAP overexpression. Together, these new data identify EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and strengthen our conclusion that deregulated EPHB2-SRC signaling drives the hyperproliferative phenotype associated with SLAP loss.

      To address the relevance of our findings in CRC patients, we also extended our analyses to human datasets (new Figure 5D). We observed a significant inverse correlation between SLAP expression and a colorectal cancer stem cell-like activity score in TCGA tumours. In addition, co-expression of SLAP and SLAP2 with EPHB2 was associated with improved disease-free survival in microsatellite-stable (MSS) CRC patients, whereas no such association was observed in microsatellite instability (MSI) tumours. These findings support the clinical relevance of the SLAP-EPHB2 signaling axis and are consistent with a role for SLAP in restraining EPHB2-dependent CSC signaling in CRC.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers have reported that the study of the EPHB2-SLAP-Src axis is not very developed, as most results are derived from using inhibitors. A key question is whether EPHB2 is activated by SLAP depletion and whether it is critical to SRC activation. Addressing these questions would greatly improve the paper.

      We thank the Reviewing Editor for this important suggestion. We would be happy for the editors to assess the revised version without involving the reviewers again. In the revised manuscript, we have substantially strengthened the mechanistic link between SLAP loss, EPHB2 activation, and SRC signaling. We show that Slap deletion increases both EPHB2 expression and tyrosine phosphorylation, enhances EPHB2-pSRC association in colonic epithelial cells. Importantly, pharmacological inhibition of EPHB2 suppresses SRC activation and rescues the hyperproliferative phenotype of Slap-deficient organoids. We further demonstrate that EPHB2 inhibition selectively impairs growth of CRC tumoroids with low SLAP expression, whereas this effect is largely abolished upon SLAP overexpression. An SH2- or SH3-inactivating point SLAP mutant failed to suppress tumoroid growth and restored sensitivity to EPHB2 inhibition, further supporting EPHB2 as a critical target of SLAP-mediated tumour suppression. Together, these new biochemical and functional data establish EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and significantly reinforce the central conclusions of the study.

      Reviewer #1 (Recommendations for the authors):

      (1) Evidence of SLAP expression in the colon is an important basis for these studies and could be moved to the main Figure 1 rather than being in the supplementary material.

      We thank the reviewer for this suggestion. While we agree that documenting SLAP expression in the colon is important, we have retained these data in Figure S1 to maintain a concise main figure set, consistent with the recommended format for Short Reports.

      (2) Define AOM/DSS and briefly describe the model at first mention. In addition, the model in 3A includes TAM treatment at 45 days, but this is not mentioned in the text. Why is this done?

      We have better defined the AOM/DSS protocol at first mention in the revised manuscript and specified the rationale for tamoxifen administration at day 45, which is required to maintain efficient SLAP deletion throughout the duration of the experiment (90 days).

      (3) Evidence that the SRC inhibitor decreased phospho-tyrosine levels in addition to inhibiting the growth of organoids should be included.

      We included data showing the inhibitory effect of the used SRC inhibitor on global phospho-tyrosine levels in organoids in the revised manuscript.

      (4) Further experiments investigating the involvement of EphB2 in colonic tumor formation are of interest and would increase the significance of this work.

      We included data showing that SLAP modulation affects the response of tumoroids derived from cell lines to EphB2 inhibition, providing complementary mechanistic insights.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should confront their findings with data obtained in normal and pathological tissues: are SLAP, SRC, and EphB2 co-expressed, at the single cell level, in normal colon and CRC? In which cell populations? Is loss of SLAP associated with poor prognosis in CRC patients?

      We thank the reviewer for this important suggestion. We agree that assessing the relevance of the SLAP-EPHB2-SRC axis in human CRC is important. However, transcriptomic datasets have inherent limitations in this context, as SLAP primarily regulates signaling at the post-transcriptional level and SRC activity cannot be reliably inferred from mRNA expression.

      To address the clinical relevance of our findings, we performed additional analyses of CRC patient datasets. We found a significant inverse correlation between SLAP expression and a colorectal cancer stem cell-like activity score in TCGA tumours. Furthermore, co-expression of SLAP and SLAP2 with EPHB2 was associated with improved disease-free survival in microsatellite-stable (MSS) CRC patients. These findings are consistent with a role for SLAP in restraining EPHB2-dependent signaling in CRC. Finally, while our data identify EPHB2 as a critical upstream regulator of SRC signaling controlled by SLAP, we do not exclude the possibility that additional receptor tyrosine kinases contribute to the effects of SLAP loss during colorectal tumorigenesis.

      (2) In Figure 4A, total EphB2 levels are increased in the absence of SLAP in colonic crypts. However, the level of EphB2 phosphorylation is not shown. This is an important point to address. Which ligand(s) may activate EphB2?

      We agree that assessing EPHB2 activation is important. To address this point, we have included new data in the revised manuscript showing that Slap deletion increases EPHB2 tyrosine phosphorylation in isolated colonic epithelial cells, providing direct evidence that EPHB2 signaling is enhanced in the absence of SLAP. In addition, we show that loss of Slap increases EPHB2-pSRC association and that pharmacological inhibition of EPHB2 reduces SRC activation and suppresses the hyperproliferative phenotype of Slap-deficient organoids. Together, these findings establish EPHB2 as a critical upstream regulator of SRC signaling following SLAP loss.

      Regarding EPHB2 activation, previous studies have shown that EPHB2 is primarily activated by ephrin-B ligands expressed within the intestinal crypt compartment (Batlle et al., 2002). EPHB2 signaling may also be reinforced through cooperation with other Eph receptors, particularly EPHB3, which is highly expressed in intestinal stem and progenitor cells (Holmberg et al., 2006; Genander et al., 2009). In addition, SRC has been reported to phosphorylate EPH receptors, raising the possibility of bidirectional signaling that could further amplify EPHB2-SRC pathway activity (Leroy et al., 2009; Hochgräfe et al., 2010).

      (3) In the absence of SLAP, inhibitors of EphB2 should also decrease SRC activity as EphB2 lies upstream of SRC (Figure S3C). Does this occur in organoids?

      We now show that EphB2 inhibition reduces SRC activity in SLAP-deficient organoids.

      (4) What is the status of EphB2 and SRC (total, phosphorylated) in SW620 and HT29 CRC cells in the absence of SLAP?

      SW620 and HT29 cells are SLAP-low CRC models. Given that SLAP expression is already minimal in these cells, further depletion is unlikely to provide meaningful additional insight.

      (5) Expression of SLAP is associated with a decrease in the stem cell compartment in CRC cell lines (Figure S2). Is there a stem cell signature associated with low SLAP levels in CRC?

      We analyzed TCGA colorectal cancer datasets using a published colorectal cancer stem cell (CSC) signature. We found a low but significant inverse correlation between SLAP expression and the CSC-like activity score, supporting our experimental observations that SLAP restrains stem cell properties in CRC cells and organoids.

      (6) Does overexpression of the mutant form of SLAP (SLAPmut) limit SLAP effects in SW620 and HT29 CRC cells in Figure S2?

      We have now performed the requested experiments and found that, unlike wild-type SLAP, SLAPmut failed to inhibit tumoroid growth in CRC cells. These results are consistent with our previous findings showing that SLAPmut lacks tumour suppressor activity in CRC cells (Naudin et al., Nat Commun, 2014) and further support the requirement of SLAP signaling functions for the regulation of CRC stem-like properties.

      (7) Total SRC is missing in Figure 2B

      Total SRC levels are now included in the revised figure.

    1. eLife Assessment

      This useful study describes a physical mechanism for the emergence of spiral patterns in the outer epithelial layer of the mammalian cornea independent of pre-patterning or guidance cues, using an agent-based model of self-propelled particles with alignment. The well constructed model show that spiral patterns can emerge from the interaction between the limbus position, cell division, extrusion, and collective cell migration. While the conclusions are solid, some significant questions related to the importance of topological defect remain, and the comparison between the model and data are so far mostly qualitative.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Kostanjevec et al. investigates the mechanism behind spiral pattern formation in the cornea. The authors demonstrate that the spiral motion pattern on the mammalian corneal surface emerges from the interaction between the limbus position, cell division, extrusion, and collective cell migration. Using LacZ mosaic murine corneas, they reveal a tightening spiral flow pattern and show that their cell-based, in silico model accurately reproduces these patterns without global guidance cues. Additionally, they present a continuum model that extends the XYZ hypothesis to describe cell flux on the cornea, offering a quantitative explanation for tissue-scale processes on curved surfaces.

      Strengths:

      The manuscript is well-written, with a systematic approach that clearly explains experimental setups, model construction, assumptions, parameter selection, and predictions. The discussion also provides insightful perspectives on the broader implications of the results for both physics and biology.

      Weaknesses:

      The authors emphasize polar alignment as a key feature of the spiral pattern based on simulation results. However, they do not provide experimental evidence for this polar alignment.

    3. Reviewer #2 (Public review):

      In K. Kostanjevec et al., the authors study a possible mechanism for the formation of spiral patterns in the cornea. First the authors analyze an inferred velocity field, which is deduced from images of fixed corneas, and then determine the position-dependent spiral angle of this velocity fields. Next, the authors analysed two possible markers of cell polarity: the direction of the centrosome-nuclei and the axis of mitosis. Then the authors introduce a stochastic agent-based model of self-propelled particles with over-damped dynamics and with aligning interactions to the orientation of the nearest neighbors and to the particle's velocity. The authors claim to be able to reproduce the equal-time autocorrelation function and the velocity Fourier spectrum. Then the authors introduce the geometry of the cornea by constraining the dynamics on a spherical cap and show that their model can reproduce a typical trajectory in experiments. Finally, the authors produce a phase diagram of the states at a fixed time point as a function of the spherical cap radius and the strength of the coupling aligning constant. Finally, the authors propose an interpretation of the cell fluxes based on the equation of mass conservation.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      The manuscript by Kostanjevec et al. investigates the mechanism behind spiral pattern formation in the cornea. The authors demonstrate that the spiral motion pattern on the mammalian corneal surface emerges from the interaction between the limbus position, cell division, extrusion, and collective cell migration. Using LacZ mosaic murine corneas, they reveal a tightening spiral flow pattern and show that their cell-based, in silico model accurately reproduces these patterns without global guidance cues. Additionally, they present a continuum model that extends the XYZ hypothesis to describe cell flux on the cornea, offering a quantitative explanation for tissue-scale processes on curved surfaces.

      Strengths

      The manuscript is well-written, with a systematic approach that clearly explains experimental setups, model construction, assumptions, parameter selection, and predictions. The discussion also provides insightful perspectives on the broader implications of the results for both physics and biology.

      We thank the reviewer for their positive assessment of the manuscript. We are pleased that the reviewer found the work well written and systematic, and that the experimental design, model construction, assumptions, parameter selection, predictions and broader discussion were clearly presented. We have aimed to preserve these strengths in the revised manuscript while substantially expanding the discussion and analysis in response to the reviewer’s concerns.

      Weaknesses

      The central premise of the manuscript, that the spiral patterning of epithelial corneal cells occurs without guidance cues, is not fully supported. The authors overlook the potential role of axons in guiding epithelial cells, despite clear evidence of spiral axon patterns in their own Fig. 1b. Previous literature indicates that axon patterning precedes epithelial cell patterning, suggesting that epithelial migration might be influenced by pre-existing neural structures (e.g., Leiper et al. 2002, IOVS 2013). The authors need to address this point, possibly by exploring whether axonal patterns serve as a template for epithelial cell migration, or by providing experimental evidence to rule out axon-based guidance.

      The reviewer raises an important point that we now address in the revised manuscript. We respectfully disagree with the assertion that the premise of our work is faulty. At no point do we claim that global guidance cues are absent or ignore the fact that the corneal nerves swirl; rather, our results show that such global or contact-mediated cues are not required to explain the observed swirling patterns of epithelial cell radial migration in the adult cornea.

      Although our work shows that a swirling prepattern of nerves is not required to obtain radial patterns of epithelial cell migration, we now address why nerve swirling occurs and how it affects interpretation of our data. Previous literature indicates that nerve swirling is visible from approximately 3 weeks, before epithelial striping patterns become apparent in transgenic LacZ and GFP reporter mosaics at about 5 weeks. However, both experimental observations and our modelling indicate that spiral epithelial cell migration proceeds for many days before reporter stripe patterns become visible. Thus, epithelial migration may already be underway before epithelial striping is detectable. If axons are following epithelial migration, they would therefore be expected to become radially aligned before the reporter epithelial stripes are evident.

      We also considered the alternative possibility that epithelial cells follow an axonal prepattern. However, several experimental observations argue against the hypothesis that epithelial migration is primarily guided by axonal projections. In situations of genetic mutation or corneal injury, epithelial cell migration can progress independently of, or ahead of, corneal axon extension. In chimeric Pax6+/− LacZ<sup>+</sup> ↔ Pax6<sup>+/+</sup> LacZ<sup>-</sup> mice, where the normally disrupted radial migration of Pax6+/− epithelial cells is restored, the underlying nerves may continue to exhibit abnormal projection patterns. These findings suggest that axonal organization is at least partly dependent on epithelial behaviour rather than the reverse.

      Importantly, we do not exclude the possibility that additional cues, including axonal contact, neurotrophins or other environmental signals, may modulate or refine epithelial migration. We have therefore expanded the revised manuscript to discuss these possibilities and the relevant literature more fully.

      While the model is well-constructed, it currently falls short of its stated goal of elucidating the mechanisms of spiral formation. Key questions remain unanswered:

      Is the curvature of the cornea necessary for spiral formation, or would a simpler disk geometry suffice?

      What role do boundary conditions play?

      How well do the model's predictions quantitatively match experimental data?

      The current comparisons in Fig. 4c-f lack quantitative agreement, and this discrepancy should be discussed with possible explanations.

      We thank the reviewer for identifying these points, which we have now addressed in the revised manuscript.

      First, we have examined the role of geometry more systematically. The spiral pattern also appears in a disk geometry, and the same qualitative migration pattern is observed across a broader range of simulated geometries, including different curvatures, cap angles, a prolate ellipsoid, an oblate ellipsoid and a disk. The precise shape of the spiral depends on geometric features such as curvature and cap-angle opening, but spiral formation is robust across convex cornea-like geometries. These results are now discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and shown in the revised Fig. 10.

      Second, we have clarified the role of boundary conditions, particularly the role of limbal epithelial stem cell proliferation. Without limbal stem cells, and with all other parameters unchanged, the cornea fails to produce the radial striping pattern. At the alignment strength where robust spiral formation normally appears, the simulated tissue flow is disordered and resembles our in vitro calibration simulations. At higher alignment strengths, spiral formation is still not recovered; instead, defects become anchored to the boundary. These findings show that ordered influx from the limbus is important for promoting the spiral state. They are now discussed in the same new section and shown in revised Fig. 9.

      Third, we have revised the quantitative comparison between model and experiment. We agree that the original comparison was limited. We have therefore reanalysed both experimental and simulation data, focusing on the time-averaged hydrodynamic velocity field rather than short-range fluctuations amplified by divisions in the numerical model. We also replaced the Fourier-space velocity correlation functions with spatial velocity correlation functions, which are more directly interpretable. The revised analysis shows that the experiments have mesoscale spatial and temporal correlations, of the order of 5-6 cell sizes in space and about one hour in time, and that these are well captured by the simulations for both plastic and explant substrates. We have also added representative experimental and simulation snapshots in revised Fig. 4g-j.

      For the full cornea, we acknowledge that direct quantitative comparison remains limited by the available experimental data. We can compare with inferred migration direction fields and resurfacing timescales, but we do not yet have direct live measurements of the full corneal velocity field.

      The authors emphasize polar alignment as a key feature of the spiral pattern based on simulation results. However, they do not provide experimental evidence for this polar alignment. The manuscript includes discussions of polar and nematic symmetries that, without supporting data, feel somewhat distracting. If direct experimental evidence for polar alignment is not available, the authors could instead quantify nematic alignment as the spiral forms. This would also allow them to explore potential crosstalk between nematic cell orientation and the polar alignment of self-propulsion, especially considering recent studies showing alternative mechanisms for vortex formation in similar systems.

      We thank the reviewer for pointing out that the discussion of polar and nematic alignment was confusing. We have substantially revised this part of the manuscript.

      We agree that we do not have direct experimental evidence for polar alignment. However, several observations support the interpretation that the system is dominated by substrate-based polar motility with weak polar alignment. We have now quantified nematic alignment of cell orientations and find no evidence of significant elongation or local nematic order in corneal epithelial cells, as shown in the new Fig. 3d.

      Our computational model begins from uncorrelated, substrate-based polar active cell migration. We tested whether polar alignment is necessary and found that the in vitro data are inconsistent with the complete absence of alignment: the flocking order parameter is too high, and the spatial and temporal correlations are larger than expected from persistent driving alone.

      We also discuss that cell-cell stress patterns in the epithelium may remain nematic, but any such effect must be sufficiently weak not to dominate the observed axon motion. We have further revised the discussion to include recent related work showing spiral formation in substrate-based cell migration models with polar dynamics, as well as recent work indicating that nematic-like phenomenology can arise from minimal ingredients such as uncorrelated polar activity and cell deformability. This revision is intended to make the interpretation clearer and to avoid overemphasising unsupported claims.

      Reviewer #2 (Public review):

      In K. Kostanjevec et al, the authors study a possible mechanism for the formation of spiral patterns in the cornea. First the authors analyze an inferred velocity field, which is deduced from images of fixed corneas, and then determine the position-dependent spiral angle of this velocity fields. Next, the authors analysed two possible markers of cell polarity: the direction of the centrosome-nuclei and the axis of mitosis. Then the authors introduce a stochastic agent-based model of self-propelled particles with over-damped dynamics and with aligning interactions to the orientation of the nearest neighbors and to the particle's velocity. The authors claim to be able to reproduce the equal-time autocorrelation function and the velocity Fourier spectrum. Then the authors introduce the geometry of the cornea by constraining the dynamics on a spherical cap and show that their model can reproduce a typical trajectory in experiments. Finally, the authors produce a phase diagram of the states at a fixed time point as a function of the spherical cap radius and the strength of the coupling aligning constant. Finally, the authors propose an interpretation of the cell fluxes based on the equation of mass conservation.

      We thank the reviewer for their careful assessment of the manuscript and for recognising the work as a solid theoretical study. We have revised the manuscript substantially in response to the reviewer’s major concerns, particularly regarding the terminology of topological defects and stagnation points, the comparison with experiments, and the role of corneal geometry.

      Regarding the terminology of topological defects, we have clarified the distinction between a velocity-field stagnation point and the topological classification of the velocity direction field. Stagnation points can be assigned a topological index when one considers the direction of the vector field away from the core, while ignoring the magnitude. We agree that the physical origin of interactions in a velocity field differs from that in a director field, and we have revised the text to avoid confusion. The revised manuscript now includes Box 1, which summarises the relevant topological concepts and caveats.

      Regarding the comparison with experiments, we have expanded and clarified the validation of the inferred velocity field. The LacZ reporter system was designed for lineage tracing and therefore reports coarse-grained cell motion and growth patterns. Direct live imaging of the full cornea remains technically difficult because of the macroscopic size, curved surface and long resurfacing time. However, live fluorescent reporter systems are consistent with our in vivo model and support the interpretation that inferred velocity fields recapitulate epithelial migration in vivo. We have also clarified the interpolation procedure using the XY model: approximately 30% of the corneal surface is covered by directly estimated velocity directions from stripe edges, rising to more than 50% near the central spiral. These measured directions provide sufficient boundary conditions for the annealing procedure to converge to a slowly varying field consistent with the observed stripe geometry.

      We have also revised the comparison between simulation and experiment. For in vitro data, we now compare spatial velocity correlation functions of the time-averaged hydrodynamic velocity field, rather than relying on Fourier-space correlations. The revised comparison shows good agreement between experiments and simulations for both plastic and explant substrates. For the full cornea, we acknowledge that the available quantitative data are limited to inferred migration direction fields and resurfacing timescales.

      Regarding the role of geometry, we have now simulated a wider range of substrate geometries, including spherical caps with different curvatures and cap angles, prolate and oblate ellipsoids, and a flat disk. The spiral migration pattern is robust across these convex geometries, although the detailed spiral shape depends on geometric features. We have also clarified that the boundary condition of inward limbal influx is crucial: without limbal stem cells, the radial striping pattern does not form, and defects may instead anchor to the boundary. These results are now shown in revised Figs. 9 and 10 and discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells.”

      Overall, these additions clarify that the proposed mechanism relies on the combination of polar motility, weak alignment, limbal influx, cell division and extrusion, and cornea-like confinement, while also acknowledging the current experimental limitations.

      Recommendations for the authors:

      Reviewing Editor:

      The authors could strongly improve the manuscript by following the recommendations given below.

      Reviewer #1 (Recommendations for the authors):

      There are, however, substantial shortcomings that authors need to address to make their claims supported by enough evidence:

      Major concerns:

      (1) Neglect of Potential Axon Guidance

      As it stands the premise of the paper is unfortunately faulty. The authors overlook the potential role of axons in guiding epithelial cells, despite clear evidence of spiral axon patterns in their own Fig. 1b. Previous literature indicates that axon patterning precedes epithelial cell patterning, suggesting that epithelial migration might be influenced by preexisting neural structures (e.g., Leiper et al. 2002, IOVS 2013). The authors need to address this point, possibly by exploring whether axonal patterns serve as a template for epithelial cell migration, or by providing experimental evidence to rule out axon-based guidance.

      See for example:

      from: https://iovs.arvojournals.org/article.aspx?articleid=2126612:

      It is therefore not clear why the authors clearly show the axon vortex in Figure 1, then talk about prepatterning - and never mention the axon vortex again.

      The reviewer raises an important point that we now address in the revised manuscript. We respectfully disagree with the assertion that the premise of our work is faulty. At no point do we claim that global guidance cues are absent or ignore the fact that the corneal nerves swirl; however, our results show that such global or contact-mediated cues are not required to explain the observed swirling patterns of epithelial cell radial migration in the adult cornea. Although our work shows that a swirling ‘prepattern’ of nerves is not required to obtain radial patterns of epithelial cell migration, we address below why it occurs and how it affects our data.

      An intuitive interpretation of corneal anatomy would be that sensory axons follow the path of least resistance between migrating epithelial cells. We believe this is the most likely explanation in the normal wild-type cornea; however, the reviewer correctly notes that previous literature indicates that axon patterning precedes epithelial cell patterning. Nerve swirling is observed from approximately 3 weeks, earlier than the epithelial striping patterns that become apparent in transgenic LacZ and GFP reporter mosaics at about 5 weeks (e.g. Collinson et al., 2002 [PMID: 12203735]; Iannaccone et al., 2012 [PMID: 22347498]; McKenna and Lwigale 2011 [PMID: 20811061]). We now address this in the revised manuscript and below.

      The early appearance of radial axonal projections can be readily explained even if axons are following the epithelial cells. Experimental observations (Collinson et al., 2002 [PMID: 12203735]) and our modelling (Fig. 6a,b) both indicate that spiral epithelial cell migration proceeds for many days before stripe patterns become visible in reporter mosaics. Thus, epithelial migration is already underway before the striping pattern becomes detectable. If axons are following the epithelial migration, they would be expected to become radially aligned before the reporter epithelial stripes were evident. The apparent precedence of radial axonal projections over epithelial striping is therefore fully consistent with the biological scenario in which axons follow the migrating epithelial cells. Corneal epithelial basal cells have been shown to wrap around individual and grouped subbasal axons, acting as surrogate glia (Stepp et al., 2016, Investigative Ophthalmology & Visual Science 57, 1292), which represents a plausible mechanism to allow migrating epithelial cells to shepherd axons in their direction of movement.

      We considered that epithelial cells may follow an axonal prepattern, but several experimental observations argue against the hypothesis that epithelial migration is guided by axonal projections. In situations of genetic mutation or corneal injury, epithelial cell migration can progress independently of, or ahead of, the extension of corneal axons (Leiper et al., 2009 [PMID: 19029029]; Song et al., 2004 [PMID: 14744881]). Furthermore, in chimeric Pax6<sup>+/−</sup> LacZ<sup>+</sup> ↔ Pax6<sup>+/+</sup> LacZ<sup>-</sup> mice, where the normally disrupted radial migration of Pax6<sup>+/−</sup> epithelial cells is restored, the underlying nerves may continue to exhibit abnormal projection patterns (Leiper et al., 2009 [PMID: 19029029]). These results suggest that axonal organization is at least partially dependent on epithelial behaviour rather than the reverse.

      Importantly, none of this evidence excludes the possibility that additional cues (including axonal contact or other environmental signals) may modulate or refine epithelial migration. For example, Walczysko et al. 2016 [PMID: 27563231] showed that isolated epithelial cells cultured on de-epithelialised and de-nervated corneal stroma still migrate with a small but significant radial bias, indicating that epithelial cells can respond to physical features of their environment. Our model likewise includes a component of alignment with environmental structure. Since the central radial striping in many of our simulations is somewhat less ordered than in vivo, it is plausible that additional biological guidance cues or axonal neurotrophins help refine the pattern.

      We now discuss these issues and the relevant experimental evidence more fully in the revised manuscript.

      (2) Model Validation and Complexity

      While the model is well-constructed, it currently falls short of its stated goal of elucidating the mechanisms of spiral formation. Key questions remain unanswered:

      Is the curvature of the cornea necessary for spiral formation, or would a simpler disk geometry suffice?

      To address the Reviewers’ concerns about the role of geometry, we first note that the spiral pattern also appears in a disk geometry. In fact, the geometric parameters of the cornea vary across mammals, including humans, yet the spiral pattern is preserved (Dua, et al., 1993 [PMID: 8325424]; Zander and Weddell, 1951 [PMID: 14814019]). Similar variations arise in disease states, e.g. in keratoconus the cornea becomes elongated.

      On a spherical cap, and all shapes with the same topology, the boundary winding number fixes the interior index, so ongoing limbal influx maintains a total index of 1. To explore the role of geometry more systematically, we simulated a broader range of geometries, including different curvatures, cap angles, a prolate and an oblate ellipsoid, and finally a disk, and compared the resulting patterns with published data across mammals and with disease states.

      For all of these shapes, we find the same qualitative migration pattern – a central spiral, although its precise shape depends on geometric features such as curvature and cap-angle opening. These results are discussed in a new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and shown in Fig. 10 of the revised manuscript.

      What role do boundary conditions play?

      In the revised manuscript, we address the role of boundary conditions, in particular the presence of limbal epithelial stem cell proliferation. Without limbal stem cells, and with all other parameters kept unchanged, the cornea fails to make the radial striping pattern. We observe two distinct changes: First, at J=0.1, the amount of alignment at which robust spiral formation appears normally, the simulated tissue flow is instead disordered, in fact very similar to our in vitro calibration simulations (revised Figure 4). This shows that the boundary cue of an ordered influx from the limbus promotes flocking when it otherwise would not (yet) appear. Second, when we increase the alignment to J=0.15 or J=0.2, we still do not observe spiral formation. Instead of the expected central vortex shape, we have anchoring of defects to the boundary, facilitated by the effective compressibility of the tissue because of the density feedback in the division/extrusion rates.

      These findings are discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and in Fig. 9 in the revised manuscript.

      How well do the model's predictions quantitatively match experimental data?

      The current comparisons in Fig. 4c-f lack quantitative agreement, and this discrepancy should be discussed with possible explanations.

      Regarding the in vitro cell data and matching simulations: We are aware that we have limited data to work with, and our match is intended as a rough estimate of physical parameters.

      For the revision, we have reanalysed both experimental and simulation data carefully. We realised that certain details of our numerical model amplify short-range fluctuations, namely the way cell divisions induce stress dipoles, and the fact that we did not include cell-cell friction forces. This is not merely a guess but emerged from related theoretical work by some of us (Keta and Henkes [PMID: 40556485]; Kammeraat et al, arXiv:2508.01046 (2025)).

      We therefore compared the time-averaged velocity field excluding divisions to the experiment, the same quantity that we already introduced as hydrodynamic velocity for the corneal surface. We also carried out further simulations that included cell-cell friction forces for comparison. They led to very similar results, albeit with a transition to flocking at somewhat lower alignment strength J.

      Instead of the Fourier-space velocity correlation functions that are hard to interpret, we switched to spatial velocity correlation functions. Our experiments have ‘swirly’ velocity fields with mesoscale spatial and temporal correlations, of the order of 5-6 cell sizes in space and one hour in time. As can be seen in the revised panels 4c,d for the spatial correlations of the hydrodynamic velocity, the match between experiment and simulations is in fact good for both plastic and explant substrates. There are also systematic changes in length and time scales between the two experiments that emerge without fine-tuning from simulation. We furthermore have included snapshots of both experimental in vitro conditions and matching simulations as new panels 4g-j, showing that the mesoscale correlations appear in the hydrodynamic velocity.

      Ultimately the parameter values and length and time scales that we infer for our systems (see Table 1) are quantitatively consistent with three other estimates of in vitro epithelia (Henkes et al. 2020 [PMID: 32179745]; Saraswathibhatla et al, Extreme Mechanics Letters 48, 101438 (2021), Kammeraat et al, arXiv:2508.01046 (2025)).

      For cell flows on the full cornea, regrettably we lack further quantitative data to compare to beyond the migration direction fields inferred in Figure 2, and the time scales of corneal resurfacing (Figure 6).

      (3) Importance of polar alignment.

      Polar alignment is put forward as one main feature for the observed patterns based on the simulation results. However, no experimental confirmation for such polar alignment is presented. The authors instead present various scattered discussions about polar versus nematic symmetry, which at times reads rather unnecessary and distracting from their main message.

      If they are not providing direct experimental evidence on the polar alignment, at least they could quantify nematic alignment of cell orientation as the spiral forms. One possibility is that nematic alignment of cell orientations has a crosstalk with the polar alignment associated with the self-propulsion of the cells. This is important because, as acknowledged by authors, several recent works in the context of in vitro epithelial under disk confinement have revealed alternative mechanisms for spiral vortex formation.

      We thank the Reviewer for pointing out the confusion, and we have therefore completely revised our discussion of alignment. While the evidence remains indirect, the following observations lead us to conclude that our system is dominated by substrate-based polar motility and weak polar alignment:

      We have investigated nematic alignment of cell orientations. As can been seen in new Figure 3d, corneal epithelial cells do not show evidence of significant elongation in any direction, i.e. we find no local nematic order.

      Our computational model starts from uncorrelated, substrate-based polar active cell migration. We have carefully investigated if polar alignment is a necessary ingredient and found that the in vitro data are inconsistent with the absence of alignment: the flocking order parameter is too high, and the spatial and temporal correlations are larger than those expected from persistent driving only.

      Still, cell-cell stress patterns in the epithelium may remain nematic. While one cannot exclude anything, the effect must be sufficiently weak to not affect axon motion. Their naturally long, thin shapes would strongly react to nematic stresses, but we do not see ±1/2 defects in their growth patterns, only polar ±1 defects.

      We were recently made aware the work of Lång et al. (2024) [PMID: 38630812], in which the authors also observe spiral formation in cell migration on a substrate, with +1 topological defects. They explain their observations using a polar model, and our results are consistent with their observations and model.

      In the active matter community, the ‘active nematic cell sheet’ paradigm is currently undergoing a revision. Notably, nematic-like phenomenology, in particular ±1/2 defects, can also arise from the minimal ingredients of uncorrelated polar activity and cell deformability (Chiang et al. 2024 [PMID: 39302997]).

      Minor comments:

      - It would help the reader to see some representative images of the cells, explants, and the simulation (maybe with vectors overlaid), as it could be hard to imagine what exactly is happening and the other relevant properties to compare.

      Please see new panels Fig. 4g-j, and new supplementary videos S3 and S4.

      - Add a colorbar that represents direction to Fig. 1a.

      We assume the Reviewer meant Fig. 2a. We have added a circular orientation colour chart to the image.

      - When do additional +1,-1 defects appear? The authors just say it is unlikely, but it is not clear when they are observed. Disease state?

      The reviewer is correct about pathology. We now cite (in conclusion) Collinson et al. 2004 [PMID: 15037575], which shows disruption and discontinuity in Pax6-mutant corneas with chronic corneal degeneration. We also have a publication in prep that shows extra +1 and -1 discontinuities occur during wound healing. We don’t want to include these data in this manuscript, but cite Sagga et al. (in prep).

      Reviewer #2 (Recommendations for the authors):

      Overall, the manuscript presents a solid theoretical work. However, I have a few major concerns on this manuscript. (1) The authors use the concept of topological defect to refer to stagnation points of the velocity field (2) The comparison to experiments remains qualitative and it is unclear whether the proposed mechanism is at work, (3) The role of geometry remains unclear.

      Major points:

      (1) On page 3 the authors claim that topological defects exist in velocity fields, however this statement is incorrect and can confuse readers. Unlike the director field of an ordered phase, their velocity field is a vector in R^2 with a norm that is not fixed, and therefore the velocity field has no topological defects. In fluid dynamics, these special points in the velocity field are called stagnation points, fluid sinks or sources. This distinction is important because the physical nature of the interaction forces between two "topological defects" in a velocity field is fundamentally different to the interaction forces between two topological defects in a director field. For these reasons, it can confuse readers to mix the two concepts (stagnation points vs topological defects). Note that if the authors address this concern, many parts of the main manuscript should be rewritten.

      Stagnation points are topological defects since the topological classification ignores the magnitude and looks only at the direction of the vector field away from the core. There are some caveats related to the role of boundary conditions, since in the case of a fluid, if boundary conditions are not fixed, one can eliminate the defect by setting the flow field to zero everywhere. Mathematically, a pair of point sources or sinks in an incompressible potential flow interacts through the same Green’s function that gives the elastic interaction between two-point disclinations in a 2D nematic director field, so their long-range pair potentials are formally identical (~ln(r) in 2D). The physical mechanism behind these forces is, as correctly pointed out by the reviewer, quite different. In addition, our coarse-grained velocity field is effectively compressible, which has consequences for the hydrodynamic equations we (can) write, see below.

      In the revised version, we clarify these points by adding Box 1, which summarizes the idea of topological defects.

      (2.1) How did the authors check that the velocity field that is inferred from the stripe edges matched the coarse-grained cell velocity field in a life sample? How did the authors validated the interpolation of the inferred velocity field using an XY model? What is the scale of the inferred velocity field? Can the authors clarify also this point?

      The murine cornea LacZ reporter system was specifically designed to allow for lineage tracing, i.e. following coarse-grained cell motion and growth patterns. The combination of macroscopic size (3.6 mm diameter), two-week resurfacing time and the curved surface however stymied our early attempts to directly measure the cell velocity field on the cornea using confocal time-lapse microscopy. However live sample fluorescent reporter systems are fully consistent with our in vivo model and shown conclusively that inferred velocity fields are recapitulated in vivo (Park et al., 2019 [PMID: 31843909]).

      For the XY model inference: We first note that approximately 30% of the corneal surface is covered by directly estimated velocity directions from the stripe edges (red arrows, Fig. 11f). This rises to more than 50% near the central spiral (red arrows, Fig. 11h) due to the way the stripes narrow near the centre due to cell extrusion. Therefore, the inferred areas are only slightly more than half of the cornea, and in regions where we expect the flow field to be largely uniform with no defects. The red arrows provide sufficient boundary conditions that a simple annealing simulation (or equivalently an energy minimisation) of the XY model rapidly converges to a slowly varying field consistent with those boundary conditions.

      Furthermore, in simulation we observe that stripe edges and the macroscopic velocity field correlate strongly with each other once the spiral has fully formed (Fig. 6b-c).

      The scale of the inferred velocity field is the distance between arrows in our digital version of the cornea, approximately 40 concentric rings over a 70° cone angle for a R = 1800 μm micron cornea, resulting in a spacing of 55 μm between velocity arrows. This is about at the scale of the in vitro velocity correlations. We are not able to obtain velocity magnitudes using this procedure.

      Furthermore, the spiral angle profile reported in Fig. 7b-c appear to be different to that found in experiments Fig. 2b. Can the authors clarify if their theoretical framework reproduces the spiral angle profiles?

      First, we note that empirically, simulations with the largest two radii (R = 1000 μm, R = 1500 μm), approaching the full experimental size, match the observed angle profile best. They both consist of a radially inward profile α(0) = 0° at the edges, only increasing, corresponding to a tightening spiral, below about θ = 20°. Note that very near the corneal centre, few cells contribute to data, and additionally the central defect position fluctuates somewhat. Therefore, the angle profile not reaching α(0) = 90° in the simulations is due to fluctuations and lack of statistics. As these simulations were run with our best fit experimentally matched parameters, this convergence is meaningful and cannot be scaled out.

      Second, our partial model (eq. 4, reproduced below) links the spiral profile angle α(θ) with the macroscopic velocity magnitude v(θ) and the net cell loss rate A(θ), and it depends explicitly on the corneal radius R.

      That is one equation for three radial fields. If we make the reasonable assumption that simulations at fixed 𝐽 that differ only in R have the same constitutive law A(ρ), and would follow the same continuum velocity equation that ultimately sets v(ρ), the radius R still appears explicitly in the equation. This is consistent with the different radial profiles for different R that we observe in Fig. 7c. Furthermore, in Fig. 7b we observe that above the flocking threshold, different J lead to very similar profiles. This indicates that the system enters a fully polar phase.

      The same is true in experiment: We expect that cell mechanics and planar cell polarisation coordination are local effects that will set J, A(ρ), and v(ρ). Thus, we do predict that the radius R will explicitly affect the spiral profile consistent with the equation above, but we would need more direct measurements of J, A(ρ), and v(ρ) to go any further.

      Properly answering this question would first require simulations of a non-dimensionalised model with different radii, boundary influx, alignment strengths and constitutive laws for the division / extrusion dynamics. Then one would want to construct a matching equation for the polarisation and / or velocity field, going beyond the Malthusian flock approximations of constant magnitude v(θ) and net cell loss rate A(θ) = 0. This is well beyond the scope of this publication, and there is certainly no experimental data to compare to.

      (3.1) It appears that the formation of the spiral velocity pattern results from a combination of a radial flow of cells due to cell division at the outer boundary and apoptosis at the geometrical center and azimuthal flow of cells due to flocking in confined geometries. Is this correct? Can the authors explain how does the 3d geometry of the spherical cap modify each of these flow fields?

      The Reviewer is broadly correct; however, the true picture is more subtle. Almost all of the divisions and extrusions occur in TA cells, neither at the limbus or at the corneal centre (please see the A(θ) profiles in Figure 8). Flow is then a spiral flock that is partially radial and azimuthal, and with a variable velocity magnitude. It ultimately all has to follow the flux equation 4, in steady state.

      For the modification of the flow field due to 3d geometry: Broadly, they simply change the amount of corneal surface available at different angles from the limbus when switching to isomorphic surfaces like, e.g. the disk. Thus, we still observe spirals, but of somewhat modified shapes. Please see the reply to Reviewer 1 above, and new Figure 10 for simulations of different geometries.

      A full theory of radially symmetric corneal shapes would again need additional velocity and constitutive equations and is beyond the scope of this publication.

      (3.2) In their model, it appears that the geometry is introduced by constraining the dynamics of agents, and it has not direct influence on the alignment of agents. Can the authors explain how the geometry of the spherical cap influences the emergent spiral states? Can the authors show whether their results are robust to changes in the substrate geometry? For example, by changing the substrate geometry from a spherical cap to another convex shape. Can the authors identify differences between the spiral patterns on a spherical cap vs that on a flat disk?

      Please see the response to Reviewer 1, above, and new Figure 10. Briefly, the results are robust to changes in the corneal geometry as long as they are other convex shapes. There are differences in details of the spiral shape.

      Minor comments:

      - On page 3, the authors claim that the Euler characteristic of a spherical cap or a disk is 1. Unfortunately, this statement is incorrect. The Euler characteristic of a spherical cap or a disk is determined by the winding of the vector field around the open boundary. In their case the Euler characteristic is 1 because the velocity field is oriented towards the top of the spherical cap, which give a winding of +1. The authors explain this correctly on page 13.

      The Euler characteristic is a topological invariant of the surface and does not depend on the tangent vector field. A disc or spherical cap has Euler characteristic 1. What depends on the boundary winding is the index formula for a vector field on a surface with boundary: the winding of the field along the boundary determines the corresponding boundary contribution, and hence the sum of interior indices. Thus, while the reviewer is correct about the role of boundary winding in computing the field’s index, this does not alter the Euler characteristic of the surface itself.

      We note that the orientation of the velocity field does indeed matter in our simulations: In new Figure 10, we show that in the absence of the boundary condition of inward flux at the limbus, we do not observe a spiral robustly, and we do see anchoring of defects to the boundary, with winding numbers that are now different.

      - On page 5, the authors claim the clockwise and counterclockwise -oriented spirals are equally likely, however no quantification is provided to support this claim. Can the authors clarify this point?

      The approximately equal likelihood is as described in Collinson et al. (2002) [PMID: 12203735].

      - On page 8 the authors state "Dipolar active force cannot cause a single cell to migrate, and the flow is an emergent collective phenomenon", here I was confused, because a bacteria swimming in a Newtonian fluid can self-propel by exerting a dipolar active force on the fluid. See for instance work by E. Lauga or I. Aronson. Can the authors clarify this point?

      The Reviewer is correct about how bacteria can swim using dipolar active forces. However, and unfortunately, that language was straight ported to very different conditions, that of cells migrating on a frictional substrate with no induced flow, with the equation of motion . Unless there is a net force arising from the stress profile, the cell cannot move, and with microscopic models where the active stress is a single value per cell, that statement is always true. Of course, more detailed cell models with active stresses exist, but their motion is still due to the net force arising from them. We have clarified the statement in the revised manuscript.

      - On page 16, the expression for the velocity seems to miss a parenthesis.

      Fixed. Thanks!

      - On page 16, the authors claim that the flux profiles in the simulations and in the XYZ model are in good agreement. However, there are clear differences for theta> 50 deg. Can the authors discuss the possible explanation for these differences? Note that one may expect a between agreement between two theoretical approaches.

      These disagreements are due to imperfections in the observed spiral patterns in simulations. They are not perfectly radially symmetric, and the defect is not always in the dead centre of the cornea. Thus, the theoretical predictions are not quite accurate. Numerically, what happens for θ > 50° is that the radial bins sometimes include part of the limbal zone with strong proliferation, and the corneal edge itself. Both are not described by the flux prediction. We decided not to crop out this region and rather explain where it comes from in the text.

    1. eLife Assessment

      This important study contributes to our understanding of how epithelial cells establish polarity by identifying a hierarchy in which Par3 acts upstream of centrosome positioning and apical membrane initiation. The evidence supporting the main conclusions is convincing, and the authors addressed the unresolved questions about microtubule organization and the need for clearer integration of quantitative and conceptual points raised in review. The work will be of interest to cell and developmental biologists.

    2. Reviewer #1 (Public review):

      Summary:

      Wang, Po-Kai et al., utilized the de novo polarization of MDCK cells cultured in Matrigel to assess the interdependence between polarity protein localization, centrosome positioning and apical membrane formation. They show that the inhibition of Plk4 with Centrinone does not prevent apical membrane formation, but does result in its delay, a phenotype the authors attribute to the loss of centrosomes due to the inhibition of centriole duplication. However, the targeted mutagenesis of specific centrosome proteins implicated in the positioning of centrosomes in other cell types (CEP164, ODF2, PCNT and CEP120), as well as the use of dominant negative constructs to inhibit centrosomal microtubule nucleation did not affect centrosome positioning in 3D cultured MDCK cells. A screen of proteins previously implicated in MDCK polarization revealed that the polarity protein Par-3 was upstream of centrosome positioning, similar to other cell types.

      Strengths:

      The investigation into the temporal requirement and interdependence of previously proposed regulators of cell polarization and lumen formation is valuable. The authors have provided a detailed analysis of many of these components at defined stages of polarity establishment and well demonstrate that centrosomes are not necessary for apical polarity formation, but are involved in the efficient establishment of the apical membrane.

      Weaknesses:

      Key questions remain regarding the structure of the intracellular cytoskeleton following depletion of centrosomes, centrosome proteins, or abrogation of centrosome microtubule nucleation. The authors strengthen their model that centrosomes are positioned independently of microtubule nucleation using dominant negative Cdk5RAP2 and NEDD-1 constructs, however, the structure of the intracellular microtubule network remains unresolved and will be an important avenue for future investigation.

    3. Reviewer #3 (Public review):

      Here the Wang et al resubmit their manuscript describing the events in the establishment of polarity in MDCK cells cultured in vitro. As with the original version, the description is throughout and is important to the field to report as it establishes a hierarchy of events in polarization, placing Par3 upstream of centrosome positioning and apical membrane component trafficking. Unfortunately, in the revised version, the authors addressed almost none of my points. They did a cursory job of responding in the rebuttal letter but made little attempt to actually address what was being asked or to incorporate any of my suggestions into the manuscript. The particularly egregious examples are cited below:

      Comments on revisions:

      (1) My original main experimental concern was not addressed: I had originally asked what role microtubules play in the process of polarization (either centrosomal or non-centrosomal). An obvious model is that Gp135, Rab11, etc. are delivered to the AMIS on centrosomal microtubules. Centrosomes might also be pulled to the AMIS via cortically derived microtubules as is the case in the C. elegans intestine where the centrosome moves apically on apical microtubules via dynein directed transport to the cortically anchored minus ends. The authors do not explore the role of microtubules in the revision, citing that it was not possible to observe the microtubules directly or to perform nocodazole experiments during polarization. Instead, the authors use a relatively new genetic tool to disrupt centrosomal microtubules. They appear to succeed in displacing centrosomal g-tubulin using this tool, but without being able to observe microtubules, a remaining caveat of this experiment is that it is still unclear whether the authors have removed centrosomal microtubules. Compounding this issue is that this tool has never been used in MDCK cells. The authors conclude "we found that cells lacking centrosomal microtubules were still able to polarize and position the centrioles apically.", but they have not shown this, instead the data suggest this conclusion and the authors should acknowledge the caveat that they have no idea whether centrosomal microtubules are abolished. Similarly, the authors also state: "Additionally, although PCNT knockout cells show reduced microtubule nucleation ability, they still recruit a small amount of γ-tubulin". Where are the data that show that microtubule nucleation is reduced in these PCNT knock out cells?

      (2) Many of my comments were addressed in the rebuttal, but not in the text.

      The non-centrosomal GP135 in Figure 2 is not acknowledged or explained.

      That the polarity index does not actually measure polarity, but nuclear-centrosome distance is not acknowledged or explained in the paper.

      I still don't believe that the quantification in Figure 3D matches the images I am being shown in Figure 3A. In the centrinone treatment condition, there is certainly an enrichment of GP135 at the AMIS that is not detected in the quantification. The method described in the rebuttal might miss this enrichment if it is offset from line drawn between the centroid of the two nuclei.

      Cell height changes in the centrosome depleted cysts are still referenced in the text ("the cell heights of the centrosome-depleted cysts are less uniform"), but no specific data or image is called out. Currently, Figure 3G is referenced, but that is a graph of GP135 intensity.

      In my original review, I called on the authors to comment on the striking similarity of the mechanisms they documented in MDCK cells to what has been shown in in vivo systems. The authors did not do this, instead restating in the rebuttal some features of what they found. But the mechanisms shown here are remarkably similar to the polarization of primordia that generate tubular organs in vivo. Perhaps most striking is the similarity to the C> elegans intestine where Par3 localizes to the cortex at the site of an apical MTOC that pulls the centrosome to the apical surface via dynein (Feldman and Priess, 2012). Instead of discussing this similarity, the authors state: "Par3 is likely to regulate centrosome positioning through some intermediate molecules or mechanisms, but its specific mechanism is still unclear and requires further investigation." Given the acetylated tubulin signal emanating from the Par3 positive patch in Figure 5E and F, I suspect similar mechanisms to the C. elegans intestine are at play here. Such a parallel should be noted in the Discussion.

      I had originally commented that "I find the results in Figure 6G puzzling. Why is ECM signaling required for Gp135 recruitment to the centrosome. Could the authors discuss what this means?" The authors responded that "The data in Figure 6G do not indicate that ECM signaling is required for the recruitment of Gp135 to the centrosome". In Figure 6G, the localization of GP135 to the centrosome appears significantly delayed compared to its localization to the centrosome in images where cells were cultured in Matrigel. Indeed, the authors argue that the centrosomal localization precedes and contributes to its localization to the AMIS. In the absence of ECM, GP135 localizes to the membrane before it localizes to the centrosome and its localization to the centrosome appears significantly reduced. Thus, my original and current interpretation is that ECM signaling is somehow required for the centrosomal targeting of GP135. One could make a competition argument, i.e. that the cortex in the absence of ECM is somehow a more desirable place to localize than the centrosome, but this experiment also argues that the centrosome does not need to be a source of this material in order for it to end up on the cortex.

      (3) There needs to be precision in the language used in many places:

      I don't understand this line in the abstract: "When cultured in Matrigel, de novo polarization of a single epithelial cell is often coupled with mitosis." If a cell has divided, it is no longer a single cell.

      The authors state in the Introduction "Because of its strong ability to nucleate microtubules, the centrosome functions as the primary microtubule organizing center", but then state ""In polarized epithelial cells, the centrosome is localized at the apical region during interphase, which contributes to the construction of an asymmetric microtubule network conducive to polarized vesicle trafficking". In the latter statement, I assume the authors are describing the well-characterized apical microtubule network in epithelial cells that is non-centrosomal. Thus, the latter sentence is at odds with the former.

      The authors continually refer to Par3 as a tight junction protein. "Par3, which controls tight junction assembly to partition the apical surface from the basolateral surface". To my knowledge, PARD3 is an apical protein with similar localization to C. elegans PAR-3 and Drosophila Bazooka. PARD3B is a junctional protein. I assume that the antibody that the authors are using is to PARD3 and not PARD3B? Can the authors please clarify this in the text.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Wang, Po-Kai et al., utilized the de novo polarization of MDCK cells cultured in Matrigel to assess the interdependence between polarity protein localization, centrosome positioning and apical membrane formation. They show that the inhibition of Plk4 with Centrinone does not prevent apical membrane formation, but does result in its delay, a phenotype the authors attribute to the loss of centrosomes due to the inhibition of centriole duplication. However, the targeted mutagenesis of specific centrosome proteins implicated in the positioning of centrosomes in other cell types (CEP164, ODF2, PCNT and CEP120), as well as the use of dominant negative constructs to inhibit centrosomal microtubule nucleation did not affect centrosome positioning in 3D cultured MDCK cells. A screen of proteins previously implicated in MDCK polarization revealed that the polarity protein Par-3 was upstream of centrosome positioning, similar to other cell types.

      Strengths:

      The investigation into the temporal requirement and interdependence of previously proposed regulators of cell polarization and lumen formation is valuable. The authors have provided a detailed analysis of many of these components at defined stages of polarity establishment, and well demonstrate that centrosomes are not necessary for apical polarity formation, but are involved in the efficient establishment of the apical membrane.

      Weaknesses:

      Key questions remain regarding the structure of the intracellular cytoskeleton following depletion of centrosomes, centrosome proteins, or abrogation of centrosome microtubule nucleation. The authors strengthen their model that centrosomes are positioned independently of microtubule nucleation using dominant negative Cdk5RAP2 and NEDD-1 constructs, however, the structure of the intracellular microtubule network remains unresolved and will be an important avenue for future investigation.

      We thank the reviewer for raising this important point. We agree that understanding the organization of the intracellular microtubule network following centrosome depletion, disruption of centrosomal proteins, or inhibition of centrosomal microtubule nucleation will be important for further mechanistic insight. However, a detailed analysis of cytoskeletal architecture under 3D culture conditions would require super-resolution or other advanced microscopy techniques and therefore falls beyond the scope of the current study. Nevertheless, several previous studies conducted under conventional 2D culture conditions provide relevant mechanistic context for interpreting our findings.

      (1) Centrosome depletion by centrinone treatment

      Previous studies have shown that upon centrosome loss induced by the Plk4 inhibitor centrinone, alternative microtubule-organizing centers (MTOCs), particularly the Golgi apparatus, can compensate by nucleating non-centrosomal microtubules (Chen et al., 2022; Gavilan et al., 2018; Martin, Veloso, Wu, Katrukha, & Akhmanova, 2018; Wu et al., 2016). Consistent with these findings, our microtubule regrowth assays in 2D MDCK cells (see Author response image 1) demonstrated that centriole depletion markedly altered microtubule organization. Control cells displayed the typical radial microtubule array emanating from a centralized centrosome, whereas centrinone-treated cells exhibited a dispersed microtubule network with enhanced microtubule growth from Golgi-associated sites.

      Author response image 1.

      Staining of control or centrinone (CN) treated MDCK cells for α-tubulin (magenta), Golgi GM130 green) and DNA (DAPI, blue) 1 min after nocodazole washout. Z-maximum projections of confocal images. The boxed cells in the overview images are magnified, and the microtubule regrowth regions are further enlarged. Scale bars: 50 μm (overview images), 10 μm (magnified cells), and 5 μm (enlarged microtubule regrowth regions).

      (2) Disruption of centriole or centrosomal proteins

      Previous studies have reported that depletion of the subdistal appendage protein ODF2 reduces centrosome–microtubule interactions and destabilizes centrosomal microtubules (Hung, Hehnly, & Doxsey, 2016; Ibi et al., 2011; Tateishi et al., 2013). In addition, in cells lacking the PCM protein pericentrin (PCNT), AKAP450 has been shown to partially compensate for centrosomal microtubule nucleation activity, although at a somewhat reduced level compared with wild-type cells (Figure 4—figure supplement 3A, B) (Gavilan et al., 2018).

      (3) Inhibition of centrosomal microtubule nucleation

      For dominant-negative Cdk5RAP2 and NEDD1 constructs, a previous study demonstrated that these constructs displace γ-tubulin from centrosomes and impair centrosomal microtubule nucleation (Vinopal et al., 2023). Consistent with this report, our MDCK cells expressing dominant-negative Cdk5RAP2 or NEDD1 also exhibited reduced γ-tubulin localization at centrioles (Figure 4—figure supplement 3C, D, and G).

      Together, these findings support the interpretation that centrosomal microtubules are not strictly required for polarized vesicle trafficking, centrosome migration, or epithelial polarization, but instead enhance the efficiency and robustness of these processes. We agree with the reviewer that future studies using advanced imaging approaches under 3D culture conditions will be important to resolve the spatial organization and dynamics of the intracellular cytoskeleton during epithelial polarization.

      Reviewer #3 (Public review):

      Here the Wang et al resubmit their manuscript describing the events in the establishment of polarity in MDCK cells cultured in vitro. As with the original version, the description is throughout and is important to the field to report as it establishes a hierarchy of events in polarization, placing Par3 upstream of centrosome positioning and apical membrane component trafficking. Unfortunately, in the revised version, the authors addressed almost none of my points. They did a cursory job of responding in the rebuttal letter but made little attempt to actually address what was being asked or to incorporate any of my suggestions into the manuscript. The particularly egregious examples are cited below:

      Comments on revisions:

      (1) My original main experimental concern was not addressed: I had originally asked what role microtubules play in the process of polarization (either centrosomal or non-centrosomal). An obvious model is that Gp135, Rab11, etc. are delivered to the AMIS on centrosomal microtubules. Centrosomes might also be pulled to the AMIS via cortically derived microtubules as is the case in the C. elegans intestine where the centrosome moves apically on apical microtubules via dynein directed transport to the cortically anchored minus ends. The authors do not explore the role of microtubules in the revision, citing that it was not possible to observe the microtubules directly or to perform nocodazole experiments during polarization.

      Instead, the authors use a relatively new genetic tool to disrupt centrosomal microtubules. They appear to succeed in displacing centrosomal g-tubulin using this tool, but without being able to observe microtubules, a remaining caveat of this experiment is that it is still unclear whether the authors have removed centrosomal microtubules. Compounding this issue is that this tool has never been used in MDCK cells. The authors conclude "we found that cells lacking centrosomal microtubules were still able to polarize and position the centrioles apically.", but they have not shown this, instead the data suggest this conclusion and the authors should acknowledge the caveat that they have no idea whether centrosomal microtubules are abolished.

      We appreciate the reviewer’s important comments regarding the role of microtubules during epithelial polarization. We agree that determining how centrosomal and/or non-centrosomal microtubules contribute to apical trafficking and centrosome positioning represents an important mechanistic question.

      We previously attempted to directly visualize microtubules during live imaging using SPY-tubulin labeling (see Author response image 2). However, under 3D Matrigel culture conditions, MDCK cells rapidly become rounded and densely packed, substantially reducing image contrast and making the majority of intracellular microtubule networks difficult to resolve, except for spindle microtubules and the cytokinetic bridge. In addition, nocodazole treatment caused mitotic arrest under our experimental conditions, thereby preventing de novo polarization from proceeding and precluding interpretation of polarity establishment.

      Author response image 2.

      Time-lapse maximum-intensity z-projections of MDCK cells expressing EGFP-PACT (yellow; centrosome marker) and H2B-mCherry (magenta; nuclei) embedded in Matrigel. Microtubules were labeled with SiR-tubulin (cyan) before live-cell imaging. Images show a representative dividing cell. Time stamps indicate hours and minutes relative to anaphase onset (0:00). Scale bar, 10 μm.

      To partially address the role of centrosomal microtubules, we used dominant-negative Cdk5RAP2 and NEDD1 constructs that have previously been shown to displace γ-tubulin from centrosomes and impair centrosomal microtubule nucleation (Vinopal et al., 2023). Consistent with this study, we observed substantial loss of γ-tubulin from centrioles in MDCK cells expressing these constructs (Figure 4— figure supplement 3C, D, and G). However, as the reviewer correctly points out, we were unable to directly visualize centrosomal microtubules under our 3D imaging conditions. Therefore, we cannot definitively conclude that centrosomal microtubules were completely abolished. We have revised the manuscript to clarify this limitation and to more cautiously state that our data suggest centrosomal microtubules may not be strictly required, for apical polarization and centrosome positioning under these conditions.

      Similarly, the authors also state: "Additionally, although PCNT knockout cells show reduced microtubule nucleation ability, they still recruit a small amount of γ-tubulin". Where are the data that show that microtubule nucleation is reduced in these PCNT knock out cells?

      We thank the reviewer for this comment and apologize for not sufficiently presenting these data in the previous revision. To directly assess microtubule nucleation activity in PCNT-KO cells, we performed microtubule regrowth assays in MDCK cells following nocodazole washout (Figure 4— figure supplement 3A, B). Compared with wild-type cells, PCNT-KO cells showed reduced centrosomal microtubule regrowth, indicating impaired microtubule nucleation capacity.

      Importantly, microtubule nucleation was not completely abolished in PCNT-KO cells, consistent with previous reports showing that AKAP450 can partially compensate for the loss of pericentrin and maintain residual centrosomal microtubule nucleation activity (Gavilan et al., 2018).

      (2) Many of my comments were addressed in the rebuttal, but not in the text.

      We sincerely thank the reviewer for the valuable suggestions. We have carefully considered all comments and incorporated many of the recommended revisions into the revised manuscript. However, we found that including every additional analysis, experiment, and discussion in the main manuscript would substantially reduce its coherence and readability. Therefore, while not all new analyses and experimental results are included in the revised manuscript, we have addressed every comment comprehensively in this response letter. Where necessary, we performed additional experiments and analyses to obtain the requested data, and the corresponding results and explanations are provided in our responses. We hope the reviewer will understand our effort to thoroughly address all comments while preserving the clarity and overall flow of the manuscript.

      The non-centrosomal GP135 in Figure 2 is not acknowledged or explained.

      We apologize for not sufficiently addressing the non-centrosomal Gp135 signal in Figure 2. We have now revised the manuscript to explicitly describe and discuss this point in the text (Page 5, Paragraph 4).

      That the polarity index does not actually measure polarity, but nuclear-centrosome distance is not acknowledged or explained in the paper.

      We have revised the manuscript to explicitly state that the “polarity index” represents the distance between the nucleus and the centrosome, which we use as a quantitative indicator of the degree of cell polarity (Page 5, Paragraph 1).

      I still don't believe that the quantification in Figure 3D matches the images I am being shown in Figure 3A. In the centrinone treatment condition, there is certainly an enrichment of GP135 at the AMIS that is not detected in the quantification. The method described in the rebuttal might miss this enrichment if it is offset from line drawn between the centroid of the two nuclei.

      We thank the reviewer for this comment. To better address this concern, we have now included a 3D view of the corresponding image data (see Author response image 3 and Author response image 4). This analysis clarifies that, in the centrinone-treated condition, the Gp135 signal is not localized at the geometric center of the cell doublet, but is instead offset from the axis used in our line-scan quantification. As a result, the enrichment visible in the projection image was not fully captured by the original quantification method. SiR-DNA Gp135

      Author response image 3.

      Author response image 4.

      3D reconstructions of p53-KO control and centrinone-treated MDCK cell doublets expressing EGFP-Gp135. Images are shown after 90° rotations about the x-axis (or y-axis) to visualize the spatial distribution of Gp135. Fluorescence intensity profiles of EGFP-Gp135 were measured along the line connecting the two nuclei. White arrows indicate the central fluorescence intensity value used to quantify Gp135 accumulation at the apical membrane initiation site (AMIS) (a.u., arbitrary units). Time stamps indicate hours and minutes.

      Cell height changes in the centrosome depleted cysts are still referenced in the text ("the cell heights of the centrosome-depleted cysts are less uniform"), but no specific data or image is called out. Currently, Figure 3G is referenced, but that is a graph of GP135 intensity

      We have revised the manuscript to indicate representative images, ensuring that the text and figures are consistent (Page 7, Paragraph 3).

      In my original review, I called on the authors to comment on the striking similarity of the mechanisms they documented in MDCK cells to what has been shown in in vivo systems. The authors did not do this, instead restating in the rebuttal some features of what they found. But, the mechanisms shown here are remarkably similar to the polarization of primordia that generate tubular organs in vivo. Perhaps most striking is the similarity to the C. elegans intestine where Par3 localizes to the cortex at the site of an apical MTOC that pulls the centrosome to the apical surface via dynein (Feldman and Priess, 2012). Instead of discussing this similarity, the authors state: "Par3 is likely to regulate centrosome positioning through some intermediate molecules or mechanisms, but its specific mechanism is still unclear and requires further investigation." Given the acetylated tubulin signal emanating from the Par3 positive patch in Figure 5E and F, I suspect similar mechanisms to the C. elegans intestine are at play here. Such a parallel should be noted in the Discussion.

      We thank the reviewer for this insightful suggestion. In the revised manuscript, we have expanded the Discussion section to compare our findings with epithelial polarization mechanisms described in in vivo systems, including the C. elegans intestine and other tubular epithelial tissues (Page 13, Paragraph 2).

      We agree that the hierarchical relationship we observe between Par3 localization, centrosome positioning, and apical membrane formation bears important conceptual similarities to mechanisms reported in the C. elegans intestine (Feldman & Priess, 2012). We also considered the possibility that Par3 may regulate centrosome positioning through dynein-dependent mechanisms.

      To examine this possibility, we performed immunofluorescence staining for the dynein cofactor dynactin subunit p150<sup>Glued</sup>. However, we did not observe enrichment of p150<sup>Glued</sup> at the center of cell doublets during the cytokinetic pre-abscission stage (see Author response image 5), suggesting that dynein is not strongly concentrated together with Par3 near the AMIS under our conditions.

      In addition, pharmacological inhibition of dynein resulted in cytokinesis failure and the formation of binucleated cells, preventing reliable assessment of centrosome migration and polarity establishment.

      We would also like to clarify that the acetylated tubulin signal observed in Figure 5E and F does not emanate from the Par3-positive patch. Rather, this signal corresponds to the cytokinetic bridge, adjacent to which Par3 accumulates during cytokinesis. Consistent with this interpretation, γ-tubulin was not detected at the Par3-positive region (Figure 1A).

      Author response image 5.

      Single MDCK cells after 12 h of culture in Matrigel. Immunostaining signals of the indicated markers are shown: centrosome marker PACT-mKO1, dynactin subunit p150Glued, Gp135, and DAPI. Single confocal sections through the middle of a cyst are shown. The order of polarization is arranged from single cell (1-cell), metaphase (Meta), telophase (Telo), cytokinetic pre-abscission (Pre-Abs), post-cytokinesis (Post-CK), to lumen open (LO). Scale bar: 5 μm.

      I had originally commented that "I find the results in Figure 6G puzzling. Why is ECM signaling required for Gp135 recruitment to the centrosome. Could the authors discuss what this means?" The authors responded that "The data in Figure 6G do not indicate that ECM signaling is required for the recruitment of Gp135 to the centrosome". In Figure 6G, the localization of GP135 to the centrosome appears significantly delayed compared to its localization to the centrosome in images where cells were cultured in Matrigel.

      Indeed, the authors argue that the centrosomal localization precedes and contributes to its localization to the AMIS. In the absence of ECM, GP135 localizes to the membrane before it localizes to the centrosome and its localization to the centrosome appears significantly reduced. Thus, my original and current interpretation is that ECM signaling is somehow required for the centrosomal targeting of GP135. One could make a competition argument, i.e. that the cortex in the absence of ECM is somehow a more desirable place to localize than the centrosome, but this experiment also argues that the centrosome does not need to be a source of this material in order for it to end up on the cortex.

      We agree that the absence of ECM substantially alters the trafficking behavior of Gp135.

      Our interpretation is that ECM primarily promotes the endocytosis and internal trafficking of Gp135, thereby enabling its redistribution to membrane domains lacking ECM contact and facilitating AMIS formation (Buckley & St Johnston, 2022; O'Brien et al., 2001; Yu et al., 2005).

      Under ECM-free conditions, a larger fraction of Gp135 remains associated with the plasma membrane, resulting in reduced internalized Gp135 available for centrosome-associated trafficking. During anaphase to telophase (Figure 6G, 0:05–0:20), Gp135 predominantly redistributes along the plasma membrane toward the cleavage furrow. Only after cytokinesis initiation (Figure 6G, 0:30) do we observe a small amount of internalized Gp135 associated with centrosomes near the center of the cell doublet.

      Importantly, we agree with the reviewer that these findings suggest centrosomal trafficking is not absolutely required for Gp135 to localize to the plasma membrane. Rather, our data support a model in which centrosome-associated trafficking contributes specifically to the efficient and spatially restricted delivery of Gp135 to the AMIS during epithelial polarization.

      We have revised the manuscript to clarify this interpretation and to avoid overstating the role of centrosome-associated Gp135 trafficking.

      (3) There needs to be precision in the language used in many places:

      I don't understand this line in the abstract: "When cultured in Matrigel, de novo polarization of a single epithelial cell is often coupled with mitosis." If a cell has divided, it is no longer a single cell.

      We have revised the sentence to: “When cultured in Matrigel, de novo polarization of a single epithelial cell is often coupled with cytokinesis (Page 1, Paragraph 1).” This indicates that polarization happens as the cell divides. We thank the reviewer for this helpful suggestion, which has improved the readability of the sentence.

      The authors state in the Introduction "Because of its strong ability to nucleate microtubules, the centrosome functions as the primary microtubule organizing center", but then state ""In polarized epithelial cells, the centrosome is localized at the apical region during interphase, which contributes to the construction of an asymmetric microtubule network conducive to polarized vesicle trafficking". In the latter statement, I assume the authors are describing the well-characterized apical microtubule network in epithelial cells that is non-centrosomal. Thus, the latter sentence is at odds with the former.

      We did not intend to refer to the apical non-centrosomal microtubule network present in mature, fully polarized epithelial cells. Rather, we were referring to the off-center centrosome functions as an off-center MTOC, creating an asymmetric microtubule network during the early stages of epithelial polarization, as mentioned in previous review papers (Meiring, Shneyer, & Akhmanova, 2020).

      The apical non-centrosomal microtubule network is a feature of mature, fully polarized epithelial cells. In fact, its formation is also driven by the release of microtubule minus-ends from the off-centre centrosome, which are then transferred to the apical membrane (Goldspink et al., 2017; Moss et al., 2007; Sanchez & Feldman, 2017).

      The authors continually refer to Par3 as a tight junction protein. "Par3, which controls tight junction assembly to partition the apical surface from the basolateral surface". To my knowledge, PARD3 is an apical protein with similar localization to C. elegans PAR-3 and Drosophila Bazooka. PARD3B is a junctional protein. I assume that the antibody that the authors are using is to PARD3 and not PARD3B? Can the authors please clarify this in the text?

      The antibody used for PARD3 staining was the Merck Millipore rabbit polyclonal antibody (Cat. No. 07-330), generated against a GST-tagged recombinant fragment corresponding to 288 amino acids from the internal region of mouse PAR-3.

      Canine PARD3 and PARD3B are encoded by distinct genes located on chromosomes 2 and 37, respectively. In our study, we used two independent shRNA constructs specifically targeting canine PARD3, both of which reduced the immunoblot signal detected by this antibody, supporting the conclusion that the antibody primarily recognizes PARD3 rather than PARD3B.

      However, in immunofluorescence staining, the signal was mainly localized at tight junctions, and we did not observe significant signal at the apical membrane. We will further clarify in the revised manuscript that this antibody targets PARD3 rather than PARD3B (Page 10, Paragraph 1).

      Reference

      Buckley, C. E., & St Johnston, D. (2022). Apical-basal polarity and the control of epithelial form and function. Nat Rev Mol Cell Biol, 23(8), 559–577. doi:10.1038/s41580-022-00465-y

      Chen, F., Wu, J., Iwanski, M. K., Jurriens, D., Sandron, A., Pasolli, M., . . . Akhmanova, A. (2022). Self-assembly of pericentriolar material in interphase cells lacking centrioles. Elife, 11. doi:10.7554/eLife.77892

      Feldman, J. L., & Priess, J. R. (2012). A role for the centrosome and PAR-3 in the hand-off of MTOC function during epithelial polarization. Curr Biol, 22(7), 575–582. doi:10.1016/j.cub.2012.02.044

      Gavilan, M. P., Gandolfo, P., Balestra, F. R., Arias, F., Bornens, M., & Rios, R. M. (2018). The dual role of the centrosome in organizing the microtubule network in interphase. EMBO Rep, 19(11). doi:10.15252/embr.201845942

      Goldspink, D. A., Rookyard, C., Tyrrell, B. J., Gadsby, J., Perkins, J., Lund, E. K., . . . Mogensen, M. M. (2017). Ninein is essential for apico-basal microtubule formation and CLIP-170 facilitates its redeployment to noncentrosomal microtubule organizing centres. Open Biol, 7(2). doi:10.1098/rsob.160274

      Hung, H. F., Hehnly, H., & Doxsey, S. (2016). The Mother Centriole Appendage Protein Cenexin Modulates Lumen Formation through Spindle Orientation. Curr Biol, 26(6), 793–801. doi:10.1016/j.cub.2016.01.025

      Ibi, M., Zou, P., Inoko, A., Shiromizu, T., Matsuyama, M., Hayashi, Y., . . . Inagaki, M. (2011). Trichoplein controls microtubule anchoring at the centrosome by binding to Odf2 and ninein. J Cell Sci, 124(Pt 6), 857–864. doi:10.1242/jcs.075705

      Martin, M., Veloso, A., Wu, J., Katrukha, E. A., & Akhmanova, A. (2018). Control of endothelial cell polarity and sprouting angiogenesis by non-centrosomal microtubules. Elife, 7. doi:10.7554/eLife.33864

      Meiring, J. C. M., Shneyer, B. I., & Akhmanova, A. (2020). Generation and regulation of microtubule network asymmetry to drive cell polarity. Curr Opin Cell Biol, 62, 86–95. doi:10.1016/j.ceb.2019.10.004

      Moss, D. K., Bellett, G., Carter, J. M., Liovic, M., Keynton, J., Prescott, A. R., . . . Mogensen, M. M. (2007). Ninein is released from the centrosome and moves bi-directionally along microtubules. J Cell Sci, 120(Pt 17), 3064–3074. doi:10.1242/jcs.010322

      O'Brien, L. E., Jou, T. S., Pollack, A. L., Zhang, Q., Hansen, S. H., Yurchenco, P., & Mostov, K. E. (2001). Rac1 orientates epithelial apical polarity through effects on basolateral laminin assembly. Nat Cell Biol, 3(9), 831–838. doi:10.1038/ncb0901-831

      Sanchez, A. D., & Feldman, J. L. (2017). Microtubule-organizing centers: from the centrosome to noncentrosomal sites. Curr Opin Cell Biol, 44, 93–101. doi:10.1016/j.ceb.2016.09.003

      Tateishi, K., Yamazaki, Y., Nishida, T., Watanabe, S., Kunimoto, K., Ishikawa, H., & Tsukita, S. (2013). Two appendages homologous between basal bodies and centrioles are formed using distinct Odf2 domains. J Cell Biol, 203(3), 417–425. doi:10.1083/jcb.201303071

      Vinopal, S., Dupraz, S., Alfadil, E., Pietralla, T., Bendre, S., Stiess, M., . . . Bradke, F. (2023). Centrosomal microtubule nucleation regulates radial migration of projection neurons independently of polarization in the developing brain. Neuron, 111(8), 1241–1263 e1216. doi:10.1016/j.neuron.2023.01.020

      Wu, J., de Heus, C., Liu, Q., Bouchet, B. P., Noordstra, I., Jiang, K., . . . Akhmanova, A. (2016). Molecular Pathway of Microtubule Organization at the Golgi Apparatus. Dev Cell, 39(1), 44–60. doi:10.1016/j.devcel.2016.08.009

      Yu, W., Datta, A., Leroy, P., O'Brien, L. E., Mak, G., Jou, T. S., . . . Zegers, M. M. (2005). Beta1-integrin orients epithelial polarity via Rac1 and laminin. Mol Biol Cell, 16(2), 433–445. doi:10.1091/mbc.e04-05-0435

    1. eLife Assessment

      The authors show that innate defensive behavior in mice is shaped by threat intensity, reward value, and social hierarchy, highlighting how value and social context influence instinctive decisions. The authors provide a valuable characterization of escape behavior which approximates naturalistic conditions. Despite minor methodological limitations, the work provides a solid foundation for future investigation of how reward and social context interact to influence behavior.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments in their discussion of the limitations that were raised in the previous round of review.]

      This study by Li and colleagues examines how defensive responses to visual threats during foraging are modulated by both reward level and social hierarchy. Using a semi-naturalistic paradigm, the authors test how the availability of water or sucrose, with sucrose being more rewarding than water, shapes escape behavior in mice exposed to looming stimuli of different intensities, which are used to probe perceived threat level and defensive responses. In parallel, the study compares dominant and subordinate animals to assess how social rank biases the trade-off between reward seeking and threat avoidance. By combining behavioral analyses with computational modeling, the work addresses how reward level and social context jointly influence escape decisions in an ethological setting.

      Across the different experimental conditions, perceived threat level is the main determinant of behavior. The authors show that looming stimuli associated with higher threat (contrast) consistently elicit faster and more robust escape responses than lower threat stimuli. This effect is particularly evident during early exposures, when animals are highly vigilant and have not yet habituated to the looming stimulus (learned that it is not dangerous). Later they described that as animals gain experience and habituate, behavior becomes more flexible, and reward level begins to exert a graded modulation of the escape response. Importantly, the authors show that under high threat conditions increasing reward value leads to more frequent and faster escape rather than greater reward pursuit, specifically in dominant mice. This finding is particularly relevant, as it suggests that highly valued rewards can heighten vigilance and thereby enhance responsiveness to threat, highlighting that reward does not simply compete with defensive behavior but can also reshape it depending on the perceived level of danger, in contrast to low threat conditions, where threat can be more easily outweighed by reward. However, it is worth noting that the authors use an extremely low contrast for the low threat condition (20%), which may to some extent be insufficient to reliably trigger escape responses. Thus, an important conceptual contribution of the study is the introduction of vigilance as a useful framework to interpret these effects. Vigilance is treated as a behavioral state reflecting heightened attention to potential danger. In line with what is known from natural foraging, mice initially maintain high vigilance when confronted with an innate threat. This perspective helps clarify a finding that might otherwise appear counterintuitive. One might expect higher rewards to motivate animals to tolerate risk, explore more, and habituate faster in any scenario. Instead, the data suggest that highly rewarding outcomes can elevate vigilance, making animals more responsive to threat and leading to faster or more frequent escape under high threat conditions. In this sense, reward does not simply compete with threat but can also amplify sensitivity to it, depending on the internal state of the animal.

      The social results are particularly interesting in this context as well. Dominant mice consistently prioritize avoidance over reward, showing stronger escape responses and slower habituation than subordinates. This behavior is well captured by the vigilance framework proposed by the authors: dominant animals appear to maintain higher vigilance, which biases decisions toward threat avoidance. The authors further suggest that stable social relationships sustain high vigilance and slow habituation, framing this as an evolutionarily conserved strategy that may enhance survival. This interpretation provides a valuable perspective on how social structure shapes defensive behavior beyond immediate physical interactions. At the same time, there are important limitations to this interpretation. All experiments were conducted in male mice, and it is possible that the relationship between social hierarchy, vigilance, and defensive behavior would differ substantially in females. In addition, the idea that stable social relationships sustain elevated vigilance should be interpreted carefully, as it does not fully align with broader views of social stability as protective against anxiety and stress and generally beneficial for mental health and resilience. These points do not undermine the findings but suggest that the social effects described here should be interpreted with caution and within the specific context of the task and sex studied.

      Another important limitation is that the neural mechanisms underlying these effects remain highly speculative. Although the manuscript includes an extensive discussion of candidate circuits, particularly involving the superior colliculus and downstream structures, these interpretations go far beyond the data presented in the study and are not directly supported by experimental evidence within the paper itself. The discussion gives substantial weight to potential circuit mechanisms based primarily on previous literature rather than on findings from the current study. Given the complexity and distributed nature of the circuits likely involved in integrating vigilance, reward, social context, and defensive behavior, the present work is better viewed as providing a strong behavioral framework rather than direct mechanistic insight into the underlying neural substrates. In this context, some references discussing how animals learn to suppress defensive responses to repeated looming threats and the neural mechanisms supporting this process could further strengthen the discussion (Salay et al 2021; Fratzl et al. 2021; Conway et al. 2025; Mederos et al. 2025).

      Methodologically, the behavioral paradigm is well suited for studying escape decisions in socially housed animals, and the machine learning based classification of defensive responses is a strength. The computational model provides a useful formalization of how threat level, reward level, and vigilance interact and may be valuable for other laboratories studying escape, approach avoidance, or conflict situations, particularly as a way to classify behavioral outcomes after pose estimation. More generally, the work will be of interest to the neuroethology community for its detailed characterization of escape behavior under naturalistic conditions. At the same time, some statements in the discussion slightly overstate the novelty of the methodological approach. For example, the claim that the study differs from earlier work by using machine learning rather than manual annotation overlooks that several previous studies have already implemented automated or semi-automated strategies to classify looming evoked defensive behaviors beyond manual scoring alone.

      Given the ethological nature of the study and the high inter individual variability reported by the authors, clarity and precision in the methods are especially important for reproducibility. While the revised manuscript addresses many earlier concerns, some aspects remain slightly difficult to follow. For example, the main text states that animals were not water deprived to minimize differences in internal state across conditions, whereas parts of the methods describe experiments in which animals were water deprived. This distinction is not always clearly explained across the different experimental sections, despite internal state being central to the interpretation of the behavioral findings. A clearer separation and description of these conditions would further strengthen confidence in the work. In addition, it was somewhat surprising that the low contrast (20%) looming condition was still sufficient to trigger robust escape responses, and additional clarification or discussion regarding stimulus saliency at this contrast level could help readers better contextualize these findings.

      Overall, this study provides a rich analysis of how reward level and social hierarchy modulate defensive behavior through changes in vigilance. It offers a useful conceptual advance for thinking about escape behavior in semi-naturalistic settings and lays a solid foundation for future work aimed at linking these behavioral states to underlying neural circuits.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Li and colleagues examines how defensive responses to visual threats during foraging are modulated by both reward level and social hierarchy. Using a semi-naturalistic paradigm, the authors test how the availability of water or sucrose, with sucrose being more rewarding than water, shapes escape behavior in mice exposed to looming stimuli of different intensities, which are used to probe perceived threat level and defensive responses. In parallel, the study compares dominant and subordinate animals to assess how social rank biases the trade-off between reward seeking and threat avoidance. By combining behavioral analyses with computational modeling, the work addresses how reward level and social context jointly influence escape decisions in an ethological setting.

      Across the different experimental conditions, perceived threat level is the main determinant of behavior. The authors show that looming stimuli associated with higher threat (contrast) consistently elicit faster and more robust escape responses than lower threat stimuli. This effect is particularly evident during early exposures, when animals are highly vigilant and have not yet habituated to the looming stimulus (learned that it is not dangerous). Later they described that as animals gain experience and habituate, behavior becomes more flexible, and reward level begins to exert a graded modulation of the escape response. Importantly, the authors show that under high threat conditions increasing reward value leads to more frequent and faster escape rather than greater reward pursuit, specifically in dominant mice. This finding is particularly relevant, as it suggests that highly valued rewards can heighten vigilance and thereby enhance responsiveness to threat, highlighting that reward does not simply compete with defensive behavior but can also reshape it depending on the perceived level of danger, in contrast to low threat conditions, where threat can be more easily outweighed by reward. However, it is worth noting that the authors use an extremely low contrast for the low threat condition (20%), which may to some extent be insufficient to reliably trigger escape responses. Thus, an important conceptual contribution of the study is the introduction of vigilance as a useful framework to interpret these effects. Vigilance is treated as a behavioral state reflecting heightened attention to potential danger. In line with what is known from natural foraging, mice initially maintain high vigilance when confronted with an innate threat. This perspective helps clarify a finding that might otherwise appear counterintuitive. One might expect higher rewards to motivate animals to tolerate risk, explore more, and habituate faster in any scenario. Instead, the data suggest that highly rewarding outcomes can elevate vigilance, making animals more responsive to threat and leading to faster or more frequent escape under high threat conditions. In this sense, reward does not simply compete with threat but can also amplify sensitivity to it, depending on the internal state of the animal.

      We agree that the low-contrast condition (20%) represents a relatively weak threat signal by design. This level was intentionally chosen to fall within a regime that biases behavior toward freezing and escape after assessment rather than direct escape, allowing us to examine graded decision-making across threat intensities. The higher escape probability observed under low-contrast conditions relative to previous studies (Evans et al., 2018; Fratzl et al., 2021) is likely attributable to the linear arena used here, which generally promoted escape behavior over freezing.

      The social results are particularly interesting in this context as well. Dominant mice consistently prioritize avoidance over reward, showing stronger escape responses and slower habituation than subordinates. This behavior is well captured by the vigilance framework proposed by the authors: dominant animals appear to maintain higher vigilance, which biases decisions toward threat avoidance. The authors further suggest that stable social relationships sustain high vigilance and slow habituation, framing this as an evolutionarily conserved strategy that may enhance survival. This interpretation provides a valuable perspective on how social structure shapes defensive behavior beyond immediate physical interactions. At the same time, there are important limitations to this interpretation. All experiments were conducted in male mice, and it is possible that the relationship between social hierarchy, vigilance, and defensive behavior would differ substantially in females. In addition, the idea that stable social relationships sustain elevated vigilance should be interpreted carefully, as it does not fully align with broader views of social stability as protective against anxiety and stress and generally beneficial for mental health and resilience. These points do not undermine the findings but suggest that the social effects described here should be interpreted with caution and within the specific context of the task and sex studied.

      We thank the reviewer for this important comment and agree that findings obtained in male mice may not necessarily generalize to female mice. We have acknowledged the limitation in the Discussion and note that future studies will be required to determine the extent to which the present findings extend to female mice.

      We would also like to clarify our interpretation regarding vigilance and social buffering. Vigilance should not be conflated with stress or anxiety, and the slower habituation observed in group-housed mice reflects sustained sensitivity to repeated threat exposure rather than elevated anxiety. Thus, our data do not contradict the concept of social buffering; rather, they are consistent with it, as pair-housed mice exhibited reduced defensive responses compared with individually housed animals (Lenz et al., 2022). Furthermore, in an independent study (Li, Gao, and Li, 2026; eLife 15:RP109571), we directly compared responses to looming stimuli when mice were tested alone versus in the presence of a social partner and found clear evidence of social buffering. These findings suggest that social interactions can attenuate defensive responses while prolonging vigilance during repeated threat exposure. We have revised the Discussion accordingly.

      Another important limitation is that the neural mechanisms underlying these effects remain highly speculative. Although the manuscript includes an extensive discussion of candidate circuits, particularly involving the superior colliculus and downstream structures, these interpretations go far beyond the data presented in the study and are not directly supported by experimental evidence within the paper itself. The discussion gives substantial weight to potential circuit mechanisms based primarily on previous literature rather than on findings from the current study. Given the complexity and distributed nature of the circuits likely involved in integrating vigilance, reward, social context, and defensive behavior, the present work is better viewed as providing a strong behavioral framework rather than direct mechanistic insight into the underlying neural substrates. In this context, some references discussing how animals learn to suppress defensive responses to repeated looming threats and the neural mechanisms supporting this process could further strengthen the discussion (Salay et al 2021; Fratzl et al. 2021; Conway et al. 2025; Mederos et al. 2025).

      We agree that the proposed neural mechanisms remain speculative and that the circuits involved in integrating internal state, reward, and social context are likely far more complex. We have revised the manuscript to acknowledge this limitation and have included a discussion of learning-dependent suppression of defensive responses to repeated looming threats and the underlying circuit mechanisms.

      Methodologically, the behavioral paradigm is well suited for studying escape decisions in socially housed animals, and the machine learning based classification of defensive responses is a strength. The computational model provides a useful formalization of how threat level, reward level, and vigilance interact and may be valuable for other laboratories studying escape, approach avoidance, or conflict situations, particularly as a way to classify behavioral outcomes after pose estimation. More generally, the work will be of interest to the neuroethology community for its detailed characterization of escape behavior under naturalistic conditions. At the same time, some statements in the discussion slightly overstate the novelty of the methodological approach. For example, the claim that the study differs from earlier work by using machine learning rather than manual annotation overlooks that several previous studies have already implemented automated or semi-automated strategies to classify looming evoked defensive behaviors beyond manual scoring alone.

      We have revised the manuscript to moderate our claims regarding the novelty of the machine-learning–based classification approach.

      Given the ethological nature of the study and the high inter individual variability reported by the authors, clarity and precision in the methods are especially important for reproducibility. While the revised manuscript addresses many earlier concerns, some aspects remain slightly difficult to follow. For example, the main text states that animals were not water deprived to minimize differences in internal state across conditions, whereas parts of the methods describe experiments in which animals were water deprived. This distinction is not always clearly explained across the different experimental sections, despite internal state being central to the interpretation of the behavioral findings. A clearer separation and description of these conditions would further strengthen confidence in the work. In addition, it was somewhat surprising that the low contrast (20%) looming condition was still sufficient to trigger robust escape responses, and additional clarification or discussion regarding stimulus saliency at this contrast level could help readers better contextualize these findings.

      To improve clarity, we have revised the Methods section to clearly distinguish between experimental conditions that involved water deprivation and those that did not.

      Overall, this study provides a rich analysis of how reward level and social hierarchy modulate defensive behavior through changes in vigilance. It offers a useful conceptual advance for thinking about escape behavior in semi-naturalistic settings and lays a solid foundation for future work aimed at linking these behavioral states to underlying neural circuits.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      - Clarify which experiments involved water deprivation and which did not in methods section.

      We have revised the Methods section to clearly specify which experiments involved water deprivation and which were conducted without water deprivation.

      - Add discussion on why the 20% low contrast stimulus is still sufficient to trigger escape responses.

      We have expanded the Discussion regarding the 20% contrast stimulus to clarify why it was sufficient to elicit escape responses.

      - Tone down statements regarding the novelty of the machine learning based behavioral classification.

      We have revised the manuscript to moderate claims regarding the novelty of the machine-learning–based classification approach.

      - Include references related to learning-dependent suppression of looming responses and circuit mechanisms

      We have incorporated the suggested references and expanded the Discussion of learning-dependent suppression of looming responses and the underlying circuit mechanisms.

      - Moderate the discussion of candidate neural circuits, particularly the SC related interpretations, as these are not directly tested in the study.

      We have revised the discussion of candidate circuit mechanisms.

      - Clarify that the interpretation linking stable social relationships to elevated vigilance may be specific to this ethological context.

      We have revised the Discussion to clarify this interpretation.

    1. eLife Assessment

      This important study addresses a discrepancy between population-level growth laws and single-cell correlations. It shows, for flagellar and synthetic genes in E. coli, that while gene expression of certain genes reduces population-average growth, expression levels positively correlate with growth at the single-cell level. The measurements are convincing, and the proposed mechanism - inheritance of growth factors such as ribosomes during asymmetric division - is consistent with the data.

    2. Reviewer #1 (Public review):

      Summary:

      Garcia-Alcala, Kratz and Cluzel investigate to what extent our understanding of bacterial physiology in bulk experiments can be applied to single-cell observations. They find that intrinsic noise may be powerful enough to even inverse the trends found in the bulk. The authors hypothesize that asymmetric distribution of ribosomes to daughter cells during the cell division plays the dominant role in the intrinsic noise and is able to generate the observed phenomenon. They do not show it directly, but the data and its agreement with the model suffice to support this claim.

      Strengths:

      The experimental part is convincing: the positive correlation between the elongation rate and promoter activity of unnecessary protein is clear, as well as the negative correlation between the mean values while changing the promoter strength. This was demonstrated in both rich and poor media. The causality between the growth rate and the promoter activity was shown using the negative lag time of the cross-correlation function. A simple, reasonable model accounts well for the data. This paper demonstrates an interesting phenomenon and provides a plausible theory for it, advancing our understanding of bacterial physiology on the single-cell level.

      Weaknesses:

      (1) Mean-reversion timescales were assumed to be longer than the simulation time and much longer than the cell cycle time. It is not clear whether the results robust in case mean-reversion timescales become of the order of cell cycle or smaller. If not, is there an argument for such practically infinite reversion timescales?

      (2) It is not easy to understand the simulation part unless one reads Ref. [14]. Is k(t) assumed to follow Eq. (1) from ref. [14]? Is this crucial that the ribosome noise appears only at the division? The ribosome noise strength \sigma_R=0.06 - is it lower or higher than the naively expected binomial division?<br /> Also, more intuitive explanation of the Simpson paradox would help the reader.

      (3) It would be useful for the reader to see the raw data and not only the filtered one to appreciate the measurement noise level.

      (4) Negative lag time of the cross-correlation function is visible, but consider adding statistical test for it.

      (5) Can you make similar cross-correlation plots using the model? Can you infer using it whether the data agrees better with the assumption that ribosomes noise appear only at division or continuous fluctuations during the cell cycle?

      Comments on revised version:

      The authors addressed the five comments listed above.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by Garcia-Alcala et al. reports an interesting paradox: the cost of gene expression slows the population-average growth rate, whereas at the single-cell level, expression levels from these genes positively correlate with the growth rate. The effect is observed in the expression of flagellar genes and a gene under a synthetic promoter in E. coli. The findings are explained by the inheritance of growth factors, including ribosomes, during asymmetric division.

      Strengths:

      (1) The manuscript adds strength to an emerging body of literature showing that the population-level bacterial growth laws do not match correlations based on single-cell data. The evidence presented here is more striking than in previous works (such as Pavlou et al., Nat. Commun. 2025), as the trends in population-level data and single-cell data are reversed.

      (2) A relatively simple model correctly explains the trends in the data.

      Weaknesses:

      (1) The differing behavior of the MG1655 and MC4100 strains remains a lingering question concerning the generality of the conclusions. It appears unlikely that ribosomes or other growth factors partition significantly differently in the MC4100 strain than in the MG1655 strain. Furthermore, based on Fig. S15, it is still unclear to what extent MC4100 exhibits growth-rate fluctuations, as stated in the text, rather than primarily size fluctuations, as shown in the figure. It is also unclear why such very slow fluctuations would lead to qualitatively different behavior, given that the proposed mechanism appears to be rather fundamental. It would be helpful for the authors to discuss these two points.

      (2) It is unclear what fraction of the total proteome mVenus represents in different measurements. Adding this information would strengthen the conclusions.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Garcia-Alcala, Kratz and Cluzel investigate to what extent our understanding of bacterial physiology in bulk experiments can be applied to single-cell observations. They find that intrinsic noise may be powerful enough to even inverse the trends found in the bulk. The authors hypothesize that the asymmetric distribution of ribosomes to daughter cells during cell division plays the dominant role in the intrinsic noise and is able to generate the observed phenomenon. They do not show it directly, but the data and its agreement with the model are sufficient to support this claim.

      Strengths:

      The experimental part is convincing: the positive correlation between the elongation rate and promoter activity of unnecessary protein is clear, as well as the negative correlation between the mean values while changing the promoter strength. This was demonstrated in both rich and poor media. The causality between the growth rate and the promoter activity was shown using the negative lag time of the cross-correlation function. A simple, reasonable model accounts well for the data. This paper demonstrates an interesting phenomenon and provides a plausible theory for it, advancing our understanding of bacterial physiology on the single-cell level.

      Weaknesses:

      (1) Mean-reversion timescales were assumed to be longer than the simulation time and much longer than the cell cycle time. It is not clear whether the results are robust in case mean reversion timescales become of the order of the cell-cycle or smaller. If not, is there an argument for such practically infinite reversion timescales?

      Due to an error in the simulation code, the incorrect mean-reversion timescales were reported in the manuscript and have now been corrected. Instead of 1000 h and 100 h for 𝜏<sub>R</sub> and 𝜏<sub>U</sub>, respectively, they are 6.25 h and 0.0625 h, which are both much shorter than the total simulation time (75 h). The error had no effect on simulation results and was purely a timescale conversion mistake. We appreciate the referee’s careful review of the manuscript which allowed us to catch this error.

      Given the correct values, 𝜏<sub>U</sub> is well within the cell-cycle time, as is expected since we assume the source of noise in unnecessary protein expression is from stochastic gene expression. In contrast, 𝜏<sub>U</sub> is still significantly longer than the cell-cycle time and needs to be in order to generate the experimentally observed behavior (i.e., the increase in growth rate with an increase in unnecessary protein expression at the single-cell level). This is consistent with our biological hypothesis that some cells inherit ribosomal surpluses from their mothers which enable bursts in protein production. 𝜏<sub>U</sub> less than the cell-cycle time would correspond to a case where ribosomal composition quickly decays to the population average, meaning that daughter cells would never have time to capitalize on the benefit of receiving a ribosome surplus. Furthermore, 𝜏<sub>U</sub> being longer than the cell-cycle is biologically justifiable as proteins such as ribosomes are passed from mother to daughter at division, thus allowing for memory to persist over longer timescales than a single generation.

      (2) It is not easy to understand the simulation part unless one reads Ref [14]. k(t) is assumed Equation (1) from Reference [14]? Is it crucial that the ribosome noise appears only at the division? The ribosome noise strength σ<sub>R</sub> =0.06 - is it lower or higher than the naively expected binomial division? Also, a more intuitive explanation of the Simpson paradox would help the reader.

      𝜅(𝑡) is indeed Eq. (1) from Ref. [14]. To make the computational results clearer, the methods section has been updated to include the full set of equations used to perform the simulations, and the code used to produce the computational figures is now on GitHub. The relative standard deviation expected by modeling ribosome distribution at division by a binomial distribution with equal probability of being inherited by either daughter cell is in the range of 1-3% (0.01-0.03), assuming N~10<sup>3</sup>-10<sup>4</sup> ribosomes. Thus, 0.06 is reasonable as it is in the same order of magnitude as what is predicted by binomial division. Furthermore, if clustering is present as already demonstrated in [39, 40, 43], we would expect the relative standard deviation to increase as clustering reduces the effective number of proteins which are distributed between daughters.

      (3) It would be useful for the reader to see the raw data and not only the filtered one to appreciate the measurement noise level.

      We have included the direct calculations of cell size and promoter activity in the time-lapse plots in Fig. 1. In addition, we included Fig. S2, which shows typical time traces of Class-2 activity and elongation rate, displaying both the direct measurements and the corresponding smoothed traces for the flagellar reporter strains.

      (4) Negative lag time of the cross-correlation function is visible, but consider adding a statistical test for it.

      We have included a statistical analysis of the lags of maximum correlation of elongation rate and activity for the flagellar reporter strains in Fig. S10. For each strain, we now show an overlay of the cross-correlation functions for all lineages and the distribution of the lag corresponding to the maximum correlation. In addition, we have added a section in Materials and Methods describing in detail how the cross-correlation and the strain-averaged lag were computed.

      (5) Can you make similar cross-correlation plots using the model? Can you infer by using it, whether the data agrees better with the assumption that ribosomal noise appears only at division or continuous fluctuations during the cell cycle?

      The model in its current form is unable to capture the observed cross-correlation (the correlation is sharply peaked at zero). This is because the model coarse-grains transcription and translation into one process of protein production and thus lacks any delay or memory mechanisms which would make a significant positive or negative correlation.

      Both continuous fluctuations and noise from division could in principle contribute to our observation of Simpson’s paradox. We are unable to use the model in its present form to dissect the contributions from both mechanisms. However, our model simulations clearly show that noise at division by itself is sufficient to explain the observed effect with parameter values that are biologically plausible. As Chao et al., [40] has previously shown experimentally that a significant source of ribosomal noise comes from unequal distribution at division, which is the hypothesis we retained for the model.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Garcia-Alcala et al. reports an interesting paradox: the cost of gene expression slows the population-average growth rate, whereas at the single-cell level, expression levels from these genes positively correlate with the growth rate. The effect is observed in the expression of flagellar genes and a gene under a synthetic promoter in E. coli. The findings are explained by the inheritance of growth factors, including ribosomes, during asymmetric division.

      Strengths:

      (1) The manuscript adds strength to an emerging body of literature showing that the population-level bacterial growth laws do not match correlations based on single-cell data. The evidence presented here is more striking than in previous works (such as Pavlou et al., Nat. Commun. 2025), as the trends in population-level data and single-cell data are reversed.

      (2) A relatively simple model correctly explains the trends in the data.

      Weaknesses:

      (1) It is not clear whether flagellar proteins are expressed proportionally to the reporter signal. Furthermore, it is questionable if E. coli bacteria in the mother machine channels are flagellated. If they are, they could potentially swim out of the channels, which is not the case when they do not carry the MotA E98K mutation. The authors should provide some evidence that E. coli expresses the actual filament proteins in the channels.

      We agree that it is important to demonstrate that our reporter reflects the production of functional flagellar structures under our experimental conditions. To this end, we first tested the swimming capabilities of our strains before introducing the MotA E98K mutation, using a standard soft‑agar motility assay. For strains in which only the Class‑1 promoter was modified (Pro2, Pro4, Pro5), after 12 h of inoculation, the diameters of the swarming rings followed the expected order based on promoter strength: Pro2 showed the smallest ring, followed by Pro4, WT, and Pro5. In contrast, control strains carrying Pro4 together with either MotA E98K (non‑rotating motors) or ΔfliC (no flagellin filament) did not form rings, consistent with their inability to swim. We also included an MG1655 strain carrying an IS5 insertion upstream of the Class‑1 promoter, which is known to enhance flagellar expression [56]; this strain showed a larger ring, as expected. A qualitative summary of ring diameters for all strains is provided in Author response table 1. We have included the figure and section “Experimental validation of functional flagella expression in reporter strains” on Supplementary Material.

      Author response table 1.

      We also tested whether cells assemble functional flagella on the mother‑machine. We compared MG1655 WT and MG1655 carrying the MotA E98K mutation under identical microfluidic growth conditions. Many WT cells left the channels during the experiment (∼35% of lineages over ~20 h after the onset of exponential growth inside the device), consistent with active swimming, whereas the non‑motile MotA E98K strain did not leave the channels. Because the only difference between these two strains is the MotA E98K mutation, which disables motor rotation but not flagellar assembly, this result indicates that (i) cells do express functional filaments and motors in the mother‑machine environment, and (ii) the strain used for our main experiments is immobilized by MotA E98K.

      Both experiments are included in the Supplementary Information section “Experimental validation of functional flagella expression in reporter strains” and Fig. S3.

      (2) It is unclear what fraction of the total proteome mVenus represents in different measurements. Some quantification is needed (for example, using the Coomassie staining). Using f_U as high as 14.4% in simulations is questionable.

      We agree that we do not currently know the exact fraction of the proteome occupied by mVenus in our experiments, as we only quantified fluorescence and did not perform Coomassie staining or proteomics. The primary goal of our simulations is not to reproduce exact experimental conditions, but to illustrate how unequal ribosome partitioning affects daughter cells across a range of protein synthesis burdens. Consequently, our use of values up to F<sub>U</sub> = 14.4% in the simulations was intended as an exploratory upper range rather than as a direct estimate of the experimental condition.

      To put our experimental burden in context, we compare our growth-rate reduction to a well– characterized high-burden case in the literature. In the study by T. Hwa’s group [6], overexpression of β–galactosidase such that it constituted approximately 27% of the total proteome led to a 67% reduction in growth rate for E. coli growing in a medium similar to ours (differing only in the carbon source: glucose in their case, glycerol in ours). In our system, overexpression of mVenus from a plasmid result in a substantially smaller growth–rate reduction of 9% relative to the non–expressing control, and this is observed on the poorer carbon source (glycerol), under which burden effects are typically less pronounced [6].

      Although we cannot convert our fluorescence measurements into an exact proteome fraction, the much smaller growth defect compared to the 27% β‑galactosidase case strongly suggests that mVenus does not approach such an extreme fraction of the proteome under our conditions. Under the simplifying assumption that the qualitative relationship between unnecessary‑protein fraction and growth‑rate reduction is similar in the two systems, our data are therefore consistent with a modest fraction of unnecessary protein and make it unlikely that mVenus reaches very high fractions such as 27%. This gives us confidence that exploring F<sub>U</sub> values up to 14.4% in the simulations represents a conservative upper range relative to our experimental burden, rather than an underestimate.

      (3) The data from the MC4100 strain does not directly match the trends of MG1655. The justification for filtering out the low-frequency components of MC4100 is not particularly convincing. It appears unlikely that ribosomes or other growth factors partition significantly differently in the MC4100 strain than in the MG1655 strain. Further discussion and a plot similar to Figure 1 (Left) for this strain are warranted.

      We thank the reviewer for this helpful suggestion. We agree that the filtering analysis from the initial manuscript was not clear enough to be used as robust supplementary information, and we therefore removed it. Instead, we carried out with the ‘unfiltered’ MC4100 strain, the same single-cell analyses as with MG1655, including the binning analysis of instantaneous elongation rate versus Class-2 promoter activity, and a cross-correlation analysis between such measurements.

      Both analyses did not reveal a significant correlation between growth and flagellar promoter activity in MC4100. We now present these results in the revised manuscript (Fig. S15), where we explicitly show the lack of association in a plot directly comparable to Fig. 1.

      While ribosomes partition mechanism is likely to be the same between these two strains, MC4100 is known to exhibit very long oscillations of growth rate over 10 generations, which are absent in MG1655.

      We now present MC4100 as an explicit counterexample to highlight that the positive correlation between short timescale growth fluctuations and flagellar expression observed in MG1655 is not universal across all E. coli strains, especially in strains like MC4100 whose growth rate fluctuations are dominated by long timescales, much longer than the division time. In MC4100 these slow modes are largely decoupled from flagellar gene regulation. We have also revised the text to clarify this point.

      (4) The model needs to be described in more detail. A closed set of equations that have been simulated must be presented, along with all values of the model parameters and their sources. The authors should consider depositing their code on GitHub or another publicly accessible repository.

      The methods section has now been updated to include the full set of equations used to perform the simulations along with all parameter values and their sources. Additionally, the code used to produce the computational figures is now on GitHub. For a full derivation and biological justification of each model component, we still refer readers to ref [14] where this model was first published.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The first paragraph of Results still belongs to the Introduction Section, I feel.

      Thank you for the recommendation, we agreed with the change.

      (2) Figure 1, left, is after some filter (Savitzky-Golay)? It might be useful to see the raw points.

      We have included the direct calculations of cell size and promoter activity in the time-lapse plots in Fig. 1. In addition, we included Fig. S2, which shows typical time traces of Class-2 activity and elongation rate, displaying both the raw (direct) measurements and the corresponding smoothed traces for the flagellar reporter strains.

      Reviewer #2 (Recommendations for the authors):

      (1) Is there a statistically significant positive slope in single-cell data of the elongation rate as a function of Class-2 activity? There should be an analysis of statistical significance for the slopes in Figure 1 (Right) and Figure 4A, including both population-average and single-cell data.

      Thank you for your recommendation. We have now added statistical analyses of the correlations between activity and elongation rate at both the single-cell and population levels. The results for the flagellar strains are presented in Fig. S5, S6, and S7. Also, the corresponding analysis for constitutive Venus expression is shown in Fig. S16.

      (2) How does FlhC-YFP data compare with the CFP signal in the first set of measurements? What would Figure 1, Left look like for this signal?

      At the single‑cell level, the relationship between Class‑1 promoter activity and elongation rate within each strain is still positive, but clearly weaker than for Class‑2, as shown in Fig. S8. When we bin the data, we can observe that the binned averages don’t display a clear positive trend, even though the Pearson correlation for each flagellar reporter strain is still positive. We think this is because Class‑1 controls a much smaller part of the proteome (it only encodes the two subunits of FlhDC) while Class‑2 promoters drive many structural and export proteins. So, changes in Class‑2 activity more directly reflect shifts in global translational capacity and are more tightly linked to growth, whereas Class‑1 activity adds only a small translational load.

      (3) The information from the ER-Activity cross-correlation functions is interesting but has not been interpreted or compared with the model. Which signal precedes the other? What can explain the observed lag time on the order of Tdiv? Why do some cross-correlations show negative values in Fig. S9 while others are positive (as they should)? Can the model explain the experimentally observed cross-correlation function?

      In our cross‑correlation analysis, a negative lag at the maximum means that fluctuations in elongation rate (ER) precede fluctuations in promoter activity (A) by that lag. We now describe this explicitly in Materials and Methods and quantify lag distributions for all reporter strains in Fig. S10.

      In Fig. 13 (previously Fig. 9), we mainly observe positive correlations between ER and PA with negative or near zero peak lags, i.e., ER tends to lead PA by a lag by about a division time. The reason governing the negative lags is not immediately clear, but it is consistent with a resource driven mechanism: cells that inherit more growth factors (e.g. ribosomes) at division, use them to prioritize housekeeping processes and thus increase growth rate first, and only subsequently use these extra resources to increase flagellar promoter activity.

      As for Class–1 activity, the correlation with growth is much weaker as it was already observed in Kim et al. Averaging the cross–correlations across all lineages yields a modest peak (mean correlation 0.10, SD 0.12; see Author response image 1) at a small positive lag of 0.5 h (while average division time is T<sub>div</sub> ~ 1.7h). This indicates a weak but real positive correlation between Class 1 activity and ER at short lags. However, the lags of the individual maxima are widely distributed (SD ≈ 9 h; mean −1.8 h, median −0.28 h, mode ≈ 0), with only ~53% of the cells showing a negative time lag. Thus, delays are roughly symmetrically spread around zero with only a slight negative bias.

      Author response image 1.

      Left: cross−correlation between elongation rate and Class−1 activity for each lineage (N = 96, colored lines), and their average as a function of lag (black line). Right: distribution of the lags at which each lineage’s cross−correlation attains its maximum. The dashed line indicates zero lag, and the full line marks the lag of the peak of the mean cross correlation.

      An in-depth analysis about the sign and magnitude of the lag would require measurements from a broader set of promoters. The lag is likely to depend on what genes the promoter controls (e.g. stress response, housekeeping, or large structural modules), on its strength and regulation, and on growth conditions.

      The model in its current form is unable to capture the observed cross-correlation (the correlation is sharply peaked at zero). This is because the model coarse-grains transcription and translation into one single process of protein production and thus lacks any delay or memory mechanism that could generate a phase shift.

      Nonetheless, this memory-free formulation shows that stochastic redistribution of growth factors at division is by itself sufficient to generate the Simpson’s paradox behavior. Capturing the experimentally observed lag time would require including explicit transcription/translation delays and additional regulatory dynamics between housekeeping and flagellar genes whose information we do not have.

      (4) Figure 3, Left - it is not clear what this plot shows. Red and blue are scattered over the whole plot. How have daughter 1 and daughter 2 been assigned? Perhaps choosing one of the daughters with a higher growth rate and then plotting the data could reveal some trends.

      We agree that the original left panel of Fig. 3 was difficult to follow. In the original version, the daughter labels were assigned as “Daughter–1” for the cell at the closed end of the channel and “Daughter–2” for the cell closest to the open side. After discussing this with Dr Camilla Ulla Rang (whom we now acknowledge in the manuscript), we relabeled the daughters as “new–pole daughter” and “old–pole daughter,” following the convention used in studies of aging and ribosome distribution in E. coli.

      We now show the class–2 promoter activity comparison between new pole vs old pole, and in a separate panel, we plot the mean ratios of elongation rate, and class–1/class–2 promoter activities between the new–pole and old–pole daughters. These ratios show that the new–pole daughter tends to have both greater growth rate and flagellar gene activity than the old–pole daughter. Importantly, this new plot is in line with Rang and colleagues who demonstrated that new–pole daughters have greater ribosome density and faster growth rates than their old–pole sisters. Together, these results further support our hypothesis that excesses of ribosomes inherited at division underlies the observed growth boosts.

      (5) Flagellar activity -> activity of flagellar gene synthesis (presumably no flagellar activity in these cells).

      We corrected the terms used to reference the flagellar gene activity.

      (6) Page 7: "The strain MC4100, known to exhibit slow, long period oscillations in growth ..." - some reference is- needed here.

      We placed the reference some lines after, as such paper also includes the information of MG1655 short-term oscillations. Now the reference is [44]: Tanouchi, Y., et al., A noisy linear map underlies oscillations in cell size and gene expression in bacteria. Nature, 2015

      (7) Page 7: Figure S11B - is Figure S11C perhaps meant?

      The correct panel indeed was Fig. S11C, and we have now corrected and updated the figure label and text accordingly.

      (8) Page 11: Savitsky-Golay filter.

      We have corrected it, thank you.

    1. eLife Assessment

      This fundamental study demonstrates that polarized second-harmonic generation microscopy can be used to probe the ON/OFF states of myosin in both permeabilized and intact muscle, making this key measurement accessible to a greater number of labs. This has the potential to help with the study of disease-causing mutations and our understanding of drug function. The methodology is well defined, and the results are convincing; however, there are some limitations to the interpretation of the data.

    2. Reviewer #2 (Public review):

      Summary:

      In striated muscle, myosin motors can dynamically switch between an energy-conserving OFF state and an activated-ON state. This switching is important for meeting the body's needs under different physiological conditions, and previous studies have shown that disease causing mutations associated with cardiomyopathies can affect the population of these states, leading to aberrant contractility. Studying these structural states in muscle has previously only been possible via X-ray diffraction which requires access to a beam line. Here, Arecchi et al. demonstrate that polarized second-harmonic generation microscopy (pSGH), a technique that is more accessible, can be used to probe the ON/OFF states of myosin in both permeabilized and intact muscle.

      Comments on revised version:

      The manuscript has been significantly strengthened in the revision. The authors have addressed my concerns.

    3. Reviewer #3 (Public review):

      Summary:

      This is a very interesting paper extending the use of SHG to the study of relaxed muscle and its use to assess the order- disorder (and on /off) states of myosin heads in the thick filament. The work convincingly shows that SHG, and the parameter gamma, provide a reliable measure of the state of the myosin heads in a range of different relaxed muscle fibres, both intact and skinned and in myofibrils. In mini pig cardiac fibres the use of dATP and mavacamten increased or decreased the number of heads in the disordered state respectively. On the assumption that these treatments push myosins fully into the disordered or ordered state then this allows the fraction of ordered heads to be assessed under a wide variety of conditions. The extension of this part of the study to mouse heart and rabbit psoas samples extends the validation of the approach.

      The results with the myosin mutant R403Q support the idea that this mutation reduces the fraction of myosin heads in the ordered state and that mavacamten can recover the WT situation.

      The results from SHG were compared with parallel studies using X-rays to validate the conclusions. Independent fibre ATPase data further support the conclusions.

      The work is solid and provides a novel approach assessing the activity state of muscle thick filaments. The authors point out some of the potential uses of this approach in the future including time resolved SHG measurements. Indeed, jumps in mavacamten or dATP concentration with time resolved SHG could measure the rates of entry and exit from the ordered , off state of the filament. A measurement urgently needed in the field.

      Strengths:

      (1) The SHG signal is convincingly shown to assess the fraction of ordered/disordered myosin heads in the thick filament of a variety of muscle fibres.

      (2) The results are similar for rabbit psoas, mouse and minipig cardiac fibres, Skinning the fibres and production of myofibrils does not change the SHG signal.

      (3) Use of myosin R403Q mutant in mini pig confirms a loss of ordered myosin heads and the ordered heads can be recovered by mavacamten.

      (4) Parallel X-ray scattering and ATPase data support the conclusions.

      (5) Assuming that dATP and mavacamten generate 100% disordered vs ordered myosin heads respectively then the % ordered heads can be calculated for a variety of conditions.

      (6) The potential of extending the technique with time resolved studies and sub sarcomere variations of SHG are very exciting prospects.

      Weaknesses:

      Issues like the effect of fibre disarray on the SHG signal are not well defined.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study utilizes polarized second-harmonic generation (pSHG) microscopy to investigate myosin conformation in the relaxed state, distinguishing between the disordered, actin-accessible ON state and the ordered, energy-conserving OFF state. By pharmacologically modulating the ON/OFF equilibrium with a myosin activator (2deoxyATP) and inhibitor (Mavacamten), the authors demonstrate that pSHG can sensitively quantify the ON/OFF ratio in both skeletal and cardiac muscle. Validation with X-ray diffraction supports the accuracy of the method. Applying this approach to a hypertrophic cardiomyopathy model, the study shows that R403Q/MYH7-mutated minipigs exhibit an increased ON state fraction relative to controls. This difference is eliminated under saturating concentrations of myosin modulators, indicating that the ON/OFF balance can be pharmacologically shifted to its extremes. Additionally, ATPase assays reveal elevated resting ATPase activity in R403Q samples, which persists even when the ON state is saturated, suggesting that increased energy consumption in this mutation is driven by both a shift toward the ON state and inherently higher myosin ATPase activity.

      Strengths:

      This is a well-written and well-conducted study that clearly reveals the power of SHG microscopy. The study clearly establishes the great utility of SHG to study thick filament regulation.

      Weaknesses:

      (1) Several studies have shown that the ON state of the thick filament is sensitive to both temperature and filament lattice spacing, with a common recommendation to conduct skinned fiber experiments at temperatures above 27 °C and in the presence of dextran to better preserve physiological conditions. The authors should clarify the experimental temperature used in their skinned fiber studies, indicate whether dextran was included, and discuss whether adherence to these recommended conditions would have impacted their results.

      Additional experiments were performed on skinned psoas muscle strips to assess the influence of temperature and lattice spacing on the measured γ values. In particular, psoas strips were tested under relaxing conditions at three different temperatures (10 °C, 15 °C, and 20 °C). In a separate set of experiments, psoas strips were examined under relaxing conditions in the absence and in the presence of 5% dextran.

      These measurements confirmed previous reports indicating that the ON state of the thick filament is sensitive to both temperature and lattice compression. Specifically, increasing the temperature from 10 °C to 20 °C resulted in a shift toward lower γ values, consistent with a greater radial displacement of myosin heads. A similar trend, although less pronounced, was observed in the presence of dextran, which also produced a decrease in γ compared with control conditions.

      Figure 1 and text have been modified including these new findings (revised manuscript, pages 11-12).

      (2) On page 13, the authors report the proportion of disordered heads as approximately 30% in wild-type and 65% in R403Q fibers. They should clarify whether these values represent the percentage of total myosin heads, or rather the percentage of heads that are responsive to Mavacamten and dATP.

      The proportions reported in the original version referred to the fraction of myosin heads assumed to be responsive to Mavacamten and dATP. These estimates were obtained using a simplified model in which dATP was assumed to produce a complete (100%) shift of myosin heads toward the ON state, whereas Mavacamten was assumed to cause a complete depletion (0%).

      Since these assumptions are not fully supported by structural data, we decided to remove these values from the Results section and instead discuss these considerations in more general terms in the Discussion (page 19). Figure 3 and text have been modified accordingly (pages 14-15).

      (3) In Figure 5, regarding ATPase measurements, the content of contractile material per unit volume of muscle preparation will influence the results. Did the authors account for this variable, and if not, how might it have affected the conclusions?

      We thank the reviewers for raising this concern and agree that the content of contractile material per unit volume can influence the absolute values obtained in ATPase measurements. In our experiments, however, we primarily assessed the effects of compounds using paired measurements (e.g., the same strip measured before and after Mavacamten treatment).

      In addition, based on our previous structural data obtained from similar preparations of human myocardium (https://doi.org/10.1161/CIRCRESAHA.122.321956), the ratio of contractile tissue volume to total muscle volume is highly preserved. This supports comparisons between muscle strips from different experimental groups (e.g., WT and R403Q).

      (4) For readers primarily interested in assessing the ON/OFF state of thick filaments, could the authors list the specific advantages of polarized second harmonic generation (pSHG) microscopy compared to X-ray diffraction?

      In the present work, we deliberately chose to frame pSHG microscopy and X-ray diffraction as complementary approaches, rather than emphasizing a direct one-to-one comparison of their respective advantages and disadvantages.

      That said, we agree that for readers specifically interested in assessing the ON/OFF state of thick filaments, a clearer delineation of the respective strengths and limitations is useful. We have therefore slightly revised the Discussion (page 21) to better clarify the advantages and constraints of both approaches.

      (5) Given that many data points were derived from the same fiber or myocyte, how did the authors address the risk of type I errors due to non-independence of measurements? Was a nested or hierarchical statistical approach used?

      We thank the reviewer for raising this important point and agree that the original statistical analysis did not fully account for the non-independence of measurements derived from the same fibre.

      To address this issue, we have now reanalyzed the entire dataset with the support of Prof. Francesco Sera, who has been included as a co-author. A hierarchical (mixed-effects) statistical model was applied, explicitly accounting for the nested structure of the data, repeated measurements, and unbalanced group sizes.

      Importantly, this revised analysis substantially confirms the original results, with no major changes in the observed trends or in the statistical significance of the findings. All figures have been modified accordingly, and the statistical methods used are now described in the Methods section (page 8).

      Reviewer #2 (Public review):

      Summary:

      In striated muscle, myosin motors can dynamically switch between an energyconserving OFF state and an activated ON state. This switching is important for meeting the body's needs under different physiological conditions, and previous studies have shown that disease-causing mutations associated with cardiomyopathies can affect the population of these states, leading to aberrant contractility. Studying these structural states in muscle has previously only been possible via X-ray diffraction, which requires access to a beam line. Here, Arecchi et al. demonstrate that polarized second-harmonic generation microscopy (pSGH), a technique that is more accessible, can be used to probe the ON/OFF states of myosin in both permeabilized and intact muscle.

      Strengths:

      (1) There is an outstanding need in the field to better understand the regulation of the ON/OFF states of myosin. Currently, this is studied using X-ray diffraction, meaning that it is accessible to only a few labs. The authors demonstrate that pSGH can be used to probe the ON/OFF states of myosin both in intact and permeabilized muscle. This is a significant advance, since it makes it possible to study these states in a standard research laboratory.

      (2) The authors demonstrate that this approach can be employed in both skeletal and cardiac muscle. Importantly, it works with both porcine and mouse cardiac muscle, which are two of the most important animal models for preclinical studies.

      (3) The authors manipulate the ON/OFF equilibrium using both drugs and a genetic model of hypertrophic cardiomyopathy that has been shown to modulate the ON/OFF equilibrium. Their results generally agree with previous studies conducted using X-ray diffraction as well as biochemical measurements of myosin autoinhibition.

      Weaknesses:

      (1) While the application of pSGH to the ON/OFF equilibrium is an important advance, there are limited new biological insights since the perturbations used here have been extensively characterized in previous studies.

      We acknowledge that the biological insights provided in this study largely confirm previous findings reported in the literature. However, the primary aim of our work was to demonstrate the applicability and robustness of pSHG in probing the ON/OFF equilibrium of thick filaments. In this context, obtaining results that are consistent with established knowledge represents an important validation of the technique.

      We therefore believe that the combination of these findings with the advantages of pSHG makes our work noteworthy and supports the broader utility of this approach in future biological and physiological studies.

      (2) SGH has previously been applied to study the nucleotide-dependent orientation of myosin motors in the sarcomere (PMID: 20385845). The authors have previously interpreted the value of gamma as being a readout of lever arm position, but here, it is interpreted as a measure of ON/OFF equilibrium. When this technique is applied to intact muscle, it is not clear how to deconvolve the contributions of lever arm angle from the ON/OFF population (especially where there is a mix of states that give rise to the gamma value). This is an important limitation that is not discussed in the manuscript.

      We thank the reviewer for this insightful and important observation. The SHG signal arises from the coherent contribution of peptide bonds within the myosin molecule and therefore reflects an average angular distribution rather than a single structural state. In the study cited (PMID: 20385845), by reconstructing each contributing second-harmonic emitter within the actomyosin motor array at the atomic scale, we were able to establish a direct relationship between myosin conformation and the γ value. In particular, this analysis revealed an overall trend: γ increases as the average angle of myosin relative to the thick filament backbone becomes larger, and decreases as this angle is reduced. Based on these observations, we proposed that γ could serve as a proxy sensitive to changes in the ON/OFF equilibrium.

      However, we fully acknowledge that, especially in intact muscle, it is not possible to disentangle the respective contributions of lever arm orientation and population shifts between different structural states. The reviewer’s point is therefore well taken.

      To address this, we have revised the Discussion (pages 18-19) to more clearly acknowledge this limitation and to better articulate the interpretative framework underlying our analysis.

      Finally, we would like to emphasize that, in the present study, we aimed to minimize the contribution of actomyosin-bound states in intact trabeculae by performing experiments under conditions expected to strongly favor relaxation (low stimulation frequency, 0.1 Hz, and low temperature, 21 °C). Under these conditions, the number of strongly bound actomyosin cross-bridges during the diastolic phase is expected to be minimal, and variations in γ are therefore primarily associated with changes in the relaxed myosin population.

      (3) The R403Q mutation has previously been shown to cause an increase in ATP usage. Here, the authors measure an elevated basal ATPase rate under relaxing conditions, and they interpret this as showing increased myosin ATPase activity intrinsic to the motors; however, care should be used in interpreting these results. Work from the Spudich lab has shown that the R403Q mutation can appear as increasing motor function in some assays but depressing motor function in others (see PMID: 32284968, 26601291). Moreover, the actin-activated ATPase rate is an order of magnitude higher than the basal ATPase rate, and thus, small changes in the basal ATPase rate are unlikely to be important for physiology.

      We thank the reviewer for this thoughtful comment, and for pointing out the complexity of interpreting the functional consequences of the R403Q mutation. We agree that the functional effects of the R403Q mutation on myosin motor activity have been reported to vary depending on the experimental system and assay used, with studies showing both reduced and enhanced motor performance (PMID: 32284968, 26601291, 20560002).

      Consistent with this complexity, previous biochemical and physiological studies have reported altered energetic properties associated with the R403Q mutation in cardiac muscle. In a previous study, the relationship between cross-bridge kinetics and energetics was investigated in single cardiac myofibrils and multicellular cardiac muscle strips from human HCM samples with and without the R403Q mutation. In those experiments, cross-bridge relaxation was faster in R403Q samples and correlated with an increased energetic cost of tension generation. Basal ATPase activity measured in human samples was also elevated (4.4 ± 0.5 vs 6.6 ± 1.2 μmol L<sup>-1</sup> s<sup>-1</sup> in HCMsn and R403Q, respectively), supporting the idea that the mutation affects energetic balance in the relaxed state (PMID: 24928957).

      We also agree that actin-activated ATPase activity is substantially higher than basal ATPase activity. However, cardiac muscle spends a large fraction of the cardiac cycle in the relaxed (diastolic) state, during which myosin heads are predominantly detached from actin. Even relatively small changes in ATP turnover during this phase could therefore influence the overall myocardial energetics, and potentially contribute to the activation of signaling pathways involved in pathological remodelling.

      To address the reviewer’s concern, we have now expanded the Discussion (pages 19-20) to clarify the heterogeneous results reported in the literature for the R403Q mutation, and to more cautiously interpret the physiological implications of the observed resting ATPase activity.

      (4) The authors interpret some of their data based on the assumption that the high concentrations of drugs cause the myosin to either adopt 100% OFF or ON states. This assumption is not validated, limiting the ability to interpret the fraction of myosins in the ON/OFF states.

      We fully agree: these assumptions are not fully supported by structural data. We consistently decided to remove these analysis from the Results section and instead discuss these considerations in more general terms in the Discussion (page 18). Figure 3 and text have been modified accordingly (pages 14-15).

      (5) The ATPase measurements are innovative but hard to interpret. dATP and ATP do not have identical ATPase kinetics, meaning that it is hard to deconvolve whether the elevated ATPase rate with dATP is due to changes in the ON/OFF population and/or intrinsic ATPase activity. Similarly, mavacamten reduces the rate of phosphate release from myosin, and this effect is not strictly coupled to the formation of the OFF state (e.g., see PMID: 40118457). As such, it is difficult to deconvolve drug-based changes in the inherent ATPase kinetics of the myosin from changes in the OFF-state population.

      We thank the reviewer for this important comment. We recognize that dATP and ATP do not exhibit identical ATPase kinetics, and that Mavacamten slows steps in myosin nucleotide release that are well documented for ATP and not for dATP (for both S1; PMID: 28808052 and, more markedly, HMM; PMID: 30018063). At the same time, both compounds have been shown to perturb the regulatory state of myosin heads along the thick filament. These effects are, in turn, mediated by shifts in the distribution of ATPase intermediate states, which alter the likelihood of myosin adopting autoinhibited conformations and thereby biasing the system toward OFF or ON states (see, e.g., PMID: 39444161; PMID: 30018063). This configures a dual mechanism of action for small molecules, which we now describe more clearly in the Discussion (page 20). As suggested by the referee, we also highlight the intrinsic difficulty of disentangling compound-induced changes in the intrinsic ATPase kinetics of myosin from shifts in the population of the OFF state. Consequently, we have now better focused results and discussion section considering mainly differential effects of the drugs on SHG and ATPase measurements.

      Moreover, under the strongly relaxing conditions used in our experiments (pCa 10), the population of actomyosin-bound cross-bridges is expected to be extremely small (<0.1%), thereby minimizing any dATP-mediated activation of myosin arising from enhanced electrostatic interactions with actin (see, e.g., PMID: 31110001). This is supported by the absence of a significant effect of the nucleotide on the resting tension of myofibrils (new Fig. S1).

      We have now better clarified these points in both result and discussion section of the revised manuscript.

      Reviewer #3 (Public review):

      This is a very interesting paper extending the use of SHG to the study of relaxed muscle and its use to assess the order-disorder (and on /off) states of myosin heads in the thick filament. The work convincingly shows that SHG and the parameter gamma provide a reliable measure of the state of the myosin heads in a range of different relaxed muscle fibres, both intact and skinned, and in myofibrils. In mini pig cardiac fibres, the use of dATP and mavacamten increased or decreased the number of heads in the disordered state, respectively. On the assumption that these treatments push myosins fully into the disordered or ordered state, then this allows the fraction of ordered heads to be assessed under a wide variety of conditions. It is unfortunate that dATP treatment was not used (as mavacmten was) on rabbit psoas and mouse samples to further test this hypothesis.

      The results with the myosin mutant R403Q support the idea that this mutation reduces the fraction of myosin heads in the ordered state and that mavacamten can recover the WT situation.

      The results from SHG were compared with parallel studies using X-rays to validate the conclusions. Independent fibre ATPase data further support the conclusions.

      The work is solid and provides a novel approach to assessing the activity state of muscle thick filaments. The authors point out some of the potential uses of this approach in the future, including time-resolved SHG measurements. Indeed, jumps in mavacamten or dATP concentration with time-resolved SHG could measure the rates of entry and exit from the ordered, off state of the filament. A measurement is urgently needed in the field.

      Strengths:

      (1) The SHG signal is convincingly shown to assess the fraction of ordered/disordered myosin heads in the thick filament of a variety of muscle fibres.

      (2) The results are similar for rabbit psoas, mouse, and minipig cardiac fibres. Skinning the fibres and production of myofibrils do not change the SHG signal.

      (3) Use of myosin R403Q mutant in mini pig confirms a loss of ordered myosin heads, and the ordered heads can be recovered by mavacamten.

      (4) Parallel X-ray scattering and ATPase data support the conclusions.

      (5) Assuming that dATP and mavacamten generate 100% disordered vs ordered myosin heads respectively, then the percentage of ordered heads can be calculated for a variety of conditions.

      Weaknesses:

      (1) Issues like the effect of fibre disarray and lattice spacing on the SHG signal are not well defined.

      We thank the reviewer for raising this important point. Regarding fibre disarray, we took advantage of the spatial resolution of SHG imaging in thick samples to selectively analyse regions of the preparation in which myofibrillar organization was preserved. This approach proved particularly useful in samples such as those from HCM, where a pronounced global disarray is present but locally well-organized regions can still be identified and reliably analysed.

      Concerning lattice spacing, we have performed additional experiments on psoas muscle in which lattice spacing was modulated using dextran. This allowed us to directly assess the sensitivity of the technique to changes in inter-filament spacing.

      Figure 1 and the corresponding text have been revised accordingly to incorporate and clarify this point (page 11-12).

      (2) The, now well-defined heterogeneity of thick filament structure is not acknowledged.

      We agree that the heterogeneity of thick filament structure is an important aspect that should be acknowledged. In the revised manuscript, we have now explicitly addressed this point in the Discussion (page 21). In particular, we highlight that the capability of pSHG to probe the ON/OFF state with sub-sarcomere spatial resolution offers the future opportunity to investigate spatial heterogeneity in thick filament organization. This aspect is now clearly acknowledged and discussed in the context of the potential applications of the technique.

      (3) dATP was only used on minipig cardiac fibres. The effect of dATP on rabbit psoas and mouse cardiac fibres would be a useful comparison and would help validate the calculation of % ordered heads.

      We agree that, in the original version of the manuscript, there was a methodological imbalance in the use of dATP across preparations. To address this point, we have performed additional experiments in which the effect of dATP was also evaluated in rabbit psoas and mouse cardiac fibres.

      Importantly, the inclusion of these data has also proven useful in the discussion of the ON/OFF equilibrium across different muscle types and species. The corresponding results and discussion have been added to the revised manuscript (page 18-19).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      In addition to addressing the points in the Public Review, please also address the following points.

      (1) There are some issues with the calculated ionic strengths of the solution (or some details are missing). For example, on p. 4, it is stated that there is a 200 mM ionic strength solution that contains both 100 mM KCl and 2 mM MgCl2, meaning that the ionic strength is higher than 200 mM. There are other similar issues in other places.

      We thank the reviewer for pointing this out. We agree that there was an error in the reported ionic strength calculations and that some details were not sufficiently clear.

      We have now corrected the ionic strength values throughout the manuscript and revised the Methods section to provide a clearer and more consistent description of the solution composition (pages 6-7).

      (2) Please report standard deviations rather than standard errors.

      We thank the reviewer for this suggestion. Following the recommendations of Reviewer 1, we have reanalyzed the entire dataset using an updated statistical framework. As part of this process, we carefully evaluated the most appropriate measure of variability and have now consistently reported the corresponding error estimator throughout the manuscript.

      (3) For the statistical testing, please mention what tests were done for normalcy. It appears that some data is not normally distributed, in which case nonparametric tests should be used.

      We have now reanalysed the entire dataset with the support of Prof. Francesco Sera, who has been included as a co-author. A hierarchical (mixed-effects) statistical model was applied, explicitly accounting for the nested structure of the data, repeated measurements, and unbalanced group sizes. As part of this updated statistical framework, normality tests were performed for all datasets. When the assumption of normality was not met, appropriate nonparametric or model-based approaches were used.

      (4) Please discuss sex as a biological variable.

      Sex as a biological variable was not specifically investigated in the present study, and we agree that this represents a limitation. This point has now been explicitly acknowledged in the revised manuscript (page 5).

      (5) Please add an explicit section on limitations.

      We thank the reviewer for this suggestion. In the revised manuscript, the Discussion has been expanded to more clearly highlight the limitations of the technique and to better contextualize them in comparison with other approaches (page 21). We believe that integrating these aspects within the Discussion provides a more coherent and balanced presentation, and we have therefore chosen not to include a separate, dedicated limitations section.

      (6) Please discuss the limitations of using saturating concentrations of the drug. For example, the sensitivity of this method to detect changes in ON/OFF equilibrium at physiological concentrations is likely lower.

      The Discussion has been implemented to highlight that the sensitivity of this technique to the ON/OFF ratio could be further explored across species by employing a range of concentrations of Mavacamten and dATP, thereby better capturing physiologically relevant conditions (pages 18-19).

      (7) It is stated on p. 13 that R403Q has a higher sensitivity to mava versus dATP. I'm not sure this is supported by the data. The R403Q starts at a higher percentage of ON, and thus the effect size will be larger with mava, but this doesn't imply anything about sensitivity (which implies concentration dependence).

      Our statement was not intended to imply a difference in sensitivity in terms of concentration dependence, but was instead based on the statistical outcome of our measurements. Specifically, while a significant response was observed in the presence of mavacamten, no statistically significant response was detected upon dATP application in the R403Q condition.

      We acknowledge that this does not constitute evidence of differential sensitivity per se, and we have revised the text accordingly by removing the concept of sensitivity to avoid potential misinterpretation.

      (8) Please add a discussion of the potential contributions of RLC phosphorylation to ON/OFF regulation.

      We agree with the reviewer that the potential involvement of RLC phosphorylation in ON/OFF regulation is an important aspect to consider in future work, and we have now acknowledged this point in the revised Discussion.

      Reviewer #3 (Recommendations for the authors):

      Some things that need to be clarified:

      (1) The comparison of skinned and intact muscle. These appear to give an unaltered SHG gamma signal (Figure 2), but a change in lattice spacing is expected between skinned and intact fibres. There is no mention of the lattice spacing of the samples or whether this was controlled. The implications are that SHG is independent of lattice spacing and/or the fraction of ordered myosin heads is independent of lattice spacing - each of which would be a useful result.

      We thank the reviewer for this important observation and agree that the role of lattice spacing is a relevant factor in the interpretation of the SHG signal. To address this point, we have performed an additional series of experiments on rabbit psoas muscle in which lattice spacing was modulated using dextran. These measurements allowed us to directly assess the sensitivity of the SHG signal to changes in interfilament spacing. We observed relatively small effects, but in the expected direction.

      The lack of appreciable differences between skinned and intact preparations may therefore be explained by the limited sensitivity of the technique to detect the relatively small variations in lattice spacing associated with these conditions.

      (2) In comparing the WT and mutant mini pig cardiac data, the authors note that the mutant fibre has more disarray. A comment on the effect of disarray on the SHG signal would be helpful. How much of the difference between WT and mutant could be due to this disarray? What happens if more vs less disordered areas of the fibre are compared?

      Myofibrillar disarray is indeed a characteristic feature of HCM tissue and is more evident in the R403Q minipig samples compared to WT. To minimize potential polarization artefacts related to structural disorganization, the entire field of view was first examined to identify regions of interest (ROIs) where sarcomeres showed minimal local disarray. Data acquisition and analysis were restricted to these locally well-aligned regions, as described in the Methods section. Therefore, the comparison between WT and mutant samples was performed on locally well-oriented regions rather than on highly disorganized areas.

      We have now clarified this point in the revised manuscript to better explain how ROI selection, restricted to locally aligned regions, minimizes the potential contribution of myofibrillar disarray to the pSHG measurements (page 14).

      (3) I note in Figure S3 that there is a bigger dispersion in the minipig data than mouse or rabbit. The minipig may also show a non-normal distribution. Has this been considered?

      Following the reviewer’s suggestion to include dATP measurements in additional species, we have extended the interspecies analysis and revised Figure S3 to provide a direct comparison across rabbit psoas, mouse cardiac, and minipig cardiac samples under the investigated conditions.

      From this more comprehensive dataset, no substantial differences in data dispersion are apparent among the different muscle types. Moreover, as part of the updated statistical framework, normality tests were performed for all datasets

      (4) A note about the heterogeneity in the regulation of thick filaments along their length should be added. There is significant evidence for a difference in regulation between the MyBP-C regions and the rest from single molecule studies (Kad lab) and interference X-ray signals (London Kings group). This is probably beyond the current resolution of the SHG, but the complexity should be acknowledged. Heterogeneity in the thick filament is also apparent from cryo-EM images of relaxed muscle and thick filaments.

      This important point is now acknowledged the revised Discussion (page 21).

      (5) Interpretation of the effects of mavacamten on the ATPase data ae complicated by the observation that mava is an inhibitor of myosin independent of the effect on the order- disorder of thick filaments. I.e. mavacamten will inhibit myosin S1.

      This point has now been addressed (see response to Reviewer #2, point 5 weaknesses).

      Minor issues:

      (1) P7 line 2: where/were purchased from Sigma.

      (2) P7 last but one line: R403Q/R4303Q.

      (3) P8 last line: 2-deaoxyATP 2-deoxyATP.

      (4) Results, p9: in the section title and first line, replace psoas with rabbit psoas.

      (5) Line 5: ROI is not defined, only in the Figure legend.

      (6) Mavacamten: 50 uM used in psoas and 10 uM in cardiac fibres. A note on why would help those not familiar with this literature. Similarly, in Figure S3, presumably the mava concentration was saturating in each case.

      (7) P14, last paragraph, line 3: Figure 4 should be Figure 5.

      (8) P17, 5 lines from the end: as effective at 2-deoxyATP at/as.

      (9) P19, lt line: fibber/fiber.

      All these minor points have been fully addressed. We thank the reviewer for noting them.

    1. eLife Assessment

      This useful study uses creative scalp EEG decoding methods to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. The evidence supporting the conclusions is solid, although it could be further strengthened with a modified experimental design. This paper would be of interest to researchers studying cognitive control and adaptive behavior.

    2. Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has an innovative design and the analytical choices are appropriate and the evidence supporting the conclusions are overall solid after taking consideration the control analyses the authors have provided, although within limits of their study design.

      Strength:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      Comments on revised version.

      I appreciate the tireless effort the authors have presented to provide extra control analyses, however, on the other hand I do wish to remind them that sometimes limitations of one study is better addressed by a new study with improved design. There is very good reason why orthogonalization is critical to separating confounding influencing factors which the current study did not completely achieve. Reviewer 1's suggestion on alternative designs is certainly worth considering, and I encourage the authors to continue working on perfecting the experimental design in future work.

    3. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

      We thank the editors and the reviewers for their positive assessment and constructive feedback on our work. We also thank the editor for communicating with reviewer #1 regarding our thoughts on their comments. We hence revised the manuscript based the new feedback from reviewer #1. Please see below our responses to each comment raised in the reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. In particular, two types of learned associations are studied. One association involves associations between stimulus features and control states (SC), and the other involves stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates.

      The authors conclude that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing that SC and SR correlates are (1) not entirely overlapping in cross-decoding, (2) simultaneously observed on average over trials, (3) independently correlate with RT, and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Thank you for acknowledging the strength in EEG decoding and design. We have addressed all your concerns raised below point by point.

      In my view, the ability to address this issue with additional analyses is relatively limited. I cannot think of a solid way to escape this issue in the present design. Adding a nuisance regressor to their RSA regression, which I suggested in my previous response, may reduce the bias, but the efficacy of this would be limited to controlling only a certain kind of phase-dependent confound (a 'main effect' component of study phase, i.e., one that is constant across other conditions; see next message for discussion).

      Rather than new analyses, I think a more straightforward revision might be to modify the conclusions advanced in the paper, so that they are more solidly supported by the design and evidence. In my opinion, this design is ill-posed to solidly identify SC and SR representations. As a result, I think that any framing that alleviates pressure on this design to yield 'solid' evidence for identification of SC and SR representations, and instead emphasizes stronger areas of this work, would be constructive.

      For example, one potential framing is to advance the idea of SC and SR representations, and discuss an idealized design that could identify them by orthogonalizing stimulus features from ISPC, which are genuinely novel and useful ideas. The current design could then be presented as an opportunistic or initial case study of testing this question, while acknowledging its limitation upfront. The goal here would be to frame the study in a way that allows for the results to be presented with an appropriate grain of salt, while also illustrating the authors thoughtfulness and creativity in devising analyses to test for latent associative representations. In this case the 'solid' label would reference the authors' reasoning and analyses rather than design and conclusions.

      I only intend this example as an illustration; there may be several ways of framing this paper so that it is on more 'solid' ground, and I don't want to dictate how exactly authors should write their paper.

      Nonetheless I think this issue is important, not only to avoid faulty inference, but also to avoid establishing counterproductive precedents in this field. For example, if students read a paper whose conclusions are labeled "solid" but that nevertheless has critical flaws in its design, then those students may be misled in their own work. But if the flaws were discussed transparently and critically, and the strength of the conclusions were de-emphasized relative to other aspects of the paper, students may not only be inspired by the ideas developed in the paper, but also come away with knowledge about the issues of experimental design.

      Discussion of the weaknesses in the conclusions and an additional potential analysis:

      The key goal of this study is to identify SC/SR representations, which requires decoupling stimulus features from item-specific proportion congruency (ISPC), but the study-phase confound contaminates this decoupling. I think this impacts both their cross-phase decoding and RSA analyses.

      In their results (lines 139-144):

      "This standard ISPC manipulation can test whether neural representations of controlled and non-controlled information are wrapped on the same trial by combing with the following EEG analysis (See Methods). However, it potentially mixes color identity with SC, word identity with SR, and the ISPC between SC and SR. To deconfound these factors when estimating SC and SR association representations on each trial, we modified this paradigm by flipping the ISPC contingencies across different phases of the task."

      SC/SR representations are higher-order conjunctive classes, formed by the interaction of lower order variables (Color, Word, and ISPC). This means that successfully identifying these representations relies on demonstrating that each class can be reliably individuated from every other class in a manner that cannot be explained by (1) representation of shared lower-order features, such as stimulus color or response, and (2) trivial nuisance factors such as study phase. However, in this design, the trivial factor of study phase is strongly confounded with the ISPC contingency flip.

      Regarding RSA: If study phase indeed leads to trivial separability, it seems that similarity among conditions within the same phase would be inflated because the study phase does not appear in the RSA model (Figure 12). In which case, the SC/SR coefficients would be inflated, as the SC and SR models are entirely within-phase.

      Nevertheless, an additional control analysis may be possible here. Looking at this regressor set, it seems possible to me to fit a model where the predominant phase is entered as an additional covariate. (If I am reading this correctly, phase would appear as a 2x2 block-diagonal matrix). I suggested this in my last review, but in their recent letter, authors refused. I do not understand why, as I think this regressor set should be identifiable, but perhaps I am wrong here.

      That being said, I do not think adding this nuisance covariate would fully solve the issue. The lower-order regressors (Color, Word, ISPC) are defined as the EEG responses shared across phases of the study. To interpret the higher-order SR/SC coefficients, the lower-order components must be fully partialled out. But the putative impact of study phase contaminates this interpretation, as study phase could trivially decrease the similarity of lower-order terms (e.g., decreasing Color similarity between phase 2 vs 3). In which case, the partialling would be expected to be incomplete.

      This incomplete partialling is the RSA analogue of the issue in interpreting cross-phase decoding discussed above. As there, so too here I do not see a solid way around it in the present design. This is why in my previous review I referred to the addition of a phase covariate in the RSA regression as a "band-aid": it controls for some problems (main effect of phase), but not all (interactions of phase and lower-order terms).

      To summarize, I think that RSA would offer an additional opportunity to control for this potential confound, albeit in a limited sense (study phase effects that are consistent across conditions). But in a more general and rigorous sense, to me, the SC and SR terms in the RSA regression also seem susceptible to the same weakness as the cross-phase decoding analysis.

      We thank the reviewer for taking the additional time and effort to provide the new comments. Following the reviewer’s suggestion, we revised the language regarding the conclusions of this project (page 2,5,27-28) and explicitly discussed the weaknesses of the design on decoding and outline a possible solution for future studies in the Discussion section:

      “A limitation of the current design is that in theory temporally structured noise (e.g., autocorrelation in EEG data) may bias the decoding accuracy due to the blocked design. Although the present data provided no evidence that the decoding results in this study were biased by temporally structured noise, future studies should aim to develop experimental designs that eliminate this potential confound at the source. One potential solution would be to introduce additional phases flipping ISPC manipulations. At the same time, enough trials must be included in each phase to ensure the strength of the ISPC effect within each phase. A careful balance between session number and length will be helpful to optimize the duration of such a design.”

      As eLife also publishes review report, below we also summarize the three control analyses we ran and our reasoning of how they (would) address the issue of temporally structured noise for interested readers to assess:

      We acknowledge the theoretical issue of temporally structured noise (TSN) in our design when classes were decoded across different phases. However, the key question for the current data is whether there is empirical evidence that the decoding results were actually driven by TSN. To clarify this issue, we summarize several lines of evidence suggesting that the decoding results were not attributable to TSN:

      (1) Split-half cross-validation. We split the EEG data from Phase 2 and the combined Phases 1 and 3 into chronological first and second halves. Phases 1 and 3 were combined because they shared the same MC and MI assignments. This resulted in four possible combinations, each consisting of eight classes drawn from different phases: combination 1 included the first half of Phase 2 and the first half of Phase 3; combination 2 included the first half of Phase 2 and the second half of Phase 3; combination 3 included the second half of Phase 2 and the first half of Phase 3; and combination 4 included the second half of Phase 2 and the second half of Phase 3. We trained the decoders on one combination and tested them on another and then averaged the decoding results across all possible training-test assignments. The similar decoding patterns observed across these analyses (Fig. 6a,b) further confirmed that the decoding results were not driven by TSN.

      This analysis is conceptually similar to the “cross-phase” decoding analysis suggested by the reviewer in the first round of review. We also performed an additional distance-based control analysis (see below) to further test whether the decoding results could be explained by TSN, without imposing the constraint used in the split-half cross-validation that trials from the two phases had to fall within a 400-trial window.

      (2) Distance analysis. We predicted that if a test trial is closer to a training trial of the same trial type, the higher similarity in TSN between the training and test data would more strongly inflate the decoding accuracy of the test trial, resulting in a negative correlation between distance between a test trial and its closest training trial of the same type and the test trial’s decoding accuracy. However, we did not observe such a negative pattern (Fig. 6c). Note that this distance was defined with respect to trials of the same type, rather than absolute chronological time.

      (3) Shuffled analysis. If the decoding results were primarily driven by TSN, either at a short-term or long-term timescale, then shuffling the condition labels within each mini block should preserve the TSN structure present in the real data. In that case, the decoding results from shuffled data should not differ from those observed from real data. However, we found the significant difference between real data and shuffled data as shown in Author response images.

      Author response image 1.

      Shuffling analyses with stimulus-locked data support separable SC and SR subspace. (a) Group average decoding accuracy of all 16 experimental conditions as a function of time after stimulus onset. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (b) Group average t values of representational strength for each factor over time. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (c) SC and SR association results from Fig. 1b.

      Author response image 2.

      Shuffling analyses with response-locked data support separable SC and SR subspace. (a) Group average decoding accuracy of all 16 experimental conditions as a function of time after stimulus onset. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (b) Group average t values of representational strength for each factor over time. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (c) SC and SR association results from Fig. 2b.

      We thank the reviewer for the suggestion on RSA with phase. There are some concerns for this analysis:

      First, we think that the suggested analysis may be difficult to interpret. Because the SC and SR conditions differ across phases. Regressing out phase in RSA could also remove SC and SR information.

      Second, based on the reviewer’s comment, we understand that the suggested analysis may still not provide a clear falsifiable criterion for determining whether the results could be driven by the theoretical TSN issue inherent in the design.

      Third, the three control analyses we have performed examine this issue from different perspectives and collectively provide no evidence that the results were driven by TSN.

      Other readers may, like me, be puzzled by the selection of this particular experimental design to test this question of SC and SR coding, given the temporal confound among SC/SR classes, and given that there would seem to be many possible designs that are less confounded. For example, why not use a design where ISPC was swapped/shuffled several more times within each subject, so that PHASE is more orthogonal to long-timescale noise? Isn't ISPC learning fast enough to support learning phases shorter than 700 trials? Such readers would likely appreciate a frank discussion of this dilemma, and a motivation for the choice of the present design, within the manuscript.

      Thank you for your suggestion regarding the design. It is possible that the (re-)learning of ISPC can be fast. That said, enough trials are required to obtain a robust ISPC effect for each phase after the ISPC flips. Given that the EEG scanning (not including capping) in current design was about 1.5 hours, it is impractical to have both more sessions for a more orthogonal design and long sessions for robust within-session ISPC effects. We chose to maximize the latter because flipped behavioral ISPC effect in each session is the basis for the following EEG analysis. We have included the reviewer’s suggestion as a potential design solution for future studies in the Discussion section mentioned above on page 27.

      Pre-stimulus coding:

      To explain the apparent pre-stimulus coding of several task variables, the newest version of the manuscript proposes that subjects were proactively coding these variables via predictive mechanisms. This is an interesting account of item-specific control. It is also surprising, given that item-specific control mechanisms are typically conceptualized as reactive or stimulus-driven phenomena. But I think support for a proactive control account was incomplete. The mechanistic logic was not presented, and no hypotheses under this account were developed or tested. So I would suggest pinning down some hypotheses here and actually putting this account to the test.

      Thank you for raising this important point. Although ISPC effects are considered reactive, in our design the long sessions may create a temporal context for the participants to differentiate the current control demand linked to each color. The maintenance of such contextual information needs to span across trials, leading to pre-stimulus coding that proactively guides the control demand for each color. This claim is not central to this manuscript, which investigates whether SC and SR representations simultaneously guide behavior. Additionally, we do not think the current design is well-equipped to test this hypothesis because the pre-stimulus onset is the only supporting evidence. In the revised manuscript, we discussed this as a future research direction and proposed a design that aims at better isolating proactive control signal on page 25.

      Random slopes were omitted due to convergence failure, but this can inflate false positive inferences (e.g., Barr et al. 2013), and doesn't really motivate a minimal model. I'd suggest trying a slightly reduced model (e.g., drop correlations via `slope || subject`) using buildMer automated selection, or switching to brms.

      We indeed tried both the full model of random effects (i.e., considering covariance between all slopes and intercept) and a reduced model without any covariance (i.e., listing each random slope separately without intercept in lme4). However, neither model converged at all time points.

      Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a solid design and the analyses are appropriate for the research questions.

      Strengths:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      We thank you for acknowledging our work on design and brain-behavior analysis. We have addressed your concerns raised below.

      Weaknesses:

      I still have some doubts on the effectiveness of the experimental manipulation on Phase 2: although the ISPC effect is still present, it is much weaker in comparison, suggesting the participants did not learn the contingency statistics in Phase 2 as well as they did in the other phases, due to either the lingering effect of the previous phase or an inherent bias towards one color pairs. Perhaps by separately plotting the earlier and later blocks of Phase 2 any difference can be revealed if it exists. This behavioral difference could result in unequal levels of SC/SR representation across phases, which may raise problems when data were combined for analyses that assume the neural effects are equivalent.

      Thank you for your concern about this important issue. We agree with the reviewer that the true SR/SC levels may not be equivalent between Phase 2 and Phase 1/3. Nevertheless, because the manipulation of ISPC is binary, the decoders were trained to test whether the neural signals represent the two levels of ISPC (i.e., a higher vs. a lower level) differ systematically. The decoding analysis does not require that the neural effects of SC and SR must be numerically equivalent between phases (i.e., it is not necessary that the two levels are equidistant from the center point of SC/SR. Indeed, the decoding analysis only requires that the two levels are different). For example, if the ISPC level ranges from -1 to 1 and the EEG signals can reliably decode ISPC levels of -0.5 and 0.7 (i.e., two unequal levels), it can still be treated as supporting evidence that ISPC levels are encoded in the EEG signals. The same logic applies to the representational subspace analysis. As to the RSA, as can be seen in Fig. 12, the regressors are also binary, encoding whether two experimental conditions share the same SC/SR level without assuming equivalent neural effects. Thus, we argue that the reported decoding and RSA can still test the encoding of SR and SC. We discussed this issue on page 24.

      Following the reviewer’s comment, we plotted the ISPC effects in first and second half of Phase 2 separately (the figure below). We also tested whether the ISPC effects differ qualitatively between the two halves using a 3-way ANOVAs separately on RT and Error rate. The results showed that the time (the first half vs. the second half) × Congruency × ISPC interaction was not significant for either RT data (F<sub>(1,39)</sub> = 3.40, p > 0.05) or error rate (F<sub>(1,39)</sub> = 2.49, p > 0.05), suggesting that ISPC effect did not systematically change over time in Phase 2 See Supplementary Figure 9.

    1. eLife Assessment

      This work of fundamental significance introduces a novel statistical model of spiking activity that incorporates continuous-time gain modulation. The authors provide exceptional evidence that the model outperforms earlier approaches and alternative candidates in capturing spiking responses across multiple visual areas in the macaque. Beyond its methodological contribution, the study offers new insights into how stimulus-driven variability and internally generated gain fluctuations evolve over time and between brain areas. The framework is likely to find broad application beyond the datasets examined here.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Rupasinghe and co-authors introduce a new statistical model for spiking neurons. Building on earlier work, they propose to model spikes as arising from a Poisson process whereby the firing rate is the product of stimulus drive and a stimulus-independent gain signal. The critical innovation of this work is that the gain signal is modeled in continuous time. Earlier explorations of this statistical construction treated the gain-signal as constant within a trial. This innovation is elegant and important. It makes the model richer, more plausible, and more broadly applicable. The authors show that the model parameters are recoverable from realistic amounts of data and then apply the framework to previously studied datasets. They show that the new model outperforms earlier models and alternative candidates in capturing spiking data across four visual areas of the macaque monkey. Analysis of the model parameters replicates some earlier findings and uncovers several new insights. The model and fitting methods can be broadly applied to partition different types of signals and noise from spiking data and are likely to be widely adopted in the systems neuroscience community.

      Strengths:

      (1) Through clever use of advanced statistical techniques, the authors manage to infer critical information from single trial single cell data.

      (2) The question of which aspect of a spike train is signal and which is noise is omnipresent in neuroscience. By improving our ability to characterize the distinct factors that shape spiking activity, this work makes a fundamental contribution to the literature.

      Weaknesses:

      (1) The work is entirely focused on single cell data. While this is a great starting point, expanding the approach to spiking activity in neural populations is an important future goal. The discussion lays out a roadmap towards this goal.

      Comments on revised version.

      I thank the authors for their sincere engagement with the reviews. They have addressed all issues I had raised. I found the first version of the manuscript already impressive. The revised version is a bit clearer about the exact relationship to some prior work and now documents additional new findings that validate the successful partitioning of signal and noise and directly connect stimulus-induced variability quenching to the stabilization of the latent gain signal. This makes it a really great paper.

    3. Reviewer #2 (Public review):

      Summary:

      Neurons have varied responses to external stimuli that cannot be explained by naive Poisson models. Previous work has quantified and partitioned higher-than-Poisson variability in the brain into different components. The authors improve on these methods to infer how both the stimulus drive and internal gain dynamics impact neuronal variability continuously in time. The clean and well-reasoned model is rigorously developed and then applied to neural data across the visual hierarchy. This lends new insights into how variability is partitioned, agreeing with and extending previous work on how that variability changes from early visual areas (LGN, V1) through to higher, motion-sensitive areas (area MT). Another key contribution is that this partitioning can be fully addressed as a continuous-time process, which allows for dissection of how the timescale of fluctuations in these two components changes across the brain's processing arc.

      Strengths:

      (1) The model is cleanly derived and thoroughly documented, including useable code shared in a GitHub repo. This makes the method immediately portable to other neural systems.

      (2) The figures and writing are clear and understandable and all pieces of the derivations are included in the main text and supplementary information.

      (3) Comparisons to other models, particularly the one from Goris et al., 2014 shows how this Continuous Modulated Poisson (CMP) model outperforms previous work.

      (4) New insights about how variability partitioning changes across the visual stream from LGN to MT are revealed, including how the gain fluctuates on longer timescales in higher visual areas. Another key result about the anticorrelation between the variance in stimulus drive and gain fluctuations comports with theories about how neurons maintain efficient, reliable encoding.

      (5) In addition to the results reported here, this work will serve as an excellent tutorial for students and postdocs first delving into the sources of variability in the brain.

      Weaknesses:

      (1) The work builds off previous studies of the partitioning of variability in the brain, but provides important new extensions as noted above. Sub-poisson variability cannot be addressed in the current framework, but ideas for extensions are included in the Discussion.

      Comments on revised version.

      The revisions have thoroughly addressed my previous comments and concerns and the paper's clarity and scope have improved.

    4. Author response:

      The following is the authors’ response to the original reviews.

      In the revised manuscript, we have expanded the real-data analyses, clarified the relationship between CMP and prior modulated Poisson models, and added discussion of model limitations and future extensions. In summary, the major changes include:

      (1) We revised the Introduction, Results, Methods, and Goris-model appendix to clarify the relationship between CMP and prior modulated Poisson models. In particular, we now emphasize that the key distinction is CMP’s continuous-time stochastic gain process.

      (2) We moved the simulation-based recoverability analysis from Appendix 3 into the main Results section (Figure 3 in the revised manuscript), making the validation of the inference procedure more visible to readers.

      (3) We added new analyses of the inferred gain process. Specifically, we now show the cross-trial gain mean and cross-trial gain variance in Figure 4A to assess whether gain captures stimulus-locked structure, and we added an analysis of pre- versus post-stimulus cross-trial gain variability in Figure 4C to test for gain-variability quenching during stimulus presentation.

      (4) We clarified the definitions and implementation of the Baseline Poisson, Poisson-GP, and Goris-style comparison models, including the role of the smoothness prior on the stimulus drive.

      (5) We expanded the Discussion to describe future extensions to population recordings, including a GPFA-inspired extension with low-dimensional shared gain activity across neurons.

      (6) We added a Discussion paragraph clarifying that the current CMP model captures Poisson and super-Poisson variability, but not sub-Poisson variability, and outlined possible extensions using spike-history terms, renewal-process likelihoods, or alternative count distributions.

      eLife Assessment

      This work of fundamental significance introduces a novel statistical model of spiking activity that incorporates continuous−time gain modulation. The authors provide exceptional evidence that the model outperforms earlier approaches and alternative candidates in capturing spiking responses across multiple visual areas in the macaque. Beyond its methodological contribution, the study offers new insights into how stimulus−driven variability and internally generated gain fluctuations evolve over time and between brain areas. The framework is likely to find broad application beyond the datasets examined here.

      We sincerely thank the Senior Editor, Reviewing Editor, and both reviewers for their careful evaluation and constructive feedback. We are encouraged by the positive assessment of the work and by the recognition of its methodological and conceptual contributions. We especially appreciate the acknowledgement that the continuous-time formulation provides a useful framework for modeling gain modulation in spiking activity, improves upon earlier approaches in capturing responses across multiple visual areas, and offers new insights into how stimulus-driven variability and internally generated gain fluctuations evolve over time and across brain regions.

      In the revised manuscript, we have addressed the reviewers’ comments by clarifying the relationship between CMP and prior modulated Poisson models, strengthening the presentation of the simulation-based recoverability analysis, adding new validation analyses of the inferred gain process, and expanding the Discussion of model scope, limitations, and future directions. In particular, we now more clearly distinguish the continuous-time gain process in CMP from Goris-style models with constant or piecewise-constant gain, move the simulation recoverability analysis into the main Results, examine trial-averaged inferred gain and gain-variability quenching, clarify the definitions of the baseline and comparison models, and discuss extensions to population recordings and sub-Poisson variability.

      We believe these revisions improve the clarity, rigour, and scope of the manuscript. Below, we address each reviewer comment in turn and describe the corresponding changes made in the revised manuscript.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, Rupasinghe and co−authors introduce a new statistical model for spiking neurons. Building on earlier work, they propose to model spikes as arising from a Poisson process whereby the firing rate is the product of stimulus drive and astimulus−independent gain signal. The critical innovation of this work is that the gain signal is modeled in continuous time. Earlier explorations of this statistical construction treated the gain−signal as constant within a trial. This innovation is elegant and important. It makes the model richer, more plausible, and more broadly applicable. The authors show that the model parameters are recoverable from realistic amounts of data and then apply the framework to previously studied datasets. They show that the new model outperforms earlier models and alternative candidates in capturing spiking data across four visual areas of the macaque monkey. Analysis of the model parameters replicates some earlier findings and uncovers several new insights. The model and fitting methods can be broadly applied to partition different types of signals and noise from spiking data and are likely to be widely adopted in the systems neuroscience community.

      Strengths:

      (1) Through clever use of advanced statistical techniques, the authors manage to infer critical information from single−trial single−cell data.

      (2) The question of which aspect of a spike train is signal and which is noise is omnipresent in neuroscience. By improving our ability to characterize the distinct factors that shape spiking activity, this work makes a fundamental contribution to the literature.

      We sincerely thank the reviewer for the thoughtful and detailed evaluation of our manuscript. We are pleased that the continuous-time formulation and its methodological contributions were viewed as elegant, important, and broadly applicable. We also appreciate the reviewer’s recognition that the framework provides a useful way to separate stimulus-driven and modulatory components of neural variability from single-trial, single-cell data. The reviewer’s comments helped us improve the precision of our framing, clarify the relationship between CMP and prior modulated Poisson models, and strengthen the validation of the inferred gain process. Below, we respond to each point in turn and describe the revisions made in the manuscript.

      Weaknesses:

      Overall, I find the work impressive and important. I have a couple of questions and suggestions.

      (1) The work is entirely focused on single−cell data. While this is a great starting point, expanding the approach to spiking activity in neural populations is an importantfuture goal.

      We thank the reviewer for this important suggestion. We agree that extending the CMP framework to population recordings is a natural and important direction for future work. In the present study, we focus on single-neuron responses to establish the continuous-time model, validate the inference, and characterize how stimulus-driven activity and stochastic gain fluctuations can be separated at the level of individual cells. However, the same modeling principles could be extended to simultaneously recorded neural populations by introducing shared latent structure across neurons. For example, one natural direction would be to combine CMP with ideas from Gaussian Process Factor Analysis [Keeley et al., 2020], using low-dimensional shared gain activity to capture population-wide fluctuations, while retaining neuron-specific stimulus-driven components. Such an extension would allow the model to capture correlated variability and shared modulatory dynamics across neural ensembles. In the revised manuscript, we have expanded the Discussion to describe this possible future extension to population recordings.

      To address this comment, we expanded the Discussion (Page 14: lines 473-478) to describe a possible GPFA-inspired extension of CMP to population recordings.

      (2) Line 49−53: These statements seem incorrect to me. The modulated Poisson model , as introduced in Goris et al (2014), is a process model that can perfectly be used to generate spike trains (within a trial, spiking emerges from a Poisson process, which canbe homogeneous or inhomogeneous). Moreover, the model contains a parameter thatrepresents the duration of the counting window (delta t). The dependency of over− dispersion on the size of the time bins for real neurons is shown in Figure 1b (inset plot) of that paper (and shown to resemble the model prediction). This time− dependency was further explored by the same authors in Goris et al (2018 − Journal ofVision) and also in Henaff et al (2020 − Nature Communications). I suggest that the authors rephrase this argument (here and at some later points in the paper). They could just say that the Goris model makes the simplistic and implausible assumption that, within a given trial, gain does not fluctuate. This is clearly an important limitation and the key difference with the continuous model introduced here.

      We sincerely thank the reviewer for identifying this lack of clarity in our original description. We agree that our original description was not sufficiently precise. The modulated Poisson model introduced by Goris et al. (2014) is indeed a generative process model and can be used to generate spike trains, with spiking arising from a Poisson process that may be homogeneous or inhomogeneous within a trial. We apologize for implying otherwise.

      Our intended point was that, in the original formulation, the modulatory gain is represented as a scalar random variable associated with a counting window or trial, and therefore does not explicitly model gain as a continuously time-varying process within a trial. Thus, the key limitation addressed by CMP is not the use of a Poisson process, but the assumption that gain is constant or piecewise constant over the relevant interval.

      In the revised manuscript, we have rephrased the Introduction to clarify this distinction. We now describe the Goris model more accurately as a modulated Poisson framework in which gain is constant over the counting window, and we emphasize that CMP extends this framework by replacing this assumption with a continuous-time stochastic gain process. We have also added discussion of related time-dependent analyses and extensions [Goris et al., 2018, H´enaff et al., 2020], as thoughtfully suggested by the reviewer.

      In addition, we revised the Results and Methods to clarify how the Goris-style baselines were implemented in our comparisons. Specifically, all Goris-style results reported in the main model comparisons use versions with a smoothness prior on the stimulus drive, where the stimulus-dependent firing rates are set to the smooth firing-rate estimates obtained from the Poisson-GP model. This ensures that the comparisons focus on different assumptions about the temporal structure of the gain process, rather than differences in stimulus-drive estimation. We also clarified the comparison to Goris-style variants without this smoothness prior, in which the stimulus-drive parameters are estimated directly under the corresponding Goris-style likelihood (Figure 5 - figure supplement 2). These results show that the smoothness prior on the stimulus drive substantially improves model performance. Finally, we revised the Figure 1 caption and the Goris-model appendix to make these distinctions explicit.

      To address this comment, we revised the Introduction (Pages 2-3: Lines 49-77), Results (Page 9: Lines 263-266, 273-277, Page 11: Lines 319-326), Methods (Page 21), Figure 1 caption, and Goris-model appendix to clarify that CMP extends the Goris framework by modeling gain as a continuously time-varying process within trials.

      (3) Line 54−55: I think the first part of the claim is a bit misleading. There is nothing in the Goris model that would inherently limit it to homogeneous Poisson processes, as seems to be implied by this description. The model is built on the assumption thatspike generation within a trial arises from a Poisson process. This may very well be an inhomogeneous Poisson process (i.e., a stimulus−dependent time−varying firing rate). Homogeneous and inhomogeneous Poisson processes both give rise to Poisson distributed spike counts (and thus a mixture of Poisson distributions across trials in the Goris model). I suggest the authors clarify this description a bit. Note that the two model variants illustrated in Figure 1b and c were also explored in Henaff et al (2020 − Nature Communications).

      We thank the reviewer for this helpful clarification. We agree that the Goris model is not limited to homogeneous Poisson spiking and can incorporate a stimulus-dependent, time-varying firing rate within trials. We did not intend to imply otherwise, and we have revised the relevant text to avoid this misunderstanding.

      Our intended point was that, in formulating continuous-time extensions of the modulated Poisson framework, we explicitly model the time-varying stimulus drive using a smoothness prior, as in the CMP framework, and then consider different assumptions about the temporal structure of the gain process, including constant gain and independently resampled gain across time bins. This highlights the distinction between piecewise-constant gain assumptions and the fully continuous gain process introduced in CMP.

      In the revised manuscript, we have clarified this distinction in the Introduction, Results, and Methods. We now state that the Goris-style variants use stimulus-dependent, time-varying Poisson firing rates, and that the main difference between these variants and CMP lies in the temporal structure assumed for the gain process. We have also acknowledged related variants explored in Goris et al. [2018] and H´enaff et al. [2020], and clarified that our continuous-time formulations of the Goris model differs by imposing a smoothness prior on the stimulus drive. This allows us to estimate a regularized time-varying stimulus component while comparing different assumptions about gain dynamics, ensuring that the comparison focuses on the temporal structure of the gain process rather than differences in stimulus-drive estimation. We also highlight in Figure 5 - figure supplement 2 that even for the Goris-style models, versions that use a smoothness prior on the stimulus drive outperform versions that do not, which are closer to the original modulated Poisson formulation.

      To address this comment, we revised the Introduction (Pages 2-3: Lines 49-77), Results (Page 9: Lines 263-266, 273-277, Page 11: Lines 319-326), Methods (Page 21) to clarify that the Goris-style variants allow stimulus-dependent time-varying firing rates and differ from CMP primarily in their assumptions about gain dynamics. We also added citations to related time-dependent extensions of the modulated Poisson framework.

      (4) The extension to the continuous case is very elegant!

      We thank the reviewer for the positive comment and are pleased that the continuous-time formulation was viewed as elegant.

      (5) I find the result shown in Appendix 3 critically important. The recoverability of the model for realistic amounts of data is foundational for the rest of the paper. I wouldconsider including this analysis in the main results section. Not all readers may check Appendix 3, but they should know about this result.

      We thank the reviewer for emphasizing the importance of this result. We agree that demonstrating parameter recoverability is foundational to the paper and should be visible to readers in the main Results section. In the revised manuscript, we have moved the simulation-based validation from Appendix 3 into the main Results. This section now describes the synthetic CMP dataset, the inference procedure used to estimate the latent stimulus-drive and gain processes, and the comparison between true and inferred GP hyperparameters. These results show that the proposed inference framework can accurately recover the ground-truth stimulus drives, gain processes, and hyperparameters from realistic amounts of simulated data.

      To address this comment, we moved the simulation-based recoverability analysis from Appendix 3 into the main Results section (Page 6: Lines 200-211 and Figure 3).

      (6) Figure 3: I am wondering whether the inferred gain is capturing some response fluctuations that originate from the cell’s phase−selectivity. Could the authors compute the trial−averaged inferred gain (ideally, aligned to stimulus−phase at the start of the trial if this experimental parameter varied across repeats)? If they have successfully partitioned the response variance, the trial−averaged gain should have no systematic temporal structure. If it has a sinusoidal modulation, it may partially capture stimulus−drive. This could be an interesting test to run on all model fits to further validate that the partitioning into a signal and noise component succeeded as intended.

      We thank the reviewer for this insightful suggestion. We agree that verifying that the inferred gain does not capture stimulus-driven structure is an important validation of the model. In the revised manuscript, we have added the trial-averaged inferred gain to Figure 4A for the example neuron. This analysis shows that the trial-averaged inferred gain is relatively flat and neither resembles the inferred stimulus drive nor exhibits clear stimulus-locked temporal structure. This suggests that trial-specific gain fluctuations largely average out across repeats, consistent with the interpretation that the gain process captures random trial-to-trial variability rather than stimulus-driven activity.

      We also note that a direct comparison of this inferred gain trace across methods is not possible for the Goris-style baselines, because these models do not infer a continuous trial-specific gain process. Instead, they marginalize over scalar or time-bin-independent gain variables when computing likelihoods and Fano factor curves. Thus, the trial-averaged gain diagnostic is specific to the CMP model, where the posterior over the continuous-time gain process is explicitly inferred.

      To address this comment, we added the trial-averaged inferred gain to Figure 4A and clarified that it does not show a clear stimulus-locked temporal structure (Page 7: Lines 229-236).

      (7) One common observation that is currently not explored is the quenching of neuronal response variability following stimulus onset (Churchland et al 2010 − NatureNeuroscience), which was suggested to reflect a quenching of gain variability in Goris et al (2024 − Nature Reviews Neuroscience). Building on the previous suggestion, the authors could compute the temporal evolution of cross−trial gain variability from the inferred gain traces. Do they recognize a reduction in gain variability following stimulus onset? If so, it would be worthwhile to show this.

      We sincerely thank the reviewer for this valuable suggestion. We agree that examining whether gain variability decreases following stimulus onset provides an important test of the inferred gain process. In the revised manuscript, we have added an analysis of the temporal evolution of cross-trial gain variability before and after stimulus onset.

      First, in Figure 4A, we now show the cross-trial variance of the inferred gain for the example neuron. This trace shows larger gain variability during the stimulus-off period and a reduction following stimulus onset, suggesting that the inferred gain captures a stimulus-related quenching of trial-to-trial variability. To quantify this effect across the population, we also added a pre- versus post-stimulus comparison in Figure 4C. Following the approach of Churchland et al. [2010], we compared gain variability in two matched 400-ms windows: a pre-stimulus window ending at stimulus onset and a stimulus-period window beginning 100 ms after stimulus onset. For each neuron and stimulus condition, we computed the cross-trial variance of the inferred gain at each time bin, averaged this quantity within each window, and then compared the pre- and post-stimulus values across neuron-stimulus pairs.

      This analysis revealed a significant reduction in inferred gain variability following stimulus onset (one-sided paired Wilcoxon signed-rank test, p≤ 10<sup>−15</sup>), consistent with gain variability quenching [Churchland et al., 2010, Goris et al., 2024]. We now report this result in the main text and illustrate it in Figure 4A and Figure 4C. This provides additional evidence that the inferred CMP gain captures meaningful trial-to-trial variability and its temporal modulation around stimulus presentation.

      To address this comment, we added the cross-trial gain variance trace to Figure 4A and a population-level pre- versus post-stimulus gain-variability quenching analysis (Page 7 and 8: Lines 239-248) in Figure 4C.

      (8) Line 543−565: I want to make sure I understand the Baseline Poisson model and Poisson−GP correctly. For the baseline model, I had imagined that the authors would simply use the stimulus−conditioned PSTH as an estimate of the time−dependent firing rate, coupled with an inhomogeneous Poisson process assumption. But they additionally assume a Gamma prior on the firing rate to compensate for the sparsenessof the data (sometimes only 5 repeats per condition). The Poisson−GP includesexactly the same model components, but now the time−dependent firing rate is modeled by a Gaussian process. Doing this massively improves the goodness−of−fit (Fig 4A). Do I understand this correctly?

      We thank the reviewer for this careful reading. Yes, this understanding is broadly correct, and we have revised the manuscript to clarify the relationships among the Baseline Poisson, Poisson-GP, and Goris-style models. The Baseline Poisson model estimates a stimulus- and time-dependent firing rate independently for each stimulus condition and time bin, using a Gamma-Poisson formulation to regularize the estimate when the number of repeats is limited. The Poisson-GP model uses the same conditionally Poisson observation model, but replaces these independent time-bin-wise rate estimates with a smooth stimulus-specific Gaussian process model for the log firing rate.

      We have also clarified how the Goris-style models were implemented. All Goris-style results reported in the main model comparisons use versions with a GP prior on the stimulus drive. In these versions, the stimulus-dependent firing rates are set to the smooth firing-rate estimates obtained from the PoissonGP model, and the gain parameters are then fit under either the independent-gain or constant-gain assumptions. We used these GP-smoothed versions as stronger baselines. In Figure 5, Figure Supplement 2, we additionally compare these models to Goris-style variants without the GP prior on the stimulus drive, in which the stimulus-drive parameters are estimated directly under the corresponding Goris-style likelihood. This comparison shows that adding a GP smoothness prior to the stimulus drive substantially improves held-out model fit. Together, these analyses clarify that the GP-smoothed stimulus drive improves the Goris-style baselines, while the continuous-time gain process in CMP provides an additional improvement by capturing temporally structured trial-to-trial variability.

      To address this comment, we clarified the definitions of the Baseline Poisson, Poisson-GP, and Goris-style models (Pages 8-9: Lines 255-259, 263-266, 273-277), and revised the text (Page 11: Lines 319-326) describing Figure 4 - figure Supplement 2 to make explicit how this existing comparison isolates the effect of the GP prior on the stimulus drive.

      Reviewer #2 (Public Review):

      Summary:

      Neurons have varied responses to external stimuli that cannot be explained by naive Poisson models. Previous work has quantified and partitioned higher−than−Poisson variability in the brain into different components. The authors improve on these methods to infer how both the stimulus drive and internal gain dynamics impact neuronal variability continuously in time. The clean and well−reasoned model is rigorously developed and then applied to neural data across the visual hierarchy. This lends new insights into how variability is partitioned, agreeing with and extending previous work on how that variability changes from early visual areas (LGN, V1) through to higher, motion−sensitive areas (area MT). Another key contribution is that this partitioning can be fully addressed as a continuous−time process, which allows for the dissection of how the timescale of fluctuations in these two components changesacross the brain’s processing arc.

      Strengths:

      (1) The model is cleanly derived and thoroughly documented, including usable code shared in a GitHub repo. This makes the method immediately portable to other neural systems.

      (2) This is a clear and well−presented piece of work. The figures and writing are clear and understandable, and all pieces of the derivations are included in the main text and supplementary information.

      (3) Comparisons to other models, particularly the one from Goris et al., 2014 shows how this Continuous Modulated Poisson (CMP) model outperforms previous work.

      (4) New insights about how variability partitioning changes across the visual stream from LGN to MT are revealed, including how the gain fluctuates on longer timescales in higher visual areas. Another key result about the anticorrelation between the variance in stimulus drive and gain fluctuations comports with theories about how neurons maintain efficient, reliable encoding.

      (5) In addition to the results reported here, this work will serve as an excellent tutorial for students and postdocs first delving into the sources of variability in the brain.

      We sincerely thank the reviewer for the thoughtful and positive assessment of our work. We are pleased that the model development, empirical analyses, and presentation were viewed as clear, rigorous, and useful for the broader neuroscience community. We also appreciate the reviewer’s recognition that the continuous-time formulation meaningfully extends prior variability-partitioning approaches by allowing stimulus drive and internal gain dynamics to be characterized across temporal scales. The reviewer’s comments helped us further clarify the positioning of the work, expand the Discussion of model scope and limitations, and better articulate future extensions. Below, we address the specific suggestions raised by the reviewer and describe the revisions made in the manuscript.

      Weaknesses:

      The work is somewhat incremental, building on previous studies of the partitioning of variability in the brain, but it provides important new extensions, as noted above.

      Regarding the comment on incremental contribution, we agree that our framework builds directly on previous variability-partitioning approaches, especially the modulated Poisson framework of Goris et al. However, the main goal of this work is to move this class of models from a count-based formulation to a continuous-time spike-train framework. This extension is important because it allows us to model gain as a temporally structured latent process, characterize how variability depends on the timescale over which spikes are counted, and infer the temporal covariance structure of stimulus-independent fluctuations. In addition, the CMP framework provides analytic expressions for the Fano factor as a function of bin size, introduces the EPL covariance function for slowly decaying gain dynamics, and enables direct comparisons of gain amplitude and timescale across visual areas. In the revised manuscript, we have clarified this positioning and emphasized how CMP extends prior variability-partitioning models while preserving their interpretability.

      To address this comment, we revised the Introduction (Pages 3-4: Lines 108-111 and Lines 118-121) and Discussion (Page 13: Lines 425-429) to clarify better how CMP builds on prior variability-partitioning models while extending them to continuous-time spike-train data.

      The only major gap I would suggest addressing in the Discussion is the observation of sub−Poisson variability in the brain. It seems clear that this model can extend to sub− Poisson variability and its partitioning and perhaps even show how that varies in real time, with an animal’s attentional state. That is, of course, beyond the scope of the current work, but could be mentioned in the Discussion.

      We thank the reviewer for this suggestion. We agree that sub-Poisson variability is an important phenomenon observed in neural data. Because the CMP model uses a conditionally Poisson observation model with stochastic gain modulation, it naturally captures Poisson and super-Poisson variability but does not generate sub-Poisson spike count statistics in its current form. In the revised manuscript, we have clarified this limitation in the Discussion and outlined possible extensions that could address sub-Poisson variability, including spike-history terms, renewal-process likelihoods, and alternative count distributions [Truccolo et al., 2005, Paninski et al., 2007, Aghamohammadi et al., 2024]. We also note that such extensions could allow future models to examine how sub-Poisson and super-Poisson components vary with behavioral state, attention, or arousal.

      To address this comment, we added a Discussion paragraph describing the current model’s limitation for sub-Poisson variability and possible extensions to capture it (Page 14: Lines 460-471).

      References

      Stephen Keeley, Mikio Aoi, Yiyi Yu, Spencer Smith, and Jonathan W Pillow. Identifying signal and noise structure in neural population activity with gaussian process factor models. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 13795–13805. Curran Associates, Inc., 2020.

      Robbe L. T. Goris, Corey M. Ziemba, J. Anthony Movshon, and Eero P. Simoncelli. Slow gain fluctuations limit benefits of temporal integration in visual cortex. Journal of Vision, 18(8):8–8, 08 2018. ISSN 1534-7362. doi: 10.1167/18.8.8. URL https://doi.org/10.1167/18.8.8.

      Olivier J H´enaff, Zoe M Boundy-Singer, Kristof Meding, Corey M Ziemba, and Robbe L T Goris. Representation of visual uncertainty through neural gain variability. Nat. Commun., 11(1):2513, May 2020.

      Mark M Churchland, Byron M Yu, John P Cunningham, Leo P Sugrue, Marlene R Cohen, Greg S Corrado, William T Newsome, Andrew M Clark, Paymon Hosseini, Benjamin B Scott, David C Bradley, Matthew A Smith, Adam Kohn, J Anthony Movshon, Katherine M Armstrong, Tirin Moore, Steve W Chang, Lawrence H Snyder, Stephen G Lisberger, Nicholas J Priebe, Ian M Finn, David Ferster, Stephen I Ryu, Gopal Santhanam, Maneesh Sahani, and Krishna V Shenoy. Stimulus onset quenches neural variability: a widespread cortical phenomenon. Nat. Neurosci., 13(3):369–378, March 2010.

      Robbe L T Goris, Ruben Coen-Cagli, Kenneth D Miller, Nicholas J Priebe, and M´at´e Lengyel. Response sub-additivity and variability quenching in visual cortex. Nat. Rev. Neurosci., 25(4):237–252, April 2024.

      Wilson Truccolo, Uri T. Eden, Matthew R. Fellows, John P. Donoghue, and Emery N. Brown. A point process framework for relating neural spiking activity to spiking history, neural ensemble, and extrinsic covariate effects. Journal of Neurophysiology, 93(2):1074–1089, 2005. doi: 10.1152/jn.00697.2004. URL https://doi.org/10.1152/jn.00697.2004. PMID: 15356183.

      Liam Paninski, Jonathan Pillow, and Jeremy Lewi. Statistical models for neural encoding, decoding, and optimal stimulus design. In Paul Cisek, Trevor Drew, and John F. Kalaska, editors, Computational Neuroscience: Theoretical Insights into Brain Function, volume 165 of Progress in Brain Research, pages 493–507. Elsevier, 2007. doi: https://doi.org/10.1016/S0079-6123(06)65031-0. URL https://www.sciencedirect.com/science/article/pii/S0079612306650310.

      Cina Aghamohammadi, Chandramouli Chandrasekaran, and Tatiana A. Engel. A doubly stochastic renewal framework for partitioning spiking variability. bioRxiv, 2024.

    1. eLife Assessment

      This valuable study uses repeated experience sampling to track the relationship between depressive symptoms and perceptual confidence. The central claim that depression impairs the integration of momentary confidence into a global sense of confidence is supported by solid, robust evidence. However, the largely subclinical, highly compliant sample with limited mood variability restricts the generalisability of the conclusions. This study will be of interest to researchers studying mood disorders and metacognition, as it suggests more ecologically valid ways to investigate how mental health influences metacognition.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate how depressive symptoms relate to metacognitive confidence across multiple levels of a metacognitive hierarchy. Participants completed twice-daily assessments of depressive symptoms and bi-daily assessments of a perceptual confidence task for eight weeks. The study replicates prior findings that depression is associated with lower confidence and extends these findings by suggesting that trait depression weakens the temporal persistence of local confidence signals, thereby impairing their accumulation into global confidence.

      The study addresses an important question in computational psychiatry: how disturbances in confidence contribute to persistent negative self-beliefs in depression. The longitudinal design and repeated assessments represent advances over prior cross-sectional studies. The attempt to bridge momentary confidence fluctuations with broader self-beliefs through a hierarchical metacognitive framework is novel.

      The principal contribution is the finding that trait depression moderates the lagged relationship between local and global confidence. The authors interpret this as evidence that positive fluctuations in local confidence decay more rapidly in individuals with higher depression, limiting their integration into higher-order confidence beliefs. This contribution is potentially important.

      However, several aspects of the interpretation warrant caution. The moderated mediation analysis remains correlational and does not establish that impaired local confidence persistence causally produces global under-confidence. Alternative explanations, including stable individual differences in response styles, latent third variables, or measurement properties of the confidence scales, remain plausible. The manuscript occasionally adopts language suggesting mechanistic or causal conclusions that exceed the inferential scope of the analyses.

      The temporal resolution of the study also complicates interpretation. The absence of cross-lagged effects between depression and confidence may reflect a mismatch between the timescales over which mood and metacognition interact. The authors acknowledge this possibility, but it substantially limits conclusions regarding temporal precedence.

      Furthermore, the sample size is not justified, and the sample differs considerably from populations typically studied in depression research. Participants were older, predominantly female, and self-selected citizen scientists with low and relatively stable depression scores. Consequently, it remains unclear whether the observed dynamics generalise to clinically depressed populations, where symptom severity and variability may differ substantially.

      Overall, this is a thoughtful and technically sophisticated study that provides valuable new data on the temporal organisation of confidence in relation to depression. The central findings are interesting and likely to stimulate future work, although the mechanistic interpretations would benefit from greater caution.

      Strengths:

      (1) Innovative use of dense longitudinal sampling to investigate metacognitive processes.

      (2) Large number of repeated observations per participant and good adherence over eight weeks.

      (3) Integration of EMA, multilevel vector autoregression, Bayesian modelling, and computational modelling.

      (4) Novel proposal that depression weakens the persistence of local confidence signals and their integration into global confidence.

      (5) Careful consideration of local versus global metacognitive processes.

      Weaknesses:

      (1) The causal and mechanistic claims may exceed what can be inferred from the data.

      (2) No justification is provided for the sample size, and the sample is older, predominantly female, and largely non-clinical, limiting generalisability.

      (3) The sampling intervals may not be optimally suited to detect temporal relationships between mood and confidence.

      (4) Several modelling decisions require additional justification and sensitivity analyses.

      (5) The moderated mediation framework assumes a temporal ordering that cannot be conclusively established.

    3. Reviewer #2 (Public review):

      Summary:

      Phon-Amnuaisuk and colleagues address an important question in cognitive psychology and computational psychiatry: how does the established relationship between depression and metacognitive under-confidence unfold over time? The study is grounded in a hierarchical view of metacognition, in which local confidence in individual decisions contributes to global estimates of performance. To test this, the authors used an intensive longitudinal design in which participants repeatedly reported depressive symptoms and completed a gamified perceptual decision-making task over eight weeks. This allowed them to examine whether depressive symptoms and confidence fluctuate together within individuals, whether one predicts the other over time, and whether trait depression alters the way local confidence is carried forward and integrated into global confidence. The main findings are that higher trait depression is associated with lower local and global confidence, that within-person mood fluctuations show limited temporal precedence over confidence at the two-day timescale, and that trait depression is linked to weaker temporal persistence of local confidence and reduced carry-over into later global confidence.

      Strengths:

      A major strength of the study is its repeated-measures design across a large sample. Participants completed up to 28 metacognitive task sessions over eight weeks, alongside repeated ratings of depressive symptoms. This allows the authors to separate stable between-person differences from within-person changes over time, which is a clear advantage over standard cross-sectional studies. The analytical strategy is also appropriate: the authors use multilevel vector autoregressive models to examine temporal, contemporaneous, and between-person associations; Bayesian ordinal models to test interactions; and computational modelling to examine how confidence is formed.

      The strongest results concern stable individual differences. Participants with higher average depressive symptoms reported lower local and global confidence. While this pattern is consistent with prior work showing reduced confidence in depression, the computational model further suggests that higher depression is associated with a more conservative confidence criterion: these participants required more evidence before reporting high confidence. This adds nuance by suggesting that under-confidence in depression may not reflect poorer metacognitive sensitivity, but rather a bias in how confidence is reported.

      The study also proposes a novel temporal account. Trait depression was associated with weaker autocorrelation of local confidence across sessions, which in turn reduced the extent to which local confidence carried over into later global confidence. This indirect pathway was supported by a moderated mediation analysis and was consistent across most individual depressive symptoms. The finding is theoretically interesting and supported by strong statistical evidence.

      Weaknesses:

      The main limitation concerns the interpretation of the temporal and mechanistic claims. The strongest effects are observed at the trait level, whereas the within-person temporal associations between mood and confidence are weak or absent. The study therefore provides stronger evidence that people with higher average depressive symptoms are generally less confident than evidence that momentary changes in depressive mood drive later changes in confidence, or vice versa.

      The sample also constrains the conclusions. Depression scores were strongly concentrated near the lower end of the scale, and the final sample was highly selected, with many of the initial 976 participants excluded because they did not complete enough task sessions. This means that the study may be better suited to detecting stable individual differences than dynamic mood-confidence processes.

      The temporal spacing of the metacognitive assessments is a further constraint. Because the metacognition task was administered every two days, the cross-lagged analyses could only test mood-confidence dynamics across this interval. If depressive mood and confidence influence each other over shorter timescales, such effects may have been missed. The absence of cross-lagged effects should therefore be interpreted with caution: it shows that temporal precedence was not detected at the two-day lag in this sample, but it does not rule out shorter-term directional effects or effects in more symptomatic clinical populations.

      Related to this, the manuscript moves between several related but distinct terms - "mood," "depressive mood," "depression," "depressive symptoms," and "trait depression" - without always clarifying whether these are intended as interchangeable or as conceptually distinct constructs. This matters for a study whose central claims concern temporal precedence and trait-versus-state distinctions, and it is compounded by a sample with generally low depressive symptom levels, where the boundary between transient low mood and a trait-like depressive disposition is harder to draw.

    4. Author response:

      Reviewer #1:

      We thank the reviewer for their comments. They raised an issue with the correlational nature of the analysis, suggesting alternative explanations, including stable individual differences in response styles, latent third variables, or measurement properties of the confidence scales. We agree with the reviewer’s comments that a latent third variable may confound our findings and higher-order metacognitive beliefs (which we did not assess) are a possible candidate. We will broaden our discussion to include other plausible confounds that may jointly relate to depression and confidence dynamics. Regarding stable individual differences in response styles, our analysis decomposed confidence into within and between-person components and standardised the within-person component by each participant’s own variability. This allowed the interaction and moderated mediation analyses to separate within-person associations from stable between-person differences (e.g., range of confidence scale used; Epskamp et al., 2018). With regards to measurement properties, we will conduct additional analyses investigating whether trait depression is associated with altered response patterns (e.g., non-linear or more variable mapping of latent evidence into confidence reports).

      The reviewer also noted that the temporal resolution we selected may not be optimal to measure co-fluctuation of depression and confidence. We agree this remains a challenge for the depression research (Jamalabadi et al., 2025; Tamm et al., 2024), and metacognitive confidence research (da Fonseca et al., 2023; Wright et al., 2024). We chose the interval as it allowed us to balance retention and data quality across an 8-week study (Eisele et al., 2022) and address a gap in the literature – frequent but shorter (~1-week) EMA-based assessments of existing metacognitive confidence studies (da Fonseca et al., 2023; Wright et al., 2024) precludes insights into the dynamics of mood and metacognitive confidence across longer timescales. We have explored one-day and half-day lagged associations but did not find significant effects across most symptoms for both local and global confidence. This was omitted from the original manuscript as the analyses could not symmetrically control for local/global confidence (as these were only assessed bi-daily) but will be included in the revised manuscript’s supplement.

      The reviewer noted that the temporal ordering of the moderated mediation analysis cannot be conclusively established. We agree that there is a lack of research investigating alternative orderings. We selected the local-to-global ordering because prior work has similarly modelled and demonstrated global confidence estimates as a cumulative integration of local confidence estimates across the block (Cavalan et al., 2023; Katyal et al., 2025; Lee et al., 2021; Rouault et al., 2019, 2022). Nevertheless, we modelled an alternative process (i.e. whether global confidence interacted with trait-level depression in predicting subsequent local confidence) but did not find a significant frequentist interaction effect. We will elaborate on our rationale for the current ordering and include analyses of this alternative process in our revised manuscript and supplement.

      Finally, we agree with the reviewer’s comments that the sample constraints generalisability. The revised manuscript will more clearly discuss the generalisability of our findings to clinical populations and the implications of self-selection and limited within-person variability in depressive symptoms for interpretation and future work. We will also tone down the causal language in the manuscript.

      Reviewer #2:

      We thank the reviewer for their comments. We agree with the reviewer that our strongest effects live at the trait level, whereas our cross-lagged findings provided insufficient evidence for depression driving underconfidence or vice versa (at least with a two-day lag). However, our moderated mediation finding centres on a different angle - trait level depression moderating how within-person confidence is integrated into more global beliefs. Nevertheless, we agree that the findings provide a plausible account for why underconfidence (and possibly metacognitive beliefs) remain persistent in depression, but not how either underconfidence or depression arose in the first place. We will make this clearer in our revised manuscript.

      Consistent with reviewer #1’s comments, we agree that the non-clinical nature of sample does constrain generalisability and inference. We will conduct a sensitivity analysis with a larger sample (more lenient inclusion criteria) and discuss its implications in more detail in our revised manuscript. We also agree with the reviewer’s comments that our findings do not preclude the possibility of temporal precedence and will further clarify in our revised manuscript with reference to shorter time lags. The reviewer noted that our use of terminologies (e.g., “depression”, “depressive symptoms”, “depressive mood”) was not clearly defined. Our revised manuscript will make this point clearer whilst also making the usage of terminologies consistent.

      References

      Cavalan, Q., Vergnaud, J.-C., & de Gardelle, V. (2023). From local to global estimations of confidence in perceptual decisions. Journal of Experimental Psychology: General, 152(9), 2544–2558. https://doi.org/10.1037/xge0001411

      da Fonseca, M., Maffei, G., Moreno-Bote, R., & Hyafil, A. (2023). Mood and implicit confidence independently fluctuate at different time scales. Cognitive, Affective, & Behavioral Neuroscience, 23(1), 142–161. https://doi.org/10.3758/s13415-022-01038-4

      Eisele, G., Vachon, H., Lafit, G., Kuppens, P., Houben, M., Myin-Germeys, I., & Viechtbauer, W. (2022). The Effects of Sampling Frequency and Questionnaire Length on Perceived Burden, Compliance, and Careless Responding in Experience Sampling Data in a Student Population. Assessment, 29(2), 136–151. https://doi.org/10.1177/1073191120957102

      Epskamp, S., Waldorp, L. J., Mõttus, R., & Borsboom, D. (2018). The Gaussian Graphical Model in Cross-Sectional and Time-Series Data. Multivariate Behavioral Research, 53(4), 453–480. https://doi.org/10.1080/00273171.2018.1454823

      Jamalabadi, H., Koosha, T. A., Stocker, E., Jansen, A., Ebner-Priemer, U. W., Proppert, R. K. K., Rieble, C. L., Tutunji, R., & Fried, E. I. (2025). Optimizing the frequency of ecological momentary assessments using signal processing. Psychological Medicine, 55, e358. https://doi.org/10.1017/S003329172510264X

      Katyal, S., Huys, Q. J., Dolan, R. J., & Fleming, S. M. (2025). Distorted learning from local metacognition supports transdiagnostic underconfidence. Nature Communications, 16(1), 1854. https://doi.org/10.1038/s41467-025-57040-0

      Lee, A. L. F., de Gardelle, V., & Mamassian, P. (2021). Global visual confidence. Psychonomic Bulletin & Review, 28(4), 1233–1242. https://doi.org/10.3758/s13423-020-01869-7

      Rouault, M., Dayan, P., & Fleming, S. M. (2019). Forming global estimates of self-performance from local confidence. Nature Communications, 10(1), 1141. https://doi.org/10.1038/s41467-019-09075-3

      Rouault, M., Will, G.-J., Fleming, S. M., & Dolan, R. J. (2022). Low self-esteem and the formation of global self-performance estimates in emerging adulthood. Translational Psychiatry, 12(1), 1–10. https://doi.org/10.1038/s41398-022-02031-8

      Tamm, J., Takano, K., Just, L., Ehring, T., Rosenkranz, T., & Kopf-Beck, J. (2024). Ecological Momentary Assessment versus Weekly Questionnaire Assessment of Change in Depression. Depression and Anxiety, 2024, 9191823. https://doi.org/10.1155/2024/9191823

      Wright, A. C., Palmer-Cooper, E., Cella, M., McGuire, N., Montagnese, M., Dlugunovych, V., Liu, C.-W. J., Wykes, T., & Cather, C. (2024). Experiencing hallucinations in daily life: The role of metacognition. Schizophrenia Research, Hallucinations: Neurobiology and Patient Experience, 265, 74–82. https://doi.org/10.1016/j.schres.2022.12.023

    1. eLife Assessment

      This valuable study presents evidence that the human brain encodes perceptual absence and numerical absence through distinct neural codes, while confirming that symbolic and non-symbolic forms of zero share a common representation. The evidence supporting this dissociation is convincing, drawing on neural decoding techniques and Bayesian analyses, although some concerns remain about task design and decoding method, which could partly explain the lack of shared representation. The work will be of interest to neuroscientists and psychologists studying numerical cognition and the boundary between perception and higher-level cognition.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate whether the brain uses the same neural representations for "absence" when it comes to seeing nothing versus thinking of zero. To do this, they recorded MEG while participants performed two types of tasks: (1) a perceptual detection task where subjects reported the presence or absence of a faint visual stimulus, and (2) numerical comparison tasks where subjects saw streams of numbers or dot patterns (both including "0" or empty sets) and decided which of two color-coded streams had a larger average. Multivariate decoders were trained to distinguish "present" vs "absent" in the detection task and "zero" vs "non-zero" in the numerical tasks. As a sanity check, the authors first replicate that symbolic (digit "0") and non-symbolic (empty dot sets) zeros share a common neural code (cross-format generalization). Crucially, they find that this numerical "zero" code does not overlap with the code for perceptual absence: cross-decoding between the detection task and number tasks yields Bayes factors strongly favoring distinct representations. In other words, the brain's pattern for "no grating was seen" cannot decode the pattern for "the number zero was shown," and vice versa. A small brief cross-decoding effect around 300 ms was observed, which the authors attribute to low-level visual confounds (and which they test with an additional control decoder for stimulus presence), but overall the evidence supports a dissociation.

      Strengths:

      The authors investigate whether the brain uses the same neural representations for "absence" when it comes to seeing nothing versus thinking of zero. To do this, they recorded MEG while participants performed two types of tasks: (1) a perceptual detection task where subjects reported the presence or absence of a faint visual stimulus, and (2) numerical comparison tasks where subjects saw streams of numbers or dot patterns (both including "0" or empty sets) and decided which of two color-coded streams had a larger average. Multivariate decoders were trained to distinguish "present" vs "absent" in the detection task and "zero" vs "non-zero" in the numerical tasks. As a sanity check, the authors first replicate that symbolic (digit "0") and non-symbolic (empty dot sets) zeros share a common neural code (cross-format generalization). Crucially, they find that this numerical "zero" code does not overlap with the code for perceptual absence: cross-decoding between the detection task and number tasks yields Bayes factors strongly favoring distinct representations. In other words, the brain's pattern for "no grating was seen" cannot decode the pattern for "the number zero was shown," and vice versa. A small brief cross-decoding effect around 300 ms was observed, which the authors attribute to low-level visual confounds (and which they test with an additional control decoder for stimulus presence), but overall the evidence supports a dissociation.

      Weaknesses:

      My main concern is whether the perceptual and numerical tasks are truly matched aside from their "absence" content. The perceptual task is a simple yes/no detection of a faint grating, whereas the numerical tasks involve holding two streams of 5 items in working memory and comparing their average. These tasks differ in many ways (stimulus complexity, decision rule, cognitive load), so it is possible that the lack of cross-decoding is due to general task differences rather than a fundamental "absence vs zero" dissociation. The authors do partially address this by showing that other shared aspects (like color) can cross-generalize, but one might still worry that an "absence" decision in a detection task engages different attentional or decisional mechanisms than a "zero" decision in a numerical context.

      The numerical averaging task closely resembles that used by Spitzer et al. (2017), who reported that both behavioral weighting and neural representational geometry exhibit anti-compression, with disproportionately stronger representations for larger numerosities. In contrast, the present manuscript interprets its decoding results as reflecting an ordered numerical continuum. It is therefore unclear whether the current analyses are sensitive only to ordinal structure or whether they also preserve the nonlinear representational geometry reported previously. This distinction is important because the interpretation of zero as part of a numerical continuum depends on the geometry of that continuum. The authors should clarify whether their representational analyses are compatible with the anti-compressed neural number line described by Spitzer et al., or explain why the two studies yield different conclusions.

      Spitzer, B., Waschke, L., & Summerfield, C. (2017). Selective overweighting of larger magnitudes during noisy numerical comparison. Nature Human Behaviour, 1(8), 145. https://doi.org/10.1038/s41562-017-0145

      The authors attempt to account for non-numerical visual information by controlling for Total Dot Area and Density. While this is an important control, these two variables do not exhaust the visual dimensions that covary with numerosity in dot displays. A large body of work has demonstrated that multiple continuous features, including average item area, total surface area, convex hull (field area), and density, are inherently intercorrelated and cannot all be independently controlled simultaneously (e.g., Piazza et al., 2004; Gebuis & Reynvoet, 2004; Castaldi et al., 2019; Karami et al., 2025). Consequently, controlling only two features does not fully establish that the decoded signal specifically reflects numerosity. To better characterize the stimulus space, I encourage the authors to report the correlation matrix among the principal visual features of the dot arrays (average item area, total surface area, convex hull/field area, density, and numerosity). In addition, it would be informative to quantify the unique contribution of each feature to the neural data using a multiple-regression RSA or semi-partial correlations RSA, similar to the analyses employed by Castaldi et al. (2019) and more recently by Karami et al. (2025). Such analyses would provide a more rigorous assessment of whether the decoded representations uniquely reflect numerosity after accounting for correlated visual properties.

      Piazza, M., Izard, V., Pinel, P., Bihan, D. L., & Dehaene, S. (2004). Tuning curves for approximate numerosity in the human intraparietal sulcus. Neuron, 44(3), 547-555. https://doi.org/10.1016/j.neuron.2004.10.014

      Gebuis, T., & Reynvoet, B. (2011). The interplay between nonsymbolic number and its continuous visual properties. Journal of Experimental Psychology General, 141(4), 642-648. https://doi.org/10.1037/a0026218

      Castaldi, E., Piazza, M., Dehaene, S., Vignaud, A., & Eger, E. (2019). Attentional amplification of neural codes for number independent of other quantities along the dorsal visual stream. eLife, 8. https://doi.org/10.7554/elife.45160

      Karami, A., Castaldi, E., Eger, E., & Piazza, M. (2025). Distinct neural representational geometries of numerosity in early visual and association regions across visual streams. Communications Biology, 8(1), 1029. https://doi.org/10.1038/s42003-025-08395-z

      The manuscript reports predominantly diagonal temporal generalization for non-symbolic numerosity, implying a rapidly evolving neural code. However, a recent study using time-resolved decoding of numerical representations (Karami et al., 2025) reported substantial off-diagonal temporal generalization, consistent with a temporally stable representational format. Although methodological differences between the studies may account for this discrepancy, the apparent contrast deserves discussion. In particular, it would be useful for the authors to clarify whether the differences arise from task demands, stimulus characteristics, preprocessing and decoding procedures, or from theoretical differences in what is being decoded. More generally, these findings raise the possibility that the temporal stability of numerical representations is task-dependent rather than fixed. If so, it would be interesting to discuss whether task demands might also influence the relationship between perceptual and conceptual representations of absence. Such a possibility could help explain why cross-decoding was not observed in the present study and suggests an interesting direction for future research.

      Karami, A., Castaldi, E., Eger, E., Hebart, M., & Piazza, M. (2025). Numerosity Is Directly Sensed and Dynamically Transformed in the Human Brain: Evidence from MEG-MRI Fusion. bioRxiv (Cold Spring Harbor Laboratory). https://doi.org/10.1101/2025.11.15.687894

      Throughout the manuscript, the authors appear to treat non-symbolic numerosity as a conceptual representation and contrast it with perceptual absence. I find this interpretation insufficiently justified. A substantial body of behavioral (Anobile et al., 2013; Cicchini et al., 2016) and neuroimaging (Piazza et al., 2004; Castaldi et al., 2019; Karami et al., 2025) research has argued that non-symbolic numerosity is represented as a perceptual attribute extracted relatively early in the visual processing hierarchy, even if its precise computational origin remains debated. Consequently, it is not immediately clear why non-symbolic numerosity should be regarded as a conceptual representation comparable to symbolic number or the concept of zero. This distinction is important because it directly affects the interpretation of the negative cross-decoding results. If both perceptual absence and non-symbolic numerosity are primarily perceptual representations, the absence of cross-decoding cannot be taken as evidence that perceptual and conceptual absence are represented differently. Rather, it may simply indicate that these two perceptual representations encode different visual attributes. I therefore encourage the authors to clarify their theoretical position regarding the representational status of non-symbolic numerosity and to discuss how their interpretation relates to influential theories of numerical cognition that conceptualize non-symbolic numerosity as an early perceptual representation rather than an abstract conceptual one.

      Anobile, G., Cicchini, G. M., & Burr, D. C. (2013). Separate mechanisms for perception of numerosity and density. Psychological Science, 25(1), 265-270. https://doi.org/10.1177/0956797613501520

      Cicchini, G. M., Anobile, G., & Burr, D. C. (2016). Spontaneous perception of numerosity in humans. Nature Communications, 7(1), 12536. https://doi.org/10.1038/ncomms12536

      Piazza, M., Izard, V., Pinel, P., Bihan, D. L., & Dehaene, S. (2004). Tuning curves for approximate numerosity in the human intraparietal sulcus. Neuron, 44(3), 547-555. https://doi.org/10.1016/j.neuron.2004.10.014

      Castaldi, E., Piazza, M., Dehaene, S., Vignaud, A., & Eger, E. (2019). Attentional amplification of neural codes for number independent of other quantities along the dorsal visual stream. eLife, 8. https://doi.org/10.7554/elife.45160

      Karami, A., Castaldi, E., Eger, E., Hebart, M., & Piazza, M. (2025). Numerosity Is Directly Sensed and Dynamically Transformed in the Human Brain: Evidence from MEG-MRI Fusion. bioRxiv (Cold Spring Harbor Laboratory). https://doi.org/10.1101/2025.11.15.687894

      I have two related concerns regarding the discussion of Paul et al. (2022). First, I think it would be helpful to describe more explicitly what was measured in that study. To my understanding, Paul et al. quantified the aggregate Fourier power (AFP) of the stimuli. Moreover, AFP has primarily been discussed in the context of dot arrays with constant dot size within each stimulus. In the current manuscript, it is not entirely clear from the Methods whether dot sizes vary within displays. I therefore encourage the authors to explicitly describe how dot sizes were generated and varied across stimuli. If AFP is correlated with numerosity in the present stimulus set, it would also be helpful to explain how the analyses dissociate neural representations of numerosity from those potentially driven by AFP. Second, I am not entirely convinced by the argument that training a classifier to distinguish Hits from Correct Rejections is sensitive to aggregate Fourier power. It would be helpful if the authors could explain more explicitly why this decoding contrast should be expected to be sensitive to AFP. As currently written, the logical connection between the AFP hypothesis and the proposed control analysis is not entirely clear.

    3. Reviewer #2 (Public review):

      The authors tackle the question of whether conceptual absence is neurally encoded in the same way as perceptual absence. On one hand, the neural bases of perceptual absence have been largely investigated, as exemplified by the study of neural correlates of perception of aware vs. unaware stimuli, and on the other hand, the overlapping neural encoding of symbolic ('0') and non-symbolic (number of dots) formats of numerical absence has been previously established (Barnett & Fleming, 2024). However, the direct comparison of neural representations between numerical and perceptual absences remained to be investigated.

      This article fills this gap by designing a Magneto-EncephaloGraphy (MEG) study using multi-voxel pattern analysis (MVPA) and temporal generalization to probe the similarity of neural patterns across the representation of perceptual absence (lack of stimuli), symbolic ('0'), and non-symbolic (number of dots) formats of numerical absence. They confirmed previously obtained evidence for shared neural representation across both formats of numerical absence. They report evidence for an absence of shared representation between both formats of numerical absence on one hand and perceptual absence on the other hand, while controlling for the confounding effect of low-level visual features. Their results overall support the conclusion that conceptual and perceptual absence are neurally encoded in a distinct way and speak in favour of a boundary between the representation of the concepts and the percept of absence.

      Major strengths:

      (1) Behavioral and neuroimaging results convincingly demonstrate that neural encoding of symbolic and non-symbolic absences is shared and situated on a graded, abstract neural number line, replicating previous results, notably from the authors themselves (Barnett & Fleming, 2024).

      (2) They show that neural encoding of perceptual and numerical absence do not generalise across each other, while controlling for spurious confounds due to visual stimuli that are commonly shared in the cases of non-symbolic numerical absence and perceptual absences.

      (3) They adequately use Bayesian analysis to distinguish absence of evidence vs. evidence of absence to support their claim.

      (4) The interpretation of numerical absence as a representation of the concept of "nothingness" is adequate, although it might be further discussed by distinguishing the concept of "zero" on a number line from the concept of nothingness and that of an empty set (Nieder, 2016).

      (5) The discussion about development and metacognition paves an interesting road for further investigation on the acquisition of the concept of zero, especially in light of debates on the progressive development of metacognitive abilities in children (Goupil & Kouider, 2019).

      (6) Data and code are published with open-source access, allowing the community to further investigate the points as major weaknesses evoked below, if desired.

      Major weaknesses:

      (1) Interestingly, restricting the neural decoding method to the alpha band shows distinct representations across formats of numerical absence. This begs for providing more details on how neural representations of perceptual, symbolic, and non-symbolic absences differ at the source and frequency level and to report the decoding weights to better assess what drives the performance of the neural decoding algorithm in each case and whether they overlap with each other.

      (2) Task-demands between perceptual (present vs absent) and numerical (lower vs higher) are different, raising concerns about whether these aspects of experimental design could drive, at least partially, the shared representational patterns across numerical representation of absence and their distinction from perceptual representation of absence.

      (3) The same argument can also be raised for the way that the neural decoders of perceptual and numerical absences are trained and tested. Both formats of numerical absence are built using the same procedure (zero vs. rest) and differ from the way the decoder is built for perceptual absence (hits vs misses), which might possibly drive the difference observed here.

      (4) It is thus unknown whether the claim supporting the evidence of absence holds as long as other counterfactual hypotheses that might drive these results are not excluded, such as the nature of the decoded features, the effect of task demands, or the way the neural decoder is trained, as mentioned above.

      Overall, I was pleased by the quality of the methods and the clarity with which the question of the boundary between cognition and perception is addressed for the case of numerous and perceptual absence. While the methods used are well established in the field, the choice and rigor of their analysis and the controls provided stand as a convincing methodology to test their hypotheses, although they do not fully exclude alternative interpretations nor explore the wider extent of possibilities that may provide exhaustive evidence for showing that perceptual and numerical absences are distinctly encoded in the brain.

      This work will be of great appeal to neuroscientists interested in comparing the representation of percepts and concepts across different formats, to psychologists interested in the origin of number representation, and to philosophers interested in debates on the boundary between cognition and perception.

      Nieder, A. (2016). Representing something out of nothing: The dawning of zero. Trends in Cognitive Sciences, 20(11), 830-842.

      Goupil, L., & Kouider, S. (2019). Developing a reflective mind: From core metacognition to explicit self-reflection. Current Directions in Psychological Science, 28(4), 403-408.

      Barnett, B., & Fleming, S. M. (2024). Symbolic and non-symbolic representations of numerical zero in the human brain. Current Biology, 34(16), 3804-3811.

    1. eLife Assessment

      This valuable study presents a comparative dataset on crab locomotion to investigate the evolution of sideways walking. The evidence supporting the authors' claims is convincing. This work will be of broad interest to researchers in animal locomotion and evolutionary biology.

    2. Reviewer #1 (Public review):

      Summary:

      This is an interesting and well-written manuscript in which the authors set out to answer a simple, longstanding question with a modern comparative approach. Namely where in crab evolution did sideways walking arise, how often has it been lost or regained, and is its evolution plausibly associated with the ecological and taxonomic success of true crabs. To address these questions the authors recorded locomotion from 50 live species, quantified the predominant direction of locomotion, and mapped these behavioral states onto a recent crab phylogeny to reconstruct the likely evolutionary history of sideways walking. The revised manuscript also includes analyses of the underlying movement-angle distributions and tests whether the evolutionary conclusions depend on the original behavioral classification scheme.

      Strengths:

      The strongest part of the study remains the dataset itself. Comparable behavioral measurements across dozens of crab species are rare, and obtaining and recording live representatives from this range of taxa required substantial field, aquarium, and husbandry effort. The overall pattern that emerges, in which most true crabs are strongly biased toward sideways locomotion while several specialized lineages move predominantly forward, is interesting and likely to be useful to researchers studying animal locomotion, functional morphology, and behavioral evolution.

      The revised analyses substantially strengthened the manuscript. In the original version, I was concerned that the main behavioral classification depended too strongly on first assigning individual movements to forward or sideways bins using a fixed angular boundary. The authors have now analyzed the underlying continuous movement-angle distributions and have shown that, although mixed directional tendencies are present in some species, most taxa have a dominant directional preference. They also derived a separate, data-informed boundary from the distribution of dominant movement directions. I favor this alternative approach and it produces the same classification of species as the original index-based method, providing useful evidence that the main evolutionary reconstruction is not simply an artifact of the original 60{degree sign} cutoff.

      The authors also responded appropriately to the limitation that locomotion was measured from one individual per species. This sampling design cannot establish the full extent of within-species, ontogenetic, or size-dependent variation, but the revised manuscript now states this limitation clearly and restricts its conclusions to broad interspecific patterns in predominant locomotor direction. This is a more appropriate interpretation of the available sampling.

      The phylogenetic analysis provides a reasonable framework for addressing the main evolutionary question. Taken together, the behavioral and phylogenetic results support the conclusion that sideways locomotion likely arose once within the lineage leading to true crabs and was followed by multiple reversions toward predominantly forward locomotion in specialized groups. The manuscript therefore makes a convincing case that sideways walking is not simply an inevitable consequence of possessing a crab-like body plan.

      Weaknesses:

      My main remaining reservation concerns the interpretation of forward and sideways locomotion as two discrete biological modes. The revised analyses convincingly show that species can be classified according to their predominant direction of locomotion and that this classification is robust to alternative analytical approaches. However, this does not necessarily demonstrate that forward and sideways locomotion represent two intrinsically discrete or mutually exclusive behavioral modes. Indeed, the new analyses show that many taxa are better described by two-component movement-angle distributions, even though most of these have one dominant component. This is consistent with strong directional preferences, but it also indicates that mixed movement strategies are common. The supplementary circular distributions similarly show considerable variation in the shape and breadth of directional preferences among taxa. Having said that, I do not think this substantially weakens the central evolutionary conclusion. The phylogenetic analysis requires a defensible classification of predominant locomotor direction, and the revised analyses now provide one. The evolutionary story remains interesting whether the underlying behavioral variation consists of two sharply discrete modes or a broader continuum of directional strategies with strong clustering toward forward and sideways movement.

      A second limitation is that the proposed relationship between sideways locomotion and diversification remains necessarily correlational. The revised manuscript handles this more cautiously than the original version and now frames sideways locomotion as a plausible key innovation whose emergence is associated with the exceptional diversity of true crabs, rather than as a demonstrated causal driver of diversification. This distinction is important because differences in species richness among lineages can also reflect ecological opportunity, extinction history, and other lineage-specific factors. The revised framing is therefore better aligned with the strength of the evidence.

      Final assessment: Overall, this is a valuable comparative study with an unusually broad behavioral dataset. The revisions have addressed the principal methodological concerns raised in the original review, particularly by analyzing continuous movement directions and demonstrating that the main phylogenetic classification is robust to an alternative, data-informed approach. The evidence now convincingly supports the central conclusion concerning the evolutionary origin and repeated reversal of predominant locomotor direction in crabs, although the stronger interpretation that forward and sideways locomotion represent two strictly discrete biological modes remains less certain.

    3. Reviewer #2 (Public review):

      Summary:

      The current work investigates the evolution of sideward locomotion in Brachyura in light of a single evolutionary origin. To this end, the authors first analysed the mode of locomotion in 50 crab species and observed mutually exclusive presence of sideways vs. forward movement. The phylogenetic analysis confirmed that there is indeed a single evolutionary origin for sideways movement, which was sometimes followed by several reversions to forward locomotion. This way, authors demonstrate how locomotor movement modes shape evolutionary diversification in animals by showing that species richness is much higher in side-ways-moving crabs than in the nearest groups. This is an interesting work that integrates behavioural analysis and phylogenetic relations, capitalising largely on crabs.

      Original questions/suggestions:

      Firstly, I think the paper spends too much time on a straightforward analysis of the mode of locomotion. I was also wondering whether the phylogenetic analysis could be simply achieved by maximising an objective function in which the modes of movement are inversely coded for two putative groups, with all values calculated at all possible nodes.

      Unfortunately, I find that the authors did not sufficiently discuss differences in the ecological niches of species with forward vs. sideways locomotion modes (including challenges of locomotion and substrate).

      Likewise, what are the anatomic correlates of forward vs. sideways locomotion? For instance, how are the advantages assumed for sideways movement associated with a flattened body? Is it possible that the mode of motion is secondary to flattened/narrow body structure, which basically limits the distance between legs and thus makes the forward movement difficult - under this logic, the mode of movement would be a secondary phenomenon to body shape traits. How can one differentiate between this alternative and the one that puts the mode of movement in the centre of the story? On a related note, how do different modes of movement relate to the ability to fit into tight spaces - how does it relate to differences in leg joints?

      Is it possible that the sideways movement maximises the scanned visual field per unit time/displacement, which may be beneficial for mostly forward-moving predators?

      Briefly, although I find the study interesting, the presented complexity may not be necessary given the endpoints; it can be achieved much more simply. Furthermore, the degree to which the conceptual analysis of different modes of locomotion was exercised was limited. The general approach may serve as a good model for the evolutionary analysis of other traits. The demonstration of traceability of the relations in question is a major contribution of the work.

      Comment on revised version:

      I am not fully convinced by the authors' handling of the complexity of the paper, but this seems like a moot point.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting and well-written manuscript in which the authors set out to answer a simple, old question with a modern toolkit: where in crab evolution did sideways walking arise, how often has it been lost or regained, and is it plausibly linked to the ecological and taxonomic success of true crabs. To do this, they record locomotion from 50 live species, convert each species' movements into a quantitative index that compares forward versus sideways bouts, and then map the resulting states onto a recent crab phylogeny to infer the most likely evolutionary history of locomotor direction.

      We thank the reviewer for this positive summary of the study and for recognizing the value of our comparative behavioral dataset and phylogenetic approach.

      Strengths:

      The strongest part of the study is the dataset itself. Comparable behavioral measurements across dozens of crab species are rare. The authors have done the field and husbandry work needed to make this possible. The overall pattern they recover, that most true crabs are strongly biased toward sideways movement (while a smaller set of lineages move predominantly forward), is interesting and likely to be useful to others. The phylogenetic mapping is also a reasonable way to address the "how many times" question (although this is peripheral to my expertise). The manuscript makes a convincing case that sideways locomotion is not simply a trivial byproduct of a crab-like body plan.

      We appreciate the reviewer’s recognition of the dataset and the overall value of the study. We have revised the manuscript to make the conclusions more robust and better aligned with the strength of the evidence.

      (1) Where I am less convinced is in how strongly the authors describe the discreteness of the behavioral categories and the absence of intermediates. The manuscript states that the Forward-Sideways Index shows a clear separation between two locomotor types with little evidence for intermediates, and it cites a statistical test rejecting a single peak in the distribution. However, the histogram in Figure 3 appears structured within each labeled category, with subclusters inside both the forward and sideways groups rather than a single tight peak per group. This matters because the index is built by first placing each movement bout into "forward" versus "sideways" bins using a fixed angle boundary and then collapsing the result into a single ratio. That approach is simple and transparent enough, but it can also hide mixed strategies. For example, a species that produces substantial amounts of both forward and sideways walking can still end up with a strongly positive or negative index, and therefore be classified as a pure "type," even though the underlying behavior is mixed. In that context, rejecting a single peak in the across-species distribution does not, by itself, justify the stronger claim that intermediates are rare or absent.

      Related to this, a key methodological choice is the use of 60 degrees as the cutoff between forward and sideways bouts. This boundary may be reasonable as a convention, but the paper does not explain why it is the right place to draw the line, and there is a plausible biological concern that a fixed angular cutoff does not mean the same thing across taxa.

      Crabs vary in body shape and in how the legs are arranged around the body. In my own comparative work, for example, some species show an elliptical stance pattern elongated along the preferred direction of travel, while others show a more circular leg arrangement, and the latter can express more mixed forward and sideways behavior. When limb arrangement and body geometry differ across species, the same measured angle can correspond to different underlying mechanics and different functional "degree of sidewaysness." The practical implication is that the reported binary separation may partly reflect the imposed classification rule, rather than a sharp biological divide.

      We thank the reviewer for this important point. We agree that the across-species distribution of FSI values alone does not justify a strong statement that intermediate or mixed locomotor tendencies are absent. We also agree that reducing continuous bout-angle distributions to a single index could potentially obscure mixed directional strategies. We have therefore revised the manuscript to avoid implying a strict absence of intermediates and have added an additional analysis of the underlying continuous angle distributions (Abstract, lines 27-29; Results, lines 191-208; Table S2).

      Specifically, we fitted one- and two-component mixture models to the continuous bout-angle distributions of each taxon and examined the supported number of components, peak locations, and mixture weights (Results, lines 196-204; Table S2). This analysis showed that 14 taxa were best described by a one-component model, whereas 36 taxa were best described by a two-component model. Importantly, among the 36 taxa best described by a two-component model, 33 had a dominant component explaining at least 70% of the distribution, whereas only three taxa showed relatively balanced two-component distributions. Thus, although some taxa do show mixed directional tendencies, most taxa are dominated by a primary directional component rather than showing an even mixture of forward and sideways locomotion.

      We also clarified the rationale for the 60° threshold used in the FSI calculation (Methods, lines 125-130). This threshold was not intended to represent a taxon-specific biological boundary between forward and sideways locomotion. Rather, it was used to divide the 360° space into three equal directional sectors: forward, sideways, and backward. This equal partitioning provides a consistent reference under a null expectation of uniformly distributed movement directions.

      To further assess whether our classification depended on the original FSI-based classification, we performed an additional data-driven check based on continuous bout-angle distributions (Results, lines 204-208; Fig. S2). We extracted the dominant peak location from each taxon’s continuous bout-angle distribution (Table S2) and estimated a boundary from the distribution of these dominant peak locations using a Gaussian mixture model. This yielded a data-informed cutoff of approximately 49.4°. The resulting peak-based classification was identical to the original FSI-based classification, with 15 forward-moving and 35 sideways-moving taxa. Thus, no taxon changed category under this independent classification approach.

      (2) Another limitation that affects interpretation is the decision to use one individual per species. I understand the logistics, and for some questions, a single representative individual can be a reasonable first pass. But it is not strong support for negative claims about intermediates, especially in a group where individuals can change substantially with growth and allometry. Crabs can grow dramatically, often with pronounced allometric shifts in limb proportions that can alter the center of mass location. Size alone can alter the kinematics and choice of locomotor behaviors in crustaceans. In species where appendage proportions change with size, or where certain legs become disproportionately large (or calcified), it is plausible that locomotor direction and the distribution of movement angles shift across ontogeny. That makes it hard to treat a single individual as a complete description of a species-level strategy, particularly for species that fall closer to the boundary between categories.

      We thank the reviewer for raising this important limitation. We agree that using one representative individual per species cannot capture the full range of within-species variation, including ontogenetic, size-dependent, or allometric changes in locomotor behavior. We also agree that this limitation is particularly relevant to strong claims about the absence of intermediates.

      As described in our response to Comment #1, we have therefore toned down statements implying a strict absence of intermediates and added analyses of the underlying continuous angle distributions (Abstract, lines 27-29; Results, lines 191-208; Table S2). These additional analyses showed that some taxa do exhibit mixed directional tendencies, although most taxa were dominated by a primary directional component.

      We have also revised the manuscript to clarify the scope of our conclusions. Specifically, we now state that our single-individual sampling design does not capture possible ontogenetic, size-dependent, or allometric variation within species (Methods, lines 105-107). We also clarify that our conclusions are intended to identify broad interspecific patterns in the predominant direction of locomotion across major brachyuran lineages, rather than to describe the full range of locomotor variation within each species (Methods, lines 107-108). Thus, we no longer treat a single individual as providing a complete description of species-level behavioral variation, but instead use it as a standardized representative observation for broad comparative and phylogenetic analyses.

      In sum, this is a valuable and useful behavioral comparative study with a dataset that many in the field will appreciate. The main conclusions about the likely evolutionary placement of sideways walking are plausible, but several of the stronger claims about discrete locomotor types, the absence of intermediates, and the relationship to diversification would be more convincing if the analysis were less dependent on a fixed angular cutoff and on single individuals per species, or if the manuscript framed those points more cautiously so the conclusions track the strength of the evidence.

      We thank the reviewer for this constructive summary and for recognizing the value of our behavioral comparative dataset. We have addressed these concerns in detail in our responses above and revised the manuscript to make the main claims better aligned with the strength of the evidence.

      Reviewer #2 (Public review):

      Summary:

      The current work investigates the evolution of sideward locomotion in Brachyura in light of a single evolutionary origin. To this end, the authors first analysed the mode of locomotion in 50 crab species and observed mutually exclusive presence of sideways vs. forward movement. The phylogenetic analysis confirmed that there is indeed a single evolutionary origin for sideways movement, which was sometimes followed by several reversions to forward locomotion. This way, authors demonstrate how locomotor movement modes shape evolutionary diversification in animals by showing that species richness is much higher in side-ways-moving crabs than in the nearest groups. This is an interesting work that integrates behavioural analysis and phylogenetic relations, capitalising largely on crabs. I have a few suggestions and questions.

      We thank the reviewer for the positive assessment of the study and for recognizing the value of integrating behavioral analysis with phylogenetic relationships. We address the specific suggestions and questions below.

      (1) Firstly, I think the paper spends too much time on a straightforward analysis of the mode of locomotion.

      We agree that the final classification of taxa into predominantly forward- and sideways-moving groups is conceptually simple. However, because our study compares locomotor behavior across a broad range of crab taxa and then uses these behavioral data for phylogenetic reconstruction, we considered it important to describe the behavioral quantification in a transparent and reproducible way. The purpose of this section is therefore not to make a simple endpoint unnecessarily complex, but to show how discrete locomotor states were derived from raw trajectory data using a standardized procedure. For this reason, we retained the current analytical description.

      (2) I was also wondering whether the phylogenetic analysis could be simply achieved by maximising an objective function in which the modes of movement are inversely coded for two putative groups, with all values calculated at all possible nodes.

      The proposed objective-function approach may be useful for identifying a node that best separates two predefined locomotor groups. However, in the present study, we aimed not only to locate a possible boundary between forward- and sideways-moving lineages, but also to reconstruct the evolutionary history of locomotor transitions under an explicit phylogenetic model.

      For this reason, we used standard ancestral state reconstruction and stochastic character mapping rather than maximizing an ad hoc objective function across possible nodes. This approach allowed us to compare alternative transition-rate models (ER and ARD), estimate uncertainty in ancestral states at internal nodes, and quantify the posterior distribution of gains and reversals. We therefore retained the current phylogenetic framework, as it provides a model-based and more informative reconstruction of locomotor evolution across true crabs.

      (3) Unfortunately, I find that the authors did not sufficiently discuss differences in the ecological niches of species with forward vs. sideways locomotion modes (including challenges of locomotion and substrate).

      Likewise, what are the anatomic correlates of forward vs. sideways locomotion? For instance, how are the advantages assumed for sideways movement associated with a flattened body? Is it possible that the mode of motion is secondary to flattened/narrow body structure, which basically limits the distance between legs and thus makes the forward movement difficult - under this logic, the mode of movement would be a secondary phenomenon to body shape traits. How can one differentiate between this alternative and the one that puts the mode of movement in the centre of the story? On a related note, how do different modes of movement relate to the ability to fit into tight spaces - how does it relate to differences in leg joints?

      Is it possible that the sideways movement maximises the scanned visual field per unit time/displacement, which may be beneficial for mostly forward-moving predators?

      We thank the reviewer for this helpful comment. We agree that the previous version did not sufficiently address the possible relationship between locomotor mode and body shape, especially the alternative explanation that sideways locomotion may be secondary to carapace flattening. In response, we added a new morphological analysis using two carapace shape indices: relative carapace length (CL/CW) and relative carapace depth (CD/CS) (Methods, lines 178–185). In the revised Results, we report that relative carapace length differed significantly between forward- and sideways-moving taxa (phylogenetically informed ANOVA: F = 26.90, p < 0.001), whereas relative carapace depth did not differ significantly between the two groups (F = 1.18, p = 0.403) (Results, lines 209–214; Fig. S3). We also added this interpretation to the Discussion, noting that locomotor mode is associated with some aspects of carapace shape but is not explained by simple carapace flattening alone (Discussion, lines 318–325).

      We also revised the Discussion to address the reviewer’s suggestions about possible functional advantages of sideways locomotion beyond rapid bidirectional escape. Specifically, we now mention that other possible advantages may include movement through confined spaces and visual-field sampling during locomotion (Discussion, lines 310-318).

      Finally, we retained the existing discussion of ecological specializations in forward-moving lineages, including coordinated collective movement in soldier crabs, decoration and concealment in majoid crabs, and life inside confined host spaces in pea crabs. This discussion supports the broader point that the adaptive value of sideways locomotion may depend on ecological context.

      (4) It is really difficult to decipher the information contained in the nodes (circles) in the printed black-and-white version of the manuscript.

      We have changed the color scheme and strengthened the outlines of the node pie charts so that the ancestral-state probabilities can be more easily distinguished (Fig. 5). We also applied the same revised color scheme to Figure 3 and Figure S4 for consistency across the manuscript.

      (5) Briefly, although I find the study interesting, the presented complexity may not be necessary given the endpoints; it can be achieved much more simply. Furthermore, the degree to which the conceptual analysis of different modes of locomotion was exercised was limited. The general approach may serve as a good model for the evolutionary analysis of other traits. The demonstration of traceability of the relations in question is a major contribution of the work.

      We thank the reviewer for this constructive summary and for recognizing the broader value of our approach. We have addressed the methodological and conceptual points raised here in our responses to the specific comments above.

      Strengths:

      The research question and the novel combination of different data types.

      We thank the reviewer for highlighting the research question and the novel combination of different data types as strengths of the study. We have revised the manuscript to further strengthen this integrative framework.

      Weaknesses:

      The complexity of the methods used, along with a limited discussion of the potential dynamics that may underlie the evolution of the sideways movement mode.

      We have addressed these concerns in our responses to the specific comments above, particularly by clarifying the rationale for the behavioral quantification and expanding the discussion of morphology, ecological context, and functional hypotheses.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Unbiased analysis of angle distribution. The authors already extract continuous bout angles prior to binning. I recommend using these distributions directly to assess modality at the species level (e.g., unimodal vs bimodal, peak locations, mixture weights) before collapsing behavior into the Forward-Sideways Index. Even a simple circular density estimate or mixture model would clarify whether species classified as "forward" or "sideways" are behaviorally pure or mixed, and would provide a quantitative basis for claims about intermediacy.

      We added mixture-model analyses of the continuous bout-angle distributions, including modality, peak locations, and mixture weights (Results, lines 196-204; Table S2).

      (2) Justify or stress-test the 60° cutoff. The manuscript should either provide a clear biological or data-driven justification for using 60° as the boundary between forward and sideways bouts, or demonstrate that the main conclusions are robust to reasonable alternative cutoffs (e.g., 45°, 75°). A brief sensitivity analysis in the supplement would be sufficient and would greatly strengthen confidence in the classification. Alternatively (my preference) would be to let the data inform the cutoff.

      We clarified the rationale for the 60° sector definition used to calculate FSI (Methods, lines 125-130; Fig. 2) and added an independent data-driven boundary analysis based on dominant peak locations, which yielded the same forward/sideways classification (Results, lines 204-208; Fig. S2).

      (3) Sampling justification. I recommend explicitly acknowledging that sampling a single individual per species limits the ability to detect ontogenetic, size-dependent, or allometric variation in locomotor strategy. If feasible, adding even limited replication across size classes or individuals for a small subset of taxa (particularly those near the classification boundary) would substantially strengthen the conclusions; otherwise, the manuscript should more clearly delimit which claims do and do not rely on the assumption of within-species invariance.

      We clarified that our conclusions concern broad interspecific patterns of predominant locomotor direction, rather than the full range of within-species variation (Methods, lines 102-108).

      (4) I would suggest toning down or reframing statements about "no intermediates". If additional analyses are not added, I recommend revising statements that imply a strict absence of intermediates to language that reflects what is directly shown (e.g., bimodality in an index derived from binned data). This would better align the claims with the current evidence.

      We revised the manuscript to avoid implying a strict absence of intermediates and now acknowledge that some taxa show mixed directional tendencies (Abstract, lines 27-29; Results, lines 191-208).

      (5) Framing and claims about diversification. The discussion of sideways locomotion as a key innovation would benefit from clearer separation between observed correlations and causal inference. If trait-dependent diversification analyses are not added, I suggest consistently framing this section as a hypothesis supported by comparative patterns rather than a demonstrated mechanism.

      We revised the Discussion to more clearly frame sideways locomotion as a possible key innovation associated with diversification, rather than as a demonstrated causal mechanism (Abstract, lines 31-34; Discussion, lines 294-309).

    1. eLife Assessment

      This important study investigates how the size of an LLM may influence its ability to model the human neural response to language recorded by ECoG. Overall, solid evidence is provided that larger language models can better predict the human ECoG response. This study will be of interest to both neuroscientists and psychologists who work on language comprehension and computer scientists working on LLMs.

    2. Reviewer #1 (Public review):

      Summary:

      The authors perform an analysis of the relationship between the size of an LMM and the predictive performance of an ECoG encoding model made using the representations from that LMM. They find a logarithmic relationship between model size and prediction performance, consistent with previous findings in fMRI. They additionally observe that as the model size increases, the location of the "peak" encoding performance typically moves further back into the model in terms of percent layer depth, an interesting result worthy of further analysis into these representations.

      Strengths:

      The evidence is quite convincing, consistent across model families and complementary to other work in this field. This sort of analysis for ECoG is needed and supports the decade-long enduring trend of the "virtuous cycle" between neuroscience and AI research, where more powerful AI models have consistently yielded more effective predictions of responses in the brain. The lag analysis showing that optimal lags do not change with model size is a nice result using the higher temporal resolution of ECoG compared to other methods like fMRI.

      Comments on revised version.

      After the latest revision, I am pleased to remove my previous remarks about weaknesses of the paper, as I believe the additional data scaling analysis, discussion of layerwise trends, and other additional commentary makes the paper a compelling addition to the literature.

    3. Reviewer #2 (Public review):

      Summary:

      This paper investigates whether large language models (LLMs) of increasing size more accurately align with brain activity during naturalistic language comprehension. The authors extracted word embeddings from LLMs for each word in a 30-minute story and regressed them against electrocorticography (ECoG) activity time-locked to each word as participants listened to the story. The findings reveal that larger LLMs more effectively predict ECoG activity, reflecting the scaling laws observed in other natural language processing tasks.

      Strengths:

      (1) The study compared model activity with ECoG recordings, which offer much better temporal resolution than other neuroimaging methods, allowing for the examination of model encoding performance across various lags relative to word onset.

      (2) The range of LLMs tested is comprehensive, spanning from 82 million to 70 billion parameters. This serves as a valuable reference for researchers selecting LLMs for brain encoding and decoding studies.

      (3) The regression methods used are well-established in prior research, and the results demonstrate a convincing scaling law for the brain encoding ability of LLMs. The consistency of these results after PCA dimensionality reduction further supports the claim.

      Comments on revised version.

      I thank the authors very much for their efforts in addressing my comments. One remaining concern is the extent of the paper's conceptual advance. Several recent studies have made broadly similar claims regarding the increasing alignment between large language models and human language processing, although using fMRI data (Antonello et al., 2023; Gao et al., 2025). I would therefore encourage the authors to more clearly articulate what additional insights are gained from using ECoG. Clarifying this point would help better establish the novelty and contribution of the present study.

      Antonello, R. J., Vaidya, A. R., & Huth, A. G. (2023). Scaling laws for language encoding models in fMRI. Advances in Neural Information Processing Systems, 36, 21895-21907.

      Gao, C., Ma, Z., Chen, J., Li, P., Huang, S., & Li, J. (2025). Increasing alignment of large language models with language processing in the human brain. Nature Computational Science, 5(11), 1080-1090.

    4. Reviewer #3 (Public review):

      This manuscript studies the connection between neural activity collected through electrocorticography and hidden vector representations from autoregressive language models, with the specific aim of studying the influence of language model size on this connection. Neural activity was measured from subjects that listened to a segment from a podcast, and the representations from language models were calculated using the written transcription as the input text. The ability of vector representations to predict neural activity was evaluated using 10-fold cross-validation with ridge regression models.

      The main results are that (as well summarized in section headings):<br /> (1) Larger models predict neural activity better.

      (2) The ability of language model representations to predict neural activity differs across electrodes and brain regions.

      (3) The layer that best predicts neural activity differs according to model size, with the "SMALL" model showing a correspondence between layer number and the language processing hierarchy.

      (4) There seems to be a similar relationship between the time lag and the ability of language model representations to predict neural activity across models.

      Strengths:

      (1) The experimental and modeling protocols generally seem solid, which yielded results that answer the authors' primary research question.

      (2) Electrocorticography data is especially hard to collect, so these results make a nice addition to recent functional magnetic resonance imaging studies.

      Weaknesses:

      (1) The interpretation of some results seems unjustified, although this may just be a presentational issue.

      a) Figure 2B: The authors interpret the results as "a plateau in the maximal encoding performance," when some readers might interpret this rather as a decline after 13 billion parameters. Can this be further supported by a significance test like that shown in Figure 4B?

      b) Figure S1A: It looks like the drop in PCA max correlation is larger for larger models, which may suggest to some readers that the same trend observed for ridge max correlation may not hold, contra the authors' claim that all results replicate. Why not include a similar figure as Figure 2B as part of Figure S1?

      (2) Discussion of what might be driving the main result about the influence of model size appears to be missing (cf. the authors aim to provide an explanation of what seems to drive the influence of the layer location in Paragraph 3 of the Discussion section). What explanations have been proposed in the previous functional magnetic resonance imaging studies? Do those explanations also hold in the context of this study?

      (3) The GloVe-based selection of language-sensitive electrodes (at least to me) isn't explained/motivated clearly enough (I think a more detailed explanation should be included in the Materials and Methods section). If the electrodes are selected based on GloVe embeddings, then isn't the main experiment just showing that representations from larger language models track more closely with GloVe embeddings? What justifies this methodology?

      (4) (Minor weakness) The main experiments are largely replications of previous functional magnetic resonance imaging studies, with the exception of the one lag-based analysis. Is there anything else that the electrocorticography data can reveal that functional magnetic resonance imaging data can't?

      Comments on revised version.

      I reread the manuscript, my previous review, and the authors' response to it. I thank the authors for clarifying any misunderstanding from my end (e.g. the different LLM tokenizers) and feel that the authors addressed my concerns very carefully.

    5. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors perform an analysis of the relationship between the size of an LMM and the predictive performance of an ECoG encoding model made using the representations from that LMM. They find a logarithmic relationship between model size and prediction performance, consistent with previous findings in fMRI. They additionally observe that as the model size increases, the location of the "peak" encoding performance typically moves further back into the model in terms of percent layer depth, an interesting result worthy of further analysis into these representations.

      Strengths:

      The evidence is quite convincing, consistent across model families, and complementary to other work in this field. This sort of analysis for ECoG is needed and supports the decade-long enduring trend of the "virtuous cycle" between neuroscience and AI research, where more powerful AI models have consistently yielded more effective predictions of responses in the brain. The lag analysis showing that optimal lags do not change with model size is a nice result using the higher temporal resolution of ECoG compared to other methods like fMRI.

      We thank the reviewer for their thoughtful assessment! We agree that the “virtuous cycle” between neuroscience and AI research has been, and will continue to be, a driving force in advancing our understanding of brain function through more powerful predictive models. We are especially pleased that the reviewer appreciated the lag analysis, as we view this as a valuable complement to the existing fMRI work.

      Weaknesses:

      I would have liked to have seen the data scaling trends explored a bit too, as this is somewhat analogous to the main scaling results. While better performance with more data might be unsurprising, showing good data scaling would be a strong and useful justification for additional data collection in the field, especially given the extremely limited amount of existing language ECoG data. I realize that the data here is somewhat limited (only 30 minutes per subject), but authors could still in principle train models on subsets of this data.

      We thank the reviewer for their valuable suggestion. For the revised manuscript, we performed a new analysis where we trained encoding models using subsets of the data (randomly sampling contiguous chunks of 50%, 25%, and 10% of all words in each of the training folds) and tested these models on all words in the test fold. As expected, we found that encoding performance increases as the training dataset size increases, suggesting that model performance scales with data quantity even within the constraints of our relatively small dataset. This result reinforces the importance of collecting dense ECoG data. We have added the following text to our Results section: “We also built encoding models using subsets of the data and found that encoding performance increases as the volume of training data increases (Fig. S6)” and included the results as a supplementary figure 6 in the revised manuscript.

      Separately, it would be nice to have better justification of some of these trends, in particular the peak layerwise encoding performance trend and the overall upside-down U-trend of encoding performance across layers more generally. There is clearly something very fundamental going on here, about the nature of abstraction patterns in LLMs and in the brain, and this result points to that. I don't see the lack of justification here as a critical issue, but the paper would certainly be better with some theoretical explanation for why this might be the case.

      We thank the reviewer for this insightful comment. The general inverted U-shaped trend of encoding performance across layers has been a frequently observed phenomenon in studies comparing LLM representations to brain activity (Goldstein, Ham, et al., 2025; Schrimpf et al., 2021). A potential explanation is the existence of a “two-phase abstraction process” within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the initial layers, models begin by processing relatively low-level input features. As layers get deeper, representations become increasingly abstract and richly contextualized in semantic features relevant for understanding language. These intermediate layers often show the highest correlation with brain activity in language areas, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. These observations suggest that it is primarily the abstractive, contextual features developed in the intermediate layers of LLMs that drive their alignment with brain activity. As models become more potent at prediction, their most predictive layers (for the LLM’s natural language task) and their most generalizable layers (for brain activity) can diverge.

      A key finding in our study is that the initial processing phase does not scale and take up more layers as models scale up in size and layers. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Consequently, the prediction phase may begin relatively earlier in these larger models, and the later layers could develop highly specialized representations that are increasingly divergent from the more general linguistic processing captured in brain activity. For example, these layers may specialize in capturing very specific patterns of language (thus lowering their perplexity) that do not actually occur often or at all in our naturalistic dataset. It is also possible that the later layers of larger models are overall underutilized and do not contribute as much to linguistic processing and next-word prediction (Csordás et al., 2025).

      We have added the following text to our Discussion section:

      “The inverted U-shaped trend of encoding performance commonly found in previous research is likely due to a "two-phase abstraction process" within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the early and intermediate layers of the model, a composition phase occurs, where low-level input features become increasingly abstract and contextualized. The intermediate layers of the model show the highest correlation with brain activity, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers of the model, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. Our results indicate that the initial composition phase does not take up more layers as models scale up in size. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Thus, as LLMs increase in size, the later layers of the model may contain representations that are increasingly divergent from the more general linguistic processing captured in brain activity. It is also possible that the later layers of larger models are overall underutilized and may not significantly contribute to benchmark performances during inference (Csordás et al., 2025; Fan et al., 2024; Gromov et al., 2024).”

      Lastly, I would have wanted to see a similar analysis here done for audio encoding models using Whisper or WavLM as this is the modality where you might see real differences between ECoG and other slower scanning approaches. Again, I do not see this omission as a fundamental issue, but it does seem like the sort of analysis for which the higher temporal resolution of ECoG might grant some deeper insight.

      We appreciate this suggestion. In a separate project, we focused on multimodal audio-to-speech-to-language large language models (LLMs), building encoding models using Whisper embeddings (from both the encoder and decoder stacks) to predict electrocorticographic (ECoG) signals during naturalistic conversations (Goldstein, Wang, et al., 2025). The higher temporal resolution of ECoG enables us to trace the temporal flow of information from the superior temporal gyrus (STG) and somatomotor areas (SM) to the inferior frontal gyrus (IFG) during speech comprehension. Conversely, during speech production, encoding in IFG peaked significantly earlier than in the STG and SM. We agree that scaling encoding models using multimodal approaches and our ECoG conversation datasets could yield valuable insights, and we look forward to exploring this in future work. However, we feel that the added complexity of multimodal encoding models falls beyond the scope of this paper.

      We have modified the following text to our Discussion section:

      “Since we exclusively employ textual LLMs, which lack inherent temporal information due to their discrete token-based nature, future studies utilizing multimodal LLMs integrating continuous audio or video streams, like Whisper or WavLM may better unravel the relationship between model size and temporal dynamic representations in LLMs (Goldstein, Wang, et al., 2025; Millet et al., 2023; Vaidya et al., 2022).”

      Reviewer #2 (Public review):

      Summary:

      This paper investigates whether large language models (LLMs) of increasing size more accurately align with brain activity during naturalistic language comprehension. The authors extracted word embeddings from LLMs for each word in a 30-minute story and regressed them against electrocorticography (ECoG) activity time-locked to each word as participants listened to the story. The findings reveal that larger LLMs more effectively predict ECoG activity, reflecting the scaling laws observed in other natural language processing tasks.

      Strengths:

      (1) The study compared model activity with ECoG recordings, which offer much better temporal resolution than other neuroimaging methods, allowing for the examination of model encoding performance across various lags relative to word onset.

      (2) The range of LLMs tested is comprehensive, spanning from 82 million to 70 billion parameters. This serves as a valuable reference for researchers selecting LLMs for brain encoding and decoding studies.

      (3) The regression methods used are well-established in prior research, and the results demonstrate a convincing scaling law for the brain encoding ability of LLMs. The consistency of these results after PCA dimensionality reduction further supports the claim.

      We thank the reviewer for their thoughtful and positive feedback.

      Weaknesses:

      (1) Some claims of the paper are less convincing. The authors suggested that "scaling could be a property that the human brain, similar to LLMs, can utilize to enhance performance", however, many other animals have brains with more neurons than the human brain, making it unlikely that simple scaling alone leads to better language performance.

      We thank the reviewer for this insightful comment. We agree that simply having more neurons does not automatically confer more complex or human-like cognitive or linguistic capabilities. This suggestion deserves a more nuanced treatment than we had included in the original manuscript.

      Research in comparative neuroscience has argued that human cognitive abilities emerge from scaling up the primate brain (Herculano-Houzel, 2012). However, the critical aspect is not merely the number of neurons, but how these neurons contribute to computational power within a specific evolutionary and cultural context. The uniqueness of human cognition has been argued to result from a global adaptation for increased information processing capacity (Cantlon & Piantadosi, 2024). Moreover, the language network in humans is likely grounded in the evolution of particular structural networks in the primate brain (Friederici & Becker, 2025). This suggests that the way brain regions are connected and the expansion of certain pathways are critical, not just the overall scale. Furthermore, the specialized structure must be tuned by its learning environment and training data. For example, both humans and LLMs learn from language data generated by other humans, which reflects world knowledge that has accumulated over many generations.

      We have modified the following text in the Introduction:

      “Research in comparative neuroscience has suggested that uniquely human cognitive abilities emerge from scaling up the primate brain (Herculano-Houzel, 2012).”

      We also added a caveat to the Discussion on this point:

      “As in the human brain, while scaling alone may yield emergent cognitive abilities (Cantlon & Piantadosi, 2024; Herculano-Houzel, 2012), specialized architectural features likely also play a critical role (Friederici & Becker, 2025).”

      Additionally, the authors claim that their results show 'larger models better predict the structure of natural language.' However, it remains unclear to what extent the embeddings of LLMs capture the "structure" of language better than the lexical semantics of language.

      We appreciate the reviewer's point about how well LLM embeddings capture the "structure" of language versus just lexical semantics. It's true that distinguishing these aspects is complex. From our perspective, a model's ability to predict/produce natural language entails that the model captures various levels of linguistic structure, including morphology, syntax, semantics, and contextual dependencies. We use "structure" inclusively in this sense. A model cannot achieve high predictive accuracy without representing, to some extent, all of these structural elements (Linzen & Baroni, 2021; Manning et al., 2020; Pavlick, 2022). There is a very active field of research into understanding exactly how these models represent these different structures of language (Ameisen et al., 2025; Chemla et al., 2024; Elhage et al., 2021, 2022; Hewitt & Manning, 2019). Our results confirm the core trend that larger models tend to better reproduce the various structures of language (i.e., yield lower perplexity; Fig. 2A).

      In previous work, we have shown that LLM embeddings better predict neural activity during natural language processing than lexical embeddings (e.g., GloVe) that do not contain other elements of linguistic structure (Goldstein et al., 2022; Kumar et al., 2024; Zada et al., 2024). In response to the following comment, we also compare LLMs to simpler models capturing specific speech and language features (see next comment). To clarify our intended use of the word “structure”, we’ve added a brief explanation in the Methods section:

      “In this study, we use the term “structure” to refer to a variety of linguistic patterns (e.g., morphology, syntax, semantics, context) that LLMs encode in order to better predict natural language.”

      (2) The study lacks control LLMs with randomly initialized weights and control regressors, such as word frequency and phonetic features of speech, making it unclear what the baseline is for the model-brain correlation.

      We’ve added several supplementary analyses to the revised manuscript to address these concerns. To establish a baseline, we extracted embeddings from each layer of the SMALL model with randomly initialized weights and constructed encoding models. The encoding performance is significantly higher for pretrained SMALL than for untrained SMALL for every layer (Fig. S4). For the untrained model, the performance is the highest for the 0th layer and decreases in subsequent layers. This is because at the 0th layer, every instance of the same word receives an identical, albeit random, embedding (See Supplementary Figure 4).

      We also compared the encoding performance of LLMs with more classical speech/language features and static GloVe embeddings (Goldstein, Wang, et al., 2025; Kumar et al., 2024). First, we extracted features capturing lower-level speech features. Using the stimulus transcript as input, we created one-hot vectors for phonetic and articulatory features. Phoneme classes (39 total classes) were obtained from the Carnegie Mellon Pronouncing Dictionary (The CMU Pronouncing Dictionary, n.d.). We further classified the phonemes based on their place of articulation (9 classes), manner of articulation (9 classes), and voiced or voiceless status (3 classes), according to the general American English consonants of the International Phonetic Alphabet. Given that each word consists of multiple phonemes, we averaged the one-hot vectors for all phonetic and articulatory features for each word.

      Second, we extracted linguistic features using spaCy (Honnibal et al., 2020), including part of speech (17 classes), tag (50 classes), function or content word (3 classes), dependency (45 classes), whether the word is an alpha character (binary), and whether the word is a stop word (binary). We also extracted prefix (30 classes) and suffix (44 classes) information using the Cambridge Dictionary. We constructed one-hot vectors for each multi-class feature and one-dimensional vectors for each binary feature.

      Third, for each word, we obtained word frequency from the Google Web Trillion Word Corpus (Brants & Franz, 2006) and from our own dataset.

      Fourth, we generated static word embeddings of dimension 50 using GloVe (Pennington et al., 2014).

      We then built encoding models in the same way as the contextual embeddings for each of the three categories of speech features, all speech features concatenated, and the GloVe embeddings. To control for the different dimensions of the embeddings, we also standardized all embeddings to the same size (50 dimensions) using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression. For both ridge and OLS encoding, our contextual embeddings from LLMs showed significantly better performance than the classic speech features and GloVe embeddings.

      We have added the following text to our manuscript and updated our Figures S4, S5, Table S1, and the methods section:

      “To establish a general baseline for encoding performance, we built encoding models using embeddings from the SMALL model with randomly initialized weights. The trained SMALL model exhibits significantly higher encoding performance across all layers compared to the untrained SMALL model (Fig. S4). We also assessed the encoding performance of contextual embeddings from LLMs against classic speech features and static GloVe embeddings (Table S1). The SMALL and XL embeddings achieved markedly higher encoding correlations than the speech features and GloVe embeddings (Fig. S5).”

      (3) The finding that peak encoding performance tends to occur in relatively earlier layers in larger models is somewhat surprising and requires further explanation. Since more layers mean more parameters, if the later layers diverge from language processing in the brain, it raises the question of what aspects of the larger models make them more brain-like.

      We thank the reviewer for this insightful comment; this point was also highlighted by Reviewer 1. We agree that this result is somewhat surprising, and we aim to provide a more detailed explanation in the revised manuscript. The general inverted U-shaped trend of encoding performance across layers has been a frequently observed phenomenon in studies comparing LLM representations to brain activity (Goldstein, Ham, et al., 2025; Schrimpf et al., 2021). A potential explanation is the existence of a “two-phase abstraction process” within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the initial layers, models begin by processing relatively low-level input features. As layers get deeper, representations become increasingly abstract and richly contextualized in semantic features relevant for understanding language. These intermediate layers often show the highest correlation with brain activity in language areas, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. These observations suggest that it is primarily the abstractive, contextual features developed in the intermediate layers of LLMs that drive their alignment with brain activity. As models become more potent at prediction, their most predictive layers (for the LLM’s natural language task) and their most generalizable layers (for brain activity) can diverge.

      A key finding in our study is that the initial processing phase does not scale and take up more layers as models scale up in size and layers. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models.

      Consequently, the prediction phase may begin relatively earlier in these larger models, and the later layers could develop highly specialized representations that are increasingly divergent from the more general linguistic processing captured in brain activity. For example, these layers may specialize in capturing very specific patterns of language (thus lowering their perplexity) that do not actually occur often or at all in our naturalistic dataset. It is also possible that the later layers of larger models are overall underutilized and do not contribute as much to linguistic processing and next-word prediction (Csordás et al., 2025).

      We have added the following text to our Discussion section:

      “The inverted U-shaped trend of encoding performance commonly found in previous research is likely due to a "two-phase abstraction process" within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the early and intermediate layers of the model, a composition phase occurs, where low-level input features become increasingly abstract and contextualized. The intermediate layers of the model show the highest correlation with brain activity, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers of the model, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. Our results indicate that the initial composition phase does not take up more layers as models scale up in size. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Thus, as LLMs increase in size, the later layers of the model may contain representations that are increasingly divergent from the more general linguistic processing captured in brain activity. It is also possible that the later layers of larger models are overall underutilized and may not significantly contribute to benchmark performances during inference (Csordás et al., 2025; Fan et al., 2024; Gromov et al., 2024).”

      Reviewer #3 (Public review):

      This manuscript studies the connection between neural activity collected through electrocorticography and hidden vector representations from autoregressive language models, with the specific aim of studying the influence of language model size on this connection. Neural activity was measured from subjects who listened to a segment from a podcast, and the representations from language models were calculated using the written transcription as the input text. The ability of vector representations to predict neural activity was evaluated using 10-fold cross-validation with ridge regression models.

      The main results are that (as well summarized in section headings):

      (1) Larger models predict neural activity better.

      (2) The ability of language model representations to predict neural activity differs across electrodes and brain regions.

      (3) The layer that best predicts neural activity differs according to model size, with the "SMALL" model showing a correspondence between layer number and the language processing hierarchy.

      (4) There seems to be a similar relationship between the time lag and the ability of language model representations to predict neural activity across models.

      Strengths:

      (1) The experimental and modeling protocols generally seem solid, which yielded results that answer the authors' primary research question.

      (2) Electrocorticography data is especially hard to collect, so these results make a nice addition to recent functional magnetic resonance imaging studies.

      We thank the reviewer for their thoughtful and positive feedback.

      Weaknesses:

      (1) The interpretation of some results seems unjustified, although this may just be a presentational issue.

      (a) Figure 2B: The authors interpret the results as "a plateau in the maximal encoding performance," when some readers might interpret this rather as a decline after 13 billion parameters. Can this be further supported by a significance test like that shown in Figure 4B?

      We agree that this could be a subjective interpretation, so we conducted an additional analysis. We performed paired two-sided t-tests between best layer encoding performances averaged across electrodes (df = 159 electrodes), comparing all models with larger models. We found that after 13 billion parameters, only the encoding performance for OPT-66B, the largest model in the OPT family, is significantly worse than the encoding performance of some other smaller models, supporting the claim that the maximal encoding performance declines after 13 billion parameters. However, we did not find conclusive statistical evidence of a decline in encoding performance for other model families.

      We have added the statistical results as Supplementary Figure 1.

      We have also modified the following text in the manuscript:

      “We also observed a plateau in the maximal encoding performance, occurring around 7 billion parameters (Fig. 2B), with a decline in performance for the OPT-66B model (Fig. S1).”

      (b) Figure S1A: It looks like the drop in PCA max correlation is larger for larger models, which may suggest to some readers that the same trend observed for ridge max correlation may not hold, contra the authors' claim that all results replicate. Why not include a similar figure as Figure 2B as part of Figure S1?

      PCA is an unsupervised dimensionality reduction technique and may discard model features with small eigenvalues that nonetheless contribute to encoding performance. Ridge regression, a supervised method, can capitalize on these features. We suspect that this is why there appears to be a larger drop in model performance for larger models with PCA than with ridge regression. We replicated the logarithmic relationship between model size and encoding performance using PCA and ordinary least-squares (OLS) regression encoding models. We have updated Supplementary Figure 2.

      (2) Discussion of what might be driving the main result about the influence of model size appears to be missing (cf. the authors aim to provide an explanation of what seems to drive the influence of the layer location in Paragraph 3 of the Discussion section). What explanations have been proposed in the previous functional magnetic resonance imaging studies? Do those explanations also hold in the context of this study?

      We suspect that the increased expressivity of larger models - that is, their improved sensitivity to nuanced structure in natural language - yields improved alignment to brain activity (given large enough samples of brain activity) (Antonello et al., 2023). This effect persists even when dimensionality is tightly controlled in our PCA-based analysis, indicating that the improved alignment with the brain is not a modeling artifact of dimensionality alone, but results from the structural representations learned by these larger models.

      We have added the following text to our Discussion section:

      “We suspect that the improved alignment with brain activity in larger models is driven by their increased expressivity and sensitivity to nuanced linguistic structure present in large-scale naturalistic datasets (Antonello et al., 2023).”

      (3) The GloVe-based selection of language-sensitive electrodes (at least to me) isn't explained/motivated clearly enough (I think a more detailed explanation should be included in the Materials and Methods section). If the electrodes are selected based on GloVe embeddings, then isn't the main experiment just showing that representations from larger language models track more closely with GloVe embeddings? What justifies this methodology?

      We selected electrodes based on previously established methods (Goldstein et al., 2022). Our use of GloVe embeddings for electrode selection does not imply that larger language model representations are simply more closely aligned with GloVe embeddings. On the contrary, contextual embeddings from LLMs, which incorporate the word’s previous context, consistently outperform static embeddings like GloVe or word2vec (Fig. S3). Selecting electrodes using LLM embeddings would likely result in a slightly different, potentially larger set of electrodes (Goldstein et al., 2022), but would be more circular (Kriegeskorte et al., 2009). The GloVe-based electrode selection represents a more conservative approach by identifying words encoding linguistic content without biasing the selection directly toward any LLMs.

      We have added the following text to our Method section:

      “We used GloVe embeddings for electrode selection to avoid biasing our main results toward a particular LLM.”

      (4) (Minor weakness) The main experiments are largely replications of previous functional magnetic resonance imaging studies, with the exception of the one lag-based analysis. Is there anything else that the electrocorticography data can reveal that functional magnetic resonance imaging data can't?

      We thank the reviewer for this thoughtful question. While we agree that a key contribution of our work corroborates previous fMRI findings, we would argue that using ECoG is not merely a replication but a crucial validation and extension of that work. It is important to validate these effects across distinct measurement modalities. In our work, we further observed a novel trend where the peak encoding performance tends to occur in relatively earlier layers for larger models. This is supported by recent studies suggesting that later layers of large LLMs may not significantly contribute to benchmark performance (Csordás et al., 2025). While scaling has been an effective method to improve LLM performance, including in encoding models, future research should explore the potential underutilization of the later layers as models scale.

      Furthermore, ECoG data offers temporal resolution on the order of milliseconds, far superior to fMRI’s. Although we did not observe a relationship between model size and temporal lags in this study, future work should investigate the temporal dynamics of encoding that are accessible with ECoG (Goldstein, Ham, et al., 2025; Goldstein, Wang, et al., 2025).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Thank you to the authors for the fun and personally useful read.

      I see in Supplementary Figure 1 the authors show a comparison of the performance between OLS vs. Ridge regression. Is the OLS model the only one that is working over PC features, or are both models using PC features? The current text is a bit unclear. My current understanding is that the comparison is between (OLS + PCA) and (Ridge with no PCA), but I am not sure.

      The OLS model is the only one that works over PC features, following previous methods (Goldstein et al., 2022).

      We have added the following text to our Results and Methods section for clarity:

      “To control for the different embedding dimensionality across models, we standardized all embeddings to the same size using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression, replicating the logarithmic relationship but with significantly lower encoding performance overall (Fig. S2). The PC features are used by the OLS models only.”

      Clarification in the text would be appropriate. If this is the correct understanding, the authors should note in the main text that the ridge approach is more effective than the PCA approach, which is still the dominant approach to building linear encoding models in the field for some unjustifiable reason.

      We thank the reviewer for pointing out the confusion. We have updated Supplementary Figure 2.

      How were the alpha values for ridge regression determined? Do you use the same ridge parameter for all electrodes or fit a different parameter for each electrode? This is not mentioned anywhere.

      The alpha values are determined by cross-validation using the “RidgeCV” method from the “himalaya” package (Dupré la Tour et al., 2022). Specifically, we perform a grid search over cross-validation folds in the training data to find the best-performing alpha. The alpha parameter is specific to each ridge regression model, meaning each fold, lag, and electrode has a different alpha parameter.

      We have added the following text to our manuscript:

      “For each ridge regression model (for each fold, lag, and electrode), the alpha parameter is determined by cross-validation using the “RidgeCV” method from the “himalaya” package (Dupré la Tour et al., 2022).”

      It's not entirely clear to me how the authors handle tokens that do not terminate in words (such as the "there" + "'s" example in the text). My current reading of the text is that authors essentially ignore these half-word embeddings, doing one forward pass per word, rather than per token, but the current text is somewhat ambiguous.

      If a word is tokenized into several tokens, like “there” and “‘s”, we average the token embeddings to get a word embedding.

      We have added the following text to our Method section:

      “To facilitate a fair comparison of the encoding effect across different models, we aligned all tokens in the story across all models. We averaged the token embeddings if a word is split into multiple tokens, resulting in one embedding per word for each model.”

      The authors describe the scaling relationship they find as a "log-linear" relationship. I believe this is a misnomer derived from the original paper describing this relationship in fMRI as log-linear (Antonello et al.) The correct term is simply "logarithmic", and for what it's worth, the authors of the original fMRI work have made this correction as well.

      Thank you! We have made this correction.

      Is the data publicly available? If not, there should be some basic justification as to why (consent reasons, etc.).

      We have recently made the data publicly available (Zada et al., 2025). We have also provided tutorials for preprocessing the data and training encoding models: https://hassonlab.github.io/podcast-ecog-tutorials. For this specific project, the analysis code is available at https://github.com/hassonlab/247-pickling/tree/scaling-paper-0 and https://github.com/hassonlab/247-encoding/tree/scaling-paper-1.

      The authors assert that ECoG has "superior spatiotemporal resolution". While this is unquestionably true for temporal resolution, the story is a bit more complicated for spatial resolution, where ECoG has far less cortical coverage than fMRI. Perhaps this sentence should be revised.

      Thank you for pointing out the typo! We have changed it to “superior temporal resolution”.

      Minor Points:

      The bolded title of Figure 3 probably shouldn't be bolded, as this is just actually the title of Figure 3A.

      Fixed.

      Figure 4d is has a typo: "Best Encoidng Layer".

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      The authors could consider adding control regressors such as word rate, word frequency, phonetic features, and syntactic features like node counts, as well as control LLMs of comparable size to serve as baselines. The authors could also include correlation analyses of the embeddings from different layers of the same LLM to further illustrate how distinct the layers are within the models.

      We have added untrained LLM embeddings as a baseline and included a comparison of encoding models between LLM contextual embeddings and classical speech features. We have also performed some preliminary correlation analyses of embeddings. In some models, we found evidence of the “two-phase abstraction process” (Cheng & Antonello, 2024). However, the result is inconclusive across different LLM families. Since each LLM layer accesses and modifies the residual stream (Elhage et al., 2021), the embeddings across layers are inherently correlated. Future work could instead explore the isolated transformations within each layer to illustrate the distinct information across layers (Kumar et al., 2024).

      The analysis codes and data should be made available.

      We have recently made the data publicly available (Zada et al., 2025). We have also provided tutorials for preprocessing the data and training encoding models: https://hassonlab.github.io/podcast-ecog-tutorials. For this specific project, the analysis code is available at https://github.com/hassonlab/247-pickling/tree/scaling-paper-0 and https://github.com/hassonlab/247-encoding/tree/scaling-paper-1.

      Reviewer #3 (Recommendations for the authors):

      Most of my concrete recommendations are in the public review. Below are some additional minor ones:

      (1) Introduction: "Remarkably, these models learn from much the same shared space as humans: from real-world language generated by humans."

      I think this is an extremely strong claim due to e.g. the different nature of child-directed speech vs. written text corpora, the lack of multimodality and grounding in language models, etc. I might suggest re-wording this sentence or removing it entirely.

      We thank the reviewer for their suggestion! We have removed the sentence from the manuscript.

      (2) Introduction: "EleutherAI, n.d." reference for GPT-Neo

      GPT-NeoX-20B has an associated paper, which the authors might cite instead: https://aclanthology.org/2022.bigscience-1.9

      Thank you! We have added the reference for GPT-NeoX-20B (Black et al., 2022).

      (3) Figure 4D: Encoidng -> Encoding

      Fixed.

      (4) Materials and Methods, Contextual embeddings: "except for GPT-Neox-20b, which assigns additional tokens to whitespace characters."

      What do the authors mean by "additional tokens to whitespace characters?" The tokenizer for GPT-NeoX-20B works in much the same way as that of GPT-Neo, just with a different vocabulary set.

      >>> t1 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-125M")

      >>> t2 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")

      >>> t1.convert_ids_to_tokens(t1("The quick brown fox jumps over the lazy dog.").input_ids) ['The', 'Ġquick', 'Ġbrown', 'Ġfox', 'Ġjumps', 'Ġover', 'Ġthe', 'Ġlazy', 'Ġdog', '.']

      >>> t2.convert_ids_to_tokens(t2("The quick brown fox jumps over the lazy dog.").input_ids)

      ['The', 'Ġquick', 'Ġbrown', 'Ġfox', 'Ġjumps', 'Ġover', 'Ġthe', 'Ġlazy', 'Ġdog', '.']

      If the authors are referring to Ġ as the "additional token to whitespace characters," then these are in all other tokenizers as well (not only that for GPT-Neo, but also those for GPT-2 and OPT).

      We agree that “additional tokens to whitespace characters” is an oversimplification. The GPT-Neo model family, which includes the 125M, 1.3B, and 2.7B models, utilizes the same Byte Pair Encoding (BPE) tokenizer as GPT-2. This common tokenizer has a vocabulary size of 50,257 tokens, providing compatibility and seamless integration across the models.

      The GPT-NeoX-20B model introduces a modified tokenizer to address limitations observed in the GPT-2 tokenizer (Black et al., 2022). As detailed in Section 3.2, this new tokenizer incorporates a few key improvements:

      (1) New BPE tokenizer: A more general-purpose BPE tokenizer was trained using the Pile dataset.

      (2) Space Delimitation: Unlike the GPT-2 tokenizer, which treats tokenization at the start of a string as a non-space-delimited token, the GPT-NeoX-20B tokenizer applies consistent space delimitation regardless. This change resolves inconsistencies related to the presence of prefix spaces in the tokenization input.

      (3) Whitespace Handling: The tokenizer includes tokens for repeated space characters (up to 24 consecutive spaces), enhancing efficiency in tokenizing text with substantial whitespace, such as program source code or LaTeX documents.

      These modifications result in the GPT-NeoX-20B tokenizer representing the Pile validation set with approximately 10% fewer tokens than the GPT-2 tokenizer. This efficiency gain is particularly beneficial for processing texts with extensive whitespace.

      In our analysis, we extracted embeddings by setting `add_prefix_space = True` to all tokenizers, so space delimitation does not result in tokenizer differences. We highlight here examples of the other two tokenizer differences using the Huggingface `AutoTokenizer`:

      >>> t1 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-125M")

      >>> t2 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")

      >>> t1.convert_ids_to_tokens(t1("The Downing Street.").input_ids) ['The', 'ĠDowning', 'ĠStreet']

      >>> t2.convert_ids_to_tokens(t2("The Downing Street.").input_ids)

      ['The', 'ĠDown', 'ing', 'ĠStreet']

      >>> t1.convert_ids_to_tokens(t1("Hello !").input_ids)

      ['Hello', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ!']

      >>> t2.convert_ids_to_tokens(t2("Hello !").input_ids)

      ['Hello', ' ', '!']

      More examples showing the differences between the GPT-2 tokenizer and the GPT-NeoX-20B tokenizer can be found in Appendix F: Tokenizer Analysis (Black et al., 2022).

      We have added the following text to our manuscript for simplicity:

      “All models within the same model family adhere to the same tokenizer convention, except for GPT-Neox-20B, which utilizes a different tokenizer (Black et al., 2022).”

      References

      Ameisen, E., Lindsey, J., Pearce, A., Gurnee, W., Turner, N. L., Chen, B., Citro, C., Abrahams, D.,  Carter, S., Hosmer, B., Marcus, J., Sklar, M., Templeton, A., Bricken, T., McDougall, C.,  Cunningham, H., Henighan, T., Jermyn, A., Jones, A., … Batson, J. (2025). Circuit Tracing:  Revealing Computational Graphs in Language Models. Transformer Circuits Thread. https://transformer-circuits.pub/2025/attribution-graphs/methods.html

      Antonello, R., & Huth, A. (2024). Predictive coding or just feature discovery? An alternative account of why language models fit brain data. Neurobiology of Language (Cambridge, Mass.), 5(1), 64–79.

      Antonello, R., Vaidya, A., & Huth, A. G. (2023). Scaling laws for language encoding models in fMRI. NeurIPS 2023. https://doi.org/10.48550/ARXIV.2305.11863

      Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell,  K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., & Weinbach, S. (2022). GPT-NeoX-20B: An Open-Source Autoregressive Language Model.  Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models. Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models, virtual+Dublin. https://doi.org/10.18653/v1/2022.bigscience-1.9

      Brants, T., & Franz, A. (2006). Web 1T 5-gram Version 1 [Dataset]. Linguistic Data Consortium. https://doi.org/10.35111/CQPA-A498

      Cantlon, J. F., & Piantadosi, S. T. (2024). Uniquely human intelligence arose from expanded information capacity. Nature Reviews Psychology, 3(4), 275–293.

      Chemla, E., D’Ascoli, S., Diego-Simón, P., King, J.-R., & Lakretz, Y. (2024). A Polar coordinate system represents syntax in large language models. Advances in Neural Information Processing Systems 37, 105375–105396.

      Cheng, E., & Antonello, R. J. (2024). Evidence from fMRI supports a two-phase abstraction process in language models. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2409.05771

      Csordás, R., Manning, C. D., & Potts, C. (2025). Do language models use their depth efficiently?  In arXiv [cs.LG]. https://doi.org/10.48550/ARXIV.2505.13898

      Dupré la Tour, T., Eickenberg, M., Nunez-Elizalde, A. O., & Gallant, J. L. (2022). Feature-space selection with banded ridge regression. In bioRxiv. https://doi.org/10.1101/2022.05.05.490831

      Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., & Olah, C. (2022). Toy Models of Superposition. Transformer Circuits Thread.

      Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A.,  Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A.,  Kernion, J., Lovitt, L., Ndousse, K., … Olah, C. (2021). A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread.

      Fan, S., Jiang, X., Li, X., Meng, X., Han, P., Shang, S., Sun, A., Wang, Y., & Wang, Z. (2024). Not all Layers of LLMs are Necessary during Inference. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2403.02181

      Friederici, A. D., & Becker, Y. (2025). The core language network separated from other networks during primate evolution. Nature Reviews. Neuroscience, 26(2), 131–132.

      Goldstein, A., Ham, E., Schain, M., Nastase, S. A., Aubrey, B., Zada, Z., Grinstein-Dabush, A.,  Gazula, H., Feder, A., Doyle, W., Devore, S., Dugan, P., Friedman, D., Brenner, M., Hassidim, A., Matias, Y., Devinsky, O., Siegelman, N., Flinker, A., … Hasson, U. (2025). Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models. Nature Communications, 16(1), 10529.

      Goldstein, A., Wang, H., Niekerken, L., Schain, M., Zada, Z., Aubrey, B., Sheffer, T., Nastase, S. A., Gazula, H., Singh, A., Rao, A., Choe, G., Kim, C., Doyle, W., Friedman, D., Devore, S., Dugan, P., Hassidim, A., Brenner, M., … Hasson, U. (2025). A unified acoustic-to-speech-to-language embedding space captures the neural basis of natural language processing in everyday conversations. Nature Human Behaviour. https://doi.org/10.1038/s41562-025-02105-9

      Goldstein, A., Zada, Z., Buchnik, E., Schain, M., Price, A., Aubrey, B., Nastase, S. A., Feder, A.,  Emanuel, D., Cohen, A., Jansen, A., Gazula, H., Choe, G., Rao, A., Kim, C., Casto, C., Fanda, L., Doyle, W., Friedman, D., … Hasson, U. (2022). Shared computational principles for language processing in humans and deep language models. Nature Neuroscience, 25(3), 369–380.

      Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P., & Roberts, D. A. (2024). The Unreasonable Ineffectiveness of the Deeper Layers. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2403.17887

      Herculano-Houzel, S. (2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences of the United States of America, 109 Suppl 1(supplement_1), 10661–10668.

      Hewitt, J., & Manning, C. D. (2019). A Structural Probe for Finding Syntax in Word Representations. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North (pp. 4129–4138). Association for Computational Linguistics.

      Honnibal, M., Montani, I., Van Landeghem, S., & Boyd, A. (2020). spaCy: Industrial-strength Natural Language Processing in Python.

      Kriegeskorte, N., Simmons, W. K., Bellgowan, P. S. F., & Baker, C. I. (2009). Circular analysis in systems neuroscience: the dangers of double dipping. Nature Neuroscience, 12(5),  535–540.

      Kumar, S., Sumers, T. R., Yamakoshi, T., Goldstein, A., Hasson, U., Norman, K. A., Griffiths, T. L., Hawkins, R. D., & Nastase, S. A. (2024). Shared functional specialization in transformer-based language models and the human brain. Nature Communications, 15(1), 5523.

      Linzen, T., & Baroni, M. (2021). Syntactic Structure from Deep Learning. Annual Review of Linguistics, 7(1), 195–212.

      Manning, C. D., Clark, K., Hewitt, J., Khandelwal, U., & Levy, O. (2020). Emergent linguistic structure in artificial neural networks trained by self-supervision. Proceedings of the National Academy of Sciences of the United States of America, 117(48), 30046–30054.

      Millet, J., Caucheteux, C., Orhan, P., Boubenec, Y., Gramfort, A., Dunbar, E., Pallier, C., & King, J.-R. (2023). Toward a realistic model of speech processing in the brain with self-supervised learning. NeurIPS 2022. https://doi.org/10.48550/ARXIV.2206.01685

      Pavlick, E. (2022). Semantic structure in deep learning. Annual Review of Linguistics, 8(1),  447–471.

      Pennington, J., Socher, R., & Manning, C. (2014). Glove: Global vectors for word representation.  Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar. https://doi.org/10.3115/v1/d14-1162

      Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences of the United States of America, 118(45), e2105646118.

      The CMU Pronouncing Dictionary. (n.d.). Retrieved May 27, 2025, from http://www.speech.cs.cmu.edu/cgi-bin/cmudict

      Vaidya, A. R., Jain, S., & Huth, A. G. (2022). Self-supervised models of audio effectively explain human cortical responses to speech. ICML 2022. https://doi.org/10.48550/ARXIV.2205.14252

      Zada, Z., Goldstein, A., Michelmann, S., Simony, E., Price, A., Hasenfratz, L., Barham, E., Zadbood,  A., Doyle, W., Friedman, D., Dugan, P., Melloni, L., Devore, S., Flinker, A., Devinsky, O., Nastase, S. A., & Hasson, U. (2024). A shared model-based linguistic space for transmitting our thoughts from brain to brain in natural conversations. Neuron, S0896627324004604. Zada, Z., Nastase, S. A., Aubrey, B., Jalon, I., Michelmann, S., Wang, H., Hasenfratz, L., Doyle, W.,  Friedman, D., Dugan, P., Melloni, L., Devore, S., Flinker, A., Devinsky, O., Goldstein, A., & Hasson, U. (2025). The “Podcast” ECoG dataset for modeling neural activity during natural language comprehension. Scientific Data, 12(1), 1135.

    1. eLife Assessment

      This study demonstrates that toll-like receptor 4 in non-myeloid cells is a central regulator of leptomeningeal inflammation in the context of neonatal E. coli meningitis. The data, derived from conditional gene deletions in mice, and from cultured endothelial cells, are convincing and support the authors' conclusions. This work is important because it advances our understanding of the host cellular and molecular factors underlying meningitis pathogenesis, especially in the context of the vascular barrier breakdown.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have adequately addressed how the percentage of fibroblasts was derived.]

      Summary:

      In this manuscript, Seegren and colleagues demonstrate that in a mouse model of neonatal E. coli meningitis, loss of toll-like receptor 4 (TLR4) in VE-cadherin+ endothelial cells and a subset of meningeal fibroblasts leads to a marked decrease in transcriptional dysregulation across multiple leptomeningeal cell types, a decrease in vascular permeability, and a decrease in macrophage abundance. In contrast, loss of macrophage TLR4 had less pronounced effects. Using cultured wildtype and TLR4-knockout endothelial cells, the authors further demonstrate that TLR4 signaling leads to reversible internalization of the tight junction protein claudin-5, establishing a potential mechanism of increased vascular permeability. Authors also show that claudin-5 internalization is independent of NF-κB. Finally, the authors use RNA-sequencing of wildtype and TLR4-knockout endothelial cells to define the TLR4-dependent cell-autonomous transcriptional response to E. coli.

      Comments on previous version:

      The authors have considerably improved and strengthened the work through the addition of new experimental data, new data analyses, and modifications to their interpretation. Notably, the authors used additional Cre-reporter mice to clarify that Cdh5-CreER is active in endothelial cells and some meningeal fibroblasts, and thus revised nomenclature and interpretation to acknowledge that the Tlr4fl/-;Cdh5-CreER cKO (Tlr4-VEKO) is not exclusively endothelial. The authors also demonstrated that Tlr4-VEKO does not affect peripheral E.coli burden, but acknowledge that changes to periphery-derived signals (e.g., cytokines) may contribute to observed leptomeningeal phenotypes.

      The authors added PCA plots to show similarity in gene expression shifts across biological replicates (mice). This provides support for the claim that Tlr4-VEKO attenuates infection-associated transcriptional changes. With respect to differential expression analysis, I agree with authors that characteristics of individual cells (e.g. heterogeneity) are of interest. I remain concerned, however, that the formal differential analysis strategy appears to consider cells as independent experimental units, which they are not because a single cell cannot be randomly assigned to an experimental group (control or cKO, uninfected or infected). The mouse is the correct experimental unit for a comparison across these groups because it can be randomized. I appreciate that many of the gene expression changes appear consistent across mice (e.g. Figure 1 - Figure supplement 7) and that there are clear infection- and genotype-associated phenotypes in other assays. I would simply caution that the authors' analysis strategy likely leads to a larger number of type I errors (false positives) than is generally accepted; a mixed (hierarchical) model or pseudo-bulk approach would be more appropriate for future studies.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use a postnatal mouse model of E. coli bacterial meningitis and a mouse brain endothelioma cell line combined with cell type specific gene deletion to study the function of endothelial TLR4, a cell surface receptor that recognizes gram positive bacterial wall components, in the local leptomeningeal (LPM) response with a focus on endothelial barrier breakdown mediated by TLR4. Single cell transcriptional profiling and imaging studies using wholemount preps of the LPM support that LPM endothelial, CD206+ local macrophage and LPM fibroblast and arachnoid barrier cell inflammatory response and is abrogated in endothelial specific KO of TLR4, pointing to a role for endothelial TLR4 in local LPM response. Culture studies using Bend3.1 cells (a mouse brain endothelioma cell line) support a direct role for TLR4 in the bacteria-mediated inflammatory response and in internalization of Cldn5 via the endosomal-lysosomal pathway, resulting in loss of barrier integrity

      Strengths:

      The local LPM cell response in meningitis and the role of specific LPM cells in inflammation and CNS barrier breakdown has not been extensively studied, despite ample evidence for primary immune response in the meninges in human patients and in animal models. The authors employ a robust, multi-model approach using both in vivo and in vitro models with cell-type specific knockout to study the function of TLR4 in brain endothelial cell response. The authors nicely combine functional barrier assays with IF for junctional localization in their experimental design and they delve into potential mechanisms of Cldn5 internalization using markers of endosomal-lysomal pathway localization. The authors also describe a new type of barrier assay using a streptavidin-coated plates upon which barrier forming cell cultures can be plated, this could be a very useful alternative or complement to other size-selective barrier assays and presumably could work for other barrier forming cell types, like epithelial cells.

      Comments on previous version.

      In their revision, the authors addressed prior noted weaknesses with new data and analysis. They now show that TLR4-VE-cad cKO mice have a largely similar disease progression as control mice, including increased bacterial burden in the LPM and brain. This underscores that that the reduced vascular leakage and blunted inflammatory response is due to loss of TLR4 response to bacteria on VE-cad recombined cells and not because the mice are protected from meningitis. The authors also performed additional experiments to show that Cldn5 internalization via the endosomal-lysosomal pathway is independent of NFKB signaling. The authors also added in important discussion points about how their results fit into the broader literature on TLR4 in BBB endothelial cell junctional protein localization and prior work on meningitis in global TLR4.

    4. Reviewer #3 (Public review):

      Summary:

      This study investigates the molecular underpinnings of immune responses in the leptomeninges in neonatal bacterial meningitis. Bacterial meningitis is a major disease burden, particularly for neonates, and it has previously been noted that the meningeal immune environment in infants is permissive to opportunistic infection (Kim et al., Sci Immunol, 2023). There is less known about the contribution of the stromal compartment to meningeal immune responses. Seegren et al. interrogate the role of leptomeningeal endothelium in host defense in E. coli infected neonatal mice using mouse genetic tools to delete the LPS receptor Tlr4 from either endothelial cells/stromal cells (using Cdh5-CreER) or myeloid cells (using LysM-Cre). The authors use snRNAseq, cleared cortical mounts, and in vitro work to define the impact of E. coli infection on leptomeningeal endothelial cells. This study uses a range of innovative techniques to probe the role of the stromal compartment in meningitis. With additional experiments to confirm the specificity of their Cre models, this strengthens the interpretation of the study significantly. The only major weakness is the inability to confirm TLR4 knockout in myeloid cells.

      Strengths:

      This study makes excellent use of cleared cortical mounts to examine the biology of the leptomeninges, in particular, changes to the endothelium, with unprecedented detail. In combination with high-quality sequencing data provide new insights into the impact of meningitis on the leptomeninges. The data presented by the authors is of very high quality.

      The authors have also done substantial work to address my two major comments regarding 1) the specificity of their Cre systems and 2) peripheral impacts of the interventions.

      (1) The authors identified and acknowledged some impacts in the leptomeningeal stroma (the relatively high level of recombination in ECs vs FBs presumably reflects a single low dose being given, where other groups have done more aggressive tamoxifen regimens that drive recombination in FBs as well). Given the incomplete recombination in the leptomeningeal FBs, I agree with their conclusion that it is probably endothelial driven. Acknowledging the contributions of other myeloid cells with the L. The Cre-NLS experiments with nuclear markers provided excellent data and had beautiful staining.

      (2) The authors did not observe differences in bacterial burden in peripheral organs in either CKO model, suggesting that CNS impacts are not downstream of peripheral bacterial control.

      Weaknesses:

      (1) While the inducible Cre lines used by the authors target both peripheral and CNS tissues, this potential confound is mitigated by the lack of impact on peripheral disease burden.

      (2) The authors were not able to confirm TLR4 knockout in myeloid cells, and this caveat is acknowledged. The lack of response in TLR4 VEKO mice strongly suggests successful conditional knockout.

      (3) The cell line model (bEnd.3) is a relatively low fidelity model of BBB endothelial cells. The authors acknowledge this, and it is likely that endothelial cell responses to LPS are highly conserved.

      (4) It is perhaps not surprising that Tlr4 is required for meningitis responses with E. coli. However, it is unclear if these findings can be generalised to other, more common, meningitis infections (streptococcal/pneumococcal).

    5. Author response:

      The following is the authors’ response to the previous reviews

      We are grateful to you and the reviewers for your careful and positive assessment.  In response to the comment below from reviewer #2, we have expanded Figure 1- figure supplement 1 to show the results (images plus quantification) of ERG and PU.1 immunostaining together with DAPI staining that underly the calculation of the percent on non-endothelial cells that are recombined by Cdh5-CreER.

      However, I found their explanation as to how they arrived at only 18% of the fibroblasts being recombined confusing. Mostly because it looks like there are many GFP+/ERG- cells in panels A & D, these would be the recombined fibroblasts and AB cells. Considering that Cdh5 gene is pretty broadly expressed across LPM, the prediction is the recombination rate would be higher (recognizing mice strains vary and this is inducible Cre).

      Ideally, they would perform recombination analysis with a TF expressed by all fibroblasts in combination with Erg (like Foxc1). However, a potentially simpler approach could be to quantify this using DAPI and Erg in existing images, of the total DAPI+, what are GFP+/DAPI+/Erg- (fibroblasts) vs GFP+/ERG+/DAPI+ (endothelial). The figure would be improved by adding DAPI to one set of panels with GFP/ERG, this would show a lot of DAPI+/GFP- cells, encompassing CD206+ cells and non-recombined fibroblasts.

      Thank you for overseeing this manuscript.

    1. eLife Assessment

      This study employs state-of-the-art quantitative imaging and genomics approaches to address a fundamental question regarding the establishment of Polycomb domains during Drosophila embryogenesis. The authors identify the maternal-to-zygotic transition, rather than earlier developmental stages, as the critical window for Polycomb domain establishment, providing significant clarification of the timing of this process and they further investigate the roles of Zelda and GAGA factor, demonstrating that Zelda, but not GAGA factor, contributes to Polycomb domain establishment. Overall, the conclusions are well supported by the experimental evidence presented. These compelling findings advance our understanding of how Polycomb domains are established during early embryogenesis and have broad implications for chromatin regulation and developmental biology.

    2. Reviewer #3 (Public review):

      Gonzaga-Saavedra et al report an analysis on genomic binding of Polycomb group proteins, and of H2Aub1 and H3K27me3 domain formation in the early Drosophila embryo. Using carefully stage embryos during the nuclear cycles (NC) leading up to the cellular blastoderm stage, the authors provide compelling evidence that H3K27me3 domains at PcG target genes are only established during NC14 and do not exist in NC13. In contrast, H2Aub1 domains already start to appear during NC13. The authors show that E(z), the catalytic subunit of the H3K27 histone methyltransferase PRC2, is readily detected in interphase nuclei during the rapid nuclear divisions in pre-blastoderm embryos. In contrast, the DNA-binding proteins Pho, Cg and GAF that are known (Pho) or have been postulated (Cg, GAF) to anchor PRC2 and PRC1 to Polycomb Response Elements (PREs) in Polycomb target genes only start to show nuclear localization from NC10 onwards with gradually increasing nuclear concentrations, reaching a maximum during NC14. These data strongly corroborate the simple straightforward view that targeting of PRC2 and PRC1 to PREs by sequence-specific DNA-binding proteins is a pre-requisite for the formation of H3K27me3 and H2Aub1 domains at Polycomb target genes.

      The authors then explore the potential role of GAF/Trl in this process. They find that in embryos depleted of GAF/Trl, H3K27me3 domain formation is largely unperturbed.

      The authors also depleted the pioneer factor Zelda (Zld) and found that removal of Zld results in a more complex outcome. Zelda appears to counteract accumulation of H3K27me3 at the Polycomb targets eve and zen but also appears to be required for effective H3K27me3 domain formation at Polycomb targets such as amos or atonal.

      This is a very thorough study that reports data of superior technical quality that are highly relevant for the field. The study by Gonzaga-Saavedra et al extends and strengthens previous work from the labs of Eisen (Li et al, eLife 2014) and Zeitlinger (Chen et al, eLife 2013) to convincingly demonstrate that Polycomb domain formation in the early embryo occurs during ZGA but that such domains do not exist prior to ZGA. This should now finally put to rest earlier claims by the Iovino lab (Zenk et al, Science 2017) that H3K27me3 domains present in the zygote nucleus would be propagated and partially maintained during the rapid nuclear cleavage cycles and serve as seeds for H3K27me3 domain formation during ZGA.

      The experiments analyzing H3K27me3 domain formation in embryos depleted of GAF/Trl or Zelda will be of great interest to the field.

      Comments on revised version.

      In the revised version, the authors have addressed the comments and suggestions raised by this reviewer and added the missing references to earlier work.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This well-conceived manuscript investigates the mechanisms that shape the chromatin landscape following fertilization, using the Drosophila embryo as a model system. Importantly, the authors revisit conflicting data using new approaches and analysis to show that the silent H3K27me3 mark deposited by PRC2 is established de novo in the embryo in coordination with the slowing of the nuclear division cycle and activation of zygotic transcription. Unexpectedly, they demonstrate that the transcription factor GAF is not required for the deposition of this mark, but that the well-studied pioneer factor Zelda, which is required for widespread gene expression, is required for H3K27me3 deposition at a subset of regions. The experiments are rigorously performed, and interpretations are clear. Strengths of this manuscript include the rigor of the experimental design, careful analysis, and well-supported conclusions. Some additional citations, analysis, and broadening of the Discussion section to include additional models and data would further strengthen this manuscript.

      We appreciate the reviewer’s positive assessment and in revision we have revised the Discussion to clarify some of the mechanistic insights of the work as well as including references to related studies in the zebrafish model system.

      Reviewer #2 (Public review):

      Summary:

      Epigenetic silencing of target genes by the Polycomb pathway is central to maintenance of cell fates during development and depends on repressive chromatin states involving Polycomb complexes and histone modifications. However, the mechanisms by which these chromatin states are built at the earliest stages of development are unclear. Here, Gonzaga-Saavedra and colleagues use the premier experimental system for studying Polycomb gene regulation, Drosophila development, to investigate when Polycomb domains emerge and how they are assembled. Using a combination of CRISPR gene editing, imaging, and genomic profiling, they determine that while H3K27me3 is initially present in the first nuclear cycles, it quickly dissipates and does not re-emerge until mid-nuclear cycle 14, during the major wave of zygotic genome activation (ZGA). This finding helps resolve current discrepancies in the field, informs potential mechanisms of transgenerational inheritance, and indicates that repressive Polycomb domains are built de novo on target genes in embryogenesis. The authors then set out to examine how Polycomb domains are built. Through live imaging and immunofluorescence, they determine that the histone H3K27 methyltransferase, E(z), is present in nuclei at high levels throughout cleavage and blastoderm stages. By contrast, they determine that several Polycomb proteins that bind PREs (cis elements that demarcate Polycomb targets in the genome) are absent from early cleavage nuclei and progressively increase following nuclear cycle 10. These findings suggest that the absence of H3K27me3 in early embryos may be due to failure to assemble functional Polycomb complexes at target genes. Lastly, the authors test the requirement of two transcription factors with important roles in ZGA, GAF, and ZLD. Despite binding to many PREs and regulating chromatin accessibility in early embryos, they find that GAF is largely dispensable for the emergence of H3K27me3 domains. On the other hand, they find that the pioneer factor ZLD is required for proper H3K27me3 emergence; in its absence, some Polycomb domains accumulate greater levels of H3K27me3, whereas other Polycomb domains accumulate less H3K27me3.

      Strengths:

      The strengths of this study are manifold. It studies an important topic with broad interest to the chromatin and epigenetics fields. It is well-written with detailed method descriptions. In addition, the experimental design and rigor of execution are exceptional despite working with very small amounts of biological material. Example strengths include that the Polycomb proteins studied were tagged with the same epitope, permitting direct quantitative comparisons in imaging and in genomics experiments. Microscopy studies are quantified and performed both via live imaging and via immunofluorescence. The microscopy studies reinforce and extend conclusions made via ChIP. Sophisticated loss-of-function analyses allow for direct mechanistic tests of Polycomb domain emergence.

      Weaknesses:

      Overall, the study is quite strong already, but it can be further strengthened in several ways. First, several conclusions should be refined based on the data presented. Second, the extent to which ZLD is important for initiating Polycomb domain formation should be made clearer. Third, additional genomic profiling experiments are needed to provide insight into models explaining why H3K27me3 is absent prior to NC14.

      We are grateful for the reviewer’s thorough and supportive comments. We have revised certain assertions and conclusions for objectivity. For the point about providing “insight into models explaining why H3K27me3 is absent prior to NC14,” we have a separate study that addresses this issue directly (Degen, Gonzaga-Saavedra, and Blythe, bioRxiv 2025, in press). In summary, we find evidence that a maternal PcG imprint is indeed maintained through cleavage divisions, albeit through lower-order methylation states (maximally, H3K27me2). We chose not to include these additional results in this manuscript to maintain the focus of this study on ZGA. Our revision of the manuscript includes a reference to this associated study in the Discussion.

      Reviewer #3 (Public review):

      Gonzaga-Saavedra et al report an analysis on genomic binding of Polycomb group proteins, and of H2Aub1 and H3K27me3 domain formation in the early Drosophila embryo. Using carefully staged embryos during the nuclear cycles (NC) leading up to the cellular blastoderm stage, the authors provide compelling evidence that H3K27me3 domains at PcG target genes are only established during NC14 and do not exist in NC13. In contrast, H2Aub1 domains already start to appear during NC13. The authors show that E(z), the catalytic subunit of the H3K27 histone methyltransferase PRC2, is readily detected in interphase nuclei during the rapid nuclear divisions in pre-blastoderm embryos. In contrast, the DNA-binding proteins Pho, Cg, and GAF that are known (Pho) or have been postulated (Cg, GAF) to anchor PRC2 and PRC1 to Polycomb Response Elements (PREs) in Polycomb target genes only start to show nuclear localization from NC10 onwards with gradually increasing nuclear concentrations, reaching a maximum during NC14. These data strongly corroborate the simple, straightforward view that targeting of PRC2 and PRC1 to PREs by sequence-specific DNA-binding proteins is a prerequisite for the formation of H3K27me3 and H2Aub1 domains at Polycomb target genes.

      The authors then explore the potential role of GAF/Trl in this process. They find that in embryos depleted of GAF/Trl, H3K27me3 domain formation is largely unperturbed.

      The authors also depleted the pioneer factor Zelda (Zld) and found that removal of Zld results in a more complex outcome. Zelda appears to counteract the accumulation of H3K27me3 at the Polycomb targets eve and zen, but also appears to be required for effective H3K27me3 domain formation at Polycomb targets such as amos or atonal.

      This is a very thorough study that reports data of superior technical quality that are highly relevant for the field. The study by Gonzaga-Saavedra et al extends and strengthens previous work from the labs of Eisen (Li et al, eLife 2014) and Zeitlinger (Chen et al, eLife 2013) to convincingly demonstrate that Polycomb domain formation in the early embryo occurs during ZGA but that such domains do not exist prior to ZGA. This should now finally put to rest earlier claims by the Iovino lab (Zenk et al, Science 2017) that H3K27me3 domains present in the zygote nucleus would be propagated and partially maintained during the rapid nuclear cleavage cycles and serve as seeds for H3K27me3 domain formation during ZGA.

      The experiments analyzing H3K27me3 domain formation in embryos depleted of GAF/Trl or Zelda will be of great interest to the field.

      We thank the reviewer for recognizing the strength of our data and conclusions, and we agree that our results help settle conflicting claims in the field. We have emphasized Zelda’s context-dependent effects more clearly in the revised manuscript.

      Recommendations for the authors:

      Reviewing Editor Comments:

      It would strengthen the manuscript to more fully acknowledge and discuss related work in the field. Addressing the caveats in the functional analyses, either through editorial clarification or additional experiments, would also improve the study.

      For recommendations to the authors, comments from each reviewer are listed below.

      Reviewer #1 (Recommendations for the authors):

      Some additional analysis would clarify the relationship between pPREs and transcription factors.

      Because of the limitations of depleting Pho, the model that nuclear levels of Pho, Cg, and GAF regulated E(z) activity is purely correlative, as GAF depletion did not change H3K27me3 distribution. As it stands, it is possible that this NC14 nuclear enrichment of these factors is not relevant. The statement on lines 295-297 regarding the correlation between re-establishment of the modification state and nuclear localization of Pho, Cg, and GAF is true, but a bit misleading since there is no evidence to support the necessity of these factors for H3K27me3 establishment. Other models remain possible, and the discussion should be toned down to account for this. Furthermore, the only factor that is shown to influence H3K27me3 is Zelda, which does not show an increase in nuclear localization at NC14.

      We thank the reviewer for highlighting this issue. As the reviewer indicates, we lack definitive mechanistic evidence that limited nuclear localization of any recruitment/nucleating factor is limiting for H3K27me3 deposition. However, our data are, we feel, definitive in terms of demonstrating that nucleation from PREs and PRE-like regions arises at mid-NC14 for H3K27me3, and in late cleavages for H2Aub. Ideally, we would test each known nucleating factor for a necessary role in mediating this activity. We have chosen not to include such measurements in this manuscript because a proper mechanistic treatment of each factor (Pho, Cg, and others) would need to be extensive, and we feel better suited for an independent study. We have added language to the “Limitations of the Study” section to reflect the remaining need to identify the key nucleating factors responsible for establishment of the zygotic PcG landscape.

      We re-read the lines the reviewer suggested were misleading and we respectfully disagree. The lines (“Taken together, these observations are consistent with a model where…re-establishment of this modification state is restricted to late cleavage divisions by limiting nuclear localization of nucleating factors such as Pho, Cg, and GAF.”) are expressing a hypothesis/model, and in our opinion these lines are suitably framed as to not be misleading.

      Additional analysis/discussion regarding the relationship between Zelda, H3K27me3, CBP, and paused polymerase would provide further clarity into how Zelda might promote this methylation. It would be useful to discuss how the various classes defined correlate with enhancers versus promoters, and also the gene expression of the underlying gene. It would be clarifying to discuss how the H3K27me3 at these Zelda-dependent regions relates to gene silencing since many of the genes depend on Zelda for expression. Are these genes expressed prior to NC14 and then silenced at this time point? How do H3K27ac levels, which Zelda promotes through recruitment of CBP, relate to the various pPRE classes? The authors should also consider the report from the Mannervik lab that CBP is instrumental in promoting H3K27me3 (Hunt, Boija, Mannervik et al. Mol Cell 2022 82:3580-3597) and the relationship with paused RNA Pol II. Given this possible connection, it could be useful to consider that paused polymerase is also first evident at NC13/14. Overlaying the pPREs with paused polymerase from Chen et al. 2013 eLife (current citation 30) could be informative. At a minimum, a discussion of these additional mechanisms and the implications of the data in Hunt et al. would strengthen the Discussion section.

      We thank the reviewer for this comment. We agree that additional analysis of the relationship between Zelda, H3K27me3 and CBP would be essential for providing mechanistic clarity on the role of Zelda for putting these loci into play for apparently either positive or negative regulation. At present, we feel that extensive additional analysis, including work with CBP, PRC1, and nucleating factors such as Pho would be better suited for a future study.

      H2Aub is clear earlier in development than H3K27me3, and in mice and zebrafish, it promotes PRC2-mediated H3K27me3 (Hickey et al. eLife 2022 doi: 10.7554/eLife.67738, citations 86, 89). As such, it remains possible that this mark is instructive for the H3K27me3 deposition observed. As such, a bit more analysis of where H2Aub is deposited and how it overlaps with pPREs might help determine whether similar mechanisms could be important in Drosophila.

      We suspect that similar mechanisms are important in Drosophila as well. The omission of the Hickey…Cairns reference was an oversight in the original document. We have revised this sentence to refer to both mouse and zebrafish and have added the citation.

      Prior work has noted the sudden increase in GAF nuclear concentration. Please cite Dima and Reeves. Development 2025 152:dev204460 in support of the observations shown in Figure 3D.

      Thank you. Yes, this article was published shortly after we submitted this manuscript for review and we have now added it to reflect its independent replication of the GAF nuclear concentration result.

      It is not clear that Figure 4 warrants an entirely new figure, since the conclusions drawn are similar to/the same as Figure 3. Perhaps change to a supporting figure?

      We agree that this figure was repetitive. We have now moved it to a figure supplement of Figure 3.

      The authors have developed a powerful modified ChIP protocol that enables them to perform the experiment on the equivalent of 10 embryos! This is not highlighted in the manuscript, despite the vast improvement this provides. The authors should highlight this in the manuscript, unless a separate manuscript describing this technique is being written/published. Regardless, this is a very exciting protocol.

      Thank you. A methods paper describing this approach is in preparation.

      Minor:

      (1) Line 49-53: clarify that, as opposed to the mechanisms described earlier in the paragraph, these mechanisms are specific to Drosophila.

      Done.

      (2) Line 61: cite Sun et al. and Schulz et al. (current citations 60 and 61) since these papers demonstrated the pioneering function of Zelda.

      Done.

      (3) Line 95: a word seems to be missing. Perhaps "approach (STAN) to identify a set of PcG "domains" from our"?

      We have made this revision.

      (4) Line 217: Calling an embryo a "specimen" is odd. Can you just say all embryos?

      Ok.

      (5) Line 316: GAF is encoded by Trithorax-like, not Trithorax-related.

      Revised.

      (6) Figure 5: The arrowheads and asterisk are so small that they are nearly impossible to see when the figure is printed. Please make it larger and perhaps use a color that stands out more.

      We have made the arrowheads and asterisk larger and changed the color to yellow to improve visibility.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding weakness 1, the title states "nucleation-dependent propagation". This is an overstatement of the study's conclusions because these features are inferred and not directly tested here. This is an easy fix with text revisions.

      We acknowledge that we have not directly tested the role of specific nucleating factors for the process of nucleation-dependent propagation. This is reflected in the Discussion text. We have added to the “Limitations of the Study” section the following text: “Finally, we acknowledge that further loss-of-function analysis will be necessary to determine the key nucleating factors responsible for the initial establishment of the zygotic H2Aub and H3K27me3 landscape.” However, we feel that we have demonstrated definitively that, by mid-NC14, nucleation of H3K27me3 sites is first detected genome-wide. As such, we have left the title as-is.

      (2) Also, regarding weakness 1, the abstract makes additional overstatements. These are easy fixes with text revisions.

      (a) "A large subset of targets requires ZLD..." (line 24). This phrase makes it sound like the majority of Polycomb domains depend on ZLD, which I am not sure is accurate.

      Thank you for pointing this out. We agree with this assessment and have revised the abstract to remove the word “large” so that now the sentence reads “; a subset of targets requires Zelda…”

      (b) "...requires ZLD not for PcG factor recruitment" (same sentence). This sentence should state "E(z)" instead of "PcG factor" because E(z) was the only one tested in ZLD mutants.

      We agree as well and have made the requested change to exchange “PcG factor” to “E(z)”.

      (c) "to license a loaded PRE" (same sentence). Whether PREs are fully loaded in ZLD mutants was not tested. In addition, see comments below for feedback on the use of the term "license."

      We have revised this to read “to license an E(z)-loaded PRE,” to keep consistent with the above suggested change. We also removed the mention of “H2Aub” because now the PRE-loading acknowledges we have only measured E(z), although we see effects on both modifications. We acknowledge, in light of Reviewer 3’s comment, that we have not directly measured whether PRC1 still can bind to certain PREs in the absence of Zelda like we see with E(z)/PRC2.

      (3) Also regarding weakness 1, the authors set up a dichotomy for ZLD's role in initiating Polycomb domain formation, either acting as a pioneer or as a licensing factor. However, based on the data presented (e.g. the atonal browser shot), it appears that chromatin accessibility is lost in both sub-classes of H3K27me3 domains that depend on ZLD, meaning that ZLD's role as a pioneer would explain both sub-classes, and that a licensing role is not supported by the data. To assess whether there are pioneer-independent roles of ZLD in the emergence of Polycomb domains, it may help to examine the role of chromatin accessibility changes directly (e.g., is there a substantial fraction of changes in H3K27me3 domains or E(z) peaks that cannot be attributed to changes in chromatin accessibility in ZLD mutants?). The authors should refine their language or provide additional support for the non-pioneering role of ZLD.

      We acknowledge that the pioneer/licensing distinction is still not clear from a mechanistic perspective and that future work will be needed to address this issue. We do not mean to imply that the licensing is necessarily independent of pioneering, rather that at sites like atonal, clearly chromatin accessibility is not the only job that Zelda performs. There are at least two possible mechanisms: 1) it is all about pioneering, and in the absence of accessible chromatin, some other factor required for stimulating H3K27me3 deposition is unable to bind; or 2) it is all about Zelda, and in mutants Zelda both does not confer accessible chromatin, and also does not stimulate H3K27me3 deposition through whatever means. For either mechanism, the remarkable feature is that E(z) still localizes to its genomic target and requires additional information to deposit H3K27me3. This is distinct from the other class (represented by amos in Figure 6 in the final manuscript) where loss of accessibility correlates with loss of E(z) –and presumably PRC2– binding. To explain this activity, we have invoked the term “licensing” which we feel captures the effect of Zelda (whether it be direct or indirect). The possibility of indirectness is a significant caveat, so we have added clarifying text to indicate this possibility. We have added to paragraph 2 of the Discussion the sentence, “As such, we emphasize that Zelda-dependent licensing could stem either directly or indirectly from Zelda function.”

      (4) Regarding weakness 2, as currently written, it seems like a small fraction of H3K27me3 domains depend on ZLD. Is this accurate? Some of my confusion may stem from alternating use of bins and runs. Can the fraction of domains that depend on ZLD be made more explicit through text revisions and additional bioinformatics? More comprehensive bioinformatics analyses can be performed with the ZLD mutant datasets by incorporating ATAC and E(z) peaks. For instance, what fraction of E(z) peaks inside and outside of Polycomb domains are affected in ZLD mutants, and do these E(z) changes correlate with H3K27me3 changes? Similarly, what fraction of PREs change in accessibility in ZLD mutants, and are these accessibility changes correlated with H3K27me3 changes? When do PREs become accessible during wild-type embryogenesis? Are there unique features of ZLD-dependent H3K27me3 domains or E(z) peaks? The authors' perspective on the extent to which ZLD is required for Polycomb domain initiation should also be added to the Discussion.

      We find that 38 PcG domains are sensitive to Zelda (9 have increases, 29 have decreases in H3K27me3), and this is stated in the Results. This accounts for 16% of domains, which is a small fraction of the total. We have added a sentence to the results reporting this fraction. An earlier draft of this manuscript included an analysis of accessibility: both in terms of the timing of when PREs gain accessibility and their dependency on Zelda function. This section was omitted from the submitted manuscript because it did not add clarity the distinction between classes that we report. As for the final request that we add to the Discussion our perspective on the extent to which Zelda is required for PcG domain initiation, we now address this in the second paragraph of the Discussion.

      (5) Regarding weakness 3 (why is H3K27me3 missing pre NC14?), it is suggested that the absence of H3K27me3 is due to the short duration of nuclear cycles relative to the rate of me2->me3 catalysis. And although the authors suggest that Polycomb complex assembly occurs at pPREs without H3K27me3, their microscopy studies indicate that binding of Polycomb proteins to PREs may be regulated via nuclear accumulation. Therefore, it remains unclear whether (and when) the lack of H3K27me3 is due to incomplete Polycomb complex assembly. Additional genomic profiling experiments are needed to directly test when Polycomb complexes assemble on chromatin relative to the emergence of H3K27me3 domains. These experiments would also support claims of nucleation.

      (a) ChIP of E(z) and a PRE binding protein (Pho, Gc, GAF) should be performed at NC13 to test whether Polycomb proteins (especially PRC2) are bound at PREs prior to H3K27me3 emergence.

      (b) ChIP of E(z) should also be performed at NC10 when the PRE binding proteins appear to be absent but when E(z) is hyperabundant in the nucleus. Can E(z) bind PREs in the absence of these "nucleators"? Or does E(z) promiscuously interact with chromatin, helping to explain the broad, low-level H3K27me1 enrichment profile?

      We thank the reviewer for this comment. We have addressed this issue in a separate study (preprinted and accepted for publication as of this writing). Although H3K27me3 is not detectable on cleavage-stage chromatin, lower-order H3K27me2 is maintained on chromatin and detected throughout cleavages by immunostaining and by ChIP. The maintenance of cleavage-stage H3K27me2 depends on both E(z) and Esc. Notably, this lower-order state reflects maintenance of a maternally supplied H3K27 methylation state: H3K27me2 is only detected on maternal (not paternal) chromatin during the period (prior to NC10) when Pho/Cg/GAF do not localize to nuclei. Overall, these observations underscore the limiting nature of early cleavages to support de novo establishment of H3K27 methyl states, but also demonstrate the competency of the PcG system during this time to engage in some degree of H3K27 maintenance.

      Minor Comments:

      (1) It is interesting that the great majority of E(z) peaks (72%) are outside of H3K27me3 domains at NC14. What are these sites? Do these sites correspond to Polycomb domains later in embryogenesis (ie, do they become marked by H3K27me3 later)? Or, do these E(z) peaks disappear at later stages of embryogenesis?

      We agree that this observation is interesting. Figure 2 shows that the majority of these sites are concurrent with transcription start sites. Visual comparison between ChIP datasets generated here and any of the various publicly available datasets generated at later stages (e.g., stage 16 embryos) reveals that the set of PcG domains observed at ZGA is fairly consistent with the set of PcG domains observed later on. Therefore, on a bulk level, these extra-domain E(z) sites observed at ZGA do not represent later-onset PcG domains. However, at this level of resolution, we cannot rule out that in some cell type these sites are converted to a cell-type-specific PcG domain. We have not addressed the perdurance of these E(z) peaks at later stages. The function and the fate of these extra-domain binding sites remains an open question for future investigation.

      (2) Figure 3A live imaging indicates that E(z) nuclear signal intensity diminishes over subsequent nuclear cycles. Can the authors expand on whether photobleaching may contribute to this decrease? The E(z) signal in IF experiments should be quantified to test whether it also diminishes over successive nuclear cycles.

      The imaging conditions for this experiment were controlled to minimize photobleaching in the EGFP channel. While we cannot rule out some small contribution of photobleaching to the overall signal intensity, the overall trend reported in the quantification of Figure 3A reflects, to the best of our abilities to measure, a biological effect. This effect is also evident in the immunofluorescence imaging without need for additional quantification: From NC10 to NC14, when nuclei are presented on the embryo surface, a clear decrease in staining intensity is observed (Figure 3-figure supplement 1A in the final manuscript). Our observations indicate that E(z) decreases in concentration with increasing nuclear content during the cleavage divisions.

      (3) Can the authors provide further interpretation of the H3K27me1 signal profile? There is very little difference in signal amplitude inside relative to outside domains in Figure 1C, and across a 20kb window surrounding E(z) peaks in Figure 1B. Is all this chromatin considered to be H3K27me1-enriched, or is there a high level of noise? The anticorrelation with H3K27me3 at NC14 is compelling and seems to argue that the broad H3K27me1 signal is real.

      The H3K27me1 signal profile reflects a likely broad distribution of H3K27me1 that includes not only canonical “PcG domains” (i.e., regions that ultimately gain high-level H3K27me3) but also non-canonical domains. Consistent with observations in other systems (e.g., PMID: 24289921), H3K27me1 is broadly distributed across the Drosophila genome. We have added the following sentence to the Results section to contextualize our description of the H3K27me1 domains: “The broad genome-wide distribution of H3K27me1, including in regions outside of canonical PcG domains, resembles profiles previously measured in mammalian tissue culture.” (citing the Ferrari et al study).

      Reviewer #3 (Recommendations for the authors):

      My suggestions for changes are mainly of an editorial nature.

      My comments below are not in order of priority, but grouped into different types of suggestions for changes.

      Comments on the presentation of results:

      (1) The relationship between the genomic regions shown in the heat map in Figure 1A and Figure 2A is not clear. In Figure 2A, 1264 regions with E(z) peaks in PcG/H3K27me3 domains are shown. In Figure 1A, the text (line 97/98) states that 237 PcG/H3K27me3 domains contain 1 or more E(z) peaks, and on the left of the H3K27me3 heat map panel, it says "PcG domains". Are these then the 237 regions that are shown in Figure 1A? Only the top ones seem to have high levels of H3K27me3. Please clarify.

      We could have been more clear about this. As the reviewer points out, there are different numbers of peaks shown in the heatmaps in Figures 1A and 2A. Figures 1A and B plot one E(z) peak per domain, determined by finding the one maximal E(z) peak per and plotting it. This is stated in the figure legend: “One representative E(z) peak per domain was selected for plotting.” Figure 2A plots all E(z) peaks within Domains (n = 1264). For the heatmaps and the average plots in Figures 1A and 1B, we found that selecting the maximal E(z) peak yielded a more accurate representation of the accumulation of H3K27me1/3, compared with plotting this for all E(z) peaks within domains, presumably because some called peaks (e.g., many of the minor E(z) peaks shown in Figure 1E) do not appear to be the primary sites of nucleation within the domain. For Figure 2, we wished to compare all E(z) peaks, inside and outside of domains, and for this we plotted heatmaps over all of the peaks.

      (2) In general, in the text, the figure panels should be discussed in order of appearance. For example, it is not ideal that after describing the data in Figure 1A, the text jumps to discuss results shown in Figure 2A and only then goes back to discuss Figure 1B and C. This should be easy to resolve by reorganizing the text or perhaps re-arranging figure panels (e.g., perhaps moving elements to additional supplemental figures?).

      We acknowledge that this is inconvenient and we apologize for this. However, we feel the flow of the manuscript and the figures themselves benefit from this somewhat awkward order of discussion and we have chosen to leave them as-is.

      (3) In general, I felt that the text could be improved by putting some of the more detailed technical procedures and result descriptions into Materials and Methods or Figure legends in order not to disrupt the flow of the text.

      Comments for improving discussion:

      (4) When citing and discussing the previous literature that claimed a function for GAGA factor (GAF/ Trl) in Polycomb repression (i.e., refs. 12, 14, 49-52) on page 15/16 and again on page 25, the authors may also want to cite earlier studies that failed to observe a function of GAF/Trl in Polycomb repression. Specifically, previous studies (Brown et al, Development 2003) had investigated a possible role of GAF/Trl in Polycomb repression in larvae using stringent tests (analysis of HOX gene expression in Trl null mutant cell clones in imaginal discs, removing Trl in a sensitized pho null mutant background, mutation of GAF/Trl binding sites in HOX-LacZ reporter genes) and had found no evidence for a role of GAF/Trl in Polycomb repression in larval tissues. It is fair to say that in most of the studies cited (i.e., references 12, 14, 49-52), the effect of GAF/Trl had been analyzed using PRE-miniwhite reporter gene activity as read-out, a much less stringent assay.

      Thank you for pointing this out. Brown et al (reference 15) was overlooked when entering citations to these sections. We have added a sentence to the discussion to highlight the lack of necessity for Trl/GAF for imaginal disc silencing of Hox targets.

      (5) The authors report that removal of GAF/Trl had almost no impact on H3K27me3 domain formation. Although this supports a model where PRC2 recruitment to PREs is unaltered after GAF/Trl depletion, it does not eliminate the potential caveat that PRC1 binding may be affected. Do the authors have binding profiles of PRC1 subunits or the H2Aub1 profile in embryos lacking GAF/Trl? It seems that the current paper would be a great opportunity to include such data and get them published. There is no need to generate PRC1 or H2Aub1 profiles if they don't already exist.

      Unfortunately, we do not as of yet have any PRC1 reagents that are compatible with ChIP. Following the lack of effect of GAF on H3K27me3, we did not pursue the measurements of H2Aub in these knockdown conditions, given the likely dependence of H3K27me3 on H2Aub deposition at ZGA.

      (6) Previous studies found that Polycomb group protein complexes bind to PREs in HOX genes both in cells where genes are OFF but also in cells where genes are ON but that the H3K27me3 profile is different in the two states, with H3K27me3 decorating the gene in the OFF but not in the ON state (Papp and Müller, Genes Dev 2006; Bowman et al, eLife 2014). In this study, the H3K27me3 profiles at target genes represent the sum of ChIP signals coming from cells where the gene is OFF and cells where the gene is ON. Is it known how eve and zen are deregulated in Zelda-depleted embryos? Could it be that the increased H3K27me3 enrichment at eve and zen is an indirect effect caused by loss of eve and zen expression in a large fraction of cells and consequently a gain of H3K27me3 ChIP signal at the gene in those cells? It may be worth at least discussing such scenarios in light of the studies mentioned above, and referring to them.

      Thank you for raising this question. We have a manuscript in preparation that specifically addresses this issue, namely the relationship of on/off states to H3K27me3 deposition in the early embryo. In short, it is likely that the increase in H3K27me3 at eve reflects significantly reduced eve expression in zelda mutant embryos. We have added a clarifying sentence to the Discussion and added the indicated references.

    1. eLife Assessment

      The study presents valuable findings on early behavioral phenotypes that arise in an ADHD-associated adgrl3.1 mutant zebrafish, using a new behavioral approach with a closed-loop OMR assay. The evidence supporting the claims of the authors is solid; however, the validation of the method and analysis is incomplete and would benefit from more rigorous approaches and validation practices. This work will be of broad interest to developmental biologists and behavioral neuroscientists.

    2. Reviewer #1 (Public review):

      Summary:

      Reynolds and colleagues provide a deep phenotypic analysis of behavior in adgrl3.1 mutant zebrafish at larval stages using a closed-loop optomotor response (OMR) assay. The analyses conducted are interesting and extract new locomotor phenotypes with possible relevance to the role of adgrl3.1 in ADHD. Reduced interbout interval (both in the OMR assay and in dark rest periods) and increased distance moved provide greater resolution on hyperactivity phenotypes already described in these mutants. Reduced variation in interbout interval and reduced variation in swim speeds throughout the assay provide new insights into how behavior is altered; the authors suggest that these findings reflect more stereotyped, less flexible behavior in adgrl3.1 mutant animals. Analyses of task performance are interesting and could help understand how / whether animals maintain vigilance over time in the OMR assay and reveal trends in adgrl3.1 mutants relative to siblings, but ultimately do not identify significant phenotypes for adgrl3.1 mutants. While methods are extremely clear and analyses and phenotypic insights are solid, the authors do not provide sufficient support for assertions that their paradigm separates anxiety from locomotor activity or extracts phenotypes central to ADHD (impulsiveness, attention, etc). In some instances, interpretation of behavioral phenotypes in the context of disease is difficult to follow or not well supported with citations, etc.

      Strengths:

      (1) Deeper phenotypic analysis of adgrl3.1 locomotor phenotypes reveals changes to bout timing/initiation of locomotion as potentially causative for broader hyperactivity phenotypes previously reported.

      (2) Interesting dissection of OMR performance over time and variability in locomotor parameters, assessment of OMR performance in high- and low-contrast.

      (3) Methods are clearly described and considered to be rigorous.

      Weaknesses:

      (1) The introduction does not clearly spell out why the closed-loop OMR assay is expected to capture phenotypes central to ADHD (impulsiveness, attention, etc). Similarly, it's stated in the discussion that hyperactivity is driven by shorter inter-bout intervals and longer bout lengths...reflecting a reorganization of locomotor timing," and that "such fine-scale insights are not possible in standard light/dark paradigms." But in fact, each of these parameters was examined in the dark periods and could be assessed in a standard light/dark assay. As explained at the end of the discussion, this work provides a detailed analysis of locomotion and extracts new and interesting phenotypes, but the assertion that this is a function of the assay / that the assay is uniquely relevant to ADHD is not well-supported. The final statement of the introduction more accurately captures the advantages of the assay used: "this allowed us to assess whether loss of adgrl3.1 alters not only overall locomotor drive...but also specific visuomotor behavioral responses under different stimulus demands."

      (2) The statement early in the results that "this approach extends beyond classical locomotor assays conducted in static light / dark environments, where locomotor activity may conflate with anxiety-related responses" and later that the closed-loop OMR assay "disentangles hyperactivity from anxiety-related responses" are not well-supported. Anxiety states could influence performance on OMR (Braun et al., 2024, Molec Psychiatry).

      (3) Some interpretations of the phenotypes are not well supported by citations and may be overstated. For example, "adgrl3.1 elevates baseline arousal...producing a phenotype of heightened but less exploratory visuomotor activation." Since bout duration is increased alongside reduced interbout interval and increased total distance traveled, reduced exploration is not well-supported by the data. Later in the results, it's suggested that the increase in distance traveled reflects "over compensatory hyperactivity under ambiguous sensory conditions, consistent with attentional deficits." It's not clear what this means - references would be helpful to create links between hyperactivity and detection of ambiguous sensory conditions, and also between hyperactivity under these conditions and attentional deficits.

    3. Reviewer #2 (Public review):

      In the study by Reynolds et al., the authors propose a new behavioral approach for ADHD evaluation using a genetically modified zebrafish model. The study is interesting and has potentially important implications for the field. However, there are several methodological, analytical, and validation-related issues that should be addressed before the study can be considered scientifically rigorous.

      Comments:

      (1) Introduction section

      What is the epidemiological evidence supporting the prevalence of ADGRL3 dysfunction in the human population? The authors should consider adding this information, as well as clarifying where ADGRL3 mutations rank among other genetic variants associated with ADHD.

      I suggest reconsidering the sentence "quantifiable behavioural repertoires that complement rodent approaches." Zebrafish studies do not simply complement rodent studies; they can serve as independent pharmacological and toxicological tools that may be used in parallel with rodent models.

      The statement "forced light/dark (FLD) locomotion test" is too broad. Are the authors referring to the Visual Motor Response Test? If so, this is a robust assay that can evaluate not only anxiety-like responses but also locomotor state, arousal, decision-making, and potential cognitive impairment in larvae. A clearer description of the assay is necessary, especially to justify the statement that "while useful for detecting overall activity differences, it cannot determine whether increased movement reflects hyperactivity, altered arousal, disrupted behavioural control, or anxiety-like responses." In contrast, subtle behavioral changes across light and dark phases can be highly informative when velocity, time moving, and anxiety-like responses are analyzed together.

      (2) Methods section

      The zebrafish husbandry section lacks several essential details. The authors should include fundamental information, such as the embryo medium used, how embryos were obtained, the age of the breeding adults, and how larval age was determined in hours post-fertilization. These details are necessary for proper interpretation and reproducibility of the data.

      Why was the mutant DNA not sequenced? Although agarose gel electrophoresis can provide useful preliminary evidence of mutation, it cannot precisely determine the number or nature of base-pair changes. This information is essential because different mutations can have distinct impacts on gene function.

      The sentence "All statistical analyses and tests were completed on Prism10 (GraphPad)" is insufficient. It should be specified which statistical tests were used, including assumptions tested, post hoc comparisons, correction methods, and how experimental replicates or batch effects were handled.

      Overall, the methods section lacks sufficient information to support the scientific rigor of the study. It is unclear whether every individual evaluated behaviorally was injected and then only a subset was genetically confirmed, or whether stable breeding matrices were generated and all experimental individuals were derived from these parents. The manuscript mentions "2-6 parent batches per experiment, with batches collected and run on separate days," but the genetic origin and validation of these batches remain unclear.

      If embryos were injected for each batch, how did the authors ensure that the mutation was homogeneous enough across individuals to consider them equivalent? How did the authors confirm that the mutation was homozygous or present across all relevant cells? Zebrafish embryos remain at the single-cell stage for only a short period before mitosis begins. Without detailed information regarding breeding timing, embryo collection, injection timing, and sequencing validation, it is difficult to determine whether the injected embryos developed homogeneous mutations or mosaic patterns.

      Additionally, to claim a knockout model, protein-level validation, such as Western blotting or another protein expression assay, should be provided. At present, there appears to be some confusion between a knockout and a knockdown model.

      To validate a new behavioral protocol, the authors should compare their assay with an established gold-standard behavioral paradigm using the same experimental batch. They should clarify why this comparison was not performed.

    4. Reviewer #3 (Public review):

      Summary:

      The study provides an in-depth phenotyping of a novel zebrafish larval model of ADHD. This topic is interesting, and the model and the approach are relevant and well-justified. While the paper has a massive amount of high-quality data, the general structure and presentation of this material lack focus and a clearly articulated rationale.

      Strengths:

      The paper is methodologically sound, well-presented, and well- illustrated. It has a clear logical rationale and reasonable experimental design.

      Weaknesses:

      The amount of high-quality data is impressive, yet the general structure and presentation of this material lack focus and a clearly articulated rationale.

      (1) First, it is unclear why VR is necessary here. It needs a better explanation in both the abstract and the intro section of the manuscript.

      (2) Second, data need to be better presented (most important things first, least important - shorter or move to the Supplementary materials). Currently, it is too much to be clear and easy to follow.

      (3) Discussion needs to better state the novelty and the significance of these findings. What does the study offer that is new? Why was it important to perform? What big questions does it address?

      (4) The authors should better discuss the study limitations and future directions of research.

      (5) There should be a stronger conclusion with a take-home message to emphasize what new information the study brings and why it is important.

      (6) The overall style of the paper should be improved. Currently, it reads like a dry bulleted CRO report, not a usual scholarly paper.

      (7) Optimize the text flow. Currently, the overall flow of the discussion needs to be smoother - it now reads as a selection of bulleted paragraphs, with few connections between them.

    5. Author response:

      We would like to thank the editorial team and the reviewers for their thoughtful assessment of our manuscript. We are highly encouraged that the reviewers found our closed-loop OMR virtual reality assay to be a valuable new behavioural paradigm, and that our findings regarding the early behavioural phenotypes in adgrl3.1 mutant zebrafish provide solid, high-quality data of broad interest to the community. We also appreciate your constructive feedback regarding the structural flow of our manuscript, the need for tighter conceptual framing, and the request for more rigorous methodological explanation and improved presentation. In our upcoming revision, we plan to directly address the suggestions raised to ensure clarity of our work.

      For our revised manuscript, we will ensure to focus on fully contextualising our conceptual rationale and unique utility of our behavioural assay, ensuring a clear link to clinical relevance. This means we will provide further clarity on the rationale and relevance of the paradigm and interpretations of behaviours, while ensuring we are not overstating the absolute separation of anxiety-like states from hyperactivity and contextualising our findings within the broader literature. In addition, as suggested by the reviewers, we will improve the Abstract to explicitly articulate the unique advantages of the closed-loop OMR virtual reality paradigm over standard static assays and provide definitions. In addition, we will also revise our interpretations in the Discussion, such as around swim speed, exploration, and visual sensitivity, to ensure they are fully grounded and supported with clear flow.

      Equally important to our revision is to clarify genetic validation, methods and statistical rigour as recommended by the reviewers. Therefore, we will update the Methods to explicitly detail further husbandry details, such as the embryo medium used, breeding protocols, and the exact timings in hours post-fertilisation. To resolve any ambiguity surrounding our genetic model, we will provide further explanations and include details on genotype and sequencing protocols. Where needed, we will also update the statistical analysis and expand the Methods accordingly to detail all statistical reporting for clarity.

      Finally, as suggested, we will also improve the overall structure, flow, and accessibility. In the Results, we will ensure the core behavioural phenotypes take primary focus throughout the writing. Furthermore, we will also improve the Discussion to ensure it is a unified, cohesive narrative that smoothly integrates our main behavioural findings with genetic model validation, limitations, and future directions. Additionally, we will update all visual presentations, figure formatting, and labelling to ensure accessibility and improved readability.

      We are incredibly grateful to the reviewers for their insightful recommendations. We are confident that by integrating their revisions, we will substantially strengthen the clarity of the work and better demonstrate its impact.

    1. eLife Assessment

      This valuable cross-sectional longitudinal study leverages high-definition transcranial direct current stimulation to the left dorsolateral prefrontal cortex to examine its effect on procrastination behavior over an extended time span. The cross-sectional longitudinal study a testing of competing models provided solid evidence for how stimulating DLPFC impacts reveal-world procrastination behavior. Whether these results generalize to a larger population will be a significant future direction. This work will be of interest to those interested in cortical function, procrastination, and related states.

    2. Reviewer #1 (Public review):

      Summary:

      The authors report the results of a tDCS brain stimulation study (verum vs sham stimulation of left DLPFC; between-subjects) in 46 participants, using an intense stimulation protocol over 2 weeks, combined with an experience-sampling approach, plus follow-up measures after 6 months.

      Strengths:

      The authors are studying a relevant and interesting research question using an intriguing design, following participants quite intensely over time and even at a follow-up time point. The use of an experience-sampling approach is another strength of the work.

      Comments on revised version.

      With the last round of revisions, the authors have now addressed my concerns.

    3. Reviewer #4 (Public review):

      Summary:

      The current study tested the effects of repeated sessions of tDCS targeting the DLPFC on procrastination behavior. The main outcome is that anodal versus sham DLPFC tDCS reduces procrastination behavior on both a short-term and a long-term scale up to six months after the stimulation sessions.

      Strengths:

      The current study tests competing models of procrastination with state-of-the-art high-definition transcranial electric stimulation. The study assesses stimulation effects on procrastination on both a short-term and a long-term scale, suggesting that repeated stimulation of the prefrontal cortex reduces procrastination on a time scale of up to six months.

      Weaknesses:

      The manuscript has already been reviewed and revised before, and it seems that the quality of the manuscript has substantially improved as a result of this revision process. I agree with the other reviewers that one must be cautious with drawing conclusions regarding the cognitive mechanisms underlying this effect, as many different cognitive functions are implemented by the DLPFC.

      One aspect of the current results that puzzles me is the strength of the current stimulation effects. Meta-analyses suggest that tDCS shows only small-to-moderate effect sizes (with Cohen's d around 0.5). While the authors report no effect sizes for their statistical models, the small p values, in combination with the unusually small sample size of 18 participants per group, suggests that the effect size must be rather large. Can the authors provide an estimate of the effect size of their stimulation effects? If they are considerably larger than to be expected, could the authors give an explanation for why their stimulation setup is showing much stronger effects than comparable high-definition tDCS studies on cognition or decision making?

      Regarding the strengths of the stimulation effects, I moreover found remarkable that the post-test procrastination rate was 100% in all (!) participants in the DLPFC group (figure 3F). I admit that it is hard to trust results that have no individual variation at all. This means that all participants are perfect responders to tDCS, which is again at variance what one typically expects for tDCS (where one usually has many non-responders). Do the authors have an explanation for this?

      In any case, I am surprised by the rather small sample size. Due to the small effect sizes for tDCS, it is common to have a minimum of 30 subjects per group in between-subject designs. According to G*Power, a between-subject design with 17 subjects per group could detect only relatively large effect sizes of Cohen's d = 0.99 (alpha = 5%, power = 80%, independent-samples t-test). As explained above, this is far above the effect size that can be expected for tDCS. In addition, small samples bear the risk that results strongly depend on outliers in the data, which might explain the strong effect size observed in the current study. The small sample size should be discussed as a major limitation of the current study and that the results need to be replicated by studies with larger sample sizes. Moreover, to rule out that the results are driven by outlier in the data, the authors should show individual data points in all plots showing empirical data.

      Related to this, in the figure showing individual data points (3B/F), I count only around 10 data points per tDCS group for the 18 participants per group. I ask the authors to modify the plot that the data points from all participants can be seen (for example, by adding some noise on the x-axis for participants with the same value on the y axis).

      Another surprising aspect of the data is that repeated sessions of tDCS change procrastination behavior up to six months after stimulation. Do the authors think that their tDCS setup leads to such long-lasting neuroplastic changes, and if yes, can they cite prior work where similar dosages of tDCS also showed such long-lasting effects? Or could the results be explained by learning effects, for example because participants in the DLPFC group learned during the repeated tDCS sessions that it feels internally rewarding to finish one's tasks instead of procrastinating them, and they still benefit from this kind of "learned industriousness" 6 months later? In any case, in my view it is important to be more specific about how seven sessions of tDCS can affect behavior half a year later.

      Lastly, the link to the data repository works, but I could not inspect the data because I was asked to request access to the data, which I did not do in order to remain anonymous.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors report the results of a tDCS brain stimulation study (verum vs sham stimulation of left DLPFC; between-subjects) in 46 participants, using an intense stimulation protocol over 2 weeks, combined with an experience-sampling approach, plus follow-up measures after 6 months.

      Strengths:

      The authors are studying a relevant and interesting research question using an intriguing design, following participants quite intensely over time and even at a follow-up time point. The use of an experience-sampling approach is another strength of the work.

      Comments on revisions:

      Overall, I think the authors made many improvements to their manuscript. There are, however, still a number of concerns that first need to be addressed, since it is still not currently possible to fully evaluate the analyses, results, and conclusions presented in the paper. I list these points below:

      (1) The authors still use causal language where they must not use causal language. This is true for many places in the manuscript; I am highlighting here just a few places, but the authors nevertheless have to go carefully through the whole manuscript to change these instances.

      We sincerely thank the reviewer for this critical and well-taken point. We fully agree that our design (manipulating DLPFC excitability while measuring procrastination, task value, and aversiveness) does not directly measure or manipulate self-control, nor does it rule out alternative neurocognitive mechanisms. Accordingly, we have conducted a comprehensive, line-by-line revision of the entire manuscript to systematically replace causal claims with cautious, hypothesis-consistent language. In response, we have replaced all the wordings that may imply causal inferences, such as impair, cause, boost, by association-consistent phrasing. Furthermore, as you clearly raised below, we explicitly reframed self-control as a hypothesized theoretical construct rather than an empirically verified mediator throughout the whole re-revised manuscript. Please see specific revisions below:

      Abstract Section (Page 2, Line 53-59)

      “... a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement. In conclusion, these findings are consistent with the hypothesis that enhancing DLPFC function may reduce procrastination by selectively amplifying the valuation of future rewards, not by simply reducing negative feelings about the task.”

      Introduction Section (Page 3, Line 83-86)

      “... Even worse, chronic procrastination has been consistently associated with poor general health conditions, such as immune system disruption, gastrointestinal disturbance, hypertension and cardiovascular disease (Sirois, 2015; Sirois, 2016).”

      Introduction Section (Page 4, Line 143-145)

      “Consistent with this framework, the left dorsolateral prefrontal cortex (DLPFC)—a region frequently implicated in value-based decision-making and top-down regulation—has been associated with procrastination. ...”

      Introduction Section (Page 4, Line 135-139)

      “... Also, given the integrative nature of prefrontal regulatory functions, we hypothesize a third pathway whereby both decreased task aversiveness and increased task-outcome value may jointly contribute to reduced procrastination, potentially reflecting coordinated downstream effects on valuation and affective processing.”

      Introduction Section (Page 5, Line 183-186)

      “... Thus, this study aims to clarify the brain-behavior association between DLPFC neuromodulation and procrastination, and to test whether observed changes in task valuation and aversiveness are consistent with theoretical models of top-down regulation.”

      Results Section (Page 11, Line 534-536)

      “... Thus, these findings are consistent with the view that neuromodulation of the left DLPFC is associated with reduced task aversiveness and increased task-outcome value.”

      Results Section (Page 11, Line 579-584)

      “... In summary, these findings identified a statistical pathway consistent with the theoretical model: neuromodulation of the left DLPFC was associated with increased task-outcome value, which in turn was associated with reduced procrastination.”

      Discussion Section (Page 12, Line 617-620)

      “On balance, our findings provide evidence consistent with the hypothesis that neuromodulation of the left DLPFC is associated with reduced procrastination, primarily through increasing task-outcome value rather than merely reducing task aversiveness. ...”

      Discussion Section (Page 13, Line 664-669)

      “... Building on this foundation, among several theoretical interpretations and cognitive pathways, our study showed the one plausible neurocognitive mechanism of procrastination: the cortical excitability of the DLPFC produced by active neuromodulation may engage prefrontal regulatory networks to increase task outcome value, which in turn is associated with reduced procrastination behavior, statistically supporting the theoretical accounts of temporal decision model (TDM, Zhang et al., 2019).”

      Discussion Section (Page 14, Line 755-762)

      “... Moreover, this study did not collect data for assessing participants' self-control at either baseline or post-neuromodulation. Accordingly, we explicitly note that self-control was not directly measured or manipulated in this study; the observed associations between DLPFC neuromodulation, task-outcome value, and procrastination are consistent with theoretical models positing a role for top-down regulatory processes, but do not constitute direct evidence that self-control mechanisms were engaged. This limitation precludes definitive conclusions about the unique contribution of self-control-related pathways versus alternative neurocognitive mechanisms.”

      Some examples:

      (a) In response to my comment (1) in the previous round, where the authors adjusted their text, the authors still use causal language in their last sentence "... procrastination behavior has been observed to impair general health..." Unless the cited study truly allowed causal conclusions, the causal language should be removed here as well.

      Thank you for pointing out this inappropriate phrasing. As you kindly suggested, we have reworded it as “Even worse, chronic procrastination has been consistently associated with poor general health conditions, such as immune system disruption, gastrointestinal disturbance, hypertension and cardiovascular disease” (Introduction Section, Page 3, Line 83-86).

      (b) The authors still make (causal) claims about the involvement of self-control in their observed results. To reiterate from the previous round of revisions: The authors cannot make any strong claims about the role of self-control processes because they do not directly measure self-control nor do they directly manipulate self-control or have a design that would rule out alternative mechanisms other than self-control. Therefore, their claims about self-control have to be toned down. It is laudable that the authors have added a statement towards the end of their discussion about not being able to make strong conclusions about the role of self-control. But the authors need to use similar careful wording not just at the end of the discussion but throughout the manuscript.

      We appreciate you reiterating this concern. In the re-revised manuscript, we have thoroughly removed or rewritten all the statements implying causal inferences, and have substantially toned-down claims for the roles of self-control in procrastination reduction from this neuromodulation. Please see instances for what we have replied to the Comment #1.

      (i) In the abstract, the authors use the formulation "...conceptualized roles of self-control on procrastination..." -- this wording is still too strong, suggesting that you actually studied self-control.

      Thank you for providing this specific instance. This inappropriate sentence has been removed.

      (ii) In the introduction (page 4, lines162-169), the way the authors formulate these sentences suggests that they directly measured self-control. Again, the authors need to make it explicit that they are not directly measuring self-control but its hypothesized down-stream consequences on valuations/behavior.

      Many thanks. This statement has been removed, and we have reworded it as “... Thus, this study aims to clarify the brain-behavior association between DLPFC neuromodulation and procrastination, and to test whether observed changes in task valuation and aversiveness are consistent with theoretical models of top-down regulation.” (Introduction Section, Page 5, Line 183-186).

      (iii) In the discussion, for example, on page 11, lines 555 and following, the authors write: "One major contribution this study has made is to disentangle the neurocognitive mechanism of procrastination by demonstrating that self-control could increase task-outcome value so as to reduce procrastination."

      As you kindly instructed, we have rewritten this statement as “One contribution of this study is to provide empirical evidence partially consistent with the temporal decision model (TDM), showing that increased task-outcome value—rather than decreased task aversiveness—was statistically associated with reduced procrastination following DLPFC neuromodulation.” (Introduction Section, Page 12, Line 625-628), which no longer implies any conclusions for the role of self-control in this study.

      Again, please be aware that you are NOT demonstrating that self-control does anything, since you only measure procrastination rates, outcome values, and task aversiveness. It is possible that mechanisms other than self-control might be relevant for this. Perhaps neuromodulation directly increases outcome values, without involvement of self-control processes. You simply cannot know that and therefore you cannot make those claims in the form that you are making them. You can write that the observed results are consistent with the idea that neuromodulation might have had an effect on self-control and this in turn might have affected outcome values. But you also need to make it explicit that, to substantiate these claims, you would need more direct evidence that indeed self-control was involved. These more careful formulations would not at all reduce the value of your work, but indeed they would rather demonstrate your carefulness in interpreting the results you obtained.

      We sincerely thank the reviewer for this exceptionally clear and constructive guidance. We fully agree that our study design does not measure or manipulate self-control, and therefore we cannot demonstrate that self-control processes are causally involved in the observed effects. As you correctly note, it is entirely possible that neuromodulation directly modulates outcome valuation or engages alternative neurocognitive pathways (e.g., attentional allocation, feedback learning, or affective processing) without invoking self-control mechanisms.

      In direct response, as we replied above, we have completely rewritten the whole revised manuscript to remove any assertions that we “identified” a role of self-control. The revised text now explicitly states as follow: (1) our findings merely are consistent with the theoretical hypothesis that DLPFC neuromodulation might engage prefrontal self-regulatory functions, which in turn influence outcome valuation; (2) we explicitly acknowledge that substantiating this specific pathway would require more direct evidence. Rather than single sentence, we have applied this careful, hypothesis-consistent framing systematically across the Abstract, Introduction, Results, and Discussion. As you suggested, these revisions more accurately reflect the interpretative boundaries of our data and demonstrate our commitment to rigorous, transparent scientific reporting. Please see specific cases for this revision above.

      (2) I am still puzzled by the power analysis. In the text, you write that a sample size of 18 participants (i.e., 9 per group) would be sufficient to achieve 80% power. I still feel this seems far too optimistic and hard to believe, but that is not my point here. While in the text, you write that you need 18 participants, the G*power output seems to suggest a sample size of 34, not 18. Why this contradiction? Or is it not contradictory? If it is not, then please explain it more fully.

      We appreciate you pointing out this critical typo. In the last round of revision, we mean that 18 participants per group are required to achieve at least 80% statistical power, rather than a total sample size, as shown by the GPower software. We are sorry for this critical typo to confuse you. As you correctly pointed out, the GPower indicated that the minimum sample size to reach 80% power is 34 (i.e., 17 per group). Thus, we selected 36 (i.e., 18 per group) participants as minimum sample size in case of potential drop-out. We have thoroughly corrected this typo, and double-checked no such numeric issues:

      Methods Section (Page 5, Line 234-237)

      “... statistical power was predetermined by G*Power at a relatively medium effect size (1-β err prob = 0.80, f = 0.25), indicating the total sample size at 34 (17 per group) to reach acceptable power. To account for potential attrition, we determined to recruit 36 participants, at least.”.

      (3) I have several comments about the mixed-effects analysis.

      First of all, I want to thank the authors for adding more details, things have become much clearer now. However, I still have a few questions and comments related to these analyses:

      (a) The variable Emotions was within-subjects, as far as I understood. Accordingly, Emotions should most likely be modelled with random slopes varying over participants (in addition to being modelled as a fixed effect).

      We thank you raising this reasonable concern on the mixed-effect linear modeling. Yes, the Emotions reflect daily baseline affect, which is modeled as covariates of no interests to adjust for daily emotional fluctuation (if any). In this vein, this baseline emotion score is included for each participant across all the sessions. Therefore, it should be modeled with random slopes as you assumed indeed.

      In response, we have remodeled this mixed-effects analysis by including the daily baseline emotion as random slopes varying over participants. After centering the variables, we estimated the revised models as “Procrastination Rate ~ Group * Treatment day + Age + Gender + SES + Emotions + (1 + Treatment day + Emotions || SubjectID)” and “Task execution willingness ~ Group * Treatment day + Age + Gender + SES + Emotions + (1 + Treatment day + Emotions || SubjectID)”. Notably, as you correctly assumed, fitting this complicated random-effect structure is likely to result in convergence failure, given the limited sample size in the present study. Therefore, we hypothesized the independence among random effects for model simplification. Consistent with this assumption, model comparisons indicated that the simplified models fit better than original ones (Procrastination Rate model, ∆AIC = -1.0, ∆BIC = -11.8, LRT, χ<sup>2</sup>(3) = 5.04, p = .17; Task-execution willingness model, ∆AIC = -5.5, ∆BIC = -16.4, LRT, χ<sup>2</sup>(3) = 0.51, p = .91).

      Taken together, as you kindly suggested, we have rebuilt the mixed-effect models by adding daily baseline emotion as random slopes varying over participants, and have demonstrated the consistent findings with the original one:

      Methods Section (Page 8-9, Line 412-419)

      “... Given the risks of convergence failure with the two correlated random-effects structure (i.e., treatment days and self-reported emotions), we hypothesized that the random effects are independent, leading to model simplification. Consistent with this assumption, model comparisons favored the simplified independent structure over the full correlated structure for both outcomes. For the actual procrastination model, the simplified model showed lower AIC (∆ = -1.0) and BIC (∆ = -11.8), with a non-significant likelihood ratio test (χ<sup>2</sup> (3) = 5.04, p = .17). For the task-execution willingness model, the simplified model was also preferred (∆AIC = -5.5, ∆BIC = -16.4; LRT: χ<sup>2</sup> (3) = 0.51, p = .91).”

      Results Section (Page 9-10, Line 469-489)

      “For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001; Fig. 3A). In the post-hoc simple effect analysis, it demonstrated a significantly increased task-execution willingness (i.e., decreased procrastination willingness) after neuromodulation in the active neuromodulation group (NM-before: 35.65 ± 30.21, NM-after: 80.43 ± 19.92, Mean Diff = 41.79, SE = 7.58, DF = 103.4, t.ratio = 5.51, p < .0001, Tukey correction), but no such effects were identified in the sham control group (SC-before: 37.57 ± 26.46, SC-after: 47.35 ± 30.49, Mean Diff = 2.58, SE = 7.56, DF = 96.8, t.ratio = 0.34, p = .73, Tukey correction) (Fig. 3B-C). A linear uptrend for task-execution willingness was further observed across multiple sessions in the active NM group, indicating gradually increasing neuromodulation effects (Fig. 3D; p < .01, Mann-Kendall test). For actual procrastination behavior, changes to actual procrastination rates across all the sessions have been detailed in the Fig. 3E. Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group (NM-before: 56.74 ± 39.10, NM-after: 0.00 ± 0.00, Mean Diff = 44.40, SE = 9.36, DF = 110.0, t.ratio = 4.74, p < .0001, Tukey correction), but no such prominent changes found in the sham control group (SC-before: 46.47 ± 40.76, SC-after: 33.35 ± 37.82, Mean Diff = 7.53, SE = 9.28, DF = 102.0, t.ratio = 0.81, p = .42, Tukey correction) (Fig. 3F-G).”

      (b) The analyses still cannot fully be evaluated as I cannot access the scripts and data. The authors mention that the scripts and data should be available via a link they provide (https://doi.org/10.57760/sciencedb.35140). However, when I try to access these materials via this link, no page opens; it seems the link is dead?

      Thank you very much for bringing this case to us. We checked this link and found it to be still active.

      To ensure accessibility for your evaluation, we have uploaded scripts and data into this online submission system. Please do let us know if you are still unable to access them. We are glad to send them to you by other available pathways. This link is a private access to you, and the repository would be openly available for other users upon the final publication.

      (c) What are the results and conclusions if you do not include the covariates of no interest? I.e., please re-run your main models without age, gender, SES, Emotions.

      Thank you for raising this question. As you clearly instructed, we have rerun main models without all those covariates. As shown in the table below, the results for the key predictors of interest (Group, Treatment day, and their interaction) remained largely unchanged in terms of effect size, direction, and statistical significance:

      Author response table 1.

      Comparison to statistics derived from model with covariates (i.e., age, gender, SES, Emotions) and without covariates

      (d) The authors mention that they use GLMMs, which would suggest generalized mixed-effects models, but they do not describe what family/distribution they used. Since they mention lmerTest and seem to report F-tests, my guess is that they used Gaussian models. However, both their DVs (procrastination rates and their ratings) are bounded variables and at least procrastination rates hit the lower boundary. That can mean that their analyses suffer from inflated Type 1 and/or Type 2 rates. Therefore, please repeat the analyses with an appropriate generalized mixed-effects model (perhaps a beta regression type of model?).

      We are very grateful to you for raising this crucial statistical point. As you correctly pointed out, we used the Gaussian distribution in estimating this model. We are sorry to confuse you due to the absence of reporting family/distribution we used. In the original manuscript, we meant “general” linear mixed-effect model, rather than “generalized” one. As you clearly and correctly raised, procrastination rates and willingness are technically bounded, and that procrastination rates frequently reached the lower boundary (0%) in the present study, which are in high risks to be inflated for Type 1 and/or Type 2 error.

      Thus, as you kindly suggested, a beta family distribution with logit function is used to reanalyze those main effects of interest. Results are tabulated in Author response table 2.

      Author response table 2.

      These convergent results confirm that the critical main effect (i.e., Group and Treatment day) and their interaction remain statistically significant across distributional specifications, and that our primary conclusions are not artifacts of the Gaussian assumption. Taken them together, as you kindly suggested, we have repeated the analyses with beta regression family distribution, which replicated our main findings, potentially supporting their statistical robustness.

      Following your suggestion, we have added those results derived from such sensitivity analyses into the revised manuscript:

      Methods Section (Page 9, Line 430--436)

      “... To examine whether our findings were sensitive to the distributional assumptions of the dependent variables, we re-analyzed the main models using an alternative distributional specification. Given that both procrastination rates (ranging from 0% to 100%) and task-execution willingness (measured on a 0-100 visual analog scale) are bounded continuous outcomes, and that procrastination rates frequently reached the lower boundary (0%) in the present study, a Beta regression model with a logit link function was employed for a sensitivity analysis.”

      Results Section (Page 10, Line 506-512)

      “... Furthermore, as a sensitivity analysis, we reran the main LMMs using Beta regression distribution with a logit link function, which is appropriate for the both bounded outcomes mentioned above (i.e., procrastination rate and procrastination willingness). The main effects (i.e., Group and Treatment day) and their interaction remained significant for both procrastination rate and willingness (see SI Results and Tab. S5), confirming that our findings are robust to alternative distributional assumptions.”

      (e) When reporting the results of the mixed-effects models, the authors report the regression coefficient, standard error, DFs and p value, but not the actual test statistic. Please add the information about the test statistic and report all degrees of freedom (in case of F tests that would be the degrees of freedom of the test and the residual degrees of freedom).

      We truly thank you for this nuanced reminder. As you suggested, we have added actual test statistics, including t-values and all degrees of freedom (DF). Please see specific instances below:

      Results Section (Page 9-10, Line 469-489)

      “For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001; Fig. 3A). In the post-hoc simple effect analysis, it demonstrated a significantly increased task-execution willingness (i.e., decreased procrastination willingness) after neuromodulation in the active neuromodulation group (NM-before: 35.65 ± 30.21, NM-after: 80.43 ± 19.92, Mean Diff = 41.79, SE = 7.58, DF = 103.4, t.ratio = 5.51, p < .0001, Tukey correction), but no such effects were identified in the sham control group (SC-before: 37.57 ± 26.46, SC-after: 47.35 ± 30.49, Mean Diff = 2.58, SE = 7.56, DF = 96.8, t.ratio = 0.34, p = .73, Tukey correction) (Fig. 3B-C). A linear uptrend for task-execution willingness was further observed across multiple sessions in the active NM group, indicating gradually increasing neuromodulation effects (Fig. 3D; p < .01, Mann-Kendall test). For actual procrastination behavior, changes to actual procrastination rates across all the sessions have been detailed in the Fig. 3E. Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group (NM-before: 56.74 ± 39.10, NM-after: 0.00 ± 0.00, Mean Diff = 44.40, SE = 9.36, DF = 110.0, t.ratio = 4.74, p < .0001, Tukey correction), but no such prominent changes found in the sham control group (SC-before: 46.47 ± 40.76, SC-after: 33.35 ± 37.82, Mean Diff = 7.53, SE = 9.28, DF = 102.0, t.ratio = 0.81, p = .42, Tukey correction) (Fig. 3F-G).”

      (f) Thank you for adding the analysis where you remove the last two sessions. But currently you present them in the manuscript without explaining/motivating why you do this. Please add this motivation, as otherwise it will be puzzling for the reader why you conduct these analyses.

      Thank you for this very practical requirement to clarify the motivation of reanalyzing main models from removing the last two sessions. Please see specific explanation as follow:

      Results Section (Page, Line 499-506)

      “... To systematically test whether such effects are biased by extreme data points or patterns, we reran the main LMMs by iteratively removing data from the last two sessions, which showed extraordinarily high effectiveness from neuromodulation (e.g., all the participants in the neuromodulation group had no actual procrastination behavior in session #6 and #7). Results showed the significant group*neuromodulation sessions interaction effects across all those nested models (removing session #6, #7 or both, all p < .05; see SI Results and Tab. S3-4), potentially indicating a statistical robustness from the data pattern.”

      (4) Mediation analysis

      In your manuscript, you present some mediation analyses. Please be aware that such mediation analyses cannot establish causality and they suffer from extremely high Type 1 error rates (see, e.g., https://datacolada.org/103). My suggestion would be to completely remove all mediation analyses. However, if you want to keep them, then you need to be extremely careful in how you present the results. You need to explicitly mention that you cannot derive any causal conclusions from them and that simulation studies have shown that such mediation analyses suffer from extremely high Type 1 errors.

      We sincerely thank you for this exceptionally important methodological guidance. We fully agree that mediation analyses, especially those based on observational measures rather than experimentally manipulated mediators, cannot establish causal pathways and are susceptible to inflated Type 1 error rates, as rigorously demonstrated in recent simulation studies (https://datacolada.org/103).

      As you kindly suggested, please allow us to retain those mediation analyses upon explicitly highlighting that this mediation statistical model cannot generate any causal conclusions. In response, we have systematically replaced all instances of “causal mediation” with “statistical mediation” or “exploratory mediation analysis”, and removed causal verbs (e.g., “depends on”, "drives", "explains") in favor of association-consistent phrasing (e.g., “is statistically mediated”, “aligns with the hypothesis that”) throughout the abstract, introduction, methods, results and discussion sections. Furthermore, in the Discussion section, we explicitly reiterated the limitations of extending this mediation associations to causal conclusions. Please see specific modifications underneath:

      Abstract Section (Page 2, Line 52-56)

      “... While the intervention is significantly associated with both decreased task aversiveness and increased perceived task outcome value, a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement.”

      Methods Section (Page 9, Line 448-453)

      “... As these mediation analyses are based on observational measures rather than experimentally manipulated mediators, they do not establish causal pathways. Simulation studies have shown that such analyses can suffer from inflated Type 1 error rates. Results should therefore be interpreted as hypothesis-generating and statistically consistent with the proposed theoretical model, rather than as confirmatory evidence of causal mechanisms.”

      Methods Section (Page 9, Line 438-440)

      “... the Quasi-Bayesian mediation analysis was used to model the association between the effects of tDCS, task aversiveness/outcome and decreased procrastination.”

      Results Section (Page 11, Line 568-573)

      “As an exploratory analysis, results indicated that increased task outcome value was associated with changes in the task-execution willingness (δ = 21.73, p < .01; ζ = 11.25, p = .07, ρ = 32.99, p < .01, simulation = 1,000; see Fig. 5A) and real-world procrastination (δ = 30.75, p < .01; ζ = 3.05, p = .52, ρ = 33.81, p < .01, simulation = 1,000; see Fig. 5B), in the context of ms-tDCS neuromodulation. ...”

      Results Section (Page 11, Line 582-584)

      “... Nevertheless, all mediation findings are now explicitly labeled as “exploratory” and framed as quantifying statistical associations consistent with the TDM pathway, not causal mediation.”

      Discussion Section (Page 14, Line 747-755)

      “Notably, we explicitly acknowledge that exploratory Quasi-Bayesian mediation analyses, based on observational measures rather than experimentally manipulated mediators, cannot establish causal pathways and are susceptible to inflated Type 1 error rates as demonstrated in recent simulation studies. These findings should be interpreted strictly as hypothesis-generating and statistically consistent with the proposed theoretical model, rather than as confirmatory evidence of causal mechanisms. Substantiating the precise neurocognitive pathway will require future studies employing stronger causal designs, such as experimental manipulation of task valuation or longitudinal cross-lagged modeling. ...”

      As an example (but the mediation results are mentioned in several places, for example, also in the abstract): On page 10, lines 501-503: What you can causally conclude is that neuromodulation affects your measured variables (outcome values, procrastination rates, task aversiveness), but you cannot conclude that the effect of neuromodulation on procrastination rates causally operates via outcome values. Thus, please adjust the formulation accordingly. The same applies to the mediation section that follows right afterwards (page 10, lines 505-522).

      Thank you for offering those specific instances. As we replied above, they have been revised accordingly:

      Results Section (Page 11, Line 557-559)

      “... Collectively, these findings provide statistical evidence consistent with the hypothesis that the outcome-value pathway may contribute to procrastination reduction.”

      Results Section (Page 11, Line 563-582)

      “Increased task outcome value is specifically associated with reduced procrastination in the context of neuromodulation

      To explore the potential neurocognitive pathways of procrastination, the Quasi-Bayesian mediation analysis was undertaken, with increased task outcome value as a statistically mediated variable. As an exploratory analysis, results indicated that increased task outcome value was associated with changes in the task-execution willingness (δ = 21.73, p < .01; ζ = 11.25, p = .07, ρ = 32.99, p < .01, simulation = 1,000; see Fig. 5A) and real-world procrastination (δ = 30.75, p < .01; ζ = 3.05, p = .52, ρ = 33.81, p < .01, simulation = 1,000; see Fig. 5B), in the context of ms-tDCS neuromodulation. To ensure the statistical robustness and specificity of these findings, the sensitivity analysis was implemented by changing sampling parameters and outcome variables. By doing so, those findings were validated statistically robust, as shown by replicated observations across bootstrapping sampling subsets (see SI Results and Tab. S6-7). Moreover, the results of the control analysis further validated the specificity of these findings by showing a null statistically mediated effect of this model to predict one’s task aversiveness (see SI Results and Tab. S8). In summary, these findings identified a statistical pathway consistent with the theoretical model: neuromodulation of the left DLPFC was associated with increased task-outcome value, which in turn was associated with reduced procrastination.”

      (5) In the introduction, the authors introduce several theoretical procrastination frameworks (TMT, mood repair, TDM). Do the results of the current paper help to decide which framework might be the most appropriate, at least for the authors data set? It might be of interest to address this explicitly.

      We do thank the reviewer for this insightful theoretical question. We agree that explicitly positioning our findings within the broader theoretical landscape strengthens the conceptual contribution of our work. Upon careful consideration, we believe that our results provide the strongest empirical support for the TDM over alternative frameworks (TMT, mood repair). TDM uniquely posits procrastination as contingent on the dynamic trade-off between task aversiveness and task-outcome value. In the present study, neuromodulation was identified to be associated with both pathways but only increased outcome value statistically predicted reduced procrastination. This aligns precisely with TDM’s hypothesis that value-based processes may dominate aversiveness-avoidance processes in driving behavioral change. Neither TMT (which emphasizes temporal discounting of utility per se) nor the mood repair perspective (which prioritizes short-term affect regulation) explicitly predicts this dissociable pattern. As you kindly suggested, we have extended the Discussion section to explicitly contextualize our findings into this theoretical landscape (Discussion Section, Page 13, Line 669-687).

      (6) The language is sometimes hard to understand and seems in quite some places grammatically incorrect. Thus, I think the paper would profit very much from thorough English proofreading.

      We sincerely thank you for this practical suggestion. We fully agree that the original manuscript contained grammatical inaccuracies and awkward phrasing that could hinder readability. In response, we have engaged a professional academic editing service to thoroughly proofread and polish the entire manuscript. All sentences have been revised for grammatical correctness, syntactic clarity, and academic tone, while strictly preserving the original scientific meaning and technical terminology. We believe this language improvements have substantially enhanced the readability and overall quality of the paper.

      Reviewer #2 (Public review):

      Summary:

      Chen and colleagues conducted a cross-sectional longitudinal study, administering high-definition transcranial direct stimulation (HD-tDCS) targeting the left DLPFC to examine the effect of HD-tDCS on real-world procrastination behavior. They find that seven sessions of active neuromodulation to the left DLPFC elicited greater modulation of procrastination measures (e.g., task-execution willingness, procrastination rates, task aversiveness, outcome value) relative to sham. They show that HD-tDCS reduces task aversiveness and increases task-execution willingness on real-world tasks as quantified by intensive experience sampling methods, providing causal evidence for the role of DLPFC in modulating contextual features to delaying or completing one's goals.

      Strengths:

      • This is a well-designed protocol with rigorous administration of high-definition transcranial direct current stimulation across multiple sessions. The intensive experience sampling approach which probes and assesses self-relevant task goals is innovative and aims to address an important question regarding the specific role of DLPFC in modulating specific features of chronic procrastination behavior (e.g., task-execution willingness, task aversiveness).

      • The quantification of task aversiveness through AUC metrics is a clever approach to account for the temporal dynamics of task aversiveness, which is notoriously difficult to quantify.

      Weaknesses:

      • While the findings that neurostimulation reduces procrastination behavior is compelling, there remain several alternative interpretations for these effects. For example, it could be that the task-execution willingness isn't increased per se, but rather that the goal completion becomes more valuable as participants learn from feedback or become more aware of their successful attainment of or failure to complete task goals. It is unclear whether the effects could be driven by improved working memory or attention to the reported tasks (and this limitation is addressed by the authors). In short, it is also difficult to examine the temporal dynamics of how these goals are selected across time.

      We sincerely thank you for raising these thoughtful and methodologically important points. We fully agree that the observed reductions in procrastination could reflect multiple neurocognitive pathways beyond the value-based mechanism emphasized in our primary analysis.

      In response, we have thoroughly removed claims on the “unique mechanism” of value-based pathways to procrastination reduction, and fully substituted language implying exclusive mediation by “value amplification” with more cautious phrasing (e.g., “statistically consistent with a value-based pathway”; “one plausible mechanism among several processes”). In the revised manuscript, we reiterated that the pattern of results, showing increased outcome value predicting reduced procrastination while decreased aversiveness did not, aligned with the Temporal Decision Model, yet does not rule out concurrent contributions from attention, learning, or executive processes. To clearly bring this interpretative boundary of our primary findings for audiences, we have explicitly warranted such cautions in the Discussion Section. Please see specific modifications underneath:

      Discussion Section (Page 13, Line 664-669)

      “... Building on this foundation, among several theoretical interpretations and cognitive pathways, our study showed the one plausible neurocognitive mechanism of procrastination: the cortical excitability of the DLPFC produced by active neuromodulation may engage prefrontal regulatory networks to increase task outcome value, which in turn is associated with reduced procrastination behavior, statistically supporting the theoretical accounts of temporal decision model (TDM, Zhang et al., 2019). ”

      Discussion Section (Page 13, Line 676-687)

      “... Despite statistically supporting the TDM, we acknowledge that alternative neurocognitive mechanisms could contribute to the observed reductions in procrastination. For instance, repeated exposure to the experience-sampling protocol may have enhanced participants’ awareness of task progress or facilitated feedback-based learning, thereby increasing the subjective value of goal completion independent of DLPFC neuromodulation. Similarly, improvements in working memory for task maintenance, attentional allocation to reported goals, or strategic shifts in goal selection across sessions could plausibly mediate the intervention effects. While our double-blind, sham-controlled design and inclusion of daily emotional covariates help mitigate some non-specific confounds, the present study did not incorporate direct measures of these alternative processes. Consequently, we cannot definitively isolate the value-based pathway posited by the TDM from concurrent contributions of attention, learning, or executive functions.”

      • It is unclear whether the current evidence support long-retention of this neurostimulation intervention. The study includes one 6-month timepoint after the study to examine the long-term retention of the neural stimulation effect. Future studies that evaluate the long-term effects across multiple time points would strengthen the evidence for the robustness of this intervention.

      We genuinely appreciate you for this insightful and methodologically reasonable point. We fully agree that a single 6-month follow-up assessment, while valuable, provides only preliminary evidence for long-term retention, and that multiple follow-up timepoints would substantially strengthen claims about the durability of neuromodulation effects. To carefully address this point, we have rephrased the “long-term retention” as “long-term after-effects” throughout the whole revised manuscript, and overall toned down the claims on the retention effects. Moreover, this limitation has been explicitly elaborated in the Discussion section:

      Abstract Section (Page 2, Line 49-50)

      “... we assessed the effect of anodal HD-tDCS on real-world procrastination behavior at offline after-effect (2-day interval) and long-term after-effect (6-month follow-up).”

      Results Section (Page 12, Line 606-608)

      “... Therefore, beyond short-term effects, the benefits of ms-tDCS neuromodulation on reducing procrastination were still detectable at a 6-month follow-up, providing preliminary evidence consistent with long-term after-effects.”

      Discussion Section (Page 14, Line 721-725)

      “... Thus, the detectable effects at 6 months are consistent with the hypothesis that repeated neuromodulation may induce neuroplastic changes in the DLPFC that support sustained behavioral change. However, we explicitly note that a single follow-up timepoint cannot establish the stability or trajectory of these effects; future studies with multiple longitudinal assessments are required to substantiate claims about long-term retention.”

      Discussion Section (Page 15, Line 771-775)

      “... Finally, while our 6-month follow-up provides preliminary evidence for sustained effects, the use of a single follow-up timepoint limits our ability to characterize the temporal trajectory of retention. Future studies incorporating multiple follow-up assessments (e.g., 1-month, 3-month, 6-month, 12-month) would strengthen evidence for the robustness and durability of this intervention.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please see my detailed comments above (6 points; several of them with subpoints a, b, c, etc).

      Thank you for listing those specific and helpful recommendations above. Please see our detailed response posed above, point-by-point.

    1. eLife Assessment

      This study presents a valuable human organoid platform for investigating neuron-immune interactions in Alzheimer's disease and modeling the interplay between innate and adaptive immunity in the context of amyloid pathology. The system has the potential to advance the field by enabling the study of human-specific neuroimmune interactions that are difficult to recapitulate in rodent models; however, the evidence supporting several central conclusions remains incomplete. Key claims, including the causal role of microglia-derived chemokines in T cell recruitment and the existence of a neuroinflammatory feedback loop, rely primarily on correlative observations and lack the mechanistic experiments necessary to establish causality; in addition, a major confounding factor, the non-autologous nature of the cellular components, is not adequately addressed. Despite these limitations, the study will be of considerable interest to researchers in neuroimmunology and Alzheimer's disease.

    2. Reviewer #1 (Public review):

      Summary:

      A growing body of evidence indicates that Alzheimer's disease is not simply a disease of neurons accumulating toxic protein aggregates, but one in which the immune system, both its resident brain component and its circulating peripheral arm, plays an active and sustained role. Understanding how these two immune compartments interact with one another and with diseased neural tissue has been hampered by the fact that the mouse immune system differs fundamentally from the human one in ways likely to matter for disease progression. The authors set out to address this gap by building a modular laboratory model that brings together three human cell types in a three-dimensional setting: brain organoids derived from human stem cells to provide a neural substrate, stem cell-derived brain immune cells (microglia) to represent the resident immune compartment, and circulating immune cells (CD8-positive T cells) harvested from human blood to represent the peripheral adaptive immune response. By exposing this tri-cellular system to a toxic form of amyloid protein, the hallmark aggregating molecule of Alzheimer's disease, the authors aimed to dissect, step by step, how microglia respond to amyloid stress, what inflammatory signals they release as a consequence, and whether those signals are sufficient to attract T cells into the neural environment. They further aimed to test whether blocking the molecular receptors that guide T cell movement could interrupt this process, with the broader goal of positioning the platform as a tool for human-relevant drug screening.

      Strengths

      The conceptual architecture of the platform is one of its clearest strengths. The decision to add immune components in a stepwise, modular fashion, first characterising the neural response to amyloid, then adding microglia, then adding T cells, makes it possible to attribute observed changes to specific cellular contributions in a way that a more complex all-at-once model would not allow. This staged design is well thought-through, and its logic is clearly communicated. The combination of single-cell transcriptional profiling, calcium imaging for real-time functional readouts, transwell migration assays, and protein secretion measurements gives the study a genuinely multi-modal character that goes beyond what purely transcriptomic or purely imaging-based approaches can offer. The observation that T cells failed to migrate toward amyloid-treated organoids in the absence of microglia is a clean and conceptually important result, clearly supporting the idea that the resident immune response acts as an intermediary between amyloid pathology and the recruitment of peripheral immune cells. The identification of specific chemokine receptor pathways mediating T cell movement and the demonstration that pharmacological blockade of those receptors reduces migration and provide a degree of mechanistic resolution useful for thinking about future therapeutic strategies.

      Weaknesses

      Despite these strengths, several aspects of the work as presented substantially limit the confidence one can place in its conclusions.

      The most consequential issue concerns the origin of the cells used in the model. The three cellular components: the brain organoids, the microglia, and the T cells are derived from genetically unrelated individuals. The T cells, in particular, come from healthy blood donors unrelated to the stem cell lines used to generate the neural tissue. This means the immune cells and the tissue they are interacting with carry different molecular identity markers (the proteins that the immune system uses to distinguish self from non-self). In this setting, any T cell activation or directed movement could reflect a generic rejection-like response to foreign tissue rather than a disease-relevant, chemokine-directed recruitment process. This is not a subtle concern: it represents a fundamental ambiguity at the heart of the model's central finding, and it is not acknowledged anywhere in the manuscript. For the transwell migration data to be interpretable as a model of Alzheimer's disease rather than of immune incompatibility, the authors would need to demonstrate that migration is driven by the specific chemokine environment and not by the genetic mismatch between cells, for example, using cells from the same donor or from matched donors, or by showing that blocking identity-marker recognition does not alter migration.

      A related concern is that the T cells used are from healthy individuals, whereas T cells from people with Alzheimer's disease are known to differ in their activation state, surface receptor expression, and functional behaviour. The platform cannot yet claim to model the specific T cell biology of Alzheimer's disease until disease-relevant T cells are incorporated.

      Beyond this foundational issue, the study frequently describes findings in causal terms that the experimental design does not support. The resident immune cells are said to "drive" T cell recruitment and "establish" a feedback loop. These are strong mechanistic claims. The evidence presented indicates that when microglia are present, more T cells migrate, and that blocking T cells receptors reduces migration. What is missing is direct evidence that the specific molecules measured, particularly the chemokines CCL4 and CCL5, are the agents responsible, as opposed to other signals also present in the conditioned environment. No experiment directly neutralises these chemokines to test whether their removal is sufficient to abolish T cell recruitment. Without such an experiment, the receptor-blocking data show only that the receptors matter, not that the measured ligands are the ones activating those receptors.

      The abstract describes one particular molecule, CXCL10, as a contributor to T cell recruitment, but the data in the paper itself show no significant change in CXCL10 levels between conditions. This discrepancy between the abstract and the results is misleading to readers who may not read the figures in detail.

      The single-cell sequencing data, which form the basis for claims about changes in cell populations following amyloid treatment or microglia addition, are presented without validation of the cell type labels against established reference datasets from human brain tissue. The proportional shifts in cell populations between conditions (Figures 1H and 3E) are described as significant findings but are shown without any statistical test appropriate for this type of compositional data. Comparisons of cell-type proportions derived from single-cell sequencing require specialised statistical approaches that account for the interdependence of proportions and the variability between samples; standard tests are not appropriate here, and none are applied.

      There is also an unresolved inconsistency in the age at which the organoids were analysed by single-cell sequencing: the text states day 90, while the figure legend states day 60, and the methods section contains a passage describing experimental conditions (including a cholesterol treatment and a drug called semaglutide) that are entirely unrelated to this study and appear to have been copied from a different manuscript. These issues raise concerns about the rigour of the manuscript preparation and should be corrected.

      Finally, the sample sizes underpinning several key conclusions are small (typically three to four organoids per group), particularly for the protein-secretion measurements used to identify the inflammatory signals responsible for T cell recruitment. While organoid studies are inherently limited in scale, the strength of the mechanistic claims made here would benefit from larger sample size or independent experimental replication.

      Conclusion:

      The authors have built a platform that is conceptually well-conceived and generates data consistent with a role for microglia in bridging amyloid pathology and T cell recruitment. In that sense, they have made meaningful progress toward their stated aims. However, the platform, as described, cannot yet deliver the human-specific mechanistic insight it claims to provide, primarily because the non-autologous configuration of the model introduces an uncontrolled variable that confounds the interpretation of the immune interaction data. The claim to have provided "the first human-specific mechanistic demonstration" of microglial activation as a bridge between amyloid pathology and adaptive immune recruitment is not supported by the evidence presented. The data are consistent with this interpretation but do not establish it.

      The general approach, building increasingly complex human neural-immune models by adding components in a controlled, stepwise manner, is a valuable direction for the field and one that other groups working on neuroinflammation will find useful to consider. The combination of live calcium imaging and transcriptional profiling in the same experimental system is a practical contribution that demonstrates the kind of multi-modal readout this class of model can support. If the autologous confound is resolved in future iterations and if the mechanistic claims are grounded in more direct experimental evidence, this type of platform could become a genuinely useful tool for investigating human neuroimmune biology and for screening candidate therapeutic compounds in a human-relevant context. As currently presented, however, readers and researchers considering adopting this approach should be aware that the immune interaction data may reflect genetic mismatches between cell sources rather than disease-specific biology, and that the causal conclusions drawn from the chemokine and migration data go beyond what the experiments can support.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors developed a human forebrain organoid model incorporating both iPSC-derived microglia and CD8⁺ T cells, enabling them to recreate and investigate multicellular aspects of AD pathology in a human-relevant system.<br /> Their findings show that microglia help clear amyloid-β deposits, but they also promote inflammatory responses. Activated microglia recruit CD8⁺ T cells by releasing the chemokines CCL4, CCL5, and CXCL10, which signal through the receptors CCR1/CCR5 and CXCR3. Pharmacological inhibition of CCR5 or CXCR3 prevents T-cell recruitment and alters autophagy pathways in a microglia-dependent manner.

      Strengths:

      The study presents a versatile human organoid platform for investigating neuron-immune interactions in Alzheimer's disease. It highlights the critical role of microglia-driven recruitment of CD8⁺ T cells in sustaining neuroinflammation and identifies CCR5 and CXCR3 signaling pathways as promising therapeutic targets for neuroinflammatory conditions.

      This study is interesting and presents novel findings supported by state-of-the-art approaches, including single-cell RNA sequencing, a three-dimensional cerebral organoid model, and co-culture systems involving two distinct immune cell populations.

      Weaknesses:

      Several aspects of the study require clarification and further improvement. For example:

      (1) Figure 1H is missing statistical analyses.

      (2) The scRNA-seq analysis shows a reduction in the proportion of cells occupying transcriptional states associated with later pseudotime values, which the authors interpret as evidence that Aβ treatment inhibits neuronal maturation. However, the data presented do not appear sufficient to support this conclusion. An alternative explanation is that Aβ preferentially affects the survival of more mature neuronal populations, leading to their depletion, consequently, an apparent enrichment of cells at earlier pseudotime states. Therefore, the observed pseudotime shift does not necessarily demonstrate impaired maturation per se. The authors should revise the interpretation of these results in the first paragraph and either provide additional evidence supporting a maturation defect or discuss alternative explanations such as selective loss of mature neurons.

      (3) A similar concern applies to the scRNA-seq data presented in Figure 3. The authors interpret the shift toward later pseudotime states in the presence of microglia as evidence of enhanced neuronal maturation. However, the data do not exclude alternative explanations. For instance, microglia may preferentially promote the survival of more mature neuronal populations or protect them from cell death, thereby increasing their relative abundance in the dataset. Consequently, the observed pseudotime distribution cannot be taken as direct evidence of enhanced maturation. The authors should revise their interpretation accordingly and discuss the possibility that the observed effect reflects differential survival rather than accelerated neuronal maturation.

      (4) In Figures 4A-E, the authors should report the levels of the secreted proteins in pg/mL instead of relative values, as this would better reflect the actual amounts produced. In Figure 4H, the inhibitor-treated control T-cell samples should be included. Furthermore, it should be explicitly stated that the inhibitor-treated data points currently shown refer to T cells cultured in the presence of myeloid Aβ.

    1. eLife Assessment

      This valuable study presents a comparative analysis of the transcriptomic features underlying C. elegans longevity, providing insights into how different changes in gene expression can promote longevity. The authors present solid evidence with analysis and selected functional validation showing that some long-lived animals share common changes while others appear to use opposing strategies. The datasets and analyses contained within and the user-friendly website developed will be of interest to researchers interested in complicated transcriptomic analyses and/or the biology of aging.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      This manuscript by Rudich ZD et al. systematically profiled the transcriptomic changes in nine long-lived C. elegans mutants and presented a careful and informative comparative analysis of these aging-related changes. In addition to these valuable datasets and bioinformatics analyses, the authors performed a large-scale RNAi screen to assess the role of the differentially expressed genes (DEGs) in these mutants and identify several potential targets to promote healthy aging. Moreover, the authors have provided a user-friendly website to examine genes of interest in those longevity mutants from their datasets.

      Strengths:

      Compared to previous transcriptomic analyses of these mutants in different reports, this study minimized the technical variations and benefitted from the advances in RNA-Seq technology and bioinformatics tools. Therefore, it should provide a more consistent and comprehensive view of the molecular mechanisms underlying the longevity of these mutants. The datasets in this manuscript are valuable to other researchers in the biology of aging.

      Weaknesses:

      Meanwhile, since these mutants have been extensively studied, the advance of this study in unknown ageing mechanisms remains limited.

      Comments on revised version.

      The authors addressed the concerns successfully.

    3. Reviewer #2 (Public review):

      Summary:

      In the manuscript titled "Multiple Molecular Pathways to Longevity: Opposing Gene Expression Programs Define Distinct Aging Strategies", the authors investigated diverse genetic pathways that contribute to lifespan extension in Caenorhabditis elegans and aimed to identify shared and distinct molecular mechanisms among various longevity mutants. Through comprehensive RNA sequencing of different longevity mutants representing seven distinct pathways, the authors showed that these mutants cluster into three primary groups based on their gene expression profiles. This transcriptomic analysis revealed that while some longevity genes are commonly regulated across multiple pathways, others exhibit opposing expression patterns, suggesting that distinct molecular strategies can lead to increased lifespan. Specifically, they identified a set of 196 genes that are consistently upregulated in most longevity mutants, many of which are involved in innate immunity and stress defense. By performing RNAi-based screening, the authors further validated the functional roles of several candidates, including C08F11.7, ugt-62, and K05C4.9, supporting their contributions to longevity and stress resistance. The authors conclude that longevity is mediated through multiple molecular pathways and provide a public online tool to study these complex transcriptomic landscapes.

      Significance:

      This study provides a systematic, side-by-side transcriptomic comparison of nine genetically distinct long-lived C. elegans mutants, revealing that lifespan extension arises from both shared and opposing gene expression programs. By identifying three distinct longevity groups and demonstrating that key pathways can be modulated in opposite directions to achieve long life, the work challenges the notion of a single universal transcriptional signature of aging. Importantly, functional validation shows that select commonly regulated genes can directly modulate lifespan and stress resistance, highlighting actionable molecular targets for promoting healthy aging.

      Comments on revised version:

      The authors addressed my concerns successfully.

    4. Author response:

      The following is the authors’ response to the original reviews

      Reviewer #1 (Public review):

      In the revised manuscript, the authors have addressed most of my concerns. In the text of this manuscript, the authors should still include more discussion on why osm-5 and daf-2 are categorized into two different groups. 

      According to this suggestion, we have expanded our discussion to discuss why osm-5 and daf-2 worms fall into different longevity groups despite the fact that disruption of DAF-16 decreases both of the their lifespans. Please see lines 363-376.

    1. eLife Assessment

      This important study uses an elegant visual-anagram approach to test whether perceived animacy shapes visual working memory and guides visual attention while tightly controlling for lower- and mid-level image properties. The evidence is convincing and provides a rigorous demonstration that perceived animacy influences visual cognition beyond the contribution of its typical visual correlates. The findings will be of broad interest to researchers studying high-level vision, attention, and working memory.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      This is a very cool paper that casts light on a persistent problem in the psychology and philosophy of visual representation: is there high-level perception? Every vision scientist agrees that low-level features such as shape, color, texture, motion and spatial frequency are represented in visual perception, but there is a great deal of controversy about the representation of high-level properties such as causation, faces, agency and animacy. Animacy is especially problematic because there are large differences in line curvature between stimuli that represent animate and inanimate items.

      This article uses a novel approach-visual "anagrams" that are exactly the same image, except one is rotated 90 degrees relative to the other. They found persistent differences in visual processing between animate and inanimate stimuli. (Of course, the stimuli aren't animate-they represent animate items.). For example, there were processing differences between changes between animate and inanimate items (rabbit to boot) that were not present in rabbit to dog. They also showed such differences in two kinds of visual search tasks.

      Of course, there are feature differences that exploit orientation. A classic example is the difference between a square and a diamond that is produced from the square by rotating it 45 degrees.

      They addressed an aspect of this challenge having to do with some features using silhouettes. There was no search advantage for silhouetted stimuli.

    3. Reviewer #2 (Public review):

      Summary:

      The authors present a creative approach using visual anagrams matched on low-level image statistics to isolate animacy from low-level visual features and report consistent effects of animacy on visual working memory and attention.

      Strengths:

      (1) An important methodological advance in controlling low-level confounds that have historically complicated the study of animacy.

      (2) The converging effects across multiple experiments, together with the pre-registered design, strengthen the reliability of the reported findings.

    4. Reviewer #3 (Public review):

      This study makes clever use of generative AI to create stimuli that are pixel-for-pixel identical but which have radically different meanings depending on their orientation, to investigate the perception of animacy while retaining control over low-level image features (so-called 'anagram' stimuli).

      The authors present seven elegantly designed experiments in a commendably compact format.

      Experiments 1 and 2 involved a working memory paradigm in which participants had to spot which of five objects in an array changed after a pause. Importantly, the changed object was an anagram stimulus that in one orientation matched the animacy/inanimacy of the changed object, and in the other orientation was the opposite (e.g., a rabbit is replaced by either a dog or a boot, where the dog and boot stimuli are actually identical, just rotated by 90 degrees). They found a difference in accuracy depending on whether the animacy of the objects matched.

      Experiments 3 and 4 used a visual search task in which the participants had to localize the target, and the distractors were anagrams that either matched the target in terms of animacy or did not. There was a significant cost in terms of response time when the animacy of the target was the same as that of the distractors. Experiments 5 and 6 also used a similar visual search design, except that the task was to determine if the target was present or absent from the display, and the distractors again either matched or differed from the target in terms of animacy. Again, the authors found slower responses when the distractor arrays matched the animacy of the target than when they differed.

      An obvious potential concern about the studies is addressed by Experiment 7. It is unclear if the observed effects are related to the specific orientations of the target and distractor stimuli selected in each condition. For example, it could be that all the animate versions of the anagrams involved tall and skinny shapes, while all the inanimate versions involved wide and short objects, due to the 90-degree rotational difference between the two versions of the stimuli. To control for this, the authors repeated the visual search experiment but with convex-hull silhouettes of each of the stimuli. In other words, all targets and distractors from each trial were replaced by a black splotch with approximately the same overall outline (envelope) as the corresponding stimulus. Importantly, in contrast to the anagram stimuli, the silhouettes had had no meaningful semantic interpretation, and their animacy did not change depending on their orientation.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study uses an elegant visual-anagram approach to test whether perceived animacy structures visual working memory and attention while controlling for many low-level image properties. The evidence is solid, with converging results across seven preregistered experiments, but the central claim that animacy itself is represented independently of visual features should be tempered, as residual mid-level configural cues, ensemble or category structure, and broader semantic differences may also contribute to the effects. The work will be of interest to researchers studying high-level visual representation, attention, and working memory.

      We thank the Editors and Reviewers for this careful and informed assessment. We appreciate that every Reviewer found our approach to be elegant, our findings to be solid, and our question to be of broad interest. We respond to each Reviewer’s specific comments in more detail below; but we thought to summarize some of the highlights - especially the specific comments that come up in this Assessment - here.

      (1) The Reviewers make the insightful point that, even if our stimuli effectively control for many low-level features, there may be other high-level features that explain performance in our experiments (R2: “Although the anagram paradigm effectively controls low-level visual features […] these stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories.”). We are happy to embrace this possibility. If our results were explained by high-level visual representation of the natural vs. manmade distinction, rather than the animate vs. inanimate distinction, this would still be an appeal to a (not altogether unrelated) high-level property being represented independently from its lower-level features, which was the primary motivation for our study. We framed our work specifically around animacy given the persistent debates regarding perceived animacy, as well as the fact that our stimuli do quite saliently vary along that dimension; but we are certainly open to other nearby high-level categories being at play. We also think this is an empirical question that could be tested in future work. For example, objects like rocks and lakes are natural but inanimate. If they behave more like dogs than like boots in our paradigms, then Reviewer #2 may be right that naturalness was the relevant property all along; but if they behave more like boots than like dogs, then perhaps it really was animacy doing the work. We now discuss this explicitly in our paper, and we appreciate the opportunity to not only clarify our claims but also spur discussion for future work.

      (2) Multiple Reviewers raise the question of whether semantic factors that go beyond the images themselves may be driving our effects. Reviewer #3 raises a particularly interesting question along these lines: “if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results?” We have now taken this question quite literally and run this experiment exactly as described. Of course, much research already explores cognitive processing of animate/inanimate words, finding (for example) stronger memory for animate objects than inanimate ones (e.g., Nairne et al., 2013; Nairne et al., 2017). However, such tasks do not invoke effects of visual processing, whereas the question at issue here is specifically whether the visual system prioritizes animacy independent of its lower-level features. To this end, we conducted a new, pre-registered experiment (now Experiment 8) where participants search for animate/inanimate words on some trials, and animate/inanimate pictures on others. Given the nature of visual search tasks, we should expect to find no search advantage for words (as their meanings are not processed in vision per se) — and we should also expect to replicate (once again) our search advantage for pictures. This is exactly what we found. In other words, linguistic stimuli alone failed to produce the effect, while anagrams did produce the effect. We believe this rules out the strongest form of the semantic labeling account.

      (3) Finally, Reviewers #2 and #4 raise some concerns regarding residual mid-level features such as configural shape and ensemble statistics, which lie somewhere between animacy itself and more basic properties like contrast or spatial frequency. In our paper, we now clarify each of these concerns in greater detail. In short: We think that our stimuli and experiments indeed control for these residual cues. For example, rotating an image preserves its configural shape; and, as we argue below, the specific ensemble statistics argument fails to get off the ground without appeal to animacy itself.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Evidence for visual representation of animacy.

      Strengths:

      This is a very cool paper that casts light on a persistent problem in the psychology and philosophy of visual representation: is there high-level perception? Every vision scientist agrees that low-level features such as shape, color, texture, motion and spatial frequency are represented in visual perception, but there is a great deal of controversy about the representation of high-level properties such as causation, faces, agency and animacy. Animacy is especially problematic because there are large differences in line curvature between stimuli that represent animate and inanimate items.

      This article uses a novel approach-visual "anagrams" that are exactly the same image, except one is rotated 90 degrees relative to the other. They found persistent differences in visual processing between animate and inanimate stimuli. (Of course, the stimuli aren't animate-they represent animate items). For example, there were processing differences between changes between animate and inanimate items (rabbit to boot) that were not present in rabbit to dog. They also showed such differences in two kinds of visual search tasks.

      Of course, there are feature differences that exploit orientation. A classic example is the difference between a square and a diamond that is produced from the square by rotating it 45 degrees.

      They addressed an aspect of this challenge having to do with some features using silhouettes. There was no search advantage for silhouetted stimuli.

      Weaknesses:

      I thought this was an excellent submission. I have two suggestions for revision:

      We are glad to hear this Reviewer recognizes the broad challenge we are tackling in this work (separating high-level from low-level visual features) and found our submission to be “excellent”.

      (1) I thought that experiment 7 should have been described in more detail, with the upshot explained better. What exactly do the authors take it to show?

      Sorry for the lack of clarity here. We think the Reviewer actually gets this right earlier in their review; many feature differences exploit orientation, and our silhouettes control (Experiment 7) shows that those differences alone fail to explain our effects. For example, one might worry that our search effects merely reflect oddities in the aspect ratio or center of mass of the images. Converting the anagrams into silhouettes preserves these features. Thus, the fact that we found no search advantage with silhouettes suggests that these features on their own fail to produce the relevant effects; put the other way around, the effects we observed earlier must go beyond those features. We now discuss this in greater depth in our paper.

      (2) There should be a candid discussion of what the loose ends are and how they might be addressed. It would be good to have some examples like the square/diamond case with some indication of what would address such challenges.

      We agree with this, though we are somewhat limited by the space constraints of the Short Report format. A primary loose end we see is the possibility that high-level properties other than animacy explain our results (as raised by other Reviewers). We have added some discussion of this possibility to the paper.

      We would like to thank this Reviewer for their thoughtful feedback.

      Reviewer #2 (Public review):

      Summary:

      The authors present a creative approach using visual anagrams matched on low-level image statistics to isolate animacy from low-level visual features and report consistent effects of animacy on visual working memory and attention. While this is a thoughtful design and is well executed across seven pre-registered experiments, it remains unclear whether the reported effect is truly driven by animacy, as opposed to broader differences in ensemble statistics or semantic structure across the "mixed animacy" versus "uniform animacy" conditions. As such, the interpretation of a "pure" animacy effect may be overstated.

      Strengths:

      (1) An important methodological advance in controlling low-level confounds that have historically complicated the study of animacy.

      (2) The converging effects across multiple experiments, together with the pre-registered design, strengthen the reliability of the reported findings.

      We are glad to hear this Reviewer found our work to be “creative” and believes it offers an “important methodological advance”.

      Weaknesses:

      (1) Specificity of the animacy effect vs. category-level ensemble structure

      The central claim is that animacy itself drives the observed effects. However, the key manipulation ("mixed animacy" versus "uniform animacy") also introduces differences in category-level ensemble structure. For example, in Experiments 1-2, cross-category change detection (e.g., dog to chair) may be easier not because of animacy per se, but because of a change in overall ensemble statistics (Brady & Alvarez, 2011, 2015). In addition, since each display contains five objects (two in one category and three in the other category), cross-category changes may also alter category balance in a way that further facilitates detection. In contrast, within-category changes preserve both ensemble structure and category composition, making them more difficult to detect.

      Brady, T. F., & Alvarez, G. A. (2011). Hierarchical encoding in visual working memory: Ensemble statistics bias memory for individual items. Psychological Science.

      Brady, T. F., & Alvarez, G. A. (2015). Contextual effects in visual working memory reveal hierarchically structured memory representations. Journal of Vision.

      We appreciate the opportunity to clarify our claims and the support for them. Our claim is indeed that animacy (or a closely related high-level property; see below) drives our effects, over and above its lower-level correlates — i.e., that the explanation for differences in change detection or search across conditions will invoke a high-level property of the images. As we understand the Reviewer’s concern(s), they either (a) are already addressed by our novel methodology, or (b) would still fall perfectly in line with our claim as stated above.

      Consider the Reviewer’s concern that cross-category change detection “may be easier not because of animacy per se, but because of a change in overall ensemble statistics”. Which ensemble statistics change across categories in our stimulus set? Take as an example the case depicted in our figure, where a rabbit changes into either a dog (within-category) or a boot (cross-category). The dog and the boot are the very same image, just rotated; thus, they have the same luminance, curvature, area, spatial frequency, and so on. So if the change from rabbit to dog changes the array’s ensemble statistics with respect to any of those properties, it does so in the very same way as the change from rabbit to boot — and yet detection is still better for rabbit → boot than for rabbit → dog. Indeed, for nearly any ensemble statistic, the difference between the rabbit-display and the dog-display will be identical to the difference between the rabbit-display and the boot-display. To engage with the specific cases discussed in the two cited papers (Brady & Alvarez, 2011, 2015): The dog and the boot are the same size (because they are the same image), so average size is identical (just as average luminance, curvature, area, and spatial frequency are identical). And the very few properties left over (e.g., aspect-ratio) are addressed by later experiments.

      To be clear: We are not saying that there are no differences in ensemble statistics between the rabbit-display and the dog-display; across those displays, we replace one image with a different image, so there are likely all kinds of corresponding differences in ensemble statistics. The key question is whether that change in ensemble statistics differs across trial types in ways that might explain our effect - i.e., whether there is any difference between the rabbit-display and dog-display that is not also present between the rabbit-display and the boot-display. We don’t see how the answer could be yes, at least with respect to the statistics typically considered. A similar logic applies to the search tasks, with the silhouette control (Experiment 7) providing especially strong evidence that certain ensemble statistics or lower-level features cannot explain our effect.

      Now, it’s possible the Reviewer is referring to properties other than the low-/mid-level properties we mention above. Perhaps, for example, many animate stimuli on a display at one time have a striking collective appearance (all these animals are looking at me!) that lots of inanimate stimuli do not (this might be related to the Reviewer’s concern about “category balance”). But as we see it, this explanation just invokes animacy all over again, and so is the sort of explanation we would embrace.

      We now say more about this concern in the paper to be as clear as possible about our claims.

      (2) Limited stimulus set and potential learning effects

      The relatively small stimulus set (six anagram pairs) and repeated exposure raise the possibility of learning or familiarity effects. Does performance change over time? e.g., are there meaningful differences between early and late trials (e.g., first 10% vs. last 10%)? If such differences are present, they could suggest the development of task-specific strategies or increased efficiency with repeated exposure, rather than stable effects driven by the experimental manipulation itself.

      This is an interesting question, and we recognize this analysis absent from our initial submission. To be fair, stimulus sets of this size are not unusual in change-detection and search tasks, which often involve red, green, and blue squares repeated over the course of several hundred trials. Still, we certainly take the Reviewer’s point here and also embrace their analytical approach to addressing it. We’ve now run the “familiarity effects” analyses the Reviewer suggests (as well as some they did not suggest). The top-level headline is that learning or familiarity effects cannot explain our results, and if anything most of these analyses not only fail to support this alternative account but actively point against it. Below are more details.

      First, we worry that the Reviewer’s concern about “the development of task-specific strategies … rather than stable effects driven by the experimental manipulation itself” isn’t actually addressed by the suggested analysis of comparing the last 10% of trials to the first 10%. One reason for this is simply that it’s possible that both mechanisms are at play - i.e., that there is a baseline difference even without any familiarity that is then enhanced by some learning mechanism. (There are other issues as well: For example, one might imagine that participants get quite good at the task during the middle 80% of trials, but then get fatigued at the end. If this were true, then comparing the first 10% to the last 10% of trials could make it seem like there is no learning or familiarity, even if there were such effects. And on top of all this there is just the issue of statistical power, since far fewer trials go into these analyses than into our primary, pre-registered analyses). Nevertheless, we ran the Reviewer’s proposed analyses (using the first and last 10 trials of each type, which offers the best chance to find the pattern the Reviewer is concerned about). If anything, this analysis points in the opposite direction to the Reviewer’s prediction: 4/6 experiments (Experiments 1, 3, 4, and 5) revealed numerically weaker effects at the end of the task than the start, while only 2/6 experiments (Experiments 2 and 6) revealed numerically stronger effects at the end of the task than the start. Moreover, most of these results were non-significant, with only one marginal result (Experiment 5, p< = 0.08) and one significant result (Experiment 2, p = 0.01), and this is before any correction for multiple comparisons, which would make all of these results non-significant. So even though our account could easily accommodate learning effects, it’s not clear that they even exist here in any consistent or reliable way.

      Second, however, we think a more informative way to answer the Reviewer’s question is to ask not about learning over the course of the experiment but rather whether the key effects arise very early in the task. If they do, then any learning effects arising later couldn’t fully account for our results. Now, again, these tests are underpowered and only exploratory (to do this analysis properly, we would want to run entirely new experiments designed for this purpose), but we in fact did find evidence that our key effects arise early. In 5/6 experiments (Experiments 1, 3, 4, 5, and 6), the key effect was significantly (or in one case marginally) present even at the beginning of the experiment (Experiment 1, p = 0.07; Experiment 3, p = 0.01; Experiment 4, p < 0.001; Experiment 5, p < 0.001; Experiment 6, p < 0.01), and most of these results would survive correction for multiple comparisons. (In only one experiment, Experiment 2, was there a numerical disadvantage, but it was not significant; p = 0.34.) So even though our experiments were not designed or powered for this purpose, they do seem to suggest that the effects arise even without much familiarity at all.

      All told, we think these analyses suggest quite strongly that learning alone fails to explain our key effects. There is no evidence that the effects in general are stronger at the end of the experiment than the beginning (if anything it is the opposite); and there is evidence that most of the effects we investigated can be detected even very early in the experimental sessions. We have added discussion of these new analyses to our manuscript.

      (3) Role of semantics

      Although the anagram paradigm effectively controls low-level visual features, it still relies on high-level semantics (e.g., "dog" vs. "boot"). These stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories. From a semantic standpoint, it remains unclear whether the observed effects can be uniquely attributed to animacy or whether they reflect broader conceptual distinctions.

      We agree with the Reviewer here. While we feel comfortable interpreting our effects in terms of a high-level property like animacy as opposed to a lower-level property like curvature, it remains possible that the observed effects reflect some other, closely related high-level distinction (like natural vs. manmade). Our primary concern was to tease apart high-level properties from low-level features, which the Reviewer’s question does not threaten — if attention and memory are sensitive to the natural/artificial distinction, that’s interesting too, and a near neighbor of our actual claim. Still, we agree that this could be addressed, and we even see it as an empirical question testable in future work. Perhaps the most relevant departures between animate/inanimate and natural/manmade include objects like clouds, plants, and rocks — objects that are natural but not “animate” in the sense often used in this literature. If something like our paradigm revealed that rocks behave more like dogs than like boots, that would suggest that naturalness, rather than animacy, was driving the effects; but if rocks behave more like boots than like dogs, that would point to animacy even more strongly. We remain open-minded about this possibility, but it would of course require multiple new experiments with a brand new stimulus set and so goes beyond the present contribution. In any case, we have added a discussion of this issue to the paper and have adjusted our claims accordingly.

      Reviewer #3 (Public review):

      Summary:

      This study makes clever use of generative AI to create stimuli that are pixel-for-pixel identical but which have radically different meanings depending on their orientation, to investigate the perception of animacy while retaining control over low-level image features (so-called 'anagram' stimuli).

      The authors present seven elegantly designed experiments in a commendably compact format.

      Experiments 1 and 2 involved a working memory paradigm in which participants had to spot which of five objects in an array changed after a pause. Importantly, the changed object was an anagram stimulus that in one orientation matched the animacy/inanimacy of the changed object, and in the other orientation was the opposite (e.g., a rabbit is replaced by either a dog or a boot, where the dog and boot stimuli are actually identical, just rotated by 90 degrees). They found a difference in accuracy depending on whether the animacy of the objects matched.

      Experiments 3 and 4 used a visual search task in which the participants had to localize the target, and the distractors were anagrams that either matched the target in terms of animacy or did not. There was a significant cost in terms of response time when the animacy of the target was the same as that of the distractors. Experiments 5 and 6 also used a similar visual search design, except that the task was to determine if the target was present or absent from the display, and the distractors again either matched or differed from the target in terms of animacy. Again, the authors found slower responses when the distractor arrays matched the animacy of the target than when they differed.

      An obvious potential concern about the studies is addressed by Experiment 7. It is unclear if the observed effects are related to the specific orientations of the target and distractor stimuli selected in each condition. For example, it could be that all the animate versions of the anagrams involved tall and skinny shapes, while all the inanimate versions involved wide and short objects, due to the 90-degree rotational difference between the two versions of the stimuli. To control for this, the authors repeated the visual search experiment but with convex-hull silhouettes of each of the stimuli. In other words, all targets and distractors from each trial were replaced by a black splotch with approximately the same overall outline (envelope) as the corresponding stimulus. Importantly, in contrast to the anagram stimuli, the silhouettes had had no meaningful semantic interpretation, and their animacy did not change depending on their orientation.

      Strengths:

      The main strength is the elegant use of stimuli that control almost perfectly for low-level image features.

      Thank you for this kind feedback. This summary perfectly captures both our empirical contribution and the claims we are making.

      Weaknesses:

      My only real concern about the study is whether the findings truly provide evidence for a high-level visual representation of animacy independent of the low-level stimulus characteristics, or whether, instead, the effects are essentially semantic priming, which is independent of visual processing per se. For example, if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results? Words can also access semantic representations of the animacy of objects, and also don't suffer from low-level visual confounds. It would be helpful to add a discussion of this possibility to the article.

      Wow, we love this question! And so we’ve now conducted exactly the experiment the Reviewer suggests here. In a new pre-registered study (Experiment 8), we presented participants with a present/absent search task (as in Experiments 5–7). One half of trials consisted of the anagram stimuli (such that we could, once again, replicate the mixed-animacy search advantage); but the other half of trials consisted of the words describing the anagrams (e.g., “dog”, “boot”, “sheep”, “car”, etc.). The experiment worked beautifully: We found no effect with the words, but replicated the search advantage with the pictures — and also found a significant difference between the effects elicited by the two stimulus types.

      We agree with the Reviewer that this now rules out the possibility that semantic representations alone explain these visual effects. Thank you! 

      Reviewer #4 (Public review):

      In this article, the authors investigate whether perceived animacy influences visual processing independently of lower-level visual features by using "visual anagrams." Across seven experiments, they test whether animacy, isolated from many lower-level visual properties, structures visual working memory and guides visual attention. The central claim is that the visual system may represent animacy itself, rather than animacy emerging solely from associations among low-level visual properties.

      I find this investigation compelling. The experiments described provide strong control over several lower-level visual features, including curvature, texture, and related image properties. However, the visual anagrams are not pixelwise-identical across orientations. Because the images are rotated, the retinal configuration of pixels and the spatial organization of some low- to mid-level shape features also change. As a result, the configural arrangement of mid-level visual features may still contribute to perceived animacy.

      We are glad to hear the Reviewer finds our investigation “compelling”.

      I encourage the authors to discuss how independent perceived animacy is in this context from the contribution of mid-level visual features, such as configural shape cues that are diagnostic of animacy. This distinction would help sharpen the interpretation of the results and more precisely define the level of visual representation isolated by the visual-anagram approach.

      This is a helpful point, and it also echoes a sentiment expressed by Reviewer #2. While configural shape is diagnostic of animacy writ large, it can’t account for our observed effects here because rotating an image does not vary its configural shape. We now mention this in our work, and we agree that it helps sharpen the interpretation of our studies.

      Additionally, previous studies have argued that low- and mid-level curvilinear features may contribute to animate/inanimate categorization, and may in some cases be sufficient to support such distinctions (e.g., PMID: 33798259; PMID: 28654965). I encourage the authors to clarify how these previous findings on curvilinearity and rectilinearity fit with the overarching claim of the current study, namely that the visual system may represent animacy itself rather than animacy emerging solely from associations among lower-level visual properties.

      Yes, many studies from exactly that corner of the field actually motivated the present work, which is why we cited them in our submission. In a way, we are approaching this issue from the other side of the equation. Whereas the papers the Reviewer points to (along with many others) ask whether mid-level features (such as curvilinearity and rectilinearity) are sufficient to support perceived animacy, we ask whether these and other features are necessary to support perceived animacy. Prior work is relatively split on this issue, leaving the question wide open. We take our work to show that differences in curvature are not necessary for differences in perceived animacy, because our anagrams have identical curvature yet differ in animacy — and the visual system capitalizes on that difference. Put the other way around, representation of animacy can and does go beyond representation of its low- and mid-level correlates. Thank you!

    1. eLife Assessment

      This study provides important insights into the neural mechanisms linking sleep and long-term memory consolidation. By combining behavioural, genetic, imaging, and connectomic approaches in Drosophila, it identifies a target neural circuit that will be of broad interest to researchers studying sleep, memory, and neural circuits. The evidence supporting the involvement of the identified circuit in the regulation of sleep and memory is solid and represents a substantial advance in the field. Nevertheless, there is limited evidence to support the mechanistic claim that this circuit directly links sleep and memory consolidation within the available data, and some results should therefore be interpreted with appropriate caution.

    2. Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behaviour, imaging and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      The investigation follows a logical strategy to collect huge dataset of sleep, appetitive memory and live imaging. The authors identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation dependant manner. The author also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the new data provide TRIC-LUC provided better temporal resolution of neural activity correlates for PAMalpha1-DPM inhibition. Importantly, the author demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP at the memory consolidation period after training.

      Weaknesses:

      Although the revised version carries arguments to satisfy the reviewers' concern, the writing is now less cohesive. Crucially an explanation however remains required for the following experimental contradiction: the central observation of the study indicates that PAM alpha1 activation cause DPM inhibition which disrupt sleep and memory consolidation. Therefore, one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found the opposite is true from now enhanced TRIC-LUC dataset. The authors indicate this data reinforce the inhibitory nature of PAM-alph1-DPM, but it does not explain why such a reduced DPM activity is observed after training.

    3. Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain incompletely understood. The authors examined a specific subset of PAM dopaminergic neurons, PAM-α1, and DPM neurons in Drosophila. These neurons have previously been implicated in memory, and DPM neurons have also been linked to sleep. The study explores whether this circuit provides a mechanistic link between sleep and memory consolidation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-α1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly during the night. The authors further show that perturbation of PAM-α1 and DPM neurons impairs sleep and appetitive memory consolidation under starvation conditions, and that pharmacological sleep induction during the night rescues the LTM defects. Together, these findings suggest that PAM-α1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important observations that advance our understanding of the circuits regulating sleep and memory consolidation.

      Weaknesses:

      Some claims require additional evidence or clarification.

      (1) Previous studies linking impaired memory to reduced sleep have primarily examined conditions involving severe sleep deprivation. In contrast, this manuscript argues that relatively modest decreases in total sleep, accompanied by sleep fragmentation, are sufficient to impair memory consolidation. It remains unclear whether sleep fragmentation of this magnitude is itself critical for LTM consolidation. An independent method for inducing comparably mild sleep loss and fragmentation would be needed to directly test this interpretation.

      (2) It is unclear why both activation and inactivation of PAM-α1 neurons produce similar effects on sleep and memory. In addition, MB299B-labeled neurons exert stronger effects on memory than MB043B-labeled neurons, whereas MB043B-labeled neurons have stronger effects on sleep. If sleep disruption is the primary driver of impaired memory consolidation, a stronger correspondence between the sleep and memory phenotypes might be expected. The authors speculate that MB043B may affect sleep through non-PAM neurons, but without identifying the relevant neurons, this remains speculative.

      (3) The complex schematic model (Fig. 12), with parallel circuits and unidentified neuronal groups, underscores the difficulty of interpreting the current data. In the "less activity" arm of the model, distinct circuits are proposed to regulate sleep and LTM, respectively, and DPM neurons are not included. This makes it difficult to reconcile the model with the central claim that the PAM-α1-to-DPM microcircuit links sleep and LTM consolidation.

      (4) The TRIC-LUC reporter system is not ideal for resolving dynamic changes in neuronal activity. Activity-dependent Ca²⁺ signaling must first reconstitute the TRIC transcriptional system, which then drives luciferase transcription, translation, and accumulation. The original characterization of TRIC indicates that TRIC signals accumulate and decay over several hours. Thus, the kinetics of the TRIC-LUC reporter should be interpreted cautiously, particularly when inferring transient or precisely timed changes in neuronal activity.

      (5) Including data from training under fed conditions would provide a more complete understanding of state-dependent neural activity and would help distinguish starvation-specific effects from more general circuit mechanisms.

    4. Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence how dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strength:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuit interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of the role of dopamine in cognitive functions. Additional experiments investigating the contribution of MBON-α1 to the circuit, connectomic analysis together with a dissection of dopamine receptor function further strengthen the proposed circuit motif and its biological relevance.

      Weaknesses:

      While the study is well designed, presents compelling findings and has been further strengthened by additional experiments, some aspects remain unclear. The role of DPM neurons in memory consolidation seems not yet fully resolved, as different genetic approaches yield variable results. Furthermore, some manipulations impair memory without affecting sleep fragmentation - or vice versa, suggesting that the observed memory deficits cannot be explained solely by impaired sleep-dependent consolidation. It would also have been interesting to discuss potential mechanisms by which dopamine receptor-mediated cAMP signaling could lead to a reduction in Ca²⁺ signals. I am confident that these questions can be addressed in future studies.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopaminergic circuits interact to regulate memory consolidation in Drosophila and may reveal general principles underlying the neural regulation of memory.

    5. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This study approaches an important topic providing insight into the neuronal circuitry that interconnects memory consolidation and sleep. The data were collected and analysed using a solid methodology, contributing new findings for neurobiologists working on how memories are stored and the roles of sleep. However, the data is incomplete to support the proposed role of the PAM-DPM circuits as the link between sleep state and long-term memory consolidation.

      We sincerely appreciate the editor and reviewers’ thoughtful and constructive comments on our study. Your insightful feedback has not only affirmed the significance of our work on the interplay between memory consolidation and sleep, but also provided valuable inputs for improving the clarity, rigour, and impact of our study.

      We have carefully addressed all the comments raised by the reviewers and revised the manuscript accordingly. We have also streamlined the paper with the goal of making it more accessible to readers. We feel this revised version strengthens our conclusion that the PAM-DPM circuits as the link between sleep and memory consolidation.

      The main improvements in terms of data addition are three complementary sets of circuit-specific experiments:

      (1) To better characterize the dynamics of the PAM-DPM circuit following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. This experiment addresses the activity of the microcircuit in a much more relevant time frame than the CRTC data in the previous version of the paper which looked only at the first hour after training. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, both PAM-α1 and DPM neurons displayed similar neural activity over the 3-hour recording period. The first hour was characterized by a synchronous decrease in activity for both neuron types, with hours 2 and 3 achieving a stable baseline. Since the decrease in the first hour is also seen in the trained condition, we think it is likely a reflection of the animals becoming acclimated to the recording tubes.

      Notably, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit in the LTM consolidation time window. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 8C-F). These findings not only support the hypothesized role of the inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics.

      (2) To further support the functional connectivity of the PAM-DPM microcircuit, we conducted in vivo experiments to complement the dissected brain prep P2X2 data. Optogenetic activation of PAM neurons in intact flies via the red light-gated cation channel CsChrimson (Klapoetke NC et al., 2014) resulted in a significant reduction in GCaMP signals within DPM neurons (revised Figure 2B). These findings strongly confirm that PAM neurons exert direct inhibitory control over DPM neurons in the intact brain.

      (3) Further, we investigated how dopamine signaling to the DPM inhibits its activity, and issue which has not been investigated previously. We conducted a series of experiments:

      Firstly, we verified which dopamine receptors (Dop1R1, Dop1R2, DopEcR, and Dop2R) express on the DPM neurons via double-labeling with gene-embedded GAL4 lines. We found that DPM neurons have expression of both Dop1R1 and Dop1R2 (revised Figure 10A).

      Secondly, to clarify which receptors on DPM neurons respond to dopamine and how they signal, in addition to EPAC experiments in the first submission, we recorded neural activity changes when we knocked down Dop1R1 and Dop1R2 in DPM neurons. DPM neurons exhibited a significantly reduced GCaMP level with DA application, regardless of whether Dop1R1 or Dop1R2 was intact or knocked down knockdown in comparison to the no-DA control condition (revised Supplemental Figure 3C-E). These data suggest that either residual Dop1R1 and Dop1R2 remaining in the RNAi condition is sufficient or that the two receptors may coordinate to mediate the inhibition of neural activity.

      Finally, we investigated the behavioral contributions of Dop1R1 and Dop1R2 in DPM neurons to sleep and memory processes (revised Figure 10C-H). Dop1R1 knockdown resulted in a marked reduction in daytime sleep and a significant impairment of 24 h memory expression. In contrast, Dop1R2 knockdown selectively compromised 24 h memory without affecting sleep.

      When integrated with our EPAC assay findings from the initial submission, which demonstrated, that Dop1R1 is the primary receptor mediating dopamine-induced cAMP elevation, these new data collectively delineate a more complex mechanistic framework: dopamine signaling in DPM neurons coordinates the dual regulation of sleep and memory predominantly via Dop1R1. Meanwhile, Dop1R2 are engaged in the selective modulation of memory.

      All newly generated experimental datasets, comprehensive statistical analyses, and their corresponding figure panels (revised Figures 2B, 10, 11 and Supplemental Figure 3) have been fully incorporated into the revised manuscript.

      In addition to adding the experiments described above, we have reorganized and streamlined the paper. First, the CRTC data have been replaced by the Tric-luc data. The CRTC data were taken in the first hour after training and do not shed light on the bulk of the consolidation window. Since the behavioral and sleep effects we see with manipulation of the PAM/DPM microcircuit all occur with a time delay, examining later times in consolidation is more relevant. Additionally, the first hour post-training is quite complex since there are sensory changes and STM processes overlaid on the processes we want to study. Second, we have moved the data in Figure 8 to supplemental (revised Supplemental Figure 2) since they are basically a control for the experiments in Figure 7 validating known requirements for appetitive LTM.

      We have also substantially expanded the Discussion section to contextualize the PAM-DPM circuit within the broader framework of well-characterized memory-regulatory pathways, such as the intrinsic circuits of the mushroom body, and to explicitly delineate the hierarchical interplay between sleep-dependent synaptic plasticity and LTM consolidation.

      We contend that these complementary experimental assays and targeted revisions markedly strengthen the causal evidence underscoring the role of the PAM-DPM circuit as a pivotal regulatory node bridging sleep states and LTM consolidation. We are confident that these revisions essentially address the concerns raised by the reviewers.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behavior, imaging, and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      From a Drosophila sleep researcher's perspective, the investigation follows a clear and logical strategy to collect a huge dataset of sleep, appetitive memory, and live imaging. The authors clearly identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation-dependent manner. The authors also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the authors applied a new way of sleep statistics to demonstrate hour-by-hour changes between treatment and genotypes. Importantly, the authors demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP during the memory consolidation period after training.

      Weaknesses:

      Two investigatory gaps relate to the misalignment between circuital activity and behaviors, due to the nature of large circuital functional analysis like this. Firstly, the central observation of the study indicates that PAM alpha1 activation causes DPM inhibition which disrupts sleep and memory consolidation. Therefore one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found that the endogenous CRTC::GFP reported neuronal activity for PAMalpha1 and DPM are both increased after memory training (Figure 9). This can be due to the difficult functional demarcation among the 14 PAMalpha1 projections. Secondly, the authors acknowledged the contradicting finding that memory defect is detected in PAMalpha1 inactivation (Figure 7C), yet suggested a tight link between sleep and memory consolidation; it is clear loss of PAM subset activity can disrupt memory consolidation without affecting sleep (cf Figure 7C and 7I).

      Thank you for your insightful analysis and the relevant possibilities you've raised. We agree that given that memory consolidation and sleep are time-dependent processes, the 1-hour window employed to capture neural activity changes via the CRTC::GFP reporter may not fully reflect the overall dynamics of neural activity in this microcircuit. To better characterize the dynamics of the PAM-DPM circuit in the consolidation window following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, PAM-α1 neurons displayed stable neural activity over the entire 3-hour recording period. However, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 8C-F). These findings not only support to the hypothesized role of inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics.

      Regarding the second question, the core finding underlying the link between sleep and memory elucidated in the present study lies in the whole PAM-α1-DPM microcircuit rather than the specific DANs alone. MB299B and MB043B, the two split-GAL4 drivers employed to target PAM-α1 neurons, were originally characterized previously (Aso et al., 2014). However, these drivers also exhibit non-specific labeling of additional cells, and we can not rule out the possibility that such off-target labeling may have masked the subtype-specific necessity in sleep or memory processes.

      Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain poorly understood. The authors present novel findings implicating two small neuronal groups with inhibitory connections, PAM-a1 to DPM, in sleep regulation and LTM consolidation. However, whether the PAM-a1 to DPM microcircuit promotes LTM consolidation through sleep regulation requires further investigation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-a1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly at night. Interestingly, disruption of PAM-a1 and DPM neurons impairs sleep and appetitive memory consolidation only under starvation conditions, and pharmacological induction of sleep during the night rescues the LTM defects. These findings suggest that PAM-a1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important findings that advance our understanding of the link between sleep and memory consolidation.

      Weaknesses

      Some claims lack sufficient evidence or clarity:

      (1) All sleep experiments are conducted under the "training" (temperature-change) condition. While genotypic controls are helpful, additional no-training controls are required to confirm that the observed differences are due to training rather than unknown genotype-related factors. The fact that experimental genotypes exhibit significantly altered sleep even before "training" (e.g., Figs. 7H, J, K, 8A, B, D) highlights the necessity of these controls.

      Thank you for raising this important question. We have re-examined the sleep profiles recorded over two acclimation days and one day of baseline sleep, which preceded the implementation of the “training” paradigm (temperature manipulation) and thus served as a valid no-training control. As shown in Author response images 1-4, subtle yet discernible genotype-dependent differences were indeed observed under baseline conditions. However, when animals were subjected to starvation, the experimental manipulations (activation or inactivation of the target cells) elicited marked, statistically significant alterations in sleep patterns that cannot be accounted for by the baseline genotype differences. Collectively, these data confirm that the observed sleep phenotypes are attributable to the “training” intervention, rather than to confounding, pre-existing genotype-related factors.

      Author response image 1.

      Baseline and manipulation day sleep profiles following PAM activation and PAM/DPM inactivation under starvation conditions.

      Author response image 2.

      Baseline and manipulation day sleep profiles following PAM activation and PAM/DPM inactivation under non-starvation conditions.

      Author response image 3.

      Baseline and manipulation day sleep profiles following PAM- α1 activation and inactivation under starvation conditions.

      Author response image 4.

      Baseline and manipulation day sleep profiles following PAM- α1 activation and inactivation under non-starvation conditions.

      (2) Previous studies on disrupted memory due to sleep reduction have primarily examined conditions with severe sleep deprivation. In contrast, this report claims that relatively small decreases in total sleep accompanied by sleep fragmentation are responsible for impaired memory consolidation. It remains unclear whether sleep fragmentation at this level is truly critical for memory consolidation. The authors should cause sleep loss and fragmentation of similar magnitude through other means and determine whether it can impair LTM.

      We appreciate the reviewer’s insightful suggestion. While alternative assays for inducing sleep loss or sleep fragmentation are indeed available, this line of investigation lies beyond the core scope of the present study. We will certainly take this valuable suggestion into consideration for the future studies.

      (3) The authors employed a neural activity reporter to show that starvation increases the basal activity of PAM-a1 but not DPM neurons in untrained flies (Figures 9C-E). They observed small increases in the activity of both neuron groups immediately after training but not one hour later. Given the inhibitory connection from PAM-a1 to DPM, it is unclear why both neuron groups show increased activity after training. Additionally, as the authors acknowledge, it is puzzling how the inactivation of PAM-a1 produces similar effects on sleep and memory as DPM inhibition and PAM-a1 activation. Further experiments are needed to clarify these findings, such as manipulating PAM-a1 activity during the one-hour post-training period and evaluating the effect on DPM activity. Including data from training under fed conditions would provide a more comprehensive understanding of state-dependent neural activity. Even if certain experiments are not feasible, these issues warrant further discussion. It is also important to clarify that the term "synchronized" does not imply single-spike-level synchrony.

      Thank you for raising these critical questions. To deepen our understanding of these issues, we have conducted additional experiments and have incorporated them into the revised manuscript. Below are our specific responses to each of your points:

      (1) Regarding the contradiction between "PAM-α1 inhibition of DPM" and a transient increase in the activity of both neurons immediately after training:

      PAM/PAM-α1 neurons are well-documented to respond to reward signals (Liu et al., 2012, Ichinose et al., 2015), while DPM neurons have been shown to respond to both olfactory stimuli and electric shocks, and to form delayed olfactory memory traces (Yu et al., 2005). Thus, the concurrent increase in the activity of PAM-α1 and DPM neurons immediately following training is likely a response to the olfactory and/or sucrose stimuli in the assay. Given that memory consolidation and sleep are time-dependent processes, the 1-hour window employed to capture neural activity changes via the CRTC::GFP reporter likely does not fully reflect the overall dynamics of neural activity in this microcircuit. Additionally, this time window overlaps with the period in which the animals are adapting to the new tubes and is likely contaminated with other sensory information.

      To better characterize the dynamics of the PAM-DPM circuit following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, PAM-α1 neurons displayed stable neural activity over the entire 3-hour recording period. Notably, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 9F-I). These findings not only support to the hypothesized role of inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics post-training. We have replaced the CRTC data with this more relevant data set.

      (2) Regarding the state-dependent neural activity:

      We agree that investigating state-dependent neural activity would be an interesting extension of our study. However, this falls beyond the scope of the current study and will be considered in future research. Our primary findings, including sleep disruptions and the associated memory impairments, were specifically observed under starvation conditions, which align with the appetitive memory paradigm employed here. Delving into neural activity changes under non-starvation state would not yield direct evidence to support the core conclusions of the present work, as the study’s focus is on the starvation-dependent interplay between sleep, neural circuitry, and appetitive memory consolidation.

      (3) Regarding the terminology of “synchronization”:

      We believe that the use of the term “synchronization” in our study is appropriate. In the context of neural circuitry, synchronization refers to the process by which distinct neurons or neural populations achieve temporal alignment of their activity, a phenomenon that supports neural communication and information integration. In the present work, this specifically describes how PAM-α1 and DPM neurons exhibit phase-related temporal coordination of their activity to regulate the interplay between sleep and memory consolidation.

      (4) The authors considered that PAM-a1 and DPM might function in parallel, independent pathways for sleep and LTM. They rejected this possibility based on the lack of additive effects when both neuronal groups were simultaneously inactivated. However, they found that MB299B-labelled neurons exert stronger memory effects than MB043B-labelled neurons, while MB043B neurons have stronger sleep effects. If sleep is a primary driver of memory consolidation, a stronger correlation between memory and sleep effects would be expected. This observation merits further discussion.

      We appreciate the reviewer’s constructive suggestions. We have performed additional experiments to explore a well-characterized memory-related PAM-α1 recurrent loop in sleep regulation. The new data, along with further discussion, have been incorporated into the revised manuscript.

      The two split-GAL4 drivers (MB299B and MB043B) used to target PAM-α1 neurons were originally characterized previously (Aso et al., 2014). However, these drivers exhibit non-specific labeling of additional neuronal populations, a technical limitation that may have masked the subtype-specific functional requirements of PAM-α1 in sleep and memory processes.

      In addition, we assessed sleep and LTM following the thermoactivation of DPM neurons (revised Supplemental Figure 1), and no significant changes were observed in either phenotype.

      PAM-α1 has previously been demonstrated to drive appetitive LTM formation and consolidation via a recurrent loop with MBON-α1 (Ichinose et al., 2015). To investigate whether MBON-α1 also participates in sleep regulation, we activated or inactivated MBON-α1 neurons under both starvation and non-starvation conditions. Our results revealed that inhibition of MBON-α1 under both starvation and non-starvation conditions resulted in a significant reduction in sleep and a reduced arousal threshold (revised Figure 11B, D), suggesting that MBON-α1 participates in regulating sleep in a state-independent manner. However, no significant changes were observed upon activation of MBON-α1 neurons (revised Figure 11A, C). Combined with our observation that inhibition of MBON-α1 during the memory consolidation phase also impaired 24 h LTM, these new data indicate that MBON-α1-mediated sleep is necessary for effective memory consolidation. Notably, while activation of MBON-α1 during consolidation phase similarly impaired LTM, this manipulation did not alter the sleep profile, suggesting a dissociation between MBON-α1’s mechanistic roles in sleep regulation and LTM processing.

      Taken together (see Author response table 1 and the new schematic diagram of revised Figure 12), these findings reveal a dedicated hierarchical, modular regulatory network that mediates sleep-LTM coupling via an activity-dependent mechanism. Within this network, activation of PAM-α1 acts as an upstream modulator to inhibit the activity of DPM, a downstream integrative hub that coordinates the execution of sleep and memory processes via recruiting different signaling cascades mediated by distinct dopamine receptors. MBON-α1, which is likely inhibited by PAM-α1, serves as parallel pathway to suppress sleep and impair LTM. Conversely, inactivation of PAM-α1 relieves its inhibitory control over MBON-α1, leading to MBON-α1 activation; MBON-α1 then functions as a signal amplifier that further exacerbates the reduced activity of PAM-α1, ultimately resulting in LTM impairment. Inactivation of PAM-α1, together with non-PAM-α1 neurons labeled by MB043B, contributes to the regulation of sleep. Sleep and memory are highly intertwined within this circuit, where distinct neuronal populations exhibit specialized yet interdependent functional roles, with overlapping and divergent regulatory contributions to sleep and LTM. The inherent complexity of this regulatory network thus merits further dedicated investigation in future studies.

      Author response table 1.

      (5) Given prior knowledge that PAM neurons are heterogeneous and that the R58E02 driver is broadly expressed, data in Figures 1-5 concerning PAM are outdated. The use of more restricted PAM-a1 drivers from the outset would make the manuscript easier to read and interpret.

      We sincerely appreciate the reviewer’s point of view regarding the selection of PAM drivers. While we acknowledge the well-characterized heterogeneity of PAM neurons and the broad expression profile of the R58E02 driver, and fully agree that employing subtype-restricted drivers enhances the precision of functional interpretation, this set of experiments serves as an essential foundational step and logical basis for subsequent subtype-specific investigations and thus merits retention in the manuscript. As detailed above, the more specific drivers also have some drawbacks in terms of additional expression, making the broad driver critical for setting the stage.

      (6) Some figures lack relevant data, certain experiments are missing necessary controls, and anomalies are present in some data sets.

      We sincerely appreciate the reviewer’s detailed suggestions, and we have revised the manuscript comprehensively in accordance with them.

      Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence that dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strengths:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuits interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of dopamine's role in cognitive functions.

      Weaknesses:

      While the study is well-designed and presents compelling findings, some aspects require further clarification. The interpretation of dopamine receptor signaling remains incomplete, particularly regarding inhibitory pathways. The role of DPM in memory consolidation is not entirely conclusive, as different genetic approaches yield variable results. Additionally, some inconsistencies in neuronal activity patterns and experimental variability, especially regarding sleep patterns or pharmacological rescue, should be addressed to strengthen the mechanistic framework.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopamine circuits interact to regulate memory consolidation. While the findings are compelling, addressing the points above-particularly receptor signaling and the specific role of DPM and its activity patterns within the microcircuit would further solidify the study's conclusions.

      We sincerely appreciate the reviewer’s constructive feedback and useful suggestions, which have been instrumental in enhancing the rigour and completeness of our study.

      To address these points, we have performed a series of additional experiments that we believe strengthen the mechanistic framework of our work. The key new findings are summarized below:

      (1) Regarding the dopamine receptor signaling

      To define the dopamine receptor (DAR) signaling mechanisms underlying DPM neuron activity and its regulatory roles in sleep and memory, we first characterized DAR expression profile of the DPM neurons. Using double-labeling assay, we detected robust expression of Dop1R1 and Dop1R2 in DPM neurons, whereas no detectable colocalization was observed for DopEcR and Dop2R (revised Figure 10A). Accordingly, we refined our FRET-based EPAC data by removing the DopEcR knockdown group, and now present cAMP changes in DPM neurons following Dop1R1 and Dop1R2 knockdown, in direct comparison with the intact receptor control group (revised Figure 10B). These data conform that Gαs-coupled Dop1R1 is the primary receptor mediating DA-dependent cAMP elevation in DPM neurons.

      To further identify the DARs responsible for transducing DA-induced inhibitory effect on DPM neural activity, we quantified GCaMP levels in DPM neurons with targeted knockdown of individual DARs. Knockdown of either Dop1R1 or Dop1R2 failed to abolish DA-induced Ca<sup>2+</sup> decrease; only Dop1R2 knockdown exhibited a trend toward attenuating this Ca<sup>2+</sup> decrease (revised Supplemental Figure 3C-E), suggesting that the two receptors cooperate to modulate DPM neural activity.

      Finally, to dissect the specific contributions of DARs in DPM neurons to sleep and/or memory regulation, we performed sleep monitoring and memory assays in animals with DPM-specific knockdown of distinct DARs (revised Figure 10C-H). Knockdown Dop1R1 in DPM neurons resulted in statistically significant sleep reduction, decreased arousal threshold, and impaired 24 h LTM memory (revised Figure 10C-E). In contrast, knockdown Dop1R2 in DPM neuron selectively impaired 24 h LTM memory with no effect on sleep (revised Figure 10F-H). Collectively, these findings demonstrate that coupling sleep and LTM requires Dop1R1 in DPM neurons through the modulation of both cAMP signaling and neuronal activity, while Dop1R2 specifically mediates LTM regulation, likely through modulating DPM neural activity alone.

      (2) We have additionally characterized the role of MBON-α1 in sleep, which has been previously shown as a PAM-α1-related recurrent feedback loop in the regulation of memory formation and consolidation (Ichinose et al., 2015).

      To investigate whether MBON-α1 also participates in sleep regulation, we activated or inactivated MBON-α1 neurons under both starvation and non-starvation conditions (revised Figure 11A-D). Our results revealed that inhibition of MBON-α1 under both starvation and non-starvation conditions resulted in a significant reduction in sleep and a reduced arousal threshold (revised Figure 11B, D), suggesting that MBON-α1 participates in regulating sleep in a state-independent manner. However, no significant changes were observed upon activation of MBON-α1 neurons (revised Figure 11A, C). Moreover, inhibition of MBON-α1 during the memory consolidation phase significantly impaired 24 h LTM (revised Figure 11E-F). These results indicate that MBON-α1mediated sleep is necessary for effective memory consolidation. Notably, while activation of MBON-α1 during consolidation phase similarly impaired LTM, this manipulation did not alter the sleep profile, suggesting a dissociation between MBON-α1’s mechanistic roles in sleep regulation and LTM processing.

      Taken together (see Author response table 1 and the new schematic diagram of revised Figure 12), these findings reveal a dedicated hierarchical, modular regulatory network that mediates sleep-LTM coupling via an activity-dependent mechanism. Within this network, activation of PAM-α1 acts as an upstream modulator to inhibit the activity of DPM, a downstream integrative hub that coordinates the execution of sleep and memory processes via recruiting different signaling cascades mediated by distinct dopamine receptors. MBON-α1, which is likely inhibited by PAM-α1, serves as parallel pathway to suppress sleep and impair LTM. Conversely, inactivation of PAM-α1 relieves its inhibitory control over MBON-α1, leading to MBON-α1 activation; MBON-α1 then functions as a signal amplifier that further exacerbates the reduced activity of PAM-α1, ultimately resulting in LTM impairment. Inactivation of PAM-α1, together with non-PAM-α1 neurons labeled by MB043B, contributes to the regulation of sleep. Sleep and memory are highly intertwined within this circuit, where distinct neuronal populations exhibit specialized yet interdependent functional roles, with overlapping and divergent regulatory contributions to sleep and LTM. The inherent complexity of this regulatory network thus merits further dedicated investigation in future studies (See Author response table 1).

      We have modified the schematic diagram in the revised manuscript to illustrate the mechanistic framework (revised Figure 12).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here I listed details for potential clarification or further investigation related to the weaknesses:

      (1) Line 145-147: I suspected the authors used previously verified RNAi lines, but it would be informative to include a citation or their own validation for the effectiveness of these RNAi lines.

      We sincerely appreciate the reviewer’s suggestion. As our double-labeling assays confirmed that only Dop1R1 and Dop1R2 are colocalized with DPM neurons (see our responses to the public review from Reviewer #2 and #3), we refined the revised data to focus exclusively on these two DARs. Corresponding revisions have been made to the Materials and Methods, Results and Discussion sections. Additionally, we conducted qPCR analysis to verify the knockdown efficiency of these DARs, providing further support for our findings that Dop1R1 and Dop1R2 are functionally required in DPM neurons for the regulation of sleep and memory (revised Supplemental Figure 3).

      (2) Line 169: Moving from describing Figure 3A/B to Figure 3C, it is not immediately clear from 3C-H, the authors follow the training paradigm of 3B?

      To enhance clarity, we have added the referenced figure citations in the “Memory assay” section: “For all 24 h sucrose-odour memory, a single training session of sucrose paired with an odour for 2 min was employed (Figure 3A).”

      (3) Line 179: before the PAM inactivation data are shown in Figure 7, the authors seem to be getting ahead of themselves by stating "suggesting that activity of PAM neurons is necessary for the consolidation window or that heterogeneity in the subsets of PAM neurons masks any phenotype." when the Figure 3 data collectively indicate that "suppression" of PAM is necessary.

      This statement is based on our observation that inactivation of the majority of PAM neurons labeled by R58E02 results in sleep disruption but leaves memory intact, and we stand by this conclusion.

      (4) Line 186-188: The labelling of DP1 is not entirely aligned between the figures and the text for a reader to follow which time period is described, as DP1 is embedded within the dark phase in the figures.

      These experiments spanned two consecutive days. LP1 and DP1 denote the light and dark periods on the first day, respectively, whereas LP2 designates the light period on the second day. As only one full dark phase was monitored across the experimental interval, we characterized the relevant phenotype using the general terms dark phase or nighttime, rather than specifying DP1. We thank the reviewer for this thoughtful observation; nonetheless, we consider the original description correct and unambiguous, and thus appropriate for inclusion in the manuscript.

      (5) Line 321: The statistics for Figure 9 CRTC::GFP measurement is crucial for interpretation but the referring and labelling for this on Figure 9 is poor: it is not apparent which comparisons are indicated. There is inconsistency between Table 1 Figure 9D and Table 3 Figure 9D entries: no significant between train and untrain indicated in Table 1 but it is described as significant in the text and Table 3?

      Thank you for this observation. As described above, we have removed these data from the paper and replaced them with Tric-Luc data that capture the consolidation window more completely.

      (6) Line 423: The starvation-mediated sleep suppression is not clear in this manuscript, can the author comment on this? The response to this may also alter the summary concept cartoon.

      This is an important point. To directly address the reviewer’s question regarding starvation-mediated sleep suppression, we have generated a representative response figure comparing sleep duration under starvation versus non-starvation conditions (Author response image 5). This figure clearly demonstrates that sleep is suppressed under starvation, providing straightforward evidence to address this concern.

      Author response image 5.

      Examples of starvation-mediated sleep suppression.

      However, the key focus of our study is that changes in neuronal activity disrupt sleep under starvation conditions but not under non-starvation conditions. To emphasize this critical distinction, we have incorporated additional discussion focused specifically on this point.

      “It is well established that starvation induces sleep suppression (MacFadyen, 1973; Thimgan et al., 2010; Melnattur and Shaw, 2019; Keene et al., 2010; He et al., 2020; Yangkyun et al., 2022), and our results are consistent with these previous findings: all genotypes exhibited less sleep under starvation than under fed conditions (i.e. Figures 4A-B, 5A-B and 11). Under normal appetitive memory training, starvation-induced sleep loss does not necessarily impair memory processing (Thimgan et al., 2010; Chouhan et al., 2021). PAM-α1 neuronal activity is higher in starved, trained flies than in fed or untrained flies (data not shown), suggesting that these neurons act as a critical node for integrating internal motivational and arousal states, as well as conveying positive valence for the normal appetitive memory process, independently of starvation-induced sleep loss. While DPM neurons are less sensitive to starvation, they still exhibit training-induced elevated activity (data not shown), indicating coherent responsiveness to upstream signaling. In the present study, we found that under fed conditions, sleep remained intact even when excessive changes in neural activity occurred within the PAM(-α1)-DPM circuit; in contrast, under starvation conditions, significant sleep reduction and fragmentation were observed. These observations indicate that starvation may trigger a transition from a physiologically normal brain state to an unstable, abnormally active state, which consequently elicits negative behavioral outputs.”

      (7) Line 1121: The data points for Figure 9 D-E are surprisingly low considering there are 14 PAMalpha1 labelled, the data presented here indicated potentially only 1-2 neurons were counted per fly brain. Can this contribute to the large variation and the contradiction of PAM's memory-suppressing role?

      We sincerely appreciate the reviewer’s critical comments regarding the sample size of labeled PAM-α1 neurons in Figure 9D–E. We have revisited our raw data, incorporated additional brain samples, and reanalyzed the dataset. For this updated analysis, we included all clearly distinguished neurons, excluded overlapping ones, and calculated a single NLI per brain for statistical analysis. The key conclusions remain consistent with those in the original submission, confirming the robustness of the observed phenotype.

      Memory consolidation is a time-dependent process. To further elucidate the link between neural activity and behavioral outputs, we performed additional experiments with an extended recording period. A detailed response to this point is provided in the response to public review, and we therefore do not reiterate the details here.

      (8) Line 345: the effect size and data spread of THIP restored memory is different from the controls in Figure 10, perhaps warranting a more conservative interpretation of the role of sleep in memory consolidation.

      We appreciate this critical comment. We fully agree that the role of sleep in memory consolidation requires cautious interpretation, a point we have integrated into the revised manuscript.

      Drug treatment in Drosophila, particularly for group-based assays, can introduce substantial variability at both the individual and group levels. To account for this, we employed a statistically valid sample size for our analyses to ensure robust conclusions. While minor quantitative discrepancies exist in the data, this technical consideration does not significantly alter the core conclusions of the study.

      Reviewer #2 (Recommendations for the authors):

      (1) As mentioned in the public review, all data using the broad PAM-DAN driver should be removed. Concerns regarding the experiments involving the broad driver are not included here.

      A detailed response to this point is provided in the response to public review, and we therefore do not reiterate the details here.

      (2) In GCaMP experiments (Figure 9B), the ΔF/F traces for the AHL and AHL+ATP conditions start diverging before the addition of ATP. The quantification shows they are not significantly different in the first 30 seconds, but the fact that in two separate experiments (2A and 9B), they diverge in the same direction makes me wonder whether the AHL condition is different from the +ATP condition even before the ATP treatment. Also, the traces should include standard errors.

      We observed the same diverging trend in the first 30-second baseline as the reviewer. We reviewed the raw data for each sample and found that this divergence is likely attributable a small number of outliers. Given the absence of a statistically significant difference, this divergence does not affect our conclusions.

      We have also added standard errors to the revised figures.

      (3) Figure 9B. The authors need to show data for a control genotype. +>P2X2; VT064246-LexA > GCaMP6f that does not include MB299B-Gal4 is crucial to demonstrate that expression of P2X2 in PAM-α1 is responsible for the inhibitor effect, as LexA-P2X2 may be leaky.

      One of the UAS-P2X2 lines was found to exhibit leaky expression, so we instead used a non-leaky UAS-P2X2 line for all related experiments. To address the reviewer’s comments and further validate our findings, we have added complementary experiments with a control genotype. In addition, we also added a control to confirm the non-leaky expression of LexA-P2X2 under the driver of R58E02-LexA. As shown in revised Figures 8B, application of ATP in the absence of MB299B-GAL4 failed to induce a significant inhibitory effect. These data strongly and convincingly support our conclusion.

      (4) Figure 9B. Some of the individual data show values lower than -100% ΔF/F0. By definition, ΔF/F cannot be less than -100%, as this would require negative fluorescence, which is physically impossible. The calculation of fluorescence changes using ΔF/F should be carefully reconsidered.

      We thank the reviewer pointing out this potential confusion. We used a standard method of calculating the change in fluorescence over time using △F/F = (Fn-F<sub>0</sub>) / F<sub>0</sub>×100% as we previously described (Liu et al., 2019). Changes of greater than +100% of △F/F would not be unusual, since the reported value is a ratio to the initial level of fluorescence, not a subtraction of the baseline value from the signal (which obviously could not go below 100%). We have included a sentence in the results explaining this (page 7): “As previously described, we used the percent change in fluorescence over time as a ratio to the initial level, △F/F = (Fn-F0)/F0×100% for quantification (Liu et al., 2019).” And we have carefully reviewed our raw and processed data and confirmed that our analysis was correct.

      (5) Figure 2B. The number of UAS transgenes should be controlled, as Gal4 could be diluted with 3 UAS constructs in experimental conditions compared to only 1 UAS construct in controls. Are Dop1R2 and DopEcR significantly different from wt? Why do they present an average ΔF/F in 2A and a maximum in 2B?

      We appreciate the reviewer’s careful observations and valuable comments.

      As the reviewer noted, the EPAC imaging experiment utilizes three UAS transgenes, which enable Gal4 enhancement via Dicer, targeted manipulation of dopamine receptor expression levels, and neural activity monitoring in DPM neurons. All other imaging experiments in the study employ only one or two UAS transgenes. Given the robustness of the observed phenotypes, the potential dilution effect is not a major concern. Knockdown of Dop1R2 and DopEcR showed no significant differences relative to the WT control group; the maximum values presented in Fig. 2B are included solely to illustrate statistical significance. While the EPAC (CFP/YPF) signal reflects an obvious cAMP elevation, no differences were detected in the averaged signal across groups.

      Notably, in the revised manuscript, our double-labeling assays confirmed that only Dop1R1 and Dop1R2 are colocalized with DPM neurons (see our responses to the public review from Reviewer #2). Accordingly, we have refined our data analysis to focus exclusively on these two DARs.

      (6) Figures 7H, J. Why is almost every MB299B>TrpA1 fly sleeping at ZT0?

      To align the starvation protocol for sleep analysis with that used in the memory assay, MB299B>TrpA1 flies and their genetic controls were transferred to fresh sleep tubes containing starvation food during the ZT0–1 time window. This transfer resulted in no detectable locomotor activity during this period, a pattern indicative of sleep in all flies.

      (7) The number of episodes and P(wake) should be presented for all sleep data.

      We have added these two parameters as new panels to all relevant sleep figures. The corresponding statistical analyses have also been included in the supplemental tables.

      Reviewer #3 (Recommendations for the authors):

      The study's findings provide compelling insights into the neural circuits connecting sleep and memory and the role of dopamine in general. While the anatomic dissection of the microcircuit and its overall involvement in sleep and memory is convincing and of high interest to the field and beyond, some statements of the study need further clarification, particularly the interpretation of receptor signaling and the role of DPM.

      Major Points

      (1) Figure 2: cAMP Imaging and Dopamine Receptor Involvement

      The authors present calcium and cAMP imaging to support the inhibitory connection between PAM and DPM neurons. While using both sensors is a robust approach, I am not entirely convinced that cAMP imaging is the ideal approach for identifying the dopamine receptors involved. To my knowledge, only Dop1R1 is classically linked to Gs-mediated cAMP signaling. Dop1R2 is typically coupled to Gq (PLC and DAG), while DopEcR is non-canonical and can engage both pathways. Additionally, these receptors are classically excitatory, yet the authors did not analyze Dop2R, the primary inhibitory dopamine receptor - which would represent the most relevant candidate for an inhibitory PAM-DPM connection.

      We have addressed this point in our response to the public comments, so will not reiterate here.

      While dopamine receptor functions can vary by neuronal context, I would appreciate clarification on the following points:

      (a) Why was Dop2R not tested? Was it omitted or found to have no effect?

      We sincerely appreciate the reviewer’s critical questions. This point has been addressed in our response to the public comments. Briefly, Dop2R is not colocalized with DPM neurons; instead, only Dop1R1 and Dop1R2 are detected in DPM neurons, which is why we focused exclusively on these two receptors in the revised manuscript.

      (b) Why was cAMP imaging chosen for receptor identification? Was calcium imaging performed, and if so, what were the results?

      This is an excellent point, and we sincerely appreciate the reviewer’s valuable input, which has helped to strengthen the logical framework of our analysis on receptor-mediated neural activity. These dopamine receptors are well-characterized as members of the Gas-coupled protein receptor family, and cAMP signaling serves as a reliable readout of their functional activity. To strengthen the logic flow of our analysis on the target inhibitory circuit, we have made the following key revisions to the manuscript: 1) defined the expression profile of dopamine receptors in DPM neurons; 2) refined our cAMP imaging data analyses based on specific receptor subtypes; and 3) assessed DPM neural activity via calcium imaging under conditions of targeted receptor knockdown. For further details, please refer to our response to the public comments.

      (c) Since the data suggest multiple receptor involvements and complex interactions, I encourage a more detailed discussion of the working hypothesis, particularly regarding the unexpected finding that classically excitatory receptors contribute to an inhibitory connection.

      We appreciate the suggestion to elaborate on our working model. Accordingly, we have revised the schematic diagram and refined the manuscript to clearly illustrate the underlying mechanistic framework. For further details, please refer to our response to the public comments.

      (2) Figure 3: DPM Involvement in Memory Consolidation

      The authors show that PAM activation and DPM inhibition during consolidation impair appetitive LTM. However, the role of DPM is critical. While the c316-GAL4 driver yields strong effects, VT064246 inhibition shows only slight significance, requiring more than twice the sample size of other experiments. Given that c316-GAL4 is not DPM-specific and also labels MB Kenyon cells, I suggest using MB-GAL80 to restrict expression - or commenting on the possibility that other neurons like MB-KCs could directly participate in the phenotype. This is particularly relevant since VT064246 efficiently modulates sleep, indicating that it is generally effective in altering behavior. These issues weaken the claim that DPM plays a crucial role in linking sleep and memory, and should be addressed. Minor comment on this Figure: In the Figure legend, the driver and "n" are not mentioned for 3C, while this is the case for all other panels. Moreover, the DPM schematic only depicts the MB, making it somewhat confusing. DPM innervates the entire MB, still, it would be helpful to shade the DPM projections more distinctly within the MB for clarity.

      We thank the reviewer for the suggestion to improve the precision of our figures.

      Regarding the expression specificity concern, in all experiments using c316-GAL4, we had eyeless-GAL80 and MB-GAL80 co-expressed to restrict GAL4-driven expression to DPMs. While complete suppression of expression of MB-KCs was not achievable, we largely eliminated the potential confounding effects from majority of these cells. VT064246-GAL4 is known to exhibit weak expression (Jenett et al., 2011; Haynes et al., 2015), but high relative specificity. Importantly, the overall conclusion derived from experiments using c316-GAL4 with GAL80s and VT064246-GAL4 are consistent, which strongly supports the role of DPM neurons in mediating the link between sleep and memory.

      As suggested, we have added sample sizes for all panels and refined the depiction of DPM projections in revised Figure 3C.

      Minor Comments

      (1) Introduction:

      The authors introduce dopamine's role in forgetting but focus on aversive rather than appetitive memories. To avoid confusion, this distinction should be mentioned explicitly (likewise in the discussion). Regarding references: Zhang et al. (line 95) do not discuss DPM or APL. Donlea et al. (line 97) do not cover dopamine - I think Pimentel et al. (2016) would be a more appropriate citation.

      This is a good point. We have removed Zhang et al. (2013) and replaced Donlea et al with Pimentel et al. 2016 as suggested.

      (2) Figure 9: DPM Activation During Consolidation:

      The authors show that PAM neurons are activated by starvation and further enhanced by appetitive training. Surprisingly, DPM neurons also increase activity post-training, despite the proposed inhibitory connection between PAM and DPM. The authors state that "PAM-α1-DPM microcircuit exhibits synchronized neural activity changes during the consolidation window" (line 326), yet they do not address this apparent contradiction. If I have not overlooked key information, this should be clarified/addressed e.g. in the discussion.

      This is an excellent point. We have addressed this in our response to point (3) from Reviewer #2 in the public comments, so we will not reiterate it here.

      (3) Figure 10D/E: THIP Rescue of LTM Deficits:

      Some inconsistencies in the THIP rescue experiments need clarification:

      (a) In Figure 10D, MB299B activation with THIP appears not to significantly restore memory relative to zero, nor to differ from untreated conditions in Figures 7A or 10E.

      (b) In Figure 10E, MB299B activation +/- THIP shows a much clearer effect.

      (c) Are Figures 7A, 10E, and 10D independent experiments, or were they conducted together?

      (d) Should the left bar in 10D and the right bar in 10E be identical? If not, I do not fully understand the discrepancy and suggest discussing the variation.

      Upon revisiting the raw datasets and conducting a one-sample t-test to analyze the group differences, the experimental group in Figure 7A showed no significant difference from the theoretical mean (set at zero). This group also did not differ from the two genetic controls, indicating that the restored memory was comparable to control levels. In Figure 10E, the group with MB299B activation plus THIP treatment exhibited a significant difference from the theoretical mean (one-sample t-test) and from the non-THIP control group, confirming a significant restoration of memory function. Owing to our laboratory relocation, the starvation duration at the new facility was adjusted based on a recalibrated starvation curve; the higher overall 24 h memory index in Figure 10E is likely attributable to a relatively longer starvation period. However, this experimental parameter variation does not alter the study’s overall conclusions.

      (4) Sleep Phenotypes and Starvation Effects:

      Sleep scores are shown under starvation/fed conditions but not under baseline conditions (without inhibition/activation). Could the authors indicate whether they observe basal starvation-induced sleep changes? The authors frequently state that PAM-DPM effects on sleep are context-dependent, yet mild but significant changes occur under fed conditions. I suggest rewording to clarify that the effect is enhanced in a context-dependent manner rather than strictly context-dependent.

      Starvation-induced sleep reduction is a well-characterised phenotype. Our study focused on the key question of whether altered neuronal activity modulates sleep under innate starvation conditions. Accordingly, all comparisons were made between the experimental and control groups under both starvation and fed conditions. We appreciate the reviewer’s suggestion to improve clarity and have revised the text as suggested.

      (5) Starvation Duration in Methods:

      The authors use 30-46h of starvation, which is longer than the ~20h typically used in appetitive memory studies. Could the authors explain why such extended starvation times were necessary?

      Determining starvation levels via survival curves is a well-established and relatively objective method, one that has been widely adopted in prior studies. For the memory test, we standardized the total starvation duration for each genotype to the time point at which mortality reached 20%. Owing to inherent differences in to starvation resistance across distinct genotypes, the final starvation durations ranged from 20 hours to 46 hours.

      (6) Variability in PAM-α1 Sleep Effects:

      (a) The extent and timing of sleep effects differ across PAM-α1 drivers (e.g. night vs. light-period effects). Could MBON co-targeting by these drivers contribute to the variability?

      We have supplemented additional experiments to investigate the effects of MBON-α1 neurons on 24h memory and sleep. For detailed findings, please refer to our response to your public comments.

      (b) Even within the same driver, results differ (e.g., Figure 7H vs. 10A). A general comment on these differences would be important, e.g. regarding the relevance of day and night sleep for memory consolidation.

      We sincerely appreciate the reviewer’s incisive observation regarding these details. The discrepancy stems from the timing of neuronal activity inhibition, during which a laboratory relocation led to adjustments in starvation duration for memory experiments, which in turn indirectly altered sleep patterns.

      (c) Technical note: Similar y-axis scales for sleep plots (Figures 10A and B) would make comparison easier.

      We have unified the y-axis scales to the same range.

      (7) Discussion, Line 376:

      The phrase "sleep deprivation is important for memory consolidation" is misleading, as it could imply that deprivation aids memory formation. Please clarify.

      We appreciate the reviewer’s suggestion. We have revised the text to: “These results demonstrate that preserving unperturbed sleep during the critical memory consolidation window is essential for stabilizing appetitive long-term memory.

    1. eLife Assessment

      This important work examines the effects of gaze on valuation signals in the human brain as participants choose between bundles of sequentially presented items food items. The paper provides convincing analyses of how gaze affects participants choice behaviour and how this varies across time. The work will be of interest to neuroscientists working on attention and decision-making.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have updated the labelling of Figure 7. As all of the reviewer comments have now been addressed, I believe that this version of the manuscript can now be put forward as the Version of Record.]

      Summary:

      This study builds upon a major theoretical account of value-based choice, the 'attentional drift diffusion model' (aDDM), and examines whether and how this might be implemented in the human brain using functional magnetic resonance imaging (fMRI). The aDDM states that the process of internal evidence accumulation across time should be weighted by the decision maker's gaze, with more weight being assigned to the currently fixated item. The present study aims to test whether there are (a) regions of the brain where signals related to the currently presented value are affected by the participant's gaze; (b) regions of the brain where previously accumulated information is weighted by gaze.

      To examine this, the authors developed a novel paradigm that allowed them to dissociate currently and previously presented evidence, at a timescale amenable to measuring neural responses with fMRI. They asked participants to choose between bundles or 'lotteries' of food times, which they revealed sequentially and slowly to the participant across time. This allowed modelling of the haemodynamic response to each new observation in the lottery, separately for previously accumulated and currently presented evidence.

      Using this approach, they find that regions of the brain supporting valuation (vmPFC and ventral striatum) have responses reflecting gaze-weighted valuation of the currently presented item, where as regions previously associated with evidence accumulation (preSMA and IPS) have responses reflected gaze-weighted modulation of previously accumulated evidence.

      A major strength of the current paper is the design of the task, nicely allowing the researchers to examine evidence accumulation across time despite using a technique with poor temporal resolution. The dissociation between currently presented and previously accumulated evidence in different brain regions in GLM1 (before gaze-weighting), as presented in Figure 5, is already compelling. The result that regions such as preSMA response positively to |AV| (absolute difference in accumulated value) is particularly interesting, as it would seem that the 'decision conflict' account of this region's activity might predict the exact opposite result. Additionally, the behaviour has been well modelled at the end of the paper when examining temporal weighting functions across the multiple samples.

      In response to reviewer comments, the authors have explicitly tested for the effects of gaze-weighting over and above any main effect of value, and convincingly shown that these effects are both present in the main regions of interest - namely |SV| and gaze-weighted |SV| in the vmPFC, alongside |AV| and |AV_gaze| in the pre-SMA. This provides clear evidence in support of the notion of gaze-weighting of value signals in these regions.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper the authors seek to disentangle brain areas that encode the subjective value of individual stimuli/items (input regions) from those that accumulate those values into decision variables (integrators) for value-based choice. The authors used a novel task in which stimulus presentation was slowed down to ensure that such a dissociation was possible using fMRI despite its relatively low temporal resolution. In addition, the authors leveraged the fact that gaze increases item value, providing a means of distinguishing brain regions that encode decision variables from those that encode other quantities such as conflict or time-on-task. The authors adopt a region-of-interest approach based on an extensive previous literature and found that the ventral striatum and vmPFC correlated with the item values and not their accumulation whereas the pre-SMA, IPS and dlPFC correlated more strongly with their accumulation. Further analysis revealed that the pre-SMA was the only one of the three integrator regions to also exhibit gaze modulation.

      The study uses a highly innovative design and addresses an important and timely topic. The manuscript is well-written and engaging, while the data analysis appears highly rigorous.

      Weaknesses:

      With 23 subjects the study has relatively low statistical power for fMRI although the within-subjects design and relatively high trial count reduces these concerns.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study builds upon a major theoretical account of value-based choice, the 'attentional drift diffusion model' (aDDM), and examines whether and how this might be implemented in the human brain using functional magnetic resonance imaging (fMRI). The aDDM states that the process of internal evidence accumulation across time should be weighted by the decision maker's gaze, with more weight being assigned to the currently fixated item. The present study aims to test whether there are (a) regions of the brain where signals related to the currently presented value are affected by the participant's gaze; (b) regions of the brain where previously accumulated information is weighted by gaze.

      To examine this, the authors developed a novel paradigm that allowed them to dissociate currently and previously presented evidence, at a timescale amenable to measuring neural responses with fMRI. They asked participants to choose between bundles or 'lotteries' of food times, which they revealed sequentially and slowly to the participant across time. This allowed modelling of the haemodynamic response to each new observation in the lottery, separately for previously accumulated and currently presented evidence.

      Using this approach, they find that regions of the brain supporting valuation (vmPFC and ventral striatum) have responses reflecting gaze-weighted valuation of the currently presented item, where as regions previously associated with evidence accumulation (preSMA and IPS) have responses reflected gaze-weighted modulation of previously accumulated evidence.

      A major strength of the current paper is the design of the task, nicely allowing the researchers to examine evidence accumulation across time despite using a technique with poor temporal resolution. The dissociation between currently presented and previously accumulated evidence in different brain regions in GLM1 (before gazeweighting), as presented in Figure 5, is already compelling. The result that regions such as preSMA response positively to |AV| (absolute difference in accumulated value) is particularly interesting, as it would seem that the 'decision conflict' account of this region's activity might predict the exact opposite result. Additionally, the behaviour has been well modelled at the end of the paper when examining temporal weighting functions across the multiple samples.

      In response to reviewer comments, the authors have explicitly tested for the effects of gaze-weighting over and above any main effect of value, and convincingly shown that these effects are both present in the main regions of interest - namely |SV| and gazeweighted |SV| in the vmPFC, alongside |AV| and |AV_gaze| in the pre-SMA. This provides clear evidence in support of the notion of gaze-weighting of value signals in these regions.

      We thank the reviewer for their comments.

      Reviewer #2 (Public review):

      Summary:

      In this paper the authors seek to disentangle brain areas that encode the subjective value of individual stimuli/items (input regions) from those that accumulate those values into decision variables (integrators) for value-based choice. The authors used a novel task in which stimulus presentation was slowed down to ensure that such a dissociation was possible using fMRI despite its relatively low temporal resolution. In addition, the authors leveraged the fact that gaze increases item value, providing a means of distinguishing brain regions that encode decision variables from those that encode other quantities such as conflict or time-on-task. The authors adopt a region-of-interest approach based on an extensive previous literature and found that the ventral striatum and vmPFC correlated with the item values and not their accumulation whereas the preSMA, IPS and dlPFC correlated more strongly with their accumulation. Further analysis revealed that the pre-SMA was the only one of the three integrator regions to also exhibit gaze modulation.

      The study uses a highly innovative design and addresses an important and timely topic. The manuscript is well-written and engaging, while the data analysis appears highly rigorous.

      Weaknesses:

      With 23 subjects the study has relatively low statistical power for fMRI although the within-subjects design and relatively high trial count reduces these concerns.

      We thank the reviewer for their comments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Something seems to have gone slightly wrong in (I think) the labelling of new figure 7 (the correlation matrix between the different regressors). There are five variables on both the x- and y-axes of the figure, and they are the same five variables - meaning the diagonal of the matrix would be all equal to 1 (being the correlation of a regressor with itself - e.g. |AVgaze| with |AVgaze|). In the figure legend, six variables are mentioned, including lagged |deltaAVgaze| - but this doesn't appear on the x or y-axes. I suspect that the authors may need to check that this matrix has been calculated correctly, and isn't being mislabelled?

      We thank the reviewer for noticing this issue. We have now corrected Figure 7.

    1. eLife Assessment

      This important study reports the development of the first tankyrase degrader and demonstrates its enhanced ability to inhibit β-catenin signaling compared to conventional tankyrase inhibitors. The evidence supporting the conclusions is comprehensive and convincing, based on rigorous biochemical and cellular analyses. The findings will be of broad interest to researchers studying Wnt signaling, protein degradation, and cancer biology.

    2. Reviewer #1 (Public review):

      [Editors' note: the second round of revision has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the minor comments raised in the previous round of review.]

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation-thereby impairing β-catenin degradation-the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Comments on previous version:

      I had a favorable opinion of the manuscript in the first round of review. I don't have additional comments on the revised manuscript. The manuscript looks fine to me.

    3. Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Comments on previous version:

      I thank the authors for responding to the queries raised in the original review, most of which have now been addressed. This further strengthens this well-conducted study and well-presented manuscript. I congratulate the authors for this interesting and insightful work.

    4. Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

    5. Reviewer #4 (Public review):

      From the Reviewing Editor:

      This important study reports the development of the first PROTACs targeting the ADP-ribosyltransferases tankyrase 1 and 2, with the goal of inhibiting Wnt/β-catenin signaling more completely than is possible with catalytic tankyrase inhibitors. The work addresses a significant limitation of existing tankyrase inhibitors: although catalytic inhibition stabilizes AXIN1/2 and suppresses Wnt signaling, it also stabilizes tankyrase itself, potentially enhancing non-catalytic scaffolding functions and promoting accumulation of degradasome-like puncta.

      The evidence is convincing. The authors use appropriate and well-validated approaches, including chemical biology, cellular assays, and proteomic profiling, to show that PROTAC-mediated degradation of tankyrase avoids tankyrase accumulation while still stabilizing AXIN and inhibiting Wnt/β-catenin signaling. The data support the conclusion that degradation of tankyrase can separate pathway inhibition from the confounding effects of stabilized tankyrase protein and may therefore offer advantages over conventional catalytic inhibitors.

      A strength of the study is the clear mechanistic comparison between tankyrase degradation and catalytic inhibition. The manuscript provides convincing evidence that the PROTAC and catalytic inhibitors act through distinct mechanisms, with the PROTAC targeting both catalytic and scaffolding roles of tankyrase. The study is well conducted and clearly presented, and the authors have addressed most concerns raised during review.

      A remaining limitation is that the therapeutic potential of the compound is not tested in vivo, for example in APC-mutant colorectal cancer models, APCmin mice, or patient-derived xenografts. Such experiments would strengthen claims about practical efficacy, although they are not essential for the main mechanistic conclusions of the manuscript.

      Overall, this is an important and insightful contribution. It advances the tankyrase and Wnt signaling fields by providing a new chemical strategy to suppress tankyrase function more completely than catalytic inhibition alone, and it offers a useful framework for future therapeutic exploration of tankyrase degradation.

    6. Author response:

      The following is the authors’ response to the previous reviews

      We thank the Reviewers for the favourable feedback. There is no additional comments from Reviewers 1, 3, and 4, and we address the minor concerns from Reviewer 2 as follows.

      (1) I appreciate the authors acknowledge that testing the physical properties of the degradasome puncta is necessary to explore whether they indeed represent condensates. The term "condensates" implies liquid-liquid phase separation (rightly or wrongly). However, this question has not yet been resolved in the case of degradasomes. I therefore suggest the term "condensates" to be avoided. A simple morphological description as "puncta" may suffice.

      We have revised our manuscript to state: “The DC has been proposed to exist as biomolecular condensates.” Additionally, we have included a time-lapse image showing that AXIN1-GFP puncta exhibit dynamic fusion behaviour in cells (Fig. S8), suggesting that the DC may be liquid-like, at least with AXIN1 overexpression.

      (2) I thank the authors for including the additional data comparing tankyrase binding by IWR and IWRPOMA. I agree that using the BRET signal of IWR-POMA is informative. Adding the IC<sub>50</sub> values directly to the figure panels (S3E, S3G) would help the reader to quickly assess binding. The comparison between these two panels is insightful.

      Added.

      (3) Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue? Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue?

      We now specify in the Introduction that TNKS1/2 are encoded by TNKS/TNKS2, and are collectively referred to as TNKS in this manuscript.

    1. eLife Assessment

      This revised paper provides solid evidence for HGF-induced trafficking of the HGF receptor, MET, together with metalloprotease MT1-MMP into invadopodia and the role of this trafficking in triple-negative breast cancer (TNBC) cell invasion in vitro. The evidence could be improved with additional experimental controls and analysis, but the findings will be valuable for cancer cell biologists investigating TNBC.

    2. Reviewer #1 (Public review):

      Summary:

      This study identifies a mechanism responsible for the accumulation of the MET receptor in invadopodia, following stimulation of Triple-negative breast cancer (TNBC) cells with HGF. HGF-driven accumulation and activation of MET in invadopodia causes the degradation of the extracellular matrix promoting cancer cell invasion, a process here investigated using gelatine-degradation and spheroid invasion assays.

      Mechanistically, HGF stimulates the recycling of MET from RAB14-positive endodomes to invadopodia, increasing their formation. At invadopodia, MET induces matrix degradation via direct binding with the metallo protease MT1-MMP.

      The delivery of MET from the recycling compartment to invadopodia is mediated by RCP which facilitates the colocalization of MET to RAB14 endosomes. On this compartment, HGF induces the recruitment of the motor protein KIF16B promoting the tubulation of the RAB14-MET recycling endosomes to the cell surface.

      This pathway is critical for the HGF-driven invasive properties of TNBC cells as it is impaired upon silencing of RAB14.

      Strengths:

      The study is well organized and executed using state of the art technology. The effects of MET recycling in the formation of functional invadopodia are carefully studied taking advantage of mutant forms of the receptor that are degradation-resistant or endocytosis-defective.

      Data analyses are rigorous and appropriate controls are used in most of the assays to assess the specificity of the scored effects. Overall, the quality of the research is high.<br /> The conclusions are well supported by the results and the data and methodology are of interest for a wide audience of cell biologists.

      Previous Weaknesses:

      The role of the MET receptor in invadopodia formation and cancer cell dissemination has been intensively studied in many settings including Triple Negative breast cancer cells. The novelty of the present study mostly consists in the detailed molecular description of the underlying mechanism based on HGF-driven MET recycling. The question of whether the identified pathway is specific for TNBC cells or represents a general mechanism of HGF-mediated invasion detectable in other cancer cells is not addressed or at least discussed.

      Comments on revised version:

      The authors have partially replied to my previous concerns.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Khamari and colleagues investigate how HGF-MET signaling and the intracellular trafficking of the MET receptor tyrosine kinase influence invadopodia formation and invasion in triple-negative breast cancer (TNBC) cells. They show that HGF stimulation enhances both the number of invadopodia and their proteolytic activity. Mechanistically, the authors demonstrate that HGF-induced, RAB4- and RCP-RAB14-KIF16B-dependent recycling routes deliver MET to the cell surface specifically at sites where invadopodia form. Moreover, they report that MET physically interacts with MT1-MMP - a key transmembrane metalloproteinase required for invadopodia function- and that these two proteins co-traffic to invadopodia upon HGF stimulation.

      Although the HGF-MET axis has previously been implicated in invadopodia regulation (e.g., by Rajadurai et al., Journal of Cell Science 2012), studies directly linking ligand-induced MET trafficking with the spatial regulation of MT1-MMP localization and activity have been lacking.

      Overall, the manuscript addresses a relevant and timely topic and provides several novel insights.

      Comments on revised version:

      I appreciate the authors' efforts to revise the manuscript and address the reviewers' comments. While the revised version includes additional experiments and several improvements in data presentation, the major methodological and conceptual concerns raised in the initial review remain largely unresolved. In my opinion, these issues critically undermine the central mechanistic conclusions of the study.

      (1) Inappropriate experimental design for studying MET trafficking

      A major concern remains the use of prolonged HGF stimulation times (2-6 hours) to study MET endocytosis and recycling. This is not an appropriate experimental design for investigating receptor tyrosine kinase trafficking dynamics. Ligand-induced internalization of MET occurs within minutes, with maximal endosomal accumulation typically observed within 5-15 minutes, whereas recycling occurs over approximately 15-60 minutes.

      Importantly, the authors have not included short stimulation time points or any kinetic analysis that would allow a proper assessment of MET internalization or recycling. The additional surface biotinylation experiment does not address this issue, as it still does not provide temporal information regarding receptor trafficking.

      Therefore, the current data do not support the conclusions regarding MET endocytosis or recycling, and this major methodological concern has not been adequately addressed in the revised manuscript.

      (2) Insufficient validation of antibody specificity in immunofluorescence

      The validation of antibody specificity for MET, phospho-MET, and MT1-MMP in immunofluorescence experiments remains insufficient. While the authors demonstrate knockdown efficiency by immunoblotting and show some reduction in fluorescence signal, they do not provide rigorous evidence that the immunofluorescence signal is specifically abolished upon gene silencing under identical imaging conditions. Such validation is essential, particularly because the manuscript relies heavily on imaging-based localization and colocalization analyses. Without these controls, it cannot be excluded that the observed signal represents non-specific staining.

      Importantly, the authors attempt to justify antibody specificity primarily by citing previous publications that used the same antibodies. However, this is not an adequate substitute for experimental validation within the current study. Previous reports do not guarantee specificity under the present experimental conditions, particularly in immunofluorescence, where staining patterns can be strongly influenced by fixation procedures, antibody concentrations, imaging settings, and cell type. Moreover, those studies may themselves lack sufficiently rigorous validation of antibody specificity. Therefore, antibody specificity should be demonstrated directly in the experimental system used in this manuscript, especially given that the principal conclusions rely extensively on the subcellular localization of MET, phospho-MET, and MT1-MMP.

      (3) Questionable MET localization in TIRF microscopy

      The presence of punctate MET signal in TIRF microscopy under unstimulated conditions raises additional concerns. Under basal conditions, MET is generally expected to exhibit a predominantly diffuse distribution at the plasma membrane, whereas prominent punctate structures are typically associated with ligand-induced clustering, endocytosis, or trafficking events.

      The observation of numerous MET-positive puncta in unstimulated cells, together with the insufficient validation of antibody specificity, raises the possibility that at least part of the observed signal represents non-specific staining or imaging artefacts rather than bona fide MET localization. This concern is further compounded by the lack of rigorous immunofluorescence antibody validation discussed above and significantly undermines the interpretation of all TIRF-based trafficking analyses presented in the manuscript.

      (4) The evidence supporting a MET-specific role in invadopodia remains unconvincing

      The authors argue that the role of MET in invadopodia formation is validated using three independent approaches: shRNA-mediated knockdown, SMARTpool siRNA-mediated knockdown, and pharmacological inhibition with PHA665752. However, I do not agree that these constitute three independent orthogonal validations of MET function.

      First, the shRNA-mediated knockdown presented in this study achieves only modest depletion of MET protein. The authors themselves acknowledge this limitation and therefore selected cells with visibly reduced MET staining for imaging. Consequently, the shRNA experiments cannot be considered a robust or independent validation of MET function.

      Second, although pooled SMARTpool siRNAs are widely used to improve knockdown efficiency, they cannot exclude off-target effects, as each individual guide RNA contributes its own potential off-target profile. Therefore, pooled siRNAs cannot by themselves establish that an observed phenotype is specifically attributable to depletion of the intended target and do not replace validation using independent individual siRNAs or rescue experiments.

      Third, the pharmacological data should also be interpreted with caution. Throughout the manuscript, PHA665752 is presented as a MET inhibitor supporting the specificity of the observed phenotype. However, there is essentially no such thing as a truly selective receptor tyrosine kinase inhibitor. PHA665752 inhibits multiple kinases in addition to MET, particularly at concentrations commonly used in cell-based assays. Consequently, the inhibitor cannot be considered an independent validation of MET-specific function.

      Importantly, the newly added siRNA experiments do not resolve my original concern regarding the role of MET in invadopodia formation. Although siRNA-mediated MET depletion is substantially more efficient than the shRNA-mediated knockdown presented in the original manuscript, this marked difference in MET depletion is not accompanied by a correspondingly stronger inhibition of invadopodia formation or ECM degradation. If MET were indeed the principal driver of the observed phenotype, one would expect the magnitude of the biological effect to correlate with the efficiency of MET depletion. This inconsistency raises the possibility that the observed phenotype is not solely attributable to MET depletion and calls into question the specificity of the proposed mechanism.

      Taken together, the three perturbation approaches used by the authors cannot be regarded as independent orthogonal validation of MET function. One approach provides only modest target depletion, another relies on pooled RNAi reagents that cannot exclude off-target effects, and the third employs a multi-kinase inhibitor rather than a MET-specific compound. Collectively, these limitations substantially weaken the conclusion that the reduction in invadopodia formation is specifically attributable to loss of MET. A convincing demonstration of MET-specific function would require rescue experiments or another truly orthogonal validation strategy.

      (5) Weak evidence for MET-MT1-MMP interaction

      The evidence supporting a physical interaction between MET and MT1-MMP remains unconvincing. The newly added co-immunoprecipitation experiment does not reveal a convincing MET-MT1-MMP interaction, and I am unable to appreciate a specific co-immunoprecipitated MT1-MMP signal in the presented blot. As presented, these data do not convincingly demonstrate a specific or functionally relevant interaction. Given that this interaction constitutes a central component of the proposed mechanistic model, this remains a major weakness of the study.

      (6) Overinterpretation of the data

      Taken together, the study proposes a mechanistic model linking MET trafficking to MT1-MMP localization and invadopodia function. However, the experimental evidence largely supports correlative observations rather than demonstrating a direct mechanistic relationship.

      Specifically, MET endocytosis and recycling are not properly demonstrated because of the inappropriate temporal resolution of the trafficking experiments; the localization data remain uncertain owing to insufficient validation of the immunofluorescence reagents; and the proposed interaction between MET and MT1-MMP is not convincingly demonstrated. Consequently, the manuscript establishes correlation rather than causality, and the central mechanistic conclusions appear to be substantially overstated relative to the presented data.

      Conclusion:

      While the manuscript addresses an interesting and biologically relevant question, the current experimental evidence does not adequately support the proposed mechanistic model. The combination of inappropriate experimental design for trafficking studies, insufficient validation of key imaging reagents, questionable interpretation of the localization data, lack of convincing evidence for the proposed MET-MT1-MMP interaction, and the absence of a clear relationship between the degree of MET depletion and the biological phenotype substantially limits the reliability of the conclusions.

      In my opinion, these issues cannot be addressed by further revision of the current manuscript, as they require substantial additional experimentation, including appropriately designed trafficking assays with short kinetic time points, rigorous validation of antibody specificity for immunofluorescence, and stronger mechanistic evidence linking MET trafficking to MT1-MMP-dependent invadopodia function.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editor and the reviewers for their comments on the revised manuscript. Based on the comments, we decided to go for a minor revision which will address all the comments of reviewer 1.

      Towards the comments of the reviewer 2, we would like to state that we already provided the results from the new experiments and the reasons why we did not perform some of the suggested ones. Interestingly, we have not deviated from standard practices in the field in our approaches. Yet, the reviewer is not convinced and raised concern about the robustness of our observations. We therefore decided to carry out a few more control experiments which in our opinion are redundant as they already were carried out multiple times by us as well as by the other field experts under identical conditions and using the identical cell lines.

      Reviewer #1 (Public review):

      Summary:

      This study identifies a mechanism responsible for the accumulation of the MET receptor in invadopodia, following stimulation of Triple-negative breast cancer (TNBC) cells with HGF. HGF-driven accumulation and activation of MET in invadopodia causes the degradation of the extracellular matrix promoting cancer cell invasion, a process here investigated using gelatine-degradation and spheroid invasion assays.

      Mechanistically, HGF stimulates the recycling of MET from RAB14-positive endodomes to invadopodia, increasing their formation. At invadopodia, MET induces matrix degradation via direct binding with the metallo protease MT1-MMP.

      The delivery of MET from the recycling compartment to invadopodia is mediated by RCP which facilitates the colocalization of MET to RAB14 endosomes. On this compartment, HGF induces the recruitment of the motor protein KIF16B promoting the tubulation of the RAB14-MET recycling endosomes to the cell surface.

      This pathway is critical for the HGF-driven invasive properties of TNBC cells as it is impaired upon silencing of RAB14.

      Strengths:

      The study is well organized and executed using state of the art technology. The effects of MET recycling in the formation of functional invadopodia are carefully studied taking advantage of mutant forms of the receptor that are degradation-resistant or endocytosisdefective.

      Data analyses are rigorous and appropriate controls are used in most of the assays to assess the specificity of the scored effects. Overall, the quality of the research is high.

      The conclusions are well supported by the results and the data and methodology are of interest for a wide audience of cell biologists.

      Previous Weaknesses:

      The role of the MET receptor in invadopodia formation and cancer cell dissemination has been intensively studied in many settings including Triple Negative breast cancer cells. The novelty of the present study mostly consists in the detailed molecular description of the underlying mechanism based on HGF-driven MET recycling. The question of whether the identified pathway is specific for TNBC cells or represents a general mechanism of HGFmediated invasion detectable in other cancer cells is not addressed or at least discussed.

      Comments on revised version:

      The authors have partially replied to my previous concerns.

      We sincerely thank the reviewer for careful evaluation of our manuscript and recognizing the strength of our study. We are grateful for the positive assessment about the well-executed methods, rigorous data analysis, usage of appropriate controls, and for acknowledging that we have addressed, at least in part, the concerns raised in the previous round of review. We appreciate the reviewer’s constructive comments, which have helped us further clarify the scope and significance of our findings.

      Reviewer #1 (Recommendations for the authors):

      The authors have partially replied to my previous concerns. The following points still have to be addressed.

      (1) Despite many TNBC tumours present high expression of the EGFR, trials with EGFR inhibitors have been very disappointing (please read PMID: 41651315). EGFR inhibition has been extensively attempted with negative outcome, and this is well known.

      In the clinical practise, Triple Negative breast cancer patients are commonly treated with chemotherapy, not with EGFR or MET inhibitors.

      Line 31-38 are misleading at best and should be removed and the incipit of the study modified. The biology of this study is sound there is no need to push the clinical relevance with instances that are notoriously not applicable.

      We are thankful to the reviewer for bringing up this point. We will modify the manuscript as per the reviewer’s suggestion.

      (2) In reply to point 4, the authors claim that they checked MMP2 and the experiment is shown in FigS5I, which is not......

      We sincerely apologise for uploading an incorrect file. The correct file will be uploaded.

      (3) I noticed that, in the previous version of the manuscript in Fig. S4A the panel showing the mutation frequency was erroneously indicated in the legend as referring to MET.

      The authors replied that they fixed this but actually the legend is still wrong...

      We thank the reviewer for bringing this error into our attention. We will attentively correct the legend in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Khamari and colleagues investigate how HGF-MET signaling and the intracellular trafficking of the MET receptor tyrosine kinase influence invadopodia formation and invasion in triple-negative breast cancer (TNBC) cells. They show that HGF stimulation enhances both the number of invadopodia and their proteolytic activity. Mechanistically, the authors demonstrate that HGF-induced, RAB4- and RCP-RAB14KIF16B-dependent recycling routes deliver MET to the cell surface specifically at sites where invadopodia form. Moreover, they report that MET physically interacts with MT1MMP - a key transmembrane metalloproteinase required for invadopodia function- and that these two proteins co-traffic to invadopodia upon HGF stimulation.

      Although the HGF-MET axis has previously been implicated in invadopodia regulation (e.g., by Rajadurai et al., Journal of Cell Science 2012), studies directly linking ligandinduced MET trafficking with the spatial regulation of MT1-MMP localization and activity have been lacking.

      Overall, the manuscript addresses a relevant and timely topic and provides several novel insights.

      Comments on revised version:

      I appreciate the authors' efforts to revise the manuscript and address the reviewers' comments. While the revised version includes additional experiments and several improvements in data presentation, the major methodological and conceptual concerns raised in the initial review remain largely unresolved. In my opinion, these issues critically undermine the central mechanistic conclusions of the study.

      We thank the reviewer for critically re-evaluating our manuscript and acknowledging the novel insights and relevance of our study.

      We thank the reviewer for pointing out the study by Rajadurai et al., Journal of Cell Science, 2012, which provided crucial evidence about the role of MET signaling in invadopodia formation [1]. However, the experimental system largely used by Rajadurai et al. is fundamentally different from the receptor trafficking mechanism investigated in the present study. The group have mostly used overexpression of Tpr-MET, which is a cytosolic MET mutant, that does not undergo the canonical ligand-induced RTK endocytosis and subsequent degradation or recycling. Our study did not only establish another link between MET signaling and invadopodia formation; rather, we identified a trafficking-dependent mechanism whereby HGF stimulation regulates the spatial redistribution and recycling of full-length MET to invadopodia, thus providing more physiologically relevant insights.

      We also respectfully disagree that the methodological concerns raised critically undermine our mechanistic conclusions. Although other approaches as suggested by the reviewer could provide complementary information, we believe that the methods used in our study are appropriate for assessing invasive behaviour of the TNBC cells. These methods have been used by us and other research groups in the filed as reflected from the existing literature [2–6].

      In this context, we would also like to add that the study referred by the reviewer above, has used a more off target prone approach (SiGenome Smartpool) compared to ONTARGETplus (chemically modified SiRNA pool for minimizing off target effect) in addition to the same small molecule inhibitor used in our study. Additionally, we also used a shRNA-based silencing to verify the phenotype. Since the silencing was not as pronounced as the siRNA-mediated knockdown, the reviewer has expressed concern.

      We would like to point out that the antibody we used in IF for MT1-MMP (MMP14) have been published in multiple peer-reviewed journals by various research groups using the identical cell lines (MDA-MB-231, ATCC- HTB-26) [6–8]. MET antibodies used in the study (CST, D1C2 XP & L6E7) has been validated in MDA-MB-231 and MET-depleted cells [9-12]. Since these antibodies have been utilized for IF since a long time across various research groups under identical laboratory/experimental conditions, the exercise of validation suggested by the reviewer is surprising.

      (1) Inappropriate experimental design for studying MET trafficking

      A major concern remains the use of prolonged HGF stimulation times (2-6 hours) to study MET endocytosis and recycling. This is not an appropriate experimental design for investigating receptor tyrosine kinase trafficking dynamics. Ligand-induced internalization of MET occurs within minutes, with maximal endosomal accumulation typically observed within 5-15 minutes, whereas recycling occurs over approximately 15-60 minutes.

      Importantly, the authors have not included short stimulation time points or any kinetic analysis that would allow a proper assessment of MET internalization or recycling. The additional surface biotinylation experiment does not address this issue, as it still does not provide temporal information regarding receptor trafficking.

      Therefore, the current data do not support the conclusions regarding MET endocytosis or recycling, and this major methodological concern has not been adequately addressed in the revised manuscript.

      We want to clarify that, our prime objective is to determine how MET trafficking is regulated at the time points at which we observe the functional effects of HGF on invadopodia formation and matrix degradation. Since our functional assays were performed following 2-3 h of HGF stimulation, we specifically examined MET localization and trafficking at these same time points.

      Though shorter time points could provide information on the kinetics of MET trafficking, but their absence does not invalidate our conclusions regarding the role of MET trafficking in HGF-induced invasive function. In other words, our conclusions are made for the time points for which we have conducted the experiments. It is needless to add that different cargo molecule will show different kinetics.

      In summary, we wanted to study the MET trafficking at the late hours in accordance with our functional assays and accordingly designed our experiments. Also, we have supported our results through biochemical methods which is considered to be one of the gold standards in the field.

      (2) Insufficient validation of antibody specificity in immunofluorescence

      The validation of antibody specificity for MET, phospho-MET, and MT1-MMP in immunofluorescence experiments remains insufficient. While the authors demonstrate knockdown efficiency by immunoblotting and show some reduction in fluorescence signal, they do not provide rigorous evidence that the immunofluorescence signal is specifically abolished upon gene silencing under identical imaging conditions. Such validation is essential, particularly because the manuscript relies heavily on imaging-based localization and colocalization analyses. Without these controls, it cannot be excluded that the observed signal represents non-specific staining.

      Importantly, the authors attempt to justify antibody specificity primarily by citing previous publications that used the same antibodies. However, this is not an adequate substitute for experimental validation within the current study. Previous reports do not guarantee specificity under the present experimental conditions, particularly in immunofluorescence, where staining patterns can be strongly influenced by fixation procedures, antibody concentrations, imaging settings, and cell type. Moreover, those studies may themselves lack sufficiently rigorous validation of antibody specificity. Therefore, antibody specificity should be demonstrated directly in the experimental system used in this manuscript, especially given that the principal conclusions rely extensively on the subcellular localization of MET, phospho-MET, and MT1-MMP.

      We understand the reviewer’s concern regarding antibody specificity and agree that appropriate validation is important for imaging-based analyses. However, we strongly disagree with the statement that the MT1-MMP, MET antibody were not adequately validated. The specificity of the MT1-MMP antibody was independently validated by both siRNA- and sgRNA-mediated gene silencing, where we observed a substantially diminished MT1-MMP signal by immunoblotting (Fig S5I, M’). Moreover, the antibody has been used for IF in the same cell line by multiple research groups [6–8]. So, in our opinion, this validation is completely redundant.

      We have validated the MET staining/ signal using the antibody in gene-silenced cells by immunoblotting and immunofluorescence (Fig S1I, K, L), as also acknowledged by the reviewer in comment-4. In addition, we would also like to clarify that the references cited in support of antibody specificity were not selected simply because they used the same antibodies. They include studies that provide experimental validation of the antibodies by gene silencing.

      To further confirm the antibody specificity, we will add immunofluorescence images of MET or MT1-MMP silenced cells stained with respective antibodies. However, we may not want to add these data to the manuscript as they do not carry any additional values to the manuscript.

      (3) Questionable MET localization in TIRF microscopy

      The presence of punctate MET signal in TIRF microscopy under unstimulated conditions raises additional concerns. Under basal conditions, MET is generally expected to exhibit a predominantly diffuse distribution at the plasma membrane, whereas prominent punctate structures are typically associated with ligand-induced clustering, endocytosis, or trafficking events.

      The observation of numerous MET-positive puncta in unstimulated cells, together with the insufficient validation of antibody specificity, raises the possibility that at least part of the observed signal represents non-specific staining or imaging artefacts rather than bona fide MET localization. This concern is further compounded by the lack of rigorous immunofluorescence antibody validation discussed above and significantly undermines the interpretation of all TIRF-based trafficking analyses presented in the manuscript.

      We would like to highlight the apparent similarities between Fig 2A, B and the unstimulated condition in Fig. 2H. In figure 2A, B, MET is detected using an anti-MET antibody, whereas in Figure 2H, GFP-MET is imaged under live cell condition. We believe the reviewer would agree that imaging GFP-MET in live cells avoids fixation- and antibody-related artifacts. The comparable localization observed using these two independent approaches therefore provides additional support that the MET distribution shown in Fig. 2A, B reflects genuine receptor localization rather than an imaging or staining artefact.

      (4) The evidence supporting a MET-specific role in invadopodia remains unconvincing

      The authors argue that the role of MET in invadopodia formation is validated using three independent approaches: shRNA-mediated knockdown, SMARTpool siRNA-mediated knockdown, and pharmacological inhibition with PHA665752. However, I do not agree that these constitute three independent orthogonal validations of MET function.

      First, the shRNA-mediated knockdown presented in this study achieves only modest depletion of MET protein. The authors themselves acknowledge this limitation and therefore selected cells with visibly reduced MET staining for imaging. Consequently, the shRNA experiments cannot be considered a robust or independent validation of MET function.

      Second, although pooled SMARTpool siRNAs are widely used to improve knockdown efficiency, they cannot exclude off-target effects, as each individual guide RNA contributes its own potential off-target profile. Therefore, pooled siRNAs cannot by themselves establish that an observed phenotype is specifically attributable to depletion of the intended target and do not replace validation using independent individual siRNAs or rescue experiments.

      Third, the pharmacological data should also be interpreted with caution. Throughout the manuscript, PHA665752 is presented as a MET inhibitor supporting the specificity of the observed phenotype. However, there is essentially no such thing as a truly selective receptor tyrosine kinase inhibitor. PHA665752 inhibits multiple kinases in addition to MET, particularly at concentrations commonly used in cell-based assays. Consequently, the inhibitor cannot be considered an independent validation of MET-specific function.

      Importantly, the newly added siRNA experiments do not resolve my original concern regarding the role of MET in invadopodia formation. Although siRNA-mediated MET depletion is substantially more efficient than the shRNA-mediated knockdown presented in the original manuscript, this marked difference in MET depletion is not accompanied by a correspondingly stronger inhibition of invadopodia formation or ECM degradation. If MET were indeed the principal driver of the observed phenotype, one would expect the magnitude of the biological effect to correlate with the efficiency of MET depletion. This inconsistency raises the possibility that the observed phenotype is not solely attributable to MET depletion and calls into question the specificity of the proposed mechanism.

      Taken together, the three perturbation approaches used by the authors cannot be regarded as independent orthogonal validation of MET function. One approach provides only modest target depletion, another relies on pooled RNAi reagents that cannot exclude off-target effects, and the third employs a multi-kinase inhibitor rather than a MET-specific compound. Collectively, these limitations substantially weaken the conclusion that the reduction in invadopodia formation is specifically attributable to loss of MET. A convincing demonstration of MET-specific function would require rescue experiments or another truly orthogonal validation strategy.

      We had adopted three independent approaches to validate the phenotype. All three approaches are well practiced in the field. The small molecule inhibitor has been used in the study by Rajadurai et al, J Cell Science, 2012 and it is very much accepted in studying cellular kinases [1].

      We agree that even though the smart pool has always chance of off-target effects, the OnTargetPlus Smart pool has the minimum chance of off-target effects because of the patented chemical modifications, compared to the individual oligos and SiGenome SMARTpool, which was used by Rajadurai, et. al. in their study [1].

      The shRNA mediated silencing resulted in less reduction in the MET level (~50-60%) but is it scientifically not acceptable, particularly when it showed similar phenotype over n=3 sets of experiments?

      The arguments made by the reviewer in this context seems to be harsh. However, we decided to carry out MET silencing using two independent oligos from the SMART pool.

      (5) Weak evidence for MET-MT1-MMP interaction

      The evidence supporting a physical interaction between MET and MT1-MMP remains unconvincing. The newly added co-immunoprecipitation experiment does not reveal a convincing MET-MT1-MMP interaction, and I am unable to appreciate a specific coimmunoprecipitated MT1-MMP signal in the presented blot. As presented, these data do not convincingly demonstrate a specific or functionally relevant interaction. Given that this interaction constitutes a central component of the proposed mechanistic model, this remains a major weakness of the study.

      We have detected the interaction in both GFP pulldown assay in 4 different cell lines and the corresponding reverse His-Ni-NTA pulldown assay, providing complementary evidence for their physical association (Fig 6E, S5F, G). We have also clearly stated in the manuscript that this interaction is weak in nature and have not claimed it to be a strong interaction. While we acknowledge that the signal is modest, disregarding reproducible positive results would not be an appropriate interpretation of the pulldown assays. We believe the reproducibility of these findings supports a genuine MET and MT1-MMP association, and we have reported it accordingly.

      (6) Overinterpretation of the data

      Taken together, the study proposes a mechanistic model linking MET trafficking to MT1MMP localization and invadopodia function. However, the experimental evidence largely supports correlative observations rather than demonstrating a direct mechanistic relationship.

      Specifically, MET endocytosis and recycling are not properly demonstrated because of the inappropriate temporal resolution of the trafficking experiments; the localization data remain uncertain owing to insufficient validation of the immunofluorescence reagents; and the proposed interaction between MET and MT1-MMP is not convincingly demonstrated. Consequently, the manuscript establishes correlation rather than causality, and the central mechanistic conclusions appear to be substantially overstated relative to the presented data.

      We agree that there is scope for further investigation of MET and MT1-MMP cotrafficking, however, we have provided preliminary evidence demonstrating the cotrafficking of MET and MT1-MMP at the cell surface (Fig 6F). Importantly, MET and MT1-MMP co-trafficking represents only one component of the manuscript and not a central mechanistic conclusion. As the title of the study indicates, the major component of the study is focused on MET trafficking and its implication in invadopodia-associated TNBC invasion, for which we have provided direct experimental evidence. Therefore, we believe that describing the overall study as primarily overstated and correlative underestimates the extent of the experimental evidence supporting our mechanistic conclusions.

      Conclusion:

      While the manuscript addresses an interesting and biologically relevant question, the current experimental evidence does not adequately support the proposed mechanistic model. The combination of inappropriate experimental design for trafficking studies, insufficient validation of key imaging reagents, questionable interpretation of the localization data, lack of convincing evidence for the proposed MET-MT1-MMP interaction, and the absence of a clear relationship between the degree of MET depletion and the biological phenotype substantially limits the reliability of the conclusions.

      In my opinion, these issues cannot be addressed by further revision of the current manuscript, as they require substantial additional experimentation, including appropriately designed trafficking assays with short kinetic time points, rigorous validation of antibody specificity for immunofluorescence, and stronger mechanistic evidence linking MET trafficking to MT1-MMP-dependent invadopodia function.

      We thank the reviewer for outlining the remaining concerns. We will address the points raised by performing additional antibody validation and independent oligo-mediated MET silencing experiment, providing further support for the specificity and robustness of our findings.

      However, we respectfully disagree, that the conclusions require the extensive additional experimentation suggested by the reviewer. As clarified above, our trafficking experiments were designed around the time points at which the functional invasive phenotype is observed, rather than to define the kinetics of MET internalization.

      In conclusion, we would expect that the views of the reviewer 2 towards the manuscript should change and the reliability of our manuscript to the public should improve.

      References:

      (1) Rajadurai CV, Havrylov S, Zaoui K, Vaillancourt R, Stuible M, Naujokas M, Zuo D, Tremblay ML, Park M. Met receptor tyrosine kinase signals through a cortactinGab1 scaffold complex, to mediate invadopodia. J Cell Sci. 2012 Jun 15;125(Pt 12):2940-53. doi: 10.1242/jcs.100834. Epub 2012 Feb 24. PMID: 22366451; PMCID: PMC3434810.

      (2) Sharma P, Parveen S, Vinod Shah L, Mukherjee M, Kalaidzidis Y, Joseph Kozielski A, et al. Title: SNX27-retromer assembly directs MT1-MMP trafficking to invadopodia and promotes breast cancer metastasis.

      (3) Mader CC, Oser M, Magalhaes MAO, Bravo-Cordero JJ, Condeelis J, Koleske AJ, et al. Molecular and Cellular Pathobiology An EGFR-Src-Arg-Cortactin Pathway Mediates Functional Maturation of Invadopodia and Breast Cancer Cell Invasion [Internet]. doi:10.1158/0008-5472.CAN-10-1432

      (4) Parveen S, Khamari A, Raju J, Coppolino MG, Datta S. Syntaxin 7 contributes to breast cancer cell invasion by promoting invadopodia formation. J Cell Sci. 2022 Jun 15;135(12). doi:10.1242/jcs.259576 PubMed PMID: 35762511.

      (5) Joffre C, Barrow R, Ménard L, Calleja V, Hart IR, Kermorgant S. A direct role for Met endocytosis in tumorigenesis. Nat Cell Biol. 2011 Jun 5;13(7):827–37. doi:10.1038/ncb2257 PubMed PMID: 21642981.

      (6) Monteiro P, Rossé C, Castro-Castro A, Irondelle M, Lagoutte E, Paul-Gilloteaux P, et al. Endosomal WASH and exocyst complexes control exocytosis of MT1-MMP at invadopodia. J Cell Biol. 2013 Dec 23;203(6):1063–79. doi:10.1083/jcb.201306162 PubMed PMID: 24344185.

      (7) Wenzel EM, Pedersen NM, Elfmark LA, Wang L, Kjos I, Stang E, et al. Intercellular transfer of cancer cell invasiveness via endosome-mediated protease shedding. Nat Commun. 2024 Feb 10;15(1):1277. doi:10.1038/s41467-024-45558-8

      (8) Pedersen NM, Wenzel EM, Wang L, Antoine S, Chavrier P, Stenmark H, Raiborg C. Protrudin-mediated ER-endosome contact sites promote MT1-MMP exocytosis and cell invasion. J Cell Biol. 2020 Aug 3;219(8):e202003063. doi: 10.1083/jcb.202003063. PMID: 32479595; PMCID: PMC7401796.

      (9) Duan Q, Jia HR, Chen W, Qin C, Zhang K, Jia F, Fu T, Wei Y, Fan M, Wu Q, Tan W. Multivalent Aptamer-Based Lysosome-Targeting Chimeras (LYTACs) Platform for Mono- or Dual-Targeted Proteins Degradation on Cell Surface. Adv Sci (Weinh). 2024 May;11(17):e2308924. doi: 10.1002/advs.202308924. Epub 2024 Feb 29. PMID: 38425146; PMCID: PMC11077639.

      (10) Wei J, Wang J, Guan W, Li J, Pu T, Corey E, Lin TP, Gao AC, Wu BJ. PlexinD1 is a driver and a therapeutic target in advanced prostate cancer. EMBO Mol Med. 2025 Feb;17(2):336-364. doi: 10.1038/s44321-024-00186-z. Epub 2025 Jan 2. PMID: 39748059; PMCID: PMC11822115.

      (11) Yamasaki A, Miyake R, Hara Y, Okuno H, Imaida T, Okita K, Okazaki S, Akiyama Y, Hirotani K, Endo Y, Masuko K, Masuko T, Tomioka Y. Dual-targeting therapy against HER3/MET in human colorectal cancers. Cancer Med. 2023 Apr;12(8):9684-9696. doi: 10.1002/cam4.5673. Epub 2023 Feb 7. PMID: 36751113; PMCID: PMC10166911.

      (12) Soonnarong R, Putra ID, Sriratanasak N, Sritularak B, Chanvorachote P. Artonin F Induces the Ubiquitin-Proteasomal Degradation of c-Met and Decreases AktmTOR Signaling. Pharmaceuticals (Basel). 2022 May 21;15(5):633. doi: 10.3390/ph15050633. PMID: 35631459; PMCID: PMC9145792.


      The following is the authors’ response to the original reviews.

      We sincerely thank the editor and reviewers for thoroughly evaluating the manuscript. Following the comments from the reviewers we caried out four major sets of experiments and added the results and the conclusions derived from them in the revised manuscript. We also modified the abstract and the introduction. As suggested by the reviewers, we have rewritten the discussion. All the mislabelling and typing errors have been corrected, and representative graphs has been replaced as suggested. The list of the newly carried out experiments are -

      (i) In the original submission, we carried out a microscopy-based study to investigate the recycling of MET. We now added surface biotinylation approach to show the RCP or KIF16B-mediated surface delivery of MET (Fig. 5J).

      (ii) To demonstrate the functional effect of KIF16B silencing on TNBC invasion we have performed ECM degradation assay with depleted KIF16B cells (Fig. S4K-L). Further, the MET degradation in KIF16B-silenced cells has also been investigated using immunoblotting (Fig. S4H).

      (iii) To rule out the off-target effect of the siRNA used in this study, we have validated the invadopodia-associated function using 2 individual siRNAs for RAB14 and RCP (Fig. S4K-L).

      (iv) We have now introduced MMP2 as a positive control as a substrate of MT1-MMP to show the effect of silencing of the protease on its cleavage (Fig. S5I).

      We also incorporated following changes, largely additions of new plots, data in the revised manuscript.

      (i) RAB4 and RAB14 colocalization with MET in BT-549 cell line has also been added in Fig. 4 (A, B, C).

      (ii) Data showing MET silencing using siRNA and its effect on invadopodia has been added to Fig 1D.

      (iii) Graph showing percentage of cells forming invadopodia has added to Fig. S1G.

      (iv) Line intensity plots of MET-containing invadopodia has been added in Fig. 2G’.

      (v) We have added quantification of all the blots to the figures.

      (vi) A graph representing MET degradation kinetics with HGF over 3 experiments has been added to Fig S2F.

      (vii) The full field of view of Fig 1F has been added in the Fig S1L. A quantification of the gelatin degradation has been added to Fig 1F.

      (viii) The blot for loading control of Fig S1K has been changed from Actin to Vinculin.

      (ix) The blot showing expression of MT1-MMP in the SCR and knockout cells has been added to Fig S5M’.

      (x) The survival plot for patients with altered or unaltered MET and RCP has been removed. The graph showing frequency alteration of MET has also been removed.

      (xi) Additional immunoblots associated with all the figures are now provided in a newly added supplementary figure (Fig S6).

      Reviewer #1 (Public review):

      Summary:

      This study identifies a mechanism responsible for the accumulation of the MET receptor in invadopodia, following stimulation of Triple-negative breast cancer (TNBC) cells with HGF. HGF-driven accumulation and activation of MET in invadopodia causes the degradation of the extracellular matrix, promoting cancer cell invasion, a process here investigated using gelatin-degradation and spheroid invasion assays.

      Mechanistically, HGF stimulates the recycling of MET from RAB14-positive endosomes to invadopodia, increasing their formation. At invadopodia, MET induces matrix degradation via direct binding with the metalloprotease MT1-MMP. The delivery of MET from the recycling compartment to invadopodia is mediated by RCP, which facilitates the colocalization of MET to RAB14 endosomes. In this compartment, HGF induces the recruitment of the motor protein KIF16B, promoting the tubulation of the RAB14-MET recycling endosomes to the cell surface. This pathway is critical for the HGF-driven invasive properties of TNBC cells, as it is impaired upon silencing of RAB14.

      Strengths:

      The study is well-organized and executed using state-of-the-art technology. The effects of MET recycling in the formation of functional invadopodia are carefully studied, taking advantage of mutant forms of the receptor that are degradation-resistant or endocytosis-defective.

      Data analyses are rigorous, and appropriate controls are used in most of the assays to assess the specificity of the scored effects. Overall, the quality of the research is high.

      The conclusions are well-supported by the results, and the data and methodology are of interest for a wide audience of cell biologists.

      We sincerely thank the reviewer for the positive feedback and for considering our study to be well executed and rigorous. The valuable suggestions and comments certainly improved the understanding of the role of the RAB14-RCP-KIF16B axis in MET trafficking and breast cancer invasion.

      Weakness

      The role of the MET receptor in invadopodia formation and cancer cell dissemination has been intensively studied in many settings, including triple-negative breast cancer cells. The novelty of the present study mostly consists of the detailed molecular description of the underlying mechanism based on HGF-driven MET recycling. The question of whether the identified pathway is specific for TNBC cells or represents a general mechanism of HGF-mediated invasion detectable in other cancer cells is not addressed or at least discussed

      We thank the reviewer for raising this point. We would like to clarify that in TNBCs, the overexpression of EGFR and MET in the null background of the hormonal receptors; progesterone receptor, estrogen receptor, and HER2 is considered to be very crucial in terms of prognosis and treatment (PMID: 27655711, 25368674). Hence study of MET signalling and trafficking is more relevant for TNBCs compared to other cancer cells. In the current study, we therefore focused on two TNBC cell lines. We have added this in the first paragraph of introduction. Line no: 31-38.

      Reviewer #1 (Recommendations for the authors):

      Major points

      (1) My major concern refers to the clinical data presented in this study. Different from the mechanistic findings, the quality of these analyses is too low, the description in the figure legends is scant, and is absent in the Method section. The results concerning the prognostic value of RCP are the most problematic. What is shown in the Kaplan Meier? Is it the correlation between the RCP mRNA levels and the patient's survival? More importantly, to study the correlation between genetic alteration and prognostic outcome, multivariable analyses should be performed comparing the genetic alteration with known prognostic factors (sex, age, tumor size, node status, ER/PrR, HER2, Ki67 if available, tumor grade). This is because breast cancer prognosis depends on multiple interrelated factors, and only multivariate models can adjust for confounding and identify which variables independently predict outcome-providing far more accurate and clinically useful prognostic information than univariate analysis. Furthermore, overexpression of RTKs has been extensively reported and studied in TNBCs. Similarly, the relevance of MET in cancer cell invasion has been firmly established. Therefore, the data presented here, whose quality does not match the mechanistic part of the study, can be considered unnecessary. I would, therefore, recommend removing the "clinical" data. The introduction should be modified accordingly.

      We thank the reviewer for this insightful comment. To generate the graphs, we selected studies available in the publicly accessible database cBioportal (cbioportal.org/). The graphs generated by the database, representing the genetic alterations of genes has added in the Fig. S4A. However, as suggested by the reviewer, we have removed the survival plots for MET and RCP, the gene alteration frequency of MET and modified the text accordingly.

      (2) Overexpression of KIF16B has been shown to inhibit the degradative pathway, stimulating the recycling one. In agreement, silencing of KIF16B accelerates EGFR degradation (PMID: 15882625). Does this also apply to the MET receptor? In the present setting, does the overexpression of KIF16B result in prolonged MET expression?

      We thank the reviewer for raising this question. Although we did not analyze the expression level of MET in the KIF16B overexpressed cells, we have analyzed the total MET levels in the control and the KIF16B silenced cells using immunoblotting and did not observe any significant changes in the MET protein levels (Fig S4H). Line no: 390-391. This led us to believe that the depletion or overexpression of KIF16B may not have any direct effect on MET expression.

      (3) The contribution of MMP2 and MMP9 to the degradative properties of HGFstimulated TNBC cells should be investigated and compared to MT1-MMP.

      We believe that this is a relevant note from the reviewer. However, MT1-MMP is one of the most well-established metalloproteases, known till date for its role in invadopodia-associated activities in breast cancer (PMID: 35008569, 27501444). Moreover, it is the best-known candidate protease, which is a transmembrane metalloprotease could be the model in studying membrane recycling to invadopodia (PMID: 20605060, 19692588). However, as pointed out by the reviewer, we completely agree that MMP2 and MMP9 also contribute significantly to invadopodia-associated functions in TNBCs (PMID: 23902685, 25699257). HGF is also known to promote the expression and activity of MMP2 and MMP9 (PMID: 23320110, 26259977). Interestingly, the cellular machineries involved in their enhanced activity due to HGF stimulation may be distinct from what was observed in the current study and may require a distinct, elaborated study, which is beyond the scope of the current one.

      (4) The authors appropriately tested the possible shedding effect of MT1-MMP on MET. They should repeat the experiments, adding a positive control of shedding. Furthermore, the legends referring to these experiments, shown in Supplementary Figure S6, seem to be wrong (or mislabeled).

      We thank the reviewer for the suggestion. MT1-MMP is known to proteolytically cleave and initiate the activation of MMP2 (PMID: 11161720, 15095267). We now carried out the experiment with MMP2 as a control (Fig S5I). We observe an increase in the unprocessed MMP2 level in the MT1-MMP silenced cells, whereas the MET levels are unaltered. Line no: 467-468.

      We sincerely apologize for the oversight in the mislabelling. We have now corrected it in the revised manuscript.

      (5) Does altered expression of RAB14 and/or KIF16B affect MT1-MMP delivery to invadopodia in the TNBC cell lines? Does KIF16B silencing affect invasion?

      The role of RAB14 and KIF16B in MT1-MMP delivery to podosomes, a structure similar to invadopodia in macrophages has been studied by Hey S. et al. (PMID: 37696580). The study suggests that KIF16B silencing reduces the invasion of macrophages. However, the effect is not known for TNBCs. Thus, we have conducted the ECM degradation assay in TNBC cell lines to show the effect of KIF16B gene silencing on breast cancer invasion to the revised manuscript (Fig S4K-L). In both MDA-MB-231 and BT-549 cells we observed reduced ECM degradation activity upon KIF16B depletion, corroborating the observation from Hey S. et al. Line no: 392-400.

      Minor points

      (1) I recommend authenticating cell lines and stable populations by STR profiling.

      All the cell lines used in the study have been purchased from ATCC, and STR is a standard practice followed by ATCC. Further, to avoid any alteration, cells were discontinued after 20 passages. We have added this statement to the methods section in the revised manuscript. Line no: 643-644.

      (2) In the legend to Supplementary Figure S4A, the panel is described as the frequency of alterations of MET, while, if I correctly interpret it, the bar graph refers to RCP. As mentioned above, these data could be removed.

      We thank the reviewer for pointing out the mistake. We have rectified this in the revised version.

      (3) Check for typos. Sometimes invadopodia is written with the capital: "Invadopodia", in other instances it is not. The authors should be consistent throughout the manuscript. English language editing would help.

      We sincerely apologize for the inconsistency in the writing. We have removed the unnecessary capitalization of invadopodia in the revised manuscript. Line no: 142, 159, 280, 436.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Khamari and colleagues investigate how HGF-MET signaling and the intracellular trafficking of the MET receptor tyrosine kinase influence invadopodia formation and invasion in triple-negative breast cancer (TNBC) cells. They show that HGF stimulation enhances both the number of invadopodia and their proteolytic activity. Mechanistically, the authors demonstrate that HGF-induced, RAB4- and RCP-RAB14-KIF16B-dependent recycling routes deliver MET to the cell surface specifically at sites where invadopodia form. Moreover, they report that MET physically interacts with MT1-MMP - a key transmembrane metalloproteinase required for invadopodia function- and that these two proteins co-traffic to invadopodia upon HGF stimulation.

      Although the HGF-MET axis has previously been implicated in invadopodia regulation (e.g., by Rajadurai et al., Journal of Cell Science 2012), studies directly linking ligand-induced MET trafficking with the spatial regulation of MT1-MMP localization and activity have been lacking.

      Overall, the manuscript addresses a relevant and timely topic and provides several novel insights. However, some sections require clearer and more concise writing (details below). In addition, the quality, reliability, and robustness of several data sets need to be improved.

      Strengths:

      A key strength of the study is the novel demonstration that HGF-mediated, RAB4- and RAB14-dependent recycling of MET delivers this receptor, together with MT1MMP, to invadopodia -highlighting a previously unrecognized mechanism, regulating the formation and proteolytic function of these invasive structures. Another strong point is the breadth of experimental approaches used and the substantial amount of supporting data. The authors also include an appropriate number of biological replicates and analyze a sufficiently large number of cells in their imaging experiments, as clearly described in the figure legends.

      We greatly appreciate the positive assessment from the reviewer, who also acknowledged the novelty and relevance of our study. Below, we have carefully addressed the comments/concerns raised regarding this study and that have strengthened the reliability and robustness by revisiting the data, providing additional analyses where required, and clarifying methodological details.

      Weakness

      (1) Inappropriate stimulation times for endocytosis and recycling assays. The experiments examining MET endocytosis and recycling following HGF stimulation appear to use inappropriate incubation times. After ligand binding, RTKs typically undergo endocytosis within minutes and reach maximal endosomal accumulation within 5-15 minutes. Although continuous stimulation allows repeated rounds of internalization, the temporal dynamics of MET trafficking should be examined across shorter time points, ideally up to 1 hour (e.g., 15, 30, and 60 minutes). The authors used 2-, 3-, or 6-hour HGF stimulation, which, in my opinion, is far too long to study ligandinduced RTK trafficking.

      We understand the reviewer’s concern regarding the HGF stimulation time point for endocytosis and recycling. We want to highlight that to study the recycling/surface delivery of MET in response to HGF, we performed TIRF microscopy-based imaging, where images were taken within 1h of HGF addition (Fig. 2I). Additionally, we have incorporated surface biotinylation to show the recycling of MET as suggested in comment-7 (Fig. 5J). For this experiment we have used 30 min of HGF stimulation. Line no: 382-391.

      Moreover, we have observed the functional effect of HGF on ECM (gelatin) degradation and invadopodia formation after 3 h of HGF stimulation. We were curious to know where does the MET localises with prolonged ligand stimulation. Hence, to study the localization of MET to invadopodia or the endocytic markers, the cells were stimulated with HGF for 2-3 hours.

      (2) Low efficiency of MET silencing in Figure S1I. The very low MET knockdown efficiency shown in Figure S1I raises concerns. Given the potential off-target effects of a single shRNA and the insufficient silencing level, it is difficult to conclude whether the reduction in invadopodia number in Figure 1F is genuinely MET-dependent. The authors later used siRNA-mediated silencing (Figure S5C), which was more effective. Why was this siRNA not used to generate the data in Figure 1F? Why did the authors rely on the inefficient shRNA C#3?

      We understand the concern raised by the reviewer. We want to emphasize that we have employed three different approaches to investigate the effect of MET silencing/inhibition on invadopodia formation. (i) A MET kinase inhibitor, PHA665752, which shows reduced invadopodia formation (Fig. 1E, E’). (PMID: 21973114, 41009793) (ii) Silencing with shRNA: Since the level of silencing of MET with the shRNA was not sufficient, cells were stained with MET as a readout for MET silencing, and images of the cells with reduced MET expression were captured. ECM degradation activity and invadopodia numbers were found to be reduced in the MET-depleted cells (Fig. 1F). (iii) We have now added the data showing the effect of siRNA-mediated MET depletion on invadopodia formation to the revised figure 1D. Line no: 123-125. To draw a robust conclusion regarding the role of MET on invadopodia-associated TNBC invasion, we have integrated all three complementary approaches.

      (3) Missing information on incubation times and inconsistencies in MET protein levels. The figure legends do not indicate how long the cells were incubated with HGF or the MET inhibitor PHA665752 before immunoblotting. This information is crucial, particularly because both HGF and PHA665752 cause a substantial decrease in the total MET protein level. Notably, such a decrease is absent in MDA-MB-231 cells treated with HGF in the presence of cycloheximide (Figure S2F). The authors should comment on these inconsistencies. Additionally, the MET bands in Figure S1J appear different from those in Figure S1C, and MET phosphorylation seems already high under basal conditions, with no further increase upon stimulation (Figure S1J). The authors should address these issues.

      We apologise for the unintentional omission of experimental detailing about HGF or drug incubation time, which we have incorporated into the figure legend appropriately. Regarding the decreased MET level in the drug-treated condition: literature suggests that the MET inhibitor PHA665752 also promotes MET degradation, corroborating our result shown in Fig. S1J (PMID: 15788682, 18327775). Further in Fig. S1J, the relative phosphorylation of MET when compared to the total MET level in the HGF-treated condition is higher (~2-fold). Quantification of the blot has been added now.

      Further, addition of HGF for 3 h leads to 40±15% reduction in the MET protein levels as seen in the Author response image 1 representing quantification of different immunoblot. The degradation of MET in the Fig S1J is 65% which nearly fall in the range for HGF-mediated MET degradation.

      Author response image 1.

      Quantification of immunoblots showing MET signal intensity in the presence or absence of HGF normalized with the loading control. N=6.

      Next, in the fig. S1A, K the rabbit anti-MET (CST, D1C2 XP) antibody has been used, which binds to a C-terminal motif of MET and identifies both the 170kDa as well as 140kDa protein representing the uncleaved and cleaved form of MET. In Fig. S1J, the mouse antiMET (CST, L6E7) antibody has been used, which binds to an N-terminal motif of MET and recognizes only the 140kDa protein.

      (4) Insufficient representation and randomization of microscopic data. For microscopy, only single representative cells are shown, rather than full fields containing multiple cells. This is particularly problematic for invadopodia analysis, as only a subset of cells forms these structures. The authors should explain how they ensured that image acquisition and quantification were randomized and unbiased. The graphs should also include the percentage of cells forming invadopodia, a standard metric in the field. Furthermore, some images include altered cells - for example, multinucleated cells - which do not accurately represent the general cell population.

      We thank the reviewer for raising this point. The single-cell images are shown for clarity and to visualize the subcellular features; however, the conclusions are made based on the quantitative analysis of multiple cells collected from multiple fields of view (Frames). At least 30 such frames per condition having 4-7 cells/ frame has been acquired and analysed for the quantification throughout this manuscript. We would like to highlight that the image acquisition has been done over random fields on a coverslip. In the revised manuscript, for a better representation of the population of cell-forming invadopodia, a graph showing the percentage of cells forming invadopodia have been added (Fig S1G). Line no: 118-119. The percentage of cells forming invadopodia increased upon HGF stimulation in MDAMB-231.

      (5) Use of a single siRNA/shRNA per target. As noted earlier, using only one siRNA or shRNA carries the risk of off-target effects. For every experiment involving gene silencing (MET, RAB4, RAB14, RCP, MT1-MMP), at least two independent siRNAs/shRNAs should be used to validate the phenotype.

      We would like to clarify that we are using SMARTPool siRNA, which contains 4 individual siRNAs for the target gene. Literature suggests that using a pool of siRNA has reduced off-target effects compared to using single oligos for gene silencing (PMID: 14681580, 33584737, 24875475).

      While SMARTpool siRNA minimizes the off-target effect, it does not eliminate the possibility of it. To confirm that the observed phenotypes are specifically attributable to the genes investigated in this study, we now performed functional experiments using two independent siRNAs targeting RCP and RAB14. The results have been added to figure S4K-L. Silencing of RCP or RAB14 using single oligos resulted in decrease in the degradation index comparable to SMARTpool siRNA, thus phenocopied the SMARTpool siRNA. Line no: 392-400.

      Further, RAB4 is well established to be associated with MET trafficking and it served as a positive control in our study (PMID: 21664574, 30537020). Additionally, a recent study by Hey et al. have used individual oligos for KIF16B to demonstrate the effect of KIF16B silencing on gelatin degradation, which corroborate with our observation from the KIF16B silencing using the SMARTpool siRNA (PMID: 37696580).

      For MET, we used siRNA, shRNA and an inhibitor to show the effect of MET inhibition/perturbation in the invadopodia-associated activity, which validates the observations of siRNA-mediated gene silencing (detailed in point 2).

      We did not perform any experiments using single oligos targeting MT1-MMP, since in our MT1-MMP siRNA-based study, now we have taken an appropriate positive control to validate the efficacy of MT1-MMP silencing (Fig. S1I). In addition, we have shown the effect of MT1-MMP depletion on invadopodia formation using a CRISPR-based gene knock-out study, and another study from our group has shown a similar effect using siRNA (PMID: 31820782), which supports our MT1-MMP KO cell observation.

      (6) Insufficient controls for antibody specificity. The specificity of MET, p-MET, and MT1-MMP staining should be demonstrated in cells with effective gene silencing. This is an essential control for immunofluorescence assays.

      The anti-MET antibody (CST, D1C2 XP) has been used in several studied (PMID: 41166312, 41152910, 39748059). The CST L6E7 anti-MET antibody has been used in studied by Radke et al, Wang et el., Kong et al. (PMID: 36435874, 38262412, 32214092). In our study, immunoblots demonstrating depletion of MET in the siRNA or shRNA-treated cells has been provided in Fig. S1I, K respectively. Further, we have demonstrated MET silencing using immunofluorescence. We also have now added the entire field of view in Fig S1L of showing cells treated with control or shRNA against MET. In the shRNA-treated condition, the cell at the centre shows low MET fluorescence intensity indicating depletion of the RTK, while the surrounding cells have MET staining similar to control.

      Tyr 1234/35 are present in the active site of MET kinase domain and upon binding of the ligand promotes their autophosphorylation (PMID: 17667909). Earlier studies have established the specificity of the phosphor-MET antibody by immunoblotting and immunofluorescence using MET inhibitors (PMID: 21973114, 41009793). In our study we have shown that the inhibition of MET kinase activity using PHA665752 abolished the MET phosphorylation at the Tyr 1234/35, as shown in Fig S1J which revalidates the specificity of the antibody.

      Additionally, in a previous study Joffre et al. have shown that an oncogenic mutant form of MET, M1250T is highly phosphorylated at the Tyr 1234/1235 (PMID: 21642981). Using the phospho-MET antibody, we have shown in Fig 3C, S2I the increased Tyr phosphorylation of M1250T MET mutant as reported by Joffre et al.

      The anti-MT1-MMP antibody is also a very well-established antibody reported in multiple studies (PMID:32479595, 31820782, 35762511). In our study, we have shown the specificity of the antibody using immunoblot analysis. Immunoblots showing significant depletion of MT1-MMP protein level following the SMARTpool siRNA and sgRNA-mediated gene silencing has been provided in Fig. S5I, M’, respectively. Further MT1MMP silencing has been also validated by immunofluorescence in the following studies. PMID: 22291036, 21571860, 20505159.

      (7) Inadequate demonstration of MET recycling. MET recycling should be directly demonstrated using the same approaches applied to study MT1-MMP recycling. The current analysis - based solely on vesicles near the plasma membrane - is insufficient to conclude that MET is recycled back to the cell surface.

      We appreciate the reviewer’s suggestion for an alternative approach to show MET trafficking. We have demonstrated MET trafficking using surface biotinylation, where we have shown that the RCP and KIF16B depletion affect the surface delivery of MET (Fig 5J). Line no: 382-391.

      In addition, to study the surface delivery of cargo, TIRF is a widely used reliable approach and it is highly sensitive technique for detection of surface delivery events (PMID: 24344185, 20971701). We have also tried to investigate the trafficking of MET using antibody uptake approach; however, it could not be established as the binding of the antibody hindered the ligand binding and vice versa.

      (8) Insufficient evidence for MET-MT1-MMP interaction. The interaction between MET and MT1-MMP should be validated by immunoprecipitation of endogenous proteins, particularly since both are endogenously expressed in the studied cell lines.

      We thank the reviewer for pointing out the insufficient evidence for MET-MT1-MMP interaction at the endogenous level. We now carried out the immunoprecipitation of endogenous MET to validate the interaction with MT1-MMP (Fig S5H). A light (low intensity) band corresponding to MT1-MMP was detected in the anti-MT1-MMP immunoblot for the immunoprecipitated sample. We believe that the interaction between MT1-MMP and MET may be weak in nature, resulting in limited co-immunoprecipitation of the endogenous MT1-MMP by MET. The immunoblot is now added to the revised manuscript. Line no: 460-461.

      (9) Inconsistent use of cell lines and lack of justification. The authors use two TNBC cell lines: MDA-MB-231 and BT-549, without providing a rationale for this choice. Some assays are performed in MDA-MB-231 and shown in the main figures, whereas others use BT-549, creating unnecessary inconsistency. A clearer, more coherent strategy is needed (e.g., present all main findings in MDA-MB-231 and confirm key results in BT549 in supplementary figures).

      MDA-MB-231 and BT-549 are two well-characterized TNBC cell lines that readily form invadopodia. These cell lines have been extensively used to study invadopodia-associated breast cancer cell invasion (PMID: 32697977, 35915226, 31533971). These two cell lines also show overexpression of MET, making them suitable model cell lines for our study (PMID: 36139568, 20687930, 27502396).

      Overall, most of the conclusions reported in this manuscript are derived from multiple experimental approaches using two TNBC cell lines for generalization.

      We agree with the reviewer that showing the results from one type of cell line in the main figure would have been better, and wherever possible, we now provided the observations from a single cell line in the main figures and the data from the other cell line in the supplementary figures. However, some of the overexpression studies were performed in BT-549 cells to derive robust statistically meaningful conclusions. Therefore, we could not avoid adding results from both the cell lines in some of the figures. We would like to add that the legends for these figures have been edited to clearly mention the cell lines associated with each of the figure panels to avoid any confusion or inconsistency.

      (10) Inconsistency in invadopodia numbers under identical conditions. The number of invadopodia formed in Figure 1E is markedly lower than in Figure 1C, despite identical conditions. The authors should explain this discrepancy.

      We sincerely thank the reviewer for pointing out the inconsistency in invadopodia numbers across 2 experiments. Fig. 1C has 2 conditions: UT and the HGF-treated condition. The Untreated condition has the serum-free media without any stimulation. Whereas we have added vehicle (DMSO) in Fig. 1E, E’, since the drug is resuspended in DMSO. This difference in the treatment is likely to be responsible for the decreased numbers of invadopodia in Fig. 1E. In different studies it has been shown that DMSO is not biologically inert and can affect invasive properties of cells by perturbing actin dynamics and metalloprotease activity (PMID: 33552397, 22529897, 7188610).

      (11) Questionable colocalization in some images. In some figures - for example, Figure 2G - the dots indicated by arrows do not convincingly show colocalization. The authors should clarify or reanalyze these data.

      As suggested by the reviewer, we have now re-analyzed the data for figure 2G. The apparent visual lack of colocalization is likely due to the relatively lower fluorescence intensity of MET at these structures. We have now added the line intensity plots for the indicated puncta to show the intensity of both channels at the ‘dots’ in the figure 2G’ and they show correlation in their intensity distribution.

      We would also like to elaborate that to quantify the colocalization of two channels, we have used the automated image analysis software Motiontracking (motiontracking.mpi-cbg.de) (PMID: 16143105), which has been detailed in the method section. Briefly, the algorithm works on object-based co-localization. If the fluorescence intensity distributions at a given object corresponding to any two different channels (fluorophores) show 35% or more overlap (Author response image 2), the object is considered as a multi-colour object and the overlapped area value is used to calculate the degree of co-localization. The calculation is carried out over all the objects in a given field of view (frame) and over all the field of views (frames) acquired for a given condition. Also, the apparent colocalization is corrected for random colocalization, which is the random permutation of object colocalization. This makes object-based colocalization more reliable than intensity-based colocalization.

      Author response image 2.

      Image showing the object identification and contour of the multicolour object identified by Motiontracking. The plot shows the intensity distribution of these two objects as analyzed by Motiontracking.

      (12) Abstract, Introduction, and Discussion require substantial rewriting.

      (a) The abstract should be accessible to a broader audience and should avoid using abbreviations and protein names without context.

      (b) The introduction should better describe the cellular processes and proteins investigated in this study.

      (c) The discussion currently reads more like an extended summary of results. It lacks deeper interpretation, comparison with existing literature, and consideration of the broader implications of the findings.

      We thank the reviewer for this suggestion. We have substantially modified the abstract, and the introduction following the reviewer’s suggestion. The introduction has been edited to describe the cellular processes investigated in this study and some of the key associated molecular machineries. In the discussion section, we have avoided redundant descriptions of the results but retained some of them wherever necessary for interpretation and relevant discussion in the light of existing literature.

      Reviewer #2 (Recommendations for the authors):

      (1) Quality of charts. Several charts (e.g., Figure 1B, 1C, 1F) are of poor visual quality. The authors should provide higher-resolution graphs with clearer axis labels, consistent formatting, and properly scaled data.

      We thank the reviewer for pointing out the insufficient visual quality of some of the charts. We believe the resolution of the charts/graphs have changed while converting to PDF, due to image compression. We will provide the uncompressed charts with much improved visual quality, provided they are not restricted by file size limitation.

      (2) Full protein names on first mention. Whenever a protein appears for the first time in the manuscript, its full name should be provided, if possible, before using the abbreviation.

      We have incorporated the full name of the protein while reporting for the first time in the manuscript.

      (3) Correct use of "invadopodium" vs. "invadopodia." Invadopodia is the plural form; the singular is invadopodium. The sentence "Invadopodia, an actin-rich membrane protrusion decorated with proteases, is a tool for ECM and basement membrane degradation during cancer cell invasion" should be corrected accordingly.

      We are thankful to the reviewer for pointing out the grammatical error. We corrected the error in the revised version. Line no: 44.

      (4) Unclear sentence about resistance and invasion.

      The sentence "However, often patients develop resistance to EGFR-targeted therapies due to overexpression of MET; yet, the mechanistic understanding of MET-dependent cancer invasion is unclear" is confusing because the shift from drug resistance to invasion is abrupt. The authors should revise this sentence for clarity and logical flow.

      We are thankful to the reviewer for highlighting this sentence. We have rewritten the sentence as follows “Since one of the receptor tyrosine kinases (RTK), EGFR is often amplified in TNBC patients, they are usually targeted for its treatment. However, often patients develop resistance to EGFR-targeted therapies due to overexpression of another RTK MET”. Line no: 35-38.

      (5) Incorrect figure reference. In the paragraph describing the results related to RAB proteins, there is an incorrect reference to Figure 3 instead of Figure 4. This should be corrected.

      We sincerely apologize to the reviewer for the incorrect figure reference. We have corrected the reference to figures in the revised manuscript. Line no: 251, 261.

      (6) Ambiguous sentence regarding MET activation. The sentence "MET, upon activation by HGF, triggers the activation of the RTK that induces cancer cell invasion" is unclear and should be rewritten for precision and clarity.

      We are thankful to the reviewer for highlighting the unintentional mistake. We have now added a clearer sentence “MET-HGF signalling axis are reported to promotes invasion in gastric cancer cells and melanoma cells”. Line no: 96-97.

      (7) Questionable wording of figure legend. The phrase "Immunoblotting of indicated cell lines with MET and Tubulin" is an informal shortcut. The authors should rephrase it.

      We are thankful to the reviewer for highlighting this sentence. We have modified the figure legend with appropriate text. The modified text is as follows: Lysates of MDA-MB-231, BT-549 and MCF10A DCIS were separated by SDS-PAGE and analyzed by Western blot. Membranes were probed with anti-MET and anti-Vinculin antibody.

      (8) Unnecessary capitalization. Terms such as invadopodia and cortactin should not be capitalized. The authors should correct capitalization throughout the manuscript.

      We are thankful to the reviewer for raising this point. We have modified this accordingly in the revised version. Line no: 142, 159, 280, 436.

    1. eLife Assessment

      This work establishes a valuable theoretical finding about how the spike timing dependence of inhibitory plasticity shapes recurrent network connectivity. The combination of theoretical analysis and simulations provides convincing evidence that effective inhibitory connectivity forms a so-called Mexican-hat profile when multiple inhibitory neuron types follow distinct learning rules. These mechanisms are thought to be implicated in the contextual modulation of neuronal responses to stimuli.

    2. Reviewer #1 (Public review):

      Summary:

      Festa et al. provide a detailed analysis of the outcome of spike-timing-dependent plasticity acting on inhibitory synapses for distinct shapes of the kernel that governs how pre- and postsynaptic spike times induce synaptic changes. The authors investigate symmetric and asymmetric kernels, providing a theoretical description of the ingredients that give rise to rate- or covariance-dominated plasticity based on a simplified two-neuron circuit. These analyses are confirmed via simulations of large recurrent networks with random excitatory connectivity. For excitatory connections arranged in a one-dimensional ring, the authors show that two distinct classes of inhibitory neurons (distinguished by their plasticity rules) form an effective Mexican-hat weight profile. Furthermore, the authors show that external inhibition of one of the inhibitory neuron types gives rise to the phenomenon of surround modulation.

      Strengths:

      The analytical description of the two-neuron circuit is robust and accurately captures the qualitative evolution of inhibitory weights in the recurrent network with random excitatory connectivity. The emergence of the Mexican hat from the combination of distinct inhibitory synaptic plasticity rules acting on different neuron types is an important result that reveals how such connectivity can be learned in biologically plausible networks. All the analyses are well done, and the simulation results are convincing, which supports a robust interpretation of the findings.

      Weaknesses:

      The two-neuron circuit model is a good choice for the analytics, but it may have hidden a covariance effect of the "rate-dominated" symmetric spike-based kernel that would appear when several inhibitory neurons, each sharing a different spike correlation with the postsynaptic neuron, converge onto it. The rate homeostasis achieved by the rate-dominated model arises from adjusting inhibitory weights according to their initial correlation with the output neuron, so that after learning, the weights are distributed such that these correlations are cancelled out (Vogels et al., 2011). In other words, even the rate-dominated rule is covariance-driven under the hood: with a single inhibitory input, the two-neuron circuit cannot expose this, but with several differently correlated inputs, the covariance dependence should reappear.

      It is unclear whether the distribution of inhibitory weights has stabilised after 25 minutes of simulation time (Figure 3C), given that a considerable proportion of (mutual) weights reach the maximum allowed weight while (unidirectional) weights appear to vanish. Without a maximum-weight bound, and given sufficiently long simulations, the weights might diverge to infinity or decay to zero, so the apparent stationarity may be imposed by the bound rather than reflecting a true steady state. This could also be a finite-size effect, given the small number of excitatory connections per neuron.

      The connections from excitatory neurons to the two inhibitory populations are different in the ring model (exc to PV is wider than exc to SST according to Table 3), and it is not clear whether this width difference, rather than the plasticity rules themselves, is responsible for the emergence of the Mexican hat.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates how inhibitory synaptic plasticity can stabilize recurrent neural circuits while also shaping their functional connectivity. The authors analyze inhibitory spike-timing-dependent plasticity rules and show that different temporal kernels promote distinct E/I motifs, including reciprocal E/I connectivity and lateral inhibition. Using reduced circuit analyses and larger spiking network simulations, they demonstrate that inhibitory plasticity can generate structured effective connectivity, including Mexican-hat-like interactions in ring networks, while maintaining stable activity. The work therefore extends the view of inhibitory plasticity from a primarily homeostatic mechanism to one that may contribute to computationally useful circuit organization.

      Strengths:

      A major strength of the study is that it identifies a concrete mechanism by which the temporal shape of iSTDP rules determines the structure of learned inhibitory connectivity. The comparison between rules favoring reciprocal E/I motifs and those favoring "lateral" inhibition is shown across both reduced circuit models and larger spiking networks. The ring-network simulations further connect these learned motifs to circuit-level outcomes, including Mexican-hat-like effective connectivity, surround-suppression, and modular spontaneous activity.

      Weaknesses:

      The main limitations concern the extent to which the learned motifs are fully self-organized and how broadly the results generalize. In particular, the ring-network results rely on a pre-specified ring-like excitatory architecture and on two inhibitory populations with distinct plasticity rules, making it important to clarify which aspects of the Mexican-hat effective connectivity emerge from iSTDP itself. The conclusions would also be strengthened by intermediate plasticity rules. Finally, the ring-network simulations provide an interpretable proof of principle, but the authors should clarify whether the PV/SST effects depend on this specific architecture or would also arise in a more generic recurrent or cortex-like connectivity motif.

      The authors largely achieve their aim of showing that inhibitory synaptic plasticity can provide structured stabilization of recurrent circuits. The results support this claim within the model framework by demonstrating that different temporal forms of iSTDP lead to distinct learned E/I motifs and can shape effective connectivity and cortical-like response patterns. However, the broader biological interpretation remains more suggestive because some results depend on specific assumptions for the network architecture and plasticity rules.

      The work is likely to be valuable for researchers studying inhibitory plasticity, E/I balance, cortical circuit development, and biologically plausible learning because it provides a clear theoretical link between local inhibitory learning rules and circuit-level organization. The combination of analytically tractable motifs, spiking network simulations, and publicly available code makes the framework useful for future research.

      The significance of the work lies not in showing that inhibitory plasticity can have functions beyond homeostatic stabilization, which has been established by previous theoretical and experimental studies, but in formalizing how the temporal form of iSTDP rules can bias the emergence of distinct E/I motifs. At present, the work identifies rules that are sufficient to generate these motifs in model networks, while the mapping of these rules onto specific interneuron types remains for future experimental testing.

    4. Author response:

      We thank the editors for sending our work for review and the reviewers for their thorough and constructive evaluations. In response to their feedback, we will submit a revised version of the manuscript soon. Below, we provide clarification on several of the concerns raised and outline the changes planned for the revised manuscript.

      Reviewer 1, weaknesses

      (1) The two-neuron circuit model is a good choice for the analytics, but it may have hidden a covariance effect of the "rate-dominated" symmetric spike-based kernel that would appear when several inhibitory neurons, each sharing a different spike correlation with the postsynaptic neuron, converge onto it. The rate homeostasis achieved by the rate-dominated model arises from adjusting inhibitory weights according to their initial correlation with the output neuron, so that after learning, the weights are distributed such that these correlations are cancelled out (Vogels et al., 2011). In other words, even the rate-dominated rule is covariance-driven under the hood: with a single inhibitory input, the two-neuron circuit cannot expose this, but with several differently correlated inputs, the covariance dependence should reappear.

      Reviewer 1  highlights that the rate-homeostatic rule by Vogels et al. (2011) also includes a covariance-dependent component. Thus, in a network with multiple excitatory units, differences in pre-postsynaptic correlations can drive a redistribution of weights, resulting in stronger inhibitory weights for higher correlations. Our two-neuron circuit, which contains only one plastic inhibitory input, cannot reveal this competitive effect. We note, however, that although in the Vogels rule the covariance-dependent term is non-zero, it is typically much smaller than the rate-dependent term (see Methods, Section 4.3). Consequently, covariance-dependent organization may emerge on a slower timescale (a similar effect was also shown by Lagzi and Fairhall 2024; https://doi.org/10.1126/sciadv.adi4350). In the revised manuscript, we will clarify this point and analyze small motifs with multiple excitatory units and heterogeneous correlations, focusing on both the timescale of weight redistribution and the resulting steady-state weights. We will also revise our interpretation of Supplementary Figure S9: within the simulated time window, the rate-dominated rule produces uniform inhibitory connectivity, but this does not exclude slower covariance-dependent reorganization.

      (2) It is unclear whether the distribution of inhibitory weights has stabilised after 25 minutes of simulation time (Figure 3C), given that a considerable proportion of (mutual) weights reach the maximum allowed weight while (unidirectional) weights appear to vanish. Without a maximum-weight bound, and given sufficiently long simulations, the weights might diverge to infinity or decay to zero, so the apparent stationarity may be imposed by the bound rather than reflecting a true steady state. This could also be a finite-size effect, given the small number of excitatory connections per neuron.

      The reviewer raises the possibility that the apparent stationarity in Figure 3C is influenced by the imposed weight bounds. In the revised manuscript, we will discuss this point and present extended versions of the simulations in Figure 3 where we remove the weight bounds. We will also test whether the observed behavior depends on network size or excitatory connection density. We note that, under different external input regimes, the weights stabilize without reaching the hard bound (Figure S11), suggesting that saturation is not a necessary outcome.

      (3) The connections from excitatory neurons to the two inhibitory populations are different in the ring model (exc to PV is wider than exc to SST according to Table 3), and it is not clear whether this width difference, rather than the plasticity rules themselves, is responsible for the emergence of the Mexican hat.

      We recognize that we did not fully justify the parametrization used in the ring model. The pre-existing ring architecture determines the spatial correlations available to iSTDP and therefore contributes to the learned connectivity. In the revised manuscript, we will add control simulations in which the excitatory inputs to the PV and SST populations have either identical or markedly different widths, while the plasticity rules are kept fixed. These controls will allow us to assess the relative contributions of these factors to the formation of a Mexican-hat effective-connectivity profile.

      Reviewer 2, weaknesses

      (1) The main limitations concern the extent to which the learned motifs are fully self-organized and how broadly the results generalize. In particular, the ring-network results rely on a pre-specified ring-like excitatory architecture and on two inhibitory populations with distinct plasticity rules, making it important to clarify which aspects of the Mexican-hat effective connectivity emerge from iSTDP itself. The conclusions would also be strengthened by intermediate plasticity rules. Finally, the ring-network simulations provide an interpretable proof of principle, but the authors should clarify whether the PV/SST effects depend on this specific architecture or would also arise in a more generic recurrent or cortex-like connectivity motif.

      This comment raises an important distinction between the components that are specified and those that emerge through plasticity. In our simulations, the excitatory architecture is fixed, whereas the initially weak and unstructured inhibitory-to-excitatory connections are learned through iSTDP. Thus, in the ring network, the spatial organization of excitation is prescribed, but the inhibitory connectivity and resulting Mexican hat-like effective connectivity emerge from the interaction of this architecture with the two iSTDP rules. Importantly, the central result that symmetric and antisymmetric rules promote distinct reciprocal and lateral E/I motifs is not restricted to the ring network, but is also observed in sparse randomly connected spiking networks and in networks with intrinsically generated irregular activity. The ring network is therefore used to demonstrate how these general motif-forming mechanisms can support specific circuit computations.

      Our framework parametrizes a broader family of pairwise iSTDP rules rather than relying exclusively on isolated, preselected rules, allowing the contribution of rule shape and rate-dependent terms to be understood analytically. Intermediate or “mixed” plasticity rules, including those considered by Yang and Doiron (2026) https://doi.org/10.1103/9nv2-y63v, as well as additional network architectures, are valuable directions for extending the framework. We will revise the Discussion to clarify which components are prescribed, which emerge through plasticity, and which conclusions apply beyond the ring-network implementation.

      (2) The authors largely achieve their aim [...] by demonstrating that different temporal forms of iSTDP lead to distinct learned E/I motifs and can shape effective connectivity and cortical-like response patterns. However, the broader biological interpretation remains more suggestive because some results depend on specific assumptions for the network architecture and plasticity rules.

      This comment points to an important distinction between biological generality and mechanistic insight. Our models are deliberately simplified to isolate how the temporal structure of iSTDP interacts with internally generated correlations to select distinct E/I connectivity motifs, and to make this relationship analytically tractable. The resulting predictions are then reproduced in large conductance-based spiking networks with different connectivity structures and sources of irregular activity. Thus, although we do not claim that the specific biological implementations considered here capture the full diversity of cortical circuits, the conclusions are not restricted to a single minimal model or network architecture. Adding further biological detail would introduce additional parameters and architecture-specific assumptions, but would not by itself establish greater generality or provide the same mechanistic understanding. In the revised Discussion, we will clarify the distinction between the general mechanistic principles established by our framework and the more specific biological interpretations that remain to be tested.

      (3) [...] At present, the work identifies rules that are sufficient to generate these motifs in model networks, while the mapping of these rules onto specific interneuron types remains for future experimental testing.

      Our use of the PV and SST labels is intended as a biologically motivated implementation rather than as a universal assignment of plasticity rules to these interneuron classes. The symmetric and antisymmetric kernels were motivated by in vitro measurements from PV and SST interneurons in mouse orbitofrontal cortex, respectively (Lagzi et al., 2021; https://doi.org/10.1101/2021.09.06.459211). Because inhibitory plasticity can vary across brain regions, developmental stages, and experimental conditions (Feldman, 2012; https://doi.org/10.1016/j.neuron.2012.08.001), the general conclusion of our work concerns the mapping from the temporal structure of an iSTDP rule to the E/I motif it promotes, rather than a fixed correspondence between a particular rule and an interneuron identity.

      Our models therefore establish more than the sufficiency of two isolated rules: the analytical framework explains how features of the plasticity kernel and rate-dependent terms determine whether reciprocal, lateral, or blanket inhibitory connectivity emerges. The specific association of these mechanisms with PV and SST interneurons in other circuits remains an experimentally testable prediction. We will clarify this distinction in the revised manuscript and emphasize that, once cell-type-specific plasticity rules are measured in a given circuit, the framework can predict the E/I connectivity motifs that those rules are expected to promote.

    1. eLife Assessment

      This set of studies provides important knowledge for the role of zona incerta neurons in motivation and food-seeking in mice. The evidence supporting the findings is convincing, with the use of multiple approaches and nice behavioral procedures to provide converging lines of evidence. Additional analyses and data from the current studies would further strengthen the evidence. This work will be of interest to those interested in motivation and its brain substrates.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript by Laura Korobkova and Brian Dias describes an interesting study of the role of GABAergic neurons in the zona incerta (ZI) in incentive motivation for reward.

      The authors report that DREADD inhibition of ZI neurons reduced the effort breakpoint in a progressive ratio task, which measures the intensity of incentive motivation to obtain food rewards. In other tests, chemogenetic inhibition did not alter food consumption or memory.

      Conversely, DREADD excitation of ZI neurons increased incentive motivation in the progressive ratio task, expressed as a higher breakpoint for food rewards.

      Korobkova and Dias report that prior stress exposure to a series of stressors (e.g., forced swim & water submersion, restraint, mild footshock) by itself reduced the breakpoint for food reward under vehicle, though it did not impair the ability to learn an instrumental response. However, DREADD excitation of ZI neurons in previously stressed mice increased the breakpoint to normal levels equivalent to the never-stressed group. This important finding indicates the ability of ZI stimulation to rescue the incentive motivational deficit induced by prior stress.

      In fiber photometry studies using vGAT-CRE mice to specifically identify GABA neurons, Korobkova and Dias report that ZI GABA neurons are excited by sensory signals, including neutral cues. However, after reward conditioning, ZI GABA neurons increase their activation to the CS+ cue that predicts reward, but not to the CS- cue that doesn't. ZI neurons also respond in an instrumental reward task during both lever press and reward delivery. The authors conclude that ZI neurons respond to sensory stimuli, but specifically code the motivational significance of reward-related stimuli.

      In optogenetic studies, the authors find that ZI GABA neuron stimulation during a reward CS+ enhances motivated responding to obtain reward, particularly in females, but not stimulation outside the CS+. This suggests the ZI stimulation in females may specifically enhance the incentive salience of the CS+, namely the cue's ability to trigger an increase in 'wanting' for the reward. However, that effect was not found here in males.

      Altogether, this is a fine contribution to the literature, and the authors deserve congratulations on their study and manuscript.

      Strengths:

      This is a powerful and creative set of studies that clarifies the roles of ZI neurons in sensory processing and especially in incentive motivation for rewards. The use of multiple methods and test situations to triangulate on reward motivation functions gives a well-rounded perspective on ZI function. The discovery of incentive motivation roles for ZI neurons is intriguing and improves understanding of ZI, which traditionally has been a relatively understudied brain structure. The finding that ZI stimulation may rescue stress-induced deficits in motivation is especially notable and may have therapeutic implications.

      Weaknesses:

      Minor: This version of the manuscript focuses the introduction and discussion specifically on ZI GABA neurons. The ZI may be primarily GABAergic, but also contains other neurons, and DREADD studies may have used the hSyn promoter, which would impact all types of ZI neurons. Other studies here did more specifically target GABA neurons using vGAT Cre mice and specific targeting. The manuscript might be slightly improved by distinguishing in the discussion a bit more clearly which effects implicate GABA neurons specifically, and which effects might include other neurons too, to more clearly parse out the relative roles of GABA vs broader neuronal populations in ZI.

    3. Reviewer #2 (Public review):

      Summary:

      This paper describes a study that uses a combination of observational and experimental techniques to investigate the hypothesis that the zona incerta is a neural loci where sensory information is integrated to interpret the motivational value of reward-associated cues. They show that manipulation of GABAergic neurons in this region bidirectionally modulates responding during a progressive ratio test, that activating these neurons recovers motivational deficits incurred by chronic stress, and that they fire in response to reward-associated visual or auditory cues. They also showed that activity in these neurons is not necessary for incentive salience of reward-associated cues, because inactivating them did not prevent Pavlovian-instrumental transfer. However, activating them did enhance responding during the presentation of reward-associated cues in females but not in males.

      Strengths:

      The study has a very systematic and elegant approach to assess how this region responds first to intrinsic motivation and then to motivation-enhancing effects of reward-associated cues.

      Weaknesses:

      Males and females are used throughout, but sample sizes are generally too small to make a meaningful interpretation of sex differences (which is not the focus of the study, but is worth bearing in mind). In the last experiment, the lack of discrimination between CS+ and CS- conditions across training for males confounds any interpretation of sex-differences in the outcomes.

      The ZI is known to be a region where there is notable convergence of neural inputs from a diverse and heterogenous range of sensory and other cortical inputs. To my knowledge, this is the first study that has directly tested whether it may serve to encode motivational/incentive properties of reward-associated cues. The outcomes are not definitive - it appears that they are sufficient but not necessary. However, this study represents an important first step - the ZI also has notable heterogeneity in the genetic identity of neurons, and properly dissecting the function of ZI microcircuits will likely require characterising function based on more than one molecular marker. This is addressed by the authors in the discussion.

      In summary, this study will have a significant impact on our understanding of how motivation is calculated based on complex environmental signals.

    4. Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the zona incerta in motivation and cue-reward associations. Using chemogenetic and optogenetic manipulations of the ZI, they altered motivation in cued and uncued variants of the progressive ratio task and rescued deficits in motivation induced by chronic stress. They further use fiber photometry to demonstrate that the ZI tracks the formation of cue-reward associations.

      Strengths:

      (1) The authors fill an important gap in the literature linking sensory input to motivation via the zona incerta.

      (2) The authors demonstrate that ZI tracks cue value rather than just tracking sensory input.

      (3) The authors demonstrate that the ZI excitation rescues stress-induced suppression of motivation.

      (4) The authors perform several important control tasks, demonstrating that their findings are not a result of alterations in locomotor activity, food consumption, or memory.

      Weaknesses:

      In Figure 1D and E (inhibitory vs excitatory DREADDS), the control groups in the Gi group appear to have more elevated breakpoints than the control groups in the Gq group, although a statistical comparison between the two is not reported. It is not clear if this is because the two groups were given a different reinforcement schedule, this should be made clearer.

      In Figure 1E, it is important to note that although the authors found a significant planned comparison between Gq VEH and Gq CNO, the interaction was not significant, nor were comparisons to mice injected with control virus. Thus, activation of ZI GABA neurons appears to be a relatively weak effect.

      In Figure 5, the authors see what is likely a significant difference in lever presses during acclimation between the Gi and GFP groups, which they state is an expected difference. However, it is difficult to see why this would be expected. While Gi:CNO manipulation yielded lower breakpoints in Figure 1D, it did not yield lower FR1 responding for food in Fig S3 (although this was FR1 for food dispenser visits rather than lever press). One reason I ask is that the authors highlight the differences in CS+/CS- between groups, but the biggest difference between groups appears to be in acclimation, which may be driving the group x block interaction.

      In Figure 6, the authors demonstrate that optogenetic stimulation during cue light increases the breakpoint in females, but not in males. They suggest that this may be because the males did not sufficiently discriminate the cue light before optogenetic manipulation began. If this were the case, then the authors would need to use "cue discrimination" as a factor to determine if it is a better predictor than sex.

      The authors' work demonstrates that chemogenetic inhibition of GABAergic ZI cells reduces uncued motivation for reward but enhances cued responses under extinction. The authors state that this is a paradoxical finding that suggests that the ZI operates within a redundant motivation network. However, a critical difference between the two tasks is that one measures motivation for food while the other measures persistent responding under food extinction, which are not the same process. Thus, a simpler explanation is that ZI inhibition reduces motivation and impairs extinction.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Laura Korobkova and Brian Dias describes an interesting study of the role of GABAergic neurons in the zona incerta (ZI) in incentive motivation for reward.

      The authors report that DREADD inhibition of ZI neurons reduced the effort breakpoint in a progressive ratio task, which measures the intensity of incentive motivation to obtain food rewards. In other tests, chemogenetic inhibition did not alter food consumption or memory.

      Conversely, DREADD excitation of ZI neurons increased incentive motivation in the progressive ratio task, expressed as a higher breakpoint for food rewards.

      Korobkova and Dias report that prior stress exposure to a series of stressors (e.g., forced swim & water submersion, restraint, mild footshock) by itself reduced the breakpoint for food reward under vehicle, though it did not impair the ability to learn an instrumental response. However, DREADD excitation of ZI neurons in previously stressed mice increased the breakpoint to normal levels equivalent to the never-stressed group. This important finding indicates the ability of ZI stimulation to rescue the incentive motivational deficit induced by prior stress.

      In fiber photometry studies using vGAT-CRE mice to specifically identify GABA neurons, Korobkova and Dias report that ZI GABA neurons are excited by sensory signals, including neutral cues. However, after reward conditioning, ZI GABA neurons increase their activation to the CS+ cue that predicts reward, but not to the CS- cue that doesn't. ZI neurons also respond in an instrumental reward task during both lever press and reward delivery. The authors conclude that ZI neurons respond to sensory stimuli, but specifically code the motivational significance of reward-related stimuli.

      In optogenetic studies, the authors find that ZI GABA neuron stimulation during a reward CS+ enhances motivated responding to obtain reward, particularly in females, but not stimulation outside the CS+. This suggests the ZI stimulation in females may specifically enhance the incentive salience of the CS+, namely the cue's ability to trigger an increase in 'wanting' for the reward. However, that effect was not found here in males.

      Altogether, this is a fine contribution to the literature, and the authors deserve congratulations on their study and manuscript.

      Strengths:

      This is a powerful and creative set of studies that clarifies the roles of ZI neurons in sensory processing and especially in incentive motivation for rewards. The use of multiple methods and test situations to triangulate on reward motivation functions gives a well-rounded perspective on ZI function. The discovery of incentive motivation roles for ZI neurons is intriguing and improves understanding of ZI, which traditionally has been a relatively understudied brain structure. The finding that ZI stimulation may rescue stress-induced deficits in motivation is especially notable and may have therapeutic implications.

      Weaknesses:

      Minor: This version of the manuscript focuses the introduction and discussion specifically on ZI GABA neurons. The ZI may be primarily GABAergic, but also contains other neurons, and DREADD studies may have used the hSyn promoter, which would impact all types of ZI neurons. Other studies here did more specifically target GABA neurons using vGAT Cre mice and specific targeting. The manuscript might be slightly improved by distinguishing in the discussion a bit more clearly which effects implicate GABA neurons specifically, and which effects might include other neurons too, to more clearly parse out the relative roles of GABA vs broader neuronal populations in ZI.

      The current version of our manuscript notes “Our chemogenetic manipulations targeted all GABAergic ZI neurons, and emerging evidence suggests molecular heterogeneity within this population that may map onto distinct functional roles (Arena et al., 2024; Wilt et al., 2025). Future work to first profile the molecular heterogeneity of GABAergic cells in the ZI and then using intersectional strategies to target defined ZI sub-populations will be essential to providing a more nuanced view of ZI GABAergic influences on motivation.” Our revision will make sure to discuss newer literature demonstrating cellular heterogeneity of ZI and emphasize that this heterogeneity will provide more nuanced contributions of ZI cells beyond the studied GABAergic population on motivation.

      Reviewer #2 (Public review):

      Summary:

      This paper describes a study that uses a combination of observational and experimental techniques to investigate the hypothesis that the zona incerta is a neural loci where sensory information is integrated to interpret the motivational value of reward-associated cues. They show that manipulation of GABAergic neurons in this region bidirectionally modulates responding during a progressive ratio test, that activating these neurons recovers motivational deficits incurred by chronic stress, and that they fire in response to reward-associated visual or auditory cues. They also showed that activity in these neurons is not necessary for incentive salience of reward-associated cues, because inactivating them did not prevent Pavlovian-instrumental transfer. However, activating them did enhance responding during the presentation of reward-associated cues in females but not in males.

      Strengths:

      The study has a very systematic and elegant approach to assess how this region responds first to intrinsic motivation and then to motivation-enhancing effects of reward-associated cues.

      Weaknesses:

      Males and females are used throughout, but sample sizes are generally too small to make a meaningful interpretation of sex differences (which is not the focus of the study, but is worth bearing in mind). In the last experiment, the lack of discrimination between CS+ and CS- conditions across training for males confounds any interpretation of sex-differences in the outcomes.

      The goal of our study was not to determine sex differences in motivation mediated by the ZI but rather demonstrate a general role of GABAergic cells in the ZI on motivation. As such, while all our experiments used male and female mice and our statistical analyses did not uncover any sex differences in most experiments, our sample sizes of each sex are too small to definitively make any statements about sex differences. As noted already in our Discussion, the sex-specific cue-utilization behavioral strategy in our optogenetic experiment warrants further investigation as a contributing factor to motivation that may or may not be influenced by GABAergic cells in the ZI. It bears mentioning that, to our knowledge, none of the recently published literature on the role of the zona incerta in learning, memory and appetitive behavior that is cited in this manuscript (including our own prior work) has uncovered sex differences in the contributions of the zona incerta to these behaviors.

      The ZI is known to be a region where there is notable convergence of neural inputs from a diverse and heterogenous range of sensory and other cortical inputs. To my knowledge, this is the first study that has directly tested whether it may serve to encode motivational/incentive properties of reward-associated cues. The outcomes are not definitive - it appears that they are sufficient but not necessary. However, this study represents an important first step - the ZI also has notable heterogeneity in the genetic identity of neurons, and properly dissecting the function of ZI microcircuits will likely require characterising function based on more than one molecular marker. This is addressed by the authors in the discussion.

      In summary, this study will have a significant impact on our understanding of how motivation is calculated based on complex environmental signals.

      Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the zona incerta in motivation and cue-reward associations. Using chemogenetic and optogenetic manipulations of the ZI, they altered motivation in cued and uncued variants of the progressive ratio task and rescued deficits in motivation induced by chronic stress. They further use fiber photometry to demonstrate that the ZI tracks the formation of cue-reward associations.

      Strengths:

      (1) The authors fill an important gap in the literature linking sensory input to motivation via the zona incerta.

      (2) The authors demonstrate that ZI tracks cue value rather than just tracking sensory input.

      (3) The authors demonstrate that the ZI excitation rescues stress-induced suppression of motivation.

      (4) The authors perform several important control tasks, demonstrating that their findings are not a result of alterations in locomotor activity, food consumption, or memory.

      Weaknesses:

      In Figure 1D and E (inhibitory vs excitatory DREADDS), the control groups in the Gi group appear to have more elevated breakpoints than the control groups in the Gq group, although a statistical comparison between the two is not reported. It is not clear if this is because the two groups were given a different reinforcement schedule, this should be made clearer.

      In our revision, we will be sure to insert language re-emphasizing that the Gi and Gq experiments were performed using different reinforcement schedules, FR3 and FR1, respectively.

      In Figure 1E, it is important to note that although the authors found a significant planned comparison between Gq VEH and Gq CNO, the interaction was not significant, nor were comparisons to mice injected with control virus. Thus, activation of ZI GABA neurons appears to be a relatively weak effect.

      In our revision, we will insert language to acknowledge that reducing motivation after inhibiting GABAergic ZI cell activity is stronger than increasing motivation seen after stimulating the activity of these cells.

      In Figure 5, the authors see what is likely a significant difference in lever presses during acclimation between the Gi and GFP groups, which they state is an expected difference. However, it is difficult to see why this would be expected. While Gi:CNO manipulation yielded lower breakpoints in Figure 1D, it did not yield lower FR1 responding for food in Fig S3 (although this was FR1 for food dispenser visits rather than lever press). One reason I ask is that the authors highlight the differences in CS+/CS- between groups, but the biggest difference between groups appears to be in acclimation, which may be driving the group x block interaction.

      In a revision, we will revise the language to state that the acclimation difference that is more pronounced in the GFP control group vs the Gi group is to be expected because we had already shown that inhibition of GABAergic cells in the ZI would reduce lever pressing. We will also report a planned comparison of CS+ versus CS− responding that omits the acclimation block, which shows discrimination between sound (CS+) and light (CS−) in the Gi group but not in the GFP group, confirming that the cue effect is not driven by the acclimation difference. We will also note that acclimation responding is non-reinforced and effortful, whereas FR1 dispenser visits (Fig. S3) are reinforced and low-effort, which is why the two measures dissociate.

      In Figure 6, the authors demonstrate that optogenetic stimulation during cue light increases the breakpoint in females, but not in males. They suggest that this may be because the males did not sufficiently discriminate the cue light before optogenetic manipulation began. If this were the case, then the authors would need to use "cue discrimination" as a factor to determine if it is a better predictor than sex.

      In a revision, we will add an analysis that includes cue discrimination as a factor, to test whether it is a better predictor of the optogenetic effect on breakpoint than sex.

      The authors' work demonstrates that chemogenetic inhibition of GABAergic ZI cells reduces uncued motivation for reward but enhances cued responses under extinction. The authors state that this is a paradoxical finding that suggests that the ZI operates within a redundant motivation network. However, a critical difference between the two tasks is that one measures motivation for food while the other measures persistent responding under food extinction, which are not the same process. Thus, a simpler explanation is that ZI inhibition reduces motivation and impairs extinction.

      Our revision will include discussion of this important point and tie it into our previous work (Venkataraman et al. 2019 and 2021 – cited in this version) that included extinction-like protocols, albeit in classical (not operant) conditioning protocols.

    1. eLife Assessment

      This valuable paper reports on a measure of flexible brain state engagement, derived from fMRI, as a predictor of cognitive control. One strength of the study is the use of external datasets for validation and replication. The results are solid, although some may benefit from further explanation.

    2. Reviewer #1 (Public review):

      Summary:

      This paper uses three different datasets to study the relationship between the standard deviation of dynamic brain state time series (state engagement variability or SEV) and measures of cognition. Results show associations between SEV and cognitive measures, with stronger associations in patients than controls (at least for inhibition).

      Strengths:

      Strengths include the use of innovative dynamic approaches to study cognition and the validations across three independent datasets.

      Weaknesses:

      With a highly innovative approach, it can be challenging to provide enough context for the reader to understand and interpret the results. In particular, the paper would benefit from:

      (1) More detail on the brain state calculation, multiple comparison control, and added benchmarking of the novel summary SEV measure.

      (2) Guidance on the interpretation of relatively low prediction performance, negative t-statistics, and more broadly regarding the justification for the multi-step approach going from 4 brain states to 1 SEV to a network of edges.

      (3) Removal of the moment-to-moment alignment results given the circularity of the edge time series extraction with overlapping contributions to SEV and cognitive control time series.

      (4) Adjustment of text to avoid causal interpretations and to reduce the emphasis on transdiagnostics.

      Major Points:

      While the brain states were developed in prior work, SEV is a new metric and therefore warrants careful benchmarking in terms of test-retest reliability, sensitivity to scan length/quality, and associations with demographic variables like age and sex (which do not appear to be controlled for in analyses).

      Although the external validation approach is appreciated, the prediction performance is pretty low (predicted-observed correlation 0.17-0.3). It would be good to also report other metrics of performance, such as balanced accuracy.

      The steps in the paper are somewhat convoluted by going from 4 brain states to 1 SEV, back to specific FC networks. This makes the paper a bit complex and difficult to interpret. It would be helpful to provide a clear justification for these steps and/or a figure to orient the readers.

      Many results are reported in the manuscript, and it is unclear whether/what multiple comparisons control was adopted where.

      The moment-to-moment change section tries to test whether inter-individual variation in SEV maps onto cognitive control, which is very interesting. However, both measures were operationalized using edge-timeseries calculated from the same data with shared inputs (as shown in Figure 4B). As such, the 'alignment' (i.e., correlation) between resulting time series appears somewhat circular given that it is likely driven by the shared inputs. More broadly, edge timeseries were summed across edges (and subtracted between edges with positive and negative CPM associations), which further complicates their interpretability in the context of 'cognitive control'. I would recommend removing this section or using behavioral data to quantify cognitive control.

      The descriptions of how brain states were derived are unclear. In line 466, what do 'these fMRI data' refer to? Was the least-squares regression performed across subjects (given that it results in one beta value per time point)? Was this performed as a multiple regression and - if so - what was the collinearity between brain state inputs?

    3. Reviewer #2 (Public review):

      Summary:

      A relatively new measure of flexible brain state engagement (SEV - State Engagement Variability) is used here. It simply measures time-to-time variation in brain activity in terms of how it matches pre-specified motifs of activity. This metric seems to be predictive of behavioural data measuring cognitive control abilities. This was found to be the case in two independent datasets with different (though related) behavioural measures.

      Strengths:

      Use of multiple datasets is a clear strength. The use of both replication and out-of-sample model prediction is another.

      Weaknesses:

      (1) It is not clear to me how specific the SEV metric is for telling us about brain state engagement flexibility. Resting state fluctuations have been described as quasi-periodic changes that can be mapped onto "states", but the fluctuations could easily be a reflection of vascular flow, which may indirectly correlate with cognition.

      (2) If SEV is calculated using other state descriptors (e.g. a random parcellation of the brain into 4 networks) - would the result still hold? Or are the motifs important (this would rule out, to some extent, the vascular argument from (1) above)?

      (3) Figure 1 confused me a little. Why not show all the combinations (patient v full sample), inhibition vs shift, and main vs validation? Instead, a subset of 4 was selected?

      (4) The inhibition/patient/main correlation seems to be driven by 4 patients with particularly high inhibition measures?

      (5) Why is SEV negative in some cases (e.g., Figure 1) if it's a std measure? Has it been demeaned or orthogonalised wrt another variable?

      (6) The external analysis is great, but why should the model predict a relationship between SEV and inhibition if the claim is that it is only true for patients? Why would it only be true for patients in the first place?

      (7) I can't get my head around the results shown in Figure 3. How can one have both positive and negative correlations being significant or meaningful in the same pairs of networks? I think this set of results could benefit from more explanation.

      (8) I struggled with Figure 4 analysis. What is the SEV network? How do we know that it is specific enough to the SEV concept? Looking at co-fluctuations with the cognitive network, are we not simply looking at the old anti-correlation between the default mode and the rest of the brain (I note that the correlations in the y-axes of Figure 4 are negative)?

    1. eLife Assessment

      This valuable study aimed to explore whether sensory representations in the cortex reorganize on the same timescale over which behavioral changes first emerge. Convincing evidence is presented to show that reward-dependent plasticity takes place in the mouse barrel cortex within a single recording session as mice learn a whisker detection task. However, the evidence supporting the authors' proposal that this rapid reorganization may involve spontaneous reactivation of neurons that gain stimulus responsiveness during training is incomplete. The work will be of interest to systems neuroscientists studying cortical circuits and learning.

    2. Reviewer #1 (Public review):

      Summary:

      The paper submitted by Renard et al. seeks to capture the moment when learning occurs and to identify the associated changes in neuronal activity within cortical circuits. Specifically, the study aims to test whether sensory representations in the cortex reorganize on the same timescale over which behavioral changes first emerge.

      To address this question, the authors developed a new behavioral paradigm in which mice were first trained on an auditory detection task and then introduced to whisker stimulation, which they learned to associate with reward. This design allowed mice to form a new whisker-reward association within a single behavioral session, enabling the authors to track learning-associated neuronal changes during the course of the experiment.

      Using pharmacological and optogenetic interventions, the authors first show that learning depends on the whisker somatosensory cortex. They then combined the task with longitudinal two-photon calcium imaging to examine real-time changes in neuronal representations that accompany improvements in task performance over trials within a session and across days. By applying a range of analytical approaches, they show that learning induces a rapid reorganization of sensory cortical representations over tens of trials, on the timescale of minutes. They further propose that spontaneous reactivation of neurons during the task may contribute to these representational changes during learning.

      Strengths:

      (1) Overall, the experiments are thoughtfully designed, well controlled, and clearly presented. The conclusions are generally well supported by the data. The manuscript is clearly written, and the Discussion acknowledges potential caveats while outlining future directions.

      (2) A major strength of the study is the design of a new learning paradigm in which head-fixed mice rapidly form a new sensory-motor association within a single session, on the timescale of minutes. This offers a unique opportunity to track real-time changes in neuronal dynamics associated with learning during a single recording experiment.

      (3) Taking advantage of this behavioral design, the authors show that learning induces rapid reorganization of sensory cortical representations. They also report an increase in spontaneous reactivation of neurons that gained stimulus responsiveness during training, and propose that these reactivations may contribute to rapid representational reorganization. These findings provide important insights into the neural dynamics associated with learning.

      Weaknesses:

      (1) The authors propose that spontaneous reactivation mediates rapid reorganization of neuronal representations and thereby supports rapid task learning. However, as they also acknowledge in the Discussion, the present study does not directly test a causal role for these reactivations in facilitating representational changes or behavioral improvement.

      (3) The authors show reorganization of neuronal representations even on the first day of training with the new whisker task. However, because there is no explicit control for natural representational drift, it remains unclear to what extent these changes reflect learning-related reorganization rather than spontaneous day-to-day drifts in neuronal responses.

    3. Reviewer #2 (Public review):

      Summary:

      Renard, Foustoukos and colleagues present a study of rapid sensorimotor learning in the mouse barrel cortex. Head-fixed water-restricted mice already trained on an auditory detection task are introduced to a novel C2 whisker stimulus, and the authors show that reward-paired mice acquire the whisker-lick association within a single behavioral session, with the two groups (rewarded vs non-rewarded) diverging behaviorally within ~22 whisker trials and ~14 minutes. Both pharmacological inactivation of wS1 across Days 0/+1/+2 and optogenetic inactivation on Day 0 impair whisker-guided performance, while fpS1 manipulations do not, establishing that wS1 activity is required for whisker-guided behavior during the initial learning period. Longitudinal two-photon imaging of GCaMP6f-expressing L2/3 neurons across five days (-2 to +2 relative to whisker introduction) reveals a bidirectional, reward-dependent reorganization of population responses to passive whisker stimuli: rewarded mice show enhancement, non-rewarded mice show suppression. The authors use a logistic-regression decoder trained to discriminate pre- vs post-learning passive trials and then project Day 0 active whisker trials onto this learning axis; the projection rises monotonically across Day 0 in R+ mice and is significantly correlated with behavioral performance, with no such trajectory in R- mice. Finally, the authors detect reactivation events during catch trials by template-matching to the average passive whisker response, and show that on Day 0, the neurons most positively modulated by learning (LMI-positive) participate in these reactivations more than LMI-negative neurons in R+ but not R- mice. The authors interpret this as evidence that online, reward-gated reactivations may act as an upstream selection mechanism for which neurons undergo learning-related plasticity, operating on the minutes-timescale of within-session learning. There is much to like in this paper, with some moderate-to-major concerns that could largely be addressed with re-analysis or re-framing.

      Strengths:

      The single-session learning paradigm is a key aspect of this paper, given the rapid learning observed. Coupled with the R+ and R- design, there's a lot to like with the behavioral approach. The bidirectional response change across these R+ and R- groups (enhancement vs suppression) is also a nice finding.

      The causal manipulations demonstrate that the imaged region is used during the task. By doing both pharmacological and optogenetic inactivation, each with a control in the spatially adjacent region (fpS1), the authors make a strong case that wS1 activity is necessary for whisker-guided behavior during the initial learning period (though see below about the limitations of the current approach).

      The longitudinal two-photon imaging of the same L2/3 neurons across five days underlies essentially every neural analysis in the paper and enables the single-cell LMI and population-trajectory analyses.

      The pathway-specific analysis in Figure 3 - figure supplement 2 is very interesting, but not much time is spent on it (lines 151-155). The dissociation between wS2-projecting neurons (which show learning-related enhancement in R+ and suppression in R-) and wM1-projecting neurons (which do not) is (in my opinion) a nice instance of projection specificity - it also aligns with the known routing of task-relevant whisker information through the wS1→wS2 pathway. I would encourage the authors to motivate this experiment in the main text rather than leaving it all to the discussion (lines 256-262).

      The methods are generally well documented and easy to follow.

      Weaknesses:

      (1) Conflation of de novo association learning with generalization from auditory pre-training.

      All mice have already learned a task structure with the auditory task - "detect the salient sensory cue → lick → reward". Under these conditions, the rapid emergence of licking to the whisker stimulus could reflect either de novo formation of a whisker-specific association or generalization of an instrumental policy to a novel salient cue. The manuscript frames the result as the former ("acquisition of a novel sensorimotor association"), but the experiment cannot distinguish between the two alternatives. This distinction between de novo learning and generalization may have a meaningful impact on the interpretation, though it doesn't impact the specific results. It would be helpful for the authors to discuss the two possibilities and generally consider the contribution of generalization from auditory pre-training to Day 0 performance.

      Relatedly, the R- group is introduced (lines 69-74) and later used (lines 244-247) as a passive-exposure control that rules out representational drift. While R- group is an important control for repeated whisker stimulation and task context, it does not appear to be a pure passive-exposure control: Figure 1B shows that on Day 0 the mice lick more to the R- stimulus than with no stimulus and then extinguish that licking by Day 1. Thus, one possibility is that R- mice actively learn to suppress licking to an unrewarded stimulus (whisker) in a context where other stimuli (auditory) remain rewarded. This would be a different cognitive operation (response suppression) from a purely passive exposure condition. The manuscript therefore lacks a true passive-exposure baseline, and several claims that rely on R- as such a baseline (including that bidirectional changes are reward-driven rather than reflecting passive drift, lines 244-247) need to be reframed.

      (2) The inactivation experiments establish that wS1 is necessary on Day 0, but they cannot separate detection, acquisition, and expression.

      Both the muscimol manipulation (whole session, Days 0/+1/+2) and the optogenetic manipulation (0.1 s before stimulus onset through the 1 s reporting window) silence wS1 during the moments when the whisker stimulus must be detected for a successful trial. Under these conditions, impaired performance could reflect that the animal cannot detect the stimulus, cannot express the learned response on that trial, or cannot acquire the association. These are causally distinct processes, and the manuscript currently treats them as equivalent.

      Specifically, on Day +1 of the opto experiment (light off), do mice learn at the same rate as a naive Day 0 cohort (e.g., the R+ imaging mice on Day 0), or is performance already higher than the naive group? If higher than the naïve group, this would suggest that there is learning occurring and would suggest that something that may have been acquired during Day 0 inactivation, even if it could not be expressed.

      (3) The interpretation of the LMI-participation correlation is complicated by the peaked LMI distribution and neuron-level pooling.

      Two related issues arise from the results shown in Figure 4I. First, the LMI distribution in Figure 3F (and visible in 4I) is sharply peaked near zero. The reported r = 0.24 in R+ mice is therefore difficult to interpret biologically because the distribution is dominated by near-zero-LMI neurons and the slope may be disproportionately influenced by neurons in the tails. The key claim is better tested by comparing significantly LMI-positive, LMI-negative, and non-modulated neurons. The authors do address this in Figure 4J - showing that participation rate rises across days for significantly LMI-positive R+ neurons (p = 5×10⁻⁴) but not for LMI-negative neurons (p = 0.05) - but this analysis is not the lead result. To my understanding, Figure 4J is more interpretable and should be the key piece of data supporting their claim.

      Second, the p-value of p = 1×10^-41 in Figure 4I comes from treating thousands of neurons pooled across 19 mice as independent observations. Neurons within an animal are correlated through shared behavioral state, shared imaging session, and circuit-level interactions, so it would be helpful to consider a different statistical unit of comparison (FOV, animal, etc). For example, a linear mixed-effects model with mouse as a random effect could work.

      (4) The reactivation-LMI relationship is partially circular, and the framing in the abstract could be more constrained.

      The "reactivation template" is the trial-averaged passive whisker-evoked population vector from each session, and reactivations are detected as moments in catch-trial activity that correlate with this template above a shuffled threshold. This approach is reasonable, but it means that the reactivation-LMI relationship is not fully independent of template construction, and the framing in the abstract blurs that line. LMI-positive neurons are defined as neurons whose passive whisker-evoked responses increase from pre- to post-learning. Therefore, neurons with strong whisker responses, or neurons that become stronger components of the whisker-evoked template across learning, may be more likely to contribute to template-matching events by construction. Thus, the LMI-participation relationship could partly reflect template weighting or sensory-response amplitude, rather than showing that reactivation events selectively recruit neurons for future learning-related plasticity. It would be helpful and more reassuring if the authors could control for each neuron's whisker-template weight, baseline whisker responsiveness, and overall calcium event rate when relating LMI to reactivation participation.

      A complementary unsupervised approach could also help: rather than starting from the whisker template, one can derive co-activity assemblies directly from spontaneous activity (e.g., via PCA or ICA on the catch-trial population activity), and then ask, separately, whether any of these assemblies overlap with the whisker ensemble. The interesting test is then whether whisker-like assemblies become more frequently expressed across Day 0 in R+ but not R- mice, and whether LMI-positive neurons are preferentially loaded onto these whisker-like assemblies. This logic inverts the current pipeline and can be complementary to the current analysis. By identifying structure in nominally spontaneous activity first and then comparing to the whisker response, this could help avoid the circularity in which the template both defines the events and contains the cells being tested. The Figure 4 - figure supplement 1B partial-correlation analysis is a step in this direction but addresses only spontaneous firing rate, not template coupling. Without such a complementary approach, the authors may want to clarify that the reactivation detection is anchored to a template defined in part by the same cells whose participation is being tested.

      (5) The reactivation-as-selection-mechanism interpretation is not supported by the current data.

      The Discussion (lines 278-281) acknowledges that the authors have not shown necessity, but the end of the intro and part of the discussion (Lines 275-277) frame reactivations as a "reward-gated selection mechanism" for plasticity. An equally plausible alternative is that neurons whose synaptic inputs or intrinsic excitability have been potentiated by reward-driven learning will simply co-fire more often during quiet periods - meaning reactivations would be a consequence of plasticity that has already occurred rather than a mechanism that selects which neurons to potentiate. The current data cannot distinguish these.

      A separate concern is the use of the term "spontaneous." The authors' usage is defensible in one sense - catch trials are stimulus-free, so the activity is not externally driven. However, "spontaneous" in the reactivation literature typically connotes offline, internally generated activity during quiet wakefulness or sleep, which carries different implications for plasticity than activity during active task engagement. Catch trials in this paradigm occur within the behavioral session, with the animal still engaged in the task, potentially anticipating reward or licking. The authors should either acknowledge this distinction in the text or qualify the term - "within-session" or "inter-trial" reactivations would be more accurate and would avoid borrowing the conceptual weight of the offline-replay literature.

      The authors should also clarify whether catch-trial activity around licks (false alarms, anticipatory licks) is excluded from the reactivation analysis, and whether reactivation rates depend on recent reward, recent whisker trial outcome, or behavioral state. Specificity controls - template-matching with shuffled templates and with auditory templates - would help establish that detected events reflect whisker-specific patterns rather than generic high-coactivity moments.

      (6) Motor, lick, and behavioral-state confounds in the neural analyses are not fully addressed.

      I have two specific concerns. First, for the Day 0 active-trial projection, mean whisker reaction times in Figure 1 - figure supplement 1G are around 350-500 ms, but the distributions extend into the 0-300 ms analysis window. The correlation between the projection trajectory and the behavioral learning curve (Figure 4E, lines 196-198) is the key piece of evidence that the neural shift tracks learning. However, on hit trials the lick may fall within or close to the analysis window, so a motor confound could in principle contribute to the rising projection. The authors could repeat the projection using an earlier/shorter window, exclude trials with early licks, or regress out lick timing. It would be helpful to better understand whether this effect is, in part, driven by licking activity.

      Second, the central evidence for representational reorganization (Figure 3) rests on a post-session passive epoch in which 50 whisker stimulations are delivered after "task disengagement" (lines 131, 387-389). The concern is that the brain state during this epoch is unlikely to be matched across groups or across days. R+ mice receive additional water rewards on whisker trials, whereas R- mice receive rewards only on auditory trials. This could lead to systematic differences in satiety, arousal, and disengagement state during the passive block. Because cortical sensory responses are strongly modulated by arousal, some of the apparent learning-related enhancement (R+) or suppression (R-) of passive whisker responses across days could reflect systematic state differences during the passive epoch rather than plasticity. The disengagement criterion ("stopped licking in all trial types") is also qualitative - no consecutive-miss or time-window threshold is specified - so the epoch may begin at slightly different behavioral states across mice. To resolve this, the authors could (i) specify the disengagement criterion quantitatively and (ii) compare pupil diameter and whisker self-motion (if available) across R+ vs R- and across days during the passive epoch.

    4. Reviewer #3 (Public review):

      This is a methodologically sound manuscript and provides reasonably interpretable results. While being appropriate, they do not seem to bring entirely novel concepts; nevertheless, most of my comments concern the calibration of the interpretive claims rather than the quality of the data.

      Strengths:

      (1) Longitudinal within-subject imaging:<br /> Tracking the same layer 2/3 neurons across learning allows the bidirectional effect (enhancement in R+, suppression in R-) to be measured within identified cells rather than inferred across cohorts.

      (2) Appropriate behavioural controls:<br /> The R+/R- design controls for repeated sensory exposure, and maintaining rewarded auditory trials in both groups controls for engagement and arousal, arguing against disengagement as the source of the R- effect.

      (3) Convergent causal manipulations:<br /> Muscimol and optogenetic inactivation both abolish acquisition and include an adjacent control region (fpS1); the temporally restricted optogenetic result partially addresses the concern (Hong et al., 2018) that sustained inactivation may destabilise downstream circuits.

      (4) Convergent analyses:<br /> Single-cell learning modulation indices, population similarity measures, and a trial-resolved decoder projection onto a naïve-to-expert axis provide consistent evidence that representational change is concurrent with behavioural acquisition.

      (5) Projection-specific resolution:<br /> Retrograde labelling shows learning-related changes in wS2-projecting, but not wM1-projecting neurons, consistent with preferential routing of task-relevant signals through the wS1 to wS2 pathway.

      (6) Mechanistically motivated reactivation analysis:<br /> Relating rapid, reward-dependent plasticity to spontaneous reactivations on a timescale of minutes is an original use of the single-session paradigm.

      Weaknesses and points requiring clarification

      (1) The stimulus is not strictly novel: Passive whisker stimulations were delivered on pre-training Days -2 and -1, so what changes on Day 0 is the stimulus-reward contingency rather than the stimulus itself. This resembles contingency reassignment with reversal-like properties (and possible habituation or latent inhibition) rather than de novo learning, and the licking response is already established during auditory training. The framing should be qualified accordingly.

      (2) Barrel cortex dependence should be stated more narrowly: The data show that acute wS1 suppression prevents acquisition of this task, not that whisker detection in general requires barrel cortex; cortical dependence varies with task and manipulation (Hong et al., 2018; Ryan et al., 2022 vs Miyashita and Feldman, 2013). The near-threshold explanation would require psychometric or stimulus-intensity data.

      (3) The passive block carries confounds: It is acquired after task disengagement, when satiety, arousal, and reward history differ across groups and days. The authors should report within-block response adaptation and the robustness of the main results to early versus late passive trials. Additionally, could passive presentation of the stimulus without reward delivery lead to devaluation of the stimulus, leading to additional behaviour and plasticity changes which are not addressed?

      (4) Some statistics appear to be neuron-level rather than animal-level:<br /> Very small p-values (e.g., the LMI-participation correlation r = 0.24, p on the order of 10^-41) suggest thousands of non-independent neurons treated as independent samples, risking pseudoreplication. Central claims should rest on hierarchical or animal-level statistics with effect sizes.

      (5) The cosine-similarity decrease in R- animals needs clarification:<br /> Because cosine similarity is scale-invariant, uniform suppression would leave it largely unchanged; the observed decrease therefore implies heterogeneous suppression, reduced signal-to-noise, or increased variability, and the favoured interpretation should be stated.

      (6) The decoder requires cautious interpretation:<br /> Training on passive trials and applying to active Day 0 trials could introduce a behavioural-state domain shift. The meaning of positive and negative values in Figure 4C should be defined, and near-zero early projections reflect the classifier boundary rather than a biological baseline.

      (7) The reactivation analysis is the least conclusive and is susceptible to circularity:<br /> The template and the LMI are both derived from the passive whisker response, predisposing responsive neurons to register as reactivation participants, and the Day 0 template is obtained after learning.

      Leave-one-cell-out and pre-learning templates, cell-specific templates, and tests of whether reactivations predict subsequent trial responses would strengthen the claim; causal disruption would ultimately be required.

      (8) Figure, sample-size, and specificity points:<br /> The positive LMI shift in R+ animals is less visible than the R- shift in Figure 3F; the optogenetic cohort is small (n = 6 per group); and confirming that auditory detection was preserved during wS1 inactivation would establish whisker-specificity.

      (9) The comparison to prior work is overly broad:<br /> Banerjee et al. (2020) and Chéreau et al. (2020) are reversal learning and discrimination paradigms and may not be equated with simple whisker detection; the defensible novelty claim is the trial-resolved tracking within the first session and its concurrence with online reactivations.

    1. eLife Assessment

      This valuable study describes a simple and robust approach for estimating information-limiting noise by splitting neural populations and comparing estimator values. The authors report more accurate and robust results compared to previous methods. The evidence for the robustness of the method is currently incomplete; some concerns regarding bias need to be resolved, and additional tests need to be provided on how the number of trials and neurons affect performance.

    2. Reviewer #1 (Public review):

      The authors address a difficult and well-known problem in systems/computational neuroscience: how to estimate the magnitude of "information-limiting" noise. Existing approaches (direct Fisher-information estimation, decoding + Cramer-Rao, and large-N extrapolation) are data-hungry and unstable, which has left the field with conflicting empirical estimates across systems.

      The central proposal - "split-trial analysis" - is simple and appealing. The recorded population is randomly partitioned into two non-overlapping halves; a decoder (continuous case) or classifier (binary case) is trained on each half using the same trials; and the covariance of the two halves' decoding errors is used to estimate the variance of the information-limiting noise. There is a clean mathematical derivation to support this conclusion (although there are a couple of mathematical errors in the methods section that should be fixed to avoid confusion on the part of the reader).

      They benchmark the method in simulation against three prior methods (Moreno-Bote et al. 2014; Rumyantsev et al. 2020; Kafashan et al. 2021) and report substantially better sample efficiency, lower bias, and greater robustness. They then apply the method to three datasets: (1) mouse head-direction cells (Ajabi et al.), (2) mouse V1 (Stringer et al.), and (3) macaque PFC during a saccade task (Bartolo et al.).

      This is a strong and timely contribution. The core idea is elegant, and the method appears to be more practical than existing alternatives in the finite-data regime that real experiments occupy. The three applications are well chosen, and each yields a non-trivial, biologically interpretable result. I am strongly supportive of the potential of this paper.

      That said, the paper makes several strong empirical claims - most notably that prior V1 estimates were substantial overestimates, and that PFC information-limiting noise is temporally redundant - and the central estimator rests on an independence assumption whose finite-N validity is only partially characterized. Before these claims can be considered well supported, I would like the authors to address the following:

      Major Points:

      (1) The method relies on a key independence assumption that may not always be satisfied in the regime of finite neurons and trials. The author's main idea is to decompose the residuals of two decoders as follows:

      X1 = delta + phi1<br /> X2 = delta + phi2

      The covariance is equal to the scale of information limiting noise, Var[delta], plus three terms:

      Cov[X1, X2] = Var[delta] + Cov[delta, phi1] + Cov[delta, phi2] + Cov[phi1, phi2].

      We can define phi1 as the part of X1 that is orthogonal to delta and likewise define phi2 as the part of X2 that is orthogonal to delta; thus, the cross terms evaluate to zero, and we are left with:

      Cov[X1, X2] = Var[delta] + Cov[phi1, phi2]

      Now the authors introduce an assumption that Cov[phi1, phi2] = 0. This leaves us with Cov[X1, X2] = Var[delta], but the question is: when is it justified to assume that Cov[phi1, phi2] = 0? For example, it is possible that

      phi1 = c(N) * z + e1<br /> phi2 = c(N) * z + e2

      where z is another shared noise dimension that is not information limiting and e1 and e2 are truly independent. Here, c(N) is a constant that goes to zero as the number of neurons used to train the decoder, N, goes to infinity. Thus, in the limit of having very large neural populations at hand for the analysis, the author's assumption of Cov[phi1, phi2] = 0 can be justified. If the authors agree with this analysis, it would be nice to (a) flesh it out and include it in the methods / supplementary notes, and (b) to analyze in simulation how good this approximation is in finite N regimes. I suspect that the assumption works in finite N regimes if noise is low-dimensional, but that if there are many additional dimensions of correlation (i.e. many z's above), you will need a very large number of neurons before Cov[phi1, phi2] approaches zero.

      Along these lines, another worthwhile analysis would be to report outcomes when the neural populations are sub-sampled further. Intuitively, it should fail once you subsample to only a handful of neurons, e.g. 3, but I'm curious where the breaking point is and whether the decline is graceful.

      (2) In point 1, I raised the question of how the method behaves with a finite number of neurons. Another worry is that there is a finite number of trials. In particular, if you train two decoders on the same trials, I would worry that non-information-limiting fluctuations in those trials would induce correlations in the decoders that then would show up as correlations on the held-out test set. A more conservative approach would be to split trials into three disjoint subsets: a training set for decoder A, a training set for decoder B, and a common test set used to compute Cov[X1, X2].

      As a concrete example, suppose that on the particular trials used for training, the animal happened to be more aroused when theta = 1 and less aroused when theta = 0, and that arousal added a fluctuation on top of the neural response. This arousal-related signal is not information-limiting - it would average away given enough trials - but because both decoders are fit to these same trials, each one adjusts its weights to partially discount the same spurious high-arousal/low-arousal trend. Their weights are now distorted in a correlated way, so when both are applied to the shared test set, their errors covary, and the method reads this shared-training artifact as information-limiting noise.

      I think this dynamic should be acknowledged in the text and clarified in more detail. Ideally, simulations could be done to estimate how many trials are needed to average out this sort of confound, and similar to the suggestion in point 1 above, I would be interested in seeing what happens when the authors sub-sample trials before running their analysis. Together with point 1, the feedback is that I'd like to see more about "how many neurons and how many trials" are needed in order to trust your results. Similarly, are there diagnostics or resampling methods (e.g. bootstrapping) that could be helpful for a practitioner to know if they have enough neurons/trials?

      (3) Unless I've fundamentally misunderstood something, there is an error on page 17 in the methods. There we find sigma2 = Var[delta] = ... = Cov[phi1, phi2], but I believe this is meant to be Cov[X1, X2]. Indeed, the method assumes that Cov[phi1, phi2] = 0, as discussed in point 1.

      Additionally, on page 4, the authors introduce the main quantity as Cov[\hat{theta}_1, \hat{theta}_2] instead of Cov[X1, X2]. However, if theta is changing from trial to trial, then these two quantities are not technically equal to each other, so it would be more accurate to write down the conditioning on theta. That is, assuming conditionally unbiased decoders, Cov[X1, X2] = Cov[\hat{theta}_1, \hat{theta}_2 | theta] for a fixed theta.

      More generally, I found it hard to wrap my head around the underlying math on my first read through the paper. The polarization identity, 1/4 * (Var(X1 + X2) - Var(X1 - X2)), seems like a very roundabout way to derive the method. This identity is very helpful for the deconvolution extension, but I would have thought that a simpler and more straightforward derivation would have just used the expansion, Cov[X1, X2] = Var[delta] + Cov[delta, phi1] + Cov[delta, phi2] + Cov[phi1, phi2], as I did in point 1. I suggest the authors revise the mathematical presentation for clarity.

      Minor Points

      (1) A very nice feature of the authors' method is that they make no parametric assumption on the distribution of noise. This is in contrast to Kanitscheider et al. [12]'s finite-sample bias correction using the inverse-Wishart distribution of $\hat\Sigma^{-1}$, which is derived under an assumption of multivariate Gaussianity. I think it is worth adding a sentence to highlight this feature of the model.

      (2) Statistical inference claims (across sessions and population sizes) are supported by reported s.d.'s but no formal tests or confidence-interval-based comparisons. Given that several claims are comparative (split-trial < naive; V1 < prior reports; PFC stable over windows), please add appropriate uncertainty quantification (e.g., bootstrap CIs over sessions) and, where a difference is claimed, a test or effect size.

    3. Reviewer #2 (Public review):

      Le and Wei present a novel estimation method for information-limiting correlations. Information-limited correlations are shared noise fluctuations that affect neural encoding, but they can be hard to estimate (even to detect their presence) because they can be very small and buried under other common sources of variability that do not affect encoding. The newly proposed method bypasses two central limitations of previous approaches: extrapolation or assuming the noise structure to be Gaussian. The authors proposed a split-trial analysis where the population is split into two, and the correlations between the decoding errors arising from each population are computed. These correlations provide an unbiased measure of information-limiting correlations. The method is very simple and sound, and it is shown to deliver stable estimates with sensible magnitudes across several brain data sets. Further, even if the decoders are suboptimal, the method can detect the presence of information-limiting correlations, as only shared fluctuations of the two population decoders can possibly be observed if there are correlations that limit information.

      Comments:

      (1) The name "split-trial analysis" does not seem to reflect well the nature of the method introduced. I would propose something like "split-ensemble analysis" or "split-population decoding-correlation analysis".

      (2) Previous work has proposed a related - but different - bootstrap method, which can be mentioned in the current paper (Nogueira et al, J of Neuroscience, 2020).

      (3) The authors proposed a deconvolution method to study the shape of the distribution of information-limiting noise. An alternative would be to split neural populations into 3 or more subpopulations and compute 3rd- and 4th-order correlations between the decoding errors. This would lead to estimates of higher-order moments that can be compared to Gaussian ones and test for non-Gaussian distributions. Further, this N-split-ensemble method could be used to compare the deconvolution method results to test their consistency.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Le and Wei proposed a new method to identify differential correlations in real recordings (and simulations) that is based on splitting the simultaneously recorded population of neurons into two disjoint subpopulations. The method is based on evaluating the correlation between the decoded stimulus for each sub-population across trials. The authors validate their method on simulations and find the magnitude of differential correlations on three different publicly available datasets.

      Strengths:

      We think that this is a solid and relevant study for the computational neuroscience community, especially for the originality of the method and the fact that it seems to bypass the problem of very large populations to identify differential correlations. Overall, the results are novel and significant, and it addresses an important gap in the field. The main results are presented clearly and are easy to follow.

      Weaknesses:

      However, we believe that there are some additional analyses and clarifications that should be made to increase the clarity and impact of this study. In general, we believe that the authors should make a better effort to explain how their novel method depends on the number of trials and the number of neurons. More specifically:

      Major

      (1) The authors should show a realistic case for the covariance matrix in Figure 1. Currently, they are showing only Poisson noise (Figure 1c-e), only gain + Poisson (Figure 1f-h), and only differential correlations + Poisson (Figure 1i-k). They should show these same plots with a biologically realistic non-differential correlation structure (limited-range correlations, see Kanitscheider PNAS 2015). Perhaps even show the case for limited-range + gain + differential correlations. They should do the same for Figure 2.

      (2) Throughout the manuscript, the role of population size (N) on the method is a bit confusing. Figures 1 and 2 give the impression that N is not particularly important, which is counterintuitive and surprising. We understand that that is one of the strengths of the split-trial method, but the authors should explain in much more detail in the results and methods the role of population size on their novel method. Why is large N crucial for the other methods, but not for them? There is a little bit of population-size dependency on Figures 3-5, especially on Figure 4g. The authors should explain in more detail those effects.

      (3a) For dataset [27], the stimulus density was ~12 samples per deg for uniform sampling and ~1000 samples per deg for dense sampling. Figure S12 shows an overestimate of information-limiting noise when the number of trials used was significantly downsampled, which is, first of all, in disagreement with simulation results showing "when only a small number of trials are available to infer a large d-prime, split-trial analysis exhibits an under-estimation". It is true that we are not strictly in a binary classification task setting, but we are wondering if the authors have any justification for this result for [27].

      (3b) Related to this point, the estimated info-limiting noise was 0.26 deg with all neurons and 0.6 deg with downsampling (we guess that is the first value of red lines in Figure S12). The only difference here, if we understand correctly, is the number of trials used. Otherwise, it's exactly the same neural responses used for estimation. So, a similar magnitude should be expected. If the latter is due to an insufficient number of trials used, would the same problem apply to the uniform sampling dataset? In other words, if there were more trials recorded with uniformly sampled stimuli, would the authors expect to see a further and significant decrease of sigma as well?

  2. Jul 2026
    1. eLife Assessment

      This valuable study establishes a novel genetic model for chronic, in vivo visualization of protein engulfment in brain macrophages from the naturally short-lived African turquoise killifish. The solid data presented here show that brain phagocytes exhibit transcriptional features resembling mammalian border associated macrophages and monocyte derived macrophages rather than classical microglia and that their engulfment capacity declines with age, providing a resource for future studies on brain immune cells in teleosts. This research will interest a broad audience across developmental biology, genetics, neuroimmunology, aging research, and evolutionary biology.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Nagvekar et al. studies engulfing macrophages in the killifish brain upon aging. It first describes the development of a transgenic knock-in killifish line overexpressing a secreted fluorescent protein in neurons. This becomes a tool for isolating myeloid cells that are capable of endocytosis or phagocytosis of the fluorescent protein, which seem to comprise the majority of the myeloid cells within the young adult brain. The paper then demonstrates the similarities of what they call "engulfing macrophages" to brain myeloid cell types of other species and investigates changes to this population upon aging. Overall, the study combines multiple complementary technologies to support their data, that are nicely presented and well described in the legends, while the textual description remains very concise. The findings are of interest to scientists studying brain aging, and microglia/macrophages.

      Major comments:

      (1) Although the authors describe and analyze their data from the viewpoint of engulfing macrophages, the paper would benefit from a broader perspective and a comparison to other studies on microglia in different species. Along this line, the title does not really seem to cover the data presented here very well, and the introduction lacks a proper explanation of terminology on microglia/brain macrophages and their known roles, cell types versus cell states and the current state of the art in fish versus other model species in the context of aging.

      (2) The result that nearly all myeloid cells in the killifish brain are of the engulfing macrophage type is somewhat surprising. This appears to differ from other studies in for instance zebrafish (e.g. ref 80, that describes the heterogeneity of the myeloid cells in detail). There are two questions we like to raise: (Q1) What is the evidence towards this homogeneity? and (Q2) Could there be a technical bias?

      Regarding (Q1): What is the evidence towards this homogeneity? The markers used are overlapping with markers for microglia. It would be helpful to clarify how canonical microglia populations are represented in the dataset. What is the heterogeneity of the oScarletHIGH cells? On several plots (Fig1f, Fig2d, Fig4a) this population of cells seems more heterogeneous than described. Are there different cell states or types? What is the percentage of myeloid cells that is oScarletLOW? To what extent do these cells compare transcriptionally to the oScarletHIGH cells?<br /> a. Fig1f-i depict an enriched oScarletHIGH group alongside oScarletLOW cells. This representation is a bit misleading since it seems to indicate that really all myeloid cells are of the engulfing macrophage type whereas it is the majority, but not all.<br /> b. Line 52: The authors describe that the oScarletHIGH cell group is "enriched for signatures characteristic of macrophage functions". This finding is logical, as the isolation procedure of this population of cells was based on the endocytic and phagocytic properties of the cells. This result appears more consistent with a validation of the isolation strategy than with definitive evidence for myeloid cell identity.

      Regarding (Q2): Could there be a technical bias? An alternative explanation that may warrant discussion is whether aspects of the experimental pipeline (cell dissociation, FACS, scRNA-seq) could influence myeloid cell states. For instance, it is conceivable that dissociation induces a reactive program that enhances uptake of fluorescent protein, potentially enriching for oScarletHIGH cells. As the authors use a similar experimental setup to prove uptake of dextran and ovalbumin, such a technical artefact may merit consideration. As this would influence the major conclusions of the paper, the authors might want to address this comment with additional experimental controls, such as single-nuclei RNA-seq on control young and aged brains to profile the natural myeloid population when not submitted to a cell dissociation and FACS procedure.

      (3) The authors compare the oScarletHIGH cell transcriptomes to mouse and killifish datasets. Both the mouse (Barr et al) and killifish (Nagvekar, this paper) dataset are from enriched immune cells (mouse= CD45+ cells, and the 3 cell types selected from that). Why did the authors not compare to the whole mouse CD45+ dataset? Including zebrafish (Rovira et al, 2025) here would strengthen the evolutionary comparison. I also feel that the additional comparison with young killifish (Ayana et al) might not be that solid since this dataset was initially not enriched and has a significantly lower number of myeloid cells, and thus much less power. The old age time point in that study contained more myeloid cells and might be interesting to include for cell type comparison. There are other, perhaps more unbiased ways of comparing cell types across species, for instance SAMap, developed by co-author Bo Wang. Did the authors consider using this or other methods?

      (4) Regarding the comparison with the aged brain:<br /> a. Figure 4: It would be nice to include the same comparisons as for young fish (cfr Fig.1 panels F-I).<br /> The percentage of oScarletHIGH cells in the aged condition is 8% (Fig1-suppl1) compared to 4% at young age. On the other hand, a lower number of cells was isolated at old age compared to young age (Figure4a). Can the authors elaborate on this difference? Later on, it is stated that the engulfing capacity declines with aging, but could this be linked to the lower or potentially biased recovery of cells?

      b. Figure4a: Transcriptional differences are stated between young and old (line 226), can a relevant selection be shown in e.g. a dot plot or heatmap?<br /> The UMAP clustering does seem to indicate batch effects on panels a and g. Can the authors provide sub clustering and show that young and old/ FACS sorted high and low cover similar cell types/states? The PCA plot (panel f) and marker analysis is not fully convincing, as PC1 and 2 alone do not suffice to explain all the variance in these cells, and the markers are common ones for many microglia/macrophage cell types (and thus likely to be expressed similarly).

      c. Figure 5: It would be informative to include the corresponding aged condition for panels c and e.

      Significance:

      General assessment

      Strengths: This manuscript introduces a valuable new transgenic tool to isolate and characterize myeloid cells in the brain of the fast-aging killifish (Nothobranchius furzeri), an emerging model organism in aging research. The study combines multiple complementary approaches, including transgenesis, FACS, histology, and single-cell transcriptomics, to investigate brain immune populations and their changes upon aging. The cross-species comparison and aging analyses provide useful datasets and candidate markers for the field of neuroimmunology and comparative brain aging. Overall, the data are clearly presented, the experiments are logically structured, and the manuscript provides a useful resource for future studies on brain immune cells in teleosts.

      Limitations/points for improvement: The major limitation of the study concerns a potential technical bias introduced by the experimental pipeline (cell dissociation, FACS isolation, and transcriptomic profiling), which may have influenced the observed predominance and transcriptional state of the oScarletHIGH/engulfing macrophage population. At present, it remains difficult to fully exclude whether the protocol itself contributes to the apparent homogeneity of the myeloid compartment or induces a shared reactive state. Because this issue affects some of the central conclusions, the manuscript would benefit either from additional controls (e.g., dissociation-independent approaches such as single-nuclei RNA-seq) or from a more cautious interpretation and discussion of this possibility in the text.

      Advance: The fast-aging killifish is becoming an important vertebrate model for studying aging, yet the brain immune compartment in this species remains relatively underexplored. This manuscript provides both a novel experimental tool and a transcriptomic resource for studying myeloid cells in the killifish brain. To my knowledge, the study is among the first to profile engulfing/endocytic myeloid populations in the context of brain aging in this model organism and to compare these cells across species. The advance is primarily technical and descriptive/resource-generating, while also offering conceptual insight into how brain myeloid populations may change during aging and how they compare evolutionarily across vertebrates. Although the mechanistic interpretation would benefit from additional validation, the study clearly extends current knowledge and provides a framework for future work on neuroimmune aging in fish.

      Audience: The manuscript will primarily be of interest to a specialized basic research audience, including researchers in neuroimmunology, brain aging, microglia/macrophage biology, and comparative neuroscience. It will also be relevant to scientists using killifish or other emerging vertebrate models for aging research. Beyond the immediate field, the study may be of broader interest to researchers investigating immune-brain interactions and the evolutionary conservation of myeloid cell states across species. The transgenic line and transcriptomic datasets are likely to serve as a useful resource for future comparative and functional studies.

    3. Reviewer #2 (Public review):

      Summary:

      The work by Nagvekar et. al., reports the development of a new model in the African Killifish to study the engulfment of extracellular proteins. Specifically, they expressed oScarlet with a signaling peptide under the control of a neuronal promoter/gene to induce secretion into the extracellular space. Using this model, they found that the secreted protein was predominantly taken up by brain macrophages. Leveraging this finding, they were able to conduct RNAseq on brain macrophages from young and aged fish, where they reported differences in translation and vacuolar acidification at the transcriptional level among others. Finally, they show that the engulfment capacity of brain macrophages from old killifish is reduced when compared to their young counterparts.

      Major comments:

      (1) Red fluorescent proteins are notorious for being prone to aggregation. Are oScarlet proteins being internalized by macrophages aggregates or soluble proteins? This distinction is important as the clearance of extracellular molecules could be mediated by most cells, yet aggregates could be removed specifically by macrophages. Can experiments be conducted to distinguish between these two possibilities? We realize this may be challenging. If not feasible, the discussion should be tempered to reflect this possibility.

      (2) Brain dissociation tends to generate a lot of debris, especially from sheared neurons. Therefore, the high level of oScarlet inside macrophages could be an artifact of dissociation rather than a reflection of in vivo clearance. Authors should use internalization inhibitors during dissociation (CytoD, Dynasore, and pitstop) to exclude this possibility. Alternatively, if they have a transgenic killifish that expresses another fluorescent reporter in neurons (and preferably at a similar level to that of oScarlet), authors should dissociate brains together and quantify how many oScarlet+ cells are now also positive for that other fluorescent reporter. This could give an idea of how much engulfment is occurring due to the dissociation processes. It is not ideal, as macrophage eating could be happening during dissociation but before cells are in single cell suspension. However, given that RNAseq is needed to identify macrophages, this reviewer would be satisfied by this alternative approach if the aforementioned pitfall is also presented in the discussion.

      (3) Related to the above, it appears based on the scRNAseq that dissociation heavily enriched for brain macrophages. Therefore, the claim that clearance is mostly macrophage mediated could be due to an enrichment of this population during dissociation rather than this cell type being responsible for most of the extracellular waste disposal. Authors should quantify the % of total oScarlet that is specifically in macrophages in the brain sections they already have that are stained against oScarlet and CSF1R/ApoEB transcript.

      (4) The flow cytometry strategy used does not distinguish between oScarlet protein that has been internalized versus that which is sticking to the surface of macrophages. Authors should stain non-premeabilized and permeabilized cell suspensions with a flow antibody against mCherry/RFP to get a sense of how much oScarlet is inside versus outside of the macrophage. For most antibodies this can be done on the same sample sequentially if the antibodies have a different fluorophore.

      (5) It is concerning that dextran and oScarlet are almost perfectly colocalized in the image presented (Figure 3a). It raises the possibility, among others, that dextran is sticking to potential oScarlet aggregates and then being internalized by macrophages. Therefore, it could be an artifact of the transgenic line. Authors should repeat the experiment in wildtype fish and use HCR against CSF1R/ApoEB to address this issue.

      Significance:

      We believe that this is an important finding as such a model in African Killifish lays the groundwork to study the pathways that mediate the clearance of extracellular molecules by brain macrophages, the impact that this process has on brain homeostasis, and how it changes in aging. In particular, this reviewer is excited about the future potential of this model to uncover the molecular processes behind macropinocytosis, a process that occurs frequently in brain macrophages yet the mechanisms regulating it remain elusive, and how it contributes to overall brain health.

    4. Reviewer #3 (Public review):

      Summary:

      Rahul Nagvekar et al. generated a novel genetic model (SP-oScarlet) to label brain macrophages via their engulfment activity in the naturally short-lived African turquoise killifish. They found that these brain phagocytes exhibit transcriptional features resembling mammalian BAMs/MDMs and provided evidence that their engulfment capacity declines with age. The model and topic are interesting, but some of the central conclusions require more precise calibration to match the strength of the supporting evidence.

      Major comments:

      The SP-oScarlet model enriches cells based on phagocytic capacity - by design, any phagocytic cell, including microglia, can be labeled. Only 0.5% of oScarlet<sup>LOW</sup> cells were myeloid cells, confirming that this method captures virtually the entire myeloid population. The transcriptional resemblance to BAMs/MDMs is therefore a post hoc characterization of brain phagocytes broadly, rather than evidence for a selectively labeled subset. The authors show examples of apoeb<sup>+</sup> cells near vasculature (Fig. 3b), but do not provide a comprehensive quantification of the full spatial distribution of oScarlet<sup>HIGH</sup> cells. Importantly, neither the SP-oScarlet macrophages nor previously published wild-type killifish brain macrophages could be transcriptionally separated into three subgroups analogous to mammalian microglia, BAMs, and MDMs by PCA. This suggests that fish brain macrophages may not exist as subpopulations that correspond with their mammalian counterparts. The authors should therefore describe these cells as brain myeloid cells that exhibit BAM/MDM-like transcriptional characteristics, rather than implying they are a population equivalent to mammalian BAMs/MDMs.

      The age-related decline in oScarlet fluorescence in oScarlet<sup>HIGH</sup> cells in vivo could reflect either reduced phagocytic capacity of macrophages, or reduced oScarlet secretion by neurons, as the authors have discussed (Fig. 5a). The ex vivo assay addresses this by standardizing substrate concentration, which is a strength, but an in vivo functional assessment would provide a more physiologically relevant complement. The authors have already established the methodology for in vivo substrate injection (Fig. 3a, dextran). A similar experiment comparing substrate uptake in young and old fish would circumvent potential artifacts of the ex vivo approach, such as enzymatic dissociation altering surface receptor availability, and would directly test whether engulfment declines in the native brain environment.

      Significance:

      General assessment: This study presents a novel genetic model (SP-oScarlet) for visualizing chronic engulfment by brain macrophages in a short-lived vertebrate. The finding that killifish brain phagocytes exhibit BAM/MDM-like transcriptional features is interesting. Leveraging the killifish's naturally short lifespan, the authors further provide functional evidence that brain macrophage engulfment capacity declines with age. However, the authors should exercise caution when defining these cells as a distinct population specialized for engulfment of material from the brain extracellular space, since the SP-oScarlet model labels nearly the entire myeloid population in the brain.

      Advance: This study establishes a novel genetic model for chronic, in vivo visualization of engulfment in a vertebrate brain. The conceptual insight that killifish brain phagocytes transcriptionally resemble BAMs/MDMs rather than classical microglia is novel and may reflect evolutionary differences in brain clearance strategies.

      Audience: This research will interest a broad audience across developmental biology, genetics, neuroimmunology, aging research, and evolutionary biology.

    5. Author response:

      General Statements

      Please find below a revision plan for our manuscript entitled ‘Engulfment by brain macrophages in a short-lived vertebrate’ that was reviewed by Review Commons and that we wish to be considered for publication in eLife.

      We were very happy to see that all three Reviewers were interested in our study and we are sincerely grateful to all of them for their constructive suggestions, which we believe can largely be addressed and will improve our manuscript.

      We would like to inquire whether you would consider publishing our manuscript at eLife in its current form, along with the reviews. We will then provide an updated version of the manuscript, based on our revision plan below when we are able.

      Thank you so much for your consideration.

      Description of the planned revisions

      Reviewer #1:

      Major comments

      Comment 1: Although the authors describe and analyze their data from the viewpoint of engulfing macrophages, the paper would benefit from a broader perspective and a comparison to other studies on microglia in different species. Along this line, the title does not really seem to cover the data presented here very well, and the introduction lacks a proper explanation of terminology on microglia/brain macrophages and their known roles, cell types versus cell states and the current state of the art in fish versus other model species in the context of aging.

      We thank the Reviewer for these suggestions. We will expand the introduction to add more information and references on microglia/brain macrophages across fish and other species. We will also edit the title to more closely reflect the data in the manuscript.

      Comment 2: The result that nearly all myeloid cells in the killifish brain are of the engulfing macrophage type is somewhat surprising. This appears to differ from other studies in for instance zebrafish (e.g. ref 80, that describes the heterogeneity of the myeloid cells in detail). There are two questions we like to raise: (Q1) What is the evidence towards this homogeneity? and (Q2) Could there be a technical bias?

      We agree with the Reviewer, and we were also surprised to find a fairly homogenous myeloid population even in wildtype (non-transgenic) killifish brains, especially given that all these 3 different populations of myeloid cells (microglia, non-microglial macrophages, and dendritic-like cells), could be found in the adult zebrafish brain by Rovira et al. using single-cell RNA sequencing of cd45+ cells (ref. 80).

      We will better describe our evidence towards this homogeneity among killifish brain myeloid cells in the manuscript by adding a paragraph in the text at the end of the section referring to Figure 2. We will also provide additional experiments and analyses (see also below, our detailed responses to Q1 and Q2):

      i) In our and Ayana et al.’s single-cell RNA sequencing datasets, we could not find distinct populations of myeloid cells corresponding to microglia, non-microglial macrophages, and dendritic-like cells, even in wildtype (non-transgenic) killifish young adult and old brain, and even when subclustering only the myeloid cells. However, cd45+ cells were not specifically enriched in these experiments.

      ii) When using PCA, we also could not separate killifish brain macrophages into microglia, border-associated macrophages, and monocyte-derived macrophages (whereas we could separate mouse brain macrophages into these three groups as a positive control for this approach).

      iii) As noted by the Reviewer below, our analysis of the Ayana et al. wildtype dataset was limited to the young adult timepoint from this dataset. In the revised version of our manuscript, we will also include analysis of the old timepoint from this dataset.

      iv) As described in our responses to Q1 and Q2 below, we will provide more detailed subclustering that highlights the heterogeneity we do observe among killifish brain myeloid cells. We find this heterogeneity to be subtle and likely corresponding to different cell states rather than functionally distinct cell types. We will also use hybridization chain reaction (HCR) to test whether cd74, a key marker we use to propose a monocyte-derived macrophage-like identity for almost all detected killifish brain macrophages, labels apoeb+ cells in young wildtype brains in situ (as would be predicted based on our single-cell RNA sequencing analysis). We will additionally use HCR to test whether batf3, a key marker used by Rovira et al. to identify dendritic-like cells, labels any cells in the killifish brain (based on our single-cell RNA sequencing analysis, we do not expect to find batf3+ cells in the killifish brain).

      We agree that there could be technical confounds that could lead to the observed homogeneity of the myeloid population and lack of canonical microglia and dendritic-like cells even in wildtype killifish in our analysis. These could include our/Ayana et al.’s protocol for brain dissociation, our FACS-sorting step, or the step of loading cells into the 10x Genomics chip for single-cell RNA sequencing – each of these could be a step where canonical microglia and/or dendritic-like cells could be lost. In addition, we did not enrich for cd45+ cells, so it is possible that rare populations of immune cells in the killifish brain may not be represented in our dataset. As described below in our responses to Q1 and Q2, we will perform additional analyses that we hope will help address some of these key points, and we will more extensively discuss all possible technical confounds leading to potential cell type bias in the Discussion section.

      In addition to discussing potential technical confounds, we will also better highlight biological possibilities that could explain the differences between our and Rovira et al.’s findings. For example, the observed lack of microglia and dendritic-like cells in killifish brains could reflect a species (killifish vs. zebrafish) difference. Alternatively, even though both studies analyzed young adult timepoints, there could also be differences in biological age between the young adult killifish and zebrafish analyzed, which may be additionally influenced by husbandry conditions (e.g., feeding, pathogen exposure).

      Regarding (Q1): What is the evidence towards this homogeneity? The markers used are overlapping with markers for microglia. It would be helpful to clarify how canonical microglia populations are represented in the dataset. What is the heterogeneity of the oScarletHIGH cells? On several plots (Fig1f, Fig2d, Fig4a) this population of cells seems more heterogeneous than described. Are there different cell states or types? What is the percentage of myeloid cells that is oScarletLOW? To what extent do these cells compare transcriptionally to the oScarletHIGH cells?

      The Reviewer makes a series of excellent points. To address them:

      i) We will more prominently highlight markers not expressed by canonical microglia (e.g., mrc1, cd74) that are expressed by nearly all detectable killifish brain myeloid cells, even in wildtype datasets. We could not find canonical microglia (e.g., tmem119+, sall1+) in any of our killifish datasets. We will highlight these results in the text and will move some of the relevant graphs in the main figure.

      ii) We agree that oScarlet<sup>HIGH</sup> cells are somewhat heterogeneous. To better understand the source of this heterogeneity, we will present UMAPs/Seurat FeaturePlots focusing on cell state markers (e.g., cycling [mki67+] and activated [iba1+] cells) as well as specific cell type markers from mammalian/zebrafish literature (canonical microglia, non-microglia macrophages, dendritic-like cells, etc.). We expect this analysis to show distinct cell states, but fairly uniform expression of cell type markers across all oScarlet<sup>HIGH</sup> cells in our datasets.

      iii) The percentage of detected myeloid cells that are oScarlet<sup>LOW</sup> is very low (<1%). However, as noted in the next comment from the Reviewer, our sorting protocol de-enriches for oScarlet<sup>LOW</sup> myeloid cells, so the actual percentage of all myeloid cells that is oScarlet<sup>LOW</sup> may be higher. We will include additional panels to illustrate this point.

      iv) oScarlet<sup>LOW</sup> myeloid cells are largely transcriptionally similar to oScarlet<sup>HIGH</sup> cells (which are almost all myeloid). The main differentially expressed genes in oScarlet<sup>LOW</sup> vs oScarlet<sup>HIGH</sup> myeloid cells are markers of non-myeloid genes (e.g., elavl3, mpz). These could reflect phagocytosis of non-myeloid cells by oScarlet<sup>LOW</sup> myeloid cells but could also reflect ambient RNA contamination, since oScarlet<sup>LOW</sup> myeloid cells (but not oScarlet<sup>HIGH</sup> cells) were sequenced alongside non-myeloid cells. We will include the differential expression analysis and add commentary in the text to discuss these possibilities.

      Fig1f-i depict an enriched oScarletHIGH group alongside oScarletLOW cells. This representation is a bit misleading since it seems to indicate that really all myeloid cells are of the engulfing macrophage type whereas it is the majority, but not all.

      We agree with the Reviewer. We will also include different representations that provide more accurate estimates of the proportion of all myeloid cells that are oScarlet<sup>LOW</sup>/not engulfing macrophages.

      Line 52: The authors describe that the oScarletHIGH cell group is "enriched for signatures characteristic of macrophage functions". This finding is logical, as the isolation procedure of this population of cells was based on the endocytic and phagocytic properties of the cells. This result appears more consistent with a validation of the isolation strategy than with definitive evidence for myeloid cell identity.

      We agree and we will reframe this result as a validation of the isolation strategy.

      Regarding (Q2): Could there be a technical bias? An alternative explanation that may warrant discussion is whether aspects of the experimental pipeline (cell dissociation, FACS, scRNA-seq) could influence myeloid cell states. For instance, it is conceivable that dissociation induces a reactive program that enhances uptake of fluorescent protein, potentially enriching for oScarletHIGH cells. As the authors use a similar experimental setup to prove uptake of dextran and ovalbumin, such a technical artefact may merit consideration. As this would influence the major conclusions of the paper, the authors might want to address this comment with additional experimental controls, such as single-nuclei RNA-seq on control young and aged brains to profile the natural myeloid population when not submitted to a cell dissociation and FACS procedure.

      We agree with the Reviewer that brain dissociation, FACS, and single-cell RNA sequencing could all represent technical biases that could all influence the representation of different myeloid cell types in our dataset. We will clearly note this potential confound in the Results and the Discussion.

      We agree that it would be valuable to test some of these issues experimentally. We will perform non-dissociative HCR (in situ) to test (1) whether in young wildtype brains, apoeb+ cells (macrophages) express cd74 (a key marker for our argument re: MDM-like cell type identity) and (2) whether we can find cells expressing batf3 (a key marker used by Rovira et al. to identify dendritic-like cells, which we cannot detect in the killifish brain in our single-cell RNA sequencing data). We will also acknowledge that these are only individual markers and would not provide as comprehensive a picture of cell type identity as single-nuclei RNA sequencing.

      Comment 3: The authors compare the oScarletHIGH cell transcriptomes to mouse and killifish datasets. Both the mouse (Barr et al) and killifish (Nagvekar, this paper) dataset are from enriched immune cells (mouse= CD45+ cells, and the 3 cell types selected from that). Why did the authors not compare to the whole mouse CD45+ dataset? Including zebrafish (Rovira et al, 2025) here would strengthen the evolutionary comparison. I also feel that the additional comparison with young killifish (Ayana et al) might not be that solid since this dataset was initially not enriched and has a significantly lower number of myeloid cells, and thus much less power. The old age time point in that study contained more myeloid cells, and might be interesting to include for cell type comparison. There are other, perhaps more unbiased ways of comparing cell types across species, for instance SAMap, developed by co-author Bo Wang. Did the authors consider using this or other methods?

      We thank the Reviewer for these excellent suggestions. We will do the following in our revised manuscript:

      i) We will include comparisons to all immune (CD45+) cells from mice (Barr et al.).

      ii) We will include comparisons to all immune (cd45+) cells in zebrafish (Rovira et al.).

      iii) We will also analyze the old age time point from the Ayana et al. wildtype killifish dataset.

      The percentage of oScarletHIGH cells in the aged condition is 8% (Fig1-suppl1) compared to 4% at young age. On the other hand, a lower number of cells was isolated at old age compared to young age (Figure4a). Can the authors elaborate on this difference? Later on, it is stated that the engulfing capacity declines with aging, but could this be linked to the lower or potentially biased recovery of cells?

      We appreciate the Reviewer’s point. We will elaborate on the percentage differences in oScarlet<sup>HIGH</sup> cells recovered in the different experiments, and more clearly indicate what they could originate from and whether they could influence the engulfment results:

      i) We will indicate more clearly that in each single-cell RNA sequencing experiment, we loaded the entire set of Live oScarlet<sup>HIGH</sup> cells collected from each condition onto the 10x Genomics chip. In the young vs. old comparison, more Live old oScarlet<sup>HIGH</sup> cells were loaded than Live young oScarlet<sup>HIGH</sup> cells (66K vs. 42K, based on the FACS sorter counts, though these numbers are likely overestimates). One possibility is that the lower recovery of successfully sequenced old oScarlet<sup>HIGH</sup> cells might be explained by differences with age in oScarlet<sup>HIGH</sup> cells’ ability to remain intact during loading onto the 10x Genomics chip.

      ii) In the ex vivo experiments where we find that engulfment capacity declines with age (current Fig. 5), the proportion of oScarlet<sup>HIGH</sup> cells among all Live cells slightly increased with age. We will include the data illustrating this, and indicate more clearly in the text that the differences in engulfment capacity observed in these experiments are unlikely to be explained solely by lower recovery of oScarlet<sup>HIGH</sup> cells from old brains.

      iii) We will more clearly acknowledge in the Discussion the possibility of potentially biased cell recovery with age.

      Figure4a: Transcriptional differences are stated between young and old (line 226), can a relevant selection be shown in e.g. a dotplot or heatmap?

      This is another great point from the Reviewer. We will show these genes in a dotplot/heatmap in a Supplemental Figure. We will also clearly indicate in the main text that many of the genes most strongly enriched in old oScarlet<sup>HIGH</sup> cells in our young vs. old comparison experiment were not strongly expressed in an independent old oScarlet<sup>HIGH</sup> cells dataset (which was compared to old wildtype cells, without a young counterpart). 

      The UMAP clustering does seem to indicate batch effects on panels a and g. Can the authors provide subclustering and show that young and old/ FACS sorted high and low cover similar cell types/states? The PCA plot (panel f) and marker analysis is not fully convincing, as PC1 and 2 alone do not suffice to explain all the variance in these cells, and the markers are common ones for many microglia/macrophage cell types (and thus likely to be expressed similarly).

      The Reviewer has another excellent suggestion. We will provide subclustering, which indeed shows that young and old oScarlet<sup>HIGH</sup> cells generally represent similar cell types and states. This analysis also shows a subcluster specific to old oScarlet<sup>HIGH</sup> cells, but this subcluster is far less pronounced in a second old oScarlet<sup>HIGH</sup> dataset (see our previous comment), so we will clearly indicate in the main text that this subcluster may not be robust. We will also remove the PCA plot.

      Figure 5: It would be informative to include the corresponding aged condition for panels c and e.

      We agree and we will include this.

      Minor comments

      The authors use the oScarlet fish in a heterozygous state. Is the homozygous line not viable? Or what is the reason to use the heterozygous state? Too high expression of OScarlett? Is apoptosis of neurons checked for this line? Compared to wild type?

      The homozygous SP-oScarlet line has not been characterized, and we will include this information in the Results. We did not check apoptosis of neurons in the SP-oScarlet line or in wildtype killifish, and we will indicate this in the Methods and Discussion.

      Is the secretion of OScarlett completely proven? Maybe the signal is engulfed by the clearance of cell debris from apoptotic neurons?

      This is a great point. We have not directly tested the secretion of oScarlet in the SP-oScarlet line, and it remains possible that engulfed oScarlet also comes from apoptotic neurons that are being engulfed by myeloid cells. We will acknowledge this possibility in the Discussion.

      We will also more clearly highlight references that adding a signal peptide to proteins (as we did for the SP-oScarlet line) is sufficient to induce their secretion in several contexts in different species (although this does not prove secretion in this particular case).

      We also do have another line (built for a separate project) in which oScarlet is expressed cytoplasmically in elavl3+ cells. In this line, in pilot data, we did not observe oScarlet<sup>HIGH</sup> cells by flow cytometry. While these experiments are not of publishable quality, these observations also suggest that the presence of a signal peptide on oScarlet is necessary for the engulfment phenotype.

      Line 329. What exact difference between teleosts and mice are you pointing at?

      We will edit the text to clarify that we are referring to our inability to find a cell population with a transcriptional profile like that of canonical adult mouse microglia in the adult killifish brain.

      Suggestions for Figures

      General remark IHC/HCR figures: Please add overview figures to guide the reader to ROI shown.

      Fig1.

      1.b Please add overview figures to show the overall distribution of the signal in the brain. Are there any hotspots or low-abundant regions?

      1.d/1.e Please add numbers of cells on figure panels.

      1.d: overview figure is unclear. Telencephalon seems to be missing? Can annotation be added to regions of the structure?

      Additional in vivo evidence of the homogeny/heterogeneity of oScarlet protein+/RNA- cells would be informative, e.g. double labeling (HCR) with some of the markers of panel Fig. 2a). This also would exclude location bias. One could expect myeloid cells that are not in close proximity to secreting neurons, for instance in dense neurogenic niches

      Fig. 2

      2.c Please add number of cells on graph

      2.d Please add brain regions to graphs of published datasets like was done for the own dataset.

      Subtext: Why is there a different number of cells for Barr in c and d? Please add number of cells of own killifish dataset in legend.

      Supplement fig. 2 Please add additional microglia-specific markers such as for instance HEXB, GPR34, SELL1, C1Q.

      Fig.3

      Please add overview pictures.

      Fig.4

      4.g grey color not very visible

      It would be informative to have the markers of 4.h plot on top of the UMAP to show the distribution/differences, maybe in supplementary information.

      We will implement all of these changes. We will add cd74 HCR labeling of oScarlet<sup>HIGH</sup> cells that do not express oScarlet RNA transcripts in situ. The different numbers of cells from the Barr et al. dataset in the current panels 2c and 2d do not reflect differences in brain regions from which macrophages were isolated (in both cases, the same whole-brain dataset was used). Instead, in this dataset, there is a small population of interferon-responsive macrophages that are not microglia, BAMs, or MDMs; these cells were included in the analysis in 2d but not 2c. We will clarify this and update the analysis in current panel 2d to include all immune cells from the Barr et all. dataset, as suggested by the Reviewer above.

      Reviewer #2:

      Red fluorescent proteins are notorious for being prone to aggregation. Are oScarlet proteins being internalized by macrophages aggregates or soluble proteins? This distinction is important as the clearance of extracellular molecules could be mediated by most cells yet aggregates could be removed specifically by macrophages. Can experiments be conducted to distinguish between these two possibilities? We realize this may be challenging. If not feasible, the discussion should be tempered to reflect this possibility.

      The Reviewer makes a great point, and we were also interested in this question. mScarlet and its derivatives (including oScarlet) would generally be expected to remain in the monomeric state to a greater extent than other red fluorescent proteins like mCherry (Bindels et al., Albakri et al.). We will indicate this with references in the Results section.

      But this does not preclude the possibility of oScarlet aggregation in SP-oScarlet brains. We did collect soluble and insoluble fractions from SP-oScarlet brains using a gentle lysis method that would be expected to preserve aggregates in the insoluble fraction (Avar et al.). However, we found that the levels of oScarlet in both fractions were below the limit of detection by western blot and by a plate reader-based fluorescence assay, thereby precluding us from directly testing how much oScarlet was aggregated. We will address the possibility of oScarlet aggregation, and its potential impact on engulfing macrophages, in the Results and Discussion sections.

      Brain dissociation tends to generate a lot of debris, especially from sheared neurons. Therefore, the high level of oScarlet inside macrophages could be an artifact of dissociation rather than a reflection of in vivo clearance. Authors should use internalization inhibitors during dissociation (CytoD, Dynasore, and pitstop) to exclude this possibility. Alternatively, if they have a transgenic killifish that expresses another fluorescent reporter in neurons (and preferably at a similar level to that of oScarlet), authors should dissociate brains together and quantify how many oScarlet+ cells are now also positive for that other fluorescent reporter. This could give an idea of how much engulfment is occurring due to the dissociation processes. It is not ideal, as macrophage eating could be happening during dissociation but before cells are in single cell suspension. However, given that RNAseq is needed to identify macrophages, this reviewer would be satisfied by this alternative approach if the aforementioned pitfall is also presented in the discussion.

      This is another excellent suggestion, and we have already done the following experiments:

      i) We have performed bulk RNA sequencing from FACS-sorted oScarlet<sup>HIGH</sup> cells from middle-aged SP-oScarlet brains dissociated with five inhibitors (cytochalasin D, dynasore, Pitstop 2, bafilomycin A, and wortmannin) in addition to transcription/translation inhibitors. Deconvolution analysis of these bulk RNA sequencing data showed that oScarlet<sup>HIGH</sup> cells in the presence of cytochalasin D, dynasore, Pitstop 2, bafilomycin A, and wortmannin are also almost all macrophages. While these data are not at single-cell resolution, they indicate that engulfment by macrophages is unlikely to solely occur during the dissociation process. We will include a revised Supplementary Figure with these experiments.

      ii) We have also already performed a pilot co-dissociation experiment (in a different genetic background) as suggested by the Reviewer using a ubb:GFP reporter as the second transgenic and observed very few GFP+ oScarlet<sup>HIGH</sup> cells (see Author response image 1). We will include co-dissociation data in the revised version of the manuscript.

      Author response image 1.

      We believe that the results of these experiments are consistent with the notion that oScarlet<sup>HIGH</sup> cells have engulfed oScarlet in vivo, prior to dissociation step. Nevertheless, we will also make note in the Discussion of the pitfall mentioned by the Reviewer.

      Related to the above, it appears based on the scRNAseq that dissociation heavily enriched for brain macrophages. Therefore, the claim that clearance is mostly macrophage mediated could be due to an enrichment of this population during dissociation rather than this cell type being responsible for most of the extracellular waste disposal. Authors should quantify the % of total oScarlet that is specifically in macrophages in the brain sections they already have that are stained against oScarlet and CSF1R/ApoEB transcript.

      The Reviewer’s point is well taken and we agree that experiments in intact brain sections are important to orthogonally test whether macrophages are responsible for engulfment. As suggested by the Reviewer, we will quantify the percentage of total oScarlet that is in apoeb+ macrophages in the brain sections we already have (current Fig. 3a). We note that it may not accurately reflect the actual in vivo percentage, as we have found that it can be challenging to draw cell boundaries around macrophages due to their irregular shapes.

      It is concerning that dextran and oScarlet are almost perfectly colocalized in the image presented (Figure 3a). It raises the possibility, among others, that dextran is sticking to potential oScarlet aggregates and then being internalized by macrophages. Therefore, it could be an artifact of the transgenic line. Authors should repeat the experiment in wildtype fish and use HCR against CSF1R/ApoEB to address this issue.

      We agree with the Reviewer. We have already injected dextran into non-transgenic killifish brains and observed dextran engulfment by apoeb+ cells. We will include this experiment in the revised version of the manuscript.

      Reviewer #3:

      Major comments

      The SP-oScarlet model enriches cells based on phagocytic capacity - by design, any phagocytic cell, including microglia, can be labeled. Only 0.5% of oScarlet<sup>LOW</sup> cells were myeloid cells, confirming that this method captures virtually the entire myeloid population. The transcriptional resemblance to BAMs/MDMs is therefore a post hoc characterization of brain phagocytes broadly, rather than evidence for a selectively labeled subset. The authors show examples of apoeb<sup>+</sup> cells near vasculature (Fig. 3b), but do not provide a comprehensive quantification of the full spatial distribution of oScarlet<sup>HIGH</sup> cells. Importantly, neither the SP-oScarlet macrophages nor previously published wild-type killifish brain macrophages could be transcriptionally separated into three subgroups analogous to mammalian microglia, BAMs, and MDMs by PCA. This suggests that fish brain macrophages may not exist as subpopulations that correspond with their mammalian counterparts. The authors should therefore describe these cells as brain myeloid cells that exhibit BAM/MDM-like transcriptional characteristics, rather than implying they are a population equivalent to mammalian BAMs/MDMs.

      We agree with the Reviewer and will describe killifish brain macrophages as “brain myeloid cells that exhibit BAM/MDM-like transcriptional characteristics.”

      As described in our responses to Reviewer 1’s comments, we will also provide/more prominently highlight analysis of killifish brain myeloid cell datasets showing:

      i) Markers (e.g., mrc1, cd74) that are shared between mammalian BAM/MDMs and killifish brain myeloid cells.

      ii) The heterogeneity we observe among killifish brain myeloid cells, which we believe corresponds to different cell states, but is not likely pronounced enough to correspond to functionally distinct cell types.

      iii) The presence of oScarlet<sup>LOW</sup> killifish brain myeloid cells, which are de-enriched by our sorting strategy.

      Minor comments

      The authors acknowledge that oScarlet is relatively resistant to degradation and could affect the state of cells that engulf it. The independent single-cell experiment comparing old SP-oScarlet and old wild-type macrophages addresses this concern at the transcriptional level. However, it remains possible that chronic accumulation of degradation-resistant protein in lysosomes could itself impair phagocytic capacity. If so, the age-related decline in ex vivo engulfment might be partially attributable to oScarlet burden accumulated over a lifetime, rather than to aging per se. While a direct functional comparison with old wild-type macrophages is experimentally challenging, the authors should at least acknowledge this interpretive caveat in the Discussion.

      The Reviewer makes a great point, and we will acknowledge this caveat in the Discussion.

      The transcriptional data show only modest changes in engulfment- and lysosome-related pathways with age. The authors should offer one or more specific, testable hypotheses (e.g., which phagocytic receptors or lysosomal components might be affected) to guide future investigation.

      We will provide a list of top engulfment-relevant genes that change the most with age, which also show only modest changes. We will better acknowledge that the age-related transcriptional changes in engulfment-relevant genes observed in our single-cell RNA sequencing experiment are modest and are unlikely to directly serve as a resource for testable hypotheses. We will also move our single-cell RNA sequencing data to the Supplementary Figures.

      In the text, gene names such as APOEB are capitalized, whereas the convention for fish gene nomenclature is lowercase italics (e.g., apoeb). The authors should ensure gene formatting follows species-appropriate conventions throughout the manuscript.

      We thank the Reviewer for this suggestion, and we will change killifish gene names to be lower-case italicized throughout the manuscript.

      Description of analyses that authors prefer not to carry out

      Reviewer #1:

      Regarding (Q2): Could there be a technical bias? An alternative explanation that may warrant discussion is whether aspects of the experimental pipeline (cell dissociation, FACS, scRNA-seq) could influence myeloid cell states. For instance, it is conceivable that dissociation induces a reactive program that enhances uptake of fluorescent protein, potentially enriching for oScarletHIGH cells. As the authors use a similar experimental setup to prove uptake of dextran and ovalbumin, such a technical artefact may merit consideration. As this would influence the major conclusions of the paper, the authors might want to address this comment with additional experimental controls, such as single-nuclei RNA-seq on control young and aged brains to profile the natural myeloid population when not submitted to a cell dissociation and FACS procedure.

      We believe that generating our own single-nuclei RNA sequencing dataset would be beyond the scope of this study. However, a preprint by Williams et al. from the Benayoun Lab (https://doi.org/10.64898/2026.04.09.717549) includes single-nuclei RNA sequencing data from wildtype killifish brains and could provide an excellent test of our findings. We will point readers to this preprint in the Discussion.

      Comment 3: The authors compare the oScarletHIGH cell transcriptomes to mouse and killifish datasets. Both the mouse (Barr et al) and killifish (Nagvekar, this paper) dataset are from enriched immune cells (mouse= CD45+ cells, and the 3 cell types selected from that). Why did the authors not compare to the whole mouse CD45+ dataset? Including zebrafish (Rovira et al, 2025) here would strengthen the evolutionary comparison. I also feel that the additional comparison with young killifish (Ayana et al) might not be that solid since this dataset was initially not enriched and has a significantly lower number of myeloid cells, and thus much less power. The old age time point in that study contained more myeloid cells, and might be interesting to include for cell type comparison. There are other, perhaps more unbiased ways of comparing cell types across species, for instance SAMap, developed by co-author Bo Wang. Did the authors consider using this or other methods?

      We thank the Reviewer for the SAMap suggestion. We did perform SAMap with a mouse reference from the Allen Brain Atlas. Our SAMap analysis identifed distinct microglia-, BAM-, and dendritic-like myeloid populations. However, we were not confident in these SAMap results because the differences between the populations were subtle and we thought they would be unlikely to represent differences between functionally distinct cell types. Of note, the dendritic-like cells we identified by SAMap in the killifish brain did not express canonical markers of dendritic/dendritic-like cells from mammalian and zebrafish literature (e.g., batf3). Additionally, other SAMap studies from the Wang Lab analyzing non-neural cell types in other vertebrates do not separate microglia from non-microglial macrophages (Kalakuntla et al., in preparation).

      Comment 4: Regarding the comparison with the aged brain:

      Figure 4: It would be nice to include the same comparisons as for young fish (cfr Fig.1 panels F-I).

      We agree with the Reviewer. Unfortunately, we did not sequence oScarlet<sup>LOW</sup> cells from old brains, for cost reasons, and this precludes the comparison requested by the Reviewer. We will more clearly indicate that we do not have the oScarlet<sup>LOW</sup> cells from old brain in the Results section.

      Reviewer #2:

      The flow cytometry strategy used does not distinguish between oScarlet protein that has been internalized versus that which is sticking to the surface of macrophages. Authors should stain non-premeabilized and permeabilized cell suspensions with a flow antibody against mCherry/RFP to get a sense of how much oScarlet is inside versus outside of the macrophage. For most antibodies this can be done on the same sample sequentially if the antibodies have a different fluorophore.

      The Reviewer makes an important point, and we will address this caveat in the Methods. We have performed a pilot of the flow cytometry experiment suggested by the Reviewer and, encouragingly, observed a higher proportion of cells labeled by the antibody (the same anti-mCherry antibody we used to label oScarlet in situ) in the permeabilized condition. However, we do not wish to publish these results as even in the permeabilized condition, the antibody labeled <0.2% of cells – meaning it likely did not label most oScarlet<sup>HIGH</sup> cells, making it difficult to draw strong conclusions about internalized vs. surface oScarlet in these cells.

      Reviewer #3:

      The age-related decline in oScarlet fluorescence in oScarlet<sup>HIGH</sup> cells in vivo could reflect either reduced phagocytic capacity of macrophages, or reduced oScarlet secretion by neurons, as the authors have discussed (Fig. 5a). The ex vivo assay addresses this by standardizing substrate concentration, which is a strength, but an in vivo functional assessment would provide a more physiologically relevant complement. The authors have already established the methodology for in vivo substrate injection (Fig. 3a, dextran). A similar experiment comparing substrate uptake in young and old fish would circumvent potential artifacts of the ex vivo approach, such as enzymatic dissociation altering surface receptor availability, and would directly test whether engulfment declines in the native brain environment.

      We agree with the Reviewer and performed a pilot of the suggested experiment, which did not show age-related differences in in vivo engulfment of injected dextran. However, in developing the brain injection procedure, we concluded that this procedure works well for comparisons within a sample, but not for comparisons across samples and conditions (e.g., young vs. old). This is because the injection site is determined by sight (aiming for the most medial point on the telencephalon/optic tectum border, with no standardized way to control injection depth), and not stereotactically with coordinates. With our current procedure, it is challenging to perform reproducible injections across ages due to known age-related differences in fish/brain size and skull thickness.

      We are interested in developing stereotactic injection approaches that would be compatible with comparisons across ages and other conditions, but we believe that this is beyond the scope of this manuscript. We will discuss this possibility as a future direction in the manuscript.

    1. eLife Assessment

      This important study builds on previous work from the same authors to present a conceptually distinct workflow for cryo-EM reconstruction that uses 2D template matching to enable high-resolution structure determination of small (sub-50 kDa) protein targets. The paper provides convincing evidence that the density for small-molecule ligands bound to such targets can be reconstructed without these ligands being present in the template, and it provides an insightful quantification of the level of model bias that remains in the resulting reconstructions. The convenient visualisation of small molecules bound to protein targets of a known structure will be relevant for the pharmaceutical industry.

    2. Reviewer #3 (Public review):

      Summary:

      Due to the low SNR of cryo-EM micrographs necessitated by radiation damage, determining the structure of proteins smaller than 50 kDa is exceedingly challenging, such that only a handful have been solved to date. This work aims to improve the reconstruction of small proteins in single-particle cryo-EM by using high-resolution 2D template matching, an algorithm previously used to locate and align macromolecules in situ, to align and reconstruct small proteins. This approach uses an existing macromolecular structure, either experimentally determined or predicted by AlphaFold, to simulate a noise-free 3D reference and generates whitened projections, crucially including high-spatial-frequency information, to align particles by the orientation with maximal cross-correlation. They demonstrate the success of this approach by generating a 3D reconstruction from an existing dataset of a 41.3 kDa protein kinase that had previously evaded attempts at high-resolution structure determination. To alleviate concerns that this is purely from template bias, they demonstrate clear density at two regions that were not present in the template: 6 residues in an alpha helix and an ATP in the ligand binding pocket. The latter is particularly important for its implications in determining structures of ligand-bound proteins for drug discovery. They also produce a composite omit map from 36 partial-deletion reconstructions spanning the entire protein, demonstrating a reconstruction can be obtained without template bias. Additionally, the authors provide an update to the classic calculation in Henderson 1995 to predict the minimum molecular mass of a protein that can be solved by single-particle cryo-EM.

      Strengths:

      I am in no doubt that this technique can be used to gain valuable insights into the structures of small proteins, and this is an important advancement for the field. It is complementary to single-particle cryo-EM and provides an extra tool for the experimentalist that may work better in certain cases. For cases where only a small region of the structure is of interest, such as in drug screening, this method provides a simple workflow to screen many structures.

      The claim that using high-spatial frequency information is essential for aligning small proteins is a valuable insight. A recent pre-print published at a similar time to this manuscript used high-resolution information in standard ab-initio reconstruction to generate a high-resolution reconstruction from the same dataset, supporting the claims made in the manuscript.

      The theoretical section outlined in the appendix is also theoretically sound. It uses the same logic as Henderson, but applies more up-to-date knowledge, such as incorporating dose-weighting and altering the cross-correlation based noise estimation. This update is valuable for understanding factors preventing us from reaching the theoretical limit.

      Weaknesses:

      This method is a complementary technique to determine the structure of small macromolecules to existing methods such as Blush regularization and HR-HAIR. Although the authors have demonstrated convincingly that their method selects a stack of high-quality particles, it is less clear whether it performs better than RELION when using the same stack of particles, particularly in the ATP binding pocket. As the authors discuss, systematic benchmarks comparing these methods over more targets than the one presented here, will be important for determining the utility of this method.

      The method presented here also introduces template bias. Omit maps are used to reduce template bias by removing the region of interest from the template. Producing a full reconstruction through a composite omit map is computationally expensive and can introduce artifacts at boundaries. Therefore, unless this method outperforms modern SPA methods, its major use case will likely be restricted to ligand binding studies rather than full 3D reconstructions.

    3. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      The evidence described for the claim that this technique improves the alignment of the reconstruction of small complexes compared to standard techniques is incomplete. The authors could better evaluate the effects of model bias on the reconstructed densities, as suggested by reviewer #1.

      We thank the editors for highlighting this remaining concern. To better evaluate the effects of model bias, we have performed the FSC-based analyses suggested by Reviewer 1, including a map–model FSC of the omit map and an FSC between the half-maps (though the latter is unreliable for the composite), and added the results to the revised manuscript (new Figure 5—figure supplements 1–3; detailed under Requested FSC Analysis below). We have additionally revised the text to avoid overstating the reconstruction as “unbiased” and to soften comparisons with other reconstruction methods (detailed below).

      Public Reviews:

      Reviewer #1 (Public review):

      In the revised version, the refinement of atomic occupancies in the 2DTM-generated maps has been insightful: densities only come back at values ranging from 0.55–0.80, whereas residues included in the template remain at 1, suggesting that the 2DTM-reconstruction does suffer from model bias. Their newly added Omega calculations, which are helpful, also suggest that model bias is present in the 2DTM-based reconstructions. These observations therefore contradict the first subsection heading of the Results, which claims “unbiased reconstruction of omitted residues”.

      We agree that the previous subsection heading was too absolute. We have changed the heading from “Unbiased reconstruction of omitted densities in a 43 kDa protein kinase” to “Recovery of densities omitted from the template in a 43 kDa protein kinase”. The opening sentence now states that we evaluated the ability of 2DTM to “recover omitted ligand densities” (L123–125).

      We also revised nearby statements to describe the observations without claiming that the entire reconstruction is free of template bias. The manuscript now states that, because the corresponding features were omitted from the search template, the recovered densities cannot result from direct inclusion of those features in the template (L150–153).

      For the omitted alpha-helical turn, we now state simply that its density was recovered despite its absence from the search template (L191–192). We also revised the interpretation of the occupancy-refinement results from “confirming partial, unbiased recovery” to “supporting partial recovery of density in the omitted regions” (L198–199). The same wording has been applied to the Supplementary file 1 caption on page 27.

      Finally, the summary of the ligand-deletion experiments now states that a ligand and nearby residues can be deleted to reduce template bias while retaining sufficient signal for their density to be recovered (L248–251). In the Methods, we retain the description that the composite omit map was constructed to avoid template bias at the omitted locations (L953–954).

      We have also toned down comparative statements about reconstruction accuracy, including the relevant subsection heading (“Comparison of 2DTM and RELION reconstructions from the same particle stack” at L252–253) and the surrounding discussion (L274–279).

      Requested FSC Analysis

      The measurement of how much model bias is present in this OMIT map by FSC calculations is still pending. This could be done in two ways. My original suggestion was to calculate a mapto-model FSC for the OMIT map and the full reference. This should be compared with a similar map-to-model FSC on the map where only the ligand was omitted. Alternatively, they can use the cisTEM FSC uncorr procedure on the OMIT half-reconstructions and compare the resulting curve with the one presented in Figure 1b.

      We have now completed both analyses and added them to the revised manuscript (new Figure 5—figure supplements 1–3, with accompanying Results and Methods text).

      (1) Map–model FSC. We computed the map–model FSC between the composite OMIT map and a density simulated from the full 1ATP model, and compared it with the equivalent FSC for the Figure 1 reconstruction, in which the ligand and residues 222–227 were omitted from the template (Figure 5—figure supplement 1). The composite OMIT map crossed FSC = 0.5 at 3.5 Å and FSC = 0.143 at 3.0 Å, compared with 3.0 Å and 2.4 Å, respectively, for the Figure 1 reconstruction. Because each local region of the composite map was taken from a reconstruction in which the corresponding residues were absent from the template, this agreement reflects genuine recovery rather than direct inclusion of those local features in the template. The lower FSC values relative to the Figure 1 reconstruction are expected. In the Figure 1 reconstruction, most of the protein remained in the template, whereas the composite is assembled from disjoint local omit regions. This stitched construction introduces holes and mask boundaries that affect Fourier-space agreement across the curve, in addition to the weaker, partial recovery of locally omitted density. Thus, the map–model FSC provides a conservative Fourier-space assessment of recovered omit-region density. The composite map–model FSC also shows a negative dip at the lowest spatial frequencies, which does not indicate failed recovery. Radial binning around the omitted atoms (Supplementary file 3) shows that the atom-centred shells (≤2 Å) recover positive but weakened density upon omission, whereas the peripheral shells (2–3 Å) are negative and nearly identical whether the residue is present or omitted, leading to net-negative density within the molecular envelope. We therefore interpret the low-frequency dip as a consequence of the composite construction rather than as evidence for failed recovery of omitted density.

      (2) Half-map FSC<sub>uncor</sub>. We also computed the half-map FSC of the individual OMIT reconstructions (Figure 5—figure supplement 2): the 36 individual omit reconstructions cross FSC = 0.143 at a median resolution of 2.8 Å, comparable to the Figure 1b reconstruction (∼3.0 Å). The composite map’s own half-map FSC appears higher, but this value is not a reliable resolution estimate because the two composite half-maps are assembled using the same voxel-assignment masks. This shared support introduces artificial correlations, which we demonstrate with a phase-randomization control (Figure 5—figure supplement 3). We therefore use the individual omit-reconstruction FSCs and the composite map–model FSC, rather than the composite half-map FSC, to assess Fourier-space agreement.

      Together, these analyses show that the OMIT reconstructions contain high-resolution signal in regions absent from the corresponding search templates, while also identifying the low-frequency dip and the inflated composite half-map FSC as consequences of the conservative stitched composite construction.

      Reviewer #3 (Public review):

      Nor was it compared to more recent strategies for processing SPA data from small molecules, such as Blush regularization or HR-HAIR. [...] This places this method as a complementary technique, and whether it outperforms those methods for a wide variety of molecules is yet to be determined.

      We agree that a systematic comparison with recent small-particle SPA methods such as Blush regularization and HR-HAIR will be important. We have added this point to the Discussion (L831–846). We also note that such as comparison should consider not only particle stack and molecular mass, but also the fidelity of the 2DTM template forward model. In the ideal limit of an accurate template and forward model, 2DTM should provide a strong prior for particle detection and pose determination. In practice, however, current templates remain imperfect approximations to the experimental signal because of inaccurately modelled solvent-boundary effects and atomic scattering factors, bonding and charge redistribution, and conformational mismatch. Improving template generation is therefore an important direction for extending the range of molecular targets and imaging conditions where 2DTM can be applied.

    1. eLife Assessment

      This study demonstrates a critical role of the glycolipid membrane protein insertase (MPIase) in the twin-arginine translocation (Tat) pathway, a protein translocation system that is conserved across all domains of life. The successful reconstitution of the bacterial Tat system in both bacteria-derived and artificial liposomes provides solid experimental evidence that MPIase has a broader role in membrane protein translocation than previously recognized. By revealing the essential function of a nonproteinaceous membrane component in catalyzing a core cellular process, this work offers fundamental insights into the molecular mechanisms of protein translocation.

    2. Reviewer #1 (Public review):

      Hanako and colleagues demonstrated that glycolipid MPIase is essential for the TAT system, and they successfully reconstituted the TAT system in vitro for the first time. This will facilitate the understanding of the mechanism of the TAT system.

      My major points are listed below for the authors to consider:

      (1) The authors successfully reconstituted the TAT system using the purified TatA/B/C, but the translocation efficiency was much lower than that of native INV. The authors partly attributed this to the reason that "MPIase recovery would be too low to detect the TAT activity" in the Discussion part. So, what would happen to the translocation efficiency if you added more MPIase to the reconstituted system? How about the abundance of MPIase from the INV and reconstituted proteoliposomes?

      (2) Why were only TatC levels measured in Figure 2C, whereas the expression levels of TatA were not detected? Also, from my observation, the amount of TatC in the third lane is lower than that in the previous two lanes.

      (3) The authors should explain why the TatA/B/C ratios in Figure 3C (1:1:1) and Figure 3D (10:1:1) are inconsistent.

      (4) ~30% of the fluorescence was recovered in the membrane fraction (Figure 4A) both in the functional TAT signal sequence (RR) and in the inactivating mutant signal sequence (KK), which suggests that MPIase acts as a relatively broad recognition factor. Given that MPIase does not discriminate between RR and KK, why do un-translocated substrates remain in the cytoplasm rather than non-specifically adhering to the membrane when MPIase is depleted in vivo?

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigated the relationship between the Tat system and MPIase, a glycolipid that facilitates protein integration into the bacterial cell membrane. The TAT (twin-arginine translocation) system is a unique membrane transport machinery that exports fully folded proteins containing a twin-arginine signal peptide. Using both in vivo and in vitro approaches, the authors demonstrated that a sufficient amount of MPIase is required for Tat-dependent protein translocation. Furthermore, the authors successfully reconstituted the Tat transport system by combining recombinant TatA, TatB, TatC, MPIase, and FoF1-ATP synthase.

      Strengths:

      The reconstituted system clearly demonstrated the requirement for each component, as substrate translocation occurred only when all components were present. Based on these findings, the authors proposed a mechanistic role for MPIase in facilitating Tat-mediated membrane translocation. Previous studies have shown that MPIase is involved in Sec-dependent protein translocation and membrane protein integration, as well as YidC-dependent membrane insertion. The present study further demonstrated that MPIase also plays an essential role in the Tat translocation pathway. Overall, this work highlights the central importance of MPIase in bacterial membrane protein biogenesis and provides new insights into the molecular mechanism of Tat-dependent protein transport.

      Weaknesses:

      (1) To show the importance of the Tat system in bacterial cells, it would be good to describe in the introduction how many proteins are translocated via the Tat system.

      (2) Figure 2B and D show that a sufficient amount of MPIase is important in SufI translocation. However, the reason why MPIase level was upregulated in the BL21 strain but not in the KS46 strain remains unexplained. The authors should address this point.

      (3) In Figures 4A and B, the authors explain that MPIase first works as a receptor of TorA-GFP without recognizing the RR motif. This conclusion is based on the results of the fractionation assays, where "sup" indicates the cytoplasmic and periplasmic fractions, and "ppt" indicates the membrane fraction. In Figure 4B, under the TatABC+++, (RR), +MPIase condition, the substrate is secreted most efficiently via the Tat pathway and should therefore be recovered in the periplasm fraction (sup). However, the authors point out that efficiently processed substrate was recovered in the ppt fraction rather than the sup fraction. The authors should explain why this occurred.

    1. eLife Assessment

      The authors provide convincing evidence that, during early mouse embryogenesis, distal visceral endoderm cells migrate in a coordinated, intermittent start-and-stop manner along an axis that correlates with the asymmetric morphology of the ectoplacental cone. This fundamental finding is supported by high-quality live imaging and quantitative image analysis, while computational modelling suggests that the observed intermittent behavior may be explained by the movement of a cohesive cell group through a viscoelastic tissue, rather than by fluctuations in DVE migration velocity itself.

    2. Reviewer #1 (Public review):

      Summary:

      The authors study how the migration of distal visceral endoderm (DVE) cells in early mouse embryos becomes channeled towards one direction and the corresponding movement of the epiblast on which the DVE cells migrate. To this end, they develop an analysis pipeline of an in toto live data set previously obtained by the authors, which includes superpixel motion tracking of the visceral endoderm surface and subregions thereof. They find that a morphological asymmetry of the ectoplacental cone is indicative of anterior-posterior axis orientation. Even during the phases prior to and after collective migration, DVE cell speed was larger than in the surrounding tissue. The crossover from the pre-migratory to the migratory phase relies on the alignment of DVE cell motion. During the migration phase, counter-rotating vortices appeared in the emVE as expected when a rigid body moves through an incompressible fluid. Furthermore, DVE migration exhibits what the authors term a ratchet-like behavior, where the cells alternate between bursts of collective migration and periods of essentially no net motion. This behavior could be reproduced in vertex-model simulations, where DVE cells were subjected to a constant external force in an otherwise passive environment of cells. The observed intermittent behavior results from building up stress in the surrounding tissue that is released through cell rearrangements involving T1 transitions. These findings are in line with experimental results, although in embryos, T1 transitions are not as abundant as in the simulations and are largely confined to the region ahead of the DVE. Finally, the authors report a distally directed planar motion in the anterior epiblast underlying the visceral endoderm and thus opposite to the motion of the DVE. Cell migration in the posterior epiblast was slower and more random than in the anterior.

      Strengths:

      The authors provide a detailed analysis of the cell migration patterns in the embryo and show through vertex-model simulations that some of the observed features are really consequences of the properties of incompressible fluids.

      Weaknesses:

      Naming the intermittent dynamics of DVE cells as ratchet-like seems inappropriate, as it is rather reminiscent of stick-slip dynamics.. Quantitatively, the simulations do not provide much more insight beyond providing the flow profile of the (complex) fluid behavior of the tissue surrounding the DVE. It would be interesting to identify mechanisms that underlie migration alignment of DVE cells and to study in detail the T1 transitions - why are they confined to certain regions of the tissue? Furthermore, the theoretical analysis should be extended so that it also considers the dynamics of epiblast cells.

    3. Reviewer #2 (Public review):

      Summary:

      The work provides mechanistic insights into the establishment of the anterior-posterior axis of the mouse embryo by quantifying multiple cellular and tissular parameters from high-quality live imaging data. It shows that the direction of the axis is predetermined by the embryo geometry, that the cells whose migration defines the direction of the axis (the anterior visceral endoderm) have a ratchet-like movement probably depending on transient relaxation events of the epithelial cells lying in their way, and that the adjacent cell layer (the epiblast) moves in the opposite direction.

      Strengths:

      The dataset is large, with multiple embryos from relevant reporter lines integrally imaged at high resolution for long periods of time, and the analysis tools are novel, original and powerful.

      Weaknesses:

      Since all data are obtained from wild-type unchallenged embryos, the direct causality between events may not be fully guaranteed.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors investigate the dynamics of distal visceral endoderm (DVE) migration during early anterior-posterior axis formation in the mouse embryo. Using long-term light-sheet imaging combined with geodesic projections and quantitative motion analysis, they characterize DVE migration at both the cellular and tissue levels. The study identifies three distinct phases of DVE migration, describes the intermittent "stop-and-go" nature of DVE movement, and quantifies coordinated tissue behaviors within the visceral endoderm. The authors further report a previously unrecognized posterior movement of the underlying epiblast that occurs concomitantly with anterior DVE migration. Finally, they develop a two-dimensional vertex model to investigate the mechanical basis of the observed intermittent migration, proposing that cycles of stress accumulation and T1-mediated stress relaxation within the surrounding visceral endoderm account for the observed dynamics

      Strengths:

      Overall, this is a very interesting study combining state-of-the-art live imaging with an impressive quantitative image analysis framework. The imaging quality is excellent, and the authors provide one of the most detailed quantitative descriptions of visceral endoderm (VE) dynamics to date. In particular, the combination of whole-embryo light-sheet imaging, geodesic projections and quantitative analysis provides a rich dataset that will undoubtedly be valuable for the community. The model is also informative and provides a mechanistic hypothesis for the start and stop motion of the VE.

      Weaknesses:

      (1) Clarification of the Superpixel-based image analysis

      The image analysis pipeline is impressive but could be explained more clearly for readers unfamiliar with the authors' previous work. In particular, the manuscript relies extensively on superpixel tracking, but it remains unclear what advantages this approach offers over more conventional Lagrangian particle image velocimetry (PIV). Since this paper should be self-contained, it would be helpful if the authors briefly explained the rationale for choosing superpixel tracking rather than referring readers to their previous eLife publication.

      Related to this point, the manuscript appears to use two different levels of coarse-graining. Motion is initially estimated from thousands of superpixels (1000-5000 according to the Methods), whereas the quantitative analyses are ultimately averaged over only 32 spatial sectors. The relationship between these two levels of representation is not entirely clear and would benefit from clarification. Why use such a dense superpixel seeding, which seems oversampled, if the intent is to eventually bin the result?

      Relatedly, how was the number of superpixels chosen? What is their effective size relative to the size of a VE or epiblast cell? This information is important because the analysis appears to be oversampled. This is particularly evident in Movie S14/Figure 7, where numerous superpixels appear to span a single epiblast cell. At this spatial scale, the measured motion is likely to include intracellular or subcellular movements, such as interkinetic nuclear migration or transient cell-shape changes, rather than pure tissue displacement. This may be somewhat misleading, as the visual impression is that the tissue itself is moving, whereas in some instances this reflects cellular/subcellular fluctuations. A discussion of the spatial scale of the superpixel analysis, together with a demonstration that the conclusions are robust to the degree of coarse-graining, would greatly strengthen the manuscript, especially regarding he movement of the epiblast (see point 4).

      (2) Use of the term "ratchet-like"

      We would recommend avoiding the term ratchet-like and instead using start-stop or stop-and-go migration throughout the manuscript. While these terms describe the same observed behavior, ratchet-like implicitly suggests an irreversible mechanism underlying the motion, whereas the present study primarily documents an intermittent migration pattern. In my opinion, stop-and-go is a more descriptive and mechanistically neutral terminology, leaving the mechanistic interpretation to the modelling section.

      (3) Mechanistic interpretation of the stop-and-go behavior

      The vertex model constitutes the principal mechanistic component of the study and provides an interesting explanation for intermittent DVE migration through stress accumulation followed by T1-mediated stress relaxation. However, the comparison between the model and the experimental data reveals an important discrepancy. As acknowledged by the authors, the model predicts a broader distribution of T1 transitions than observed experimentally, whereas in vivo T1 events appear largely confined to the embryonic visceral endoderm ahead of the migrating DVE.

      This discrepancy suggests that an important aspect of junctional mechanics may be missing from the current formulation. Have the authors considered whether an asymmetric constitutive description, in which junctions remodel more readily under compression than under tension, could better account for the observed spatial restriction of T1 events? Such constitutive asymmetry may provide a more biologically realistic mechanism for intermittent migration while preserving the overall framework proposed here.

      Overall, we find the modelling direction promising, but at present the model appears somewhat premature or overly simplified relative to the experimental observations. The simulations convincingly demonstrate that T1-mediated stress relaxation can generate intermittent migration, but they do not yet quantitatively, if not qualitatively, reproduce the spatial distribution of T1 events observed in vivo. Since the authors have segmented some samples, could all the cells then provide a movie with T1 annotated? That would be helpful to get an intuition on the level of performance of the model compared to experimental data.

      Related to this point, the stop-and-go behavior shown in Figure S7 is not immediately obvious. It would be helpful to display the instantaneous DVE velocity together with the timing of T1 transitions, allowing the proposed correlation to be appreciated more directly. In addition, in Figure 5E, the lower panel appears to be labelled "DVE position", whereas the text suggests that DVE velocity is intended. This should be clarified.

      (4) Motion of the epiblast

      The observation of coordinated epiblast motion is intriguing. However, it would be helpful if the authors quantified the magnitude of the net displacement. From the movies, the overall displacement appears relatively modest, perhaps on the order of one cell diameter. Is this indeed the case?

      More generally, we have some concerns regarding the quantification and representation of epiblast motion. As discussed above, the superpixel analysis appears to operate at a subcellular scale, with many superpixels spanning the apico-basal extent of individual epiblast cells. Consequently, the measured motion may partly reflect transient cell deformations, for example during mitosis or interkinetic nuclear migration, rather than displacement of the tissue itself. Finally, we wonder whether the flattened representation is the most appropriate way to present the epiblast data. Such projections are clearly helpful for analyzing the whole VE motion over a curved epithelial surface. However, the epiblast motion described here is essentially linear, and it is therefore less obvious how the flattening affects the apparent displacement. It would be helpful if the authors could also present the epiblast movement in the original, non-flattened imaging data (e.g. using an optical transverse section through the embryo). At present, the motion is only shown either as a geodesic projection or as a flattened transverse view, such that the reader never directly observes the movement in its native three-dimensional geometry.

    1. eLife Assessment

      This valuable study provides fMRI evidence for a hierarchical representation of sequence, arranged along the anterior-posterior axis of the entorhinal cortex, as well as evidence from human single-unit recordings for sequence position coding in the hippocampus and entorhinal cortex. Weaknesses include tenuous links between the single-unit and fMRI datasets, incomplete evidence for some claims, lack of discussion of relevant hippocampal literature, and potential over-interpretation of the findings in relation to rodent grid coding. This work will be of interest to researchers studying memory, sequence representations, and spatial coding in the medial temporal lobe.

    2. Reviewer #1 (Public review):

      Summary:

      Shpektor et al. propose a link between how humans learn abstract and hierarchical structures to support memory (for example, remembering the event of the first landing on the moon) and the medial temporal lobe (MTL) and grid cells in particular. Given that there is solid work on how grid cells in different modules jointly encode position in rodents, providing evidence for the existence of a similar code in humans in the non-spatial domain and in relation to memory formation, would constitute a valuable finding.

      The authors first examine a small human intracranial dataset to demonstrate that sequence position is decodable in MTL population codes. They then examine behavioral data from two larger groups of participants who passively viewed content presented in a hierarchical sequence and show that errors in recall of positions within that sequence qualitatively match hierarchical predictions. The task design enabled distinct signatures of memory representations at different levels of hierarchy. While there were no multivariate patterns in MTL or any brain region that matched these patterns reliably, a follow-up analysis in MTL revealed a gradient along the anterior-posterior axis, such that lower levels of the hierarchy tended to have representational peaks in more anterior regions of the MTL, which was consistent across the two fMRI datasets.

      Major strengths of the study include the novelty of the experimental paradigm and data.

      In particular, single cell recording in MTL from a small number of human participants during sequence learning and testing a larger group of human participants on a sequence amenable to hierarchical structure learning, and collecting fMRI data during retrieval.

      Furthermore, the paper tackles an important question and does so from both directions, using inspirations from both biology and computational science to navigate it.

      The primary weaknesses of the paper are a lack of compelling support for the overarching claim about hierarchical representation and a lack of clarity and consistency about exactly what those hierarchical representations should and do look like. My concerns regarding these weaknesses are described below, and I believe that most, if not all, of them could be addressed through additional analysis and paper revisions.

      In the first part of the paper, the authors provide single-cell recordings in MTL, and they report the existence of cells that are sensitive to position (more so than to picture). However, they don't elaborate on this result with a model for an abstract sequence code. This is an issue because one possible explanation for the sequential position decoding is that neurons just fire at the presentation of the first image and decay at different rates, or ramp up toward action or feedback. One might be able to decode the position in sequence from these cells' activity, but can hardly call this an abstract code of position in a sequence. However, the authors don't provide further investigation into what the single-cell result might suggest and move on to a completely different fMRI experiment in the second part of the paper. Being able to decode sequence position does not, in my view, necessarily imply an abstract positional code - and I felt that further analysis of the single unit data would be required to identify what representations gave rise to that decoding ability.

      The most compelling evidence that participants were encoding temporal order hierarchically came from behavioral data in the second part of the paper. However, these results were not presented clearly enough to evaluate their reliability and specificity. Figure 2i shows histograms of errors across participants with arrows pointing to bars that apparently correspond to errors of different levels of hierarchy. There are three colored bars, corresponding to errors of one unit at the first, second, or third levels of hierarchy. The first level is not diagnostic of hierarchy, but the other two colored bars appear higher than the colors nearby them. However, my understanding is that these bars correspond to situations with the same tone - which seems like an obvious reason that two positions might be confused, which in my view would weaken the argument for hierarchical encoding. Furthermore, there is no display of variability in the plot or indication of individual differences, so it is hard to tell whether the histogram is dominated by a few participants who made a lot of errors or is reflective of a general tendency across participants.

      The fMRI analyses, while creative, raise questions regarding interpretability. The authors report no representations of hierarchical position at any level, either in MTL or across the whole brain, which would typically be taken as a lack of evidence for the representations existing. Follow-up analyses revealed that what shadows of representations do exist seem to line up along the anterior-posterior gradient. But what does that mean if we can't be sure that the representations are really there? Typically, we tally up evidence supporting an overarching claim by testing multiple predictions that are all consistent with the same story - but in this case, it seems that not all such test results are consistent.

      In many cases, it was difficult to judge the strength of evidence due to somewhat minimal reporting on the exact hypotheses tested and test statistics.

      On a high level, I found the overarching story linking the two datasets together to be somewhat tenuous. While I understand that science rarely rolls out as a coherent story, presenting the authors' valuable experiments in this fashion makes it harder for the reader to digest the information and reach a conclusion. The relevance of the first section of the paper to the second is not immediately apparent. Each section provides somewhat incomplete evidence for a set of claims on its own - but my view was that combining the two studies led to more questions than answers - since the paradigms and measurements are so different.

      In conclusion, the authors propose an interesting account of how memories are formed in the human brain, by building an abstract and hierarchical code. The paper identifies a few separate findings that are suggestive of hierarchical abstract memory encoding in the MTL - yet I believe that more work would need to be done to irrefutably support that claim.

    3. Reviewer #2 (Public review):

      Overall, I think these are exciting results that make a very nice contribution to the literature. I thought the picture-tagging of sequence locations in the fMRI study was clever, and the across-sequence RSA results were especially compelling. But there are several aspects of the presentation of the results that reduced my confidence and enthusiasm.

      (1) This is an unusual paper in that there is one human intracranial study and two fMRI studies. The paradigm for the intracranial study is very different than the fMRI paradigm. The key differences are that the fMRI paradigm is hierarchical, while the intracranial is flat, with no sequence learning component, and the fMRI is auditory, while the intracranial is auditory. The justification for the switch from intracranial to fMRI was that intracranial does not allow anterior-posterior axis analysis, but there are so many differences between the studies that this feels like an awkward transition and justification. Also, anterior-posterior analysis in the MTL may not be feasible in EC with intracranial data, but it can be feasible in the hippocampus, and indeed this could be very worthwhile and relevant to pursue (see point 2).

      While the two independent fMRI datasets is a strength, the replications would have been much more compelling had the analysis for the second dataset been preregistered.

      (2) The intracranial results are pitched as a novel "abstract coordinate representation" but there is a substantial prior literature on MTL "ordinal position codes", which I believe is the same thing in this paradigm. Most of this literature is in the hippocampus, which is, of course, very relevant given the hippocampal findings here, but there is also evidence for this kind of information in EC, e.g., https://elifesciences.org/articles/45333.

      (3) Given the intracranial results in the hippocampus as well as the prior relevant literature on position coding, it was not clear why the hippocampus was not an ROI in the fMRI studies.

      (4) It wasn't until reading the Methods section carefully that I understood that the results do not hold for the right EC, only the left. This deserves more acknowledgment.

      (5) The use of one-sided t-tests with an alpha of .05 reduced my confidence in the robustness of the results.

    4. Reviewer #3 (Public review):

      Summary:

      Shpektor et al. investigate how hierarchical sequence structure is represented in the entorhinal cortex (EC) and medial temporal lobe (MTL) using a combination of single-unit recordings and fMRI. In the single-unit recordings, they find abstract representations of ordinal position within short sequences in both the EC and the hippocampus. Next, they use two fMRI datasets to examine representations of hierarchical sequence structure in EC. They find that these representations (1) are organized along a posterior-to-anterior hierarchy, with finer sequence structure represented in posterior EC and coarser structure in anterior EC, and (2) generalize across sensory features, suggesting an abstract representation of sequence position. The authors take these findings as evidence of a non-spatial hierarchical coordinate system in the human EC, analogous to grid cells in rodents.

      Strengths:

      The methodological approach presented in this study is commendable, combining single-unit recordings in the MTL with two fMRI datasets. The finding of hierarchical and abstract sequence representations in the EC is compelling and is replicated across these datasets and modalities. The manuscript addresses important questions about how the MTL abstracts across experiences that share hierarchical structure, a topic of considerable current interest. As such, the work is likely to be of broad interest to researchers studying these processes in both rodents and humans.

      Weaknesses:

      In my view, the main weaknesses concern the interpretation of the results, as well as several areas where additional analyses and methodological clarification would strengthen the manuscript. My point-by-point comments are as follows:

      (1) I found the evidence for hierarchical and abstract sequence-position representations interesting. However, I am less convinced by the stronger claim that these findings demonstrate a coordinate system analogous to grid-cell coding. The current results appear to provide stronger support for abstract sequence-position coding than for grid-like coding per se. In particular, it is not clear to me that hierarchical sequence representations necessarily imply a grid-like representational format or a coordinate system. Many neural systems exhibit gradients of representational scale along the anterior-posterior axis, both within and across brain regions, without being considered grid-like. I would encourage the authors to clarify why it should be interpreted specifically in terms of a coordinate system rather than more general hierarchical sequence representations. The manuscript would benefit either from a more explicit justification of this link to grid-cell coding or from a more cautious framing of the conclusions.

      (2) Relatedly, the emphasis on grid-cell-like coding naturally centers the story on entorhinal cortex (EC). Yet, the single-neuron results indicate that the hippocampus contained a comparable number of position-selective cells. In addition, a large body of literature has implicated the hippocampus in hierarchical representations of memories, sequences, and relational structure. For completeness, I encourage the authors to repeat the key fMRI analyses within the hippocampus, rather than focusing exclusively on EC.

      (3) I have some concerns regarding the amount of information available to distinguish representations at different levels of the sequence hierarchy. As I understand the design, each 113-tone sequence was associated with only eight images, meaning there were approximately 14 tones between successive image events. It would be helpful to provide additional detail regarding how image coordinates were assigned and selected, how many observations contributed to each hierarchical level, and how much statistical power was available to distinguish representations at different scales.

      (4) I was also uncertain about the potential influence of visual similarity in Dataset 1. My understanding is that the images were not entirely unique but instead consisted of rotated versions of the same images. If so, this visual similarity could potentially complicate the interpretation of representational structure. It would therefore be useful to clarify whether repeated images occurred within the same or different locations in the hierarchy and to provide analyses demonstrating that the reported effects cannot be explained by visual similarity. This seems particularly important given that the corresponding effects in Dataset 2 were weaker.

      (5) The rationale for using a custom orderness metric could be explained more clearly. It would be helpful to understand why a custom metric was preferred over rank-order measures such as Kendall's tau or Spearman's rho. I would be interested in seeing whether the orderness results replicate using one of these more conventional metrics.

      (6) I had difficulty reconciling the finding that sequence representation effects are stronger across rather than within sequences. Intuitively, I would have expected representations within a sequence to reflect both shared hierarchical position and sensory experience, thus yielding stronger within-sequence effects than across sequences. The opposite pattern seems somewhat counterintuitive. I would appreciate additional discussion of this pattern and what it implies about the nature of the underlying representation. It would also be informative to know whether similar effects are observed elsewhere in the brain, and why EC might preferentially express a purely abstract representation more strongly than representations that additionally share sensory features.

      (7) The authors' theory is that hierarchical representations of sequences in EC are used as a scaffold for memory, yet the current paper does not link their behavioural results to their neural ones. I think making such a link would greatly strengthen the results presented here. For example, is displacement error or sequence memory related to ordered representations of the sequence structure?

      (8) I thought the manuscript would benefit from a broader discussion of prior work on (1) sequence representations and (2) hierarchical representations in the hippocampus and related regions. As it stands, the manuscript does a good job of situating its findings within the literature on grid cells in the EC but gives comparatively little attention to the literature on sequence representations in the hippocampus. Placing the current findings within this broader body of work would help clarify which aspects of the results are specific to a grid-like interpretation and which may instead reflect more general principles of hierarchical representation in the MTL or across the brain.