10,000 Matching Annotations
  1. Last 7 days
    1. Reviewer #1 (Public review):

      Summary:

      Naina Gour and colleagues provide a detailed observational study in which they demonstrate that MRGPRX4, a human G-protein coupled receptor (GPCR), is expressed exclusively in human melanomas and, when expressed in mouse melanocytes, drives the development of melanomas in mice. These findings provide evidence that MRGPRX4 has the properties of an oncogene, at least in certain cellular environments.

      Strengths:

      A strength of this work is the nice historical note in which overexpression of MAS1, a GPCR, led to classic studies of transformed fibroblasts in culture and tumors in nude mice. Cloning of MAS1 led to the identification of the MRGPR family of receptors, now known to be key players in neuroimmune and neurosensory phenomena. Here, the story comes full circle with a member of the MRGPR family being linked to a tumor, specifically melanoma. Perhaps the story is not entirely surprising given that the neural crest serves as a precursor for both nerves and melanocytes. But it is nice to see.

      Additional strengths include the vast array of tools and techniques employed, from public databases to engineered mice, to establish firmly that MRGPRX4 is expressed in melanomas, although not in every malignant cell.

      Weaknesses:

      Given the power of the strengths of the data and story, the following comment is only sort of a weakness, as the topic is addressed while being saved for future studies. Specifically, what leads to the expression of MRGPRX4? The authors posit that it is an epigenetic phenomenon, look briefly at methylation, and rather than going down the proverbial rabbit hole of what comes first, have reasonably decided to punt.

      Another concern is that given what comes across as the initial observation of MRGPRX4 being expressed in melanoma, what do all of the additional studies add?

      For the non-cognoscenti, and to make the manuscript more accessible, the abbreviation NC/EMT, which is also inverted to EMT/NC, should be spelled out periodically as neural crest/epithelial-mesenchymal transition.

      Please explain how this study came about. Was it a result of someone deciding to look at expression in the GTEx project and compare it to a tumor database?

      A comment could be made to explain that while NSG and normal mice were used, the former are immunocompromised, and drawing conclusions without specifying these differences is a weakness.

      In Figure 1A, the p-value of -145 begs for a little explanation. I don't recall seeing such a p-value.

      Have you considered treating the murine melanomas with murine via PD-L1? I appreciate that this comment is somewhat superfluous given the inhibition of MRGPRX4 with compound 31-2, but given the human therapeutics combined with the fact that you have done 'everything else', I wonder what might happen.

      Given the basal ligand-independent signaling, might engineering variants of MRGPRX4 that do not signal be of value?

    2. Reviewer #2 (Public review):

      Summary:

      This study presents a fundamental new finding - the identification of a sensory-neuron itch receptor, MRGPRX4, as an unexpected melanoma oncogene through a mechanism of lineage-inappropriate expression rather than mutation. The evidence supporting the core observation (tumor-specific upregulation, restriction to invasive transcriptional states, and sufficiency to drive fully penetrant metastatic melanoma in vivo) is compelling, drawing on convergent human genomic datasets and a well-controlled genetic mouse model. However, several of the mechanistic and translational claims - particularly regarding causal drivers of invasion, the immunosuppressive tumor microenvironment, and in vivo pharmacological efficacy - remain incomplete, relying on correlative evidence.

      Strengths

      The authors propose that MRGPRX4, normally restricted to a subset of peripheral sensory neurons, is aberrantly re-expressed in melanoma rather than through mutational mechanisms, and that this re-expression is sufficient to drive tumorigenesis through basal, ligand-independent GPCR signaling. This is a genuinely novel model for oncogenesis, and the manuscript deploys an impressive range of approaches - bulk and single-cell transcriptomics, spatial transcriptomics, proteomics, phosphoproteomics, and pharmacology - to support it.

      Strengths:

      The claim that MRGPRX4 is selectively upregulated in melanoma and confined to neural-crest-like/invasive transcriptional states is well supported, with consistent results across multiple independent human scRNA-seq datasets. The claim that ectopic MRGPRX4 is sufficient to drive melanoma is convincingly demonstrated by the fully penetrant, metastatic phenotype in the Tyr-CreER;MRGPRX4-LSL model, with appropriate specificity controls showing that MRGPRX1, MRGPRX2, and MRGPRX3 do not phenocopy this effect.

      The claim that MRGPRX4 signals through basal, ligand-independent activity is reasonably well supported by bile-acid quantification showing endogenous ligand concentrations well below the EC50 required for activation.

      Weaknesses:

      The claim that MRGPRX4 remodels the tumor microenvironment toward an immunosuppressive state rests on flow cytometric frequency data (altered neutrophil/eosinophil ratios, increased PD-L1+ myeloid populations) but lacks any functional immune assay to demonstrate that this remodeling actually impairs anti-tumor immune responses.

      The claim that the two MRGPRX4-enriched tumor subpopulations (ECM-rich and NC-like/invasive) underlie the observed invasive and metastatic phenotype is not directly tested; the authors appropriately acknowledge this as an open question, but it is worth noting explicitly that this leaves the mechanistic link between the identified cell states and the functional phenotype (proliferation, invasion, metastasis shown in Figure 4) unresolved.

      Finally, the comparison with BRAF- and NRAS-driven GEMMs (Figure 3K-L) establishes overlap in transcriptional cell states but does not report whether these canonical models themselves upregulate endogenous Mrgprx4. This omission leaves unclear whether MRGPRX4 acts as a convergent node downstream of canonical oncogenic signaling, or represents an independent, parallel route to a similar phenotypic endpoint - a distinction that matters considerably for how broadly the finding should be interpreted.

      Overall assessment:

      The manuscript's central, most novel claim - that lineage-inappropriate expression of a sensory GPCR is sufficient to drive melanoma - is compellingly supported. The secondary mechanistic and translational claims built around this finding are convincing and consistent with the broader literature but are currently supported by correlative rather than causal or functional evidence.

    3. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Naina Gour and colleagues provide a detailed observational study in which they demonstrate that MRGPRX4, a human G-protein coupled receptor (GPCR), is expressed exclusively in human melanomas and, when expressed in mouse melanocytes, drives the development of melanomas in mice. These findings provide evidence that MRGPRX4 has the properties of an oncogene, at least in certain cellular environments.

      Strengths:

      A strength of this work is the nice historical note in which overexpression of MAS1, a GPCR, led to classic studies of transformed fibroblasts in culture and tumors in nude mice. Cloning of MAS1 led to the identification of the MRGPR family of receptors, now known to be key players in neuroimmune and neurosensory phenomena. Here, the story comes full circle with a member of the MRGPR family being linked to a tumor, specifically melanoma. Perhaps the story is not entirely surprising given that the neural crest serves as a precursor for both nerves and melanocytes. But it is nice to see.

      Additional strengths include the vast array of tools and techniques employed, from public databases to engineered mice, to establish firmly that MRGPRX4 is expressed in melanomas, although not in every malignant cell.

      We thank the reviewer for the positive assessment of our work and for noting the arc back to the original MAS1 studies, which we agree makes for a satisfying full-circle story. We would add one nuance: while the shared neural crest origin of melanocytes and nociceptors offers a plausible explanation for why an MRGPR family member could be co-opted in melanoma, we were nonetheless surprised that this property is specific to MRGPRX4. When we tested the other MRGPRX family members, some of which are also known to be expressed in sensory neurons, using the same genetic strategy (MRGPRX1, MRGPRX2, MRGPRX3; Supplementary Fig. 2B), none induced melanoma, despite arising from the same family of receptors. This suggests the oncogenic property is not simply a generic consequence of neural crest ancestry, but reflects a biology specific to MRGPRX4 itself, which we find is one of the more intriguing open questions this work raises.

      Weaknesses:

      Given the power of the strengths of the data and story, the following comment is only sort of a weakness, as the topic is addressed while being saved for future studies. Specifically, what leads to the expression of MRGPRX4? The authors posit that it is an epigenetic phenomenon, look briefly at methylation, and rather than going down the proverbial rabbit hole of what comes first, have reasonably decided to punt.

      We agree with the reviewer that understanding the precise mechanisms through which MRGPRX4 gets activated during oncogenesis is important, and we had noted this as a limitation of current findings. As we mentioned in the discussion, we hypothesize that one or more environmental modulators - UV exposure, inflammatory cues, or an aging skin microenvironment could initiate the epigenetic reprogramming underlying this expression. This will require dedicated studies and is suited for future work.

      Another concern is that given what comes across as the initial observation of MRGPRX4 being expressed in melanoma, what do all of the additional studies add?

      The expression data in human melanoma are correlative and serve only as the entry point. The subsequent studies establish MRGPRX4 as a causal, druggable driver of melanoma: it is sufficient to induce 100% penetrant metastatic melanoma in vivo, required for melanoma proliferation and invasion, signals through ligand-independent basal PI3K-AKT/MAPK activity, and is pharmacologically targetable.

      For the non-cognoscenti, and to make the manuscript more accessible, the abbreviation NC/EMT, which is also inverted to EMT/NC, should be spelled out periodically as neural crest/epithelial-mesenchymal transition.

      This is corrected in the revised version.

      Please explain how this study came about. Was it a result of someone deciding to look at expression in the GTEx project and compare it to a tumor database?

      This was a serendipitous discovery. As MRGPRs can modulate immune cell function, we were driving expression of various human MRGPRX genes in immune lineages in mice. We noticed that mice overexpressing MRGPRX4 developed spontaneous melanoma-like growths on the ear and tail. Upon investigation, we found that while the Cre driver we had used is appropriate for immune cells, it is also expressed at low levels in melanocytes. We hypothesized that our initial "immune-knock-in" strategy likely resulted in low-level knock-in in mouse melanocytes, thereby allowing MRGPRX4 to directly drive melanocytic transformation. This led us to explore MRGPRX4 expression in human malignancies, where we found MRGPRX4 was highly expressed in melanoma. Next, to directly test our hypothesis that MRGPRX4 transforms melanocytes, we overexpressed MRGPRX4 using Tyrosinase-CreER, the standard Cre driver for melanocytes. This led to a fully penetrant spontaneous melanoma phenotype (Fig. 2).

      A comment could be made to explain that while NSG and normal mice were used, the former are immunocompromised, and drawing conclusions without specifying these differences is a weakness.

      This is a fair point, and we agree the distinction deserves clarification. The two mouse systems in this study serve different purposes and are not interchangeable. The transgenic TyrCreER+; MRGPRX4LSL+/- model, in which melanoma arises spontaneously from endogenous mouse melanocytes, was studied in immunocompetent mice, allowing us to fully characterize the tumour, including the tumour immune microenvironment (Fig. 2L–O), in an intact immune setting. In contrast, the human A2058 melanoma cell studies (proliferation and metastatic seeding; Fig. 4B–J) required an immunodeficient host, since human cells would otherwise be rejected by a competent murine immune system. For these experiments, we used NSG mice, the standard immunodeficient strain used in human xenograft studies. We have updated the text to reflect this.

      In Figure 1A, the p-value of -145 begs for a little explanation. I don't recall seeing such a p-value.

      The small p-value reflects two features of the comparison: (1) both GTEx normal skin and TCGA-SKCM contain large sample sizes (hundreds of samples per group), which gives the Mann-Whitney U test (also known as the Wilcoxon rank-sum test) statistical power, and (2) MRGPRX4 expression shows minimal overlap between the two groups; it's essentially undetectable in normal skin but broadly expressed across melanoma samples. Together, these yield such a p-value. This is expected for rank-based tests applied to large, cleanly separated datasets.

      Have you considered treating the murine melanomas with murine via PD-L1? I appreciate that this comment is somewhat superfluous given the inhibition of MRGPRX4 with compound 31-2, but given the human therapeutics combined with the fact that you have done 'everything else', I wonder what might happen.

      We agree with the reviewer that testing checkpoint blockade in the MRGPRX4-driven model is a relevant direction. Our finding that MRGPRX4-driven tumours develop a PD-L1hi, neutrophil-rich microenvironment (Fig. 2M–O) makes anti-PD-L1 treatment a natural next experiment, and we would predict it to be informative both on its own and in combination with an MRGPRX4 inhibitor, given that the two target distinct compartments- the tumor-immune microenvironment versus tumour-intrinsic proliferative/invasive signaling. We consider it an important direction for future work.

      Given the basal ligand-independent signaling, might engineering variants of MRGPRX4 that do not signal be of value?

      This is a great point. A variant that disrupts basal (ligand-independent) signaling specifically would be informative. Identifying and validating such a basal-activity-disrupting variant, including confirming normal receptor expression and trafficking, is an important direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This study presents a fundamental new finding - the identification of a sensory-neuron itch receptor, MRGPRX4, as an unexpected melanoma oncogene through a mechanism of lineage-inappropriate expression rather than mutation. The evidence supporting the core observation (tumor-specific upregulation, restriction to invasive transcriptional states, and sufficiency to drive fully penetrant metastatic melanoma in vivo) is compelling, drawing on convergent human genomic datasets and a well-controlled genetic mouse model. However, several of the mechanistic and translational claims - particularly regarding causal drivers of invasion, the immunosuppressive tumor microenvironment, and in vivo pharmacological efficacy - remain incomplete, relying on correlative evidence.

      Strengths

      The authors propose that MRGPRX4, normally restricted to a subset of peripheral sensory neurons, is aberrantly re-expressed in melanoma rather than through mutational mechanisms, and that this re-expression is sufficient to drive tumorigenesis through basal, ligand-independent GPCR signaling. This is a genuinely novel model for oncogenesis, and the manuscript deploys an impressive range of approaches - bulk and single-cell transcriptomics, spatial transcriptomics, proteomics, phosphoproteomics, and pharmacology - to support it.

      Strengths:

      The claim that MRGPRX4 is selectively upregulated in melanoma and confined to neural-crest-like/invasive transcriptional states is well supported, with consistent results across multiple independent human scRNA-seq datasets. The claim that ectopic MRGPRX4 is sufficient to drive melanoma is convincingly demonstrated by the fully penetrant, metastatic phenotype in the Tyr-CreER;MRGPRX4-LSL model, with appropriate specificity controls showing that MRGPRX1, MRGPRX2, and MRGPRX3 do not phenocopy this effect.

      The claim that MRGPRX4 signals through basal, ligand-independent activity is reasonably well supported by bile-acid quantification showing endogenous ligand concentrations well below the EC50 required for activation.

      We are thankful to the reviewer for the positive feedback.

      Weaknesses:

      The claim that MRGPRX4 remodels the tumor microenvironment toward an immunosuppressive state rests on flow cytometric frequency data (altered neutrophil/eosinophil ratios, increased PD-L1+ myeloid populations) but lacks any functional immune assay to demonstrate that this remodeling actually impairs anti-tumor immune responses.

      We acknowledge the reviewer’s suggestion that functional immune assays will demonstrate that MRGPRX4-driven tumor remodeling impairs checkpoint-driven anti-tumour immune response. We consider this an important direction for future work.

      The claim that the two MRGPRX4-enriched tumor subpopulations (ECM-rich and NC-like/invasive) underlie the observed invasive and metastatic phenotype is not directly tested; the authors appropriately acknowledge this as an open question, but it is worth noting explicitly that this leaves the mechanistic link between the identified cell states and the functional phenotype (proliferation, invasion, metastasis shown in Figure 4) unresolved.

      We agree with the reviewer. To test the direct role of these subpopulations in driving metastasis in our model, we need to ablate them specifically. This is possible with an intersectional genetics strategy, wherein we could use a state-specific Dre or Flp driver (e.g., a Prrx1-DreER or Prrx1-FlpO knock-in) crossed to a dual-recombinase-dependent effector allele (e.g., a Frt- or Rox-gated diphtheria toxin receptor), and then crossed to our TyrCreER-MRGPRX4-LSL model. This quadruple-transgenic strategy would allow ablation restricted specifically to MRGPRX4-driven tumour cells occupying the target mesenchymal state, for example. These genetic lines can be created but require generating or sourcing new dual-recombinase-dependent alleles and a substantially longer breeding and validation timeline (approximately 18-24 months). Together, these represent possible, if long-term, strategies for directly testing the causal contribution of these MRGPRX4-enriched subpopulations to melanoma invasion and metastasis.

      Finally, the comparison with BRAF- and NRAS-driven GEMMs (Figure 3K-L) establishes overlap in transcriptional cell states but does not report whether these canonical models themselves upregulate endogenous Mrgprx4. This omission leaves unclear whether MRGPRX4 acts as a convergent node downstream of canonical oncogenic signaling, or represents an independent, parallel route to a similar phenotypic endpoint - a distinction that matters considerably for how broadly the finding should be interpreted.

      MRGPRX4 is a primate-specific receptor with no mouse ortholog in melanocytes; thus, we cannot assess endogenous expression of MRGPRX4 in BRAF/NRAS-driven GEMMs. Nevertheless, the following observations argue against MRGPRX4 functioning solely downstream of canonical oncogenes. First, TyrCreER<sup>+</sup>; MRGPRX4LSL mice develop fully penetrant melanoma without engineered BRAF or NRAS activation or tumour-suppressor loss. Second, MRGPRX4 loss in BRAF V600E mutant A2058 cells reduces pERK1/2, pAKT, and pS6K, showing that MRGPRX4 sustains these signaling outputs even in the presence of activated BRAF. Together, these findings support a model in which MRGPRX4 could provide a distinct oncogenic input.

      Overall assessment:

      The manuscript's central, most novel claim - that lineage-inappropriate expression of a sensory GPCR is sufficient to drive melanoma - is compellingly supported. The secondary mechanistic and translational claims built around this finding are convincing and consistent with the broader literature but are currently supported by correlative rather than causal or functional evidence.

    1. eLife Assessment

      This valuable study reveals a role for lactylation in influenza polymerase function and virus replication. The evidence supporting the claims is solid, although the direct involvement of ATAT1 and SIRT1 in this modification and the functional relevance of the identified sites can be further substantiated. These findings will appeal broadly to influenza virologists studying polymerase regulation, as well as researchers investigating the functional roles of lactylation.

    2. Reviewer #1 (Public review):

      Summary:

      This work characterizes the regulation of lysine lactylation on influenza A virus PA protein, and describes how this post-translational modification at residues K605/K609 facilitates asymmetric polymerase dimerization at the ANP32 interface. The authors identify ATAT1 as the host enzyme mediating PA lactylation and SIRT1 as the enzyme responsible for removing this modification. They present evidence that PA lactylation enhances viral polymerase activity and viral replication, while suppression of lactylation impairs viral growth, polymerase function, and viral pathogenicity in vivo.

      Strengths:

      Overall, this manuscript explores a virus-host interaction axis illustrating how host metabolic signaling modulates influenza polymerase function. These findings are likely to attract broad interest, including influenza virologists studying polymerase regulation, as well as researchers investigating the functional roles of lactylation. This mechanistic insight may also offer clues for developing host-targeted antiviral strategies.

      Weaknesses:

      The manuscript lacks direct experimental evidence connecting lactylation to the proposed functional mechanism. While lactylation is detected in virions and overexpression systems, it remains unclear whether lactylation dynamically modulates polymerase function during infection. It is also unknown what proportion of PA undergoes lactylation at distinct infection stages, and whether lactylation specifically takes place within replication-competent asymmetric polymerase dimers. Importantly, the authors have not shown that mutation of K605/K609 abrogates the functional effects induced by lactate supplementation or ATAT1/SIRT1 overexpression in viral replication assays, which would help establish a direct connection between lactylation and viral replication. Therefore, although a correlation exists between these residues and viral replication, a direct mechanistic link between lactylation and the proposed replication model has not been firmly established. At minimum, the authors are encouraged to acknowledge these key limitations and moderate (tone down) their conclusions. For example, the observations are consistent with, but do not definitively prove, a functional role for PA lactylation in viral genome replication.

      Major points:

      (1) All experiments were performed using PR8, a laboratory-adapted H1N1 strain. Although this strain is commonly used for mechanistic investigations, evidence demonstrating conservation of this mechanism in currently circulating viral strains or other subtypes of influenza viruses would substantially support the conclusion that lactylation promotes viral pathogenicity. In the absence of such data, it remains unclear whether the observed findings apply broadly or are limited to the PR8 strain. In addition, the authors are encouraged to verify these phenotypes in additional cell lines.

      (2) While the authors cite published work indicating that ATAT1 possesses lactyltransferase activity, it would be valuable to clarify whether ATAT1 directly catalyzes PA lactylation or functions indirectly as an intermediate. Similar considerations apply to SIRT1 regarding its potential role in removing lactylation from PA. Direct biochemical evidence, such as in vitro modification assays, would help strengthen the proposed mechanism.

      (3) The available data cannot rule out the possibility that phenotypic changes induced by K605/K609 mutations stem from structural or charge alterations independent of lactylation. In fact, the results in Figure 4A and 4B support this alternative explanation: the K609R mutant shows reduced lactylation without obvious alterations in polymerase activity. This observation raises the question of whether the functional effects of these residues are driven by modified lactylation status or merely charge alterations.

      (4) The proviral effect of ATAT1 appears largely independent of its enzymatic activity (Figure 2H), making it challenging to clarify whether ATAT1 functions by modifying PA to regulate polymerase activity and viral replication. Experiments examining SIRT1 on viral replication encounter similar interpretative limitations.

      (5) Several siRNA knockdown results warrant careful interpretation. In Figure 3F and Figure 2E, the knockdown efficiency of SIRT1 and ATAT1 appears limited, especially at 12 and 24 h.p.i. Additional independent experiments with improved silencing efficiency or complementary approaches (such as CRISPR knockout) would help strengthen these observations.

    3. Reviewer #2 (Public review):

      This work reveals a role for PA lactylation in influenza virus replication, and the proposed involvement of ATAT1 and SIRT1 is certainly intriguing. These observations open up new avenues for understanding how host metabolism may influence viral infection. Nevertheless, a few mechanistic issues remain to be clarified. Notably, the direct evidence for ATAT1 and SIRT1 acting as the writer and eraser of this modification is still incomplete, and the functional relevance of the identified sites could be further substantiated.

      Major Comments:

      (1) The direct evidence supporting ATAT1 as a PA lactyltransferase and SIRT1 as a PA delactylase is still lacking. It remains possible that these two molecules affect PA lactylation indirectly. Therefore, in vitro lactylation/delactylation assays to clarify whether ATAT1 and SIRT1 act directly on PA should be performed.

      (2) Although Figure 4A shows that K605A/K609A mutations reduce PA lactylation, the use of lactylation-mimetic mutants in functional complementation assays could further strengthen the conclusion.

      (3) ATAT1 and SIRT1 are known to regulate multiple substrates. Therefore, whether the viral phenotypes resulting from ATAT1/SIRT1 manipulation truly operate via PA K605/K609 remains to be demonstrated. Complementation experiments would help address this issue.

      (4) The mechanistic analysis currently focuses on polymerase dimerization. IP-MS assays comparing the host protein interaction profiles of PA WT versus K605/K609 mutants could reveal whether additional host factors are involved.

      (5) The downstream consequences of PA lactylation have not been explored in the context of host antiviral immunity, particularly type I interferon (IFN-I) signaling. We would suggest examining the expression of IFN-β, ISG56, and other ISGs upon infection with WT versus PA K605/K609 mutant viruses.

    4. Author response:

      We thank the editors and reviewers for their thoughtful and constructive assessment of our manuscript. We appreciate the reviewers’ insightful comments and suggestions, which will help strengthen the mechanistic rigor of our work. Below, we outline the key revisions we plan to undertake in the revised version.

      Response to Reviewer 1

      (1) Dynamic regulation of PA lactylation during infection.

      We plan to examine PA lactylation levels at multiple time points post-infection to assess whether PA lactylation changes dynamically during the viral replication cycle.

      (2) Epistasis experiments linking K605/K609 to lactate- or enzyme-dependent phenotypes.

      We acknowledge that multiple viral proteins undergo lactylation and that ATAT1/SIRT1 may mediate lactylation of multiple viral proteins. We will perform viral replication assays in the context of K605/K609 mutant viruses under lactate supplementation or ATAT1/SIRT1 manipulation conditions. These experiments will allow us to assess whether the effects of lactate or ATAT1/SIRT1 manipulation on viral replication are dependent, at least in part, on PA K605/K609.

      (3) Conservation across viral strains and cell lines.

      We will further examine key phenotypes in additional influenza A virus subtypes (e.g., H1N1 swine influenza and H9N2 avian influenza strains) and in additional cell lines to assess the extent to which the observed mechanism is conserved beyond the PR8 laboratory-adapted strain.

      (4) Direct biochemical evidence for ATAT1 and SIRT1 activity on PA.

      We plan to perform in vitro lactylation and de-lactylation assays using purified recombinant PA, ATAT1, and SIRT1 proteins to investigate whether ATAT1 and SIRT1 can directly modulate PA lactylation, respectively.

      (5) Disentangling lactylation from charge/structural effects.

      We acknowledge the reviewer’s point that K609R shows reduced lactylation without obvious changes in polymerase activity. We will include K-to-Q substitution mutants (e.g., K605Q/K609Q) in functional assays to further assess the functional consequences of these substitutions. Although K-to-Q substitutions do not strictly mimic lysine lactylation, these mutants may help distinguish effects related to lysine charge/chemical properties from those specifically attributable to lactylation. We will also temper our conclusions and acknowledge that charge and/or structural effects may contribute independently to the observed phenotypes.

      (6) Enzymatic activity dependence of ATAT1 and SIRT1.

      We will further investigate the enzymatic activity-dependent versus -independent contributions of ATAT1 and SIRT1 using catalytically inactive mutants, together with the epistasis experiments described above. We will also revise the text to clarify the interpretive limitations of these experiments.

      (7) Improved loss-of-function approaches.

      We will complement the existing siRNA experiments with CRISPR/Cas9 knockout cell lines for ATAT1 and SIRT1 and, where feasible, repeat key assays in ATAT1- and SIRT1-knockout cells.

      In the revised manuscript, we will also add a detailed methodological explanation in the figure legend and Methods section to clarify how the luciferase complementation system distinguishes asymmetric from symmetric polymerase dimers, show individual data points overlaid on bar graphs with error bars, and correct spelling errors throughout the manuscript.

      Response to Reviewer 2

      (1) Direct in vitro lactylation/de-lactylation assays.

      As noted above, we will perform in vitro modification assays with purified proteins to investigate whether ATAT1 and SIRT1 directly modulate PA lactylation and de-lactylation, respectively.

      (2) Functional assays with lysine-to-glutamine substitution mutants.

      Although we recognize that K-to-Q substitutions do not strictly mimic lysine lactylation, we will generate K-to-Q substitution mutants (e.g., K605Q/K609Q) and evaluate their effects on polymerase activity and viral replication to further assess the functional relevance of these sites.

      (3) Complementation experiments linking ATAT1/SIRT1 phenotypes to PA K605/K609.

      As noted above, we plan to perform viral replication assays in the context of K605/K609 mutant viruses under ATAT1/SIRT1 manipulation conditions to assess whether PA K605/K609 contributes to the effects associated with ATAT1/SIRT1 manipulation.

      (4) IP-MS comparison of PA WT versus mutant host interaction profiles.

      Our study focuses on the mechanism by which PA lactylation modulates viral polymerase activity and replication. A comprehensive host interactome analysis via IP-MS represents a broader systematic investigation beyond the scope of this focused work. We will discuss this as an important future research direction in the revised manuscript.

      (5) Exploration of host antiviral immunity downstream of PA lactylation.

      This work focuses on the direct effects of PA lactylation on viral polymerase activity and replication. As the PA mutants exhibit altered replication capacity, differences in IFN/ISG expression would be largely secondary and difficult to disentangle from the direct effects of altered viral replication. We therefore consider this question beyond the scope of the current study and will add it as a future research direction in the Discussion section.

      In the revised manuscript, we will also examine PA lactylation levels under increasing lactate concentrations to assess their relationship with the dose-dependent changes in viral titers. We will revise the text to clarify the interpretive limitations and, where feasible, perform endogenous co-immunoprecipitation experiments to further assess the interactions between PA and ATAT1/SIRT1 under physiological expression conditions.

      We believe these revisions will strengthen the mechanistic evidence and help address the core concerns raised by both reviewers.

    1. eLife Assessment

      This important study shows that self-supervised pretraining on related fungal genomes can substantially improve sequence-to-expression prediction in budding yeast. The evidence for the main conclusion is convincing, with careful comparisons to models trained from random initialization and clear gains on held-out expression data. The broader claims about the optimal evolutionary scope and transfer of regulatory grammar are less fully supported.

    2. Reviewer #1 (Public review):

      Summary:

      The authors ask whether self-supervised pretraining on related fungal genomes gives a useful prior for predicting gene expression in S. cerevisiae, where the compact ~12 Mb genome supplies too few independent windows to train a large supervised model from scratch. They pretrain a BERT-style masked DNA language model on a corpus of fungal genomes and fine-tune it to predict RNA-seq gene expression. They found a surprisingly (to me) large improvement in performance: 0.78 prediction Pearson R versus 0.67 for a randomly initialized model.

      The manuscript also presents a new experimental resource: 3,053 RNA-seq data sets with perturbations using the YETI experimental platform to upregulate specific genes.

      Strengths:

      (1) Overall, the manuscript is well-written and is likely to be impactful.

      (2) The protocol handles the train/test split of orthologous sequences well, which is nontrivial.

      (3) The authors did a good job "steelman-ing" Shorkie_Random_Init: it received its own learning-rate sweep, and the authors tested two reduced-capacity from-scratch architectures to rule out some overparameterization issues.

      (4) The public codebase is unusually well-organized.

      Weaknesses:

      I have a number of comments, none of which significantly impact the main findings.

      Two analyses appear missing from the MPRA section: (1) Does self-supervised pre-training improve MPRA models such as DREAM-RNN? and (2) Is Shorkie better than Shorkie_Random_Init at the MPRA task? The language about "correlative, non-causal" associations makes me think the authors tried this and got poor results; it would be informative to include these as a supplementary negative result. (There is also (3): Does MPRA pre-training improve genomic models? But this is clearly out of scope for this paper.)

    3. Reviewer #2 (Public review):

      Summary:

      Chao et al. study whether multi-species self-supervised pretraining learns sequence representations that can be transferred to improve downstream expression prediction in budding yeast. The rationale is that the budding yeast genome is relatively small and compact, which may not provide enough sequence variation for standard reference-genome-based supervised learning.

      This paper has two stages. They first train U-Net-style DNA language models on collections of fungal genomes with increasing phylogenetic breadth using a BERT-style masked language modeling (MLM) objective. The model with the best sequence reconstruction (i.e., lowest perplexity) on held-out budding yeast sequences is called Shorkie LM. They then use Shorkie LM's weights to initialize Shorkie, a supervised model for predicting functional genomic tracks (RNA-seq, ChIP-exo, ChIP-MNase) in budding yeast. Compared with a randomly initialized model with the same architecture, Shorkie clearly performs better on held-out expression prediction and outperforms the randomly initialized model on all three variant effect benchmarks. This is good evidence that the pretraining is beneficial. For Shorkie LM and Shorkie, the authors also use interpretation methods to analyze motifs captured across a range of loci, and the results seem to be consistent with known yeast biology.

      Overall, the experiments are well thought out and controlled, and most claims are supported with sufficient evidence. While the idea that pretraining can be beneficial for budding yeast has been reported previously (for scalar expression prediction) [1], this paper's exploration of optimal evolutionary scope and genomic language model pretraining in general still holds practical value for the field. The Shorkie model itself is also a useful resource and ranks first on many tasks in a recent public benchmark for fungal sequence-to-expression models [2]. I only have some minor concerns/suggestions.

      Strengths:

      (1) The authors put considerable effort into evaluating their modeling choices. For the pretraining corpora, they compare four phylogenetic scopes (a species, strain, order, kingdom). Each corpus is evaluated across multiple model architectures/capacities. Extra care is taken with homology filtering between training and held-out data. For supervised transfer, several candidate language models are carried forward to test whether their ranking by language model selection criterion (perplexity) predicts their ranking on the downstream task. In addition, the same-architecture randomly initialized baseline for Shorkie is itself optimized for learning rate, with additional randomly initialized CNN and U-Net models included as additional controls, giving some strong nulls to compare against.

      (2) The experimental dataset generated in this study is also a valuable resource. Using the ministat array, the authors generated ~3,000 time-resolved RNA-seq profiles following transcriptional regulator inductions. Such data are important for understanding transient regulatory events and how they shape the transcriptome over time.

      (3) Several figures include schematic panels that make the whole workflow easy to understand.

      Weaknesses:

      (1) As acknowledged by the authors, the current evolutionary "sweet spot" at 165 Saccharomycetales genomes is potentially confounded by factors like corpus size, annotation quality, and optimization difficulties. I agree that it would be hard to disentangle these factors, but a relatively simple control appears to be missing. Although unlikely, it is still possible that the "sweet spot" is driven partly by the amount of data rather than the phylogenetic scope. One simple control would be to sample 165 genomes from the broader fungal corpus and match the total number of training windows to the Saccharomycetales dataset.

      (2) L148, "These results demonstrate that Shorkie LM captures conserved regulatory grammar." The evidence presented in this section is mainly about recovery of known fungal/yeast sequence motifs, which are more like words than grammar. The latter is usually understood as relationships between motifs such as their spacing, multiplicity, and arrangement. So I suggest the authors either tone down the claim or provide additional analyses that directly test this. One possibility would be to generate nucleotide dependency maps [3] for a handful of loci with well-characterized regulatory syntax.

      (3) Figure 2E. The interpretation of t-SNE clusters might be confounded by the length of genomic elements. The five classes of elements shown in the plot have very different length distributions, but they are extracted and zero-padded to the same input length, and their embeddings are mean-pooled across the full padded sequence. As a result, the fraction of real sequence entering the mean embedding differs substantially across genomic element classes, which could contribute to observed separation independently of learned regulatory features.

      (4) Lines 216-218, "transfer learning from pan-fungal self-supervised pretraining yields generalizable representations of exon-intron structure and regulatory grammar, substantially improving in expression prediction." Similar to point 2, I think this statement is somewhat stronger than what the current evidence supports. The results clearly show that pretraining improves downstream expression prediction, but it is less clear that this improvement can be specifically attributed to the transfer of regulatory grammar and gene structure, rather than motif representations or more general learned inductive biases acquired during pretraining. This is itself an interesting question. It might be possible to get at it by selectively reinitializing the convolution layers vs the transformer blocks while keeping the rest of the Shorkie LM weights, although this could still be hard to interpret if the relevant information is distributed across the model.

      (5) The ISM for Figures 4 and 5 seems to correspond to the summed predicted coverage across all output bins (Methods, Equation 17). I was unclear about the intended goal here. For example, when mutating the ATG42 promoter in Figure 5, is the goal to quantify the predicted effect on ATG42 expression specifically? If so, should the ISM instead be computed by summing only the bins covering the target gene? Otherwise, predicted coverage from neighboring genes could also contribute to the ISM score and affect the comparison across induction time points.

      References:

      [1] Wang Y, Cai Z, Zeng Q, Gao Y, Ouyang J, Xu Y, et al. Genomic Touchstone: benchmarking genomic language models in the context of the central dogma. bioRxiv. Preprint posted online June 30, 2025. doi:10.1101/2025.06.25.661622

      [2] Schneider T. ybench: a benchmark for fungal sequence-to-expression models. GitHub. Accessed August 18, 2026. https://github.com/Tom-Ellis-Lab/yeast-seq2expression-benchmark

      [3] Tomaz da Silva P, Karollus A, Hingerl J, et al. Nucleotide dependency analysis of genomic language models detects functional elements. Nat Genet. 2025;57(10):2589-2602. doi:10.1038/s41588-025-02347-3

    1. eLife Assessment

      This important study investigates how experimentally introducing two facultative bacterial endosymbionts into the Russian wheat aphid, Diuraphis noxia, affects aphid performance, dispersal, and damage to cereal host plants. The authors provide solid evidence that the two symbionts can generate contrasting phenotypes: Rickettsiella increases plant damage while reducing wing formation and dispersal, whereas Regiella reduces aphid population growth and feeding damage. The successful establishment and stable transmission of these novel symbiont-host associations, combined with experiments spanning individual, whole-plant, population and mesocosm scales, are notable strengths of the work; however, the mechanisms underlying these effects remain unresolved, evidence for horizontal transmission is indirect, and some population-level conclusions rely on relatively small sample sizes or effects that are not consistently detected across time points, and therefore the potential application of these findings to pest management remains promising but speculative. The study will be of broad interest to researchers working on insect symbiosis, plant-insect interactions, and biologically based approaches to pest management.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors examine what happens when two facultative endosymbionts, Rickettsiella viridis and Regiella insecticola, are introduced into a novel aphid host, the Russian wheat aphid (Diuraphis noxia). They ask whether these introduced symbionts affect aphid performance, plant damage, alate production, dispersal, plant defense responses, and symbiont dynamics. The main result is that the two symbionts have contrasting effects: Rickettsiella tends to increase plant damage and reduce dispersal-related traits, whereas Regiella tends to reduce plant damage and aphid population growth, with less evidence for an effect on dispersal.

      Strengths:

      The manuscript presents successful establishment of stable transinfected populations of an agriculturally important aphid species, which is a substantial technical achievement in itself. I also appreciated that the authors examined the system across several experimental contexts, including different host plants, mixed cages at two temperatures, whole-plant assays, and a mesocosm dispersal experiment, rather than relying on a single laboratory setup. Taken together, these experiments provide a useful and reasonably convincing demonstration that novel symbiont associations can generate contrasting phenotypes in this system.

      Weaknesses:

      There are some major aspects of this paper that I thought could be strengthened. My main concern is that the manuscript feels broader than it is conceptually focused. A wide range of outcomes is measured, which gives the study breadth, but it also makes the central question harder to identify. As written, the paper reads more strongly as a proof-of-principle demonstration of ecologically relevant phenotypes than as a tightly framed test of a specific biological idea.

      A second issue is that the biological basis of the reported phenotypes remains less developed than the phenotypic description itself. The authors make a genuine effort to address mechanism through JA, JA-Ile, SA, and metabolomic profiling, but these analyses only partially explain the main results. The negative result for the canonical defense markers is informative, yet it still leaves a substantial gap between the observed variation in plant damage and the processes responsible for it.

      I also think some caution is needed in how the two symbionts are compared. The authors explain why some follow-up experiments were designed differently for Rickettsiella and Regiella, and that rationale is understandable. Still, because the downstream assays were not fully matched, the paper is strongest when each symbiont is interpreted on its own terms rather than as a strict comparison.

      Overall, I would suggest softening the Significance Statement so that it more clearly reflects what is directly shown here, namely that introduced symbionts can alter plant damage and dispersal-related phenotypes under controlled conditions, rather than implying that the study directly tests management utility in agricultural settings.

    3. Reviewer #2 (Public review):

      Summary:

      The authors generated two novel aphid-symbiont associations and examined the impact of these new symbiotic associations on plant-insect-symbiont interactions. The authors notably provide detailed phenotypic assessments of the insect hosts and host plants. They show that one introduced symbiont, Rickettsiella, increases aphid-induced damage to host plants, while the other, Regiella, ameliorates aphid damage. The authors suggest that such novel insect-symbiont pairings may be used as tools to mitigate crop damage in the future.

      Strengths:

      Although a few experiments seem to have limited sample sizes and limited statistical power, these are often complemented with highly replicated smaller-scale experiments. The combination of larger mesocosm and population-level experiments along with assessments of individual insects generally provides a comprehensive depiction of the effects of these symbionts on their hosts. The opposing impacts of Regiella and Rickettsiella infection on the aphid host plant are of broad interest. It is also surprising that the host plants did not exhibit strong differences in canonical defensive signalling, despite these differences.

      Weaknesses:

      One thing that I struggled a little with was the rapid spread of Regiella in the shared plant experiments. Possibly this could be attributed to an increased reproductive output (due to faster developmental time, and/or an increase in fecundity) or efficient horizontal transmission. However, the other experiments performed indicate a slight negative impact (Figure 4a) or no influence (Figure 4C, 4D, Figure S6) of Regiella infection on host fitness. Given these other results, it seems that Regiella must spread fairly efficiently between hosts, which comes as a surprise, and there are very few examples of horizontal transmission of Regiella like this in the literature. The manuscript would benefit from a clear and direct demonstration of horizontal transmission, rather than it being inferred indirectly. The similar spread observed in the Rickettsiella mixed cages is less surprising, because there are several examples where this has been demonstrated.

      It is also a little surprising that mesocosm-dispersal experiments were not also conducted using Regiella-infected lines. At several points throughout the manuscript, the idea of using symbiont transfections to reduce plant harm is raised. I can understand that these experiments are likely time-, space-, and resource-intensive, but that seems like these would have been relevant experiments, especially in the context of controlling damage to plants.

    4. Reviewer #3 (Public review):

      Summary:

      The authors were investigating the impact of introducing novel facultative bacterial endosymbionts into the pest aphid, Diuraphis noxia, to explore the possibility of using facultative symbionts as a crop protection tool. They successfully established the vertical transmission of both endosymbionts and performed a series of aphid performance and dispersal experiments together with measurement of aphid feeding on host plant health, growth, and metabolism. While most of the experiments revealed no effect of the endosymbionts, some significant treatment effects were found, showing that Rickettsiella reduced aphid dispersal, and Regiella reduced aphid population growth and feeding damage.

      Strengths:

      The team worked with two novel facultative symbionts (Rickettsiella viridis and Regiella insecticola) that they were able to successfully establish in D. noxia. The data were collected and analyzed using solid, well-described methodology.

      Weaknesses:

      While interpretation of the data is reasonable, the few experiments which revealed significant treatment effects rest on relatively small sample sizes.

      Measuring symbiont density is difficult. The authors use quantitative PCR to measure the "density" of endosymbionts relative to a host gene. This is a standard approach in the field; however, recent work has shown that endosymbionts like the aphid primary endosymbiont, Buchnera, are variably polyploid [1]; further the aphid cells that house the symbionts are also highly polyploid and variable in their ploidy [2]. It is important to understand that what is being measured when using qPCR is DNA copy number and not quantification of the number of symbiont cells. Alternative approaches to measuring symbiont density include flow cytometry [3], and SymbiQuant [4], a machine vision tool that can quantitatively characterize symbiont populations from DAPI-stained confocal images. These alternate approaches also have their limitations. Currently, there is no perfect approach to measuring symbiont density, which remains an important measure in experiments such as these. Put simply, it is important for a reader to be aware of the limitations of each approach and interpret data accordingly.

      [1] Komaki, K., and H. Ishikawa. 2000. Genomic copy number of intracellular bacterial symbionts of aphids varies in response to developmental stage and morph of their host. Insect Biochemistry and Molecular Biology 30:253-258.<br /> [2] Nozaki, T., and S. Shigenobu. 2022. Ploidy dynamics in aphid host cells harboring bacterial symbionts. Scientific Reports 12:9111.<br /> [3] Simonet, P., G. Duport, K. Gaget, M. Weiss-Gayet, S. Colella, G. Febvay, H. Charles, J. Viñuelas, A. Heddi, and F. Calevro. 2016. Direct flow cytometry measurements reveal a fine-tuning of symbiotic cell dynamics according to the host developmental needs in aphid symbiosis. Scientific Reports 6:19967.<br /> [4] James, E. B., X. Pan, O. Schwartz, and A. C. C. Wilson. 2022. SymbiQuant: A machine learning object detection tool for polyploid independent estimates of endosymbiont population size. Frontiers in Microbiology 13:816608.

      Impact and Significance:

      Food security and production, and pest control are major challenges facing the human population. This work contributes knowledge that will benefit the development of alternate pest control strategies in agriculture.

    1. eLife Assessment

      Using a range of complementary approaches, this study examines how type I and type II interferon (IFN) programs exert opposing effects on macrophage responses relevant to tuberculosis. The work addresses a key biological question and provides valuable mechanistic insights. However, the data to support several key conclusions is incomplete and requires further strengthening; in particular, the role of ferritin needs to be established more definitively, the TNF stimulation findings should be validated in the context of M. tuberculosis infection, and in vivo evidence is needed to support the proposed therapeutic strategy. Additionally, aspects of iron metabolism, lipid peroxidation, and the translatability of the findings between mouse and human macrophages would benefit from greater clarity and deeper mechanistic investigation.

    2. Reviewer #1 (Public review):

      Summary:

      This study examines how type I IFN and IFN-γ exert opposing effects on macrophage responses relevant to TB. Using bone marrow-derived macrophages from genetically susceptible B6.Sst1S mice, the authors describe a persistent pathological activation state induced by TNF and characterized by sustained type I IFN signalling, oxidative stress and lipid peroxidation. They show that IFN-γ priming limits several features of this state and propose altered iron metabolism as one mechanism underlying this protective effect. They then use a computational cell-state approach to identify pharmacological interventions that may mimic aspects of IFN-γ activity. In particular, CDK4/6 inhibition with trilaciclib and activation of retinoic acid signalling with ATRA appear to act through complementary mechanisms and, when combined at low concentrations, improve control of intracellular M. tuberculosis.

      Strengths:

      A major strength of the study is the combination of several complementary approaches, including genetic susceptibility, cytokine signalling, oxidative stress, iron and lipid metabolism, transcriptomics, computational modelling and pharmacological perturbation. Together, these experiments build a coherent model of macrophage dysfunction.

      The evidence that type I IFN signalling contributes to maintenance of the pathological state is particularly convincing within the TNF stimulation model. Blocking the type I IFN receptor after the phenotype has developed restores responsiveness to IFN-γ and prevents further accumulation of lipid-peroxidation products. The authors also provide evidence that persistence does not simply reflect continued TNF signalling, since blockade of the TNF receptor after 24 h does not abolish the elevated lipid-peroxidation phenotype. Another strength is that the computational analysis generates experimentally testable predictions, and two mechanistically distinct interventions identified by this approach are subsequently validated in macrophages.

      Weaknesses:

      There are, however, several limitations that affect the strength and scope of the conclusions.

      (1) First, the use of the terms "persistent" and especially "self-sustaining" would be better supported by a more complete time-course analysis.

      (2) Second, the proposed central role of ferritin-mediated iron sequestration in the protective effect of IFN-γ is not yet demonstrated directly. The data clearly link IFN-γ treatment to ferritin induction and reduced labile iron, but the causal contribution of ferritin itself remains to be established.

      (3) Third, an important limitation is the connection between the mechanistic model developed with TNF stimulation and actual M. tuberculosis infection. Most of the mechanistic analysis, including type I IFN super-induction, lipid peroxidation, ferritin induction, labile iron and HIF1α regulation, is performed in TNF-stimulated macrophages. The infection experiments show that IFN-γ improves bacterial control and that low-dose trilaciclib plus ATRA reduces intracellular bacterial burden, but they do not establish that M. tuberculosis infection induces the same pathological circuit, or that these interventions improve bacterial control by acting through that circuit. The study therefore defines a convincing TNF-driven macrophage phenotype with relevance to bacterial control, but the broader conclusion that this mechanism underlies IFN-dependent susceptibility to TB remains only partially supported.

      (4) Finally, the therapeutic implications go beyond the experimental evidence currently presented, since all of the pharmacological experiments are performed in cultured macrophages and there is no in vivo validation.

      Conclusion:

      Overall, this study proposes an interesting framework for understanding how inflammatory activation may become maladaptive in susceptible macrophages and how IFN-γ may combine antimicrobial activation with protection from oxidative damage. The convergence between IFN-γ, iron metabolism, lipid peroxidation and the pharmacological perturbations identified computationally is a clear strength. However, the causal role of ferritin, the operation of the proposed circuit during M. tuberculosis infection, and the in vivo relevance of the pharmacological strategy remain to be established. These limitations leave the mechanistic and translational evidence incomplete, while the study itself remains potentially important.

    3. Reviewer #2 (Public review):

      Summary:

      The authors have carried out extensive transcriptomic, phenotypic and modelling-based analyses to provide novel insights into the interaction of the type I and II interferon programs in the determination of macrophage activation status and resistance to Mtb infection and infection-mediated damage. They demonstrate how an antagonistic effect between the two programs goes beyond classical downstream immune signalling pathways to lipid peroxidation maintained in a sustained autocrine manner and generation of a persistent pathological activation state (pPAS). Based on these analyses, the authors propose a therapeutic strategy of boosting specific pathways that increase oxidative stress resilience to reduce inflammatory pathology without suppressing host defenses for bacterial control and resisting pPAS. The conceptual framework may prove applicable to interferonopathies and to other bacterial and viral infections, though this remains to be tested.

      Strengths:

      (1) The study uses macrophages from a disease-relevant genetic murine model in which the sst1 locus drives the formation of necrotic pulmonary granulomas resembling human TB lesions- pathology not seen in standard C57BL/6 mice. This provides a genetically defined comparison between susceptible and resistant macrophages on an otherwise identical background, allowing the authors to attribute differences in activation state to a single locus rather than to strain-level variation.

      (2) The experimental design isolates the phenomenon of interest: the TNF withdrawal and restimulation scheme allows the authors to establish that the aberrant activation state persists after removal of the initiating stimulus, rather than simply reflecting ongoing stimulation. The timed IFNAR blockade at 2, 12 and 24 h similarly separates initiation of the state from its maintenance.

      (3) Lipid peroxidation is assessed through two orthogonal readouts: 4-HNE immunostaining for accumulated adducts and linoleamide alkyne click chemistry for ongoing synthesis. These, coupled with ROS and labile iron pool measurements, isotype antibody controls, parallel B6 and B6.Sst1S comparisons, and an anti-TNFR control, help in establishing that the phenotype is independent of continued TNF signalling. The convergence of these independent measures gives confidence in the peroxidation phenotype itself.

      (4) The cSTAR analysis is applied here using regression rather than classification, generating a continuous DPD_TB score that correlates with measured Mtb fold change and thus provides a quantitative transcriptomic metric of macrophage priming state. Critically, the pathway predictions arising from this analysis (CDK4/6 inhibition and RAR activation) were tested and confirmed experimentally. The inferred network topology further predicted synergy between these two interventions, which bore out experimentally as an approximately ten-fold reduction in the effective dose of each agent in controlling Mtb during infection.

      Weaknesses:

      (1) Figure 2C is difficult to interpret as presented. The row labels ("No TNF"/"TNF") use different terminology from the corresponding conditions in panel A ("TNF withdrawal"/"TNF restimulated"), and "TNF" appears on both axes referring to different phases of the experiment; no timepoint is given on the panel itself, unlike neighbouring panels. Harmonising the labels with panel A and stating the harvest timepoint would help the reader. More substantively, the remaining lipid peroxidation readouts in this figure (panels D-G) are all at early timepoints of TNF stimulation (2-24 h) rather than during withdrawal, which limits what they can say about sustained autocrine signalling. Extending these assays to the later timepoints used in Figure 1 would considerably strengthen the claim that IFN-I maintains, rather than only initiates, the pathological state. The same applies to Figure 2G, where the contribution of itaconate to 4-HNE accumulation over time is not yet resolved.

      (2) Several of the pathways implicated here are reported to behave differently between murine and human macrophages, and between macrophage subsets (alveolar versus monocyte-derived macrophages), during Mtb infection. This does not diminish the findings in this model, but it does bear on how broadly they can be generalised.

      a) Type I interferon responses differ by species and by macrophage subset across multiple reports. Since the proposed model depends on autocrine IFN-I signalling reaching a threshold sufficient to sustain the pathological state, these differences in IFN-I output are worth keeping in mind when interpreting the wider significance of the findings.

      b) Similarly, itaconate production in murine BMDMs is 20-fold higher than in LPS-activated human monocyte-derived macrophages and 50-fold higher than in LPS-activated alveolar macrophage-like cells, and Mtb infection of these human macrophages very weakly induces ACOD1 with almost no detectable itaconate (PMID 41797714). This is relevant to the Acod1/4-OI arm of the mechanism.

      c) In a cross-species comparison of Mtb-infected macrophages, cholesterol homeostasis genes (including HMGCS1, IDI1, LSS) were significantly upregulated in human alveolar macrophages but downregulated in subcutaneous BCG-exposed murine alveolar macrophages (PMID 41208107)- the opposite direction to the lipid biosynthesis suppression treated here as a defining pPAS feature. The same group reports that murine AMs lack c-Maf and IL-10 whereas murine BMDMs express both (PMID 40073087), indicating that the autocrine anti-inflammatory brake on IFN-I responses is itself subset-dependent.

      d) Finally, the cSTAR network predictions were inferred from human THP-1 perturbation data, but tested only in murine BMDMs. Establishing how this circuit operates in human macrophages, and during Mtb infection rather than TNF stimulation alone, would be a valuable extension of the work.

      (3) The causal relationships linking IFN-I, lipid peroxidation, ROS and loss of IFNγ responsiveness could be drawn together more clearly. These elements are each established, but the connections between them are not always demonstrated directly. IFN-I appears to promote peroxidation through Acod1/itaconate and suppression of lipid biosynthesis rather than through iron, since neither IFNβ nor IFNAR blockade alters the labile iron pool. This would suggest two separable inputs to 4-HNE rather than a single pathway. This raises a further question about the persistent state itself: the labile iron pool rise appears to be TNF-driven, yet all labile iron measurements are made at 24 h in the continued presence of TNF and none under the withdrawal condition, so it is unclear whether elevated catalytic iron is sustained once the initiating stimulus is removed. Similarly, while IFNAR blockade reduces peroxidation, the reciprocal arm is not tested. An antioxidant or iron chelator could be used to ask whether peroxidation in turn drives Ifnb1 super-induction. In the absence of this information, the proposed feedback loop remains partly inferred. It would considerably strengthen the manuscript if the authors could clarify, through additional experiments or in the text, how the labile iron pool and lipid peroxidation relate to one another and what sustains each of them after TNF withdrawal.

      (4) Reading across the manuscript, the labile iron pool emerges as the variable most consistently associated with the phenotype. Every protective intervention tested converges on it. By contrast, the alternative candidate mechanisms do not track with outcome. Lipid biosynthesis genes are suppressed by IFNγ yet induced by both trilaciclib and ATRA, all three of which are protective. GPX4 is unchanged under IFNγ and trilaciclib. The Acod1/itaconate axis cannot account for it since IFNγ priming blocks 4-HNE accumulation induced by exogenous itaconate. Yet, a direct causal role for iron is never tested. Additionally, IFNβ induces 4-HNE with the labile iron pool entirely unchanged, indicating at least one route to lipid peroxidation that bypasses catalytic iron. Focusing on iron handling would make the manuscript's message more coherent and its therapeutic argument more compelling.

    4. Reviewer #3 (Public review):

      Summary:

      Araveti et al. demonstrate that Type I Interferon (IFN-I) and Interferon-gamma (IFN-γ) play opposing roles in regulating lipid peroxidation and host resistance during Mycobacterium tuberculosis (Mtb) infection. While both the two pathways drive inflammation, they differ fundamentally in cell protection. IFN-I signaling triggers a destructive, self-amplifying cycle catalysing the generation of reactive oxygen species (ROS) and lipid peroxidation, which ultimately impairs the host's ability to clear Mtb. Conversely, the authors demonstrate that IFN-γ couples antimicrobial activation with cytoprotection. It primes macrophages to fight the mycobacteria while simultaneously shielding them from oxidative stress. It achieves this by sequestering iron, which successfully prevents ROS from converting into damaging lipid peroxidation products.

      Strengths:

      Ultimately, this study highlights a critical biological distinction: IFN-γ safely balances inflammatory activation with cellular defense, whereas IFN-I promotes uncontrolled pathological damage. This divergent coupling of inflammation and cytoprotection carries major consequences for disease progression and host survival.

      Weaknesses:

      This study demonstrates all the findings in specific mouse strains. How these translate in human macrophages is not well characterised, thus raising the issue of its overall impact in tuberculosis disease.

    1. eLife Assessment

      This valuable study provides a proof of concept for the utility of transcriptomic methods to develop novel diagnostic tests for bovine tuberculosis. A comprehensive analysis using a suite of machine learning models provides convincing evidence for the use of mRNA biomarkers to distinguish between infected and disease-free animals. The analysis also provides major insights into both the levels of individual variation in expression between animals as well as systematic changes with respect to the time from infection.

    2. Reviewer #1 (Public review):

      Summary:

      The control of bovine tuberculosis in managed populations such as Ireland and Great Britain is unusual in that demonstrably sick animals are rarely, if ever, seen in herds. Control is therefore focused on the identification and removal of animals that test positive to the tuberculin skin test (the legal definition of infection). Despite over a century of study, the relationship between tuberculin test status, infection and most importantly infectiousness is still poorly quantified. Different formats of the tuberculin skin test are acknowledged to have both poor sensitivity and compromised specificity, although the characteristics of these tests are likely to vary considerably between contexts due to both biological variation and discretion in measurements by testers. There is an urgent need for new, more reliable and cheaper diagnostics to address the failures of existing control programs and to enable control in emerging markets that do not currently control the disease.

      Strengths:

      A key strength of this study is the use of samples from both naturally infected and experimentally infected animals. This data set is used to perform a careful and exhaustive evaluation of the extent to which patterns of transcriptomic expression can be used to classify between disease free animals and those infected with bovine tuberculosis.

      The experimentally infected animal samples provide evidence that expression patterns of infected animals vary with respect to the time from infection. The authors highlight that this suggests transcriptomic markers may be able to detect infection earlier than tuberculin and IGRA tests that target cell-mediated immune responses. However, these methods could potentially provide a valuable new tool for quantifying the role of individual variation and progression for a disease where the individual life-history is still frustratingly mysterious.

      Weaknesses:

      However, the high levels of individual variation - and in particular differences in patterns of expression between naturally and experimentally infected animals do raise questions about how diagnostic tests developed from these tools would be used in practice. In particular, while many of the models considered achieved high sensitivity - estimated specificity is consistently lower than current diagnostic tests and considerably lower than that necessary for screening tests given the frequency of testing carried out as part of statutory control programs.

      Expanding the number of samples may help to address these issues, but I would have liked to see some discussion of the extent to which the level of biological variation observed in this study may limit the precision of diagnostic tests developed using these tools. Given the likely characteristics of tests based on these methods, I would be interested to hear how the authors think they could fit within current statutory programs, either as supplementary or replacement tests?

    3. Reviewer #2 (Public review):

      Summary:

      This study evaluates whether peripheral blood transcriptomic profiles can be used to classify cattle infected with Mycobacterium bovis using a range of machine-learning approaches. By integrating RNA-seq datasets from naturally and experimentally infected animals, the authors develop and test predictive models capable of distinguishing infected from uninfected cattle and assess their ability to differentiate bovine tuberculosis from other infectious diseases. The study addresses an important challenge in bovine tuberculosis control and presents evidence that host transcriptional signatures may have utility as an adjunct diagnostic approach.

      Strengths:

      - The study combines data from multiple independent cohorts, including both naturally and experimentally infected cattle, which increases the biological relevance of the findings.

      - The analytical workflow is comprehensive, scientifically sound and clearly described. Multiple machine-learning approaches are evaluated and compared rather than relying on a single modelling strategy.

      - The inclusion of a held-out test set, especially because such data is limited, provides a useful assessment of model performance beyond cross-validation alone.<br /> - Thorough evaluation against datasets from cattle infected with MAP, BoHV-1 and BRSV is a valuable addition and provides useful information regarding the specificity of the identified transcriptional signatures.

      - The authors acknowledge important limitations, including batch effects and the need for additional validation.

      - All underlying data and code are made publicly available

      Weaknesses:

      - My main concern relates to generalisability. Although a separate testing dataset was used, the training and testing datasets were generated through random partitioning of samples from the same underlying studies. As a result, classifier performance in a completely independent external cohort remains unclear. Discussion of this limitation, and whether alternative validation strategies such as leave-one-study-out analyses were considered, would strengthen the manuscript.

      - The authors identify substantial study-specific batch effects following dataset integration and appropriately account for these in the modelling framework. However, given the magnitude of the reported batch structure, additional discussion regarding the potential influence of residual between-study variation on classifier performance would be helpful.

      - The manuscript is framed in the context of global bovine tuberculosis control, yet the practical implementation of a transcriptomic diagnostic approach is not discussed in great detail. Since bovine tuberculosis remains a significant challenge in many low- and middle-income settings, further consideration of the feasibility, cost, infrastructure requirements, and potential translation of these signatures into more deployable diagnostic platforms would improve the broader relevance of the study.

      - The datasets used for classifier development are derived primarily from Ireland, the UK and the United States. It would be useful to discuss whether differences in circulating M. bovis lineages, cattle populations, management systems, or co-infection pressures could influence host transcriptional responses and therefore the performance of the proposed classifiers in other epidemiological settings. This ties to the previous comment, since epidemiological settings in LMIC countries with a high burden of M.bovis disease would be vastly different from where the data was sourced.

    4. Reviewer #3 (Public review):

      Summary:

      This is an excellent piece of work which sheds greater light on the responses of cattle to both experimental and natural infection in cattle with Mycobacterium bovis infection, using data from different experimental and field groups.

      Strengths:

      The work is based on robust analysis of a range of highly relevant experimental and field sample sets, using transcriptomic approaches. It provides insight into pathogenesis and disease responses, as well as some evidence regarding potential future diagnostic advances.

      Weaknesses:

      I have some simple, but important comments on how the work is discussed. (Consequently, most comments focus on the discussion section). In particular, I suggest that the authors have confused or conflated the great progress that they have made in improving the understanding of the responses to M. bovis infection with an improved ability to practically improve the diagnosis of the infection in the field. Minor differences in specificity and sensitivity - and predictive values - of tests or assays being used can have profound effects in different prevalence settings on the farm, and I don't feel that this understanding is adequately reflected in the discussion in particular. In this respect, the authors really should, in my view, focus not on the outstanding results that come in or from their model fitting approaches, to focus on their model testing results in relation to extrapolated meaning / external validity. The text uses words like 'robust' and 'highly accurate' which are meaningless in the context of test interpretation in the field.

      I get their enthusiasm, based on really interesting findings in relation to disease progression and immune and inflammatory responses, but these indistinct claims rather devalue the quality of the rest of their work in my view.

      One great challenge in work of this nature, using natural cases from farms, is that there is no gold standard for identifying the cases that current diagnostic approaches miss and which they hope their new approaches can help with. This is not discussed or mentioned in their enthusiasm for what they have achieved. Their field datasets are based on the current, insensitively detected cases. It misses the 'occult' cases that are present but undiagnosed.

    1. eLife Assessment

      In this fundamental manuscript, Richter et al. present a thorough anatomical characterization of the Drosophila melanogaster larval pharyngeal sensory system, which is involved in taste-guided behaviors.This study fills a major gap in the larval sensory map, providing a detailed neuroanatomical foundation for future investigations into sensory circuits and behavior. The exceptional data will significantly enrich the field of Drosophila neurobiology.

    2. Reviewer #2 (Public review):

      Summary:

      The authors wanted to achieve a detailed ultrastructural reconstruction of the gustatory sensory organs in the Drosophila pharynx. Using serial EM and the associated bioinformatics tools they have achieved their goal.

      Strengths:

      Given the dataset, finding presented are solid and will be an important work of reference for the future.

      Comments on revised version.

      The authors have well responded to my previous comments and added text and figure material.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors provide a detailed ultrastructural analysis of the larval pharyngeal sensory organs, including the dorsal pharyngeal sensilla, dorsal pharyngeal organ, ventral pharyngeal sensilla, and posterior pharyngeal sensilla. Using electron microscopy and 3D reconstruction, Richter et al., present a comprehensive mapping and classification of pharyngeal sensory structures, defining the morphological type of pharyngeal sensilla based on ultrastructure and generating a neuron-to-sensillum map. These findings significantly advance our understanding of internal larval sensory systems and establish a robust framework for future functional studies in coordination with external sensory systems.

      Strengths:

      The application of high-resolution electron microscopy and 3D imaging analysis successfully overcomes technical challenges associated with visualizing deep internal structures. This enables an unprecedented level of anatomical detail of the larval pharyngeal sensory system. Thus, the study complements and completes existing maps of larval sensory circuits, contributing a comprehensive neuroanatomical characterization of larval sensory input pathways. These insights will inform future studies on larval behavior, sensory processing, and may also have applied relevance for insect control strategies.

      Weaknesses:

      While the manuscript is concise, clearly written, and methodologically rigorous, it primarily addresses a specialized readership with expertise in insect neuroanatomy.

      We thank the reviewer for the positive assessment of our study and for the helpful suggestions. In response, we have clarified the visual presentation of the pharyngeal sense organs in Figure 1, expanded the discussion of adult pharyngeal sensory systems, briefly broadened the comparison to other insect species, checked and corrected the scale bars, and added further methodological detail where appropriate.

      Reviewer #2 (Public review):

      Summary:

      This manuscript documents the structure of the pharyngeal nervous system of the Drosophila larva. The authors wanted to achieve a detailed ultrastructural reconstruction of the gustatory sensory organs in the Drosophila pharynx. Using serial EM and the associated bioinformatics tools, they have achieved their goal. The paper is written clearly and illustrated beautifully with 3D models and annotated sections. The data will significantly enrich the field of Drosophila neurobiology.

      Strengths:

      Given the dataset, the findings presented are solid and will be an important work of reference for the future.

      Weaknesses:

      Previous work, including EM, on the pharyngeal sensory organ is not sufficiently referenced and used for comparison with the data presented in this study.

      We are grateful for the reviewer’s thoughtful comments and for the suggestion to strengthen the historical and comparative context of the work. We have revised the introduction to better acknowledge and discuss the relevant previous EM-based literature on adult and larval internal gustatory sensilla, clarified the organization of the shared pore structure in T1–T3, highlighted the DPO multidendritic neurons more explicitly, and added a comparison that emphasizes the added value of the complete serial EM dataset.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) For improved clarity, highlight the pharyngeal sense organs in Figure 1B. Consider using the color schemes to differentiate between peripheral and internal sensory organs.

      We thank the reviewer for this helpful suggestion. We have revised Figure 1 to more clearly separate the pharyngeal sense organs from the external sense organs in the head region. This revision improves visual clarity and accessibility for readers.

      (2) In reference to lines 80-84, expand the discussion to address how future studies could explore the conserved morphological and functional characterization of the adult pharyngeal sensory system.

      We appreciate this suggestion and have expanded the discussion accordingly. We now briefly address how future work could compare the larval and adult pharyngeal sensory systems to examine conserved morphological and functional features.

      (3) To broaden the manuscript's appeal and emphasize its relevance beyond Drosophila, briefly discuss similarities, differences, or conserved roles of pharyngeal sensory systems in other insect species.

      Thank you for this valuable recommendation. We have added a paragraph placing the Drosophila pharyngeal sensory system in a broader insect context, including similarities, differences and potential conservation across species.

      (4) Recheck the scale bars in all figures, including the supplemental material.

      We thank the reviewer for pointing this out. We carefully rechecked all scale bars across the main and supplemental figures and corrected the missing ones.

      (5) Consider including additional details on image processing or provide appropriate citations for further reading.

      We appreciate this suggestion. We have expanded the methods section to include additional information on technical details and provide the relevant reference for further reading.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 57ff: The previous literature describes internal gustatory sensilla in considerable detail.

      (a) Adult: These sensilla form three complexes, the labral sensory organ, and the ventral and dorsal cibarial sensory organ (Nayak & Singh, 1983, 1985; Singh, 1997; Stocker & Schorderet, 1981; Kendroud et al., 2017). The work by Nayak and Sing includes TEM and presents detailed EM-based schematics. This should be referenced and discussed.

      (b) Larva: Gendre et al. 2004, describes the internal gustatory organs and relates them to their adult counterparts:

      - Dorsal pharyngeal sense organ (DPS) and dorsal pharyngeal organ DPO) are the forerunners of adult labral and ventral cibarial sensory organs

      - Posterior pharyngeal sensory organ (PPS) is the forerunner of the adult dorsal cibarial sensory organ

      - Ventral pharyngeal sensory organ (VPS), derived from the labial segment, undergoes apoptosis during metamorphosis

      This work, connecting larva and adult (and containing detailed diagrams comparing adult and larval pharyngeal sensilla) should be presented in the introduction.

      We thank the reviewer for this important comment. We have revised the introduction to better cite and discuss previous EM-based studies of internal gustatory sensilla in both adult and larval stages, and we now place our findings more explicitly in the context of this prior work.

      (2) Line 180: the relationship between the ending of T1-T3 in one shared pore, and the individually wrapped sensilla should be explained; maybe a simple diagram would help. I did not understand how it works. Normally, in a gustatory sensillum, you have one or more sensory neurons, surrounded by thecogen, trichogen, and tormogen cells. The trichogen generates the shaft with the pore at its tip. Now here, in T1-T3, you have three sets of thecogen/trichogen/tormogen. Do all three trichogen cells somehow participate in the shaft with the common pore? Or only a single one, and the other two generate no shaft? It is possible this cannot be resolved, but the authors should address the problem and suggest a possible scenario.

      We appreciate the reviewer’s concern and agree that this point required clarification. We have revised the relevant text to better explain the organization of T1-T3 and their shared pore and the organization of the support cells.

      (3) Line 205: the DPO multidendritic neurons with dendrites into the hemolymph should be shown; in Figure S4G, I could see only cell bodies. These MD neurons in the gustatory system are, I believe, a true novelty and should be emphasized more if the material allows (text figure!)

      Thank you for highlighting this point. We have revised the results and supplementary material to show these neurons more clearly and to emphasize their novelty and potential relevance to the pharyngeal sensory system.

      (4) A somewhat detailed comparison between the ultrastructure of the DPS as extracted from the serial EM stack of this study, and the conclusions of Nayak and Singh 1983 as depicted in their diagram Figure 7a would be productive. The idea being: what additional details can (only) a complete EM stack provide, compared to conventional EM.

      We appreciate this suggestion. Rather than directly comparing larval and adult structures in detail, we now emphasize what the complete serial EM dataset adds beyond conventional single-section EM, namely a more comprehensive and complete reconstruction of the sensory organs and associated cell types (multidendritic neurons, papilla sensilla, and chordotonal organs that were not described before, organization of support cells)

      (5) To round off the work and connect it to the previously published analysis of gustatory terminal arborizations and connectivity in the brain (Miroschnikow et al.,2018), it would be helpful to add an analysis of the distribution of axons from the different sensilla in the nerves. Miroschnikow analyzes the central terminations of the same sense for which the peripheral structure is described here, only that in their L1 connectome, the periphery was cut off. Do the findings of the current study match their predictions, as to the number of sensory neurons, etc? It should be possible to follow, even at the lower resolution of the dataset presented here, to follow axons of sensory neurons through the nerves to the neuropil entry, and thereby make the connection. I consider this to be of great importance for the field, for authors who want to use the data of this study, and the Miroschnikow et al analysis, for their own studies.

      We thank the reviewer for this thoughtful and constructive suggestion. We fully agree that linking the peripheral sensory anatomy described in this study to the central projections analyzed by Miroschnikow et al. would be highly valuable and of broad interest. However, a systematic analysis of axon distributions from the different sensilla through the nerves to their neuropil entry points is beyond the scope of the present work. Owing especially to the dataset’s resolution and inherent limitations, tracing the connections from sensory organs through the nerves to their projections in the brain is technically highly challenging and extremely time-consuming, since much of the process would need to be performed manually. We therefore do not include a detailed comparison with the predictions from Miroschnikow et al. in this manuscript. Nevertheless, we appreciate that such an analysis would be an important next step for the field and a useful resource for future studies.

      We are grateful for the reviewers’ thoughtful feedback, which has helped us improve the manuscript substantially. We hope that the revised version addresses the concerns raised and better conveys the significance of our work.

    1. eLife Assessment

      This study presents a potentially valuable approach to generate a systems vaccinology framework for enabling the design of optimized poxvirus-based MVA vaccines. The developed Boolean modelling framework provides a novel computational approach to predict immune response dynamics against vaccine candidates, which is validated against previously published data and used to test/predict the response for new genetically modified MVA vaccine candidates. However, the strength of evidence is currently incomplete, as key aspects of model construction, calibration, interpretation, validation, and reproducibility, as well as comparability of the used vaccine candidates, require further clarification and supporting information.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript "A predictive systems vaccinology framework enables rational optimization of MVA-based vaccines" by Deman and co-workers presents an approach to use Boolean models for the optimization of MVA for vaccinations. Different Boolean models are derived/inferred to perform in silico testing, e.g., of knock-outs.

      Strengths:

      The optimization of vaccine platforms is very important, and model-based approaches have proved a powerful framework for in silico testing. As far as I'm aware, this is the first time a comprehensive Boolean model is used for this. The authors make an effort to inform this model from available information and experimental data, using state-of-the-art calibration pipelines.

      Weaknesses:

      (1) Lines 154-158: "Because certain biological processes represented in KEGG (e.g., phosphorylation or ubiquitination) do not have direct logical equivalents, this conversion of signaling pathways into a Boolean network can lead to information loss and disconnection of nodes from the rest of the network. To mitigate this issue, we reconnected isolated nodes back to the main structure using oriented protein-protein interaction (PPI) data from 69, thereby restoring connectivity while preserving directionality of regulation." It is not clear to me how the reconnection addresses the described issue that not all processes can be represented in the selected modelling framework. In this context, I would also appreciate it if the authors could clarify the meaning of your states. Is it the presence of a protein (relating to low/high abundance), the activation status (relating to low/high phosphorylation), or a combination? Depending on this, different Boolean representations should be chosen, and different process information can be used.

      (2) Lines159-160: "To enhance immediate readability and interpretability, we connected the resulting network with the corresponding cellular population abundances analyzed by cytometry in the samples." I would appreciate it if the authors could clarify how the cellular layer and the population layers were connected. Is this related to proliferative potential?

      (3) Line 168++: It is unclear to me which parts of the Boolean network described in the section "Boolean naïve network construction" have been calibrated. Among other things, it would also be interesting to know how many logical expressions were changed by ZhegAlCal compared to the naive model and how these expressions were selected. Is there a regularization aiming to minimize the number of changes? In this context, I would also appreciate a clarification of the data processing. The current text mentions a 20% change compared to baseline, a threshold of 0.05, and a 2-means clustering strategy, yet it is unclear how they interact to obtain the binarized training and validation data.

      (4) Line 267++: The model constructed by the authors describes cell-level processes in infected cells. Yet, the data used in the study - which have previously been published in reference 47 - seem to rather capture population averages over heterogeneous, partially non-infected cells. It is unclear to me why / how this can be compared. I would appreciate a clarification, potentially including a more detailed description of the employed datasets.

      (5) Lines 758-759: "The networks generated and analysed during this study are publicly available in the CellCellective repository (MVA 3 pathways, MVA 6 pathways, YF17D)." I searched for the research but did not find it. In my opinion, it would be important to make the models as well as the implementations for calibration, etc. available. Without this, value and reproducibility are limited. I would encourage the authors to provide a detailed human-readable model description in the supplement.

      (6) Lines 783-785: The GO analysis seems to be performed in comparison to the human genome. Yet, the model contains only 200 nodes, so a substantially reduced fraction. I was wondering if this was considered in the analysis process and if the authors checked how often the enrichments for multiple pathways were driven by the same genes.

      (7) Figure 3: It appears as if the number of considered "network updates" was set to 10 (0 to 9) and that this somehow maps to the experimental time. Yet, the experimental observation times are far from uniform.

    3. Reviewer #2 (Public review):

      Summary

      Boolean network modeling is more commonly used in cancer and developmental biology than in vaccine research. Deman et al. apply this framework to a practical vaccinology problem: they aimed to build a mechanistic, executable computational framework capable of both explaining and predicting how the early innate immune response to the MVA vaccine changes when specific viral genes are altered, with the longer-term goal of using that framework to guide the rational design of improved MVA-based vaccines. The authors aimed to: (i) construct and calibrate a Boolean network model of the MVA-induced immune response against real longitudinal non-human primate (NHP) data; (ii) test whether the calibrated model, without being fit to this new data, could reproduce the outcomes of several previously published MVA gene-deletion mutants; and (iii) compare this MVA model to an analogous model of the well-established YF-17D yellow fever vaccine, in the hope of identifying specific molecular targets that could reorient the MVA response toward the durable, single-dose protection YF-17D is known to provide.

      In pursuit of these aims, the authors construct an executable Boolean network of the innate immune response to the MVA vaccine by merging three KEGG pathways (cytosolic DNA-sensing, apoptosis, NF-κB signaling) with cell-population data, and calibrate it against a small NHP dataset (n=3 macaques, 7 time points; Rosenbaum et al., ref. 47). They show the calibrated network reproduces 86-87% of the observed binarized states, and that forcing the network to mimic known MVA deletion mutants (e.g., the triple mutant ΔC6L/ΔK7R/ΔA46R) reproduces qualitative features (e.g., early IFN-β, TNF-α, and IL-6 upregulation) reported in independent published mouse and cell-line studies. They then build a second, six-pathway "consensus" network shared between MVA and YF-17D, compare the two networks' topology and dynamics, and use this comparison, together with a graph-theoretic search for "highly effective" signaling paths, to propose two previously untested MVA deletion mutants (ΔK7R and ΔF17R) predicted to shift the MVA response toward more YF-17D-like features.

      Strengths

      The overall workflow (Figure 1) is clearly described. The calibration approach-binarizing longitudinal cellular and transcriptomic data and fitting network trajectories with the ZhegAlCal algorithm (a Zhegalkin-polynomial/SAT-solving-based method for fitting Boolean trajectories to binarized time-series data)-appears to be a defensible, well-reasoned way to translate a literature-derived network into real kinetic data.

      The retrospective validation against independent published MVA mutants (deletions in C6L, K7R, A46R, and N2L, and separately an A21L point-mutant with three alanine substitutions rather than a deletion) is a genuine strength: the model's qualitative behavior (upregulation of IFN-β, TNF-α, IL-6; limited change in RIG-I) aligns with what those studies reported. We also appreciated that the authors report instances of partial disagreement alongside their successes (e.g., CCL5/RANTES) - that kind of candor about where the model doesn't quite line up is exactly what gives the parts that do line up more credibility.

      The static topological analysis - hub identification, "determinative power" and "vertex betweenness" (two complementary measures of how much a node's state constrains, or lies on paths between, the rest of the network), and "effective graphs" (a measure of how deterministically an edge's regulator sets its target's state) - is a thoughtful use of graph theory to complement the dynamic simulations.

      The authors also clearly discuss the Boolean formalism's core approximation, i.e., that binary on/off states can dilute real but subtle quantitative differences (lines 679-688), and that this work was based on modeling blood-only responses rather than those responses that occur at the vaccination site or draining lymph nodes (lines 709-714).

      Weaknesses

      A few things gave us pause as we read, which we raise here in the spirit of strengthening what already strikes us as a promising framework.

      The MVA/YF-17D comparison starts from two vaccines that are already known to differ substantially. The manuscript uses the divergence between the MVA and YF-17D Boolean networks as an entry point for identifying "MVA optimization" opportunities, but MVA and YF-17D are, on their face, very different vaccine platforms. MVA is a non/limited-replicating DNA poxvirus vector, sensed mainly through cytosolic DNA/cGAS-STING pathways, dosed intradermally, and typically requiring two doses for optimal protection. YF-17D, by contrast, is a live, replicating, attenuated RNA flavivirus, sensed through multiple TLR/RIG-I pathways, and given as a single subcutaneous dose that confers durable, often lifelong, protection (see refs 21, 31, 38-44 in the manuscript). Given this, it's not surprising that the authors themselves report "almost opposite behaviors" for several core cell populations - classical monocytes, B cells, NK cells, and CD4/CD8 T cells - between the two calibrated networks (lines 647-657).

      This stated motivation made us question how much of that divergence reflects a real, actionable difference in vaccine-induced immune programming (the paper's implicit premise) versus differences in virus biology, dosing route, or study/technical design (different sampling schedules, microarray vs. RNA-seq, n=3 vs. n=12 animals) that a Boolean network comparison can't easily tease apart. To their credit, the authors' Discussion is candid on this point-stating directly that "no clear optimization strategies for the MVA viral vector came to mind from the comparison" (lines 664-666)-a useful signal that the comparison's direct yield was limited. The two candidate mutants that emerged instead came from intersecting the model's high-impact nodes with a pre-existing, literature-curated list of MVA immunomodulatory genes (yielding five candidate deletions), which were then individually simulated and narrowed down to the two, ΔK7R and ΔF17R, that produced notable changes - not from the MVA/YF-17D comparison alone.

      The motivation for choosing Boolean modeling over ODE/PDE approaches could be clearer. The choice is motivated mainly by precedent - the authors note it "has rarely been used to model vaccine-induced immune responses, in contrast to statistical modeling or ordinary differential equation (ODE)-based modeling" (lines 100-105) - and by practical considerations, such as not requiring kinetic parameters and being tractable at the scale of networks with hundreds of nodes. What we found ourselves wanting was a more explicit account of the trade-off: ODE and PDE models already simplify the true, spatially resolved, continuously varying underlying biology, and Boolean modeling is a further simplification on top of that. A clearer statement of why this additional simplification is acceptable, or even preferable, for this application (for example: the scale of the curated network, the absence of measured rate constants for most edges, or the interpretability of discrete states) would greatly improve this work.

      We found the manuscript dense and long relative to the size of its central, generalizable findings. This is a presentation issue rather than an evidentiary one. The Results section narrates GO-enrichment interpretation update-by-update for four separate network trajectories (unperturbed MVA, perturbed MVA mutants, the MVA arm of the consensus network, and YF-17D), much of which restates what the (extensive) supplementary figures already show.

      The paper's only prospective predictions are experimentally untested. The two new mutants proposed at the end of the paper (ΔK7R, ΔF17R) are in silico predictions only; they haven't been constructed or tested experimentally in this study. We read the title's claim of a "predictive" framework and the paper's "rational design" framing as best describing a hypothesis-generation tool validated by retrodiction of previously published phenotypes, rather than a demonstration that these two newly proposed mutants will behave as predicted in vivo.

      Did the authors achieve their aims, and do the results support their conclusions?

      Looking at each aim in turn, the picture that emerges is mixed. The first aim - building and calibrating a Boolean network model of the MVA-induced immune response - is convincingly achieved: the calibrated network fits the underlying data well (86-98% of binarized states correctly reproduced, depending on which of the two networks is considered), and its successive states correspond to biologically sensible processes (early chemotaxis, then T-cell activation, then antiviral/ROS-related signatures) that track what is independently known about the innate response to poxvirus vaccination. The second aim-showing the calibrated model can reproduce, without being fit to it, the outcomes of previously published MVA mutants-is also substantially achieved, with the caveat noted above that agreement is qualitative and directional rather than exact, and not uniform across every marker tested.

      The third aim is where the results, in our reading, support the paper's conclusions least well. The stated purpose of comparing the MVA and YF-17D networks was to identify actionable strategies for reorienting MVA's response, but the authors themselves report that the comparison alone yielded no clear optimization strategy (lines 664-666); as noted above, the two candidates that are ultimately proposed came instead from that separate gene-list intersection and simulation step - not from the YF-17D comparison in the more direct way the framing implies. Given that, the paper's headline conclusion - that this framework "enables rational optimization of MVA-based vaccines" - reads to us as only partially supported by what is actually shown: the framework is well demonstrated as a tool for capturing and reproducing known immune biology, but its capacity to prospectively guide the design of a better vaccine remains, at this point, an untested hypothesis rather than a demonstrated result.

      Likely impact and utility to the community:

      The most durable contribution of this paper, regardless of how the two specific candidate mutants eventually fare in the laboratory, strikes us as methodological: the explicit workflow for merging curated signaling pathways into a large executable Boolean network and calibrating it against longitudinal experimental data (building on the authors' own previously published calibration method, reference 75) is clearly described and, together with the deposition of the calibrated networks on the public CellCollective platform, should be usable by other groups working on other vaccines or viral vectors. That reusability is a genuine and useful contribution to the systems-vaccinology toolkit, independent of whether MVA specifically benefits from it.

      Its more immediate, practical utility is harder to gauge from the paper alone. For vaccine developers specifically interested in MVA, the value of this work currently lies in the two testable hypotheses it generates (ΔK7R, ΔF17R) rather than in validated design guidance, since neither candidate has been built or tested here. It's also worth flagging that the framework's generalizability beyond MVA and poxviruses is untested within this paper - the approach is demonstrated for one vector and one comparator vaccine, so readers working on other vaccine platforms may want to treat it as a promising template to adapt and validate for their own systems, rather than as a result that has already been shown to transfer.

    1. eLife Assessment

      This study presents a valuable finding on a potential new regulatory mechanism for cellular proliferation in the mammalian cochlea, taking advantage of the organoid-forming potential of cells from the greater epithelial ridge (GER), or Kolliker's organ, a transient structure in the developing cochlea. The authors highlight three potential regulators - galectin 1, galectin 3, and Myc - providing solid evidence for potential roles in modulating the proliferation of cochlear GER cells using gene expression studies, single-cell profiling, pharmacological inhibition, and genetic overexpression both in vitro and in vivo. However, some of the analyses and interpretations are incomplete, and a better understanding of how the compounds and the gene overexpression studies affect the cells and modulate cell survival/death would aid a fuller interpretation of the data.

    2. Reviewer #1 (Public review):

      Summary:

      The overall aims of this study are a bit unclear. The first experiments use organoids derived from cochlear GER cells in combination with single-cell RNA-seq to try to identify factors that might be important in the initiation of cellular proliferation, although the definition of proliferation is a bit loose and includes the number of organoids, the size of organoids, cell viability, and/or expression of Mki67.

      Based on those results, the authors chose to focus on galectins 1 and 3 and Myc. The reasoning for these choices is a bit unclear, as their ranks in the DE gene list are 51 and 67, and the fold change for each is less than 2. Regardless, the subsequent experiments use inhibitors to examine the effects of galectins and Myc on proliferation of organoids. The results of these experiments do show an effect for inhibition of Lgals1 and Myc, although not Lgals3, but it was unclear whether the effects of these factors on growth could be separated from toxicity treatment, as both OTC008 and 10058-F4 seemed to lead to cell death.

      Next, overexpression of Lgals1, 3 and Myc was actuated in organoids using AAV viruses. The results do show an effect on proliferation, but the results are confusing in that the mRNA expression profiles for two of the transgenes are markedly different in terms of timing, which would not be predicted based on similarities in the constructs. Also, while showing comparable results in some assays, the Myc vector is apparently toxic, killing ~25% of the cells by D9 even though mRNA levels are steady between D5 and D9 in those cells.

      Finally, an in vivo model is used to kill several different types of cochlear cells followed by inhibition of Lgals1. The results of these experiments show a strong inhibition of expression of Ki67 following treatment with OTX008, which is intriguing. However, OTX008 was administered IP, and it does not appear that the ability of OTX008 to cross the blood-labyrinth or even blood-brain barrier has been examined. So it isn't clear whether the results of these experiments indicate a direct or indirect role for OTX008 and galectin-1 in cochlear proliferation. These issues need to be addressed.

      Strengths:

      The results present evidence for potential roles for galectins and myc in the modulation of proliferation of cochlear GER cells. In vitro and in vivo approaches are combined with single-cell profiling to provide a comprehensive analysis.

      Weaknesses:

      (1) Multiple transgenic mouse lines are used in this study, but there are no citations as to where these lines came from, how they were validated, and, for some inducible Cre lines, when the injections of tamoxifen were made.

      (2) Sixty-four organoids were formed per well, but from an average of how many seeded single GER cells? This is not clear (page 5, third paragraph).

      (3) Page 6: Why was cluster 7 grouped with clusters 1,2 and 3? Most cluster 7 cells are from D1.

      (3) In Figure 3A, there does not appear to be a correlation between expression of either galectin-1 or galectin-3 and expression of Mki67, which I would expect would be predicted if these markers play a role in proliferation.

      (4) Figure 3C: A more direct way to examine this would be immunofluorescence for galectin-1 and galectin-3 on cochlear tissue. This would also indicate whether galectin expression correlates with the Sox2+/Fgfr3- population of GER cells.

      (5) For the data shown in Figure 4, what were the experimental conditions? In particular, how long in culture? One interpretation of the data in 4B and E is a decreased increase in the number of organoids, but an alternative is that the treatments are toxic and the organoids are dying. Based on a comparison with the results for myc inhibition, isn't cell toxicity in response to treatment with OTX008 or GB1107 the more likely explanation?

      (6) I think the data in Figure 5 show that the inhibitor experiment demonstrates that the inhibitors, or their targets, are required for organoid survival, as the number of organoids drops to 0, which must be below the starting value.

      (7) It is suggested (page 11, third paragraph) that galectins and myc could be linked or independent effectors of organoids. But couldn't this be tested by combining the inhibitors in the same experiment?

      (8) On page 12, it seems AAV infection of the target cell population prevented organoid formation? This could be a major concern. If nothing else, doesn't this suggest that the effects observed in these experiments might be a result of induced organoid formation from other cochlear duct cells? Also, was expression of the transgenes (Lgals or Myc) confirmed in a cell type that is normally negative for those genes?

      (9) The data in Figure 6C are confusing. The rate of mRNA expression from the AAV transgene should be comparable regardless of the construct given that the promoter is the same. But the results suggest a significant difference in the behavior of the two vectors, with Myc levels reaching a 15-fold increase in just three days while the Lgals vector is at only half that level after 7 days.

      (10) An increase that is not significant is not an increase and should not be described as one (page 13 in the first paragraph).

      (11) In the AAV-Myc experiments, the overall level of mRNA for Mki67 on D9 is comparable to that in the AAV-lgals1 AAV (Figure 6B), but 25% of the cells are dead (page 13, first paragraph)? Similarly, in Figures 6E and 6F, the number of organoids in the AAV-Myc samples is significantly larger than in either control or Lgals, but are most of those cells dead, then?

      (12) Regarding the isolation process in Figure 7A, I am concerned this will also isolate cells from the stria vascularis? Do they retain a greater potential for growth that might lead to their predominance in the growth assay?

      (13) Was the Ki67creERT2 used to label a subset of cells for FACS (page 14)? If not, why was this included? If so, when was the induction made? And doesn't this bias the selection to cells that were proliferating at the time of the induction?

      (14) It is stated that "proliferation is most active at P4 with robust cycling of cells observed in the lateral GER". But then on the following page (page 15), it's stated that the single cell data indicates essentially no proliferating cells in the control, even though there are a lot of lateral GER cells. Can the authors give an explanation for this discrepancy?

      (15) In the first figures in the study, the isolation approach collected lateral GER cells and identified Lgals and Myc as important for organoid expansion (page 15). In Figure 8, there appears to be no change in Lgals or Myc expression in lateral GER cells in response to the damage. Instead, it is medial GER cells that appear to have increased Lgals1 and Myc. And from Figure 8H, are those increases significant?

      (16) A quick search of the literature suggests that there is no evidence that OTX008 can cross the blood-labyrinth or blood-brain barrier (page 15). Was this examined by the authors?

    3. Reviewer #2 (Public review):

      Summary:

      The study uncovers novel factors driving proliferation in the greater epithelial ridge (GER), a proliferative tissue in the neonatal cochlea that may hold important clues on the quest for hair cell regeneration via proliferative means in the adult cochlea.

      Strengths:

      The strengths include the use of both cochlear organoids and in vivo mouse models combined with pharmacological and genetic approaches to inhibit or overexpress proliferative targets identified in the RNA-Seq analysis from the FACS-sorted GER cells. The genetic and pharmacologic manipulation experiments are very strong and convincingly demonstrate that galectins 1 and 4 and Myc are necessary (and in some cases sufficient) to drive cell proliferation in cochlear tissue. However, the real treasure trove is the carefully generated RNA-Seq dataset itself, which offers a wealth of additional differentially-expressed genes that likely contribute to cochlear cell proliferation.

      Weaknesses:

      The primary weakness is that the initial genes studied here, galectins 1 and 3 and Myc, are all associated with tumor formation or cancer progression, so targeting these genes raises concerns about tumor formation in the cochlea.

    1. eLife Assessment

      This is an important study that uses a tripartite transdiagnostic framework to separate depression-specific, anxiety-specific, and shared psychopathology dimensions and relate them to mood variability and mood reactivity to reward prediction errors across several large non-clinical cohorts and a clinical sample. The evidence is compelling: large samples, a well-characterised gambling task, rigorous computational and psychometric analyses (i.e., split-half replication of the factor structure, convergent results with non-orthogonalised factors, an explicitly specified risk-attitude model, diagnostic breakdown and power analyses for the clinical cohort) and replication of the depression-specific blunting of reward prediction error sensitivity in patients. Anxiety-specific associations emerge reliably only when data are pooled and are likely underpowered clinically given co-morbid anxious depression, a constraint the authors now state explicitly. The work advances a mechanistic account of how distinct symptom dimensions shape reward-based mood updating.

    2. Reviewer #1 (Public review):

      Summary of strengths:

      Thank you very much for giving me the opportunity to review this very interesting paper. The research question is intriguing, allowing to address commonly observed co-morbidities between depression and anxiety and their dissociable and opposite relationship to mood fluctuations and sensitivity to reward prediction errors. The computational analyses are very in-depth, including many state of the art checks and validations. Finally, another strength is the inclusion of several large or very large samples, including a patient sample in addition to the general population sample.

      Comments on revised version.

      I want to thank the authors for taking the time to answer all my questions. Their answers were very thoughtful and well argued. I found the theoretical explanations very helpful for explaining their approach and ideas further. In particular, it was fascinating to see how including a single non-orthogonalized depression or anxiety scored show no effect, but including them in simultaneously revealed their previously observed patterns.

    3. Reviewer #2 (Public review):

      Summary:

      Despite their common co-occurrence, depression and anxiety are known to alter mood fluctuations in opposite ways. Here, the authors aimed at distinguishing depression-specific from anxiety-specific from psychopathology-general effects of reward processing on mood fluctuations, focusing on reward prediction errors (RPE) which are known to be linked to mood fluctuations. This mechanistic study aims at uncovering the process through which these psychopathologies are associated with mood modulations. The authors were able to appropriately test their hypothesis and obtained results corroborating their conclusions.

      This work provides a convincing demonstration of the relevance of computational psychiatry (Huys et al, 2016) and the use of decision neuroscience to shed light on the interplay of anxiety and depression and mood.

      Comments on revised version.

      (1) Methodological & Theoretical Framework: The authors used a tripartite model to effectively distinguish depression vs anxiety dimensions from broad psychopathology/distress.

      (2) Possible theoretical confounds: This manuscript addressed adequately the concerns one would have regarding risk-attitudes.

      (3) Computational Rigor: The computational model elegantly separates reward expectations (EV in the model) from outcome processing through RPE, which are two sequential cognitive processes, providing a fine-grained mechanistic account of mood fluctuations.

      (4) Clear & Logical Results Structure: In response to feedback provided during the previous round of review (previously cited as a recommendation for authors), the authors re-organized the Results section into three distinct, easy-to-navigate subsections (Depression, Anxiety, and Depression vs. Anxiety), which substantially improves readability and clarity.

      (5) Neurobiological Context: The Discussion has been enriched with a well-integrated overview of the neural circuits (striatal-midbrain dopaminergic, vmPFC, OFC, and anterior insula) likely underpinning RPE-driven mood updates, which is sure to improve the translational interest of this work.

      (6) Transparent Reporting: The authors had already provided a trustworthy writing approach when referring to trending statistical results. In this revised manuscript, they have been exceptionally transparent regarding study limitations, data collection timelines (addressing potential AI-related artifacts), and statistical power constraints.

      Status of Previous Weaknesses & Suggested Revisions

      (1) Clinical Sample Size and Anxiety-Specific Effects

      Previous Concern: The sample size of the clinical sample (N=116) may not be sufficient to detect anxiety-specific effects due to the high rate of comorbid anxious depression. It would be beneficial to include the number of MDD vs GAD vs anxious depression diagnoses in the clinical population as this would be likely to shine light on the power limitations.

      • Author Revision: The authors have fully addressed this point by adding a diagnostic breakdown in Table S8 which details the diagnosis, illness duration, and medication status. They also included details of the power analysis (Discussion, pages 17-18) putting into perspective their results and provided two possible literature-informed interpretations of their findings. This is also reflected in their updated Abstract.

      (2) Re-organization of Results

      Previous Suggestion: The results sections 2 (depression) and 3 (anxiety) could be improved by reducing the back and forth between factors throughout the results. It may be useful to split them into 3 sections: depression only, anxiety only, depression vs anxiety.

      • Author Revision: The Results section was restructured as suggested, cleanly isolating depression-specific associations, anxiety-specific associations, and direct statistical comparisons between the two ("differential associations" in the manuscript).

      (3) Neurocircuitry Discussion

      Previous Suggestion: In the discussion, authors could have mentioned the brain areas most likely to be involved in these processes, both cognitive and psychopathological, as previous studies (such as Cecchi et al, 2022) have aimed at identifying regions involved in RPE processing while modulating mood in health. A short section on this would be useful to the neuropsychiatric community.

      • Author Revision: A concise section was added to the Discussion mapping computational parameters onto striatal, prefrontal, and insular circuitry (see Strengths #5).

      Conclusion:

      The authors have satisfactorily resolved all minor-to-moderate issues raised in the previous review.

    4. Reviewer #3 (Public review):

      Summary:

      In this submission Wang and colleagues jointly examine the association between depression and anxiety symptoms and individuals' affective reactivity to reward prediction errors in Ruttledge et al.'s gambling paradigm. Taking a bifactor approach to anxiety and depression in several non-clinical (and one clinical sample), the authors find that anxiety-specific symptoms relate to over-reactivity of mood to reward prediction errors (RPEs) as well as heightened mood variability , while depression-specific symptoms relate to blunted mood sensitivity to RPEs. These depression-, but not-anxiety specific relationships replicated in patient samples.

      Strengths:

      I was impressed that the data-driven, transdiagnostic approach employed by the authors uncovered specific relationships between anxiety and depression-specific factors and RPE reactivity in a well characterized task and computational model, especially in a non-clinical sample. This sheds new light on how these affective processes may be perturbed-and importantly, in different ways-by anxiety and depression symptoms. Likewise, the replication of the depression-specific finding (RPE hypo-reactivity) in a clinical sample was nice to see.

      Weaknesses:

      While the anxiety- and depression-specific factors had differential effects on mood variability (Fig 2A-D) and RPE reactivity (Fig 2E-G) in all samples, such that the correlations between the two factors and these mood parameters were significantly different, the anxiety factor was not consistently (significantly) associated with either mood-related parameter across samples. However, the authors resolve anxiety-specific predictive effects when they collapse across datasets. While it is intuitive that achieving a larger effective sample size would afford the power necessary to detect such individual differences, this struck me as a major caveat for this set of results.

      The associations the authors observe between the 'common factor' of depression and anxiety and risk-aptitudes tendencies-presumably the alpha (exponent) parameter in a prospect theory-type subjective value model. But where is this analysis explained? (i.e. how was this model formulated and how were risk attitude parameters estimated?) And what is the interpretation of this finding-is there precedent for looking at risk attitudes in this task? And why would these predictive effects only be observed in relation to the common, but not unique factors of anxiety and depression?

      Comments on revised version.

      I believe the weaknesses identified in the previous round of review have been adequately addressed by the authors, and my suggestions concerning clarity of presentation have by and large been implemented by the authors.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study uses a tripartite transdiagnostic computational framework to distinguish depression-specific, anxiety-specific, and shared psychopathology dimensions, in their relationships to mood variability and mood reactivity to reward prediction errors across multiple large non-clinical cohorts and a clinical sample. The evidence is convincing overall because the study combines large samples, a well-characterized gambling task and in-depth computational and psychometric analyses, and it replicates the depression-specific association with blunted reward prediction error-sensitivity in a clinical sample. However, the anxiety-specific effects are less consistently supported across individual datasets, may be underpowered in the clinical cohort because of comorbidity, and some aspects of the factor-analytic, risk-attitude, and mediation analyses would benefit from clearer explanation. These findings advance a mechanistic account of how distinct symptom dimensions differentially shape reward-based mood updating and variability, providing a principled framework for future transdiagnostic modeling.

      We thank the editors and reviewers for this important assessment.

      Regarding inconsistent results for anxiety-related effects in healthy datasets. Although the anxiety-specific factor showed associations in the expected direction across healthy datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level. Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity in non-clinical participants (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). In addition, we conducted a mini meta-analysis, and results support that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE in non-clinical participants. We have clarified this point below:

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β<sub>RPE</sub> were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept to account for dataset-level variability. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001). These associations remained significant after controlling for gender, age, task earnings, and mood drift. In addition, we performed a mini meta-analysis on these correlation coefficients[49]. Results showed significant positive correlation for both mood variation and RPE-related mood sensitivity (mood variation: Z = 3.399, 95 % CI for correlation coefficient r [0.045, 0.166]; RPE-related mood sensitivity: Z = 2.618, 95 % CI for correlation coefficient r [0.021, 0.143]), supporting that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE in non-clinical participants.”

      We also admit that the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      Regarding be underpowered sample size in the clinical cohort. We agree that the clinical sample may have been underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). Although these covariate analyses support the robustness of the depression-related effect, they do not resolve whether the absence of the anxiety-related effect reflects limited power or true clinical discontinuity. We also revised the Discussion to explicitly acknowledge that the anxiety-related effect observed in the pooled non-clinical dataset was not replicated in the clinical sample. We now note two possible interpretations. First, this discontinuity may reflect limited statistical power in the clinical sample. Second, and more speculatively, it may reflect a disruption of mood homeostasis in affective disorders (Paulus, 2007). In non-clinical individuals, the counterbalancing associations of depression- and anxiety-related traits with mood variation may contribute to emotional equilibrium. In contrast, affective disorders may involve a loss of this regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals. We have revised the manuscript as follows:

      Pages 12-13:

      “To test whether abnormalities in RPE-driven mood fluctuations can serve as clinically relevant computational markers of depression- and anxiety-related symptom dimensions, we recruited patients with affective disorders (n = 116) to complete the same questionnaire battery and gambling task with momentary mood ratings (Figure 1). Demographic, psychological, and clinical characteristics are summarized in Table 1 and Table S8. We observed significant negative correlations between depression-specific scores and both mood variation (r = -0.239, p = 0.009) and RPE-related mood sensitivity (β_RPE; r = -0.216, p = 0.020). These associations remained significant after controlling for demographic and clinical covariates, task earnings, and mood drift (ps < 0.05). Bootstrap validation yielded consistent results. Mediation analyses further showed that reduced mood sensitivity to RPEs statistically mediated the association between depression-specific scores and lower mood fluctuations (a × b = -0.141, 95% CI = [-0.261, -0.038], p = 0.021; Figure 3). However, we did not observe significant correlation with anxiety (mood variation: r = -0.092, p = 0.327; β_RPE: r = -0.095, p = 0.311).”

      Pages 17-18:

      “Notably, the pattern of heightened RPE sensitivity observed in the pooled non-clinical dataset was not observed in the clinical sample. On the one hand, this discontinuity may reflect that the clinical sample was underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). On the other hand, it may reflect a disruption of mood homeostasis in clinical populations[41,58]. In non-clinical individuals, counterbalancing associations of depression- and anxiety-related traits with mood variation may help maintain emotional equilibrium. In contrast, affective disorders may involve a loss of such regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals.”

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      We have also revised the manuscript accordingly to make the factor-analytic, risk-attitude, and mediation analyses clear.

      Reviewer #1 (Public review):

      Summary:

      This is a very interesting paper. The research question is intriguing, allowing the authors to address commonly observed comorbidities between depression and anxiety and their dissociable and opposite relationship to mood fluctuations and sensitivity to reward prediction errors. The computational analyses are very in-depth, including many state-of-the-art checks and validations. Another strength is the inclusion of several large or very large samples, including a patient sample in addition to the general population sample.

      I have the following questions:

      (1) Factor analysis

      I found the hierarchical organization of the factors interesting. While this is a very common procedure in, for example, the field of intelligence (producing sub-scores and a general g factor), it is not yet very commonly used in the field of computational psychiatry (though it has been validated before for anxiety/depression, so it is used here with good reason). I was also impressed by the methodological depth. In particular, it was of note how thoroughly done it was (for example, repeating the EFA on the second half of the data set). I have one question though: is the sample size too small for the exploratory analyses, given the number of items? Given the stability across the half-split, I imagine it is not. Perhaps the authors could spell out how many items, what would be the recommended standard for a subject-to-item ratio, and comment on this. A very technical point, the authors should specify how they extracted the factor scores from the other data sets (is it using the Thurstone or Bartlett method)? From experience (though not doing a hierarchical factor analysis), Bartlett can be somewhat better compared to the default (Thurstone) - better as in the resulting factors more closely recapitulating the factor correlations in the original sample (and independence of responses of other participants in a sample for computing a person's factor score). Could you also comment on similarities or divergences in this hierarchical factor analysis approach from another one recently used transdiagnostically in Wise et al. (2026, Translational Psychiatry)?

      We thank the Reviewer for the positive evaluation of our hierarchical factor-analytic approach and for recognizing the methodological depth of our analyses. We are particularly grateful for the Reviewer’s constructive suggestions regarding the participant-to-item ratio, factor score extraction, and the relation between our approach and recent transdiagnostic hierarchical factor-analytic work.

      First, hierarchical organization. As the Reviewer noted, hierarchical and bifactor representations have a well-established tradition in intelligence research, where they are used to model the g factor alongside domain-specific abilities (e.g., Reise, 2012; Rodriguez et al., 2016). Crucially, the use of such hierarchical structures in the present study was motivated primarily by theory and evidence from anxiety and depression research, rather than by analogy to intelligence research alone. This tradition can be traced back to the tripartite model of anxiety and depression (Clark & Watson, 1991), which distinguished a broad shared component of general distress or negative affect from more specific anxiety- and depression-related components. Subsequent psychometric work has further supported bifactor and hierarchical representations of anxiety and depression symptoms, including models that separate a general internalizing/distress factor from symptom-specific dimensions (e.g., Simms et al., 2008). More recently, similar hierarchical symptom structures have also been adopted in computational psychiatry to relate shared and specific affective symptom dimensions to task-derived computational parameters (Gagne et al., 2020, 2022; see Wise et al., 2023 for a review). Thus, the bifactor structure used here provides a theoretically motivated way to capture both the variance shared by anxiety and depression and the symptom-specific variance relevant to our computational analyses. We have clarified this point in the revised manuscript as follows:

      Pages 3-4:

      “Recent work has used bifactor models of the tripartite model of depression and anxiety to clarify their distinct features and differential influences on decision-making[31,32]. The tripartite model of anxiety and depression proposes that these two symptom dimensions share a broad general distress or negative affect component while also including symptom-specific components: low positive affect/anhedonia is more specific to depression, whereas physiological hyperarousal is more specific to anxiety[30,33,34]. Bifactor analysis offers a way to model this structure statistically. In a bifactor model, symptoms load on a general factor reflecting their shared variance and on specific factors capturing residual variance in narrower symptom dimensions after accounting for the general factor. Although bifactor and hierarchical models have long been used in psychometrics, e.g., intelligence research[35,36], their application to anxiety and depression is grounded in the tripartite model and subsequent psychometric work distinguishing general internalizing/distress from symptom-specific dimensions. This framework has recently been extended to computational psychiatry, where shared and specific affective symptom dimensions have been linked to task-derived computational parameters. For example, Gagne et al. (2022) used bifactor analysis to show that depression was associated with weaker prior beliefs, whereas anxiety was associated with a stronger negative bias in belief updatin31.”

      Second, the participant-to-item ratio. We agree that the ratio of 450 participants to 128 items in the EFA split-half sample, approximately 3.5:1, is below some conventional sample-size recommendations for exploratory factor analysis, including the often-cited recommendation of five participants per item (Costello & Osborne, 2005). However, as the Reviewer noted, the split-half analysis showed a stable factor structure, and the independent CFA in the other split-half sample further supported the robustness of the solution. Importantly, participant-to-item ratios are only one criterion for evaluating factor recovery. De Winter, Dodou, and Wieringa (2009) demonstrated that reliable EFA solutions may be obtained even with relatively small samples when the data are well-conditioned, such as when factor loadings are high, the number of factors is small, and each factor is defined by multiple items. These conditions were largely met in our data. We have clarified this point in the revised manuscript as follows:

      Supplementary Page 3:

      “In addition, the ratio of 450 participants to 128 items in the EFA split-half sample, approximately 3.5:1, is below some conventional sample-size recommendations for EFA, including the often-cited recommendation of five participants per item[8]. However, the split-half EFA yielded a stable factor structure, and the independent CFA further supported the robustness of this solution. Moreover, reliable EFA solutions may be obtained even with relatively small samples when the data are well-conditioned, such as when factor loadings are high, the number of factors is small, and each factor is defined by multiple items9. These conditions were largely met in our data.”

      Next, factor score extraction. We followed prior work using bifactor modeling in computational psychiatry (Gagne et al., 2020) and extracted factor scores with the Anderson–Rubin method, implemented using psych::factor.scores with method = "Anderson". This approach yields standardized and mutually orthogonal factor scores, which is particularly appropriate for our subsequent correlation analyses because it produces orthogonal scores and therefore avoids multicollinearity among the general, depression-specific, and anxiety-specific factors. In addition, as suggested by the Reviewer, we extracted factor scores using the Bartlett method from an oblique bifactor model, which allowed the depression- and anxiety-specific factors to correlate. In the combined dataset (n = 1,026), anxiety- and depression-specific scores were significantly correlated when extracted using the Bartlett method (r = 0.638, p < 0.001), whereas, as expected, they were effectively uncorrelated when extracted using the Anderson–Rubin method (r < 0.001, p = 1.000). This comparison suggests that the Anderson–Rubin method is more appropriate for our analytic aim of estimating the unique associations of shared and symptom-specific components with task-derived parameters, because it separates the general distress/internalizing factor from the statistically separable residual anxiety- and depression-specific components. For this reason, we retained the Anderson–Rubin factor scores in the main analyses. We have clarified this point in the revised manuscript as follows:

      Supplementary Page 3:

      “For factor score extraction, we followed prior work using bifactor modeling in computational psychiatry[10] and extracted factor scores with the Anderson–Rubin method, implemented using psych::factor.scores with method = "Anderson". This approach yields standardized and mutually orthogonal factor scores, which is particularly appropriate for our subsequent correlation analyses because it avoids multicollinearity among the general, depression-specific, and anxiety-specific factors. As a robustness check, we also extracted factor scores using the Bartlett method from an oblique bifactor model, which allowed the depression- and anxiety-specific factors to correlate. In the combined dataset (n = 1,026), anxiety- and depression-specific scores were significantly correlated when extracted using the Bartlett method (r = 0.638, p < 0.001), whereas, as expected, they were effectively uncorrelated when extracted using the Anderson–Rubin method (r < 0.001, p = 1.000). This comparison suggests that the Anderson–Rubin method is more appropriate for our analytic aim of isolating the unique contributions of shared and symptom-specific variance, because it separates the general distress/internalizing factor from residual anxiety- and depression-specific components. Thus, we retained the Anderson–Rubin factor scores in the main analyses.”

      Finally, the similarities and differences between our hierarchical factor-analytic approach and the recent transdiagnostic hierarchical factor-analytic approach of Wise et al. (2026). Both approaches fit EFA models with different numbers of factors and use cross-level correlations to characterize hierarchical symptom structure. However, Wise et al. (2026) applied this framework to a broader symptom battery covering transdiagnostic and neurodevelopmental dimensions, identifying a hierarchy that included a general psychopathology factor and more specific dimensions such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal. In contrast, our study focused more narrowly on anxiety and depression dimensions, with the goal of deriving symptom factors that could be linked to task-derived computational parameters. Accordingly, whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-related variance remains unknown. Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026), to better determine whether the links among symptom dimensions, RPE-related mood sensitivity, and mood variability are disorder-specific or transdiagnostic. We have discussed this point in the revised manuscript as follows: Page 19:

      “Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026)[59], to better characterize whether links among symptom dimensions, RPE sensitivity, and mood variability are disorder-specific or transdiagnostic.”

      (2) Linking factors to task parameters

      As I understand it, the authors relate the orthogonalized depression/anxiety to task parameters (sensitivity to RPEs on mood and mood variations) using correlations. In order to have a better understanding of how this relates to other commonly used approaches, I would pose two questions:

      (i) What are the correlations when the full (non-orthogonalized) factor scores for depression and anxiety are used? Are the signs the same?

      (ii) What are the results when, instead of the independent correlations, the authors perform b_RPE ~ anxiety + depression (again using the non-orthogonalized factors)? I'm assuming all of these analyses should give the same results if the authors' hypothesis of opposing effects of anxiety and depression holds true.

      We thank the Reviewer for these helpful comments. Our original analyses used orthogonalized depression- and anxiety-specific factor scores because this approach is aligned with our analytic aim of separating shared and symptom-specific variance within the tripartite/bifactor framework of anxiety and depression (Clark & Watson, 1991), and has been used in prior work (Gagne et al., 2020, 2022; see Wise et al., 2023 for a review). Orthogonalization allows us to statistically separate the shared distress component from the symptom-specific components of anxiety and depression, which was central to our hypothesis regarding their opposing associations with task-derived parameters. As expected, the orthogonalized anxiety- and depression-specific factor scores were uncorrelated in the combined dataset (r < 0.001, p = 1.000; n = 1,026). By contrast, the full non-orthogonalized depression and anxiety scores retained substantial shared variance and were highly correlated (r = 0.638, p < 0.001; n = 1,026), making their separate associations less straightforward to interpret.

      Nevertheless, we agree that analyses using the full non-orthogonalized depression and anxiety scores provide an important comparison with more commonly used non-orthogonal symptom-score approaches. We therefore conducted the analyses suggested by the Reviewer. When the full depression and anxiety scores were entered separately into linear mixed-effects models predicting RPE-related mood sensitivity, with dataset included as a random intercept, the anxiety association was not significant (anxiety: b = 0.002, t = 1.130, p = 0.259; depression: b = -0.008, t = -3.732, p < 0.001). By contrast, when the full non-orthogonalized anxiety and depression scores were entered simultaneously in the same linear mixed-effects model, the original pattern was replicated: anxiety and depression showed opposing associations with RPE-related mood sensitivity (anxiety: b = 0.012, t = 4.619, p < 0.001; depression: b = -0.017, t = -5.851, p < 0.001). This pattern is consistent with a mutual suppression effect: shared variance between anxiety and depression may obscure their unique associations when examined separately, whereas the simultaneous regression model reveals their opposing symptom-specific associations.

      Together, these supplementary analyses support our original interpretation that RPE-related mood sensitivity is associated with the separable anxiety- and depression-specific components in opposite directions. We have revised the manuscript as follows:

      Supplementary Pages 3-4:

      “We further analyzed non-orthogonalized full depression and anxiety scores to assess the robustness of our results. When full depression and anxiety scores were entered in separate linear mixed-effects models predicting RPE-related mood sensitivity, with dataset included as a random intercept, the anxiety association was not significant (anxiety: b = 0.002, t = 1.130, p = 0.259; depression: b = -0.008, t = -3.732, p < 0.001). By contrast, when the full non-orthogonalized anxiety and depression scores were entered simultaneously in the same linear mixed-effects model, the original pattern was replicated: anxiety and depression showed opposing associations with RPE-related mood sensitivity (anxiety: b = 0.012, t = 4.619, p < 0.001; depression: b = -0.017, t = -5.851, p < 0.001). This pattern is consistent with a mutual suppression effect: shared variance between anxiety and depression may obscure their unique associations when examined separately, whereas simultaneous regression reveals their opposing symptom-specific associations. These results support our interpretation that RPE-related mood sensitivity is linked to the separable anxiety- and depression-specific components.”

      Minor comments:

      (1) The authors should write down when the data were collected for each study. This is because AI capabilities have massively increased since ~2020 in quite specific steps (with the public release of new AI models), meaning that AI is likely to have been able to do tasks and questionnaires without detection if data were collected recently.

      We thank the Reviewer for this important comment. We have now added the data collection periods for each dataset in Table 1. The laboratory and clinical dataset were collected in a controlled laboratory setting rather than through online testing. As shown in Table 1, all online experiments were conducted before November 2022, prior to the public release of ChatGPT and its broad entry into public awareness. Therefore, our data were unlikely to have been substantially affected by AI-assisted responding. We have clarified this point in the revised manuscript as follows:

      Page 20:

      “See Table 1 for demographic information and data collection periods. Because online data collection may raise concerns about AI-generated responses, we note that artificial intelligence tools, such as ChatGPT, became widely known to the public in November 2022, whereas all online experiments in the present study were conducted before November 2022 (see Table 1). Therefore, these data were unlikely to have been substantially affected by participants’ use of AI tools.”

      (2) The authors should include a statement in the methods section that checks for AI were done. If none yet, could you do any? Recent papers (Westwood, PNAS 2025; van der Stigchel PNAS, 2026) point to the risk since at least the release of o4-mini (used in the cited paper to create very human-like behaviour).

      We thank the Reviewer for this helpful comment. We have clarified this point in the revised manuscript as follows:

      Page 20:

      “See Table 1 for demographic information and data collection periods. Because online data collection may raise concerns about AI-generated responses, we note that artificial intelligence tools, such as ChatGPT, became widely known to the public in November 2022, whereas all online experiments in the present study were conducted before November 2022 (see Table 1). Therefore, these data were unlikely to have been substantially affected by participants’ use of AI tools.”

      (3) It would have been good to collect questionnaires of other, thought to be unrelated psychiatric traits, like compulsivity or schizophrenia symptoms, to check the specificity of the results, also under the assumption that higher scores on either of these skewed questionnaires can pick up individual differences in 'bad questionnaire completion'. The authors should comment on the absence of other questionnaires in the discussion in the limitations section.

      We thank the Reviewer for this helpful comment. We agree that the absence of broader psychiatric trait measures limits our ability to evaluate the specificity of the observed associations. Our symptom assessment focused specifically on anxiety and depression because the study was motivated by hypotheses about their potentially opposing links with RPE sensitivity and mood variability. Although previous research has shown intact mood sensitivity to RPEs in individuals with suicidal thoughts and behaviors (Wang et al., 2026), we cannot determine whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-style variance.

      Regarding the concern that higher scores on symptom questionnaires with skewed score distributions may partly capture individual differences in poor-quality questionnaire responding, we note that we implemented strict data-quality procedures for both questionnaire and task data. Four attention-check items were embedded throughout the questionnaire battery, requiring participants to select a prespecified response, for example, “Please select the second option for this item.” Similarly, four attention-check trials were embedded throughout the gambling task. For example, participants were asked to choose between a certain gain of 20 points and a gamble with possible outcomes of 35 and 55 points, for which the dominant response was to choose the gamble option. Participants who failed any of these attention checks were excluded. In addition, our behavioral and mood data reproduced key patterns reported in previous studies (Rutledge et al., 2014 & 2015) using momentary mood ratings during gambling tasks, including higher mood following gains than following losses and systematic mood drift over time (all ps < 0.001). These procedures and validation checks reduce the likelihood that the present findings were driven by poor questionnaire or task completion. We have clarified this point in the revised manuscript as follows:

      Page 19:

      “Second, our symptom assessment focused specifically on anxiety and depression. This choice was motivated by our primary hypotheses, but it limits our ability to evaluate the specificity of the observed associations. Recent work has shown that individuals with suicidal thoughts and behaviors exhibit reduced mood sensitivity to certain rewards (CR), but not to RPEs[49], suggesting that the current RPE-related effects are not driven by suicide-related processes. However, because we did not assess other psychiatric dimensions, such as compulsivity or schizophrenia-spectrum symptoms, we cannot determine whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-related variance. Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026)[59], to better characterize whether links among symptom dimensions, RPE sensitivity, and mood variability are disorder-specific or transdiagnostic.”

      Page 20:

      “Participants were excluded if 1) they failed any of the attentional checks (4 items); 2) they made the same choices for all items; 3) they responded with extreme inconsistency in two similar questionnaires (difference in z-scores out of ±2).”

      Page 21:

      “There were four items for attentional checks, which required the participants to make a specific choice and were embedded in the entire measurements, e.g., ‘please select the second option for this item’.”

      Page 22:

      “We also set 4 trials embedded in the entire task for attentional checks. For example, participants were asked to make a choice between a certain gain 20 and a gamble 35/55, where the correct response for this trial was the gamble choice.”

      Page 5:

      “Choice data (e.g., gambling rates) and mood data (e.g., initial mood, mean mood, and mood variation) showed patterns similar to those reported in previous studies measuring momentary mood during gambling tasks (Figure S2 & S3)[10,45]. We also replicated established effects on momentary mood: mood was higher following gains than following losses, and mood drifted over time (all ps < 0.001; Figure S4).”

      (4) The authors could include a more explicit sentence in the abstract stating that the anxiety result did not hold up in the clinical population.

      We thank the Reviewer for this helpful comment. We have clarified this point in the revised manuscript as follows:

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      Reviewer #2 (Public review):

      Summary:

      Despite their common co-occurrence, depression and anxiety are known to alter mood fluctuations in opposite ways. Here, the authors aimed at distinguishing depression-specific from anxiety-specific from psychopathology-general effects of reward processing on mood fluctuations, focusing on reward prediction errors (RPEs), which are known to be linked to mood fluctuations. This mechanistic study aims at uncovering the process through which these psychopathologies are associated with mood modulations. The authors were able to appropriately test their hypothesis and obtained results corroborating their conclusions.

      This work provides a convincing demonstration of the relevance of computational psychiatry (Huys et al, 2016) and the use of decision neuroscience to shed light on the interplay of anxiety, depression, and mood.

      Strengths:

      The authors used a tripartite model to distinguish depression vs anxiety, as well as a computational model distinguishing reward expectation (EV in the model) from outcome processing through RPE, which are two sequential cognitive processes.

      The manuscript adequately addresses the concerns one would have regarding risk-attitudes and regarding referring to trending statistical results.

      Weaknesses:

      The sample size of the clinical sample (N=116) may not be sufficient to detect anxiety-specific effects due to the high rate of comorbid anxious depression. It would be beneficial to include the number of MDD vs GAD vs anxious depression diagnoses in the clinical population, as this would likely shine light on the power limitations.

      We thank the Reviewer for this helpful comment. We agree that the clinical sample may have been underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (see Table S8 for diagnosis, illness duration, and medication status). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). Although these covariate analyses support the robustness of the depression-related effect, they do not resolve whether the absence of the anxiety-related effect reflects limited power or true clinical discontinuity.

      We also revised the Discussion to explicitly acknowledge that the anxiety-related effect observed in the pooled non-clinical dataset was not replicated in the clinical sample. We now note two possible interpretations. First, this discontinuity may reflect limited statistical power in the clinical sample. Second, and more speculatively, it may reflect a disruption of mood homeostasis in affective disorders (Paulus, 2007). In non-clinical individuals, the counterbalancing associations of depression- and anxiety-related traits with mood variation may contribute to emotional equilibrium. In contrast, affective disorders may involve a loss of this regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals. We have revised the manuscript as follows:

      Pages 12-13:

      “To test whether abnormalities in RPE-driven mood fluctuations can serve as clinically relevant computational markers of depression- and anxiety-related symptom dimensions, we recruited patients with affective disorders (n = 116) to complete the same questionnaire battery and gambling task with momentary mood ratings (Figure 1). Demographic, psychological, and clinical characteristics are summarized in Table 1 and Table S8. We observed significant negative correlations between depression-specific scores and both mood variation (r = -0.239, p = 0.009) and RPE-related mood sensitivity (β<sub>RPE</sub>; r = -0.216, p = 0.020). These associations remained significant after controlling for demographic and clinical covariates, task earnings, and mood drift (ps < 0.05). Bootstrap validation yielded consistent results. Mediation analyses further showed that reduced mood sensitivity to RPEs statistically mediated the association between depression-specific scores and lower mood fluctuations (a × b = -0.141, 95% CI = [-0.261, -0.038], p = 0.021; Figure 3). However, we did not observe significant correlation with anxiety (mood variation: r = -0.092, p = 0.327; β_RPE: r = -0.095, p = 0.311).”

      Pages 17-18:

      “Notably, the pattern of heightened RPE sensitivity observed in the pooled non-clinical dataset was not observed in the clinical sample. On the one hand, this discontinuity may reflect that the clinical sample was underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). On the other hand, it may reflect a disruption of mood homeostasis in clinical populations[41,58]. In non-clinical individuals, counterbalancing associations of depression- and anxiety-related traits with mood variation may help maintain emotional equilibrium. In contrast, affective disorders may involve a loss of such regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals.”

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      Reviewer #3 (Public review):

      Summary:

      In this submission, Wang and colleagues jointly examine the association between depression and anxiety symptoms and individuals' affective reactivity to reward prediction errors in Ruttledge et al.'s gambling paradigm. Taking a bifactor approach to anxiety and depression in several non-clinical (and one clinical sample), the authors find that anxiety-specific symptoms relate to over-reactivity of mood to reward prediction errors (RPEs) as well as heightened mood variability, while depression-specific symptoms relate to blunted mood sensitivity to RPEs. These depression- but not anxiety-specific relationships replicated in patient samples.

      Strengths:

      I was impressed that the data-driven, transdiagnostic approach employed by the authors uncovered specific relationships between anxiety and depression-specific factors and RPE reactivity in a well characterized task and computational model, especially in a non-clinical sample. This sheds new light on how these affective processes may be perturbed-and importantly, in different ways-by anxiety and depression symptoms. Likewise, the replication of the depression-specific finding (RPE hypo-reactivity) in a clinical sample was nice to see.

      Weaknesses:

      (1) While the anxiety- and depression-specific factors had differential effects on mood variability (Figure 2A-D) and RPE reactivity (Figure 2E-G) in all samples, such that the correlations between the two factors and these mood parameters were significantly different, the anxiety factor was not consistently (significantly) associated with either mood-related parameter across samples. However, the authors resolve anxiety-specific predictive effects when they collapse across datasets. While it is intuitive that achieving a larger effective sample size would afford the power necessary to detect such individual differences, this struck me as a major caveat for this set of results.

      We thank the Reviewer for this important comment. Although the anxiety-specific factor showed associations in the expected direction across datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level.

      Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). In addition, we conducted a mini meta-analysis, and results support that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE.

      However, we have clarified in the revised manuscript that the anxiety-related effects were less robust than the depression-related effects and require further replication in larger samples.

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β<sub>RPE</sub> were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept to account for dataset-level variability. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001). These associations remained significant after controlling for gender, age, task earnings, and mood drift. In addition, we performed a mini meta-analysis on these correlation coefficients[49]. Results showed significant positive correlation for both mood variation and RPE-related mood sensitivity (mood variation: Z = 3.399, 95 % CI for correlation coefficient r [0.045, 0.166]; RPE-related mood sensitivity: Z = 2.618, 95 % CI for correlation coefficient r [0.021, 0.143]), supporting that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE.”

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      (2) The authors observe associations between the 'common factor' of depression and anxiety and risk-attitude tendencies, presumably the alpha (exponent) parameter in a prospect theory-type subjective value model. But where is this analysis explained? (i.e. how was this model formulated and how were risk attitude parameters estimated?) And what is the interpretation of this finding - is there precedent for looking at risk attitudes in this task? And why would these predictive effects only be observed in relation to the common, but not unique, factors of anxiety and depression?

      We apologize for the unclear statement. We have added a description of the computational modeling of choice behavior. Please see our revisions below:

      Page 13:

      “Choice parameters were estimated using an established approach–avoidance prospect theory model[10,45,49], which included loss aversion, domain-specific risk attitude parameters in the gain and loss domains, and value-independent Pavlovian approach and avoidance parameters (see Supplementary Note 7 for details of the computational choice models). In this model, risk attitude was quantified by the exponent parameter α in a prospect-theory-inspired subjective value function. Lower α values reflect greater risk aversion, whereas values closer to or above 1 reflect more linear or risk-seeking valuation.”

      Supplementary Pages 8-9:

      Note 7: Computational model of gambling choice

      To quantify how different events impacted participants’ momentary moods during the gambling In line with previous studies[14,15], our choice model space included expected value model (cM1), prospect theory model (cM2)[16], and approach-avoidance prospect theory model (cM3)[14]. For cM2 (Equations 6-9), there were 3 parameters, including risk aversion (α, range: [0.3, 1.3]), loss aversion (λ: [0.5, 5]), and inverse temperature (μ: [0, 10]).

      Where V<sub>gain</sub> and V<sub>loss</sub> are the objective gain and loss from a gamble, respectively. Please note thatV<sub>gain</sub> is 0 in loss trials and V<sub>loss</sub> is 0 in gain trials. V<sub>certain</sub> is the objective value for the certain option. U<sub>gamble</sub> and U<sub>certain</sub> denote the subjective utilities of the gamble and the certain option, respectively. Choice probability for gamble (P<sub>gamble</sub>) is determined by the softmax rule. Building on cM2, cM3 decomposes the decision process into risk-attitude-driven valuation (e.g., loss and risk aversion) and value-insensitive motivational components (Equations 6-8 & 10-12). That is, choice probability for P<sub>gamble</sub> in cM3 is jointly determined by the softmax rule and approach/avoidance parameters (β<sub>gain</sub>: [-1, 1], β<sub>loss</sub>: [-1, 1]). Approach/avoidance parameters are not applied in mixed trials. Please note that a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups. Risk attitude is indeed conceptualized in economics as the curvature of the utility function (i.e., the subjective value) of the objective outcomes, with concave curves associated with risk aversion, and convex curves associated with risk seeking[17,18]. By contrast, the approach or avoidance bias apply to all the value. A possible interpretation of the approach bias is that participant approach the option with the highest possible gain (the lottery) in the gain frame; the avoidance bias would then reflect a tendency to systematically avoid the highest potential losses (the lottery) in the loss frame.

      Model comparison using BIC revealed that the winning model for each dataset was the approach-avoidance prospect theory model (cM3; mean R<sup>2</sup> = 0.51 for the laboratory dataset, 0.49 for the online dataset1, 0.54 for the online dataset 2, and 0.40 for the clinical dataset; Table S9).

      Please also see our interpretation of this finding below:

      Page 18:

      “With respect to decision-making, prior literature using risky decision-making tasks without feedback has linked pathological anxiety to greater risk aversion[58]. In line with this, our results from a risky decision-making task with feedback suggest that the common factor, rather than anxiety-specific variance per se, is more consistently associated with risk aversion. This suggests that heightened gain-domain risk aversion may be a transdiagnostic feature of internalizing psychopathology, rather than being uniquely attributable to anxiety.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Thank you very much for giving me the opportunity to review this very interesting paper.

      Recommendations:

      (1) Add more specific ethics information than "study was approved by ethics committee of Beijing normal university".

      We thank the Reviewer for this important comment. We have added approval number. Please see our revision below:

      Page 20:

      “The study was approved by the Ethics Committee of Beijing Normal University (approve number: ICBIR_A_0016_028). Written or electronic informed consent was obtained from all participants before participation.”

      (2) Add information on how participants were recruited. I think the websites listed only hosted the experiment/questionnaires?

      We thank the Reviewer for pointing this out. We have revised the relevant text as follows:

      Page 20:

      “A total of 2634 participants via online platforms (questionnaires from https://www.wjx.cn and tasks from https://www.naodao.com) took part in five experiments, including a psychometric experiment, a laboratory experiment, two online replication experiments. Participants were recruited through participant pools and study advertisement. For online experiments, interested participants accessed the study through an online link and completed the questionnaires and task remotely. For the laboratory experiment, participants completed the study in a controlled laboratory setting.”

      (3) Typo in Figure 1A, grey panel - psychometric.

      We apologize for the typo. We have corrected typographical errors throughout the manuscript.

      (4) In the Discussion, there is a section on r-to-z transformations, and I was not quite sure what in the Results this links to.

      We thank the Reviewer for pointing out the unclear statement. Please see our revision below:

      Page 18:

      “First, although anxiety- and depression-related associations differed consistently, the anxiety-specific associations themselves were less robust across datasets.”

      **Reviewer #2 (Recommendations for the authors):&&

      The Results sections 2 (depression) and 3 (anxiety) could be improved by reducing the back and forth between factors throughout the results. It may be useful to split them into 3 sections: depression only, anxiety only, and depression vs anxiety.

      We thank the Reviewer for this helpful suggestion. As suggested, we have reorganized this part into three sections: depression, anxiety, and depression versus anxiety. Please see our revision below:

      Page 12:

      “Differential associations of depression and anxiety with mood fluctuations. To directly test whether depression- and anxiety-specific factors differed in their associations with mood dynamics, we compared the corresponding correlations. These comparisons showed that depression-specific associations were significantly more negative than anxiety-specific associations for both mood variation (laboratory dataset: Z = -1.84, p = 0.033; online dataset 1: Z = -5.36, p < 0.001; online dataset 2: Z = -3.42, p < 0.001) and β_RPE (laboratory dataset: Z = -1.77, p = 0.038; online dataset 1: Z = -3.67, p < 0.001; online dataset 2: Z = -4.00, p < 0.001; Figures 2A–C and 2E–G). These results support distinct associations of depression- and anxiety-specific factors with RPE-related mood dynamics.”

      In the discussion, the authors could have mentioned the brain areas most likely to be involved in these processes, both cognitive and psychopathological, as previous studies (such as Cecchi et al, 2022) have aimed at identifying regions involved in RPE processing while modulating mood in health. A short section on this would be useful to the neuropsychiatric community.

      We thank the Reviewer for this helpful suggestion. We agree that the Discussion would benefit from a more explicit consideration of the neural systems that may support RPE-related mood updating and their relevance to psychopathology. We have revised the Discussion accordingly, as shown below:

      Page 16:

      “Although the present study did not include neuroimaging, the observed computational dissociation may map onto partially distinct neural systems involved in reward learning, mood updating, and affective psychopathology. RPE processing has been consistently linked to striatal–midbrain dopaminergic reward-learning circuits[8,50]. The integration of these reward-learning signals into subjective mood and value-based decision-making may further involve the ventral medial prefrontal cortex and orbitofrontal cortex[44]. In addition, the anterior insula may be particularly relevant for integrating feedback-related signals with affective and interoceptive states[8,44], potentially linking RPE processing to anxiety- and depression-related mood dynamics. Consistent with this view, Cecchi et al. (2022)[51] used intracranial EEG to show that feedback-related neural activity tracks mood fluctuations and risky choice. Future neuroimaging studies should test whether depression-related reductions and anxiety-related increases in RPE-related mood sensitivity are associated with altered interactions among striatal, prefrontal, and insular circuits.”

      Reviewer #3 (Recommendations for the authors):

      (1) The authors need to present a clearer definition of the terms "bifactor analysis" and "tripartite model" in the Introduction. What does tripartite mean in this context? What are the assumptions of such bifactor analyses (e.g. as used in Gagne et al. and the present work) and how, in broad strokes, are they carried out? These are important constructs to clarify for readers outside the computational psychiatry niche.

      We thank the Reviewer for this helpful suggestion. We have revised the Introduction accordingly, as shown below:

      Page 3:

      “Recent work has used bifactor models of the tripartite model of depression and anxiety to clarify their distinct features and differential influences on decision-making[31,32]. The tripartite model of anxiety and depression proposes that these two symptom dimensions share a broad general distress or negative affect component while also including symptom-specific components: low positive affect/anhedonia is more specific to depression, whereas physiological hyperarousal is more specific to anxiety[30,33,34]. Bifactor analysis offers a way to model this structure statistically. In a bifactor model, symptoms load on a general factor reflecting their shared variance and on specific factors capturing residual variance in narrower symptom dimensions after accounting for the general factor.”

      (2) Previous examinations of depression and RPE reactivity in this task paradigm, as the authors note (e.g. Rutledge et al., 2017), observed that individuals diagnosed with depression showed an intact association between RPEs and mood. In other words, there was no previously observed relationship between depression and affective reactivity to RPEs in this task context. Here, the authors find that the "unique" depression factor identified by the authors (in a non-clinical sample) is associated with blunted RPE sensitivity - this is worth commenting on specifically.

      We thank the Reviewer for this helpful suggestion. We have discussed this point in the Discussion. Please also see it below:

      Pages 15-16:

      “Our computational model not only replicates the important role of RPEs in mood dynamics but also highlights the divergent mediating roles of RPE-related mood sensitivity in the associations of depression and anxiety with mood fluctuations. The opposite associations of depression and anxiety with mood sensitivity to RPEs complement previous findings of apparently intact RPE-related mood sensitivity in depression[12,25,37]. These findings further underscore the necessity of decomposing shared and specific components of depression and anxiety in studies of mood dynamics, which can enhance our understanding of their distinct associations with emotion processing and cognitive flexibility. This point is consistent with bifactor-based work showing that shared and specific symptom dimensions can have different computational correlates. For example, Gagne et al. (2020) showed that bifactor-derived symptom dimensions differentially relate to maladaptation to environmental volatility[32], complementing previous findings that trait anxiety is associated with inflexible adjustment to volatility[32].”

      (3) There is a note (line 247) about the interpretation of the correlations in Figure 2, which attempts to explain away the inconsistent relationships between the anxiety-specific factor and mood variability as well as RPE reactivity observed in Figure 2. I can't say I understand the authors' point here about "signs of positive correlations", so I would say the authors need to clarify their logic here. More to the point, the authors only resolve anxiety-specific predictive effects when they collapse across these datasets. As discussed above (see 'weaknesses'), this is a serious limitation in my view and needs to be discussed as such in the paper.

      We apologize for the unclear statement. Although the anxiety-specific factor showed associations in the expected direction across datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level.

      Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). We have revised it to make it clear. Please see our revisions below:

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β_RPE were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect, we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis while accounting for dataset-level variability. Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis while accounting for dataset-level variability. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001).”

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      (4) I expected to see that the authors would also investigate relationships between anxiety/depression related factors and the decay (gamma) parameter in the 'Happiness equation', which is presumably estimated from the data here. While I don't have a strong intuition about directions of (or presence of) predictive relationships here, doesn't it stand to reason that different aspects of psychopathology examined here might map onto how long- (versus short-) lasting the effects of, say, RPEs are, upon mood?

      We thank the Reviewer for this important comment. In the healthy datasets, we fitted a linear mixed-effects model predicting the decay parameter (γ) from the three bifactor scores, with dataset included as a random intercept. None of the factors showed a significant association with γ (common: t = 0.708, p = 0.479; anxiety: t = 0.564, p = 0.573; depression: t = 1.146, p = 0.252). In the clinical dataset, we fitted a linear model predicting γ from the three bifactor scores and again found no significant associations (common: t = -0.036, p = 0.972; anxiety: t = -0.047, p = 0.963; depression: t = 0.052, p = 0.959). We have clarified this point in the revised manuscript as follows:

      Supplementary Page 5:

      “In the healthy datasets, we conducted a linear mixed-effect model against decay parameter (gamma) with all three factors, with dataset as a random factor. Results did not show significant effect (common: t = 0.708, p = 0.479; anxiety: t = 0.564, p = 0.573; depression: t = 1.146, p = 0.252). In the clinical dataset, we conducted a linear model against decay parameter (gamma) with all three factors and found no significant effect (common: t = -0.036, p = 0.972; anxiety: t = -0.047, p = 0.963; depression: t = 0.052, p = 0.959).”

      (5) The rationale for and interpretation of the mediation model, which presumably aims to explain the relationships between anxiety- and depression-specific factors, RPE reactivity, and mood variability was barely explained by the authors. At present, I'm not sure what the added value of this analysis is. The authors should either remove or explain/motivate the mediation more clearly.

      We thank the Reviewer for this helpful comment. The rationale for the mediation analysis is that mood variability in the task is not only a descriptive behavioral outcome, but may also arise from the degree to which momentary mood is updated by RPEs. Therefore, if depression is associated with reduced RPE-related mood sensitivity and anxiety with increased RPE-related mood sensitivity, these alterations should statistically account for their opposite associations with mood variability. The mediation model directly tested this possibility by examining whether RPE-related mood sensitivity accounted for the association between symptom-specific factors and mood variability. We have clarified this rationale in the revised manuscript as follows:

      Page 10:

      “Given the strong correlation between β<sub>RPE</sub> and mood variation (rs > 0.67, ps < 0.001), we further conducted a mediation analysis to examine whether individual differences in RPE-related mood sensitivity statistically accounted for the association between depression loading and mood variation. This analysis was motivated by the hypothesis that depression-related dampening of mood variability may arise, at least in part, from reduced mood sensitivity to RPEs.”

      (6) This submission would benefit from extensive English language copy editing. There are many passages in the paper (in fact, too many to list here) that suffer from either grammatical errors or clarity issues.

      We apologize for these mistakes. We have corrected the typographical errors throughout the manuscript.

    1. eLife Assessment

      This important study reports that Sox17 is key to the formation and function of the Sertoli valve, a transition region between the rete testis and seminiferous tubules that remains an understudied domain of testicular biology. The supporting data are convincing. This work will be of interest to developmental and reproductive biologists, as well as andrologists who work on male fertility and men's health.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript is an excellent follow-up to your 2022 study, in which Sox17 expression was localized to the rete testis and shown to be required for proper formation of the Sertoli cell valve (transition region). By using Nr5a1-Cre to drive conditional deletion of Sox17 specifically in rete testis cells, you demonstrate that testis weights remain normal at 2 weeks of age but become significantly reduced by 8 weeks in Sox17-cKO males. At the later time point, the seminiferous epithelium is severely disrupted, with apparent arrest of spermiogenesis: the epididymal lumen is essentially devoid of sperm, and most tubules lack elongated spermatids.

      Strengths:

      Clearly shows the role of Sox17 in Sertoli cells being important to the SV function. The SV (transition region) between the rete testis and seminiferous tubules remains an understudied domain of testicular biology. The present work, together with your prior study, highlights intriguing mechanisms operating in this specialized niche.

      Weaknesses:

      The available data do not fully explain either the developmental assembly of the Sertoli valve or the precise consequences of its functional disruption. These studies are nonetheless valuable precisely because they raise more questions than they answer; the conceptual implications are thought-provoking.

    3. Reviewer #2 (Public review):

      This manuscript investigates the role of SOX17 in the formation and function of the Sertoli valve (SV) at the interface between seminiferous tubules and the rete testis (RT). Building on previous work showing that rete testis-specific deletion of Sox17 disrupts SV formation, leading to defective spermiogenesis and male infertility, the authors explore how SOX17 overexpression in Sertoli cells regulate SV of rodent testes.

      Using transgenic mouse models with ectopic Sox17 expression in Sertoli cells, the study demonstrates that SOX17 is not only required but can also modulate SV formation. Ectopic expression in Sertoli cells induces expansion of the SV structure and partially rescues SV defects and spermatogenesis in RT-specific Sox17 conditional knockout animals. The data support a model in which SOX17 acts through paracrine signaling to regulate SV formation, although the precise mechanisms remain to be clarified.

      Overall, this is a well-executed study with novel and significant findings. The ability to experimentally manipulate SV size is particularly compelling and provides a valuable framework to study fluid dynamics and epithelial interactions in the testis. This work will be of broad interest to the reproductive biology and developmental biology communities.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript is an excellent follow-up to your 2022 study, in which Sox17 expression was localized to the rete testis and shown to be required for proper formation of the Sertoli cell valve (transition region). By using Nr5a1-Cre to drive conditional deletion of Sox17 specifically in rete testis cells, you demonstrate that testis weights remain normal at 2 weeks of age but become significantly reduced by 8 weeks in Sox17-cKO males. At the later time point, the seminiferous epithelium is severely disrupted, with apparent arrest of spermiogenesis: the epididymal lumen is essentially devoid of sperm, and most tubules lack elongated spermatids.

      Strengths:

      The study clearly shows the role of Sox17 in Sertoli cells as being important to SV function. The SV (transition region) between the rete testis and seminiferous tubules remains an understudied domain of testicular biology. The present work, together with the authors' prior study, highlights intriguing mechanisms operating in this specialized niche.

      Weaknesses:

      At the same time, the available data do not yet fully explain either the developmental assembly of the Sertoli valve or the precise consequences of its functional disruption. These studies are nonetheless valuable precisely because they raise more questions than they answer; the conceptual implications are thought-provoking.

      Reviewer #2 (Public review):

      This manuscript investigates the role of SOX17 in the formation and function of the Sertoli valve (SV) at the interface between seminiferous tubules and the rete testis (RT). Building on previous work showing that rete testis-specific deletion of Sox17 disrupts SV formation, leading to defective spermiogenesis and male infertility, the authors explore how SOX17 overexpression in Sertoli cells regulates the SV of rodent testes.

      Using transgenic mouse models with ectopic Sox17 expression in Sertoli cells, the study demonstrates that SOX17 is not only required but can also modulate SV formation. Ectopic expression in Sertoli cells induces expansion of the SV structure and partially rescues SV defects and spermatogenesis in RT-specific Sox17 conditional knockout animals. The data support a model in which SOX17 acts through paracrine signaling to regulate SV formation, although the precise mechanisms remain to be clarified.

      Overall, this is a well-executed study with novel and significant findings. The ability to experimentally manipulate SV size is particularly compelling and provides a valuable framework to study fluid dynamics and epithelial interactions in the testis. This work will be of broad interest to the reproductive biology and developmental biology communities.

      Reviewer #3 (Public review):

      Summary:

      These studies are based on previously published work that showed that deletion of expression of the Sox17 gene in the testis essentially deleted the formation of the Sertoli valve in the Rete testis. The authors extended this work by constructing a vector that resulted in increased Sox17 expression by Sertoli cells and enhanced formation of the Sertoli valve in both wild type and Sox17 knockout mice. The work provides strong evidence supporting the requirement for Sox17 expression to allow formation of the Sertoli valve.

      Strengths:

      The general approach was to express Sox17 from a Tg mouse that expressed Sox17 from Sertoli cells. This Tg mouse was bred into both the WT and the Sox17 KO mouse. The Sertoli valve was enhanced in both the WT/Tg mouse and KO/Tg mouse, showing that ectopic Sox17 could compensate in the Sox17 Ko and act in a concentration-dependent manner in the WT mouse. The results are strong and support the conclusions from the authors. The results were as expected from the original paper describing the KO of Sox 17. These results strengthen these conclusions and provide ideas for additional conclusions. These studies were technically challenging, and the authors provided a very solid manuscript.

      Weaknesses:

      The authors refer several times to high or low expression, but it all appears to be based on immunohistochemistry, and there is no real quantification using PCR, for example. The process used for cell quantification lacks a rationale for why certain numbers were assigned.

      We sincerely thank the reviewers for their careful evaluation of our manuscript and for their constructive and encouraging comments. We are grateful for the recognition of the significance of the Sertoli valve as an understudied transition region between the rete testis and seminiferous tubules, as well as for the positive assessment of our genetic approach and the evidence that ectopic SOX17 expression can modulate SV formation. We have carefully considered all points raised in the assessment and have revised the manuscript accordingly. The major revisions include:

      (1) Clarification of the scope and limitations of the study (Reviewers #1 and #2):

      In response to the comments that the developmental assembly of the Sertoli valve and the precise consequences of its functional disruption remain incompletely understood, we clarified the scope and limitations of the present study at the end of 7th paragraph in the Discussion. Although our findings support a model in which SOX17 regulates SV formation through paracrine signaling, the downstream effectors and precise molecular mechanisms remain to be identified. We therefore revised the Discussion to avoid overinterpretation of the molecular mechanisms and to emphasize that comprehensive mechanistic analyses, including transcriptomic analyses using the Tg mouse model, represent an important direction for future research. We also added histological analyses of the earliest detectable lesions at 4 weeks of age and low-magnification images of adult Sox17 cKO testes (new Figure S1), revealing selective sloughing of round spermatids despite preserved Sertoli cell architecture and subsequent mosaic spermatogenic defects among individual seminiferous tubules. These observations provide additional insights into the altered luminal microenvironment and suggest that spermatogenic defects may progress in a tubule-by-tubule manner.

      (2) Clarification of quantitative analysis and methodology (Reviewer #3):

      In response to concerns regarding the basis and methodology of cell quantification, we revised the Methods to provide detailed information on tissue preparation, fixation, orientation of the rete testis–Sertoli valve region, and the criteria used for quantitative analysis of SV-associated Sertoli cells (new Figure S4). We clarified that Sertoli cells were counted within the SV region extending approximately 100 μm from the RT boundary, including Sertoli cells protruding into the RT lumen, based on previously established criteria (Aiyama et al., 2015).

      (3) Clarification of the limitations of expression-level assessment (Reviewer #3):

      In response to concerns regarding the quantitative assessment of SOX17 and other SV-associated molecules, we clarified the technical limitations of selectively isolating the very small SV region and obtaining sufficient material for quantitative molecular analyses such as qPCR at the end of 7th paragraph in the Discussion. We therefore clarified that expression of SV-associated molecules in the present study was primarily evaluated using histological and immunohistochemical approaches and added the relevant text to acknowledge these limitations.

      We also made additional revisions to clarify each mouse Tg line, phenotypic descriptions, standardize gene nomenclature, improve methodological descriptions, and refine the relevant Discussion where appropriate.

      We sincerely appreciate the reviewers’ thoughtful and constructive comments. Their feedback has helped us clarify the scope of our conclusions, strengthen the methodological descriptions, and improve the overall presentation of the study. All changes have been incorporated into the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (i) Although the current paper is not responsible for interpreting the 2022 findings, both datasets show reduced spermatid production accompanied by multinucleated giant germ-cell syncytia. This phenotype has been attributed to backflow of tubular fluid and consequent microenvironmental perturbation. While this is a reasonable hypothesis, it is not entirely consistent with earlier experimental observations. Complete ligation of the efferent ductules reliably produces giant cells, whereas estrogen-receptor knockout, which also causes massive luminal fluid accumulation, does not. In addition, ligation of the testicular artery itself can induce giant-cell formation. Although this may have already been answered in the papers, can you be sure that a direct or indirect effect on the vasculature can be excluded in the Sox17-cKO model?

      We thank the reviewer for this important comment. In our models, SOX17 expression was manipulated specifically in the Sertoli cell lineage, either by SF1-Cre-mediated Sox17 deletion or by ectopic SOX17 expression under the hAMH-promoter. SOX17-expressing vascular endothelial cells were not targeted in either model, making a direct effect of Sox17 manipulation on the testicular vasculature unlikely. Moreover, the partial rescue of the Sox17 cKO phenotype by hAMH-Sox17 supports the interpretation that the phenotype primarily results from altered SOX17 function in Sertoli cells and RT epithelia.

      However, indirect effects on the vascular or interstitial environment by aberrant luminal flow cannot be completely excluded, particularly with the substantial accumulation of sloughed round spermatids (giant cells) within the rete testis. Addressing the potential for an initial luminal flow defect, we newly added histological images of 4-week-old testes (Figure S1), where selective post-meiotic germ cell sloughing occurs despite preserved Sertoli cell process architecture, suggesting an altered adluminal microenvironment that impairs Sertoli–spermatid adhesion. Furthermore, low-magnification images of adult mature Sox17 cKO testes (Figure S1B) display a mosaic pattern of spermatogenic defects across individual tubules. While 3D reconstruction was not conducted, this structural pattern supports the view that spermatogenic failure progresses on a tubule-by-tubule basis, potentially linked to the structural integrity of individual Sertoli valves.

      (ii) A related and important unresolved issue is the total number of Sertoli cells per testis in cKO males. The number of Sertoli cells per tubule cross-section is reported to be equivalent to controls; however, the substantial reduction in testis weight implies a corresponding reduction in tubule length. Under these conditions, maintenance of a normal per-cross-section count would still be compatible with an overall decrease in total Sertoli-cell number. Although it is generally accepted that murine Sertoli cells exit the cell cycle around postnatal day 15, continued growth of the testis may still occur in the Sertoli valve region, where Sertoli cells retain proliferative capacity. Your discussion of possible heterogeneity in the embryonic origin of Sertoli cells near the rete testis is therefore particularly intriguing and commendable. Should this hypothesis be substantiated, it would raise the possibility that Sertoli cells derived from the valve region, especially those that migrate into the seminiferous tubules, are intrinsically less competent to support full spermatogenesis than those of classic gonadal-ridge origin.

      To help readers appreciate the overall severity and topographic distribution of the spermatogenic defect (particularly in tubule segments distant from the rete), inclusion of a low-magnification photomicrograph of a well-fixed (Bouin's) testicular cross-section would be very useful.

      We thank the reviewer for this important comment. We agree that the maintenance of Sertoli cell numbers per seminiferous tubule cross-section does not necessarily indicate preservation of the total Sertoli cell number per testis, particularly given the substantial reduction in testis size and potential reduction in overall seminiferous tubule length. Although total Sertoli cell numbers can theoretically be estimated using stereological approaches, such analyses are technically demanding and beyond the scope of the present study.

      We also appreciate the reviewer’s insightful suggestion regarding potential heterogeneity among Sertoli cell populations. Sertoli cells associated with the Sertoli valve region may have distinct developmental origins or functional properties compared with classical gonadal ridge-derived Sertoli cells, which could potentially influence their capacity to support complete spermatogenesis. Although this hypothesis was not directly tested in this study, we have expanded the Discussion to highlight the developmental and functional heterogeneity of Sertoli cell populations associated with the Sertoli valve as an important topic for future investigation.

      In addition, as requested, we have added a low-magnification image of well-preserved testicular cross-sections in Supplementary Figure S1B to better illustrate the overall severity and topographic distribution of spermatogenic defects throughout the testis.

      Specific Comments:

      (1) Figure 3A and associated fertility/histology data. The results state that epididymal spermatozoa were detected in only 2 of 7 cKO;Tg males at 8 weeks of age, yet Materials and Methods indicate that spermatogenesis was evaluated in only 5 males. a) Were the remaining two males also examined histologically? b) It would be interesting to determine if the severity of pathological changes was the same in regions more distant from the rete testis, or possibly different tubules. See: Nakata H, Wakayama T, Sonomura T, Honma S, Hatta T and Iseki S (2015). "Three-dimensional structure of seminiferous tubules in the adult mouse." J Anat 227(5): 686-694. c) In addition, mating trials were performed with four independent cKO;Tg males, two of which sired offspring. It is unclear whether the testes of these four mating males were included among the five (or seven) animals evaluated for histology, and whether the two fertile males correspond exactly to the two individuals that retained epididymal sperm. Please clarify these relationships explicitly so that readers can correctly interpret the link between histological findings and fertility.

      We thank the reviewer for this important comment. We apologize that the relationship among the groups of animals used for histological analysis, epididymal sperm detection, and fertility assessment was not sufficiently clear in the original manuscript. Because this study focused specifically on the anatomically minute RT–SV region, our sampling strategy had to prioritize the maximal utilization of this limited tissue. In this study, the RT–SV region, the remaining testicular tissue, and the epididymis were processed separately as three tissue blocks for each animal (Figure S4) and were independently evaluated for distinct analysis sets. Briefly, the proximal quarter containing the rete testis and Sertoli valve region was used for SV analysis, whereas the remaining three-quarters of the testis were used for evaluation of spermatogenesis, and the epididymis was analyzed separately for the presence of spermatozoa. Therefore, due to these technical requirements, tissue allocation, and independent analytical evaluation, the numbers of animals used for RT–SV analysis, testicular histology, epididymal sperm detection, and fertility testing were not identical.

      For quantitative histological analyses, we also used virgin males to minimize potential variation associated with mating experience and to allow comparison with age-matched littermate controls. Therefore, these animals were not used for fertility testing. Fertility assessment was performed using an independent cohort of cKO; Tg males that were subjected to long-term mating trials with wild-type females. Thus, fertility outcomes and histological findings were not designed to be directly matched at the individual level.

      In response to the reviewer’s suggestion, we have clarified the selection of experimental animals and the relationship among fertility assessment and histological analyses in the Materials and Methods and added a schematic illustration of the sampling strategy in Figure S4. We also corrected the citation for Nakata H et al., 2015 in the revised manuscript.

      (2) Page 8, line 301 (Sertoli-cell quantification). The description of the counting method-"counted in each ... (~100 μm from the edge of the RT; Fig. 4C)"-is ambiguous.

      (a) Does this mean that cells were counted beginning at the rete boundary and extending radially outward for approximately 100 μm, or is a circumferential sampling area intended? (b) Figure 4C shows a large standard deviation, indicating substantial variability with the current approach. An alternative strategy (for example, counting Sertoli cells within standardized areas or per tubule specifically within the valve region) might reduce variability and improve reproducibility. Regardless of the method ultimately chosen, a more precise, step-by-step description of the quantification protocol is required so that it can be reliably replicated by other laboratories.

      We thank the reviewer for pointing out that the description of the Sertoli cell quantification method was not sufficiently clear. The Sertoli cell quantification was performed using the same criteria as previously described (Aiyama et al., 2015; Uchida et al., 2022), in which SOX9-positive Sertoli cell nuclei within the SV-associated region were counted.

      In the revised manuscript, we have clarified that the SV region was operationally defined as comprising (i) the terminal 100 μm segment of the seminiferous tubule immediately adjacent to the rete testis (RT) and (ii) the protruded SV extending into the RT lumen. Based on the distribution of spermatogonial stem cells, the ~100 μm region extending from the RT boundary along the seminiferous tubule toward the ST side was defined as the SV region (Aiyama et al., 2015). Only sagittal sections showing a continuous RT–SV–ST axis and sectioning the SV approximately through its mid-sagittal plane were included for quantitative analysis.

      Furthermore, to improve reproducibility, we have added a more detailed description of the tissue preparation and quantification procedures in the Materials and Methods and provided a schematic illustration of the quantification strategy in the new Figure S4.

      Reviewer #2 (Recommendations for the authors):

      (1) Phenotypic differences between transgenic lines: the phenotypic differences between the tg26 and tg27 lines are intriguing and warrant further clarification. While tg27 mice exhibit infertility and defective spermatogenesis, tg26 animals remain fertile with SV expansion. Could the authors elaborate on the underlying causes of these differences? In particular, is infertility in tg27 mice due to excessive SOX17 expression impairing Sertoli cell function? A comparison of Sox17 expression levels between tg26 and tg27 lines would be informative. In addition, it would be useful to assess whether acetylated tubulin (Ac-Tub) expression is present in the Sertoli cells of the tg27 mouse testis.

      We thank the reviewer for this highly constructive and insightful comment. We clarified in the revised manuscript that the analysis of the Tg27 mouse was performed using the F0 founder male and added an explanation that only the Tg26 line could be established as its heterogenous SOX17 expression in Sertoli cells did not impair overall fertility. We agree that the phenotypic differences between the Tg26 line and the Tg27 mouse provide important clues regarding the dosage-dependent effects of SOX17 in Sertoli cells. Unfortunately, we were unable to establish a stable, multi-generational transgenic line from this Tg27 founder (F0) male. Consequently, we could not perform detailed molecular or immunohistochemical analyses on this line beyond the initial histological evaluation of the F0 generation presented in Figure 1. For this reason, we cannot provide a quantitative comparison of Sox17 expression levels or evaluate acetylated tubulin (Ac-Tub) expression in Tg27 Sertoli cells.

      To address the reviewer's concern without overstepping the available data, we removed direct quantitative comparisons of Sox17 expression levels between the two lines from the text. Instead, we added a clear description of their contrasting cellular expression patterns - specifically, the mosaic, heterogeneous SOX17 expression in Tg26 Sertoli cells versus the ectopic, uniform SOX17 expression in the infertile #27 F0 male - in the 'Animals' section of Materials and Methods. This mosaic pattern in Tg26 testes suggests the presence of Sertoli cells with low or undetectable SOX17 levels, which may be associated with sustaining overall fertility.

      (2) Mechanism of SOX17 action: although SOX17 is a transcription factor, the author's studies indicate it regulates SV formation via paracrine and/or autocrine signaling. The underlying mechanisms remain unclear. Which downstream factors mediate this effect? The observed upregulation of RSPO1 and WNT4 is suggestive, but more direct evidence would strengthen this conclusion. For example, does SV expansion in tg26 mice depend on the activation of RSPO1/WNT signaling? Additional molecular analyses, such as bulk RNA-seq comparing control and transgenic testes, could help identify pathways regulated by SOX17 and clarify its mode of action.

      We thank the reviewer for this important and insightful suggestion. At present, comprehensive analyses, including scRNA-seq of Sox17 cKO and littermate control testes, have not identified definitive downstream targets of SOX17 (Uchida et al., 2022). As the reviewer rightly points out, the Tg26 mouse model generated in this study represents a valuable tool for investigating SOX17-dependent molecular pathways. To this end, we are currently conducting transcriptomic analyses of Tg26 seminiferous tubules to identify genes altered in SOX17+ Sertoli cells. However, determining whether these candidate genes represent direct transcriptional targets of SOX17 and whether they function specifically in the rete testis-associated region during Sertoli valve formation will require extensive functional and expression studies. Therefore, we feel it would be premature to draw definitive conclusions regarding the underlying molecular mechanisms, including the precise involvement of the RSPO1/WNT signaling pathway, in the present manuscript. Accordingly, rather than overinterpreting the available data, we have revised the Discussion to clarify this limitation (at the end of 7th paragraph in the Discussion). Furthermore, incorporating initial insights from our ongoing Tg26 transcriptomic analyses, we have added a brief discussion, supported by relevant literature, on the possibility that SOX17 may regulate Sertoli valve formation by modulating cell adhesion and extracellular matrix (ECM) organization and altering the responsiveness of SOX17-positive Sertoli cells to morphogenetic signals originating from the rete testis (new 5th paragraph in Discussion).

      (3) Minor comment: Gene nomenclature should be standardized: e.g. line 245, Sox17 and hAMH should be italicized.

      We thank the reviewer for pointing this out. All gene names have been italicized throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      No suggestions except to quantify some of the changes in concentration of agents by PCR rather than eyeball levels with immunocytochemistry. Verify the cell quantification procedure used.

      We thank the reviewer for this comment. The Sertoli valve (SV) is an extremely small transitional structure, with only approximately 20 sites per mouse testis. As a result, selective isolation of the SV region to collect sufficient material for molecular analyses, such as quantitative PCR, remains technically challenging. We have therefore added this limitation to the Discussion.

      Regarding the cell quantification procedure, we have clarified the methodology in the revised Materials and Methods and added a schematic illustration in Figure S4. Specifically, the Sertoli cell number in the SV region was quantified by counting SOX9-positive Sertoli cell nuclei within a standardized SV-associated region in RT–SV–ST sagittal sections.

  2. Sep 2026
    1. eLife Assessment

      The results of this study are important and the approach to dissect the developmental contribution of RIF1 function is convincing. The finding that replication timing and gene expression may be independently controlled is intriguing and provides a strong foundation for future research. The work will be of interest for researchers both in the transcription and the replication field, especially for scientists investigating the interplay between the two processes.

    2. Reviewer #1 (Public review):

      The authors sought to determine how Rif1 contributes to DNA replication timing (RT), transcriptional regulation, and embryonic development using zebrafish. They generated a maternal-zygotic rif1 knockout line and examined developmental phenotypes, genome-wide replication timing profiles, RNA-seq, and nascent transcription (SLAM-seq) during early embryogenesis.

      Their major findings in this manuscript are

      (1) Rif1 is not essential for zebrafish viability, unlike its partially essential role in mice.

      (2) Rif1 deficiency causes defects in female sex determination, delayed epiboly, and reduced primitive erythropoiesis.

      (3) Genome-wide RT is altered by Rif1, but developmental stage has a much larger influence than Rif1 itself.

      (4) Rif1 is required for the proper maturation ("sharpening") of the RT program during development rather than for specific developmental RT switches.

      (5) Rif1 has a much stronger effect on transcription during zygotic genome activation (ZGA) than on replication timing at these early stages.

      (6) Loss of Rif1 leads to increased expression of early zygotic genes, indicating that Rif1 normally suppresses widespread transcription during ZGA.

      Overall, the work proposes that Rif1 independently regulates replication timing and transcription, with these two functions becoming most prominent at different developmental stages.

      The major strengths of the manuscript are as follows.

      (1) the study combines multiple genome-wide approaches including whole-genome RT profiling, RNA-seq, SLAM-seq in combination with gene KO and developmental analyses.

      (2) One of the strongest points is that the authors conducted the analyses at multiple developmental stages rather than a single point.

      (3) The most important conclusion is that the Rif1 regulates transcription during development in a manner largely independent of its RT function, which was further strengthened by the additional data provided in the revised manuscript.

      On the other hand, the weakness of the manuscript includes the followings.

      (1) Limited mechanistic insight. The questions such as where Rif1 binds on the chromatin (in relation to the transcriptional promoters/ enhancers and replication origins).

      (2) Which functional domains of RIf1 are involved in regulation of transcription and replication (Is PP1 recruitment required for transcription regulation?) are not addressed.

      (3) Since Rif1 is known to be involved in chromatin organization/ nuclear architecture regulation, the studies addressing this (Hi-C, compartment analyses, ATAC seq etc) would provide important mechanistic information.

      (4) Female sex determination phenotype is intriguing, but it remains largely descriptive, and its mechanisms are elusive at the moment.

      Overall, the results support the authors' conclusions and they have successfully provided answers to the authors' original questions on developmental roles of Rif1 in RT and transcription in vertebrate.

      Comments on revised version:

      The authors responded to my comments in a largely satisfactory manner. They have conducted additional analyses and concluded that Rif1 regulates transcription during ZGA largely independently of its classical RT function, which is an important finding.

      Although authors did not examine origin firing and replication fork rate in rif1 KO cells, which I suggested in my original review, this can be saved for their future studies.

      I think the revised manuscript has been improved and provides important basic information on the functions of the conserved Rif1 protein in RT and transcriptional regulation.

      I have no further recommendation for additional experiments or data analyses.

    3. Reviewer #2 (Public review):

      This study by Masser et al. analyzes global replication timing and gene expression in rif-1 null zebrafish. This work is an extension of their previous report of the normal replication timing pattern during wild-type zebrafish development. The major valuable finding here is that Rif1 is not essential for viability in zebrafish, and - counter to expectation from studies in cultured cells and other species - late replication does not strongly depend on Rif1. Instead, the data suggest that Rif1 subtly sharpens replication timing pattern during normal development rather than function generally to delay replication timing. In the absence of Rif1, the normal pattern establishment is somewhat delayed. The authors also document some changes in expression during development with more genes being repressed by Rif1 than activated at some early stages.

      The study and analysis are generally rigorous, and the conclusions are supported by convincing data. Given the strong link between replication timing and cell type/development, studying timing in a whole developing organism is important. The experimental approach is technically challenging, particularly the bioinformatic analysis. The scientific advance here is largely confined to documenting the timing of Rif1-affected transcription, the unanticipated effect of the rif1 deletion on replication timing and on sex determination, though the latter is not explored. The difference in timing of the transcription phenotypes and replication phenotypes suggests they may be very distinct Rif1 roles. The overall study a useful set of findings and detailed data for future work.

      Loss of Rif1 did not affect viability, but it did strongly influence sex determination, resulting in a lower population of females. This effect is the strongest organismal phenotype, but the study provides no mechanistic explanation for the loss of females from the data gathered here.

      Comments on revised version:

      We are generally satisfied with the revised version of this manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In this manuscript authors examined the effect of rif1 knockout on replication timing and transcription in early embryos of zebrafish. Contrary to the expectation, genome-wide replication timing domains did not significantly change upon Rif1 knockout, although the replication timing became less dynamic in the mutant, meaning the entire genomes are replicated toward the mid S. In contrast, transcriptional profiles change by rif1 mutation throughout the embryo stage. These effects were more predominantly observed after gastrulation at the early stages of zebrafish development.

      The results presented in this manuscript provide new information on the effects of rif1 mutation on early zebrafish development, although the underlying mechanism has not been explored. The information is useful for researchers in the field of early development, with specific focus on replication and transcription regulation.

      The genome wide analyses of replication timing has been conducted and analyzed properly. The transcriptional analyses are conducted by RNA-seq and SLAM-seq (determining the nascent mRNA), and the results convincingly show the overall transcriptional patterns at different developmental stages.

      This work shows that Rif1 regulates replication timing and transcription in zebrafish embryos, while the extents of the effects vary during the developmental process. Although the data convincingly illustrate the whole picture of Rif1 KO on replication and transcription during zebrafish development, the mechanistic insight is missing. Especially, how Rif1 may or may not coordinately regulate replication and transcription during the zebrafish development has not been addressed.

      We thank the reviewer for recognizing the value of combining genome-wide replication-timing, RNA-seq, and SLAM-seq analyses across zebrafish development. We agree that the original study did not establish a molecular mechanism linking Rif1-dependent transcriptional and replication-timing effects. To address whether these effects are locally coordinated, we added a gene-centred analysis comparing replication-timing values for genes with increased, decreased, or unchanged transcript abundance at Dome (Figure 5--figure supplement 2). Differentially expressed genes did not show a clear enrichment in early- or late-replicating regions, either at Dome or at pre-MBT. These results argue against replication timing state being the primary determinant of the Dome-stage transcriptional changes. We also expanded the Discussion to explain the limitations of the current study and the need for future measurements of origin use, fork progression, chromatin state, and cell-type-specific effects. The new discussion of Nakatani et al. (2025) further places our findings in the context of evidence that Rif1-dependent replication-timing changes can be uncoupled from transcriptional changes.

      Reviewer #2 (Public Review):

      This study by Masser et al. analyzes global replication timing and gene expression in rif-1 null zebrafish. This work is an extension of their previous report on the normal replication timing pattern during wild-type zebrafish development. The major valuable finding here is that Rif1 is not essential for viability in zebrafish, and - counter to expectation from studies in cultured cells and other species - late replication does not strongly depend on Rif1. Instead, the data suggest that Rif1 subtly sharpens replication timing pattern during normal development rather than function generally to delay replication timing. In the absence of Rif1, the normal pattern establishment is somewhat delayed. The authors also document some changes in expression during development with more genes being repressed by Rif1 than activated at some early stages.

      The study and analysis are generally rigorous, and the conclusions are supported by convincing data. The manuscript is well written, though there are aspects of the presentation that could be improved for a broader scientific audience. Given the strong link between replication timing and cell type/development, studying timing in a whole developing organism is important. The experimental approach is technically challenging, particularly the bioinformatic analysis. The scientific advance here is largely confined to documenting the timing of Rif1-affected transcription, the unanticipated effect of the rif1 deletion on replication timing and on sex determination, though the latter is not explored. The work is descriptive and feels like two relatively unconnected studies, transcription and replication plus a small bit of development, and the difference in timing of the transcription phenotypes and replication phenotypes suggests they may be very distinct Rif1 roles. There isn't a lot of new insight into the mechanism of how Rif1 affects either replication timing or gene expression. As such, the overall study is an useful set of findings and detailed data for future work, but it doesn't make a big step forward in understanding the role of Rif1 or the biological processes it affects.

      Weaknesses worth addressing include the following:

      (1) Loss of Rif1 did not affect viability, but it did strongly influence sex determination, resulting in a lower population of females. This effect is the strongest organismal phenotype, but the study provides no explanation for the loss of females from the data gathered here.

      (2) The approach to distinguish nascent zygotically expressed mRNAs from maternal mRNAs is a strength. Are the differentially expressed genes related at all to regions of the genome whose replication timing is most affected? Are any of them related to the sex determination or developmental phenotypes?

      We thank the reviewer for recognizing the rigor of the analyses and the value of studying replication timing in a developing vertebrate. We revised the manuscript extensively to make the experimental logic, zebrafish developmental context, replication-timing analyses, and figure legends more accessible to a broad audience. We also quantified the gastrulation phenotype, showing an approximately one-hour delay in completion of epiboly in maternal-zygotic rif1 mutants rather than a persistent developmental arrest.

      We agree that the mechanism underlying the sex-ratio phenotype remains unresolved. The transcriptomic experiments were performed in whole embryos at stages much earlier than zebrafish sex determination and therefore cannot resolve changes in primordial germ cells or supporting gonadal somatic cells. We have avoided making a mechanistic connection between the early embryonic transcriptional changes and the adult sex-ratio phenotype and identify this as an important area for future study. To address the relationship between transcription and replication timing, we added Figure 5--figure supplement 2. Genes with increased or decreased transcript abundance at Dome were not preferentially associated with early- or late-replicating regions. Together with the distinct developmental timing of the transcriptional and replication-timing phenotypes, this supports the interpretation that Rif1 has separable roles in the two processes rather than a single local mechanism that directly couples them.

      Reviewer #3 (Public Review):

      Using the zebrafish model system, this manuscript assessed the roles of Rif1 protein in replication timing control and transcription during early development, and successfully demonstrated the differential impact of Rif1 protein in replication timing control and transcription. Moreover, the comprehensive assessments of the impacts of mutating Rif1 on animal development (including animal survival and sexual development) were assessed. Although there are works that examined Rif1's implications in replication timing and transcription separately, this work is unique in assessing all these points at once.

      The strength of this manuscript is the genomic analyses of replication timing and transcription being combined in a single model system. Consequently, this manuscript clearly demonstrates the differential impact of Rif1 in these processes during zebrafish development.

      The weakness of this manuscript is, as the authors comment in the Discussion, analyses of replication timing and transcription were performed using bulk embryos. There is a possibility that tissue-specific changes could have been masked. Tissue-specific or single-cell analysis in the future will fill the gap in the knowledge.

      Some of the findings presented in this manuscript are consistent with previous findings using different models such as Drosophila and mice, whereas other findings do not necessarily agree. I hope further studies will reveal more clearly what is common in these systems, and what is different.

      Also, the suggestion that the Rif1 protein may be implicated in a function similar to Fanconi-Anemia genes/proteins is very intriguing.

      Overall, the data presented in this manuscript sufficiently justify the authors' claims. Moreover, this manuscript provides interesting insights into Rif1's function, as well as how development could be controlled.

      We thank the reviewer for highlighting the strength of analyzing replication timing, transcription, and developmental phenotypes in the same vertebrate model. We agree that bulk-embryo measurements may mask tissue- or cell-type-specific effects. We now emphasize this limitation and the need for future tissue-specific or single-cell studies, particularly in the cell populations relevant to sex determination. We also expanded the cross-species context by discussing the recent mouse-embryo study by Nakatani et al. (2025), which supports a conserved role for RIF1 in consolidation of the replication-timing program while also indicating that replication-timing and transcriptional effects can be uncoupled. We agree that defining which Rif1 functions are conserved across zebrafish, mouse, Drosophila, and other systems, including possible relationships to Fanconi-anaemia pathways, will be an important direction for future work.

      Reviewing Editor:

      While the paper was under revision, a relevant paper from the Torres-Padilla lab was published (Nakatani et al., Developmental Cell, 2025). It complements these studies and cites the previous version of this manuscript. I suggest adding a reference in the Discussion to support the conclusions.

      We thank the Reviewing Editor for bringing the recent study by Nakatani et al. to our attention. We have added a standalone paragraph near the end of the Discussion explaining how this work complements our findings, and we have added the complete reference to the bibliography. The new Discussion text reads:

      “A recent study in mouse embryos independently identified RIF1 as a regulator of the developmental consolidation of the RT program. RIF1 depletion produced a less-defined, developmentally immature RT program, while RIF1-dependent RT changes were not correlated with transcriptional changes (Nakatani et al., 2025). Together with our findings in zebrafish, these results support a conserved role for RIF1 in sharpening replication timing during vertebrate development and indicate that its effects on replication timing can be uncoupled from changes in gene expression.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The results presented in this manuscript provide new information on the effects of rif1 mutation on replication and transcription during early zebrafish development, although the underlying mechanism has not been explored. I suggest authors consider conducting the following experiments.

      (1) Does replication timing domains have any role in Rif1-mediated regulation of transcription? It is not clear from the data presented whether transcriptionally affected genes are in the early replicating domains or late replicating domains (that appear after the shield stage). This should be examined.

      We thank the reviewer for this helpful suggestion. To address whether transcriptional effects in rif1 mutants are associated with replication timing, we assigned each gene the nearest smoothed replication timing value and compared replication timing distributions for genes whose transcript levels increased at Dome, decreased at Dome, or were not significantly changed. This analysis is now shown in Figure 5—figure supplement 2. Genes with increased or decreased transcript abundance at Dome did not show a clear enrichment for either early- or late-replicating regions relative to genes with no significant transcript change. This was also true when replication timing was examined at pre-MBT, the stage preceding the major transcriptional changes detected at Dome. These results argue against replication timing state being the primary determinant of the Dome-stage transcriptional changes observed in rif1 mutant embryos. We have revised the Results to describe this analysis and added Figure 5—figure supplement 2.

      (2) It is of interest whether the Rif1-mediated regulation of transcription and replication are mediated by a common mechanism, e.g. through alteration of chromatin structures. Close look at the data in Figure 3D indicates that some genome segments convert replication timing or undergo significant changes of replication timing. It would be informative to know whether these segments (Rif1-regulated replication domains) are associated with the genes whose expression change upon rif1 knockout.

      We thank the reviewer for this insightful suggestion. We agree that an association between Rif1-dependent replication timing changes and Rif1-dependent transcriptional changes would be informative, and we considered this analysis. We attempted to identify Rif1-regulated replication timing domains using the same approach that we previously used to define developmentally regulated timing domains. However, the effect of Rif1 loss differed qualitatively from the developmental timing switches described in our prior work. Rather than producing a limited set of discrete timing-domain transitions, Rif1 loss caused a broad reduction in the dispersion of replication timing values across the genome, consistent with a general flattening of the timing profile. Under these conditions, an unbiased domain-calling approach preferentially identifies genomic regions with the most extreme early or late timing values in wild-type embryos, because these regions show the largest shift toward the mean in rif1 mutants. Thus, the resulting “Rif1-regulated replication domains” largely reflect the strongest wild-type timing domains rather than a discrete set of Rif1-specific regulatory intervals. For this reason, we do not think that assigning differentially expressed genes to such domains would provide a meaningful test of whether Rif1 regulates transcription and replication timing through a common local mechanism. Instead, we have now added a gene-centred analysis comparing replication timing values for genes with increased, decreased, or unchanged transcript abundance at Dome (Figure 5—figure supplement 2), which directly addresses whether transcriptionally affected genes are associated with early- or late-replicating regions.

      (3) Replication is analyzed only by timing analysis. Authors need to analyze frequency of origin firing and replication fork rate by DNA fiber analyses to see whether they are affected by rif1 knockout at various stages of development.

      We agree that measuring origin firing frequency and replication fork rate would provide valuable additional information about how Rif1 loss affects the replication program. However, performing DNA fibre analyses across multiple zebrafish developmental stages and genotypes would require substantial optimization and experimental expansion beyond the scope of the current revision. The current study was designed to measure genome-wide replication timing and transcript abundance across developmental stages, rather than single-molecule replication dynamics. We therefore have not added DNA fibre experiments. Instead, we have revised the Discussion to acknowledge this limitation and to clarify that replication timing reflects the combined effects of origin usage, fork progression, fork directionality, and fork stability. We added the following text to the Discussion:

      “A further limitation of this study is that we concentrated on replication timing without directly measuring other features of the replication program that contribute to this timing. These features include origin usage, replication fork spacing, fork directionality, fork progression, and fork stability. A more comprehensive understanding of how Rif1 loss affects these parameters will be important for defining the relationship between Rif1-dependent changes in replication timing and transcription.”

      Figure 4B, D and F: I did not see the blue lines which represent preMBT in the panels shown.

      We thank the reviewer for identifying this error. The pre-MBT data were not intended to be shown in Figures 4B, 4D, and 4F. We have corrected the figure legend by removing the reference to the blue pre-MBT line.

      Line 270: Figure 4G should be Figure 6G.

      We thank the reviewer for identifying this error. We have corrected the figure reference from Figure 4G to Figure 6G.

      No description of Figure 6E and 6F in the main text.

      We thank the reviewer for noting this omission. We have added text to the Results describing Figures 6E and 6F. The revised text explains that Dome Up-DEGs are normally upregulated from pre-MBT to Shield stages but show earlier upregulation in rif1 mutant embryos, whereas Dome Down-DEGs normally decrease between Dome and Shield stages but show earlier reduction in mutant embryos.

      Reviewer #2 (Recommendations For The Authors):

      (1) This study is an extension of the lab’s previous work which established the wild-type genome-wide replication timing pattern during zebrafish development. The experimental details and analysis are described in the methods, but the general strategy is sometimes treated very cursorily. A non-expert can only understand parts of it by going back to the Seifert study.

      We thank the reviewer for pointing this out. We agree that the replication-timing strategy should be understandable without requiring readers to consult our previous study. We have revised the manuscript to explain the general logic of the assay more clearly. Specifically, we now state that replication timing was inferred from copy-number differences between S-phase and G1-phase genomic DNA: genomic regions that replicate early in S phase are enriched in S-phase DNA relative to G1 DNA, whereas later-replicating regions are less enriched. We also clarified that pre-MBT, dome, and shield embryos were treated as S-phase samples because most cells are in S phase at these stages, whereas nuclei from bud and 24 hpf embryos were sorted by DNA content to isolate G1 and S-phase fractions. These additions make the experimental design and interpretation of the replication-timing profiles clearer in the main text and Methods.

      Figure 2 is meant to document developmental delay in early embryos, but the differences between the single wt and mutant examples in 2D are poorly described and labeled. Most readers will be unfamiliar with the specifics of zebrafish development. There is also no quantification of this developmental phenotype, and that quantification should be included along with better labeling and description of 2D.

      We thank the reviewer for pointing this out. We agree that the developmental delay shown in Figure 2D required clearer explanation and quantification for readers who are less familiar with zebrafish gastrulation. We have revised the Results to explain that epiboly is the process by which the blastoderm and yolk syncytial layer move toward the vegetal pole to envelop the yolk cell, and that zebrafish gastrulation stages are commonly described by the percentage of yolk coverage. We also added quantification of this phenotype. At 10 hpf, most wild-type embryos had completed epiboly, whereas most rif1 mutant embryos had not: 18 of 24 wild-type embryos, but only 2 of 24 mutant embryos, had reached 100% yolk coverage. By 11 hpf, all wild-type and mutant embryos had completed epiboly. These revisions clarify that rif1 mutant embryos show an approximately 1-hour delay in epiboly completion rather than a persistent arrest in gastrulation.

      (3) The presentation could be greatly improved with additional information about the experimental approach and display. As written, the text and figure legends assume readers are intimately familiar with replication timing experiments, zebrafish development, and differential gene expression analysis. Most of the figure legends are not sufficient to understand the figures themselves, and the necessary information is also not always in the results. An example is Figure 3 which is not well described (other than the PCA plots); the term “lag” which is the x-axis in 3C is not defined.

      We thank the reviewer for this helpful comment. We agree that several aspects of the replication-timing analysis required clearer explanation for readers who are less familiar with replication-timing experiments. We have revised the Results to explain the logic of the replication-timing assay more clearly and have added a more detailed description of the autocorrelation analysis in Figure 3C. Specifically, we now explain that autocorrelation measures how similar replication-timing values are across increasing genomic distances along the same chromosome, providing a quantitative readout of the peak-and-valley structure of the timing profile. We also clarified that increasing autocorrelation across hundreds of kilobases reflects the progressive establishment of broader replication-timing domains during development. In addition, we changed the x-axis label in Figure 3C from “lag” to “Genomic distance (Mb).” Together, these changes should make the experimental approach and display easier to understand without requiring readers to consult our previous replication-timing study.

      Figure 4 is generally poorly described and labelled (4B, D, and F graph legends indicate preMBT in the data, but there are no blue lines on the graphs), and Figures 6 and 7 are quite busy.

      We thank the reviewer for pointing this out. We agree that the Figure 4 legend incorrectly described the data shown in panels B, D, and F. The pre-MBT data were not intended to be plotted in these panels, and we have removed the corresponding reference from the figure legend. We recognize that Figures 6 and 7 contain several analyses, but we have retained the current organization because the panels in each figure address a connected set of questions. Figure 6 summarizes how Rif1 loss affects abundance of developmentally regulated transcripts, whereas Figure 7 extends this analysis by directly measuring nascent transcription using SLAM-seq.

      Reviewer #3 (Recommendations For The Authors):

      I do not think any additional experiments are required to justify the authors’ claims. Well done! However, for readers’ benefit, I propose the following changes or adding more explanations:

      (1) Page 2, line 86: I guess “single copy” means “single copy per haploid”. Better to clarify this point.

      We thank the reviewer for this helpful clarification. The reviewer is correct that “single copy” refers to a single copy per haploid genome. We have revised the text to state that the zebrafish genome has a single copy of the rif1 gene per haploid genome.

      (2) Related to the data presented in Figure 2C, do you have an explanation for why sex determination is affected in the heterozygotes, despite the change in Rif1 expression being subtle (Figure 1C)?

      We thank the reviewer for raising this point. We agree that the reduction in whole-embryo rif1 mRNA levels in heterozygotes appears modest relative to the sex-ratio phenotype. At present, we can only speculate about the basis for this difference. One possibility is that whole-embryo mRNA measurements do not accurately reflect Rif1 abundance in the specific cell populations that influence zebrafish sex determination, such as primordial germ cells or their supporting somatic cells. We have therefore avoided making a strong mechanistic conclusion from the heterozygous phenotype.

      (3) Related to the data presented in Figure 2D, did you observe a delay in heterozygotes?

      We thank the reviewer for this question. We have not quantitatively analyzed epiboly progression in heterozygous embryos. However, we did not observe an obvious developmental delay in heterozygotes during early development. The delay shown in Figure 2D was observed in maternal-zygotic rif1 homozygous mutants.

      (4) Figure 3D: it is not easy to distinguish WT and mutant lines, particularly for the Bud stage. Please consider changing the colour schemes or other aspects. For example, making colour lines thinner may help.

      We thank the reviewer for this helpful suggestion. We agree that the wild-type and mutant profiles in Figure 3D, particularly at the bud stage, were difficult to distinguish in the original version. We have revised Figure 3D by reducing the line width of the colored profiles, which improves the contrast between the wild-type and mutant traces.

      (5) Figure 3E: Could you avoid overlapping of WT and mutant plots?

      We thank the reviewer for this suggestion. We considered separating the wild-type and mutant density plots in Figure 3E, but we have retained the overlaid format because the purpose of this panel is to directly compare the distributions of replication timing values between genotypes at each developmental stage. Overlaying the plots makes the reduced dispersion of timing values in the rif1 mutants easier to visualize relative to the corresponding wild-type distribution.

      (6) Figure 4C and 4E: the point legends (WT and mutant) do not match the points used in the graph.

      We thank the reviewer for noting this potential source of confusion. In Figures 4C and 4E, point shape indicates genotype, with open squares representing wild-type samples and open circles representing rif1 mutant samples. Point color indicates developmental stage. We used separate visual encodings for genotype and stage to avoid a large legend containing every genotype-stage combination. To make this clearer, we have revised the figure legend to state explicitly that point shape denotes genotype and point color denotes developmental stage.

      (7) Figure 4D: Very difficult to recognise 24 hr mutant line. Please improve the way there are shown.

      We thank the reviewer for this helpful suggestion. We agree that the 24 hpf mutant profile in Figure 4D was difficult to distinguish in the original version. We have revised the figure by changing the appearance of the mutant lines to make them more visible while preserving the stage color scheme.

      (8) Related to data presented in Figure 4B. Is it possible to show a statistical evaluation of all (or a reasonably large number of samples from) DARs?

      We thank the reviewer for this suggestion. Figure 4A already provides a genome-wide analysis of the DAR set shown by example in Figure 4B. Specifically, Figure 4A plots the change in replication timing from shield to 24 hpf for all 2,498 putative enhancer-associated DARs in both wild-type and rif1 mutant embryos. The strong correlation between wild-type and mutant values indicates that DAR-associated timing changes are largely preserved in rif1 mutants. Because all DARs used for this analysis are included in the scatterplot, we did not add a separate statistical analysis of selected examples from Figure 4B.

      (9) Page 8, line 220: It is unclear what “all” means. Is it all the available replication timing values genome-wide? Please clarify.

      We thank the reviewer for noting this ambiguity. In this sentence, “all” refers to all genome-wide replication timing values calculated from the genomic windows used in our replication timing analysis. We have revised the text to make this clearer.

      (10) Figures 6C and 6D: Colour labels are too dark and it is almost impossible to read texts inside. Please reconsider the colour scheme.

      We thank the reviewer for pointing this out. We agree that the labels in Figures 6C and 6D were difficult to read because of insufficient contrast. We have changed the text colour inside the colored boxes to white to improve legibility.

      (11) Related to overall transcription studies: Is there any sign that Rif1 mutation affects the transcription of genes involved in sex determination?

      We thank the reviewer for raising this interesting question. We have not specifically analyzed whether genes involved in sex determination are differentially expressed in the early embryonic transcriptome data. Because zebrafish sex determination occurs substantially later than the embryonic stages analyzed here, and likely depends on specific cell populations such as primordial germ cells and supporting gonadal somatic cells, we do not think the current whole-embryo RNA-seq data can directly resolve this question. We therefore avoid drawing a mechanistic connection between the early transcriptional changes and the adult sex-ratio phenotype. Determining whether Rif1 mutation affects transcription in the cell populations that regulate zebrafish sex determination will be an important direction for future work.

    1. eLife Assessment

      In this manuscript, the authors analyse the nanoscale localisation of α5β1 and αVβ3 integrins in integrin adhesion complexes (IAC) by dual-colour STORM and DNA-Paint and assess the spatial organisation at the nano and mesoscale of their main adaptors (paxillin, talin and vinculin). This is an important work that provides detailed analyses that reveal how elements of these complex structures are really organised at the nanoscale, an essential perspective for a better understanding of how IACs function and regulate mechanotransduction processes. The evidence presented is convincing, using complementary super-resolution imaging techniques and subsequent computational modelling that enabled a quantitative assessment of the resulting data.

    2. Reviewer #1 (Public review):

      Summary:

      In recent years, it becomes increasingly evident how beautifully intricate IAC are at the nanoscale. Studies like the one presented here that shed light on the precise inner organisation of IAC are thus quite important and relevant to obtain better in-depth understanding of IAC functioning and the contribution of different integrin subtypes to cell adhesive and mechanotransductive processes.

      Interestingly, the authors found a distinct localisation of α5β1 and αVβ3 integrin nanoclusters within focal adhesion of human fibroblasts, with α5β1 integrin nanoclusters being at the periphery of IAC and αVβ3 integrin nanoclusters randomly distributed. Furthermore, a surprisingly high percentage of inactive integrins within IAC and relatively low spatial integrin colocalisation with adaptor proteins has been shown.

      Strengths:

      This is a very thoroughly performed STORM-based assessment of the nanodistribution of α5β1 and αVβ3 nanoclusters within IAC (and outside). The image quality is outstanding, and the authors have meticulously executed the experiments and the image analyses.

      Weaknesses:

      The only weakness is maybe that the manuscript remains descriptive. However, the high quality of the "description" of the nano-organisation of IAC by this scrupulous study is really important to better understand the inner workings of IAC. It provides a very solid foundation to look deeper into the (patho)physiological implications of this organisation, see recommendations (which are rather suggestions in this case).

      Comments on revision:

      The authors meticulously addressed all my questions and suggestions. I want to thank the authors for an exemplary revision.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, dual-color super-resolution microscopy analysis was performed to study the co-operation between integrins and focal adhesion proteins in human fibroblast cells. The study focused on two integrins which have been previously found to be mainly responsible for focal adhesions, namely α5β1 and αvβ3.

      Specifically, the study tried to shed light on the nanoclustering of integrins in focal adhesions.

      In the current study, more integrin nanoclusters were observed in focal adhesions compared to other cell-matrix adhesion structures. The study revealed that both α5β1 and αvβ3 form nanoclusters and those appear segregated from each other. While αvβ3 nanoclusters organize randomly inside focal adhesions regardless of their activation state, α5β1 nanoclusters, and particularly the nanoclusters containing β1-integrin in active conformation preferentially organized at the edges of focal adhesions. The nanoclusters formed by each integrin were similar in size.

      Cytoplasmic adapter proteins appeared less in nanocluster assemblies, suggesting that integrin nanoclusters are also forming without the studied cytoplasmic adapter proteins (talin, vinculin, paxillin). Active integrins were identified with help of conformation-specific antibodies, and those enabled to study the colocalization between integrins and their cytoplasmic adapter proteins. This analysis revealed that activated integrins are strongly engaged with adapter proteins

      Strengths:

      The study stems from the thorough computational modelling of the nanoclusters, which enables quantification of the behavior of the clusters, including their mesoscale distribution.

      The study strengthens the view that α5β1 and αvβ3 have specific functions in focal adhesions, α5β1 nanoclusters localizing preferentially on focal adhesion edges. The study also revealed that nanoclusters localized at the edges of focal adhesion were enriched for talin and paxillin but not for vinculin.

      Analysis of adaptor protein nanoclusters (paxillin, talin, and vinculin) revealed that all adapter protein nanoclusters studied here close to active β1 nanoclusters are enriched on the focal adhesion edge region, whereas integrin adaptor nanoclusters far from active β1 appear to be more uniformly distributed.

      Importantly, the current study suggests that integrin subtype-specific nanoclusters are not only present at early stage of adhesion formation, but integrin nanoclusters remain segregated from each other also in mature focal adhesions, maintaining their sizes and number of molecules.

      Interestingly, the study revealed that selected cytoplasmic adaptors (paxillin, talin and vinculin), also form nanoclusters of similar size and number of single molecule localizations as the integrins, regardless of whether they locate inside or outside focal adhesions. The adapter nanoclusters are enriched in the focal adhesion "belt", colocalizing with the active α5β1 integrin nanoclusters.

      Weaknesses:

      The current study is highly dependent on the antibodies. It is possible, that antibodies, containing two binding sites for antigen, influence the nanoscale organization (and also activation) of the receptors. Control experiments to study possible contribution of antibodies for the measured outcome should be performed to verify the main findings. One possible approach could be to use fluorescently tagged integrins available. Alternatively, integrins (or adapter proteins) could be tagged with small ligand and detected using monovalent binder.

      Only a limited number of integrin adapter proteins were investigated. Given the high number of identified adapter proteins, this is an understandable choice. However, it would be fascinating to understand if the nanoclusters of inactive integrins are dominantly bound with certain adapter protein, such as tensin.

      Comments on revision:

      The authors addressed the concern related to the use of antibodies and secondary antibodies by performing DNA-PAINT experiment, which revealed highly similar results as obtained with conventional antibodies.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We have addressed all the concerns and recommendations by the reviewers, in particular, the requested control experiments using alternative super-resolution microscopy approaches and analysis of the data using Voronoi tessellation in addition to DBSCAN, as requested by reviewer 3. We also provide additional data on tensin3 as suggested by reviewer 2. Finally, to provide a first insight on the role of mechanical forces in the distribution of integrin nanoclusters inside FAs as recommended by reviewer 1, we have performed experiments at different cell seeding times where it is known that FA maturation over time requires mechanical forces.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In recent years, it has become increasingly evident how beautifully intricate IAC are at the nanoscale. Studies like the one presented here that shed light on the precise inner organisation of IAC are thus quite important and relevant in order to obtain a better in-depth understanding of IAC functioning and the contribution of different integrin subtypes to cell adhesive and mechanotransductive processes.

      Interestingly, the authors found a distinct localisation of α5β1 and αvβ3 integrin nanoclusters within focal adhesion of human fibroblasts, with α5β1 integrin nanoclusters being at the periphery of IAC and αvβ3 integrin nanoclusters randomly distributed. Furthermore, a surprisingly high percentage of inactive integrins within IAC and relatively low spatial integrin colocalisation with adaptor proteins has been shown.

      Strengths:

      This is a very thoroughly performed STORM-based assessment of the nanodistribution of α5β1 and αvβ3 nanoclusters within IAC (and outside). The image quality is outstanding, and the authors have meticulously executed the experiments and the image analyses.

      We are grateful to the reviewer for acknowledging the strengths of our study.

      Weaknesses:

      The only weakness is maybe that the manuscript remains descriptive. However, the high quality of the "description" of the nano-organisation of IAC by this scrupulous study is really important to better understand the inner workings of IAC. It provides a very solid foundation to look deeper into the (patho)physiological implications of this organisation, see recommendations (which are rather suggestions in this case).

      We thank the reviewer for their feedback and have addressed their recommendations in our updated manuscript and accompanying reply (see recommendations to the authors). In summary, we have now performed experiments at different seeding times as FA maturation requires mechanical forces, and enquired whether forces might play a role in establishing the spatial distribution of the two different integrins within more mature IACs. The results are now shown as new Fig. 2 and discussed in pages 9 and 10. In addition, in order to get a first insight into the biological implications of our findings we performed dual-colour super-resolution experiments of tensin-3 and α<sub>5</sub>β<sub>1</sub> in FAs, as tensin-3 has been implicated in fibronectin fibrillogenesis. The results are now shown in Fig. S8 and we discuss their potential implications in pages 22 and 23 of the revised manuscript (see more details in the reply to the recommendation to the authors).

      Reviewer #2 (Public review):

      Summary:

      In this study, dual-color super-resolution microscopy analysis was performed to study the co-operation between integrins and focal adhesion proteins in human fibroblast cells. The study focused on two integrins which have been previously found to be mainly responsible for focal adhesions, namely α5β1 and αvβ3.

      Specifically, the study tried to shed light on the nanoclustering of integrins in focal adhesions.

      In the current study, more integrin nanoclusters were observed in focal adhesions compared to other cell-matrix adhesion structures. The study revealed that both α5β1 and αvβ3 form nanoclusters, and those appear segregated from each other. While αvβ3 nanoclusters organize randomly inside focal adhesions regardless of their activation state, α5β1 nanoclusters, and particularly the nanoclusters containing β1-integrin in active conformation, preferentially organized at the edges of focal adhesions. The nanoclusters formed by each integrin were similar in size.

      Cytoplasmic adapter proteins appeared less in nanocluster assemblies, suggesting that integrin nanoclusters are also forming without the studied cytoplasmic adapter proteins (talin, vinculin, paxillin). Active integrins were identified with the help of conformation-specific antibodies, and this enabled us to study the colocalization between integrins and their cytoplasmic adapter proteins. This analysis revealed that activated integrins are strongly engaged with adapter proteins.

      Strengths:

      The study stems from the thorough computational modelling of the nanoclusters, which enables quantification of the behavior of the clusters, including their mesoscale distribution.

      The study strengthens the view that α5β1 and αvβ3 have specific functions in focal adhesions, α5β1 nanoclusters localizing preferentially on focal adhesion edges. The study also revealed that nanoclusters localized at the edges of focal adhesion were enriched for talin and paxillin but not for vinculin.

      Analysis of adaptor protein nanoclusters (paxillin, talin, and vinculin) revealed that all adapter protein nanoclusters studied here close to active β1 nanoclusters are enriched on the focal adhesion edge region, whereas integrin adaptor nanoclusters far from active β1 appear to be more uniformly distributed.

      Importantly, the current study suggests that integrin subtype-specific nanoclusters are not only present at an early stage of adhesion formation, but integrin nanoclusters remain segregated from each other also in mature focal adhesions, maintaining their sizes and number of molecules.

      Interestingly, the study revealed that selected cytoplasmic adaptors (paxillin, talin, and vinculin), also form nanoclusters of similar size and number of single molecule localizations as the integrins, regardless of whether they locate inside or outside focal adhesions. The adapter nanoclusters are enriched in the focal adhesion "belt", colocalizing with the active α5β1 integrin nanoclusters.

      We are grateful to the reviewer for acknowledging the strengths of our study.

      Weaknesses:

      The current study is highly dependent on the antibodies. It is possible that antibodies containing two binding sites for antigen influence the nanoscale organization (and also activation) of the receptors. Control experiments to study the possible contribution of antibodies to the measured outcome should be performed to verify the main findings. One possible approach could be to use fluorescently tagged integrins available. Alternatively, integrins (or adapter proteins) could be tagged with a small ligand and detected using a monovalent binder.

      We understand the concern of the reviewer regarding the use of antibodies for imaging. Nevertheless, we would like to clarify that antibody labelling has always been performed after cell fixation, precluding potential cross-linking artefacts due to protein mobility and avoiding unwanted receptor activation.

      Nevertheless, and although it is highly unlikely to happen in fixed cells, there could be two potential sources of antibody (Ab) labelling artefacts. As the reviewer noted, a primary Ab containing two binding sites could bind to two adjacent proteins (within ~10 nm from each other), potentially underestimating the stoichiometry of the nanoclusters, i.e., number of receptors or proteins per nanocluster. However, in our manuscript we never attempted to provide an estimation of the nanocluster stoichiometry, as it is highly challenging (and prone to artefacts) to provide quantification of the number of proteins using super-resolution-based single-molecule localisation methods which rely on the stochastic blinking of individual fluorophores.

      A second source for potential artefacts comes from the use of the secondary Ab, which (albeit unlikely) could bind to two different primary Abs. To exclude this potential artefact, we performed super-resolution imaging using DNA-PAINT as a different imaging strategy. In this case, the DNA docking site is site-specifically coupled to one camelid single-domain Ab (sdAB), having a much smaller size as compared to a secondary Ab, reducing therefore linkage error and increasing the accessibility of primary Ab-labelled proteins. These new data are included now in Fig. S4. As can be observed, no differences in terms of nanocluster sizes and/or compositions were observed for any of the proteins investigated using DNA-PAINT as compared to our initial STORM data. These control experiments thus rule out any potential artefacts introduced by the secondary Ab (for more details, please see the reply to the recommendations for authors section).

      Only a limited number of integrin adapter proteins were investigated. Given the high number of identified adapter proteins, this is an understandable choice. However, it would be fascinating to understand if the nanoclusters of inactive integrins are dominantly bound with a certain adapter protein, such as tensin.

      We fully agree with the reviewer and have now performed dual-colour super-resolution STED microscopy of α<sub>5</sub>β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, instead of being an integrin inactivator, we found that tensin-3 is also highly enriched at the FA periphery where a large fraction of active β<sub>1</sub> integrins are located, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original data) or to tensin-3 (our new data shown in Fig. S8). We provide more details of our answer in the section of “recommendation to the authors”. Additional experiments, which in our opinion fall outside of the scope of this work, would be necessary to identify other potential integrin inactivator partners, but certainly a topic of future interest to our group.

      Reviewer #3 (Public review):

      Summary:

      In their study, the authors reveal using dual-color super-resolution STORM microscopy modality and immunolabeling in fixed adherent cells, that β1 and β3 integrins as well as adaptors (paxillin, talin and vinculin) are all organized in nanoclusters of similar size (50nm) and molecular density (20 copy number) inside FAs but also outside. Using activityspecific immunolabeling of β1 and β3 integrins, they revealed that active integrin subpopulations were both clustered but in distinct exclusive nano-aggregates in agreement with Spiess et al. (2018). Once more, the "active" integrin nanoclusters displayed similar properties in terms of size and molecular density, suggesting that molecular organization in nanoclusters is an intrinsic property of integrins in plasma membrane multimerizing independently of their location (inside or outside FAs), their level of activation, or their connection to the cytoskeleton. Then the authors followed up by analyzing at the mesoscale how these "universal" nanoclustered adhesive units are distributed spatially. Inspecting the surface density of nanoclusters revealed that the density of integrin nanoclusters in FAs was 5x larger, compared to integrin nanoclusters outside adhesions. Interestingly, whereas the density of total integrin nanoclusters was 2-4x larger than adaptor nanoclusters, the density of "active" integrin nanoclusters stoichiometrically matches that of talin and vinculin nanoclusters, and was slightly outnumbered by paxillin nanoclusters. These findings suggest that inside FAs, among the total number of integrin nanoclusters, the subset of "active" integrin nanoclusters could be engaged with "adaptor" nanoclusters on a 1:1 ratio. Using analysis of the nearest neighbor distance (NND) between distinct integrin clusters and each of the adaptors, the authors report that they found negligible spatial colocalization of integrins with these adaptor proteins and that spatial segregation is essentially determined by the density of nanoclusters within the FAs. As authors reported that α5β1 and αvβ3 do not intermix at the nanoscale, the authors finally highlighted how α5β1 and αvβ3 distinct nanoclusters are differently organized and segregated inside FAs. Adapting the NND analysis in order to inspect how far the nanoclusters are from the edges of FAs they are located in, authors revealed that α5β1 but not αvβ3 integrin nanoclusters are enriched on FA edges and that similar FA edge-enriched distribution for "active" α5β1 and adaptor protein nanoclusters was found for talin and paxillin but not vinculin. The latter results suggest that FA edges could constitute multiprotein hubs for enhanced colocalization and activation for α5β1 integrin nanoclusters and adaptors such as talin and paxillin. Unfortunately NND analysis could not confirm this enhanced colocalization hypothesis.

      General Assessment:

      While the study presents some valuable findings, it reads currently as a compilation of intriguing but preliminary observations derived primarily from a single methodology (dual-color STORM and DBSCAN clustering analysis). As the initial findings often lack confirmation through additional data analysis (such as the NND analysis the authors used), there's a critical necessity to bolster the methodological approach. This should involve replicating the main findings using alternative single-molecule super-resolution techniques (such as quantitative DNA-PAINT) or employing different clustering analytical tools (such as voronoi-tessellation). Furthermore, the manuscript feels incomplete, focusing solely on describing molecular organization without offering substantial insights into how these observations correlate with the regulation, activation, and functionality of integrins at the cellular level.

      We appreciate the comment of the reviewer and have taken their recommendation to heart in order to validate our methodology. In summary, we have now performed extensive DNA-PAINT to replicate most of our initial findings obtained by STORM, as requested by the reviewer. In addition, as a different super-resolution imaging strategy, we have also used STED microscopy to confirm the nanoclustering of integrins and some of the adaptors demonstrating now, by means of three different super-resolution techniques, that both integrins and their adaptors form nanoclusters of similar size and composition, regardless of whether they are inside or outside FAs. We have included these data as Figs. S3 and S4 and discussed the results in pages 8-9 of the main manuscript.

      Regarding the use of an alternative analysis for the data, we have now used the Voronoi tessellation algorithm to re-analyse our STORM data, as requested by the reviewer. The results of the analysis, which render similar sizes and number of localizations as obtained by DBSCAN, are now included in Fig. S5 and mentioned in page 8 of the main manuscript.

      The manuscript presents extensive datasets and utilizes methodologies in which the investigators demonstrate expertise. Nevertheless, there's uncertainty regarding the novelty and broad appeal of the findings. For instance, the observation of integrin nanoclustering has been previously reported in several publications (e.g., Changede et al., Dev Cell 2015; Spiess et al., JCB 2018; Fujiwara et al., JCB 2023). Similarly, the accumulation of specific proteins at the periphery of FAs has been documented elsewhere (e.g., Sun et al., NCB 2016; Stubb et al., NatComm 2019; Nunes-Vicente TCB 2023), as well as the differential dynamic organization of α5β1 and αvβ3 integrins inside FAs (e.g., Rossier et al., NCB 2012). Beyond the universal organization of adhesive proteins, there's a need to identify novel insights that significantly advance the field. One potential avenue could involve pinpointing the molecular determinant controlling the FA edge enrichment of active α5β1 integrins and talin nanoclusters. For instance, could there be an interplay between α5β1 and αvβ3 integrin nanoclusters visible on one's organisation when suppressing the other using deletion (KO) or depletion (SiRNA)? Also, could KANK, which also exhibits enrichment and regulates talin activity (e.g., Sun et al., NCB 2016), play a role in this process? Identifying the molecular players that regulate even partially the mesoscale organization of nanoclusters of proteins would really benefit the breadth of this manuscript.

      We could not agree more with the reviewer and in fact, we are currently investigating the mechanisms that control the enrichment of α<sub>5</sub>β<sub>1</sub> and adaptors at the edges of FAs. However, considering the amount of work needed to determine the spatiotemporal organization of other molecular players using super-resolution imaging constitutes a major tour de force.

      To get a first insight into the process of active α<sub>5</sub>β<sub>1</sub> enrichment at the FA edges, we hypothesised that mechanical forces exerted by the actomyosin machinery could influence the lateral distribution of both integrin subsets (α<sub>5</sub>β<sub>1</sub> and α<sub>v</sub>β<sub>3</sub>) inside FAs. Since FA maturation and strengthening over time requires mechanical forces, we performed experiments at different cell seeding times (90 min, 3 hours and 24 hours) and used STORM imaging to follow the evolution of integrin nanoclustering in time as well as their spatial distributions inside FAs. Interestingly, while nanoclustering of both integrin sub-sets inside FAs is not influenced by seeding times, their lateral distribution was markedly different, with α<sub>5</sub>β<sub>1</sub> nanocluster distribution being already established at earlier seeding times, while α<sub>v</sub>β<sub>3</sub> nanocluster distribution appeared as rather random at earlier seeding times and progressively organized reaching a well-defined lateral spacing at 24 hours of spreading time. These initial data strongly suggest that mechanical forces might play a role in the distinct lateral distribution of both subsets of integrin nanoclusters over time. We have now included these data as new Fig. 2 of the revised manuscript and discuss the results in the associated text (pages 9 and 10). We also discuss potential avenues for further research along the directions suggested by the reviewer.

      In addition, since it has been recently shown that tensin-3 interaction with talin drives the formation of fibronectin-associated fibrillar adhesions (Atherton et al, J Cell Biol 2022) which are enriched in β<sub>1</sub> integrins, we performed dual-colour super-resolution STED microscopy of β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, our initial data show co-enrichment of both tensin-3 and active β<sub>1</sub> nanoclusters at the FA periphery, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original manuscript) or to tensin. Our current working hypothesis is that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery serves to facilitate the translocation of α<sub>5</sub>β<sub>1</sub> integrins from FAs to fibrillar adhesions, most probably in a talin-tensin-dependent manner. We have now included these data as Fig. S8 and accompanying discussion in pages 22 and 23 of the revised manuscript.

      Echoing the previous concern, the manuscript described a novel and rather surprising finding related to molecular clustering of adhesion proteins. Indeed, the fact that nanoclusters exhibit uniform size and molecular density regardless of the protein type, location, or activation level is indeed surprising and raises many questions about the methodology used to assess molecular clustering. I feel that the description and characterization of integrin nanoclusters appear incomplete and need to be expanded by comparing different analytical strategies for protein clustering. Furthermore, a lack of the manuscript in its actual form concerns the quantification of integrin numbers inside the observed nanoclusters. I agree that the path from optical microscopy to protein stoichiometry quantification is hard and full of drawbacks. But the authors do not fully address these issues that are extremely important when discussing protein nanoclustering. This quantitative aspect should be discussed.

      We appreciate the comment of the reviewer as indeed, the existence of “universal” nanoclusters is intriguing. Recently, together with Prof. S. Mayor we have written a short review in Curr. Opin. Cell Biol 2024 proposing that nanoclustering constitutes a molecular-scale organisation principle that governs cellular information flow at the plasma membrane. Our proposal is supported by an extensive number of recent papers showing that most cell membrane receptors and downstream signalling components are organized as pre-assembled nanoclusters. We posit that these nanoclusters serve as modular units whose concatenation in a specific spatiotemporal sequence leads to distinct signalling outputs. Thus, the existence of universal nanoclusters of integrin receptors and adaptors is indeed intriguing but not surprising to us.

      In any case, the concern of the reviewer is well-taken, and as mentioned above, we have used a different algorithm to detect and quantify nanoclustering, obtaining similar values using either Voronoi tessellation or DBSCAN approaches. These data are now included as Fig. S5 in the manuscript.

      Regarding the quantification of integrin numbers inside the observed nanoclusters, we agree with the reviewer that determining protein stoichiometry using single-molecule localization microscopy or STED remains a major technical challenge and is highly prone to artefacts. For this reason, we refrain from making claims about absolute protein numbers per nanocluster. Our relative comparison of nanoclustering among the different proteins investigated is thus exclusively based on the number of single-molecule localisations contained in each nanocluster which is a fair approach since we always use the same reporter fluorophore and maintain similar excitation conditions throughout our experiments. We have now included a few lines on page 9 regarding quantification of the absolute protein numbers inside the nanoclusters and further discuss in the revised manuscript the limitations of single-molecule localisation methods towards the stoichiometry determination of the nanoclusters (see page 20 of the revised manuscript).

      First, it is crucial for the authors to carefully examine and discuss in their manuscript whether there are any potential biases or limitations in the experimental techniques (dual-color STORM) or data analysis methods employed (DBSCAN). Second, the authors did not in the current manuscript, but should provide control samples to demonstrate the sensitivity and dynamic range of their experimental strategy.

      As already mentioned, we have validated the STORM data using both DNA-PAINT and STED and, validated our data analysis obtained with DBSCAN using the Voronoi tessellation algorithm. See Figs. S3, S4 and S5. In terms of sensitivity and dynamic range of our methodology: our set-up has single-molecule detection sensitivity which is demonstrated by the fact that we observe and detect discrete blinking events, a property of single-molecule fluorescence emission and key ingredient to super-resolution single-molecule localisation microscopy. The dynamic range (if we understand correctly the question of the reviewer) is given by the number of frames used to accumulate single-molecule localisations. In our case, we stop acquisition after we deplete most of the single-molecule spots in the imaging view, which typically occurred after 70,000 frames acquisition, as correctly mentioned in the material & methods section.

      In STORM images displayed in Figure S1, the authors highlighted localization clusters detected by DBSCAN as a signature for integrin nanoclusters. But the authors do not discuss the localization spots that were not detected by DBSCAN. Could they be individual integrins? And if so, they should also be considered as useful information? This brings me to another related technical question about how DBSCAN handles the case where fluorescent molecules are blinking. This is important as multiple emissions by a single fluorophore could be detected as a nanocluster of several molecules where it would be an artefact due to the photophysics of the fluorophore. Could the authors comment on these points?

      As mentioned in the original manuscript, between 20-30% of the localizations were not assigned to nanoclusters (Fig. S1H, I) since we imposed a minimum of ten localizations within the radius defined by DBSCAN to be considered as a true nanocluster. This essentially means that regions with less than 10 localizations were not considered in our nanoclustering analysis. However, we cannot be certain as to whether these lower number of localizations correspond to individual integrins, stochastic blinking of the fluorophore or small aggregates containing only a couple of integrins, for the same reasons that we cannot provide quantification of the absolute number of proteins included in each nanocluster: stoichiometry determination by means of single-molecule super-resolution methods is highly prone to artefacts.

      Regarding the concern of how DBSCAN handles fluorophore blinking, the reviewer is completely right as the photophysics of the fluorophore can influence the analysis of the data and the identification of true nanoclusters. To decouple the photophysics of the fluorophore we first assess the number of blinking events within the DBSCAN radius, i.e., number of localizations corresponding to individual antibodies sparsely distributed on the glass surface. In our case, the median values for the two activator-reporter pairs corresponded to 5 localizations for Alexa 405-Alexa647-conjugated Abs and 3 localizations for Cy3-Alexa 647-conjugated Abs (see Fig. 1E). Yet, despite these median values, the number of localizations per individual Ab naturally shows a distribution. Thus, to avoid any overestimation in the degree of nanoclustering, we impose an additional constrain to our analysis and consider true nanoclusters only those ones containing at least 10 localizations. We have now significantly extended the explanation in the main text (see page 6) as well as materials & methods so that it becomes clearer to the reader.

      Also, using isolated and stochastically physisorbed fluorophores (Ab coupled with activator /reporter pairs used in this study) on glass helped define the signature in STORM of a single isolated molecule. To obtain the signature of clustered fluorophores, the authors could use anti-donkey antibodies to cross-link those STORM-specifically labeled Ab as a means to artificially obtain clustered fluorophores. Ultimately, to avoid the bias effect of the glass surfaces on the photophysics of fluorophores and be in the same imaging conditions as for the described nanoclusters, the authors should use model systems composed of multimers of GFP vs. single GFP, immunolabeled with a GFP-binding monoclonal antibody. This will permit evaluation of the cluster signature obtained with DBSCAN analysis of STORM data for single vs. multimers of known stoichiometry. This would constitute an undisputable molecular stoichiometry ruler.

      We appreciate the suggestions of the reviewer. Regarding the potential bias effect of the glass surface on the photophysics of the fluorophores we would like to clarify that the “calibration” for the number of blinking events per individual Ab on glass were performed on the same sample containing the cells that we image, so that we maintain exactly the same experimental and imaging conditions avoiding any potential artefacts. To our understanding this approach is more accurate than performing the calibration on glass substrates and then moving to samples containing the cells. This information is now contained in page 6 of the revised manuscript and in the materials and method section. Once the number of blinking events from individual Abs on glass within the DBSCAN radius are determined, one can then determine the number of localizations within the same DBSCAN radius on other parts of the sample. More localizations within the same DBSCAN radius basically means more molecules, and thus nanoclusters. This approach has been extensively used by other experts in the field as we properly acknowledge in our manuscript (Pageon et al, Mol. Cell. Biol 2016; Spiess et al. J. Cell Biol 2022).

      Using anti-donkey antibodies to cross-link those STORM-specifically labelled Ab in order to artificially obtain clustered fluorophores, as suggested by the reviewer, is indeed a sound approach to retrieve signatures of clustering. Nevertheless, we have preferred not to use this approach because those artificially induced clusters would have very little resemblance to the real nanoclusters and would only allow us to validate the performance of DBSCAN for cluster recognition. As mentioned above, DBSCAN is a well-established algorithm and used by many different experts in the field and thus can be trusted by the community. Instead, and following the recommendation of the reviewer, we now provide results using an alternative cluster analysis algorithm (Voronoi tessellation) reaching similar conclusions regarding the existence of integrin and adaptor nanoclustering inside FAs.

      Finally, the suggestion of using monomeric vs multimeric GFPs to determine the stoichiometry of the nanoclusters is highly appreciated. Indeed, we have used this approach in the past to identify nanoclustering of the chemokine receptor CXCR4 in living T cells (Mol. Cell 2018 and PNAS 2022). However, these experiments are best performed at sub-labelling conditions, which inherently underestimate the degree of nanoclustering. Combining GFPs with PALM to enable super-resolution is another approach but also subject to artefacts regarding the photo-conversion efficiency of GFPs as we reported earlier (Nature Methods 2017) and leading to underestimation of nanocluster stoichiometry.

      In summary, providing nanocluster stoichiometry from single-molecule localisation images remains a major technical challenge and is highly sensitive to methodological assumptions. We have therefore focused here on providing robust evidence for the existence of integrin and adaptor nanoclustering, using three different superresolution approaches and two independent analytical methods for cluster determination.

      Due to the surprising finding of the nanoclusters' "universality", it is imperative for the authors to validate the findings through complementary methodologies and analytical tools. This should involve replication of results using alternative super-resolution techniques (quantitative DNA-PAINT) and exploring different clustering algorithms (VoronoïTesselation) to ensure the robustness and reliability of the observations.

      As already mentioned, we have now performed extensive DNA-PAINT to replicate most of our initial findings obtained by STORM, as requested by the reviewer. In addition, as a different super-resolution imaging strategy, we have also used STED microscopy to confirm the nanoclustering of integrins and some of the adaptors demonstrating now, by means of three different super-resolution techniques, that both integrins and their adaptors form nanoclusters of similar size and composition, regardless of whether they are inside or outside FAs. We have included these data as Figs. S3 and S4 and discussed the results in pages 8-9 of the main manuscript.

      Regarding the use of an alternative analysis for the data, we have now used the Voronoi tessellation algorithm to re-analyse our STORM data, as requested by the reviewer. The results of the analysis, which render similar sizes and number of localizations as obtained by DBSCAN, are now included in Fig. S5 and mentioned in page 8 of the main manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This work already, as is, provides significant and novel information on IAC.

      The unexpectedly low spatial colocalisation of integrins with adaptor proteins might indeed be caused by the potentially quite long extension of talin upon force exposure and the ample zones of activity of IAC proteins, imaging the involved proteins in scale, as can be seen in Barnett and Goult (2022, doi: 10.3389/fncel.2022.1014629)? In super-resolution microscopy, this spatial separation might, in fact, become apparent. It would be interesting to see whether lowering the actomyosin contraction by different concentrations of blebbistatin lowers this separation. In general, it would also be interesting to understand whether lowering the forces can disrupt the nano-organisation and the strong separation of the two analysed integrin subtypes. It is true that nascent adhesion formation is force-independent, but maybe the forces play a role in establishing the particular integrin subtype nano-organisation within more mature IAC. I am also aware that a lot of work has already gone into the conclusion of this project.

      We thank the reviewer for these thoughtful comments and suggestions. Our most recent preliminary data (not yet included in this manuscript) indeed indicate that the physical separation of integrin nanoclusters (and adaptors) inside focal adhesions (FAs) is force-dependent. We are currently reproducing these experiments using lipid bilayers of varying viscosities and controlled ligand density to explore how ligand mobility (i.e., equivalent to force exerted from the extracellular side) controls the degree of IAC nanoclustering and their spatial segregation in FAs. This approach is more amenable to super-resolution microscopy as lipid bilayers are quite thin and optically transparent, yet the experiments are still challenging, time-consuming, and thus ongoing.

      To obtain a first hint as to whether forces might play a role in establishing the spatial distribution of the two different integrins within more mature IACs as the reviewer suggests, we have performed experiments at different seeding times (90 min, 3 hours and 24 hours). Our results show that even at earlier times (90 min), when a lower number of mature FAs are established, nanoclustering of integrins and main adaptors are similar to 24 hours. In contrast, and as suggested by the reviewer, the spatial distribution of the different subsets of integrin nanoclusters inside FAs is markedly different as a function of seeding time, with α<sub>5</sub>β<sub>1</sub> nanocluster distribution being already established at 90 min, while α<sub>v</sub>β<sub>3</sub> nanocluster distribution appears rather random at earlier seeding times and progressively organizes reaching a well-defined lateral spacing at 24 hours of spreading time. As FA strengthening over time requires mechanical forces, and α<sub>v</sub>β<sub>3</sub> is preferentially involved in FA strengthening (Roca-Cusachs et al PNAS 2009), these data strongly suggest that forces play a differential role in the lateral distribution of both integrin nanoclusters over time. We have now included these data as a new Fig. 2 in the revised manuscript and discuss the results in the associated text (pages 9 and 10). We also mention in the discussion additional experiments, as suggested by the reviewer, to further substantiate this hypothesis.

      Considering the high quality of the work and the new insight about the inner organisation of IAC, maybe the summary Figure 5 should be elaborated a bit, taking into account e.g. different lengths of extended talin proteins and also the various positions of vinculins on talin proteins (depending on opened cryptic binding sites), as well as the possibility that various actin filaments might be associated with single talins. What I mean is, the authors impressively demonstrate the complexity of IAC nano-organisation, which should be paid more tribute in the concluding figure. The quality of the figure should be adapted to the quality of the work.

      We have adapted Figure 5 (now Figure 6) as suggested by the reviewer.

      I would be curious to hear a bit more about the further speculations of the authors in the discussion, e.g., about why the integrin subunits are organised in this way. Why might the α<sub>5</sub>β<sub>1</sub> be preferentially located in the periphery? What is the potential physiological relevance of this organisation? Is this organisation different in other cell types (have the authors looked at other cells)? Is the organisation lost in pathophysiological situations, such as cancer?

      Although we do not know yet what drives the preferential location of α<sub>5</sub>β<sub>1</sub> nanoclusters to the FA periphery, it is known that Kank2 also exhibits enrichment at the FA periphery, regulates talin activity and it is involved in the formation of α<sub>5</sub>β<sub>1</sub>-enriched fibrillar adhesions (Sun et al, Nature Cell Biol 2016). Thus, it is highly probable that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery is a necessary step for their translocation from mature FAs to fibrillar adhesions to then assemble fibronectin into the fibrillar networks as found and needed in connective tissues. Consistent with this idea, we have observed similar α<sub>5</sub>β<sub>1</sub> distribution on other fibroblast cell lines (MEFS), which are the primary cells that produce fibrillar adhesions. Thus, α<sub>5</sub>β<sub>1</sub> nanocluster distribution inside FAs might be physiologically important for the process of fibronectin fibrillogenesis.

      Since it has been documented that tensin is important for fibronectin fibrillogenesis (Pankov et al J Cell Biol 2000) and more recently, it has been shown that tensin-3 interaction with talin drives the formation of fibronectin-associated fibrillar adhesions (Atherton et al, J Cell Biol 2022), we thought to investigate the spatial distribution of tensin-3 and its relationship with α<sub>5</sub>β<sub>1</sub> inside FAs by means of dual colour super-resolution STED microscopy. Interestingly, our initial data on HFF cells seeded for 24 hours show both enrichment of tensin-3 and α<sub>5</sub>β<sub>1</sub> nanoclusters at the edges of mature FAs, supporting our working hypothesis that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery serves to translocate α<sub>5</sub>β<sub>1</sub> integrins from FAs to fibrillar adhesions, probably in a talin-tensin-dependent manner. While these initial data are quite exciting, many more experiments that include simultaneous super-resolution mapping of α<sub>5</sub>β<sub>1</sub>, talin and tensin in mature FAs are required to fully validate our hypothesis. Yet, because of their relevance we consider it appropriate to include these data as Fig. S8 and discussing their potential implications in pages 22 and 23 of the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform control experiments to confirm that the nanocluster size/composition is not affected by the antibodies used.

      As explained in the response to the public reviews, antibody labelling has always been performed after cell fixation, precluding potential cross-linking artefacts due to protein mobility and avoiding unwanted receptor activation. In addition, we have performed super-resolution imaging using DNA-PAINT as a different imaging strategy. In this case, the DNA docking site is site-specifically coupled to one camelid single-domain Ab (sdAB), having a much smaller size as compared to a secondary Ab, reducing therefore linkage error and increasing the accessibility of primary Ab-labelled proteins. As can be observed in new Fig S4, no differences in terms of nanocluster sizes and/or compositions were observed for any of the proteins investigated using DNA-PAINT as compared to our initial STORM data. These control experiments thus rule out any potential artefacts introduced by the secondary Ab. Finally, we would like to highlight that our results on the nanoclustering of integrins in terms of their size and number of localizations is consistent with previous results obtained by other groups around the world using similar labelling protocols as us (Spies et al, J. Cell Biol 2022), or relying on halo-tag strategies, as suggested by the reviewer (see Fujiwara et al, J. Cell Biol. 2023). The consistency of these results amongst different groups gives us further confidence that the nanocluster size/composition are not affected by the antibodies used.

      (2) Extend the study by inspecting a set of integrin adapter proteins for their association with inactive integrins, focusing on adapters associated with the maintenance of the inactive state. Possible candidates would be tensin and filamin, for example.

      We thank the reviewer for the suggestion and have now performed dual-colour super-resolution STED microscopy of α<sub>5</sub>β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, instead of being an integrin inactivator, we found that tensin-3 is also highly enriched at the FA periphery where a large fraction of active β<sub>1</sub> integrins are located, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original data) or to tensin-3 (our new data shown in Fig. S8). These results might be surprising at first, since tensin competes with talin for the same binding site to the cytoplasmic β-tail of integrins, and thus believed to act as integrin inactivator, as the reviewer indicates. Nevertheless, recent data has shown that tensin is capable to activate integrins (in particular if β<sub>1</sub> is phosphorylated) by interacting with the actin cytoskeleton, providing mechanical coupling for integrin activation (Georgiadou & Ivaska, Trends Cell Biol. 2017). We have now included these new data as Fig. S8 in the revised manuscript. Additional experiments, which in our opinion fall outside of the scope of this work, would be necessary to identify other potential integrin inactivator partners, but certainly a topic of future interest to our group.

      (3) While the methods are described in sufficient detail, it is important to ask if the findings are based on sufficient data. Table S5 provides detailed information about the number of samples studied, and it appears that only small numbers of samples were investigated for certain protein pairs. This should be discussed, and perhaps more data should be obtained to strengthen the data.

      We have now performed additional experiments using DNA-PAINT as alternative super-resolution imaging technique (as also requested by reviewer 3) which adds additional data to the whole manuscript.

    1. eLife Assessment

      This important study compares how different classes of drugs act on the SARS-CoV-2 main protease, a key antiviral target, and shows that many of them work by controlling whether the enzyme assembles into its active dimeric form. The evidence, based on a range of complementary biophysical methods, is convincing and points to the interface between the two protein protomers, including a newly found binding site, as a promising target for broad-spectrum antiviral drugs. This work will be of interest to biochemists and virologists working on treatments for coronaviruses.

    2. Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 Mpro enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the Mpro monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how Mpro inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      None. The requested mutagenesis data have been provided in the revised manuscript, and all of my previous concerns have been satisfactorily addressed.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents a sophisticated investigation into the mechanisms by which different inhibitor classes affect the SARS-CoV-2 main protease (Mpro), a pivotal antiviral drug target. This study reveals that effective inhibition can be achieved by modulating the stabilization of the essential dimeric state. It also indicates the dimer interface could be a druggable allosteric site, which may offer a strategy for developing broad-spectrum anticoronaviral agents.

      Strengths:

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      Comments on revised version:

      The authors have very nicely addressed most of the previous comments raised. But one comment remains to be clarified relating to original point 5 and the authors' response:

      "We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296-306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation."

      Two questions remain for the statements in line 376-380. First, while it is stated "residues 296-304 in the C-terminal region of Mpro were more flexible upon ebselen binding", the segment of 296-306 is shown Figure 4c. Second, the HDX change for this segment upon ebselen binding is very subtle in the figure (in contrast to the significant HDX change of the same segment in the protein upon PF-07321332 binding), thus making the strong conclusion that "This suggests that ebselen targeting C300 may induce structural changes in the C-terminal helical segment, weakening key hydrogen bonds at the dimer interface and ultimately inhibiting activity" not convincing. The reviewer would suggest the authors either delete this conclusion or largely tone it down.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 M<sup>pro</sup> enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the M<sup>pro</sup> monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how M<sup>pro</sup> inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      The major concern is the absence of mutagenesis data to support the proposed inhibitory mechanisms, particularly regarding the role of the inhibitor binding site.

      We thank the reviewer for the recognition and comments. We agree that mutagenesis is critical for validating the proposed role of C300. We therefore generated the C300S and C300F mutants and characterized their oligomeric states and proteolytic activities. C300S was designed to remove the reactive thiol group while minimally affecting M<sup>pro</sup> structure and dimerization. C300F was introduced to mimic the steric perturbation associated with C300 modification and assess its impact on M<sup>pro</sup> dimerization. Native PAGE showed that WT and C300S M<sup>pro</sup> predominantly formed dimers, whereas C300F was mainly monomeric. Consistently, C300S retained approximately 70% of WT activity, whereas C300F retained only approximately 10%. Because C300F itself strongly disrupted dimerization, C300S was used as the principal mutant to evaluate the specific contribution of the C300 thiol to ebselen action. Native MS showed that ebselen could still bind to both monomeric and dimeric C300S M<sup>pro</sup> but did not markedly shift its monomer-dimer equilibrium toward the monomeric state. In parallel, ebselen reduced WT activity to approximately 53% of the untreated control, whereas C300S retained approximately 78% activity at the same 1:3 M<sup>pro</sup>-to-ebselen molar ratio. These results provide experimental support for the contribution of C300 to ebselen-induced dimer destabilization and functional inhibition, while the residual binding and inhibition observed for C300S suggest the involvement of additional C300-independent interactions. The corresponding revisions have been made to Methods (Lines 627–648), and Results (Lines 397–453) of the manuscript, together with the newly added figures (Figures S14–S16).

      Reviewer #2 (Public review):

      Summary:

      This is a mechanistic study that provides new insights into the inhibition of SARS-CoV-2 M<sup>pro</sup>.

      Strengths

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      We thank the reviewer for the positive comments and recognition of our study.

      Weaknesses:

      The primary weaknesses relate to linking the biophysical observations more directly to functional enzymatic outcomes and providing more quantitative rigor in some analyses. While the study is overall strong, addressing its weaknesses and limitations would elevate the impact and translational relevance of the current manuscript.

      We thank the reviewer for these comments, which have helped to iM<sup>pro</sup>ve the quality and impact of our manuscript.

      (1) Correlation with Functional Activity:

      The most significant gap is the lack of direct enzymatic activity assays under the exact conditions used for MS and HDX. While EC50 values are listed from literature, demonstrating how the observed dimer stabilization (by peptidomimetics) or dimer disruption (by ebselen) directly correlates with inhibition of proteolytic activity in the same experimental setup would solidify the functional relevance of the biophysical observations. For instance, does the fraction of monomer measured by native MS quantitatively predict the loss of activity? Also, the single inhibitor concentration used in each MS experiment needs to be specified in the main text and legends. A discussion on whether the inhibitor concentrations required to observe these dimerization effects (in native MS) or structural dynamics (in HDX-MS) align with EC50 values would be helpful for contextualizing the findings.

      We thank the reviewer for these important points. To link the biophysical observations more directly to function, we compared the oligomeric states and proteolytic activities of WT, C300S, and C300F M<sup>pro</sup>. C300F was predominantly monomeric and retained only approximately 10% of WT activity, whereas C300S remained predominantly dimeric and retained approximately 70% activity. We further evaluated ebselen inhibition using a matched 1:3 M<sup>pro</sup>-to-ebselen molar ratio. Ebselen reduced WT activity to approximately 53% of its untreated control but reduced C300S activity only to approximately 78%, demonstrating that removal of the C300 thiol significantly attenuated the functional effect of ebselen. These data support a relationship between C300-dependent dimer destabilization and reduced proteolytic activity. The Methods (Lines 627–648), and Results (Lines 397–453) have been revised accordingly, with Figures S14–S16 newly added, in the revised manuscript. We did not expect a linear relationship between the monomer fraction measured by native MS and enzymatic activity loss, because ebselen can modify multiple cysteine residues, and individual modification events may have distinct effects on M<sup>pro</sup> dimerization and catalytic function. The concentrations and molar ratios used in the native MS, HDX-MS, and activity assays have now been stated in the figure legends. The ebselen concentrations used for native MS and HDX-MS were optimized for biophysical characterization and comparison, and therefore, these concentrations might not be directly related to their IC<sub>50</sub> or EC<sub>50</sub> values. In these experiments, ebselen was applied at a 3-fold molar excess relative to M<sup>pro</sup>, consistent with the enzymatic assay. The observed dimer disruption and conformational changes were consistent with functional inhibition, supporting their mechanistic relevance.

      (2) For the two Cys residues found to be targeted by ebselen, what are their respective modification stoichiometry related to the ebselen concentration? Especially for the covalent binding site C300, which is proposed in this study to represent a novel allosteric inhibition mechanism of ebselen, more direct experimental evidence is needed to support this major hypothesis. Does mutation or modification of C300 affect the M<sup>pro</sup> dimerization/monomer equilibrium and alter the enzymatic activity? If ebselen acts as a covalent inhibitor linked to multiple Cys, why is its activity only in the μM range?

      We thank the reviewer for the insightful comments. Our LC-MS/MS data identified C44 and C300 as ebselen-modified residues, but they do not permit reliable site-resolved occupancy measurements because modified and unmodified peptides can differ in digestion efficiency and MS response. We have therefore clarified that these data provide qualitative site identification rather than absolute modification stoichiometry. To obtain direct functional evidence for C300, we generated C300S and C300F mutants. C300S preserved dimer formation and substantial activity, whereas C300F was mainly monomeric and showed severe activity loss. Importantly, although ebselen-bound C300S species were still detected by native MS, ebselen did not markedly redistribute C300S toward the monomeric state, and its inhibition was reduced from approximately 47% for WT to approximately 22% for C300S. These results indicate that C300 is an important contributor to ebselen-induced dimer disruption, while residual binding and inhibition indicate additional reactive sites. Corresponding revisions have been made to the (Lines 627–648), and Results (Lines 397–453) of the manuscript, together with the newly added figures (Figures S14–S16). The moderate micromolar potency of ebselen is consistent with its heterogeneous, multi-site covalent reactivity: modification occupancy and functional consequence are site-dependent, and not every adduct produces complete inhibition.

      (3) For the allosteric inhibitor pelitinib with low-μM activity, no significant differences in deuterium uptake of M<sup>pro</sup> were observed. In terms of the binding affinity, what is the difference between pelitinib and ebselen? Some explanations could be provided about the different HDX-MS results between the two non-peptidomimetic inhibitors with similar activities.

      We agree with the reviewer that the absence of significant HDX changes for pelitinib requires clarification. Different from ebselen that forms covalent bond with multiple cysteine residues of M<sup>pro</sup>, which could lead to sustained conformational changes that are more readily detected by HDX-MS, pelitinib non-covalently binds M<sup>pro</sup> and might not induce significant perturbations in backbone dynamics that are detectable at the peptide level by HDX-MS. These points have been integrated into the revised manuscript (Lines 333-337).

      (4) Native MS Quantification: 

      The analysis of monomer-dimer ratios from native MS spectra appears qualitative or semi-quantitative. A more rigorous and quantified analysis of the percentage of dimer/monomer species under each condition, with statistical replicates, would strengthen the equilibrium shift claims. For native MS analysis of each inhibitor, the representative spectrum can be shown in the main figure together with quantified dimer/monomer fractions from replicates to show significance by statistical tests.

      We thank the reviewer for the suggestion. We have performed a quantitative analysis of the monomer-dimer equilibrium based on triplicate native MS measurements for each condition. Representative spectra, quantified monomer/dimer ratios, and statistical analyses have been added to Figures 1 and S3. The quantitative results have also been described in the Results section (Lines 158–161, 165-168, 172-174, 177-179, 199-200).

      (5) Changes of HDX rates in certain regions seem very subtle. For example, as it states 'residues 296-304 in the C-terminal region of M<sup>pro</sup> were more flexible upon ebselen binding (Figure 4c)', the difference is barely observable. The percentage of HDX rate changes between two conditions (with p values) can be specified in the text for each fragment discussed, and any change below 5% or 10% is negligible.

      We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296–306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) The study lacks validation through inhibitor binding site mutagenesis assays, especially peptidomimetic inhibitor PF-07321332 and ebselen, which would strengthen the mechanistic conclusions.

      We appreciate this suggestion. For PF-07321332, the inhibitor forms a covalent interaction with the catalytic residue C145 and inhibits M<sup>pro</sup> activity through a distinct mechanism. Previous studies have shown that mutation of C145, such as C145A, completely abolishes M<sup>pro</sup> catalytic activity (Bhandari, D, et al. Communications Biology 2025, 8, 1061), making it difficult to directly evaluate the contribution of this residue to inhibitor-induced inhibition using enzymatic assays alone. This limitation and the relevant literature have now been discussed in the revised manuscript (Lines 272–277). Therefore, we focused on C300-dependent regulation of ebselen, which represents a distinct inhibitory mechanism involving modulation of M<sup>pro</sup> structural dynamics and dimer stability.

      To validate the role of C300 in ebselen-mediated regulation of M<sup>pro</sup>, we generated C300S and C300F mutants and performed additional biochemical and structural characterization. The enzymatic assay showed that the C300F mutation significantly affected M<sup>pro</sup> activity, and the inhibitory effect of ebselen on C300S M<sup>pro</sup> was markedly reduced compared with WT M<sup>pro</sup>. Furthermore, native MS analysis demonstrated that ebselen could still bind to C300S M<sup>pro</sup> but failed to induce a significant shift in the monomer-dimer equilibrium observed for WT M<sup>pro</sup>. These results indicate that C300 is not the only site involved in ebselen binding but is critical for mediating ebselen-induced structural perturbation and dimer destabilization. The manuscript has been revised accordingly for the (Lines 627–648), and Results (Lines 397–453), with new figures (Figures S14–S16) included, further supporting the functional contribution of C300 in ebselen-mediated M<sup>pro</sup> regulation.

      (2) MR6-31-2 is an ebselen derivative and exhibits a lower EC50 (1.78 μM) compared to ebselen (4.67 μM). It would be helpful to discuss why their activities differ, probably based on the assay conditions or binding behavior.

      We agree with the reviewer that the difference in antiviral activity between MR6-31-2 and ebselen requires further clarification. The lower EC<sub>50</sub> of MR6-31-2 may result from iM<sup>pro</sup>ved cellular properties, including compound stability, permeability, intracellular exposure, and potentially altered interactions with M<sup>pro</sup> and/or iM<sup>pro</sup>ved cellular properties. Although MR6-31-2 shares the ebselen scaffold, the modified chemical structure may affect its binding behavior and biological activity. However, EC<sub>50</sub> values obtained from cellular assays cannot directly reflect the biochemical inhibition potency against purified M<sup>pro</sup>. These points have been integrated into the revised Introduction (Lines 98–101).

      (3) In Figures 2, S1, S2, S4, S6, and S11, adding the drug name under each panel would make the data much clearer for readers.

      The corresponding drug names have been added to panels to iM<sup>pro</sup>ve figure clarity.

      Minor points:

      (1) Line 62-63 refers to the "long linker loop," while Figure 1a labels it as the "long loop linker." Please keep this consistent.

      The terminology has been unified as “long loop linker” throughout the manuscript.

      (2) Table 1 should be cited at line 80, and PDB code 7BAK should be included in Table 1.

      PDB code 7BAK has been included in Table 1, and Table 1 has been cited in the context, as suggested.

      (3) Figure 1a should include the corresponding PDB code in the figure legend.

      The corresponding PDB code has been added to the Figure 1a legend, as suggested.

      (4) It would be helpful to indicate in Figure 1a that the upper structure represents the dimer and the lower structure represents the monomer.

      The upper and lower structures in Figure 1a have been indicated as dimeric and monomeric M<sup>pro</sup>, respectively, as suggested.

      (5) In the Figure S1 legend, it should mention that some inhibitor structures (like ebselen and MR6-31-2) are not fully resolved. Also, the Se atom in ebselen should be shown in Figure S1f (PDB: 7BAK).

      The Figure S1 legend has been revised to indicate that some inhibitor structures, including ebselen and MR6-31-2, are partially unresolved, and the selenium atom of ebselen has also been shown in Figure S1f, as suggested.

      (6) Pelitinib is an allosteric, non-covalently binding inhibitor. However, in Figure S3, the native MS profile shows dimer species (13+ to 15+) compared with unbound M<sup>pro</sup> (14+ to 17+). Please clarify this difference.

      We thank the reviewer for raising this good point. Protein charge-state distributions can be influenced by solution-phase conformation, conformational flexibility, solvent properties, and electrospray droplet charging (Susa AC, et al. J Am Soc Mass Spectrom 2017, 28, 332-340). The observed shift in charge state distribution in native MS might suggest that the addition of pelitinib caused changes in the protein conformation, solvent property and electrospray droplet charging. The relevant literature and discussion have been added in the revised manuscript (Lines 200–204).

      (7) Line 172: "S1are" should be corrected to "S1 are."

      Corrected.

    1. eLife Assessment

      This is a fundamental study on the sensory roles of cerebrospinal-fluid-contacting neurons (CSF-cNs) in mammals, revealing how the apical extension is used as an amplifier of chemical changes in content of the CSF. Specifically, the authors show compelling evidence that PKD2L1 is predominantly a pH-sensing channel in CSF-cNs and link its apical localization to dual phasic and sustained responses underlying CSF chemosensation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

    3. Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments-relates to their multimodal CSF-sensing function remains unclear.

      The authors confirm in the GATA3:GFP mice where these cells are labeled that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding on how these sensory cells can detect so finely the changes of pH in the CSF.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Comments on revised version:

      I thank the authors for their extensive revisions and detailed responses to the reviewers' comments. The manuscript has been substantially improved, and most of the major concerns raised in the initial review have been adequately addressed. In particular, the additional analyses of PKD2L1 channel activity, the incorporation of physiologically relevant pH conditions, the clarification of ASIC involvement, and the expanded Discussion have significantly strengthened the study.

      Major scientific concerns largely addressed:

      Quantification of PKD2L1 channel activity

      The authors appropriately addressed my previous concerns regarding the use of Po as the sole measure of channel activity. The inclusion of additional parameters such as apparent Po, open time, nmax, holding current, and membrane charge provides a more robust assessment of PKD2L1 activity and substantially strengthens the conclusions.

      Physiological relevance of pH modulation

      The inclusion of experiments at pH 6.5 and the additional analyses of holding current and resting membrane potential are valuable additions. These experiments considerably improve the physiological relevance of the study.

      ASIC contribution

      The additional pharmacological experiments using ASIC blockers are helpful and support the conclusion that the photolysis-evoked response in the apical process is predominantly mediated by PKD2L1 channels.

      Functional implications

      The expanded Discussion regarding Ca2+-dependent signaling, neurosecretion, and the potential physiological roles of CSFcNs considerably improves the manuscript.

      Remaining concerns:

      Continued overstatement regarding "exclusive" localization and function:

      Although the authors softened some statements in the revised manuscript, the term "exclusive" remains in several key locations, including the title.

      For example:

      "PKD2L1 channels segregated to the apical compartment are the exclusive dual-mode pH sensor..."

      The data clearly demonstrate strong enrichment of functional PKD2L1 channels in the apical process. However, the available evidence does not fully justify the term "exclusive," particularly because:

      - PKD2L1 immunoreactivity is still detectable outside the apical process.

      - ASIC-mediated responses are present in CSFcNs.

      - The authors themselves use more appropriate terminology such as "predominantly located" in the Discussion.

      Therefore, I recommend replacing "exclusive" with more conservative terminology such as:

      - predominant

      - predominantly localized

      - enriche

      - functionally segregated

      throughout the manuscript, including the title, Abstract, Introduction, Results, and Discussion.

      We agree with the reviewer that the world “exclusive” is misleading and should be replaced. Following the reviewer’s suggestions, we have deleted the word “exclusive from the title, which now reads: “PKD2L1 channels segregated to the apical compartment are the functional dual-mode pH sensors in cerebrospinal fluid-contacting neurons.”

      In addition, the word “exclusive” has been changed with more conservative terminology in other parts of the text: lines 80, 420, 466 and 551.

      Use of the term "tonic current"

      The manuscript continues to use the term "PKD2L1 tonic current."

      While the dibucaine-sensitive holding current is clearly present, the precise mechanism generating this current remains uncertain. Indeed, the authors themselves acknowledge in the Discussion that:

      - an alternative conducting state may exist, or

      - unresolved brief channel openings may account for the current.

      Therefore, the data support the existence of a sustained PKD2L1-associated current, but do not yet definitively establish a distinct tonic gating mode of the channel.

      I therefore recommend replacing:

      "tonic current" with a more neutral expression such as:

      - sustained current

      - PKD2L1-associated holding current

      - sustained PKD2L1-mediated current throughout the manuscript.

      Continued use of "off-current" and "off-response":

      The revised manuscript has improved considerably in this regard. However, the terms "off-current" and "off-response" still remain in portions of the text and figure legends.

      Because the manuscript itself demonstrates that the response reflects recovery from transient acidification rather than a separate OFF signaling mechanism, these terms remain potentially misleading.

      I recommend replacing them with terminology such as:

      - photolysis-evoked PKD2L1 current

      - recovery current

      - proton-removal-induced current

      throughout the manuscript, including figure legends.

      We apologize, as the word “tonic” and the terminology “off-current” should have completely disappeared after the first round of revisions. We have now replaced those all along the text. “Tonic” has been replaced by “sustained”.

      “Off-current” or “off-response” have been replaced by appropriate terms in lines: 339, 342, 544, 545, 546, 549, 552, 555, 556, 560, 576, 580, 585, 588, 807 and 929. We have nevertheless conserved the term “off-current” in line 552 as we are referring to terminology used by other authors.

      Minor editorial corrections

      Figure 1Bd Please change: "po" to "Po" for consistency with standard channel physiology nomenclature.

      Figure 1Ca Please add units (mV) to the voltage labels shown on the left side of the traces.

      Figure 3E Please change: "Norm po" to "Norm Po".

      Figure 4Fb Please replace: "sec" with "s" to conform with SI unit conventions.

      Done.

      The authors have addressed the majority of my previous concerns and the manuscript has been substantially improved. The remaining issues are primarily related to terminology and overinterpretation rather than experimental deficiencies.

      Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments-relates to their multimodal CSF-sensing function remains unclear.

      The authors confirm in the GATA3:GFP mice where these cells are labeled that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding on how these sensory cells can detect so finely the changes of pH in the CSF.

      Weaknesses:

      Not weaknesses, there are only minor requests to complete the beautiful study.

      (1) The apical extension's response to removal of acidification is nicely illustrated in Figure 4C,G. There's something puzzling there: while the response to Glutamate is immediate, the channel responses to H+ is extremely delayed by 100ms - 2s, and even sometimes came in bursts separated by few hundreds of ms. H+ diffuse even faster than glutamate. Why is that?

      I don't quite understand how the response is so delayed & how to explain the recurring bursts of channel opening in the figure panel ?

      The kinetic of the response to proton uncaging is analyzed in Figure 4E, where the charge of the current traces is plotted against time. What this analysis shows is that the response lasts a few hundred ms (τ 250 ms) and then the PKD2L1 activity increase subsides to baseline. The peak of the response is at 100 ms (Figure 4G), but the increase in activity happens as soon as the uncaging pulse ends (Figure 4D, G and H). This behavior has already been shown in expression systems, where the channel activity is blocked by protons and the blockage is released when the acid is withdrawn. In an intact cell as the CSFcNs studied here, the exact kinetics of the recovery response are probably more complex (and variable) than in expression systems. Indeed, it is known that the recovery of this current depends, for example, on pH and extracellular calcium. Also, PKD2L1 are inhibited by intracellular calcium (de Caen et al, eLife 2016) but are themselves permeable to Ca<sup>++</sup> ions. The interaction of these effects could give rise to the “bursts” that are observed in some cases. However, this is merely speculative at this point.

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      Following the reviewer’s suggestion, we have added a trace in Figure 4C (upper blue trace) showing the spontaneous activity of the cell, prior to uncaging, as it is already shown for another example in Figure 4D.

      - Could the authors use a fluorescent pH sensor to monitor pH in the extracellular space and in the cell ?

      This is an important point that was already addressed by the reviewing editors in the previous round of revisions. Indeed, we have attempted to perform pH calibrations in the setup using the pHsensitive dye pyranine (or HPTS: 8-Hydroxypyrene-1,3,6-trisulfonic acid). HPTS is a very useful tool for pH calibrations in the physiological range: its pKa value is close to 7.2 and it can be used as a ratiometric dye (its fluorescence is pH-independent at 405–410 nm and pH-dependent at 450 nm). Unfortunately, the calibration under the conditions of a real experiment is not possible because the photolysis in the slice occurs in a tiny volume (approximately 1 µm³ in a total bath volume of more than 1 ml). In these conditions, the 405 nm uncaging pulse bleaches the dye in the photolysis spot and any useful information is lost. In addition, our imaging system is not fast enough to follow the pH change. As discussed in the Materials and Methods section, subsection “Estimation of the pH drop induced by photolysis” (line 791), the fast protonation of bicarbonate indicates that the pH change induced by the photolysis recovers in the submillisecond range.

      - Could the authors investigate whether in the apical extension, PKD2L1 channels are mainly at the outer membrane in the apical extension OR whether many channels are located in inner membranes ?

      PKD2L1 channels are probably subject to a high rate of turnover, and they are certainly localized in the plasma membrane of the apical process as well as in the inner membranes. Although this is a very interesting point, we believe it is out of the scope of this work.

      (2) Suppl Fig 4 is very cool and should be moved to main figure. The coupling of Soma and AP is very tight, yet there is a clear difference in targeting of channels that respond to cues in the CSF. In the context of an intact spinal cord, we can wonder how and when the contribution from ASIC in the some would be relevant to physiology. Can the authors think of experiments with an intact central canal to test the sensitivity and condition of recruitment of pH sensing in the soma (ASIC) versus the apical extension (PKD2L1)?

      We have followed the suggestion of the reviewer and have made Supplementary Figure 4 a main figure.

      The fact that the normal interphase between the spinal cord parenchyma and the cc is lost is already acknowledged in the discussion, lines 486 to 489. As the reviewer suggests, PKD2L1 and ASIC channels seem both to be important in the response of CSFcN to pH changes. However, both channels are activated in very different physiological contexts, as is discussed in the section “The involvement of ASIC channels”. Keeping the central canal intact in order to be as close as possible to physiological conditions, as suggested by the reviewer, would be ideal. However, as CSFcNs are in the middle of the cord, it would require the use of optical techniques that allow to penetrate deep into the tissue (e.g., 2-photon microscopy) that unfortunately are not available in our labs.

      (3) The Reissner fiber is missing after slicing the spinal cord. From our observations in fish, the fiber being under tension triggers lots of activity in CSF-cNs (Bellegarda et al Elife 2023) that also relies on PKD2L1 (Bohm et al NC 2016; Sternberg et al NC 2019). Could the authors discuss the contribution of the Reissner fiber to the PKD2L1 mediated modulation of CSFcN excitability ? Could the authors conceive a way to slice along the anteroposterior axis (sagitally) the spinal cord to keep the Reissner fiber in the central canal when recording CSF-cN apical extension ?

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      As discussed in the previous point, the in vitro slice preparation has technical limitations that are mainly related to the alterations of the normal structure of the tissue. Although keeping the Reissner fiber intact in a sagittal slice seems possible, accessing the CSFcNs with electrophysiological methods would still be a challenge.

      We have now added a sentence in the Discussion, lines 561 to 564, where we discuss that CSFcN excitability is modulated by the Reissner fiber and that it remains to be explored whether in rodents the gating of PKD2L1 channels is modulated by the Reissner fibre, as has been shown in zebrafish.

    1. eLife Assessment

      This important study reports that neural activity in the auditory cortex (field L) of singing male zebra finches can be modulated by the presence of a female conspecific. These findings extend recent work showing that the activity of dopaminergic neurons in songbirds is also affected by an audience. Solid evidence is presented for the importance of singing context in modulating auditory processing during vocal production, but the study does not yet fully establish that these effects arise specifically from audience-dependent modulation of auditory feedback, as opposed to possible acoustic, temporal, or recording-related confounds. This work should be of interest to researchers studying the context dependence of sensory processing, the role of auditory feedback, and vocal communication during courtship behavior.

    2. Reviewer #2 (Public review):

      This study asks whether auditory responses in the songbird auditory pallium/field L during singing are modulated by social context. Specifically, the authors examine neural responses to delayed auditory feedback during male zebra finch song produced either alone or in the presence of a female. This is an interesting and important question because the evaluation of self-generated vocal output may differ when the song has a dedicated social function.

      The main strength of the work is that it addresses auditory feedback processing during natural vocal behavior and does so across two naturalistic contexts. The revised manuscript is strengthened by additional analyses of spike waveform similarity, response significance, response latency stability, and exclusion of motifs overlapping with female calls. These additions make the reported context-dependent response differences more credible and help address some concerns about recording stability and contamination by female vocalizations.

      The results show that some auditory pallium neurons respond differently to feedback perturbations during directed and undirected song. This finding is potentially significant because it suggests that auditory processing during vocal production is not rigid but could subserve a social function that depends on the listener. If robust, this would add an important dimension to models of song monitoring and sensorimotor control.

      However, the strength of evidence remains moderate rather than conclusive. Several alternative explanations are not fully ruled out. Directed and undirected songs may differ acoustically in ways that could influence neural responses, and it is not yet clear that relevant song features were directly compared or controlled across contexts. The experimental sequence also appears to be ordered, with undirected song recorded before directed song, which makes it difficult to fully separate social-context effects from time-dependent changes in recording quality or neural responsiveness. The added waveform analysis is useful, but does not completely establish continuous unit stability across long recording sessions. In addition, possible song changes around the delayed-feedback target point, including compensatory modifications before or after feedback, remain an important potential confound. Finally, clarification of the time-warping and spike-alignment procedures is important because condition-specific alignment could affect comparisons between directed and undirected song.

      Overall, the data support context-dependent differences in neural responses in some neurons, but do not yet fully establish that these differences arise specifically from audience-dependent modulation of auditory feedback processing rather than from acoustic, temporal, or recording-related confounds. The work is likely to be useful to researchers interested in vocal communication, auditory feedback, and social modulation of sensorimotor processing, particularly as a foundation for future experiments using counterbalanced designs and more direct controls of song structure across contexts.

    3. Reviewer #3 (Public review):

      In this study, Jones et al. examine how neural activity in auditory regions (the auditory pallium) of singing male songbirds is modulated by the presence or absence of an audience (a female conspecific). They test whether activity in auditory pallium differs between conditions in which the male is singing to a female (directed song) or alone (undirected song) and whether response to distortions of auditory feedback (DAF) differ between these conditions. Previous work has shown that in other parts of the songbird brain, sensory-motor activity can differ between directed and undirected song, and that responses to DAF are attenuated when males sing directed song versus undirected song. These prior results raise the interesting question of the extent to which such modulations of activity by the presence of an audience are already present in primarily auditory areas within the pallium. This possibility is also motivated by prior work that has shown that activity in the auditory pallium is not exclusively explained by auditory input, but can also be modulated by the bird's state - whether it is singing or not.

      Against this background, the questions asked here are of interest for two inter-related reasons:

      (1) The authors address whether the presence of an audience (a female conspecific) alters activity in an auditory region during singing. Primary songbird auditory areas such as Field L, and analogous mammalian thalamo-recipient cortical regions such as A1, are often thought of as responding very specifically to the features of sensory stimuli, but are also understood to be modulated by a variety of factors including the attentional and behavioral state of the animal. For audition, such modulation includes whether or not animals are vocalizing and listening to themselves or listening to playback of their own vocalizations. Cited works from Keller (2009) as well as Eliades and Wang (2008) have indicated that the act of vocalizing can modulate auditory responses to self-generated feedback in primary auditory areas relative to those arising from playback of the same sounds. Here, the question is whether responses to self-generated feedback differ between conditions of singing alone versus singing to a female audience. A demonstration that the presence of an audience matters to responses in auditory pallium would add to a general understanding of how it is that non-auditory factors can modulate activity within regions that are considered primarily sensory.

      (2) The authors address the possible source of an audience-dependent modulation of responses to feedback perturbation in the VTA previously reported by Goldberg and colleagues (2023). In the VTA, responses to perturbations during singing are consistently attenuated when males are singing to females versus when they are singing alone, but the underlying mechanisms of this modulation are unknown. Here, the authors test the possibility that such modulation by an audience is already present at the level of auditory pallium. The previously reported attenuation in VTA is a nice example of how neural processing can differ with varying behavioral priorities. Understanding whether this modulation of responses to DAF arises already in auditory areas would further a mechanistic understanding of an intriguing example of state-dependent modulation of sensory processing and behavior and lend broad insight into related phenomena.

      The authors report 1) that activity in the auditory pallium differs between directed and undirected singing at many individual recording sites, but that these changes are heterogeneous, with both increases and decreases in activity, so that there is no consistent change across the population and 2) that modulation of activity by DAF can differ between directed and undirected song, but that there is no consistent attenuation of response (as observed in the VTA) and instead heterogeneous increases and decreases in response to DAF so that there is no net change at the population level.

      These findings are important and of general interest; while they do not readily explain the source of the audience-dependent attenuation of auditory responses to DAF in the VTA, the demonstration of audience-dependent modulation of self-generated feedback and its disruption in the auditory pallium provides an opportunity for further investigation of how changes in social context influence brain and behavior.

      Additional comments and suggestions:

      The authors have done a good job of addressing many of the issues that were raised in the initial round of reviews. There is additional analysis that strengthens the study, including 1) applying a stability criterion to assess the quality of unit isolation, 2) shifting away from a categorical identification of units as "retuning" or not, to an analysis that presents a continuum of changes to neural firing between conditions, 3) use of non-parametric statistics for the assessment of significance of differences in response measures between conditions and 4) exclusion from analysis data from motifs during which females were observed to be vocalizing.

      The authors also have added to the text several important clarifications, and modified language in several ways that improve the presentation and interpretation of results. This includes 1) noting that differences between neural activity during singing with and without DAF does not necessarily reflect "error detection" but could instead reflect how neurons with fixed auditory receptive fields might respond differently to the distinct auditory inputs present between these conditions, 2) clarifying that the recordings were not specifically restricted to Field L, but were distributed more broadly across the auditory pallium, and 3) discussing some of the mechanisms whereby tuning might change due to various sensory, motor and internal factors associated with differences between singing alone and singing to a female.

      I only have a couple of areas of remaining concern that I think could be addressed with further analysis, or some additional discussion, according to the authors' preferences.

      (1) Stationarity of neural response

      My main residual concern relates to the issue raised in the previous round of review of how much of the observed difference in activity between morning sessions when the male is alone and later sessions when the male is singing to a female could reflect changes in neural response properties (non-stationarity) due to the passage of time (sometimes at least several hours) rather than specifically due to the presence or absence of an audience.<br /> The authors restriction of data to recordings that passed a stability criterion for unit waveforms is helpful in addressing whether the same units are 'held' over the course of the experiment. However, even with well isolated units, the response properties or tuning of units can change over time due to a variety of factors that include changes in internal state, circuit excitability, up and down states, neural plasticity, etc.

      The previous review noted several examples of data from the manuscript that illustrated this concern - instances where response properties of neurons appeared to change over time within a given condition. Any such changes in response properties that occur in the absence of a change in audience would tend to contribute to the reported "retuning" of responses.

      One thing that the authors could do to address this issue would be to discuss potential contributions of non-stationarity of responses over time as a potential confounding variable and then editorialize about why they think this seems unlikely to explain many cases in which response properties change between conditions. See comments to authors for one specific suggestions along these lines.

      Alternatively, the authors could carry out additional analyses to evaluate this issue more quantitatively. For example, by measuring the magnitude of "spontaneous" changes in responsiveness observed within conditions (such as by comparing the motif aligned activity for the first n examples within a condition against the activity during the last n examples) and comparing that with the magnitude of changes observed across conditions.

      Another approach would be to carry out some sort of "change point analysis" on the motif-related activity for each experiment in order to establish how often the most abrupt changes in activity occur at the transition between conditions versus spontaneously at other times.

      Lastly, while the experimental design didn't specifically include interleaved blocks of undirected (alone) singing and female directed singing, the methods indicate that the female directed singing data were collected by repeatedly introducing females for 10 minutes at a time. If there are even a couple of cases where the males produced song alone in the periods between female presentation, it would be worth testing whether modulation of neural firing tracked these interleaved conditions.

      Previously published work such as the interleaved recordings of Hessler and Doupe indicate a close and reversible tracking between modulation of neural activity in sensorimotor song system nuclei and switches between singing alone and singing to a female. With respect to the possibility raised in the rebuttal of whether males continue to sing 'female directed song' even after the removal of a female, these and other published data suggest that this is not likely to be the case. But if this were a concern in interpreting any data, the authors could directly assess the male's song for previously described changes in acoustic variability that also track changes in the presence of an audience.

      (2) Further discussion of how an audience might influence responses.

      With respect to the mechanisms whereby an audience might modulate neural responses, a somewhat expanded discussion of possibilities with reference to relevant literature would be helpful. This could include reference to evidence for various neuromodulatory systems participating in modulating singing related activity in song system nuclei based on presence or absence of a female - do these neuromodulatory systems project to the relevant regions of the auditory pallium or its lower-level inputs within the ascending auditory pathway such that they could concurrently act on auditory circuitry?

      In addition to possibility that the presence or absence of an audience affects auditory circuitry via a change in attention, alertness, or motivation, might efference signals associated with singing or locomotion/dancing reach and influence auditory pathways? Given that premotor activity and acoustic structure of song differ between conditions, might any singing-related efference copy activity that reached auditory regions also differ between conditions? A related interesting possibility that could be worth noting is that the presence of a female generally elicits increased locomotion and dancing on the part of the male that accompanies female directed song. Several studies have noted that general forms of locomotion can also result in efference copy signals reaching and influencing auditory regions (e.g. see Schneider and Mooney, Annual Review, 2018; Han et al. "Locomotion-induced neural activity independent of auditory feedback in the mouse inferior colliculus" iScience 2026 - the latter reference is interesting in that it appears to indicate bi-directional modulation of neural activity as observed across units in the current study).

      Minor:

      (1) The authors describe some units as showing "Activation by the absence of DAF." Because the birds in the study have extensive experience with DAF on a subset of trials, it is possible that the increased responses in the absence of DAF reflect a positive deviation from expectation of distortion (as seems to be the case for VTA neurons in previous work). But it is also possible that the broadband DAF stimulus drives inhibition of auditory responses in some cases, and the greater responses in the interleaved trials with normal feedback simply reflect the absence of that inhibition (rather than a positive deviation from a learned expectation). In keeping with the authors shift away from the use of "error detection" elsewhere in the manuscript, it might also be good to use less interpretive language here instead of "activation by absence of DAF".

      (2) At the authors discretion, it would be interesting to know if there is any relationship between the way in which changes in audience affect activity with normal auditory feedback versus with DAF. For example, if normal singing responses are attenuated in the female directed condition, are the responses to DAF also attenuated?

      (3) In figure 2, the vertical dashed lines associated with the rasters indicate the onset and offset of motifs. For several of the figure panels, the spectrograms show motifs that are not aligned with these onsets and offsets. Please clarify or modify (are the rasters from time-warped data but the spectrograms are not -time warped?).

      (4) The authors equate peaks in activity before the onsets of motifs with premotor activity: ["A previous study recording from Field L in zebra finches reported neural activations prior to the onset of singing, consistent with premotor signaling (Keller and Hahnloser, 2009). We tested for context dependent changes in premotor activity by examining peaks in neural activity aligned to motif onsets. Across the population neurons did not exhibit significant changes in the timing of motif onset-aligned activity (Figure S4)."]<br /> However, the spectrograms as shown in Figures 1 and 2 indicate that each motif is often preceded immediately by other song syllables such as introductory notes or syllables from the end of the preceding motif. Further analysis would be required in the current study to demonstrate that the activity present before motif onsets reflects premotor activity rather than auditory responses to the proceeding syllables. Please soften the claim that this reflects premotor activity or provide additional analysis or argument.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study examines the context-dependent modulation of auditory cortical neurons in response to expected sensory input, either self-generated sounds or expected perturbations of self-generated sounds. Specifically, using songbirds, the authors ask whether social context (the presence of a female conspecific) affects 1) the response of auditory cortical neurons to the bird's own song when he is singing; and 2) the response of neurons to perturbations of auditory feedback that the bird has been trained to expect.

      Strengths:

      First, the authors report that across the population, the responses of the neurons does not differ when a male bird sings alone or if he sings to a female. A fraction of auditory cortical neurons, however, do show significant differences in the firing rate, precision, and/or degree of burst firing when males sing alone vs. when they sing to females. This finding is broadly consistent with the literature showing that sensory neurons (visual, auditory, somatosensory, etc.) can be rapidly reconfigured into different "information processing modes" depending on behavioral state (e.g., quiescence vs. vigilance).

      For the perturbation experiments, the authors trained birds to expect distorted auditory feedback during a particular syllable. They found that some neurons showed greater responses during perturbation when a female was present (compared to when males were alone) while other neurons had smaller responses during perturbation when a female was present. In addition, the response of a small number of auditory cortical neurons were not affected by behavioral state. These results contrast with their prior report that the responses of midbrain dopaminergic neurons that project to the basal ganglia are "uniformly reduced" in the presence of a female, raising a question of how an evaluation signal is transformed in the circuit from the primary sensory region to the midbrain.

      Weaknesses:

      While the experiments and analysis are solid, the finding that social context can alter responses of auditory cortical neurons in a multitude of ways (increase, decrease or no change) raises several questions that can be examined with additional analysis. For example, do context-dependent differences in auditory responses derive from context-dependent differences in the songs? Are context-dependent differences present in all classes of neurons and throughout the auditory system?

      The observed heterogeneity in the firing properties of auditory cortical neurons, both in response to self-generated sounds and during perturbations of auditory feedback, raises the question of which neurons are sensitive to social context (which likely can be addressed by the authors in a revision). The authors should provide additional details about the recordings:

      (a) What are the locations of the recording sites? Prior work has shown that there is an organized map of spectrotemporal features of sounds in the auditory cortex of songbirds; spectral tuning widths change along the medial-lateral axis and temporal tuning widths differ between the input and output layers of Field L. Were the recordings primarily in Field L2 (thalamo-recipient region), L1 or L3? Were some recordings lateral to Field L in secondary auditory regions? Were the neurons that showed context-dependent changes in firing properties localized or distributed throughout Field L (i.e., were the context-dependent differences in neural responses truly brain-wide)? At a minimum, the authors should include a schematic showing the different regions of Field L and a summary of the location of the recording sites. Images of the processed tissue with electrolytic lesions would also be helpful.

      We agree that the anatomical targeting and limits of localization should be made explicit. In the original manuscript, we referred broadly to recordings from "Field L" and described targeting coordinates in the Methods. In the revised manuscript, we have softened the anatomical claim from "Field L" to "auditory pallium" where appropriate, while explicitly stating that electrodes were aimed at Field L. We also added anatomical caveats and a new supplemental figure.

      The revised title and abstract now reflect this more conservative anatomical framing. For example, the abstract now states: "Here we recorded neural activity from the auditory pallium in zebra finches practicing singing alone and directing courtship songs to females." In the Introduction, we now explicitly state both the intended target and the limitation: "We targeted our recording electrodes to Field L, a primary auditory pallial area that projects into multiple higher auditory areas that, in turn, project to VTA."

      We then added the caveat: "Field L is composed of multiple subdivisions and surrounds the interfacial nucleus and because the implanted wire bundles spread in a small radius of up to ~0.5 mm, our recordings likely included large territories of the auditory pallium (Figure S1)."

      We also added mechanistic/anatomical context for why these recordings may reflect activity shaped by broader auditory forebrain circuitry: "Although Field L is classically described as a primary auditory thalamorecipient region, its activity may also be shaped by contextual signals related to the courtship context, potentially via recurrent interactions with higher-order auditory forebrain regions such as the caudal mesopallium (CM) and caudomedial nidopallium (NCM) (Bauer et al., 2008; Figure S1C)."

      Changes made in revision: We added Figure S1, which includes anatomical subdivisions, an example histological slice showing the area where the cannula was implanted, and auditory pathway connectivity. We also revised the wording throughout the manuscript from "Field L neurons" to more conservative phrasing such as "auditory pallium neurons" or "pallial auditory neurons" when appropriate. We did not claim layer-specific localization, because the revised manuscript explicitly states that we cannot make such claims.

      (b) Was the context-dependent modulation limited to a particular class of neurons (distinguished by spike waveform shape, spontaneous firing rate, or other feature)?

      We agree that identifying whether context-dependent modulation is associated with specific neuronal classes is important. In the revision, we added analyses examining relationships between spike width and various firing characteristics. We also looked for potential relationships between mean rate and DAF response, IMCC and DAF response, and found no clear trend. Overall, we did not observe any clear relationship between DAF response, DAF response modulation, and metrics like spike width or mean firing rate.

      The revised manuscript states: "Action potential width of individual neurons has previously been used to classify putative interneurons or putative principal cells in the zebra finch auditory pallium (Calabrese and Woolley, 2015)."

      We then describe the new analysis: "We tested if DAF-response scores, mean firing rates during singing, burst fraction, IMCC, and the change in all of these between undirected and directed singing was correlated with spike half width (spike half-width measured as peak-to-trough time; Figure S5)."

      The revised result is: "DAF response in either condition, the change in DAF response across conditions, and the change in firing rate, burst fraction, and IMCC were not significantly correlated with spike width (Figure S5A-B)."

      We also report that some general firing properties did correlate with spike width: "Consistent with previous literature, firing rate was significantly correlated with spike width (Pearson's correlation, p=0.003; Figure S5C). Interestingly, burst fraction (p=8.3x10-4) and IMCC (p=5.4x10-5) were also significantly correlated with spike width (Figure S5C)."

      Changes made in revision: We added Figure S5 and associated text analyzing whether context-dependent changes in DAF response, firing rate, burst fraction, and IMCC were correlated with spike half-width. These analyses did not support the conclusion that context-dependent DAF modulation was restricted to a waveform-defined neuronal class.

      (a) Prior work has shown that songs of zebra finches differ slightly when males sing alone compared to when they sing to females: songs are faster; pitch is less variable; and the number of introductory elements is greater when males sing to females. Do some of the observed social context-dependent differences in the responses of auditory neurons reflect differences in the songs in the two conditions? Did the authors of this study also find premotor activity in Field L, and if so, did it differ between the two social contexts? Might differences in Field L responses reflect motor/song differences?

      The revised manuscript now addresses the issues of context-dependent changes in song in several ways. First, we explain why motif-aligned comparisons are meaningful: "The acoustic structure of undirected and directed motifs is highly similar in adult finches, enabling singing-related neural activity to be precisely aligned and compared across contexts."

      Second, we added analysis and discussion of premotor-related activity. Changes made in revision: New results paragraph and new Fig S4. "A previous study recording from Field L in zebra finches reported neural activations prior to the onset of singing, consistent with premotor signaling (Keller and Hahnloser, 2009). We tested for context-dependent changes in premotor activity by examining peaks in neural activity aligned to motif onsets. Across the population neurons did not exhibit significant changes in the timing of motif onset-aligned activity (Figure S4)."

      (b) For the perturbation experiments, this raises a question of whether perturbation amplitude is different when a male is alone and when a female is present. It would be useful to know if (and how much) perturbation amplitude varied depending on the location inside the cage as well as whether the sound pressure level of the underlying song was higher (e.g., Lombard effect).

      We previously calibrated the perturbation amplitude in Roeser et al., 2023, in an identical recording setup. Two speakers deliver the feedback on either side of the bird's home cage. We acknowledge the possibility that the position and orientation of the bird can affect the way the sound hits either of the bird's eardrums and thus potentially affect a neural response. However, neural activations following the absence of distortion playbacks were a major feature of the dataset and were context-dependent in some cases.

      The Methods state: "DAF was implemented with a custom LabVIEW acquisition program that analyzed song syllables in real-time and delivered syllable-targeted feedback." and "DAF (50 ms broadband noise bandpass filtered at 1.5-8 kHz to match frequency range of zebra finch song) was played over speakers in the recording chamber on top of a specific target syllable randomly on 50% of motif renditions."

      The revised manuscript also makes clear that experiments occurred in the bird's home cage: "Experiments were carried out in the male's home cage, which was inside a sound isolation chamber."

      Importantly, the revised Results show that not all DAF-related responses were simple activations to additional sound. Some neurons were activated by the absence of distortion: "Unexpectedly, some neurons were not activated by the song distortion but rather by the lack of target syllable distortion." and "These activations following undistorted renditions could also depend on the courtship context."

      Changes made in revision: We clarified the DAF stimulus and recording setup in Methods and added Figure 3 showing neurons activated following undistorted renditions. These data argue that context-dependent responses are not simply explained by DAF sound amplitude, although we do not claim that position-dependent acoustic variation was fully eliminated.

      Finally, it would be helpful if the authors could include a model and/or more discussion of how the uniform attenuation in midbrain dopaminergic neurons may arise given the heterogeneous responses in Field L.

      The revised manuscript provides evidence for context-dependent retuning upstream of VTA, but does not offer a direct mechanistic explanation for the uniform attenuation seen in dopaminergic neurons. The revised Discussion states: "Because the main goal of this study was to test if courtship-associated reduction in DAF signaling, recently observed in VTA DA neurons (Roeser et al., 2023), resulted from a local process in VTA or reflected a retuning of auditory responsiveness, we explicitly tested for changes in DAF responsiveness between alone and female-directed singing."

      It then explicitly contrasts auditory pallium and VTA: "Surprisingly, we discovered that Field L neurons could retune at the transition from lone to courtship singing in diverse ways, consistent with a more widespread process in the brain that does not fully explain the uniform DAF-signal attenuation observed in VTA."

      Changes made in revision: We expanded the Discussion to explicitly state that auditory pallium retuning is heterogeneous and therefore does not fully explain the uniform attenuation observed in VTA. We do not present a formal circuit model, but we now more clearly frame the result as evidence for broader sensory retuning that is likely transformed downstream.

      Reviewer #2 (Public Review):

      Summary:

      In the manuscript, Jones and Goldberg study auditory cortex in male zebra finches. They explore song-related responses in two different contexts, when the male is either alone or in the presence of a female. They find a heterogeneity of responses, in line with auditory cortical neurons computing the social modulation of responses found in VTA.

      Weaknesses:

      Stability of responses has not been studied: some neurons seem to have responses that slowly drift in time, which could lead to observed differences between alone and with-female conditions. Also, possible motor confounds and sound-of-audience confounds should be addressed. The language is often imprecise.

      Stability and Reversal: It is a bit unfortunate that stability of effects seemingly has not been studied by reversing experimental conditions. The work would be much stronger if authors could show that audience-dependent tuning is robust in individual cells. Did they record from some neurons during reversal back to the alone condition?

      We agree that recording stability is essential. A reversal experiment was not feasible for this dataset, as it is difficult to confirm whether song motifs produced immediately following female presence represent undirected singing or are directed to an unseen but recently present female. Instead, the revised manuscript adds a strict unit-stability criterion based on waveform similarity across conditions.

      The revised Results state: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." The exact criterion is: "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."

      Changes made in revision: The strict waveform-stability inclusion criterion and new Figure S2 directly showcase unit stability across the time course of the experiments.

      Motor responses: Does DAF playback change song? If so, especially if it applies only in one of the two conditions (audience/no audience), then the observed response differences could be motor-related rather than auditory responses.

      We agree that motor confounds must be minimized. We previously found that DAF did not affect the acoustics of the subsequent syllable (Gadagkar et al., 2016). The revised manuscript clarifies that DAF and undistorted trials were randomly interleaved and analyzed by comparing matched renditions within conditions. Importantly, we only analyzed motif-aligned activity, ensuring that all syllables within the song motif are the same.

      Changes made in revision: We clarified the DAF analysis framework and added a more conservative permutation-based analysis comparing distorted and undistorted trials within each context, then comparing those DAF-response vectors across contexts. We do not claim that all possible motor consequences of DAF are eliminated, but the analysis directly tests neural responses to randomly interleaved distorted versus undistorted renditions.

      Similarly, motif-aligned spiking activity was time warped to the median duration of undirected or directed motifs. Could the shorter motifs during directed song lead to alignment differences that would account for the different error responses in alone/with-female conditions?

      We agree this is an important technical point. The time-warping we conducted, standard in the field, compensates for the tempo differences between directed and undirected song. Importantly, our main analysis of change in error response no longer uses a 100 ms response window, but rather includes all windows in the motif.

      Changes made in revision: We clarified that the revised DAF response analysis uses motif-aligned, time-warped spike trains. Importantly, the revised analysis moves away from relying on a single scalar response window and uses bin-wise permutation tests with family-wise error correction.

      Audience versus sound of audience: Is it truly the audience that causes the difference in error responses or is it the sounds the audience makes?

      We agree that the sensory cues defining "audience" cannot be fully separated in this experiment. The reviewer raises an important point that female zebra finches occasionally call at the male. We have excluded all song motifs from analyses that include an overlapping female call.

      The revised Methods now explicitly state that motifs overlapping with female calls were excluded: "Any motifs that had overlapping time with a female call in directed motifs was excluded from analysis."

      We also revised the Discussion to treat the mechanism by which auditory pallium receives information about the female as an open question: "An open question is how auditory pallium receives information about whether a female is present, and how this information influences neural activity."

      Changes made in revision: We excluded motifs overlapping with female calls and added discussion explicitly acknowledging that how female presence is represented in auditory pallium remains unresolved. We do not claim to distinguish visual, auditory, social, or motivational components of the female-present condition.

      Reviewer #3 (Public Review):

      Summary:

      In this study, Jones et al. examine how neural activity in a primary auditory area (field L) of singing male songbirds is modulated by the presence or absence of an audience (a female conspecific). Prior work has demonstrated that the presence of an audience attenuates the responses of dopaminergic neurons to distortions of auditory feedback (DAF). Here the authors report that even in a region that is primarily considered sensory, responses to DAF are also modulated by the audience, although in a heterogeneous manner. However, to be fully persuasive, additional analyses will be required to address how much of the apparent modulation by audience may be explained by other factors such as changes in recorded neurons or their properties over time.

      (1) A central concern relates to whether the main reported effects associated with differences in singing directed versus undirected song reflect only those changes in conditions, versus contributions from changes in unit isolation or response properties over time.

      We completely agree that unit stability is critically important in this study. To address this concern, we now quantify stability and apply strict inclusion criteria adopted from a study that assessed unit stability over days (Dickey et al., 2009). Additionally, we now include average waveform overlays for all example units across conditions as supplemental Figure S2.

      Changes made in revision: We added: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." and "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."

      (2) A second concern has to do with the categorical definition of 'error neurons'. The authors define a subset of neurons as error responsive only if their responses to DAF exceed a specific threshold (2.5 standard deviations). The problem is that for some neurons categorically defined as being responsive to DAF in only one condition, there is almost certainly not a significant difference in the actual responses to DAF between conditions.

      We overhauled our analyses characterizing DAF responses. Rather than relying only on a 2.5 z-score threshold, we now use a more conservative permutation-based approach that directly tests DAF responsiveness and context-dependent changes in DAF responsiveness.

      The revised Results state: "Statistical tests defining auditory neurons as DAF-responsive or not in a binary fashion may not be suitable if the underlying population of DAF-related responses exist on a continuum from responsive to non-responsive."

      The updated result is: "This more conservative approach identified 48/147 neurons as DAF-responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."

      (3a) Some discussion of what is already known about the auditory tuning of Field L, and the extent to which responses associated with distortion of feedback may reflect the frequency tuning of Field L neurons versus something that might be construed as more specifically as detecting an error in perceived feedback.

      We agree that DAF-related changes in firing do not necessarily imply that neurons are explicitly detecting an "error" between predicted and actual feedback. Field L neurons can have spectrotemporal receptive fields and frequency tuning such that a broadband DAF stimulus could drive excitation or inhibition simply because the stimulus overlaps with excitatory or inhibitory regions of a neuron's receptive field. We therefore revised the manuscript to use more cautious language and to describe these responses as DAF-related or feedback-related signals rather than categorically as "error responses".

      Changes made in revision: The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We also added a sentence to the Discussion: "However, it is important to note that DAF-related changes in firing in auditory neurons do not necessarily imply that neurons compute sensory prediction errors. DAF-related responses could arise from ordinary auditory tuning to the broadband distortion stimulus."

      (3b) It would also be useful to discuss further previous work on differences in auditory tuning or responses between conditions when subjects are vocalizing, versus when vocalizations are played back, and to what extent efference copy signals might contribute to the processing of feedback distortions.

      We agree these are important points. Our experimental design did not include sufficient passive bird-own-song (BOS) playback trials to permit quantitative comparisons with vocalizing conditions, and we therefore cannot draw firm conclusions about the contribution of efference copy signals to the DAF responses described here. We did observe robust motif onset-associated neural activations, including some activity preceding motif onset, which were present across both social contexts (see new Figure S4). These observations are consistent with prior reports of premotor-related signals in Field L (Keller and Hahnloser, 2009), but whether such signals contribute differentially to DAF processing across contexts remains an open question that we now acknowledge in the Discussion.

      (3c) To what extent did the current study control for any vocalizations or other sounds produced by females during the directed singing, and could this have contributed to differences in Field L activity between conditions?

      Please see response R2.4 above, in which we describe the exclusion of all song motifs that overlapped in time with a female call. This exclusion criterion was applied throughout all analyses of directed singing.

      Figure 1D: In the directed condition there are no spikes at all following the first handful of motif renditions. Were the directed and undirected recordings interleaved here?

      Undirected and directed trials were not interleaved. The raster plots are presented in chronological order; however, for each behavioral condition, rows are sorted with the earliest renditions at the bottom and the most recent at the top. We have clarified this in the figure legend.

      A minor issue: the raw example trace with male alone does not seem to have a corresponding set of points in the raster plot. For panel E, I also cannot find rasters that correspond to the example recordings shown at top.

      In the original version, we randomly downsampled the condition with more trials to equalize trial counts across conditions in the example rasters, while performing all analyses on the full set of recorded trials. As a result, the example spike shown in the raw trace was drawn from one of the downsampled trials not displayed in the raster.

      Changes made in revision: For greater transparency, we now include all trials from both conditions for each example neuron in Figure 1.

      Figure 2A also shows a neuron that looks like it has non-stationarity; for the alone condition without altered feedback, the main peak has no spikes for the bottom half of the rasters.

      In the original version, example neurons were selected to illustrate the DAF-response scoring method, which in some cases highlighted neurons with less stable response profiles. In the revised manuscript, we have replaced this example with neurons that exhibit more robust and stable DAF-related responses, and we now provide a broader set of example neurons illustrating both increases and decreases in DAF responsiveness across conditions.

      Other figures show firing rate distributions that appear to be very non-Gaussian, with some motifs during which there is a lot of activity, and others in which there is little activity. Please consider applying non-parametric tests as appropriate.

      We agree. In general, some neurons exhibited non-uniform firing rate distributions across trials. All of our main analyses are now conducted using non-parametric permutation tests, which do not assume a Gaussian distribution of trial-by-trial firing rates.

      Approaches to addressing the non-stationarity issue could include more specifically indicating examples in which recordings from the alone condition and directed condition are interleaved and exhibit reversible changes in the pattern of responses.

      Unfortunately, nearly all of our undirected and directed recording periods were not interleaved, as the experimental design required a block of undirected singing followed by directed singing with female presence. We find it informative, however, that DAF-response modulation was observed in both directions, with some neurons losing DAF responsiveness during directed song and others gaining it, a pattern that is difficult to attribute to a simple unidirectional drift in recording quality. We now provide additional examples illustrating both directions of modulation in Figures 2 and 3.

      The methods and/or raster plots should include some further explanation of the time periods over which recordings were made in the alone versus directed conditions, and the extent to which they are interleaved or not.

      We have clarified this in the revised Methods. In brief, recording began when the home cage lights came on each day, with the male left to sing alone until at least 40 undirected song motifs were collected. A female was then introduced in approximately 10-minute intervals until at least 40 directed song motifs were collected. The total recording duration on a given day ranged from 0.56 to 10.27 hours, reflecting variability across birds in the time required to elicit sufficient singing in each context. We have added this information to both the Methods and relevant figure legends.

      It would be most helpful to assess the stability of waveforms and unit isolation across time.

      We now apply strict inclusion criteria based on waveform stability, as described in R3.1 above. SNR was quantified as Vpp/(2*sigma_noise), where Vpp was the peak-to-peak amplitude of each filtered spike waveform and sigma_noise was estimated from the median absolute deviation of the filtered voltage trace. This combines the peak-to-peak normalization used by Nordhausen et al. (1996) with the robust noise estimator described by Rey et al. (2015). Waveform overlays for all included example units are provided in Figure S2.

      It would be reassuring to see that significant differences between conditions are equally or more prevalent under the conditions of greatest unit isolation and recording stability.

      The average SNR of neurons ultimately included in the analysis was 9.47 +/- 3.57, with a minimum of 4.69. Neurons that exhibited significant DAF-response modulation did not have a significantly different SNR than neurons that did not exhibit significant modulation (Wilcoxon rank-sum test, p=0.38). The mean SNR for significantly modulated neurons was 8.70, compared to 9.5 for non-modulated neurons, indicating that the detection of context-dependent modulation was not systematically biased toward neurons with lower recording quality.

      One other way that the authors might be able to address the main concern would be to look at the stability of firing patterns within conditions.

      We agree that stability of firing patterns within conditions is an important consideration, and this concern directly motivated the adoption of the permutation-based analysis described above. In this framework, the observed DAF-response difference between conditions is compared to a null distribution generated by shuffling condition labels across trials. This approach inherently accounts for within-condition trial-by-trial variability and does not assume stationarity of firing rates.

      It would be helpful to have additional explanations of the criteria used for counting spikes, and assessing stability of recordings.

      Spike waveforms were visually inspected for consistency using our custom MATLAB GUI on a 12-second file basis. Interspike interval violations below 1 ms were explicitly checked as an indicator of multi-unit contamination. Detection thresholds were manually set, and each recording file included in the analysis was independently inspected. We have added a more explicit description of these procedures to the Methods section.

      For the specific examples shown in figures, it would be useful to indicate by small tick marks or otherwise which spikes were counted as single units.

      We appreciate this suggestion. In the revised figures, we have improved the clarity of the example raw voltage traces by annotating the detection threshold and, where multiple units were present on a channel, indicating the waveform amplitude range corresponding to the isolated single unit. We believe this provides sufficient transparency regarding spike identity without requiring tick marks on every individual spike, which would substantially reduce legibility of the example traces.

      What were the criteria for determining multi-unit versus single-unit activity?

      In the context of this manuscript, "multi-unit activity" refers to channels on which no single neuron could be reliably distinguished from others based on waveform shape and amplitude. Units ultimately included in the study were those for which a single, consistent waveform cluster could be identified and isolated in the custom GUI. In cases where a second distinguishable unit was present on the same channel, it was manually excluded from the sorted single-unit record. We have clarified this distinction in the Methods.

      Categorical scores: This definition results in cases where responses of 2.45 vs 2.55 are described as 'retuned', even if these responses are not significantly different. Retuning would be more persuasively demonstrated if the authors could provide a test of whether or not the responses for individual neurons differ significantly between conditions.

      We completely agree, and thank the reviewer for motivating us to develop a more rigorous statistical approach. Our revised analysis uses a non-parametric permutation test that explicitly tests for significantly different DAF responses between undirected and directed singing conditions, with correction for multiple comparisons. This replaces the previous threshold-based categorical classification and directly addresses the concern that neurons near the threshold boundary were being treated as categorically different.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Minor comments:

      (1) Please include a schematic of the brain, including the different subregions of Field L and the connections between auditory regions and the midbrain.

      Done. Figure S1 has been added, including a schematic of Field L subdivisions and auditory pathway connectivity.

      (2) The authors should include some additional information about the recordings, such as the proportion of Field L neurons that exhibited singing-related changes in firing rate. It would be helpful to include some examples of spontaneous activity when the bird is quiescent in Figs. 1-2, especially for cells that do not show firing locked to song.

      We appreciate this suggestion. Given the scope of the current revision and the primary focus on DAF-response modulation, we have elected not to add spontaneous activity examples to Figures 1-2 at this time. We agree this would be a valuable addition in future work and have noted it as a limitation in the Discussion.

      (3) Methods, p. 10: Surgery and awake-behaving electrophysiology: "The of the cannula" - this is the only mention of a cannula. Do the authors mean the ends of the probes?

      Cannula placement and wire bundle extension from the end of the cannula has been clarified in the Methods.

      (4) Bottom of p. 10: Fix reference for biorxiv paper: "ref andreas paper"

      Fixed.

      (5) Methods, p. 12: Redundant sentences regarding significant error response criteria.

      Fixed. The redundant sentences have been removed.

      Reviewer #2 (Recommendations For The Authors):

      (1) The abstract is too vaguely formulated. Authors should try to quantify the statements already in the abstract.

      We have reworded the abstract to align with the revision's more conservative claims regarding social context modulation of auditory feedback, and have added specific quantitative statements where possible.

      (2) Authors repeatedly refer to 'perceived song errors' without performing experiments or reporting on behavioral readouts of how birds perceive the jamming sounds. The wording should be changed to something more neutral, e.g. 'DAF responses'.

      We revised the manuscript throughout to use more neutral language centred on "DAF-related" or "feedback-related" responses rather than "perceived errors" or "mistakes". The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We similarly revised the abstract and all relevant passages in the Results and Discussion.

      (3) Authors write that 33 neurons were DAF responsive in both conditions. How should we interpret this overlap relative to independence and identity assumptions?

      We agree that the original presentation made the interpretation of overlap across conditions unclear. The observed overlap is greater than expected under a strict independence assumption but smaller than expected if responsiveness were identical across conditions, consistent with partial but incomplete sharing of DAF responsiveness across social contexts. In the revised manuscript, however, we have moved away from this binary classification framework because DAF responsiveness appears to vary continuously across neurons. The permutation-based analysis now directly tests for changes in DAF responsiveness across contexts without requiring categorical assignment.

      (4) Only 10 neurons were not affected by courtship state or only 10 error responsive neurons were not affected? I suggest authors do a multivariate analysis or use a mixed effect model and summarize the result as a table.

      We agree that the categorical accounting of neurons across conditions was difficult to follow in the original manuscript. In the revised manuscript, we clarified the distinction between neurons responsive to DAF within a condition and neurons exhibiting significant modulation of DAF responsiveness across conditions. We now explicitly report: "This analysis identified 71/147 neurons as DAF responsive in at least one behavioral condition, whereas 76/147 were not responsive in either condition." and "This more conservative approach identified 48/147 neurons as DAF responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."

      (5) It would help if authors could define 'z-scored difference'. Better known is d prime, is this the same?

      For each neuron, the z-scored DAF response was computed as the z-scored firing rate difference between distorted and undistorted trials. Importantly, our revised main analysis avoids any normalization such as z-scoring, and instead uses a permutation-based approach applied directly to spike counts.

      (6) Is the 'retuning' assessment a bit conservative? Neurons could also retune by showing error scores greater than 2.5 in both conditions but a shifted response time.

      We agree that neurons could retune by shifting the latency of DAF responses. Although potential latency shifts are beyond the scope of the current study, we did observe suggestive evidence of possible latency changes in some example neurons across conditions. We have noted this as an interesting direction for future analysis.

      (7) Could the stability of DAF response across trials be described? E.g. as the ratio between intra versus inter condition variability?

      We agree that stability of DAF responses across trials is an important concern. In addition to imposing strict waveform stability requirements, our permutation-based statistical test explicitly accounts for trial-by-trial variability by constructing null distributions from within-condition trial shuffles. We have also replaced the previously shown unstable example neuron with neurons that exhibit more consistent DAF-related responses across trials, and provide additional examples in Figures 2 and 3.

      Minor:

      (8) 'significant increase in burst fraction': specify effect size of t test in results section.

      We now specify in the main text: "A small but significant increase in burst fraction was observed (paired t-test, p=9.3x10-6, n=138 neurons, mean +/- SEM: 0.11 +/- 0.006 vs 0.15 +/- 0.007, Figure 1J)."

      (9) The IMCC parameter should be specified in the main text.

      The Gaussian smoothing parameter (20 ms) has now been specified in the main text.

      (10) Fig. 2: indicate the windows within which error scores are computed.

      This is no longer applicable, as the revised permutation-based analysis does not rely on scoring error responses within a fixed window.

      (11) In Fig. 2A, the neuron has an error score of -2.54 (significant), but the red and blue curves look almost the same.

      We agree that the previous error score quantification did not always capture firing rate differences in an intuitive way. This example neuron has been replaced in the revised manuscript, and the new analysis avoids scalar error scores in favor of the permutation-based approach.

      Reviewer #3 (Recommendations For The Authors):

      Minor points:

      (1) "(ref andreas paper)." Add reference here?

      Fixed.

      (2) Hessler and Doupe 1999 is a good reference for premotor signal re-tuning during courtship.

      We agree. The reference has been included in the revised manuscript.

      (3) Page 5: "discharge depended on courtship state, using" - should this be "depending"?

      The original wording was intentional: "we tested how discharge depended on courtship state." We have verified this reads correctly in context and made no change.

      (4) Page 9: "consistent with a brainwide process" - what is meant here?

      We have revised this wording. The revised manuscript replaces "brainwide process" with clearer language describing a distributed modulation of auditory responsiveness that is not confined to a single nucleus.

    1. eLife Assessment

      This important study provides mechanistic evidence for how the tea-adapted Kanzawa spider mite, Tetranychus kanzawai, overcomes the catechin-based defenses of green tea plants. The work identifies the horizontally transferred dioxygenase DOG15 as a key contributor to host adaptation and supports a two-step model involving evolutionary modification of enzyme activity together with strong inducible upregulation upon feeding on tea. The evidence is convincing because comparative behavioral and toxicological assays, transcriptomic and proteomic analyses, RNAi-mediated functional validation, and recombinant enzyme assays converge to link DOG15 activity and expression with improved performance on tea. The revised manuscript appropriately acknowledges that the products of catechin cleavage have not yet been characterized and that additional detoxification pathways may contribute to tea adaptation, providing a balanced interpretation of the otherwise strong mechanistic evidence.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates the molecular mechanisms allowing the KSM mite to infest tea plants, a host that is toxic to the closely related TSSM mite due to high concentrations of phenolic catechins. The authors utilize a comparative approach involving tea-adapted KSM, non-adapted KSM, and TSSM to assess behavioral avoidance and physiological tolerance to catechins. The main finding is that tea-adapted KSM possesses a specific detoxification mechanism mediated by an enzyme, TkDOG15, which was acquired via horizontal gene transfer. The study demonstrates that adaptation is a two-step process: (1) structural refinement of the TkDOG15 enzyme through amino acid substitutions that enhance enzymatic efficiency against catechins, and (2) significant transcriptional upregulation of this gene in response to tea feeding. This enzymatic adaptation allows the mites to cleave and detoxify tea catechins, enabling survival on a toxic host plant.

      Strengths:

      A multiomics approach (transcriptomics and proteomics) provided a compelling cross-validation of its findings. Functional bioassays, such as RNAi and recombinant enzyme assays, demonstrated that the adapted mite has higher activity against catechins via TkDOG15. Other methodologies, like feeding assay using a parafilm-covered leaf disc, were effective in avoiding contact chemosensation.

      Comments on revised version.

      The authors have satisfied all previous concerns through necessary text revisions and clarified discussions. The manuscript is now well-balanced and scientifically sound.

    3. Reviewer #2 (Public review):

      Summary:

      The fascinating topic of the host range of arthropods, including insects, and the detoxification of host secondary metabolites has been elucidated through studies of the host specificity of two closely related species. The discovery that key genes were acquired from fungi through horizontal gene transfer (HGT) is particularly significant.

      Strengths:

      (1) The discovery that the TkDOG15 enzyme, acquired through HGT from fungi, plays a key role in the detoxification of green tea catechins in the Kanzawa mite, revealing a new mechanism of plant-herbivore interactions, is highly encouraging.

      (2) The verification of this finding through various experiments, including behavioral, toxicological, transcriptomic, and proteomic analyses, RNAi-based gene function analysis, and recombinant enzyme activity assays, is also highly commendable.

      (3) By proposing a two-step model in which amino acid substitutions and expression regulation of a specific enzyme gene (TkDOG15) enable host adaptive evolution, this study contributes significantly to our understanding of the evolutionary mechanisms of speciation and plant defense overcoming.

      Comments on revised version.

      I believe the manuscript has been significantly refined since the initial draft was submitted.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study provides mechanistic evidence that tea-adapted two-spotted spider mite overcomes green tea catechin defenses via the horizontally transferred dioxygenase TkDOG15, supporting a two-step adaptation model, combining enzyme refinement and inducible upregulation. The evidence is convincing because multi-omics signals converge with functional validation (RNAi knockdown and recombinant enzyme assays) and well-controlled behavioral/toxicity assays to link TkDOG15 activity and expression to survival and feeding on tea.

      We thank the editors and reviewers for this positive assessment of the importance of our study and the strength of the evidence. We would like to point out one factual correction. The assessment describes the tea-adapted mite as the "two-spotted spider mite" (TSSM, Tetranychus urticae), but the species adapted to tea in this study is the Kanzawa spider mite (KSM, Tetranychus kanzawai). We would suggest revising "tea-adapted two-spotted spider mite" to "tea-adapted spider mite" or "tea-adapted Kanzawa spider mite" accordingly.

      Reviewer #1 (Public review):

      Summary:

      This study investigates the molecular mechanisms allowing the KSM mite to infest tea plants, a host that is toxic to the closely related TSSM mite due to high concentrations of phenolic catechins. The authors utilize a comparative approach involving tea-adapted KSM, non-adapted KSM, and TSSM to assess behavioral avoidance and physiological tolerance to catechins. The main finding is that tea-adapted KSM possesses a specific detoxification mechanism mediated by an enzyme, TkDOG15, which was acquired via horizontal gene transfer. The study demonstrates that adaptation is a two-step process: (1) structural refinement of the TkDOG15 enzyme through amino acid substitutions that enhance enzymatic efficiency against catechins, and (2) significant transcriptional upregulation of this gene in response to tea feeding. This enzymatic adaptation allows the mites to cleave and detoxify tea catechins, enabling survival on a toxic host plant.

      Strengths:

      A multiomics approach (transcriptomics and proteomics) provided a compelling crossvalidation of its findings. Functional bioassays, such as RNAi and recombinant enzyme assays, demonstrated that the adapted mite has higher activity against catechins via TkDOG15. Other methodologies, like feeding assay using a parafilm-covered leaf disc, were effective in avoiding contact chemosensation.

      Weaknesses:

      Although TkDOG15 is assumed to "detoxify" catechins by ring cleavage, the study doesn't identify or characterize the breakdown metabolic products. If the metabolites are indeed non-toxic compared to the parent catechins, that would strengthen the detoxification hypothesis. Also, the transcriptomic and proteomic analyses identified other potential detoxification enzymes, such as CCEs, UGTs, and ABC (Supplementary Tables 3-1 & 3-2), which were also upregulated. The manuscript focuses almost exclusively on TkDOG15, potentially overlooking a multigenic adaptation mechanism, where these other enzymes might play synergistic roles, although it was mentioned in the discussion section.

      Reviewer #1 (Recommendations for the authors):

      There is no need for additional experiments, but I suggest revising the discussion section to mention the weaknesses pointed out above.

      We thank the reviewer for the positive assessment and helpful suggestions. We have revised the Discussion (L276-283) to address both points as limitations. First, we now note that we did not characterize the products of TkDOG15-mediated catechin cleavage, and that confirming their reduced toxicity relative to the parent catechins would further support its detoxification role. We note this as a direction for future work. Second, we note that DOG15 in KSM on tea was the only enzyme upregulated at both the mRNA and protein levels, whereas the CCEs, UGTs, ABC transporter, and other DOGs were enriched in only one dataset. We now state that tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15 and warranting functional validation.

      Minor corrections below:

      (1) Figure 1a: For better readability, I recommend adding "KSM" and "TSSM" to the two pictures, respectively.

      Done.

      (2) L165: tetur20g01790 refers to a TSSM gene, while TkDOG15 refers to a TSM protein. Revise it accordingly. (Same for L442 and L481).

      The reviewer is correct that tetur20g01790 is the TSSM gene ID. As the KSM genome is not yet available, we identified the TkDOG15 gene, the KSM ortholog of tetur20g01790, by de novo assembly of our RNA-seq reads. We have revised L169 and L455 accordingly.

      (3) L264: Supplemental Table 3-2.

      Done (L271).

      Reviewer #2 (Public review):

      Summary:

      The fascinating topic of the host range of arthropods, including insects, and the detoxification of host secondary metabolites has been elucidated through studies of the host specificity of two closely related species. The discovery that key genes were acquired from fungi through horizontal gene transfer (HGT) is particularly significant.

      Strengths:

      (1) The discovery that the TkDOG15 enzyme, acquired through HGT from fungi, plays a key role in the detoxification of green tea catechins in the Kanzawa mite, revealing a new mechanism of plant-herbivore interactions, is highly encouraging.

      (2) The verification of this finding through various experiments, including behavioral, toxicological, transcriptomic, and proteomic analyses, RNAi-based gene function analysis, and recombinant enzyme activity assays, is also highly commendable.

      (3) By proposing a two-step model in which amino acid substitutions and expression regulation of a specific enzyme gene (TkDOG15) enable host adaptive evolution, this study contributes significantly to our understanding of the evolutionary mechanisms of speciation and plant defense overcoming.

      Weaknesses:

      While transcriptome/proteome analyses reported changes in the expression of other detoxification-related enzymes, including CCEs, UGTs, ABC transporters, DOG1, DOG4, and DOG7, it is regrettable that the contribution of each enzyme, including its interaction with TkDOG15 and the functional analysis of each enzyme within the overall catechin detoxification system, was not investigated.

      We thank the reviewer for the encouraging assessment and this comment. We agree that the contributions of the other detoxification-related enzymes, including their interaction with DOG15, remain to be investigated. As this point overlaps with a comment from Reviewer 1, we have revised the Discussion (L276-283) to note that DOG15 was the only enzyme upregulated at both the mRNA and protein levels, whereas the CCEs, UGTs, ABC transporter, and other DOGs were enriched in only one dataset. We now state that tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15, and that the functional analysis of their individual and combined contributions warrants future work.

      Reviewer #2 (Recommendations for the authors):

      The manuscript titled "Adaptation of an Herbivorous Arthropod to Green Tea Plants by Overcoming Catechin Defenses" presents a well-designed, mechanistically insightful study that advances our understanding of herbivore adaptation to plant chemical defenses. The work is scientifically sound and of potential interest to a broad readership in chemical ecology and evolutionary biology.

      However, before the manuscript can be considered for acceptance, the authors must adequately address the comments outlined below regarding clarity, presentation, and interpretation across the manuscript.

      We thank the reviewer for the positive evaluation of our study. We have carefully addressed each of the specific comments below regarding clarity, presentation, and interpretation, and we believe these revisions have substantially improved the manuscript.

      Specific comments on each section:

      (1) Abstract

      (a) The authors are encouraged to add a concise concluding sentence summarizing the broader significance of the study and indicating potential future research directions or limitations, which would strengthen the impact of the abstract.

      We have added a concluding sentence to the Abstract summarizing the broader significance of the study and indicating future directions (L38-40).

      (b) The authors may consider adding representative quantitative results to the abstract, as this would enhance clarity and increase the impact and interpretability of the study for readers.

      We have added representative quantitative results to the Abstract. Specifically, we now state that the mRNA and protein levels of DOG15 in tea-adapted T. kanzawai are up to 31.6 and 12.1 times higher, respectively, than in T. urticae fed on tea plants (L30-32). For consistency, we now refer to the gene as "DOG15" throughout the Abstract (L29, L30, and L36).

      (2) Introduction

      (a) While the paragraph is informative, it reads more like a summary of the main results than a statement of study objectives. The authors are encouraged to reframe this section to explicitly define the study's aims and hypotheses.

      We have reframed the final paragraph of the Introduction to explicitly state the study's aims and hypotheses rather than to summarize the results (L73-81).

      (b) The authors should avoid excessive citation of multiple references for a single thematic statement when one key reference is sufficient. Where appropriate, inclusion of more recent literature is encouraged.

      We have reduced multiple citations for single statements to the most representative references: Cabrera et al. (2006) for the health benefits of catechins (L46) and Grbić et al. (2011) and Dermauw et al. (2013) for the DOG gene count (L66-67).

      (3) Materials and Methods

      (a) The Materials and Methods section is comprehensive and technically sound; however, its length and density reduce overall clarity. The authors are encouraged to streamline descriptions of standard or well-established protocols and rely on appropriate citations where possible.

      We agree that clarity can be improved by removing redundancy. The Materials and Methods are intentionally detailed to allow independent replication of our protocols, so we have retained this detail and instead removed the overlapping methodological descriptions from the figure captions, where the same information was repeated (see our response to comment 6a).

      (b) Greater consistency is needed in reporting biological and technical replicates across different experiments (e.g., performance assays, transcriptomics, proteomics, and enzymatic activity assays) to enhance reproducibility.

      We have standardized the reporting of replicates across all experiments to the format "x independent experimental runs (n = y per run)." Throughout the manuscript, "independent experimental runs" denotes biological replicates, with technical replicates specified separately where applicable (three technical replicates for qRT-PCR).

      (c) The authors should provide brief justification for key methodological parameters, such as catechin concentrations, exclusion criteria in behavioral assays, and thresholds used for defining DEGs and DEPs, to improve transparency and interoperability.

      We have added brief justifications for the three parameters. 1) The catechin concentration range was chosen to encompass the individual catechin levels measured in fresh tea leaves (L340-341). 2) In the behavioral assays, inactive mites were excluded because their movement was insufficient to determine chemo-orientation behavior, and escaped mites were excluded because they did not complete the assay (L365-367). 3) The thresholds for DEGs and DEPs follow criteria commonly applied in mite transcriptomic studies (Vidal-Quist et al., 2025, newly added to the references) (L414-416) and are consistent with our previous spider mite proteomic analysis (Arai et al., 2025) (L444-445).

      (4) Results

      (a) While significant differences in survival and fecundity are reported, briefly indicating the magnitude of these differences (e.g., percentage or fold change) would improve clarity and strengthen the presentation (Lines 91-96).

      We have added the magnitude of the differences (L94-97). The revised text now states that after 10 days, almost 90% of tea-adapted KSM survived, compared with about 5% of non-adapted KSM and 33% of TSSM, and that tea-adapted KSM laid up to about 2 eggs/surviving female daily, whereas the other two populations laid almost no eggs.

      (b) The final sentences include interpretative and concluding statements regarding catechins as key metabolites and mite adaptation. These statements would be more appropriate for the Discussion section rather than the Results (Lines 127-130). Follow the same for the rest of the Results section also.

      Following the reviewer's suggestion, we have removed the interpretive and concluding statements from the end of the Results section, so that it now reports only the observations (L129-130). The interpretation regarding the multiple modes of action of catechins and the insensitivity of tea-adapted KSM is already presented in the Discussion (L206-213 and Conclusions), so we did not duplicate it there. We also reviewed the remaining Results subsections and confirmed that they report the experimental observations and their direct conclusions without broader interpretation.

      (c) The comparison among catechin classes is clear; however, briefly listing the mean concentrations of each catechin (as shown in Figure 2a) in the text would improve readability without duplicating the figure (Lines 135-139).

      We have added the approximate mean concentration of each catechin to the text (L136-137).

      (d) Please clarify in the Results whether the same exposure concentration and duration were applied for all catechins and mite species, or explicitly direct readers to the Methods section (Lines 141-142).

      We have clarified in the Results section that all four catechins were tested at the same concentration series (0, 10, 10<sup>2</sup>, 10<sup>3</sup>, 10<sup>4</sup>, and 10<sup>5</sup> ppm) and the same exposure duration (24 h) for both mite populations (L143).

      (e) The phrase "lower sensitivity" should be explicitly linked to LC<sub>50</sub> estimates to ensure that the basis of comparison is immediately clear to readers (Lines 143-144).

      Following the reviewer's suggestion, we have linked the sensitivity comparison to the LC<sub>50</sub> values (L143-147). The comparison is now stated relative to TSSM based on the LC<sub>50</sub> estimates, and for ECg and EC we note that the LC<sub>50</sub> of tea-adapted KSM exceeded the highest concentration tested.

      (f) This section clearly identifies TkDOG15 as a key gene underlying tea adaptation in KSM; however, the authors are encouraged to briefly clarify the criteria used to define "highly enriched" mRNAs and proteins (e.g., fold-change and statistical thresholds) in the Results text or by explicitly directing readers to the Methods. This would improve transparency and facilitate interpretation of the multi-omics comparisons (Lines 147-173).

      We have added the criteria used to define the enriched mRNAs and proteins (log2 fold change ≥ 1 with adjusted p-value < 0.05 for mRNA and p-value < 0.05 for protein) and referred readers to the Materials and Methods (L159-160).

      (g) The enzymatic comparison between TkDOG15 and TuDOG15 is well presented; however, the authors are encouraged to briefly discuss whether the two amino acid substitutions (Q127A and T203A) were individually or jointly responsible for the increased catalytic efficiency, or to acknowledge this as a limitation and potential direction for future functional studies (Lines 176-190).

      We have added a brief discussion of whether the two substitutions (Q127A and T203A) act individually or jointly (L254-257). We note that T203A is adjacent to the active-site residue Y202 and may contribute more directly to catalytic efficiency, and we acknowledge that dissecting their individual contributions by site-directed mutagenesis is a direction for future work.

      (5) Discussion

      (a) The authors appropriately acknowledge that the molecular basis of chemosensory insensitivity and the contribution of additional detoxification enzymes remain unresolved. To further improve clarity, these statements could be explicitly framed as hypotheses or future research directions to clearly distinguish them from experimentally supported mechanisms (Lines 205-208; 266-270).

      We have reframed the statements on chemosensory insensitivity (L209-213) and the contribution of additional detoxification enzymes (L272-274) as hypotheses and future directions, distinguishing them from the experimentally supported mechanisms.

      (b) While DOG15 is convincingly identified as a key contributor to tea adaptation, a brief clarification of its relative importance compared with other upregulated detoxification enzymes would strengthen interpretative balance, even if the roles of these enzymes remain unresolved (Lines 259-265).

      DOG15 was the only enzyme upregulated at both the mRNA and protein levels (Figure 3d,e), and the only enzyme functionally validated in this study, by RNAi silencing (Figure 3f) and recombinant enzyme assays (Figure 4c). We have established that DOG15 contributes to tea adaptation, but because the other upregulated enzymes were not functionally tested, their relative contributions cannot be determined at this stage. As we note in the Discussion, tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15 (L277-283). We therefore did not add further text, to avoid duplication.

      (c) The discussion linking host plant adaptation to reproductive isolation and ecological speciation is interesting and well contextualized; however, these evolutionary implications should be slightly tempered or explicitly framed as potential long-term outcomes beyond the immediate scope of the present study (Lines 271-281).

      We have tempered the evolutionary implications (L292-294). The revised sentence now frames the link to reproductive isolation and ecological speciation as a potential outcome over longer evolutionary timescales rather than a direct finding of the present study.

      (6) Figure captions

      (a) The figure captions (Figures 1-4) are exceptionally detailed and, in several places, repeat methodological information already described in the Materials and Methods. The authors are encouraged to shorten the captions by retaining only information necessary to interpret the figures, while referring readers to the Methods for experimental details.

      We have shortened the figure captions (Figures 1-4) by removing methodological details that are described in the Materials and Methods, retaining only the information needed to interpret each figure. Where appropriate, readers are now referred to the Materials and Methods or to Supplemental Figure 1-2 for the full experimental procedures.

      (b) Several captions contain long, multi-sentence descriptions that may hinder readability. The authors may consider simplifying the wording, grouping related panels more concisely, and removing procedural details (e.g., extraction conditions, exposure durations, and instrument settings) to improve clarity and visual accessibility.

      As described in our response to comment 6a, we have simplified the figure captions by removing procedural details such as extraction conditions, exposure durations, and instrument settings, and by grouping related panels more concisely. These details are retained in the Materials and Methods.

      (c) In Figure 1, the panel labels (a-h) do not appear in a clear sequential order. For consistency with the other figures and to improve readability, the authors should ensure that panel lettering is arranged in a logical, sequential order throughout the manuscript.

      We appreciate the reviewer's attention to panel ordering. In the current layout, the panel lettering follows the order in which the panels are first cited in the text. Arranging the panels in a strict left-to-right, top-to-bottom sequence would require reducing the size of several panels, including the HPLC chromatogram in panel (e) and the survival and fecundity time courses in panels (c) and (d), which would compromise their readability. We have therefore retained the current arrangement, in which related panels are grouped together and the larger panels are kept at a legible size. We hope the reviewer finds this acceptable.

    1. eLife Assessment

      This valuable study presents a real-time system for identifying multiple unrestrained marmosets in a home cage setting using a combination of facial features and color-coded beads. While there is solid evidence that the system has a precision comparable to human experimenters in the tested scenarios, there is limited evidence that this would generalize to unconstrained multi-animal environments

    2. Reviewer #1 (Public review):

      The manuscript by Yang, Wang, and Cléry presents a pipeline for real-time identification of common marmosets in a laboratory setting. Models were trained and evaluated on data derived from a family of three closely related adults and a set of juvenile twins. Freely moving animals entered an enclosed space fixed to the housing cage door, which permitted the entry of individual animals for data acquisition. Utilizing YOLOv8-nano, identification was improved through the introduction of uniquely colored collar beads. Analyses of facial similarity showed close morphological relatedness amongst individuals and highlighted the need for highly discriminative classification. The authors demonstrate that combining facial detection with visual markers enables adequate identity assignment under controlled laboratory conditions with minimal cross-individual misclassification.

      The main strengths are that the proposed pipeline offers a solution for real-time identity tracking in common marmosets. Its lightweight design enables deployment across a wide range of hardware configurations. Furthermore, if similar strategies are employed, this methodology is likely adaptable for other species with minimal modification. Additionally, evaluation of closely related individuals provides a necessary stress test for the discrimination of facial identity tracking. However, the main weakness is the pipeline's reliance on controlled animal isolation and small visual markers, which raises questions about the approach's generalizability to unconstrained multi-animal environments. The authors justify the use of beads, but the dependency of facial recognition on the beads needs to be described more clearly, as it is unclear how independent facial recognition performance truly was. The overall utility of this approach therefore remains to be seen.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Yang et al. develop a real-time system for automatic face detection and identification of multiple unrestrained common marmosets in a home cage setting.

      Strengths:

      The study aims to address an unmet need in behavioral neuroscience: the ability to non-invasively identify animals is crucial to the automated and rigorous study of neural behaviors; this is especially true for common marmosets, which are rapidly becoming a model system of choice for the study of complex social cognition. By using a YOLOv8 backbone, the study achieves human level performance, both in terms of precision and recall of the trained models.

      Weaknesses:

      The robustness of the system is not clear from the limited datasets presented.

      Comments on revised version.

      The authors have adequately addressed my comments from the previous round, and I have no further comments

    4. Reviewer #3 (Public review):

      Summary:

      In the revised manuscript, the authors provide additional details and evidence regarding the robustness and utility of their method.

      Strengths:

      (1) The authors provide a very precise automatic identification of marmosets in their home cage, to levels comparable to animal health professional.

      (2) This method is robust across lightning, camera angles etc but importantly is able to identify marmosets in naturalistic conditions, which can be of tremendous value to neuroscientists and to ecological or behavioral studies.

      (3) Easy to use and implement, requiring minimal settings. Phone videos can even be used.

      Weaknesses:

      While the manuscript improved tremendously from the previous version, given the nature of the paper, it is still a strenuous read.

      Comments on revised version.

      The authors did a good job of addressing my previous concerns and I don't have more comments.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary: 

      The manuscript by Yang, Wang, and Cléry presents a lightweight pipeline for real-time identification of common marmosets in a laboratory setting. Models were trained and evaluated on data derived from a family of three closely related adults and a set of juvenile twins. Freely moving animals entered an enclosed space fixed to the housing cage door, which permitted the entry of individual animals for data acquisition. Utilizing YOLOv8-nano, identification was improved through the introduction of uniquely colored collar beads. Analyses of facial similarity showed close morphological relatedness amongst individuals and highlighted the need for highly discriminative classification. Overall, the authors offer a framework for identity tracking that prioritizes real-time inference. The authors demonstrate that combining facial detection with visual markers enables adequate identity assignment under controlled laboratory conditions with minimal cross-individual misclassification. 

      Strengths: 

      (1) The proposed pipeline offers a solution for real-time identity tracking in common marmosets. Its lightweight design enables deployment across a wide range of hardware configurations. Furthermore, if similar strategies are employed, this methodology is likely adaptable for other species with minimal modification. 

      (2) Evaluation of closely related individuals provides a necessary stress test for the discrimination of facial identity tracking. 

      Weaknesses: 

      (1) The pipeline's reliance on controlled animal isolation and small visual markers raises questions about the approach's generalizability to unconstrained multi-animal environments. The provided confusion matrices (Figures 6-8) indicate that the most common misclassifications are background-related, possibly suggesting that detection specificity is the primary source of error. All things considered, these findings raise concerns about performance in its use in socially dynamic and visually complex environments. 

      Thank you for the comment. The background column of the confusion matrix can be explained by several occasions: a) the model detects an object where there is no object, b) there is more than one prediction label for the same object, or c) an object appeared in the image but not manually labeled, however the program was able to detect that object. The value of the background column does not necessarily mean that the detection is incorrect, as the precision score for the detection labels are good. We have rephrased the relevant sections for clarification to include the sources of the increased value in background columns in confusion matrices, as follows:

      “The background class of the confusion matrix showed frequently predictions as marmoset faces and collar beads for the training (Figure 6A) and validation set (Figure 6B). However, it does not necessarily indicate incorrect predictions or misclassifications. Instead, these values were mostly explained by multiple detections of the same object class. For instance, additional marmoset faces were predicted when multiple animals were present within a single video frame. The long collar structure or motion blur of the marmosets could also cause multiple detections of beads that belong to the same collar. This also corresponded to the high precision and recall scores observed across prediction classes (Figure 5D), suggesting that the increased background false positives were mainly related to the object-count discrepancies, instead of poor detection performance.”

      Prediction misclassification is one source of the background false positive. The misclassification could not be avoided in automatic prediction algorithm, but we included the manual filtering and majority-voting during our real-time classification to reduce this effect. Multiple detection of the same class may also be considered as the background, since only one object may be labelled in the ground truth, such as multiple collar beads or automatic face extraction. In addition, blurry objects were not labeled manually during training but can be detected during prediction, which also resulted in background false positive. It was clarified in the main text as follows:

      “The normalized confusion matrices showed high accuracy and consistency of most marmoset faces and collars detection in training (Figure 8A) and validation (Figure 8B) tests, with some exceptions. Particularly, the background was frequently identified as the collar of Young2 marmoset. This elevated background score was likely contributed by the multi-color design of the Young2 marmoset collar, making it more difficult to distinguish compared to collars with a single bead color. In this occasion, if one bead is occluded, blurred, or outside the field of view, the other visible collar bead could affect the prediction and lead to an incorrect identification from the ground truth.”

      (2) The manuscript claims performance comparable to that of human experimenters but provides no explicit evidence to support these claims. While it is plausible that human experimenters may be less accurate in facial recognition tasks involving closely related marmosets, the authors don't provide evidence. Moreover, while that might be the case, the color-coded beads provide a salient identity cue for the model, which complicates the interpretation of this comparison grounded in facial recognition. 

      Thank you for pointing out this concern. The aim of the facial recognition tool is to collect data from marmosets without having experimenters to check the identity continuously. The program is not aimed at outperforming the experimenters’ role but avoid having constant human intervention that can disrupt a more ecological in cage data collection. It is also essential for having more flexibility to collect data in case a specific experimenter is not here and thus to not disrupt the project. Human experimenters have extensive experience closely working with marmosets, having the unique collar beads associated to each marmosets allows human experimenters to hardly make mistakes identifying marmosets and to do it quickly. We collected identification accuracy of human experimenters by presenting 10 clips of the five marmosets involved in the manuscript (2 clips per marmosets), with 2 random clips repeated twice. The results were plotted by each experimenter. The identification accuracy of the experimenters correlates with the time spent with the animals, as the animal health technicians (responsible for daily health check and husbandry) achieved 95.83% average accuracy in identifying the marmosets. We clarified those points in the main text as follows:

      “Its automated pipeline substantially reduces the time and work required for traditional manual identity labeling, while maintaining an expert-level human performance and reproducibility across experimenters (95.83% average accuracy for animal health technicians, responsible for daily health check and husbandry while lab experiments ranges between 25 to 80% of accuracy depending on the amount of time spent with each animal, Supplementary figure 1). The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.”

      We filtered the prediction of the collars, and the identification result solely based on the faces for the 2 young marmosets was correct. The prediction results were plotted on Video 7, Video 7—video supplement 1, and Video 7—video supplement 2 and added to the Results section as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Explanation for classification of marmoset faces and collar beads in the Discussion section:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”.

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      Reviewer #2 (Public review):

      Summary: 

      In this study, Yang et al. develop a real-time system for automatic face detection and identification of multiple unrestrained common marmosets in a home cage setting. 

      Strengths: 

      The study aims to address an unmet need in behavioral neuroscience: the ability to non-invasively identify animals is crucial to the automated and rigorous study of neural behaviors; this is especially true for common marmosets, which are rapidly becoming a model system of choice for the study of complex social cognition. By using a YOLOv8 backbone, the study achieve human level performance, both in terms of precision and recall of the trained models.

      Weaknesses: 

      The robustness of the system is not clear from the limited datasets presented. The use of color-coded beads undercuts the study's premise that the system achieves truly non-invasive tracking. Although the system achieves good performance in face detection, it does not perform as well for classification using faces alone (especially when the faces are similar, as in twin animals). Here, too, the color-coded beads play a key role in identity discrimination. The stated goals of the study and the actual results presented are therefore at odds.

      Thank you for the comment. First, we would like to clarify the role of the collar beads in our system. Compared to the faces, a unique identity marker, the collar beads were not used as the main identity classifier but rather as an external visual marker. The color-coded bead was not used solely for the purpose of marmoset video classification; it was also used as an additional source of identification for one marmoset. As the marmosets usually move very fast inside the cage, it is mainly used as a visual marker for experimenters to recognize them in a distance in a short time.

      The mislabelling is more frequent with the young twins not only due to their face similarity, but also due to the limited number of images being used for the model training, as discussed in the paragraph #4 of the Discussion section. Collar beads are small and less frequently detected by the camera, since it could be occluded by the marmoset fur. In addition, it was invisible to the camera if the marmoset turned sideways or was far from the camera. Therefore, higher weight was assigned to the beads due to their small size and less frequent detections compared to face labels, such that it was only an element to confirm the identity, instead of the main classifier.

      The inclusion of the collar beads doesn’t affect the prediction results of the marmoset faces. The model achieved a good precision/recall score for the identity labeling in the manuscript. In the revision, with the majority-vote strategy, we filtered the detection of all collar beads and showed that the model was able to correctly identify the marmosets solely by their faces. The Results section has been modified as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Explanation for classification of marmoset faces and collar beads in the Discussion section:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”.

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      Reviewer #3 (Public review):

      Summary: 

      In this manuscript, Yang et al introduce a new method for automatically identifying marmosets in their home cage using a supervised deep learning method that recognizes the face and colored beads on marmoset collars. The authors show a high precision rate of identifying marmosets to levels comparable to a human experimenter. The method overall seems robust at identifying marmosets at different life stages and different settings; however, given the current form, I'm struggling to see the generalizability and experimental utility of this method. 

      Strengths: 

      (1) The authors provide a near-perfect automatic identification of marmosets in their home cage. 

      (2) This method is robust across lightning, camera angles, etc., making it potentially useful for marmoset (and other NHP) identification outside the housing cage as well 

      Weaknesses: 

      (1) Despite the almost perfect precision, in its current form, I'm failing to see how this method can be useful to other labs. 

      Thank you for your comment. This Tools & Resources paper mainly described the development of the marmoset identification program and methods. Future work will focus on extending the program application on identification from different housing conditions, in combination with various behavioral tasks such as in-cage touchscreen system or manual tasks, and in the wild that precludes handling or isolation of marmosets for collecting behavioral data. The program solely requires a camera, a computing device, and marmosets, as there are no hardware restrictions. In addition, we are currently collaborating with other labs on the marmoset identification from videos taken from other setups. The program achieved effective face extraction from the marmoset in the video, without the need for additional program modifications.

      (2) This is a nice methods manuscript, but the authors do not present results to show how their method can be used outside of identifying marmosets inside their home cages in a small field of view. 

      Thank you for your feedback. The method developed was applied in combination with other touchscreen behavioral tasks, aiming to extract data without human intervention. This approach was discussed in paragraph #6 in the Discussion section. While this manuscript focuses on the methods of close-view face identification when marmosets perform behavioral tasks, the identification and automatic face extraction program could also be applied to marmoset videos taken from a larger view, including phone cameras. Even though the marmosets are still housed in their home cage, the example videos presented the program’s application in a larger field of view. We have added examples of the videos/photos from a different experimental setup to respond to this comment in the Discussion section as follows:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10).”

      (3) Reading the manuscript is strenuous, given its repetitive nature. Consolidating and shortening the results, as well as adding some definitions to the results section, would be helpful. 

      Thank you for pointing this out and your suggestions. We have rephrased the Results section for simplification to facilitate the understanding of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The weight of the color-coded beads was increased to improve identification accuracy. From a brief look at the code provided on GitHub, the weight assigned to the beads seems substantial. This calls into question the need to use facial recognition in the identification strategy. As the code currently stands, facial identity appears to serve primarily as a fallback when bead detection fails to register. To strengthen the methodological justification, the paper would benefit from the authors providing a rationale for choosing this weighting scheme and, if available, supplemental figures showing performance across a range of different weights to demonstrate why that specific value was assigned in the algorithm.

      We have added the model prediction results without the collar beads showing that the facial recognition algorithm works even without the collar beads and that those collar beads are not the main classifier. The Methods section has been modified as follows:

      “For each detected bounding box, the scripts returned a corresponding label of marmoset face and collar bead color. We assigned the detected collar beads as the corresponding marmoset identity with a higher weight, which improved the detection confidence across frames.”

      The Discussion section has been modified as follows:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      The weight assigned to the beads in the GitHub page is the highest weight that we would suggest. The actual weight can be customized by the experimenters based on the actual experimental setup. For example, we used the weight of 2 in our real-time version of the marmoset face identification, while marmosets were presented with their corresponding tasks once identified. The GitHub page has been edited to clarify this point.

      (2) The overall utility of this approach, other than the real-time detection component, needs more clarification. It is currently unclear why this approach, and in which specific experimental or observational settings, is particularly advantageous compared to existing methods for assigning animal identity.

      In addition to the advantages mentioned in paragraph #1 of the Introduction and paragraphs #1-3 in the Discussion, we have added more details in the Discussion section:

      “While existing marmoset identification approaches usually utilize visible markers, Radio Frequency Identification (RFID), or observation, the manual works and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters to check the identity continuously, instead of outperforming the experimenters’ role.”

      (3) Although it appears that performance based on faces and color beads was evaluated separately, this was not clearly presented, leading to confusion about whether face detection performance also benefited from color beads on the animals.

      Prediction of the different labels in the same model is independent, so the prediction of color beads is not affecting the prediction results of marmoset faces. Correct identity classification could be achieved without depending on the color beads, as we have filtered out the color beads detection class. The Results section has been modified as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Reviewer #2 (Recommendations for the authors):

      (1) I found the paper quite confusing as written. The term "model" is overused and highly conflated: there are the YOLOv8 pre-trained models, the "face classification" model, and the "automatic facial and identity extraction" model. The flowchart in Figure 2A is equally confusing. The mapping from the flowchart to the results is not straightforward, and I needed several passes to grasp it. I would recommend that the authors simplify the terminology and the mapping of the methods to the results.

      Thank you for pointing this out! The YOLOv8 pre-trained models, the "face classification" model, and the "automatic facial and identity extraction" model were indeed separately trained object detection models. They all have different weights/parameters but share the same YOLOv8 architecture/backbone. We removed some of the “model” term in the manuscript and replaced them with “classifier/framework/pipeline” to avoid misunderstanding. This information has been clarified in the revised manuscript of the Methods section and Figure 2, which provides an overall clearer explanation of the workflow of the methodology of the program.

      (2) It is not clear how robust these results are, given the limited data sets analysed.

      We agree that only five marmosets were involved in this manuscript, this unfortunately limited the robustness of the prediction results. Indeed, the limited number of animals that can be used per study has been a main limitation in non-human primate research, as they are very valuable animal models. However, we included approximately 3400 images in the training dataset, which were collected across days. New videos and photos that were captured from different devices were also used in the testing to ensure that the program can be used on new marmosets, different housing cage, and from different recording devices as indicated here:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair system, and computer-based device.”

      (3) There are two paradoxes regarding the stated motives of the study:

      (a) If the objective was to truly use non-invasive methods for the identification of animals, then why use the color-coated beads?

      As mentioned previously, identity detection can be made without collar beads, still with correct prediction results as indicated here:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Video 7).”

      The color-coated beads are used for easier and quick marmoset identification during daily care, health check, for weekend staff, training or handling.

      (b) If the objective was to achieve high identification performance, and the color-coated beads are sufficient for this purpose, then why bother with faces at all?

      Collar beads are small compared to the face, and less visible due to fur occlusion and motion blur. Moreover, it is possible that some marmosets do not have collar beads due to their young age or when involved in other procedures such as imaging scans. The collars need to be checked and changed regularly in growing marmosets and it is not always convenient (some marmosets do not support the collar, some can have sensitive skin that would lead to abrasion) thus the need to develop a facial recognition system. Furthermore, marmosets who are from other labs or in the wild might not wear a collar with colored beads, thus face is the main classifier in this model to be more generally applicable. It is highlighted here:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”

      (4) I was puzzled by the face similarity results in Figure 9. It appears that the face similarity measures were stronger (higher cosine similarity, lower Euclidean distance) for the adult data set compared to the twin data set. If so, why was it more challenging for the system to handle the twin data set?

      Face similarity can only be compared within models (therefore within adults and within twins). As this is calculated from different models, the adult face similarity cannot be compared with twins’ face similarity. It has been clarified in the Methods section as follows:

      “Statistical tests were performed only within the face classifier of each marmoset family, as embedding spaces may vary in scaling, learned features, and baseline metrics making cross-model comparison of inter-individual face similarity unreliable (Bollegala, 2017).”

      And in the Results section as follows: “We performed the statistical tests only on the face classifier for the adult marmoset family, as the twin marmoset model only involved two individuals and thus not valid for within-model statistical analysis (Table 2, 3).”

      The twin dataset aims to represent a test for the program utility in new marmosets, especially for testing if the program can still distinguish between the marmosets with similar faces. Thus, the number of twin data collected is less than the adult dataset, as explained by Discussion paragraph #2 “While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition.” This explains the challenge the system faces when differentiating the twins, while increasing the training dataset is required to solve this issue.

      Reviewer #3 (Recommendations for the authors):

      Major issues:

      (1) My main issue is regarding the utility of this method in scientific experiments. This manuscript is a "methods paper" introducing a face recognition method to identify a single marmoset in their home cage in a very specific and confined field of view. This comprises a limitation on what experiments can be performed using this method. On the contrary, if (a) the authors can show that this method can be used for a bigger field of view, where the social structure/interactions can be studied for neuroethological, cognitive or social studies that will make this method significantly more robust; or (b) design an experiment that can be performed using the current method to show that this method in its current form is sufficient.

      (a) Our method worked in larger home cage (larger view) with videos taken inside the cage / outside the cage, with multiple marmosets moving around, while the camera and its fixation are also moving. A new figure (Figure 10) has been added to highlight this wide application:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair system, and computer-based device.”.

      (b) We are currently using this method to collect in-cage touchscreen data with multiple marmosets without the need to isolate such animals to acquire the data, avoiding social separation. The collection of data in nonhuman primates is still a long process, so we wanted to share the facial recognition system first, aligned with our commitment towards open science, to benefit the broader community (we have already been contacted by two labs since the publication of this preprint) while we keep collecting data for the scientific project. We have added the touchscreen application as example in the Discussion section as follows:

      “Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.”

      (2) The authors claim a longitudinal identification of marmosets, yet I think the data to fully support this are deficient. This might be a result of unclarity of this experiment. How was this experiment done? Was the training done on the 7 months and then applied to the 11 months? Are there more continuous data that track the precision of the identification in time? For example, how does the twin identification evolve in time?

      This Tools & Resources paper mainly described the development of the marmoset identification program and methods. Ongoing work in the lab, the main research focus of which is the longitudinal assessment of cognitive functions, either during neurodevelopment or in preclinical ageing model, is benefiting from such algorithms to help identifying the animals to collect in cage behavioral data. As such, we have done some testing in one young cohort. The training of the young marmosets’ identification was done only on the 7-month data, and then we applied the identification program to the videos of the same marmosets when they were 11 months old and 16 months old (for the no-collar results) as indicated as follows in the Methods section:

      “Moreover, we evaluated the model performance and its generalization across developmental stages using new videos: 1) from the adult marmosets and 2) from the same young marmosets at 11 months, which were not involved during initial program training”.

      And in the Results section as follows: “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Video 7).”

      The identification program was shown to correctly identify the marmosets; however, we found that “While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition.”

      The mislabeling was more frequent with the young twins not only due to their face similarity, but also due to the limited number of images being used for model training. The face images used in the identification model training were less compared to the adult model, which contributed to a less accurate prediction result. As the marmoset is still developing before adulthood, their face features will become more different as they age. By increasing the number of training images from different ages of the young marmosets, this could be solved as it is therefore possible to build efficient identification program for longitudinal study. Thus, instead of the current classifiers presented in this manuscript, we suggested that the method/tool could be beneficial for longitudinal studies, not restricting to the individual-based identification program mentioned in this manuscript, as “The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.”

      (3) How does this method compare to other methods that were used in the past?

      The advantages were mentioned in paragraph #1 of the Introduction and paragraph #1-3 in the Discussion. Current approaches for marmosets are usually visible markers (ear dye, collar, etc.), RFID, or observation, of which manual works and human interventions are required. These methods usually need continuous adjustment due to tighter collar, dye fading, etc. This can affect marmoset behaviors, especially during their behavioral task performance, as mentioned as follows:

      “While existing marmoset identification approaches usually utilize visible markers, Radio Frequency Identification (RFID), or observation, the manual works and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters to check the identity continuously, instead of outperforming the experimenters’ role.”

      (4) Did the authors think of adding a continuity or a space constraint? For example, video 6 shows misidentification of the twins; in this specific case, adding a probabilistic continuity or space constraint that will limit identity switches might be useful. This can also be using a retroactive correction - for example, video 3.

      We would like to thank the reviewer for this suggestion. We agree that these approaches will be valuable improvements for future offline analysis.

      The probabilistic continuity constraint can indeed help decrease identity switches. In our current application, we have implemented a temporal smoothing through majority voting across a 30-frame (1 second) window, of which the program outputs the most frequent prediction of identity. With this strategy, we could reduce the occasional frame misprediction and maintain the real-time performance. Our animals are free-moving and may appear in any location within the camera field of view and housing cage. Therefore, position is not strongly associated with the identity of individual.

      We agree that retroactive correction could improve the detection consistency for offline analysis by correcting past detection by future prediction results. However, the current pipeline is incorporated with behavioral tasks, meaning that the prediction results aim to be transmitted with minimal time delay. As additional frame analysis and extra computational power may be needed for retroactive correction, the increased latency can be limited to the utility of real-time system and task control.

      Minor issues:

      (5) In its current form, I think the manuscript can be significantly shortened and the results/figures can be consolidated (confusion matrices with validation figures for example).

      We have followed the reviewer’s suggestion and shortened the Methods and Results sections.

      (6) The term "unseen" that the authors use in their results is confusing. Are the authors referring to monkeys that are hidden from their view, or "unseen" before by the model? The video indicates the latter, but I think the term can be changed to something less confusing, like novel, new, etc.

      We have changed the term “unseen” by “new” to avoid confusion.

      (7) Can the authors add information about the relationship between the number of manually labeled images and the identification precision?

      The relationship between number of manually labeled images and identification precision has been described in the Discussion section as: “Moreover, the performance of the system is strongly dependent on the amount and variability of the training data, with identity classification improving as more marmoset images are involved in the model training.”

      This means that more manually labeled images (i.e. larger training datasets) could improve the identification precision. However, model performance will plateau regardless of training dataset size, referred in the Results: “Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs).”

      More manually labeled images could help improve the variability of model prediction, but too many of these images are also risky for overfitting. In this case, overfitted model might not be able to make valid predictions on new videos/images.

      (8) Can the authors expand a bit about the difference between YOLO Nano, small, and medium in the methods?

      Thank you for the suggestion. We have added a brief description of the pre-trained models in the methods: “These pre-trained models share the same object detection backbone but differ in number of parameters and computing power. Larger models, such as YOLOv8 medium, provide higher detection accuracy but require greater computational resources and longer inference time. In contrast, smaller models prioritize the computational efficiency.”

      (9) Clear and short definitions of what IoU, Recall, F1, and other terms represent should be added to the results section (not formulas, short sentences).

      The definitions and formula for the evaluation metrics were described in detail in the Methods section. To help readers while avoiding repetitions with the detailed methodology definitions, we have added a brief description of these terms at their first mention in the Results section: “Model performance included the precision (the proportion of correct positive predictions), recall (the proportion of corrected predicted ground-truth labels), and mAP@50–95 (the average detection accuracy across different IoU object localization thresholds; see Methods and Materials section for detailed definitions).”.

      (10) In the methods-"video collection" section, can the authors please include more information? Is this a motion-sensitive camera? Otherwise, what's the size of the data that is collected? This will help in reproducibility and system requirements. If this is not continuously collected data, discuss what can be done to make an online identification tool.

      We have included the camera as industrial color/RGB camera, which functions like any webcam and has no motion-sensitive functions. The size of data collected was described in detail in the Methods section that we slightly modified for clarity for: “Three adult marmosets from one family were recorded for 1 hour across each of the 5 recording days, with unrestricted voluntary access to the primate chair space. For the adult marmosets, the housing cage door was opened at the beginning of the recording session allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 1C). Two young marmosets were briefly isolated and recorded separately for testing and improving the automatic face extraction program. We recorded them at two developmental time points, 7 months old and 11 months old (an additional time point at 16 months old has been added for one marmoset to test the identification without collar). During video collection, a sliding panel and an in-cage box were positioned near the housing cage door to temporarily isolate individual marmosets from other family members. Individual isolation was kept brief (approximately 10 minutes) to prevent disturbance and potential stress due to family separation. “To capture sufficient variability in postures, individuals, and lighting conditions, clips were sampled throughout the adult marmoset videos (approximately 5 hours) (Figure 2B).”

      The size of the training dataset was also described in detail in the Methods section, referred as: “To minimize image computations and data storage, we created a dataset of 2498 annotated images from the three adult marmosets. All images were manually annotated to label marmoset faces, individual identities, and the collar bead colors (Figure 2C). The annotated images were used for training models of multi-marmoset face classification and the automatic identity extraction, which can automatically detect, localize, and identify marmoset faces (Figure 2A, Step 1 – 4). We created another dataset of two young marmosets at 7 months old (total images = 502) for testing the automatic facial and identity extraction (Figure 2A, Step 4 – 5). For both adult and young marmoset datasets, images were randomly divided into a training set and a validation set at a ratio of 8:2.”

      Our manuscript is not describing an online identification tool (i.e. the described program does not require connection to internet). Instead, once trained and the program is performing well with new marmoset videos, we could use the trained weights for real-time marmoset identification (no need to collect new training data) as we are doing it for our touchscreen data collection in cage. It is referred in the Discussion as follows: “Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.”

      This program can be used for online applications if you are using a camera that is connected 24/7. For now, we are only using it while using our behavioral testing chair due to limitations issues (safety recordings, removing all electrical apparatus during night, overheating of camera if use continuously).

      (11) Would increasing the size of the beads help with their identification? In the images included in Figure 2C, it's very difficult to see these beads.

      Yes, collar beads can be occluded by fur, blurred by motions, or outside the field of camera view. We have now included increasing the collar beads size in the Discussion, referred as: “An alternate experimental solution is to improve collar visibility, including using distinct color code across individuals within a family, increasing the size of the beads, or increasing collar beads number to reduce occlusion.” However, the size of the beads needs to be appropriate to avoid being inconvenient and disruptive to the animals to ensure their welfare.

      (12) It's difficult to understand the setup of the camera in regard to the housing cage. Can the marmosets go into the primate's chair at any point (from the videos, it seems so, but the dashed line in 1C might indicate otherwise), or is the primate chair there only to mount the camera? Consider redrawing 1C in a more clear way.

      The figure 1C is a simplified drawing of the photo 1A, we have edited the drawing of Figure 1C to highlight the free access when the chair is mounted to the cage as stated here: “For the adult marmosets, the housing cage door was opened at the beginning of the recording session allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 1C).”

      (13) Given the repetitiveness of the figures, an icon atop each figure specifying what's being tested will be helpful.

      We have added an icon for each figure 3-8.

      (14) I feel the supplemental figures for Figure 9 are more compelling than the main figure. Consider including some of the panels in the main figure.

      Thank you for the suggestion. We included heat maps of the adult family relationships in Figure 9, including the cosine similarities and Euclidean distances. The current legend of Figure 9 is changed as follows: “Figures 9. Across-model visualization of the face similarity between marmoset pairs. Four types of family relationships (mother-father, father-son, mother-son, and twin1-twin2) were compared, based on the training results of adult and young marmosets. The similarity was calculated using (A) cosine similarity and the (B) Euclidean distance. Heat map of the (C) cosine similarity and (D) Euclidean distance was plotted between the 3 relationship pairs in the adult family. The cosine similarity score ranged from 0 (very different) and 1 (exactly same) for marmoset faces. The Euclidean distance score ranged from 0 (exactly same) and 1 (very different) for marmoset faces.”

      (15) Related to this, I feel that the display of Figure 9 obscures the differences that the authors report.

      The display of Figure 9 has been improved following the above reviewer’s suggestions.

      (16) In Figure 9, given the large effect sizes but non-significant p-values, will adding more training points/epochs improve the differences.

      For the trained recognition programs, we have already implemented the “early stopping criteria” as shown in the Results section: “Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs).” This means that the program has already plateau with its performance.

      (17) Line 30: either "within" or "in".

      This has been corrected.

    1. eLife Assessment

      This important study investigates how distinct honeybee viruses differentially alter flight performance through interactions with octopamine signaling pathways. The combination of behavioral flight assays, pharmacological perturbation, and transcriptomic analyses provides solid evidence that virus-specific effects on flight are associated with octopamine signaling. The data presented also establish a framework for additional analyses that may strengthen the proposed mechanistic model, including quantification of endogenous octopamine levels and receptor specificity studies.

    2. Reviewer #1 (Public review):

      Summary:

      Kaku and Flenniken investigate the mechanistic pathways through which specific viral infections alter the flight capabilities of honeybees. Building on their previous discovery that DWV impairs flight while SBV unexpectedly enhances it, the authors hypothesized that these behavioral shifts are driven by interactions with the insect's octopamine (OA) signaling pathway, which is responsible for the "fight-or-flight" neurohormonal stress response and energy mobilization. To test this, the authors experimentally infected adult honeybees with DWV or SBV and pharmacologically manipulated the OA pathway using either octopamine supplementation or epinastine (EP), an OA-receptor antagonist. They then evaluated the bees' flight performance (distance, duration, and speed) on custom flight mills and profiled their gene expression using qPCR and RNA sequencing.

      Strengths:

      A major strength of this study Is the high prevalence of preexisting background DWV and SBV infections in the honeybee cohorts, which meant there were no completely "virus-free" control groups. However, the authors successfully mitigated this limitation by rigorously quantifying viral RNA copies for every individual bee via qPCR and utilizing these viral abundances as continuous variables in powerful linear mixed-effect models.

      Weaknesses:

      The primary weakness lies in the methodology used for targeted pharmacological manipulations, as well as the lack of OA quantification across different treatments. Thus, their claims are not sufficiently supported by the current data.

      Comments on revised version.

      I appreciate the authors' efforts to address the reviewers' concerns and to revise the wording of the manuscript. The revised version is more cautious than the original, and some of the discussion has been appropriately toned down. However, I remain unconvinced that the key mechanistic conclusions are sufficiently supported by the current evidence.

      (1) The specificity of epinastine remains insufficiently demonstrated.<br /> The authors argue that AmOARβ2 is the predominantly expressed octopamine receptor subtype in their RNA-seq dataset and therefore the physiological effects of epinastine are most likely mediated through this receptor. However, I do not find this argument fully convincing.

      First, relatively low transcript abundance of other OA receptor subtypes does not exclude their physiological contribution. Even receptors expressed at lower levels may play important functional roles, particularly in specific neuronal populations or flight-related tissues. Therefore, the possibility that epinastine affects multiple OA receptor subtypes cannot be excluded.

      Second, although epinastine is widely used as a pharmacological tool to inhibit octopamine signaling, its receptor pharmacology has not been comprehensively characterized. The study by Roeder et al. primarily employed radioligand binding assays, which provide information on receptor affinity but not on functional antagonism or subtype selectivity. Without systematic functional characterization across the insect octopamine receptor family, it remains difficult to exclude contributions from other OA receptor subtypes or potential off-target effects.

      A more convincing pharmacological strategy would be to demonstrate similar results using an additional chemically distinct octopamine receptor antagonist. Concordant phenotypes obtained with two independent antagonists would substantially strengthen the conclusion and reduce concerns regarding off-target effects.

      (2) The OA supplementation experiments should be interpreted more cautiously.<br /> The authors correctly acknowledge that exogenous octopamine produces only transient elevations in signaling. However, I do not find the comparison with synthetic agonists entirely appropriate.

      Although synthetic agonists such as amitraz generally produce more prolonged receptor activation than endogenous octopamine, the more fundamental difference lies in their physicochemical properties. Octopamine is a highly polar endogenous amine that exhibits limited tissue penetration and is rapidly cleared through uptake and metabolic pathways. Consequently, exogenously administered OA is unlikely to efficiently reach relevant target tissues or receptor populations in a manner comparable to endogenous neurotransmitter release. In contrast, the greater lipophilicity of amitraz facilitates its distribution into target organs and enables more sustained receptor engagement following systemic administration.

      More importantly, the observation that OA supplementation partially rescues flight behavior does NOT necessarily establish that altered endogenous OA signaling is the primary mechanism underlying the virus-induced phenotypes. Such rescue experiments demonstrate that pharmacological enhancement of octopaminergic signaling can modulate the phenotype, but they do NOT provide direct evidence that endogenous OA levels or OA signaling are altered by viral infection. Therefore, these experiments should be interpreted as supportive rather than mechanistic evidence.

      (3) Direct quantification of octopamine remains the major missing evidence.<br /> The authors acknowledge that direct measurements of octopamine and tyramine would strengthen their conclusions but argue that technical limitations and cost prevented these analyses. While these practical considerations are understandable, they do not compensate for the absence of the critical mechanistic evidence.

      Overall, I appreciate the authors' revisions and agree that the manuscript provides interesting evidence that octopaminergic signaling is associated with virus-dependent changes in honeybee flight performance. However, I do not believe that the current data are sufficient to support the stronger mechanistic claims regarding regulation of the OA pathway or the specific involvement of the AmOARβ2 receptor.

      Unless direct measurements of endogenous OA (and ideally tyramine) can be provided, I recommend that the authors substantially moderate the mechanistic conclusions throughout the manuscript, including the Abstract, Results, and Discussion. The study should be presented primarily as evidence for a pharmacological association with octopaminergic signaling rather than as definitive proof of the proposed mechanistic model.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study investigates how distinct honey bee viruses differentially alter flight performance through interactions with octopamine signaling pathways. The combination of behavioral flight assays, pharmacological perturbation, and transcriptomic analyses provides solid evidence that virus-specific effects on flight are associated with octopamine signaling. However, some of the stronger mechanistic conclusions regarding direct regulation of octopamine signaling remain incomplete without more specific validation of receptor-level effects and direct quantification of octopamine levels or signaling activity.

      We revised some of text in the manuscript, since we agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence—including gene expression analyses, OA supplementation experiments, and behavioral measurements—that collectively support a role for octopaminergic signaling in mediating the observed effects. The revised text better reflects the data included in this paper.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Kaku and Flenniken investigate the mechanistic pathways through which specific viral infections alter the flight capabilities of honey bees. Building on their previous discovery that DWV impairs flight while SBV unexpectedly enhances it, the authors hypothesized that these behavioral shifts are driven by interactions with the insect's octopamine (OA) signaling pathway, which is responsible for the "fight-or-flight" neurohormonal stress response and energy mobilization. To test this, the authors experimentally infected adult honey bees with DWV or SBV and pharmacologically manipulated the OA pathway using either octopamine supplementation or epinastine (EP), an OA-receptor antagonist. They then evaluated the bees' flight performance (distance, duration, and speed) on custom flight mills and profiled their gene expression using qPCR and RNA sequencing.

      Strengths:

      A major strength of this study is the high prevalence of preexisting background DWV and SBV infections in the honey bee cohorts, which meant there were no completely "virus-free" control groups. However, the authors successfully mitigated this limitation by rigorously quantifying viral RNA copies for every individual bee via qPCR and utilizing these viral abundances as continuous variables in powerful linear mixed-effect models.

      Weaknesses:

      The primary weakness lies in the methodology used for targeted pharmacological manipulations, as well as the lack of OA quantification across different treatments. Thus, their claims are not sufficiently supported by the current data.

      We thank Reviewer #1 for these comments.

      (1) The authors utilize Epinastine to block octopamine signaling, describing it as a highly specific OA receptor antagonist. However, pharmacological inhibitors often lack absolute specificity. Epinastine might bind to other octopamine receptor subtypes present in honey bee neural and flight muscle tissues, or it could potentially cross-react with tyramine and dopamine receptors. Without further genetic validation (e.g., RNA interference targeting specific receptors), it is difficult to definitively conclude that the altered flight performance is solely due to the blockade of the specific Oβ−2R pathway.

      We thank the reviewer for this thoughtful comment and agree that pharmacological approaches have inherent limitations with respect to receptor specificity. However, among the available octopamine receptor antagonists, epinastine is considered one of the most selective compounds for insect octopamine receptors. Roeder et al. (1998) reported that epinastine exhibits affinities for octopamine receptors that are at least four orders of magnitude greater than those for other insect biogenic amine receptors, including dopamine, tyramine, histamine, and serotonin receptors. We updated the text to include this information.

      Honey bees encode four β-adrenergic-like receptors (AmOARβ1- AmOARβ4) and one αadrenergic-like receptor (AmOARα1). Our transcriptomic analyses indicated that expression of AmOARβ2 was substantially higher than that of other octopamine receptor genes. Specifically, AmOARβ4 transcripts were not detected in our RNA-seq datasets, while AmOARβ1 and AmOARβ3 were expressed at very low levels in most samples (Supplementary Table S9; Figure S5). Although AmOARα1 transcripts were detected in some samples, expression levels were consistently lower than those of AmOARβ2. These observations support the interpretation that the physiological effects observed following epinastine treatment are primarily mediated through disruption of AmOARβ2 signaling. We updated the text to include this information.

      We agree that receptor-specific genetic approaches would provide valuable complementary evidence. RNAi-mediated knockdown of AmOARβ2 is an attractive future direction; however, RNAi efficacy in honey bees is variable and influenced by factors including transcript turnover rates. In addition, dsRNA treatments can induce sequence-independent antiviral effects that could confound interpretation in studies involving viral infection (Flenniken and Andino PONE 2013; Brutscher, Daughenbaugh, and Flenniken Sci Reports 2017). We have revised the manuscript to more explicitly acknowledge these limitations and to clarify the basis for our interpretation of the epinastine experiments.

      (2) As a natural neurotransmitter, insects have evolved highly efficient "cleanup" mechanisms. OA is rapidly cleared from the synaptic cleft via reuptake transporters and quickly inactivated by enzymes such as N-acetyltransferase (NAT) or Monoamine Oxidase (MAO). Consequently, an injection of OA produces only a transient "pulse" of activity. It is often a poor "tool" for inducing prolonged physiological effects compared to synthetic formamidines like Amitraz.

      We thank the reviewer for this important point regarding the pharmacokinetics of octopamine. We agree that octopamine is rapidly metabolised and cleared under physiological conditions and that exogenous administration is unlikely to precisely mimic endogenous signaling dynamics. Our goal was not to induce a prolonged pharmacological activation of octopamine signaling comparable to that produced by synthetic agonists such as amitraz, but rather to determine whether increasing octopaminergic signaling could mitigate the flight impairments associated with DWV infection. Octopamine was administered either by injection or through feeding (Lines 86-89), both of which resulted in significant improvements in flight performance in DWV-infected bees (Figure 2). The observation that two independent delivery methods produced similar outcomes supports the conclusion that enhanced octopaminergic signaling can partially rescue the DWV-associated flight phenotype. We have revised the manuscript to clarify this distinction and to acknowledge that exogenous octopamine administration likely produces transient elevations in signaling rather than sustained receptor activation.

      (3) The study relies heavily on transcriptomics and quantitative PCR to measure the mRNA expression of key synthesizing enzymes, namely tyrosine decarboxylase (tdc) and tyramine βhydroxylase (tβh), to infer the activation or suppression of the octopamine pathway. However, changes in enzyme synthesis at the RNA level are often insufficient to accurately reflect the true physiological levels of biogenic amines. To robustly prove the authors' hypothesis of a "feedback loop that regulates intracellular OA concentrations", direct quantification of actual octopamine and tyramine titers in the bees (e.g., using high-performance liquid chromatography or mass spectrometry) is necessary.

      We thank the reviewer for this comment and agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. Previous studies have successfully quantified OA in honey bees using HPLC-based approaches, including KayaZee et al. (2022, eLife), who measured OA in honey bee muscle tissue (both naturally occurring levels and levels post-treatment with 10 mM OA), and Cook et al. (2017, J. Exp. Bio) who quantified OA in pooled honey bee brain samples.

      Prior to submission, we inquired with our institutional mass spectrometry facility regarding the feasibility of measuring OA in individual honey bee samples. The expected concentrations of OA in our samples was below their limit of detection, so we did not pursue these analyses at that time. During the review process, we explored the possibility of analyzing a subset of samples at external facilities that may have the sensitivity required to quantify OA and tyramine in honey bee tissues. Since such analyses would require substantial resources, with estimated costs of approximately $5,000–10,000 for 12–15 samples that have been stored in the -80C since the study, rather than flash-frozen in liquid nitrogen as described by Zee et al. 2022. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence— including gene expression analyses, OA supplementation experiments, and behavioral measurements—that collectively support a role for octopaminergic signaling in mediating the observed effects. We thank the reviewer for this valuable suggestion. While these analyses are beyond the scope of this study, we will consider using this approach in future studies.

      Reviewer #2 (Public review):

      Summary:

      This highly original and well-designed study provides insight into how honeybee picorna-like viruses, Deformed wing virus (DWV) and Sacbrood virus (SBV), affect flight performance, and reveals the role of the octopamine (OA) pathway in virus-honeybee interactions. The authors used a flight mill to quantify the flight performance of bees with different levels of DWV and SBV. Bees were treated with OA and/or epinastine (EP) - an OA receptor antagonist; the study also quantified virus loads and expression of two key genes involved in OA biosynthesis.

      The results showed that reduced flight performance associated with high DWV levels could be alleviated by OA administration. In contrast, increased levels of SBV had the opposite effect, leading to enhanced flight performance. This suggests distinct physiological responses to DWV and SBV infections. Administration of EP had led to a reduction of flight performance in SBVinfected bees, indicating the involvement of the OA pathway.

      The authors also quantified levels of mRNAs of enzymes involved in OA synthesis, tyrosine decarboxylase (TDC) and tyramine beta-hydroxylase (TbH), and concluded that DWV induced expression of TbH, while SBV upregulated expression of TDC. Furthermore, the study identified upregulated and downregulated genes in response to SBV, DWV and DWV in combination with OA.

      Strengths:

      The study reported opposing effects of infections of related viruses, SBV and DWV, on honeybee flight performance, and identified the central role of the octopamine (OA) signaling pathway in the effect of viruses on honeybee flights.

      These findings were achieved by using a combination of approaches, including experimental measurement of flight distance, virus infections, and introduction of OA and EP. Experimental work with honeybees is technically challenging and requires specialized expertise, which makes the results produced in this study more valuable.

      DWV and SBV are among the most important honeybee pathogens affecting honeybee health and threatening the pollination service. Therefore, an understanding of the mechanisms underlying DWV and SBV pathogenesis has the potential to develop novel approaches to mitigate the negative impact of these viruses.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      We thank Reviewer #2 for these comments

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions for the manuscript.

      (1) L. 45-46

      Please note that not only high virus levels have a negative impact on honeybees. Low levels of DWV, typical of covert infections, can have long-term deleterious effects on honeybee foraging and survival. Please include citation (e.g., Benaets et al, 2017, Proc Biol Sci (2017) 284 (1848): 20162149. https://doi.org/10.1098/rspb.2016.2149).

      We thank the reviewer for these comments and edited the text accordingly, and apologize for our inadvertent omission of Benaets et al 2017, which is cited in our previous publication.

      (2) L. 113

      Clarify what is meant by "high DWV levels"

      "i.e., 10^8 copies / 2 ug RNA" -> "i.e., above 10^8 copies / 2 ug RNA"

      We thank the reviewer for this comment and corrected this in the text.

      (3) L.115

      "..mock infected bees.." /Figure 2A.

      Did these bees have low levels of DWV, below 10^8 / 2 mg RNA? What was the level of DWV in these bees?

      Mock-infected bees had an average of 3x10<sup>5</sup> DWV copies and 2x10<sup>3</sup> SBV copies per 2 µg RNA (reported in Lines 107-109 in revised manuscript, a few lines before in original manuscript).

      (4) Figure 2 / Legends to Figure 2

      Note that in Figure 2 legends, the grey areas show 95% confidence intervals for regression lines.

      We thank the reviewer for this comment and added this in the text.

      (5) Figure 2 / Legends to Figure 2

      Consider including correlation coefficients (R) and p-values for each of the regression lines in Figures 2A-F. (These could be included in the Figure 2 legends).

      We thank the reviewer for this suggestion and agree that providing sufficient statistical information is important for data interpretation. Because the analyses presented in Figure 2 are based on linear mixed-effects models that incorporate both fixed and random effects, the statistical outputs are more complex than those associated with simple linear regressions. For figure clarity, we chose not to include all model statistics within the figure panels or legends and the key statistical results, including p-values and model fit metrics (R<sup>2</sup> values), are reported in the main text (Lines 121+). In addition, complete model outputs, including all relevant coefficients, correlation estimates, and associated statistics, are provided in Supplemental Data Sheet S4. To address the Reviewer’s comments, we revised the figure caption to improve clarity and include key p-values.

      We believe this approach better balances accessibility in the main figures with comprehensive reporting of the statistical analyses and thank the Reviewer for this useful suggestion.

      (7) L.357-377 - virus-specific responses

      A previous honeybee transcriptome analysis study, which showed different responses to DWV and SBV, could be cited (Ryabov E. 2016. PeerJ 4:e1591 https://doi.org/10.7717/peerj.1591).

      We thank the reviewer for this point and included this citation in line 338 of original manuscript (line 349 in revised, tracked-changes manuscript).

      (7) L. 412

      "bees were collected 24 hours prior to eclosion" -> e.g. "bees were collected at pupal stage 24 hours prior to eclosion"?

      Specify if dark-eyed pupae were collected to make sure eclosion in 24 hr.

      We thank the reviewer for making this point, and we revised the methods and results text to improve clarity.

    1. eLife Assessment

      This manuscript reports an important study in which the authors apply smFRET imaging to probe HIV-1 Env conformational dynamics in the presence of antibodies. Previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Through the cutting-edge application of smFRET imaging, the study provides convincing insights into the mechanisms of action of relevant antibodies.

    2. Reviewer #1 (Public review):

      The authors have considered a panel of antibodies that target epitopes at the gp120/gp41 interface (8ANC195 and PGT151), the fusion peptide in the gp41 domain (VRC34), and the MPER region of gp41 (DH511.2_K3 and VRC42). They also investigate 10E8.4/iMab, which is an engineered bispecific antibody that targets the MPER and the CD4 receptor. On a technical note, they have applied a double amber codon-readthrough strategy to incorporate the non-natural TCO*A amino acid, which gets labeled through click chemistry. This approach should result in less disruption of the native Env structure as compared to the peptide insertion previously used for smFRET imaging of Env. Furthermore, previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Altogether, through the cutting-edge application of smFRET imaging, the study provides novel insights into the mechanisms of action of interesting and clinically relevant antibodies.

      Comments on revised version:

      The authors have nicely responded to all of my concerns. I have no further issues.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper, Xu and co-workers unveil two distinct modes of neutralisation by gp41-targeted broadly neutralizing antibodies on HIV-1 Env. So far, it was unclear as to how the mechanism of neutralisation occurred for this subset of neutralising antibodies (that can target the fusion peptide or the membrane proximal external region of the gp41 subunit). Thanks to single-molecule FRET, the authors show that the majority of broadly neutralizing antibodies stabilize the closed Env conformation (named State 1 since the original work by Munro and colleagues PMID: 25298114). Interestingly, the bivalent 10E8.4/iMab stabilized in turn a CD4-bound open state of Env. The two modes of neutralization described for these antibodies show previously unknown allosteric mechanisms that stabilize closed and open Env conformation, stressing the importance of Env conformational dynamics and its efficiency during the process of fusion.

      Strengths:

      The article is well-written, and the figures fully depict the data in a convincing way. The authors have used smFRET, which is now established in the field as a good tool to assess Env dynamics.

      Comments on revised version:

      I am very happy with the comments, answers and the way the new manuscript is shaped after revision. I have no further questions or concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This manuscript reports an important study in which the authors apply smFRET imaging to probe HIV-1 Env conformational dynamics in the presence of antibodies. Previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Through the cutting-edge application of smFRET imaging, the study provides convincing insights into the mechanisms of action of relevant antibodies.

      We appreciate this positive assessment and thank the reviewers for their time and constructive comments. We have made the following changes in the revised manuscript to address all points raised by reviewers.

      (1) Clarify the distinction between suppression efficiency and functional cost.

      (2) Add controls: smFRET experiments in the presence of monovalent 10E8.4 and iMab individually.

      (3) All of the smFRET population contour plots have been removed, as suggested.

      (4) Repeat neutralization experiments of tagged viruses (carrying nc-AA-incorporated, amber-suppressed Env), add and compare infectivity profiles between before and after click-chemistry labeling of tagged viruses.

      (5) Add a section (Complementary views from smFRET and structural studies) to the Discussion on how these approaches complement each other.

      (6) Further clarify three prefusion conformational states identified by smFRET, the relation with previously identified States 1, 2, 3, and asymmetry, the heterogeneity of Env presentations and virion morphology, and the focus of this study.

      Please find below our point-by-point responses to the public reviews and recommendations for the authors.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors have considered a panel of antibodies that target epitopes at the gp120/gp41 interface (8ANC195 and PGT151), the fusion peptide in the gp41 domain (VRC34), and the MPER region of gp41 (DH511.2_K3 and VRC42). They also investigate 10E8.4/iMab, which is an engineered bispecific antibody that targets the MPER and the CD4 receptor. On a technical note, they have applied a double amber codon-readthrough strategy to incorporate the non-natural TCO*A amino acid, which gets labeled through click chemistry. This approach should result in less disruption of the native Env structure as compared to the peptide insertion previously used for smFRET imaging of Env. Furthermore, previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Altogether, through the cutting-edge application of smFRET imaging, the study provides novel insights into the mechanisms of action of interesting and clinically relevant antibodies.

      Thank you for the positive comments!

      In validating the functionality of the S401TAG/R542TAG Env, the authors performed infectivity assays and observed 20% infectivity as compared to wild-type (Figure S2A). However, the text equates this with "20% dual-amber suppression efficiency". This would benefit from some explanation. Why do the authors interpret infectivity as reporting on amber suppression efficiency, and not the functional cost of modifying Env, which is probably unavoidable? Or a combination of both? Is there data to suggest that 100% amber suppression would leave Env 100% functional? If so, this would be valuable to show. If not, the text should be clarified.

      We acknowledge this concern and have clarified the distinction between suppression efficiency and functional cost in this revised manuscript. The observed reduction in infectivity does not translate into functional loss; instead, it more reflects the efficiency of suppression (one of the critical limitations of applying genetic code expansion in mammalian cells). To support the preservation of Env functionality, we performed dose-response neutralization experiments of tag-free and 100% dual-ncAA-incorporated Env virions by two trimer-specific neutralizing antibodies, which exhibited similar dose-dependent neutralization sensitivity (Fig. 1D), providing stronger validation than infectivity assays. We also compared infectivity between labeled and unlabeled virions and observed no significant difference (Fig. S3B).

      We have previously discussed several limitations of amber suppression in mammalian cells when combined with smFRET viral systems (PMID: 38232732; PMID: 40716060) and, more recently, in our methodology chapters (PMID: 42349953; PMID: 42349954). In brief, orthogonal tRNA/aaRS pair–mediated amber suppression (reassigning/repurposing amber stop codons to non-canonical amino acids) of the introduced ambers in the target protein (Env in our case) must compete with the cellular translation system, particularly release factors that recognize amber codons and terminate translation. Readthrough of endogenous amber codons in virus-producing cells (in our case, HEK293T) can disrupt normal protein expression and virus production. Similarly, readthrough of pre-existing amber codons in HIV-1 ORFs other than the targeted ambers in Env can disrupt virus assembly, which we addressed by generating an amber-free provirus (PMID: 38232732). Introducing two amber codons into Env further reduces efficiency, as dual suppression requires two sequential successful suppression events within the same Env molecule.

      The authors state that the contour plots in Figure 2E reveal "dynamic sampling" of the observed FRET states. Strictly speaking, as presented, the contour plots (and FRET histograms) provide no information on dynamics per se. They indicate only the relative thermodynamic stabilities of the FRET states; transitions between states are a matter of interpretation. The TDPs, shown later in Figure 5A, nicely display the dynamics. More importantly, interpretation of the contour plots is challenging, as some seem to suggest an evolution toward lower FRET states. This is especially evident in Figures 2F and 3D, which suggest that the system evolves into a stable 0.1-FRET state (CO) after about 3 sec. Unless the authors want to conclude something from this, I would suggest that they consider removing the contour plots, since their interpretations are fully supported by the FRET histograms alone.

      We agree and have removed the contour plots, as they do not add meaningful information beyond what the histograms show.

      The data indicating that Env conformation is manipulated by 10E8.4/iMab is interesting. If I understand correctly, 10E8.4/iMab is an engineered antibody with one Fab targeting MPER and the second Fab targeting CD4. In the absence of CD4, could the difference between 10E8.4/iMab and the other MPER antibodies be due to 10E8.4/iMab being monovalent with respect to MPER binding?

      We appreciate this question. To address this, we have performed important controls: smFRET experiments in the presence of 10E8.4 and iMab individually in the absence of CD4. The results are shown in Fig. S9 in the revised manuscript, which indicates that 10E8.4 behaves similarly to other MPER-directed bNAbs we tested in this study, whereas iMab does not appear to affect the conformational populations of Env. The dual effect exerted by the bivalent 10E8.4/iMab is therefore very unexpected and thus interesting, as discussed in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Xu and co-workers unveil two distinct modes of neutralisation by gp41targeted broadly neutralizing antibodies on HIV-1 Env. So far, it was unclear as to how the mechanism of neutralisation occurred for this subset of neutralising antibodies (that can target the fusion peptide or the membrane proximal external region of the gp41 subunit). Thanks to single-molecule FRET, the authors show that the majority of broadly neutralizing antibodies stabilize the closed Env conformation (named State 1 since the original work by Munro and colleagues PMID: 25298114). Interestingly, the bivalent 10E8.4/iMab stabilized in turn a CD4-bound open state of Env. The two modes of neutralization described for these antibodies show previously unknown allosteric mechanisms that stabilize closed and open Env conformation, stressing the importance of Env conformational dynamics and its efficiency during the process of fusion.

      Strengths:

      The article is well-written, and the figures fully depict the data in a convincing way. The authors have used smFRET, which is now established in the field as a good tool to assess Env dynamics.

      We appreciate these positive comments!

      Weaknesses:

      (1) The limited controls on how click chemistry affects Env (as labelled Env HIV virions were not evaluated).

      We agree. Our previous validation focused on ncAA-incorporated Env HIV-1 virions, but not the fluorescently labeled virions. To address this, we have added infectivity results for labeled virions after click-chemistry labeling, compared with those before labeling. We did not observe any measurable difference in infectivity (Fig. S3B), indicating that the labeling procedure does not impair viral infectivity.

      We also attempted to perform dose-dependent neutralization after labeling. However, as anticipated in our provisional response, this remains technically challenging because the additional labeling and centrifugation steps substantially increase sample handling time, while the dual amber suppression system already limits virion production in cells. As a result, we were not able to obtain sufficiently robust datasets for this additional functional validation.

      Nevertheless, we have previously demonstrated real-time tracking of single click-labeled Env virions during internalization and intracellular trafficking in live cells (PMID: 38232732), providing independent evidence that click-chemistry-labeled Env retains functional competence.

      (2) Photobleaching of donor and acceptor molecules occurs right after 10sec exposure.

      We acknowledge this limitation and have included it in the revision.

      (3) Other limitations are well described in the corresponding section.

      We appreciate this comment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As a means of clarifying the mechanism of 10E8.4/iMab, the authors might consider performing separate smFRET experiments in the presence of the normal 10E8.4 antibody and the normal iMab antibody (a negative control). Alternatively, they could consider imaging in the presence of the DH511.2_K3 and VRC42 Fabs (as opposed to full-length Ig) to make a cleaner comparison, although this may be less informative given the high concentrations of antibodies used.

      We thank the reviewer for this excellent suggestion. To enable a direct comparison, we performed the most informative control by examining virus-associated Env in the presence of 10E8.4 alone and iMab alone. The corresponding smFRET results are presented in Fig. S9. We found that 10E8.4 behaves similarly to other MPER-directed antibodies, whereas iMab alone does not appear to have a notable effect on the conformational propensity of Env. For transparency and to facilitate future antibody design, we have also included the Fab region sequences of the antibodies in Table S2.

      Reviewer #2 (Recommendations for the authors):

      The article is well-written, the findings are of high interest for the community. The article should be shared once the points stated below are clarified and revised by the authors.

      We appreciate this comment and the points raised by the reviewer and have revised the manuscript accordingly.

      (1) In Figure 1C, the tomographic slices showing HIV-1 WT as compared to HIV-1 decorated with EnvBG505 S401ncAA R542nCAA are quite different morphologically. The micrographs chosen show a big particle with two capsids close to a smaller one without a capsid and at least in this plane bold (no Env incorporation) for the WT; whilst for the HIV-1 decorated with EnvBG505 S401ncAA R542nCAA no capsid is apparent in both particles, one (the right one) is very small and the right one does present a number of Envs but no apparent capsid is visible here. Please comment - perhaps it would be important to average the morphological traits of both and look at average diameter, average Env incorporation, morphology of the capsid, percentage of immature particles, capsid abnormalities (as the one shown in the upper micrograph).

      We thank the reviewer for this thoughtful comment.

      The tomograms in Fig. 1C were included to demonstrate the overall size and shape of the viral particles rather than to provide a quantitative structural comparison. HIV-1 viral particles are inherently heterogeneous, and the original slides were selected as representative examples. Following the reviewer's suggestion, we replaced the representative wild-type (Fig. 1C, top panel) and tagged virus (Fig. 1C, bottom panel) tomographic slides with those that better reflect the overall quality of each sample. To further address this concern, we refer the reviewer to the nanoparticle tracking analysis (NTA) shown in Fig. S3, which shows no significant difference in particle diameter between the wild-type and tagged viruses. In the revised manuscript, we now replace "morphology" with "shape" or "size," as these terms better reflect what our results can say.

      We agree that a quantitative analysis of capsid morphology, Env spike incorporation, and the proportion of immature particles would be informative. However, such analyses would require a substantially larger cryoET dataset, which is beyond the scope of the present study; nevertheless, it is certainly in our interest to pursue a cryoET-focused study of EnvCA interactions, with Env complexed with 10E8.4/iMab. Our primary objective is to study Env conformational dynamics by smFRET rather than viral morphogenesis or capsid maturation, whose relationship to Env dynamics remains largely unexplored. It is also worth noting that the optimal particle populations for smFRET and cryoET differ. smFRET measures the conformational dynamics of individual Env trimers and therefore selectively analyzes virions containing a single dually labeled Env trimer, whereas cryoET structural analyses typically benefit from particles with higher Env spike densities. Therefore, the particle populations favored for the two techniques are not identical.

      (2) In Figure 1D, there is a difference in neutralisation with PGT151 - how different are these two curves - how does the labelling affect neutralisation for bNAbs targeting gp41? Would it be possible to assess also the infectivity, fusion and neutralisation profiles of particles where the flurophores are included? This would be without diluting the Env for single particle analysis, but just to understand how harsh the organic reaction is and how it affects Env function (as all experiments and conclusions in the manuscript are based on labelled Env).

      Again, we sincerely appreciate these questions, which have helped us improve the manuscript. Neutralization assays for the tagged viruses were performed using the ncAA-incorporated, amber-suppressed viruses, whereas the engineered wild-type is amber-free. The differences between these two dose-response curves in the original Fig. 1D are small and within the experimental variation routinely observed under even identical conditions (same virus and same bNAb). We have repeated these experiments, and the new results are shown in the revised Fig. 1D. Although minor variations remain, the overall neutralization profiles and IC50 values are highly consistent.

      To assess whether the fluorophore labeling reaction affects Env functionality, as noted above, we have included infectivity results for labeled virions after click-chemistry labeling, compared with those before labeling. We did not observe any measurable difference in infectivity (Fig. S3B), indicating that the labeling procedure does not impair viral infectivity. We also attempted to perform dose-dependent neutralization after labeling. However, as anticipated in our provisional response, this remains technically challenging because the additional labeling and centrifugation steps substantially increase sample handling time, while the dual amber suppression system already limits virion production in cells. As a result, we were not able to obtain sufficiently robust datasets for this additional functional validation. Nevertheless, we have previously demonstrated real-time tracking of click-labelled Env virions during internalization and intracellular trafficking in live cells (PMID: 38232732), providing independent evidence that click-chemistry-labelled Env retains functional competence.

      We believe that the unchanged infectivity of labeled viruses relative to their unlabeled counterparts, together with our previously observed real-time trajectories of click-labeled virions in live cells, provides strong evidence that our labeling strategy does not measurably impair Env function.

      (3) In Figure 2E and 2G, the authors employ a three Gaussian fit approach to recover the three populations (pre-triggered - pre-fusion closed - CD4 bound open). Can you please relate these with State 1, 2 and 3 from the original article (PMID: 25298114). Comment on the possibility that more than three populations could be fitted and what this could mean - pre-triggered and partially open (one gp120 asymmetrically open) could occur? Could this labelling approach account for this asymmetry?

      Thanks for this suggestion. In this study, we compared our results obtained using the gp120-gp41 structural axis with those obtained using the referenced gp120 V1-V4 structural axis to confidently assign the FRET-identified states to the previously reported three primary populations. The referenced axis is comparable to those used in the original article (PMID: 25298114) and later confirmed using the amber-click strategy (PMID: 38232732). We observe the same structural changes from these two distinct structural angles, as probed under ligand-free conditions (Fig. 2E and 2G) and CD4-triggered open conditions (Fig. 2F and 2H).

      The pre-triggered state corresponds to State 1; the pre-fusion closed state corresponds to the symmetric State 2 (which the SOSIP-based soluble Env primarily adopts; PMID: 30971821); and the CD4-bound open state corresponds to the fully open State 3. The assignment of the FRET states observed from the gp120 V1-V4 structural axis to States 1, 2, and 3 was originally reported in two studies (PMIDs: 27795397 and 29561264). In the asymmetric trimer configuration, the State 2 FRET signal originates from the free protomer, while the other one or two protomers bind CD4 and adopt the open conformation (PMID: 29561264). The asymmetric intermediate (PMID: 29561264) was identified using a heterotrimer experimental design consisting of a mixture of wild-type and CD4-binding-incompetent D368R protomers, which was not used in the present study. Therefore, our labeling approach cannot unambiguously resolve this asymmetry.

      Regarding the possibility of more than three populations, evidence from current and previous studies (PMIDs: 25298114, 27795397, 29561264, 38232732, 30971821) strongly supports the presence of three primary states of virus-associated Env, with additional substates that can be resolved under specific triggering conditions (PMIDs: 30974085, 41326374, 39640534). The assignment of such substates requires well-controlled experimental designs (PMIDs: 30974085, 41326374, 39640534).

      We have related PT, PC, and CO to States 1, 2, and 3, and added comments on multiple states and asymmetry in the revised manuscript.

      (4) When comparing smFRET with CryoET or structure, one can see that in smFRET there are always many potential conformations for big sub-populations of Env. Indeed, there is a trend, and the addition of bNAbs (Figure 4) clearly has an impact on increasing and stabilizing a particular state as defined by the authors (e.g. PT at 45% upon addition of 8ANC195, but also 32% PC and 23% CO). I assume that when analysing single particle CryoET or single virus CryoET, one needs to discard after template matching different scenarios that do not necessarily contribute to the highest resolution and this information is not always discussed. It would be interesting to address this in the discussion as the effect on Env dynamics of adding ligands (including CD4 and 17b) is not inducing in all Envs a drastic conformational change - this could be derived from the Ka of the ligands, but also from the intrinsic Env heterogeneity in both dynamics and architecture - I think that addressing these matters in the discussion could be of interest for the community. In this regard, the transition density plots are very helpful.

      We completely agree and appreciate this insightful suggestion. We have expanded the Discussion to better address the complementary insights provided by smFRET and structural approaches. In single-particle cryoEM, we do not observe the full spectrum of Env conformations for technical reasons, not because particles are intentionally discarded to obtain only the highest-resolution structures. One reason is that open Env conformations are much more sensitive to radiation damage than closed Env. Likewise, ligand-free closed Env is more sensitive to radiation damage than a bNAb-stabilized closed Env. Thus, the outcome of an SPA cryo-EM study depends strongly on the biological question being addressed and the conformational state that is preferentially preserved under the experimental conditions. Although one could hypothetically collect much larger datasets to recover lower-abundance conformations, this would be both cost-prohibitive and unlikely to faithfully represent the relative conformational populations due to differential, conformation-dependent radiation damage.

      We agree that the smFRET data highlight an important aspect of Env dynamics. Ligands, including bNAbs, CD4, and 17b, generally shift the conformational equilibrium toward particular states rather than driving all Env trimers into a single conformation. This likely reflects both differences in ligand binding properties and the intrinsic conformational heterogeneity of Env. We therefore believe that structural studies and smFRET provide complementary information. Structural methods resolve the molecular architecture of individual conformational states at atomic (by cryoEM) and near-atomic (by cryo-ET) levels, whereas smFRET quantifies their relative populations and dynamic interconversion. As the reviewer pointed out, the transition density plots are particularly valuable in illustrating these dynamic changes.

      We have incorporated these points into the revised Discussion (Subtitle: complementary views from smFRET and structural studies in general).

      (5) One of the very interesting findings of the paper is that the effect of bNabs (at least the ones tested) have an impact on the Env dynamics and how this shift can alter entry - therefore the structural view is perhaps less important - In spite of this, we still employ a structural jargon to refer to "Env open conformation stabilisation" for instance - even if the data shows that upon ligand exposure dynamics are still important but shifted. Please comment.

      This is a great point. Our smFRET data show that bNAb binding generally shifts the conformational distribution toward and stabilizes particular Env states by lowering their free energy and increasing their occupancy. We have clarified that ligand-induced stabilization reflects a redistribution of the conformational ensemble while preserving the intrinsic dynamic nature of Env.

    1. eLife Assessment

      This important study investigates the role of Nav1.7 voltage-gated sodium channels in regulating excitability of human dorsal root ganglion (hDRG) neurons. The authors characterize a previously identified Nav1.7 channel inhibitor using recombinant channels and human neurons. The study convincingly shows that inhibition of Nav1.7 channels with AM-2099 causes a modest decrease in neuronal excitability and prolongs the refractory period. This work offers new insights into the mechanism of clinically relevant pharmacological targets for pain relief.

    2. Reviewer #1 (Public review):

      Summary:

      Fujita and colleagues investigated two selective peripheral nerve voltage-gated sodium channel inhibitors targeting either Nav1.7 or Nav1.8 on excitability of human dorsal root ganglion neurons. The authors discovered that Nav1.8 inhibition is more effective at suppressing repetitive firing of DRG neurons and this may explain the greater clinical efficacy observed for suzetrigine.

      Strengths:

      The study is interesting and the findings are conceptually satisfying in that they may explain one aspect of Nav1.7 vs Nav1.8 targeting success.

      Weaknesses:

      (1) The use of postmortem human DRG neurons provides translational relevance, but the use of these cells is also a liability given their high degree of variability. Of note are the 10 to 20-fold differences in baseline properties among cells, which dwarfs the effects of the test compounds. The experiments may suffer from under sampling.

      Comments on revised version.

      The revised manuscript addresses my prior concern with reasonable effort given the limitations of human postmortem DRGs.

    3. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Fujita/Jo/Stewart/Osorno et al., investigate the contribution of Nav1.7 in regulating the excitability and firing properties of human dorsal root ganglion (hDRG) neurons in vitro. The authors characterize the effects of a previously reported Nav1.7-selective blocker AM-2099 in recombinant human Nav1.7 channels and in cultured hDRG neurons from postmortem organ donors. The authors observed modest changes in many of the properties expected by inhibiting Nav channels, including decreased action potential upstroke rate and amplitude, while increasing the voltage and current thresholds for spike generation. However, AM-2099 did not change the maximum number of APs in response to suprathreshold stimulation, leading the authors to conclude that Nav1.7 inhibition alone has limited efficacy in reducing the firing properties of hDRG neurons at the soma, and discuss that the effects of Nav inhibition may be different at distal axons.

      Strengths:

      Experiments are well-designed and executed, and the results presented are convincing. The focus on voltage-gated sodium channels in native human DRG neurons is highly relevant to recent efforts to develop safer analgesic options for chronic pain in people.

      Comments on revised version.

      The authors have done an excellent job addressing my prior critiques.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Fujita and colleagues investigated two selective peripheral nerve voltage-gated sodium channel inhibitors targeting either Nav1.7 or Nav1.8 on the excitability of human dorsal root ganglion neurons. The authors discovered that Nav1.8 inhibition is more effective at suppressing repetitive firing of DRG neurons, and this may explain the greater clinical efficacy observed for suzetrigine.

      Strengths:

      The study is interesting, and the findings are conceptually satisfying in that they may explain one aspect of Nav1.7 vs Nav1.8 targeting success.

      Weaknesses:

      (1) The use of postmortem human DRG neurons provides translational relevance, but the use of these cells is also a liability, given their high degree of variability. Of note are the 10 to 20-fold differences in baseline properties among cells, which dwarf the effects of the test compounds. The experiments may suffer from undersampling.

      We have added data from an additional 3 donors for the key results on increase in threshold and reduction of action potential upstroke, more than doubling the number of neurons for this data. We have also added a Supplementary Figure (Figure S2) that breaks out the effects on these parameters for each donor. This illustrates that there is a high degree of neuron-to-neuron variability in the effect of inhibiting Nav1.7 channels even within a single donor, even though we confined data to neurons that were verified to be capsaicin-sensitive. We also now note that there is a similar high degree of cell-to-cell variability in relative functional expression of Nav1.7 and Nav1.8 channels in capsaicin-sensitive mouse DRG neurons.

      (2) A potential confounder when using post-mortem human DRG neurons is heterogeneity of cell types. The methods clearly state that the cells selected for recording were of 'generally' small size, but specific criteria for what constitutes 'small' or other unstated selection criteria were not provided. A table of individual cell capacitance and input resistance values, along with information about individual donors (age, sex, ethnicity), is important to include. Additionally, some discussion of how DRG neuron heterogeneity impacts the findings. This relates to concern #1 about sample size determination and how cell heterogeneity factored into this calculation.

      We have added a figure (Figure S1) showing histograms and box plots of individual cell capacitance, input resistance, resting potentials, and maximum upstroke. We have also added a table with the information about donors (Table S1). As noted, we have also added a figure (Figure S2) that breaks out the effects on these parameters for each donor, illustrating that there is a high degree of neuron-to-neuron variability in the effects even within a single donor. We have added several sentences to the Discussion concerning the neuron-to-neuron variability in the effects of Nav1.7 inhibition, including the possibility that this may reflect heterogeneity of cell function.

      Reviewer #2 (Public review):

      Summary:

      The authors examine the functional role of Nav1.7 voltage-gated sodium channels in human sensory neuron electrogenesis using a Nav1.7 selective inhibitor and human dorsal root ganglion neurons obtained from organ donors. Patch-clamp electrophysiology is used at physiological temperature to measure the impact of Nav1.7 inhibition on sensory neurons' action potential firing. This is an important topic as Nav1.7 and Nav1.8 have been identified as therapeutic targets for the treatment of pain, but there has been mixed success with isoform-specific inhibitors in clinical trials. The data suggest that Nav1.7 and Nav1.8 have overlapping yet complementary functions in nociceptor neurons and that targeting both may be most effective for reducing nociception.

      Strengths:

      The data are of high quality. Action potential properties are measured at 37 degrees Celsius. Threshold is measured using brief pulses. The Nav1.7 inhibitor has been reported to be highly selective for Nav1.7 over Nav1.8 and moderately selective for Nav1.7 over Nav1.1 and Nav1.6. Data are collected using identical conditions and protocols to a previous study on the role of Nav1.8 in similar neurons.

      Weaknesses:

      The study relies on a single Nav1.7 inhibitor that has not been extensively characterized. One prior study indicates that the IC50 is around 140 nM, thus the 600 nM concentration used in this study could be predicted to reduce Nav1.7 currents by 80%. However, there is no voltage-clamp data in the current study to confirm this, and therefore, it is unclear if the batch of AM-2099 is as potent as reported in the paper that initially described its selectivity. The impact of Nav1.7 inhibition is compared to data from a previous study by this lab, and this is a minor concern. It would have been interesting to see if the combined inhibition of Nav1.7 and Nav1.8 completely blocked action potential generation in the human DRG neurons.

      We have done experiments to directly characterize the potency of the AM-2099 sample we used on both cloned human Nav1.7 channels and on native currents in the DRG neurons. Using a stable cell line expressing human Nav1.7 channels, we determined dose-response curves at both 22°C and 37°C, using an automated patch clamp instrument. These results are shown in a new Figure 1. Interestingly, we found that the IC<sub>50</sub> is substantially higher at 37°C than at room temperature. We also did experiments quantifying the effect of 100 nM and 600 nM AM-2099 on native sodium currents in the human DRG neurons, which align well with the results on the cloned Nav1.7 channels in suggesting that at 37°C, 600 nM AM-2099 inhibits Nav1.7 channels by about 85%.

      We have also added a new figure (Figure S3) showing the effects of a different Nav1.7 inhibitor, PF04856264. The effects of this inhibitor were qualitatively identical but quantitatively smaller than those of AM-2099. When we realized this, we did voltage clamp experiments on cloned Nav1.7 channels and discovered that the potency of PF-04856264 at 37°C was weaker than expected from the published IC<sub>50</sub>, which was determined at room temperature.

      We are currently doing experiments testing combined inhibition of Nav1.7 and Nav1.8 channels whenever we can obtain human neurons. Because it is of interest to examine effects of partial as well as full inhibition of each channel type, there are multiple permutations of inhibitors combined and alone that are of interest to characterize, and these studies are still in progress. We think the results in the present manuscript stand on their own and together with previous data on effects of Nav1.8 inhibitors alone provide a foundation for on-going and future studies on combinations of inhibitors by ourselves and others.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Fujita/Jo/Stewart/Osorno et al. investigate the contribution of Nav1.7 in regulating the excitability and firing properties of human dorsal root ganglion (hDRG) neurons in vitro. The authors characterize the effects of a previously reported Nav1.7-selective blocker AM-2099 in cultured hDRG neurons from postmortem organ donors. The authors observed modest changes in many of the properties expected by inhibiting Nav channels, including decreased action potential upstroke rate and amplitude, while increasing the voltage and current thresholds for spike generation. However, AM-2099 did not change the maximum number of APs in response to suprathreshold stimulation, leading the authors to conclude that Nav1.7 inhibition alone has limited efficacy in reducing the firing properties of hDRG neurons and that Nav1.7 blockers may have limited efficacy as analgesics. This is surprising, given that patients with loss-of-function mutations in Nav1.7 suffer from congenital insensitivity to pain. While it may indeed be true that pharmacological inhibition of Nav1.7 is unlikely to produce analgesia, the present study was limited to a single concentration of AM-2099. The manuscript would be significantly strengthened by a more careful and thorough pharmacological characterization of this compound, which has not been widely used or validated in native human DRG neurons.

      Strengths:

      Experiments are well-designed and executed, and the results presented are convincing. The focus on voltage-gated sodium channels in native human DRG neurons is highly relevant to recent efforts to develop safer analgesic options for chronic pain in people.

      Weaknesses:

      Only a single concentration of AM-2099 was used for all experiments. This compound was reported to be selective for cloned human Nav1.7 channels in heterologous systems, but has not been validated in other studies after the original publication in 2016. Since the original study reported a substantial statedependent block of recombinant Nav1.7 channels, more detailed pharmacological characterization of AM-2099 is needed in human DRG neurons to fully support these claims. This study would be significantly strengthened by the inclusion of dose-response curves to assess how much of the sodium current is inhibited at this concentration, confirming selectivity in hDRG, and whether maximal inhibition of Nav1.7 still has limited efficacy in reducing the firing of native human sensory neurons.

      We have added results from experiments to directly quantify the potency of AM-2099 on both cloned human Nav1.7 channels (new Figure 1) and on native currents in the DRG neurons (new Figure 2). These show that 600 nM AM-2099 produces about 85% inhibition of Nav1.7 channels at 37°C. We have added a paragraph to the Discussion explaining that we chose this concentration of AM-2099 to produce reasonably complete inhibition of Nav1.7 channels while minimizing potential inhibition of a component of non-Nav1.7 TTX-sensitive current.

      With regard to the broader point about reconciling the variable and sometimes relatively modest effects of Nav1.7 inhibition with the complete loss of pain sensation in humans with loss-of-function mutations, we have modified the Introduction and Discussion to eliminate any implication that the results in the manuscript suggest that pharmacological inhibition of Nav1.7 is unlikely to produce analgesia. Our experiments are only on action potential firing in the cell body, and it is perfectly possible that inhibiting Nav1.7 channels in the axon could disrupt generation or propagation of action potentials, either in the main axon or in the fine axon terminals in the spinal cord. We have modified the Discussion to explicitly point this out, which would reconcile the loss of pain sensation in humans with loss-offunction mutations with the incomplete effects of Nav1.7 inhibitors on excitability of cell bodies.

      Recommendations for the authors:

      Reviewing Editor comments:

      In addition to the points noted in the eLife assessment summary above, the study has several important strengths, including use of human primary neurons and recordings performed under physiologically relevant conditions (at 37 {degree sign}C using brief current injections). However, reviewers also identified several key weaknesses that must be addressed to support the central conclusions. In particular, multiple reviewers raised concerns regarding the lack of voltage-clamp data evaluating the efficacy and specificity of AM-2099 inhibition of Nav1.7 currents. A single dose of 600 nM was used based on the report of Marx (2016) in recombinant systems (Marx, 2016). Since no other studies other than the single Amgen report exist on this compound, it is important to validate its effects directly in the human DRGs used here. Additional concerns include the lack of dose-response analysis, as well as the large variability in baseline properties, which complicates the interpretation of the results. To assist the revision of the study, we outline below the key issues that should be addressed.

      Recommendations for authors:

      (1) Add voltage-clamp experiments to directly measure Nav1.7 current inhibition by AM-2099 in hDRG neurons. Given the limited previous characterization of this compound, it is important to confirm that the concentration used here (600 nM) effectively blocks Nav1.7 currents in the native system used here.

      (2) Related to point 1 above, perform a dose-response of AM-2099 on hDRGs on Nav1.7 currents in human DRGs. Since this study, at least in part, is framed as a comparative analysis of Nav1.7 vs Nav1.8 channel subtypes in DRGs, it seems important to establish pharmacological equivalence to ensure that the comparisons are made at functionally comparable levels of channel block.

      We have added results from experiments to directly characterize the potency of AM-2099 on both cloned human Nav1.7 channels and on native currents in the DRG neurons. Using a stable cell line expressing human Nav1.7 channels, we determined dose-response curves at both 22°C and 37°C, using an automated patch clamp instrument. These results are shown in a new Figure 1. Interestingly, we found that the IC<sup>50</sup> is substantially higher at 37°C than at room temperature. We also did experiments quantifying the effect of 100 nM and 600 nM AM-2099 on native sodium currents in the human DRG neurons, which align well with the results on the cloned Nav1.7 channels in suggesting that at 37°C, 600 nM AM-2099 inhibits Nav1.7 channels by about 85%.

      (3) Reviewer 1 notes that there seem to be 10-20-fold differences in baseline firing properties, which would exceed the effects of the test compound. This raises concerns about undersampling. Additional analysis or experiments would strengthen the conclusions.

      We have added data from an additional 3 donors for the key results on increase in threshold and reduction of action potential upstroke, more than doubling the number of neurons for this data. We have also added a Supplementary Figure that breaks out the effects on these parameters for each donor. This illustrates that there is a high degree of cell-to-cell variability in the effects even within a single donor, even though we confined data to neurons that were verified to be capsaicin-sensitive. Reviewer 1 made the excellent suggestion that because of the neuron-to-neuron variability in baseline properties, the effects of compounds could be better illustrated by displaying changes from baseline. Following this suggestion, we have added Tukey-style box plots displaying the data in this way. Together with the donor-to-donor breakout of data in the new Figure S2, these plots make it clear that the neuron-to-neuron variability reveals genuine differences in the channel make-up of each neuron and not experimental error.

      (4) Reviewer 2 notes an interesting experiment: does a combined block of Nav1.7 with the AM compound and Nav1.8 block action potential generation? If Nav1.7 controls threshold and Nav1.8 controls firing, then the combined inhibition should be highly effective in blocking nociceptive output, which could have therapeutic relevance.

      We are currently doing experiments testing combined inhibition of Nav1.7 and Nav1.8 channels whenever we can obtain human neurons. Because it is of interest to examine effects of partial as well as full inhibition of each channel type, there are multiple permutations of inhibitors combined and alone that are of interest to characterize, and these studies are still in progress. We think the results in the present manuscript stand on their own and together with previous data on effects of Nav1.8 inhibitors alone provide a foundation for ongoing and future studies on combinations of inhibitors by ourselves and others.

      Reviewer #1 (Recommendations for the authors):

      Concerns in addition to those in the Public Review:

      Major:

      (1) As per point 1 of the weaknesses in the Public Review, I'm concerned that the experiments suffer from undersampling. This requires a discussion of how the sample size was determined.

      We have added experiments from an additional 3 donors to the key results in Figures 3-5, more than doubling the number of neurons for these measurements.

      (3) The effect of compounds could be better displayed as a change from baseline in Figure 1C-E. Also, are the AP traces and phase plots shown in Figures 1AB and 2AB averages or representative?

      Thanks for this excellent suggestion. We have added box-plots that show changes from baseline for the various parameters. We have also clarified that the action potential traces and phase plots are from application of AM-2099 in a single representative neuron.

      Minor:

      (1) Provide source of VX-548 and report the purity of both compounds.

      We have provided the information for VX-548 and added the information on the purity of both compounds

      (2) Clinical failures of Nav1.7 blockers may not be solely due to pharmacodynamic limitations as implied by this study. Pharmacokinetic differences and toxicity (e.g., effects on the autonomic nervous system) may also have contributed.

      Thanks for raising this important point. We have added this point to the Introduction.

      Reviewer #2 (Recommendations for the authors):

      It is an interesting study, and the conclusions are reasonable. However, it would have been good to see validation of the potency of AM-2099 on native DRG sodium currents and/or recombinant human Nav1.7 channels expressed in a heterologous expression system.

      We have added results from experiments to directly quantify the potency of AM-2099 on both cloned human Nav1.7 channels (new Figure 1) and on native currents in the DRG neurons (new Figure 2).

      Minor comments:

      (1) Page 3, middle paragraph - there is a "(" missing before Renganathan.

      Thanks, corrected.

      (2) Page 4: Is anything known about AM-2099 in terms of state-dependence? It seems like Marx 2016 is the only previously published study using it, so additional information on the inhibitor would be helpful.

      We have not characterized the state-dependence of AM-2099, but we characterized its potency in voltage clamp using holding voltages similar to the average resting potentials of the cells in current clamp conditions.

      (3) Page 6 discusses that there might be differences between human and rodent DRG neurons in terms of Nav1.7 and Nav1.8. It would be nice if this were directly tested with these same Nav1.7 and Nav1.8 inhibitors.

      We have recently done such a study on mouse DRG neurons which has just been published (J Physiol. 604:6104-6127, doi: 10.1113/JP290574).

      (4) Figure 2A, right panel: I could not figure out the difference between the red and green traces. Perhaps this could be explained in the figure legend?

      Thank you for pointing out that this was confusing. These two traces showed two different subthreshold responses, one of which was slightly regenerative without generating a full-blown spike. We have simplified the figure by now showing only a single subthreshold response.

      Reviewer #3 (Recommendations for the authors):

      (1) The conclusion that Nav1.7 inhibition has limited efficacy for inhibiting the firing of human DRG neurons is not fully supported by the data. This may be true, but it cannot be concluded without a more thorough pharmacological characterization of this compound. Dose-response curves and experimental confirmation of Nav1.7 selectivity (maybe just total Nav current, TTX-sensitive and TTX-resistant components) are needed.

      We agree and have now added two new figures with this data.

      (2) How was the 600 nM concentration chosen? Given that AM-2099 was reported to exhibit state dependent block, how much of the Na current is inhibited by this concentration at the initial voltage used in current clamp experiments (~-80 mV)?

      We have added a paragraph to the Discussion recognizing the limitation that 600 nM AM-2099 produces ~85% rather than complete inhibition of Nav1.7 current and explaining that we chose this concentration of AM-2099 to produce reasonably complete inhibition of Nav1.7 channels while minimizing potential inhibition of a component of non-Nav1.7 TTX-sensitive current.

      (3) It appears that the effects of AM-2099 on the refractory period are bimodally distributed, where neurons that recovered more slowly at baseline were preferentially affected by AM-2099 (Figure 4). Do these reflect different neuronal populations (e.g. smaller or larger diameter DRG) or different resting voltages in these experiments?

      We agree that there seem to be two groups based on initial refractory period. Examining the parameters for the cells, there is no clear correlation between the effects of AM-2099 on the refractory period with resting potential or cell diameter. At this time, it is not obvious what determines the differences in refractory period. We speculate that neuron-to-neuron differences in the potassium conductances that generate the after hyperpolarization may be different in these cells but it will take further work to explore this.

      (4) How much of the sodium current is mediated by Nav1.7 in hDRG neurons? How does inhibition of both Nav1.7 and Nav1.8 affect hDRG excitability?

      The new Figure 2 shows data quantifying the AM-2099-sensitive current in the DRG neurons. With regard to combined Nav1.7 and Nav1.8 inhibition, we are currently doing experiments examining inhibition of excitability by combined Nav1.7 and Nav1.8 inhibition, which we agree is the logical next step in exploring how the two components of current control excitability. These are still in progress. Because designing and interpreting these experiments is facilitated by the current experiments with Nav1.7 inhibition alone, we believe that reporting the current results now will serve the scientific community better than waiting to obtain and interpret a body of data on dual inhibition in a sufficient number of donors, which we obtain only sporadically.

      (5) Donor information and soma diameters should be included. Capsaicin sensitivity testing was mentioned in the methods, but I was unable to find any inclusion of these data in the results. These may be useful to potentially infer effects in different cell types.

      We have added a figure (Figure S1) showing histograms and box plots of individual cell capacitance, input resistance, resting potentials, and maximum upstroke. We have also added a table with the information about donors (Table S1). We have now clarified that data were confined to cells verified to be capsaicin-sensitive and that ~95% of all cells tested were capsaicin-sensitive.

      (6) Please check statistical tests and reporting. Several graphs do not appear to have paired responses (e.g. Figure 1E, Figure 4B). As a result, two-tailed Wilcoxon tests would not be appropriate. Also, check reported p-values (e.g. p=.0002), which are identical for multiple panels in the Results section.

      We have clarified that the symbols of action potential width in control without a corresponding value after AM-2099 represent neurons in which the action potential in AM-2099 had a peak < 0 mV. These cells were not included in the data set of paired parameters used for the Wilcoxon test. We have also checked and verified all statistical tests.

    1. eLife Assessment

      A computational model potentially provides important new insights into the circuit mechanisms underlying navigational control in insects. The authors compare high speed video recordings of ants with detailed predictions from a new computational model. The similarities between model and behavioral data are striking and convincingly suggest how complex behavioral motifs can emerge from a simple neural circuit.

    2. Reviewer #2 (Public review):

      The paper by Freas and Wystrach is an interesting computational study, exploring the detailed mechanisms of how simple neural circuits could explain complex behavioral patterns observed in navigating ants. The authors compare detailed, high speed video recordings of Australian desert ants (Melophorus bagoti) with predictions made by their new computational model and find convincing similarities between the model and the behavioral data, at a level of detail not previously studied. Particularly interesting are emerging properties of the model, yielding behavioral motifs it was not designed to reproduce, but which occur in natural ant behavior.

      A strength of the study is that the model is based on previous models, without making major novel assumptions. It combines existing models of the insect central complex with a model of the lateral accessory lobe and adds a stochastic inhibition of forward velocity to the interaction of central complex and lateral accessory lobes. In essence, the central complex provides corrective steering signals when the goal direction and the current heading of the insect are not aligned, while the lateral accessory lobes provide an intrinsic oscillator underlying the behavioral oscillations shown by walking ants at all times. These background oscillations are modulated by the steering signals from the central complex. Depending on which phase of the intrinsic oscillations coincides with the corrective signals, and how fast the ant is moving forward during this time, a complex set of behaviors emerges.

      Most prominently, scanning behaviors, which are regularly carried out by the ants, are recapitulated in great detail by the model. Additionally, other behaviors, such as full loops, emerge naturally from the model. While computational models are not to be seen as definite evidence for any biological reality, they can provide strong support for particular neural implementations. The current study is an excellent example in that it provides evidence for a serial arrangement of central complex circuits upstream of the lateral accessory lobe circuits, modulated by speed regulating input. While the latter is hypothetical, it yields a clear hypothesis that can be validated by connectomics studies and functional work in the future.

      The computational model is explained in detail and information about all model parameters is provided in an accessible way. The approach is thus transparent and reproducible, leaving it to the readers to assess the assumptions made in the model and how the studied complex behaviors emerge. This also provides the possibility to combine this new model with existing models to expand the scope and to more comprehensively capture the behavioral repertoire of ants, and insects in general.

      Importantly, the study shows that even complex behavioral motifs do not require dedicated neural modules, but can rather emerge from the interplay of already known circuits - highlighting the efficiency of insect brains and possibly providing the path towards embodied hardware solutions of such circuits in autonomous agents.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Freas and Wystrach present a computational and experimental study of ant navigation. The main innovation of the computational model is the insertion of an oscillatory element between the steering signal and the motor control that results in a trajectory whose heading oscillates around a goal direction. Additionally, the model imposes periodic cessations of forward movement and inversely couples rotational speed to forward velocity. As a result the model periodically makes larger reorientations reminiscent of those seen in behaving ants.

      The behavioral data consists of two experimental sets: experienced Melophorus bagoti foragers, recorded in 2010 and inexperienced M. bagoti foragers, recorded in 2023-2024 at the same site. The behavioral data is qualitatively compared to the model in Figures 3 through 6. In figures 3-5, all ant sets are grouped together while in Figure 6 they are separated. In Figure 6, the authors should do a careful job of making sure the reader is aware that comparisons are being made between behavioral data sets captured more than a decade apart and of justifying the validity of a quantitative comparison between these sets.

      We now make explicit in the methods and figure that the two datasets were recorded at different times: experienced foragers from 2010 (Deeti et al., 2023) and inexperienced foragers from 2023. Their comparison is used to test a qualitative difference predicted by the model.

      The manuscript also describes Myrmecia ants and makes comparisons between modeled Myrmecia ants and supplemental videos of these ants (Videos 3,4). These videos are not described in the methods. While the captions describe these as ants "homing in an unfamiliar environment," the videos show tethered ants walking on a ball. Without more information and absent any analysis, it is difficult for me to understand how these videos support granular points in the text about coupling between rotation and forward velocities.

      We have added a description of Videos 3 and 4 to their respective captions and now state explicitly that these clips were recorded from ants on a tethered trackball apparatus and that they provided as qualitative examples of behaviour discussed in the text.

      Strengths:

      The manuscript's main thesis, that an oscillatory element interspersed between the control signal and the motor unit can reproduce aspects of ant navigation, appears supportable.

      Weaknesses:

      Qualitative agreement between aspects of a model and aspects of a behavioral measurement do not prove the correctness of a model. In the section (802), "An ancestral design? Striking parallels with crawling Drosophila larvae," the authors argue that behavioral data in larvae support their model, despite the larva's lack of a (known) central complex. C. elegans navigation can also be segmented into longer runs and shorter exploratory behaviors (Chen 2025), comparable to the runs and scans described here. C elegans definitively does not have a central complex. In general, multiple internal mechanisms are capable of producing the same macroscopic behavioral outcome. This fact limits the ability of behavioral data to confirm the details of a particular model; it does not imply that observation of similar behaviors in multiple species shows that a particular model is correct or generalizable.

      Here the ability of the behavioral data to confirm or constrain the model is further limited by the qualitative nature of the comparisons. Some of the comparisons are trivial (e.g. Figure 5E-F: any first order process will produce a Poisson distribution, and in the model a Poisson process was explicitly coded in with parameters chosen (1070) to match the behavioral data). Finally, the number of adjustable parameters (13) is comparable to the number of comparisons made; it is unclear that the model could not be adjusted to fit any set of behavioral measurements.

      Our model is a minimal neuro-mechanical model. It is not a mathematical model where each parameter can be optimised to a final output.

      From our 13 parameters, 10 were either taken directly from independent prior studies. 5 concern the oscillator, and have been arbitrarily chosen and simply need to produce regular oscillation (as explained in supplemental material). 4 are the necessary motor gain and noise, which scale neural activations values into movement units, note that this conversion is backed up by previous evidence in drosophila and present in previous ant models). 1 parameter specifies the width of the bump of activity in the CX, and is roughly matched to neural imaging data in flies. None of these parameters have been introduced or adjusted to back up our claim.

      Only 3 parameters have been added to produce scannings. Two of them were tuned to match local scanning-specific data (the probabilistic trigger to stop (p_stop) and the threshold for triggering a saccade (θ_CPG) enabling us to tune fixation duration). This level of parametrisation enables the model to reproduce realistic scans, but does not influence the qualitative predictions of this article. For instance, we agree that the probabilistic trigger producing the Poisson distribution of scan duration (Figure 5’s E) is used to parametrise the model to scanning data, which does not constitute an emerging prediction of the model. We do not include it as evidence (see ~490). Finally, the CX_output_gain, is a new parameter we invoke to implement our hypothesis that CX steering modulates the oscillator. From these three added parameters emerge the large array of behavioural signatures and predictions. These are emerging consequences of the model's architecture rather than curve-fits.

      While the introduction is improved, there is still room to eliminate confusion as to what aspects of the model reflect hypothesized rather than measured neural circuits. For instance, if there is data showing LAL oscillations in insects, the authors should cite it and call it out clearly.

      Alternatively they should say that the oscillator is hypothesized based on measured bistability. They should also clarify whether they are discussing neural oscillations or motor oscillations and whether these oscillations are measured, modeled, or hypothesized.

      As one example: Lines 283-284 "This oscillator [referring to the model's intrinsic oscillator described in the previous paragraph], which is widespread in insects (Cheng, 2024; Kanzaki, 2005; Kanzaki and Mishima, 1996), resides in the lateral accessory lobes (LAL)" reads as though it is known that a neural oscillator occupies the LAL. Cheng 2024 is a brief review of behavioral oscillation. Kanzaki et al. 2005 describes numerical modeling and simulation with a physical robot. Kanzaki and Mishima, 1996 demonstrates bistability (flip-flopping) in moth descending neurons. None of these show neural oscillations and none of them describe the LAL. The authors should review the paper and be scrupulously careful that the claims made in the text are supported in the cited references. These difficulties were pointed out in a previous round of review; hopefully they can be fully corrected this time.

      Kevin S. Chen, Jonathan W. Pillow*, Andrew M. Leifer*, "State-switching navigation strategies in C. elegans are beneficial for chemotaxis," arXiv:2508.00191 31 July 2025.

      We have softened the text to be more explicit about what is modelled versus what is neurally shown (labelled each as behavioural, electrophysiological, or modelled)

      Reviewer #2 (Public review):

      The paper by Freas and Wystrach is an interesting computational study, exploring the detailed mechanisms of how simple neural circuits could explain complex behavioral patterns observed in navigating ants. The authors compare detailed, high speed video recordings of Australian desert ants (Melophorus bagoti) with predictions made by their new computational model and find convincing similarities between the model and the behavioral data, at a level of detail not previously studied. Particularly interesting are emerging properties of the model, yielding behavioral motifs it was not designed to reproduce, but which occur in natural ant behavior.

      A strength of the study is that the model is based on previous models, without making major novel assumptions. It combines existing models of the insect central complex with a model of the lateral accessory lobe and adds a stochastic inhibition of forward velocity to the interaction of central complex and lateral accessory lobes. In essence, the central complex provides corrective steering signals when the goal direction and the current heading of the insect are not aligned, while the lateral accessory lobes provide an intrinsic oscillator underlying the behavioral oscillations shown by walking ants at all times. These background oscillations are modulated by the steering signals from the central complex. Depending on which phase of the intrinsic oscillations coincides with the corrective signals, and how fast the ant is moving forward during this time, a complex set of behaviors emerges.

      Most prominently, scanning behaviors, which are regularly carried out by the ants, are recapitulated in great detail by the model. Additionally, other behaviors, such as full loops, emerge naturally from the model. While computational models are not to be seen as definite evidence for any biological reality, they can provide strong support for particular neural implementations. The current study is an excellent example in that it provides evidence for a serial arrangement of central complex circuits upstream of the lateral accessory lobe circuits, modulated by speed regulating input. While the latter is hypothetical, it yields a clear hypothesis that can be validated by connectomics studies and functional work in the future.

      The computational model is explained in detail and information about all model parameters is provided in an accessible way. The approach is thus transparent and reproducible, leaving it to the readers to assess the assumptions made in the model and how the studied complex behaviors emerge. This also provides the possibility to combine this new model with existing models to expand the scope and to more comprehensively capture the behavioral repertoire of ants, and insects in general.

      Importantly, the study shows that even complex behavioral motifs do not require dedicated neural modules, but can rather emerge from the interplay of already known circuits - highlighting the efficiency of insect brains and possibly providing the path towards embodied hardware solutions of such circuits in autonomous agents.

      We thank Reviewer 2 for this assessment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The paper would benefit if the authors would take a more traditional/formal approach to the presentation and interpretation of results. They should avoid drawing conclusions or presenting interpretations in the figure captions (e.g. caption to figure 6) and avoid unnecessary modifiers (e.g. just use "supports" instead of "strongly supports"). This might help correct the tendency of the manuscript to overstate or over-interpret the correspondence between the model and the data.

      Figure captions are now revised to remove interpretive discussion. We also removed unnecessary modifiers throughout the manuscript.

      We also added clarifying text to the “An ancestral design?” section to make clear that we are not implying homologous neural implementations across taxa.

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed my comments fully and I only have a few minor, mostly editorial points:

      line 142: it appears that the references should refer to goal encoding, but both references are head direction papers (one review, one research paper). The only paper showing goal encoding in the CX is Mussels-Pires et al 2024. This should be fixed to ensure the citations are not misleading.

      Changed citations.

      line 165: maybe remove 'intrinsic' to not suggest that this reflects what it known from Biology? It is clear that in the model it is an intrinsic oscillator, but it should not leave the impression that this is an established fact for the LAL

      Removed when not discussing the model.

      line 168: remove either 'a diversity' or 'key qualitative'

      Changed to reproduce multiple key qualitative…. (~Line 170)

      section: 'Neural substrate of insect navigation':

      - '... compares to output steering commands.' grammar is misleading as 'output' might be read as an adjective rather than a verb, replace with 'generate'?

      Changed.

      - as above, the only paper showing goal directions is Mussels-Pires et al 2024. Any other paper either assumes this in models (such as Stone et al and the Honkanen review) or deal with head direction encoding. Pfeiffer and Homberg, 2014 is a general review. Please ensure that references are used more accurately.

      We revised the text to clarify the specific evidence provided by each cited reference.

      - The goal direction in the CX can be updated by various pathways....' After this, behavioral and modeling papers are cited, which is misleading. None of these papers deal with the CX or the neural representation of goals. Same with the rest of the sentence, referring to MB output and PI, which is only shown in modeling, not data.

      We now explicitly distinguish between behavioural evidence and modelling evidence.

      line 210: use CX, not central complex

      Changed.

      Figure 2: I suppose all data shown are modeling data? This should be more explicit in the figure caption (a bit unclear what 'using the neural circuit model.' in the caption heading means. Maybe rephrase to: 'Schematic of neural circuit model and its outputs across navigation relevant brain regions.' (or something like that, putting model first, not brain regions)

      Changed.

      Figure 7: Axis labels in the graphs are still much too small to be read on a printed version (ensure at least 5pt font size in the actual figure on the printed page)

      Enlarged axis labels.

      Line 654: mirror, not mirrors

      Changed.

      line 816: insert 'fly' before larva, as otherwise one might assume this refers to ant larva

      Added (~line 822).

      line 819: The CX does (as we currently know) not exist in fly larvae. At least not as a brain structure, or a set of homologous neurons. There might be equivalent circuits for action selection, but they have not yet been convincingly described. I suggest to rephrase to: 'Although no present as neuropils in fly larvae, the CX and LAL....'

      Changed as suggested (~Line 830).

    1. eLife Assessment

      This valuable study advances our understanding of habituation and related potentiation behavior in a single-celled organism, Stentor. The authors provide convincing evidence for their claims in the form of extensive data on habituation behavior. The work will be of interest to cell biologists, neuroscientists, and also scientists broadly interested in behavior in non-neural organisms.

    2. Reviewer #1 (Public review):

      This interesting paper addresses the phenomenon of potentiation in single-cell habituation in Stentor coeruleus. This is an important "hallmark" of habituation that helps to establish single-cell learning as being similar to habituation in animals. Prior studies from Wood, as well as our own results, have shown that potentiation occurs in Stentor, but I have always remained a little bit skeptical that this effect was possibly just due to incomplete recovery after the first trial. When I first read this paper and saw the habituation curves for the first and second trials, such as in Figure 5, I thought, yes, that is definitely what is happening, and so is this really potentiation?

      The authors were also clearly aware of this issue and, notably, they embraced it head-on by developing an analysis that allows potentiation effects to be detected even despite failure of the cell to fully recover after the first trial. The key is their "phase portrait" that allows the learning process to be depicted as a curve capturing how learning rates and response probability evolve over time, thus allowing the curves to be compared between trials. If my interpretation was correct that so-called potentiation was just incomplete recovery, the prediction would be that the curves for two successive trials would overlap, with the first trial curve extending beyond the second one towards higher response probabilities, which would be lost in the second trial due to failure to recover fully. But the data clearly are not consistent with that idea. I think that this result is very strong and important.

      Especially nice is the approach of Figure 7C, which uses a vertical shift in the phase portrait as an indicator of potentiation. I did, however, find Figure 6 a little hard to digest at first, and I have a few suggestions about that. First, I think it would be a good idea to explicitly say which curve is the first trial and which is the second. Second, I think it would help readers if the authors could start with a cartoon that explains visually what the curves mean. For example, show a habituation curve, indicate how the slope is calculated at different parts of the curve, and then show how the slope versus response are plotted to make the phase portrait. It is all spelled out in the text, but it would help a lot of readers to see it visually, I think.

      One question I have about Figure 6 is that it looks like the specific case of ITI 1 hour ISI 2 min has some kind of pathological behavior in the second trial, despite not seeing any indication of any 'weirdness' in Figure 5. I gather that this is meant to be due at least in part to the incomplete recovery seen after the first trial, but then I don't see why this would not also be an issue for ITI 1 hour ISI 3 min. I would not require the authors to explain every anomaly, but this one stands out, and I feel it could be telling us something interesting.

    3. Reviewer #2 (Public review):

      Summary:

      The authors address habituation and potentiation in the single-celled organism Stentor in a large data set by systematically varying stimulus frequency and recovery duration. They analyze habituation dynamics on the level of single cells within a Bayesian inference framework to map out how the response probability of individual cells decays during training. Mapping out the progression of habituation quantified by learning rate versus decaying response probability, they observe different dynamics for different stimulus frequencies and recovery durations, which they reconcile with multiple time-scales governing the memory of prior training.

      Strengths:

      The authors accumulate a systematic, broad data set of Stentor habituation and potentiation, which, in combination with the Bayesian framework they developed, unfolds its power to probe underlying habituation dynamics and challenge theoretical frameworks.

      Weaknesses:

      The interlacing of theoretical framework, existing concepts and expectation, and experimental data in their narrative may challenge readers. The Bayesian inference of habituations is very successful in concluding that their variation with stimulus frequency and recovery duration points to multiple time scales of memory are involved. However, the authors' comprehensive analysis of potentiation may need more guidance to follow the authors' conclusions.

      The combination of a dynamical systems-driven hypothesis, experimental data, and statistical analysis, as put forward in this work, is immensely powerful for uncovering the mechanisms that facilitate learning, such as habituation and potentiation, in single-celled organisms.

    4. Author response:

      We thank the reviewers and editor for their thoughtful and constructive comments. Our goal was to connect theoretical work on habituation with empirical findings on intracellular habituation in Stentor coeruleus. We developed the phase-portrait analysis to provide a more formal way to examine habituation dynamics and to address the concern that apparent potentiation might simply reflect incomplete recovery. We are glad that the reviewers found this approach useful, and we hope to build on it in future work through mechanistic modeling.

      We will submit a revised version with the following changes:

      (1) We will include a supplementary figure that provides visual intuition for the habituation curves and phase portraits.

      (2) We agree that the anomalous behavior of the 2 min ISI / 1 hr ITI condition is noteworthy. This behavior arises from a subtle difference in the fitted shape of the trial 2 habituation curve: its Hill coefficient is less than 1, so the curve has no inflection point and its initial slope has nonzero magnitude. As a result, the corresponding phase portrait begins away from the x-axis, unlike the other conditions, whose Hill coefficients are greater than 1 and whose phase portraits are U-shaped. We will discuss this explicitly in the revision.

      (3) We will reconsider the layout to make the background and results easier to follow. In particular, we will consolidate the repeated material while preserving the context needed to interpret the theoretical consequences of the empirical findings.

      (4) We will revise the conclusions and discussion to leave claims about the decay of potentiation more open-ended.

    1. eLife Assessment

      This valuable work discusses the phylogenetic conservation of the hippocampal region and primary sensory cortical regions in mammalian species. The authors propose that species-specific differences in behavior and mnemonic functions may be due to differences in cortico-hippocampal connectivity patterns. However, the manuscript, in its present form, is speculative, and the strength of evidence for this proposition is incomplete.

      [Editors’ note: the final version of this work has been published in the journal Hippocampus (https://doi.org/10.1002/hipo.70119)]

    2. Reviewer #1 (Public Review):

      The paper itself has a reasonable aim, to compare the inputs to the hippocampus from cortical regions across mammals. But for some reason, the conclusions that are reached are very limited. We know for example that the main laboratory rodents investigated, rats and mice, are nocturnal, live in underground tunnels, and have a very wide field of view with no fovea. In contrast, primates have a highly developed cortical system for vision and a fovea, and so have very different capabilities to rodents, as they have an ability to identify people or objects at a distance, and to remember where they have been seen. Despite this major difference in the visual cortical processing in these different mammals, somehow important points are missed in this paper about how the cortical processing is organised in these different mammals, and how this is reflected in the anatomy.

    3. Reviewer #2 (Public Review):

      Summary:

      The manuscript emphasizes a phylogenetic conservation of the hippocampal region and primary sensory cortical regions in mammalian species. The authors then propose that the evident species-specific differences in behavior and memory-related functions may be due to differences in type and amount of cortico-hippocampal connectivity.

      Strengths:

      The authors are well-established researchers with a long history of excellent results and publications. The question (co-influence of cortical and hippocampal connections) is potentially interesting.

      Weaknesses:

      The treatment is very broad and macro scale, ignoring the likelihood that hippocampal-cortical connectivity and behavioral outcomes result from multiple differences at a more micro-scale. The designated "mammalian" sample is also broad. Thus, it can appear incomplete as a sample, and incompletely discussed.

    1. eLife Assessment

      In this study, electroencephalography was recorded during a binocular-rivalry paradigm in which participants viewed a target and a distractor stimulus. The stimuli were independently frequency-tagged, allowing the researchers to track target and distractor processing over time. The findings suggest that parietal alpha oscillations initially segregate competing inputs, followed by frontal theta activity associated with suppression of distractor representations. The significance of these findings is valuable, although the strength of the evidence remains incomplete, because of specific experiment design choices and methodological limitations.

    2. Reviewer #1 (Public review):

      Summary:

      These authors used a binocular rivalry task with flickering stimuli in which subjects had to report the color of the target grating at the end of each trial. Target or distractor cues provided information about the orientation of the respective stimulus prior to each trial. The stated goals of this project include testing the neural mechanisms underlying strategic target and distractor processing. Behavioral enhancement was observed for target cueing, while no cost was noted for distractor cueing. These authors present evidence for reactive suppression, characterized by pronounced frontal theta activity that reduced the sensory gain (SSVEP) of the distractor. Distractor cues also increased alpha activity over parietal areas, which these authors link to attentional gating while pointing out no relationship with sensory gain.

      Strengths:

      This manuscript clearly reflects thoughtful analysis of the available data. Alongside a simple and effective task design, sophisticated methods provide good support for most of the claims made by these authors.

      Weaknesses:

      Lack of temporal precision for SSVEP effects. I would like to see how sensory gain is/isn't dynamically modulated in the moments after the initial ERP to see if there could be differences compared to the broader window used presently (1.3 to 3.1 seconds).

      These authors indicate that persistence of the neural representation of cued distractor orientations into the rivalry period is evidence against a "search-and-destroy" type mechanism where distractors are enhanced to then be suppressed reactively. This claim relies on an indirect link between the maintenance of information about distractor orientation (i.e., successful orientation decoding) and the processing of sensory representations. This claim would be backed up more substantially if the SSVEP (a measure of sensory processing) could reveal temporal dynamics on a finer scale.

    3. Reviewer #2 (Public review):

      Summary:

      The findings are conceptually useful - a sequential alpha-then-theta architecture for proactive gating and reactive distractor suppression would be a compelling contribution to the attention control literature - but the evidence is incomplete at best. The central dissociation rests on an inadequate proxy for perceptual dominance, the key alpha-behavior effect is small (d = 0.199) and confined to a single unprotected data quadrant, and the GLMM uses an inadequate random effects structure that inflates false-positive risk.

      Strengths:

      The SSVEP frequency-tagging + binocular rivalry combination is genuinely inventive for isolating sensory gain signals from the two competing stimuli simultaneously. The finding that distractor cueing enhances sensory processing of the distractor yet fails to impair behavior is a clean result that directly addresses a behavioral paradox in the attentional suppression literature. The non-phase-locked TF analysis and the use of RESS for SSVER extraction are methodologically sound.

      Weaknesses:

      The most consequential flaw in the paper is the operationalization of "perceptual dominance." The authors explicitly acknowledge in a footnote that trial categorization as "target-dominant" or "distractor-dominant" is based on which eye received the stimulus, not on participants' actual perceptual reports. Because participants were never asked to report which stimulus was dominant (only to reproduce the target's color), the assignment is an anatomical proxy, not a perceptual measure. This matters enormously for the paper's central claims. Specifically: (a) The entire two-mechanism dissociation (theta for target-dominant trials, alpha for distractor-dominant trials) is built on a trial-type categorization that may not reflect subjective perceptual experience on a given trial, and (b) Dominant-eye stimuli do typically win initial rivalry dominance, but dominance alternates, and in a 2-second window (the stimulus duration used), perceptual states likely fluctuate in many trials. The lack of button-press perceptual tracking (e.g., continuous dominance reports) means the authors cannot verify that their neural effects actually correspond to the perceptual states they claim. This is a major structural limitation of the design that can't be retroactively corrected, and it significantly weakens the consciousness/awareness framing of the findings.

      Another significant issue is that the parietal alpha effect on behavior is confined to a very specific quadrant of the data: distractor-dominant trials where both target and distractor SSVERs are weak simultaneously. The authors present this as an elegant result - "alpha helps most under high perceptual uncertainty" - but it could equally reflect insufficient statistical power for effects in the other three SSVER-strength cells (target strong/distractor weak; target weak/distractor strong; both strong). The Cohen's d for the alpha effect on target reporting probability is only d = 0.199, which is a very small effect. With N=36 and no correction for the multiple SSVER-strength subgroupings tested, there is a real risk that this specific cell-finding is a false positive, while the null in adjacent cells reflects inadequate power rather than a genuine boundary condition.

      A third major limitation is that with a design that includes 6 fixed effects and all their interactions, the random effects structure should include random slopes for at least the key predictors (cueing condition, dominance). Fitting maximal random effects models or justified reduced structures (Barr et al., 2013) is standard in within-subjects EEG research. Using only random intercepts risks inflating Type I error rates for the interaction terms that form the core of the paper's claims. The authors provide a supplementary table (Table S1) but do not describe whether model convergence was verified or alternative random effects structures were tested.

      Fourth, the paper's title and central claim are that alpha and theta dynamics operate sequentially. However, the temporal ordering (preparatory alpha -> rivalry-phase theta) is primarily shown by examining each oscillation in its respective analysis window, not by a single analysis testing whether the sequence itself predicts behavior better than either mechanism alone. A path analysis or cross-lagged model linking trial-level alpha to subsequent theta, and both to behavior, would directly substantiate the "relay" framing. Without this, the sequential architecture is more of an interpretation than a demonstrated property.

      Finally, the frontal theta cluster identified by permutation testing spans 3 to 16 Hz - a range that extends well into the alpha band. Calling this a "theta" effect while simultaneously discussing alpha as a separate mechanism is difficult to reconcile. At minimum, this frequency boundary issue warrants explicit discussion.

    4. Reviewer #3 (Public review):

      Summary:

      Interest was especially focused on how foreknowledge of the orientation of either the target or the distractor could be used to resolve the competition between these stimuli and properly report the target color. The target or distractor was pre-cued by a solid or dashed orientation cue. They were displayed with slightly different presentation frequencies, which allowed for examining their sensory processing with steady-state visual evoked responses (SSVERs). Furthermore, orientation decoding was performed, which revealed that orientation cues selectively affected processing after stimulus onset related to the dominant but not the non-dominant eye. EEG analyses additionally focused on parietal alpha activity and frontal theta, both during the anticipatory phase and the stimulus-processing phase. Cueing the distractor vs. the target induced increased right parietal alpha power during the anticipatory phase, but this did not result in direct inhibition of distractor features. During the stimulus-processing phase, cueing the distractor resulted in increased theta activity. Finally, a generalized linear model was employed wherein trial-by-trial behavior (precision in target color report) was predicted by target and distractor SSVERs, type of pre-cued stimulus (target/distractor), preparatory parietal alpha power, stimulus processing-related frontal theta power, eye dominance, and all their interactions. Performance in the case of reduced sensory processing of the target (based on SSVER) showed more deviations when sensory processing of the distractor was high, but no such effect was observed when sensory processing of the target was high. The latter effects were modulated by eye dominance and cue. Increased theta reduced distractor sensory processing but not target sensory processing. Increased alpha was only beneficial when sensory evidence for both target and distractor was low. Results were interpreted as favoring sensory gating before stimulus onset, reflected by increased parietal alpha (i.e., pro-active control), while theta activity especially seemed relevant to suppress distractor activity (i.e., reactive control) thereby favoring target-related performance.

      Strengths:

      The authors convincingly show that EEG can provide crucial information about how the human brain deals with the conflict between a target and distractor in a binocular rivalry paradigm with pre-cues signaling either the target or the distractor orientation. An important aspect of the study is the focus on precision of target color report, in combination with the possibility to assess SSVERs to the target and distractor. The strength of this study may actually also be its weakness; the question is whether the presented ideas on proactive and reactive mechanisms can be generalized to paradigms that do not employ binocular rivalry. Separation of target and distractor processing by selectively presenting them to the left/right eye increases the conflict when the target is presented at the non-dominant eye, but what happens in the absence of binocular rivalry concerning the target-distractor conflict?

      Weaknesses:

      An important aspect of the study relates to the cue manipulation. In many studies, cues are often informative but not mandatory. Couldn't one argue that in this study task performance crucially depends on cue processing, as without the cue, it becomes difficult to tell apart the target from the distractor. It could be argued that participants are able to do this based on the slight difference in flickering frequency, but I doubt whether this is possible at all. However, if this were the case, then they might use this as an alternative cue and ignore the orientation cue. What do participants experience while performing this task? As the cue can be considered to be mandatory, the question may be raised what strategy the participants actually employed. If the target was cued, they simply may have prepared for this orienting and could ignore the distractor. However, if the distractor was cued, they could use two strategies: search for the stimulus without the cued orientation, or first detect the distractor, and then orient towards the other stimulus. The ideas and results on parietal alpha and frontal theta in combination with the other findings are certainly very interesting, but recently, it has also been argued that frontal theta may be more related to action control (e.g., see Panek et al., https://doi.org/10.1093/cercor/bhaf276) and also pro-active control (Cooper et al., 2017). So, it might be that increased theta reflects suppression of the response related to the distractor, which feeds back on its sensory processing. This raises the question whether there is possibly also some evidence on functional connectivity between frontal and posterior regions that varies depending on the precise condition. Are the results also shining a new light on the relation between attentional orienting and eye dominance (e.g., see Schintu et al., 2020)?

    1. eLife Assessment

      This important study introduces a method, based on active control, for determining the functional connectivity between neurons or groups of neurons. While the theory is derived for a linear model with a perfectly observed control signal, the authors present compelling evidence that their method will work even when the system is nonlinear and the control signal is not fully observed. Furthermore, they provide a practical recipe for doing so.

    2. Joint Public Review:

      Summary:

      Inferring so-called "functional connectivity" between neurons or groups of neurons is important both for validating models and for inferring brain state, including in human patients. This study aims to enhance this inference process by using closed-loop perturbation-based approaches. To this end, the authors develop a framework based on linear dynamical models that minimizes the estimation error. Based on this framework, the authors provide a practical guide for applying it in realistic experiments. Modalities include non-invasive ones, such as fMRI, iEEG, and invasive ones, such as optogenetic perturbations combined with neuropixel probes or calcium imaging.

      Strengths:

      A main strength of this paper is the application and adaptation of an explicit error expression to system dynamics estimation from evoked neural responses, bringing a useful theoretical tool into computational neuroscience for, as far as we know, the first time. Importantly, while the analytical derivation assumes the neural dynamics is linear and the control signal is known, these assumptions do not appear to be essential: their method outperforms passive observation even when the true dynamics is nonlinear or the control input is not known perfectly. Moreover, the relative simplicity of the method makes its practical applications straightforward, as the authors illustrate in the context of brain state classification and neural control.

      Besides being of practical importance, simply pointing out that passive observation can lead to large mis-estimation of functional connectivity should serve as a wakeup call to anybody engaged in this endeavor.

      Weaknesses:

      None.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Weaknesses:

      (1) The derivation of the main error term misses some important steps, which complicates peer review at this stage. In particular, factorisation of the covariance into noise and the inverse of the observation covariance matrix needs a more thorough justification. The cited sources do not contain the derivation for a noise term with full covariance, which is essential for deriving this error term.

      The derivation of the main error term misses some important steps, which complicates peer review at this stage

      We thank the reviewers for this careful observation. This concern is associated with the error term. Thus, we first clarified the noise assumption explicitly. We assume that ξ(t) is i.i.d. over time with zero mean and an arbitrary (not necessarily diagonal) positive definite covariance matrix .

      In particular, factorisation of the covariance into noise and the inverse of the observation covariance matrix needs a more thorough justification. The cited sources do not contain the derivation for a noise term with full covariance, which is essential for deriving this error term.

      The cited sources do contain the derivation for a noise term with full covariance. We have updated the citation that directly supports Eq. (S.2): Proposition 11.1 of Hamilton (1994, TimeSeries Analysis), which establishes the asymptotic distribution The proof, given in Appendix 11.A of Hamilton (1994), proceeds via a CLT for martingale difference sequences. See Theoretical Details of the Supplementary Materials.

      (2) The practical recommendation at the end of the paper also requires clearer guidance on how the design perturbations are constructed, and how many times and for how long the system is stimulated in each iteration of the experiment.

      Thank you for this helpful suggestion. We agree that the practical implementation of the experimental design should be explained more clearly. We have addressed this concern in two ways. First, we have revised the manuscript to explicitly describe the parameter design procedure. Second, we have revised the manuscript to clearly provide a reference to the detailed experimental condition table in the supplementary material. See Results - Main Manuscript.

      (3) Finally, there is no analysis of model mis-specification. In particular, the true dynamics are unlikely to be linear; the noise is unlikely to be either Gaussian or uncorrelated across time; and the B matrix is unlikely to be known perfectly. We’re not suggesting that the authors consider a more complex model, but it’s important to know how sensitive their method is to model mismatch. If nothing can be done analytically, then simulations would at least provide some kind of guide.

      We thank the reviewer for raising this important point regarding model mis-specification. We agree that it is important to run simulations to assess the impact of these model mismatches, therefore we conducted preliminary simulations to assess the sensitivity when two primary assumptions are violated: linear state dynamics and a perfectly known input matrix B. We added these preliminary results in the Supplementary Material, and revised the main manuscript to include a mention of these results. In summary, the simulations showed:

      The model estimation error increases with the strength of the nonlinearity; however, perturbation can increase the information and lead to accurate estimation of the hidden mode even under nonlinearity.

      The matrix A can be estimated roughly even if the assumed input matrix B differs from the true matrix in some cases.

      The matrices A and B can be estimated jointly without bias, provided the stimulation pattern excites the full state space.

      In such joint estimation, preferentially exciting the hidden modes directly leads to more accurate estimation of A than stimulating other modes. See Background and Results - Main Manuscript, Experimental Conditions and Results - Supplementary Material.

      Recommendations for the authors:

      (1) Please tell us what tACS, tDCS, and TMS are, and how much control the experimenter has over them. That’s important, because they are going to be used as control signals, so we need to know how accurately u(t) can be specified, and what its range is.

      Please tell us what tACS, tDCS, and TMS are, and how much control the experimenter has over them.

      We appreciate the reviewer’s helpful comment. We agree that it is important to describe these stimulation methods with appropriate references. We have added a new section titled “Neural Stimulation as Control Inputs” to the background, which connects our theoretical framework to practical experimental settings.

      We need to know how accurately u(t) can be specified, and what its range is.

      The accuracy and range of control inputs vary substantially depending on the specific stimulation technique and experimental setup, and a thorough discussion would require a dedicated review beyond the scope of this manuscript. Instead, we have added a sentence acknowledging the gap between practical experimental implementations and the theoretical formulation, and cited relevant references for readers interested in further details.

      Background - Main Manuscript

      “Neural Stimulation as Control Inputs

      This section describes how commonly used neural stimulation techniques can be related to input signals in control theory. Their adjustable parameters vary depending on how the stimulation inputs are modulated.

      Three non-invasive electrical stimulation methods illustrate how stimulation paradigms map onto basic control inputs. Transcranial magnetic stimulation (TMS) induces brief and transient perturbations via electromagnetic pulses [19], which are naturally represented as a sequence of impulse-like inputs, where the timing and intensity of each pulse are the primary controllable parameters. Transcranial direct current stimulation (tDCS) primarily modulates neural activity through approximately constant inputs [26], which can be viewed as a step-like signal whose main controllable parameter is the amplitude of the applied current. Transcranial alternating current stimulation (tACS) delivers oscillatory inputs [7], corresponding to sinusoidal signals characterized by amplitude, frequency, and phase. In control theory, impulse, step, and sinusoidal inputs are the basic components used to characterize system responses and dynamics [22, 21].

      The control input framework extends beyond non-invasive techniques to invasive and optogenetic stimulation. Invasive electrical stimulation, including intracranial microstimulation and deep brain stimulation (DBS), enables direct delivery of electrical inputs to neural tissue [17], providing flexible control over amplitude and timing through pulse trains or temporally structured waveforms. Optogenetic stimulation allows genetically targeted activation or inhibition of specific neurons using light [5], providing fine-grained control over multiple input dimensions, including amplitude (light intensity), temporal pattern, and cell-type specificity. In particular, recent developments enable stimulation at the level of individual neurons with high temporal precision [27, 16], allowing flexible construction of spatiotemporal input patterns.

      These stimulation examples demonstrate that the theoretical framework developed in this paper connects to practical experimental settings. While a substantial gap remains between idealized control inputs in theory and experimentally realizable stimulation, the core principles established in the following sections provide a foundation that naturally extends to these practical stimulation paradigms.”

      (2) Is the solid curve in the last panel of Figure 4b the prediction? If not, are the data points consistent with the prediction in Equation 8? This should be clear.

      Thank you for this question. The solid curve shows the empirical eigenvalues of the state covariance matrix, not the eigenvalues of matrix A. As shown in Equation 9, the estimation error is proportional to the inverse of the sum of the eigenvalues of the state covariance matrix. We have clarified these points by revising the main text and the figure captions. See Results - Main Manuscript.

      (3) Page 10, "Nodes 7 and 8 have only outgoing edges". It looks like node 7 has an incoming edge from node 8 (Fig. 6b). Or are we misinterpreting something?

      Thank you for pointing out this inconsistency. You are correct. Node 7 did have an incoming edge from Node 8 and contradicted the statement in the text. Moreover, we now think that two hub nodes (Node 7 and 8) are not necessary for demonstrating our primary theory. To resolve these issues, we have revised the simulation so that the network now contains a single hub node with only outgoing edges. Please refer to Fig. 6.

      (4) Figure 6f, g: why isn't the estimation error proportional to tr[Sig_x^{-1}]?

      We thank the reviewer for this observation. In the original manuscript, the estimation error in Figure 6f, g was plotted on a logarithmic scale, which obscured the proportional relationship with . The underlying values are indeed proportional, consistent with our theoretical prediction.

      In the revised manuscript, we have substantially reorganized Figure 6 to make the theoretical reasoning more transparent by using 1/µ instead of following Eq. 9. The sum across the column in panel (g) is proportional to the estimation error shown in panel (h).

      (5) Page 11, "The simulation was conducted with a single node receiving an impulse-shaped perturbation input." What’s an "impulse-shaped perturbation input"? A delta function? Please make this clear.

      Thank you for this clarifying question. We clarified the explanation as follows.

      Results - Main Manuscript

      “Each node was individually perturbed by an impulse input. Here, an impulse input is defined as a Kronecker delta at t = 0 with fixed amplitude α = 10, with no external input applied at any subsequent time step.”

      (6) Page 11, "Fig. 6d shows that the system possesses modes with small absolute eigenvalues." According to Figure 6d, all the absolute eigenvalues are between 9.9 and 9.98. So this statement appears not to be correct. Are we missing something?

      Thank you for pointing this out. You are correct. The previous statement was inconsistent with the figure. In the revised simulation, we have redesigned the network so that it clearly contains modes with distinct damping characteristics: heavily damped modes with absolute eigenvalues below 0.8 and lightly damped modes with absolute eigenvalues close to 1.0 (see revised Fig. 6c). This makes the relationship between mode damping and perturbation effectiveness much more transparent.

      Results - Main Manuscript

      “Figs. 6c shows the damping rates |λA| of the eigenvalues of A: the mode formed by Nodes 1 and 2 (15Hz) is heavily damped, while those formed by Nodes 3–6 are moderately damped. Node 7 serves as a hub with only outgoing edges.”

      (7) Page 11, "Crucially, the eigenvectors corresponding to these rapidly decaying modes (e.g., evec 7, evec 8) have their largest components concentrated at Nodes 7 and 8." If "nodes" are the same as "eigenvalue index", then the components on node 8 are zero (Figure 6e, bottom). In any case, it should be clear what you mean.

      Thank you for this comment. We agree that the previous description regarding eigenvectors was unclear and inconsistent with the figure. We now think that explaining with eigenvectors is not necessary for demonstrating our primary theory. We have therefore replaced the plots of eigenvectors with the reciprocals of the eigenvalues of Σ<sub>X</sub>, which are directly and rigorously explained by Equation (9). These reciprocals clearly show that the modes associated with Nodes 1 and 2 are heavily damped and applying perturbations to these nodes contribute to the estimation error and the perturbations to hub node (Node 7) broadly increases all eigenvalues of Σ<sub>X</sub>, thereby reducing the estimation error across all modes. This change ensures that all simulation results are grounded in the theory presented in the paper. Please refer to the revised Fig. 6d–g for details.

      (8) It’s not clear to us what’s plotted in Figure 6e. The real part of the eigenvectors? Which would explain why some of the eigenvectors are the same (e.g., 1 and 2). But that does not seem like a good idea, since the eigenvectors can be rotated by an arbitrary complex phase. Also, eigenvectors 7 and 8 seem totally opposite, and there’s no weight at all on index 6 and index 8. Could there be a mistake in the figure? In addition, Figure 6e is explained and interpreted in two different paragraphs, discussing the same observation. It would be easier to understand if they were moved to the same paragraph. Also, in the second paragraph where plot 6e is referenced (page 11, line 19), what does ’concentrated’ mean?

      Thank you for this comment. As described in our response to Comment (7), we have removed the eigenvector-based analysis from the revised simulation to prevent confusing readers. The revised results focus on quantities that are directly explained by Equation (9). Please refer to the revised Fig. 6 for the updated results.

      (9) Page 11, "This simulation demonstrates that, given a tentative connectivity matrix, an effective perturbation input (such as TMS or tDCS) can be designed by targeting the node with the highest weighted out-degree." "Demonstrates" seems strong. There is only one simulation, and that wasn’t totally convincing: a perturbation applied to node 7 did almost the same as a perturbation applied to nodes 1-6, even though it had a much higher outdegree than those nodes. It would be very helpful if you provided theoretical reasoning for why perturbing nodes with high-weight connections (Figure 6) minimises the prediction error. In particular, which properties of a hub node make it effective as a stimulation target? Why is it just its total output weights and not also its number of edges/centrality? How is the high output weight of a node related to its alignment with other eigenvectors, and is this always the case or just in this example?

      Demonstrates seems strong.

      Thank you for this important comment. We agree that "demonstrates" was too strong and have replaced it with "illustrates."

      It would be very helpful if you provided theoretical reasoning for why perturbing nodes with high-weight connections (Figure 6) minimises the prediction error.

      We agree with this comment. We have revised the simulation and accompanying text to connect the results directly to the main theory based on . As responded to Comments (7) and (8), we have removed the eigenvector-based analysis and instead plotted the reciprocals of the eigenvalues of Σ<sub>X</sub>, which are directly related to the estimation error via Equation (9). This change allows us to provide a clear theoretical explanation for why perturbing certain nodes minimizes the prediction error.

      Is this always the case or just in this example?

      We have also explicitly stated the limitations: whether a hub node or a specific subnetwork node is more effective depends on factors such as the outgoing edge weights from the hub and the individual damping rates of each mode. The revised text emphasizes that this simulation presents one example of perturbation location design, and the optimal strategy must be evaluated case by case using the theoretical framework of Equation (9).

      Results - Main Manuscript

      “As this simulation represents only one example of location design, its limitations and the corresponding countermeasures should be stated. In actual experiments, the most effective perturbation location depends on factors such as hub-node connectivity and modal damping rates, and B itself may not always be known a priori. In such cases, approaches such as iterative optimization of (described in a later section) and joint estimation of A and B (Supplementary Material B.3) provide systematic alternatives. Nevertheless, the results presented here provide an intuitive guideline: stimulation directed at hub nodes or at nodes driving heavily damped modes effectively excites the full set of dynamical modes and minimizes estimation error.”

      (10) What are the physical units for the time scales and stimulation amplitudes? For instance, on page 13, there is an argument that "In practical experiments, such long windows are unrealistic because neural states change rapidly over time." However, it is unclear whether T=100 a.u. or T=1000 a.u. etc. is realistic. By relating it to the eigenspectrum of A, which is supposed to be physiologically realistic, one can estimate the length of the stimulation window and support the above statement. Similarly, the impulse amplitude on page 13 is alpha=10<sup>20</sup> (a.u.). Also, Table 1 in the Supplementary has an extremely wide range of stimulation amplitudes. Is 10<sup>20</sup> a.u. a feasible amplitude in practice? And finally, please tell us which nodes the input was applied to.

      Thank you for your incisive comments. We have addressed each comment as follows.

      What are the physical units for the time scales and stimulation amplitudes?

      They don’t have physical units. This study is a theoretical investigation that focuses on the relative differences between passive observation and perturbation-based approaches, rather than providing precise predictions for specific experimental settings. The time scales and stimulation amplitudes are therefore expressed in arbitrary units.

      For instance, on page 13, there is an argument that "In practical experiments, such long windows are unrealistic because neural states change rapidly over time." However, it is unclear whether T=100 a.u. or T=1000 a.u. etc. is realistic.

      We agree that the original expression “unrealistic” was not appropriate given the arbitrary units. We have revised the text to clarify that the time scales are in arbitrary units and that the main point is about the relative difference in required data length between passive observation and perturbation-based approaches, rather than making an absolute claim about feasibility.

      Results - Main Manuscript

      “To obtain estimates under the passive condition that are comparable to those derived under perturbation, it is necessary to experimentally observe extensive time-series data. Figure 7f illustrates the LDA projection and classification accuracy for different time-series lengths. The leftmost LDA plot (T = 20) corresponds to the passive condition shown in Fig. 7c, indicating that the estimation performance in the passive condition becomes comparable to that in the perturbation condition only when the time window reaches approximately T = 200, a 10-fold increase compared to T = 20. While the absolute duration depends on the interpretation of the time unit, such time windows may not be prohibitive in some experimental settings. Nevertheless, our results consistently show that passive observation requires substantially longer recordings to achieve comparable performance, highlighting the efficiency of the perturbation-based approach when the available data length is limited.”

      Is 10<sup>20</sup> a.u. a feasible amplitude in practice?

      Although we have already stated that the stimulation amplitudes are in arbitrary units, we agree that the original value of 10<sup>20</sup> was excessively large and could be misleading. We have revised the impulse amplitude from 10<sup>20</sup> to 10<sup>2</sup>, as 10<sup>20</sup> is physically unrealistic—it would imply a stimulus intensity many orders of magnitude beyond any conceivable experimental setting. The revised value of 10<sup>2</sup> is more plausible: for reference, TMS stimulation voltages exceed typical EEG amplitudes by roughly 4–6 orders of magnitude. The classification accuracy decreased slightly; however, our main conclusion regarding the efficiency of the perturbation-based approach under limited data remains unchanged. See Table B.1.

      Finally, please tell us which nodes the input was applied to.

      The stimulus location was optimized to minimize the . We have clarified this procedure in the main text and added a visual indication of the selected stimulation site (red circles) in Fig.7.

      Results - Main Manuscript

      “The neural signals were simulated under five different task conditions and two stimulation conditions: passive observation and external perturbation. Perturbation was applied as impulse-type inputs, such as TMS. The stimulus location was determined for each task condition by applying an impulse to each node and selecting the one that minimized . The resulting time-series data are shown in Fig. 7b. Using this data, we estimated the underlying dynamical system via a control-based identification approach presented in Eq. 5, which corresponds to an estimation of functional connectivity.

      (11) Page 13: "The controlled transition test was run with T = 1." Previously, T referred to the number of time steps. Is that the case here? If so, that seems hard to justify. If not, please tell us what T is (and, ideally, use a different symbol).

      Is that the case here?

      No. In this context, T does not denote the number of time steps.

      If not, please tell us what T is (and, ideally, use a different symbol).

      Thank you for your helpful comment. We have standardized the notation throughout the manuscript. In this paper, T consistently denotes the number of time steps (i.e., data length). In the sections “Neural State Classification” and “Neural State Transitions,” we had mistakenly used T to refer to time length. To resolve this inconsistency, we have added a separate column labeled “Data Length (T)” to Table B.1 for clarification and replaced the previous usage of T with “data length” where appropriate.

      Results - Main Manuscript

      “To obtain estimates under the passive condition that are comparable to those derived under perturbation, it is necessary to experimentally observe extensive time-series data. Figure 7f illustrates the LDA projection and classification accuracy for different time-series lengths. The leftmost LDA plot (T = 20) corresponds to the passive condition shown in Fig. 7c, indicating that the estimation performance in the passive condition becomes comparable to that in the perturbation condition only when the time window reaches approximately T = 200, a 10-fold increase compared to T = 20. While the absolute duration depends on the interpretation of the time unit, such time windows may not be prohibitive in some experimental settings. Nevertheless, our results consistently show that passive observation requires substantially longer recordings to achieve comparable performance, highlighting the efficiency of the perturbation-based approach when the available data length is limited.”

      Results - Main Manuscript

      “The controlled transition test was run with T = 50. The control objective was to set nodes3 and 4 to 25 while keeping all other nodes at 0 without any movement. See Table B.1.”

      (12) Page 14: "where the matrix A (Fig. 10a) is designed to have 16 modes." What do you mean by has "16 nodes"?

      Thank you for pointing this out. By “16 modes,” we refer to 16 dynamical eigenmodes (i.e., 16 eigenvalue pairs). In the real-valued state-space representation used in the simulations, each complex conjugate pair corresponds to a 2-dimensional real block, resulting in a 32-dimensional system (32 nodes). We revised the wording to clearly distinguish between the number of dynamical modes and the dimensionality (number of nodes) of the state vector, to avoid confusion.

      Results - Main Manuscript

      “Iterative refinement of both the perturbation design and the estimation process progressively improves the accuracy of A. The time-series data is collected from 32 points, where the matrix A (Fig. 10a) is designed to have 16 oscillatory modes (i.e., 16 complex-conjugate eigenvalue pairs, yielding 32 eigenvalues in total).”

      (13) In Figure 10b, the y-axis should start at zero; otherwise, it’s a bit misleading how much the active perturbation helps. This will make it clear that the estimation error drops by about 33%. It would be worth commenting on whether this is typical; after all, potential users of this method would want to know how much improvement they’re likely to see.

      It would be worth commenting on whether this is typical; after all, potential users of this method would want to know how much improvement they’re likely to see.

      Thank you for this insightful comment. We agree with the reviewer that quantifying the expected improvement would be valuable for experimental practice. However, this simulation is a theoretical demonstration. Its primary purpose was to show that an iterative active perturbation approach can progressively converge to an optimal perturbation design even without prior knowledge of the true system, rather than to quantify a universally expected improvement rate.

      The y-axis should start at zero; otherwise, it’s a bit misleading how much the active perturbation helps.

      We believe that the y-axis should start at the accuracy with optimal perturbation because the primary purpose of this simulation was to demonstrate that an iterative active perturbation approach can progressively converge. If we had started the y-axis at zero, the message would be visually obscured.

      Potential users of this method would want to know how much improvement they’re likely to see.

      We acknowledge that this is one example of the application of our method, and the magnitude of improvement depends on various factors such as network structure, noise level, and stimulation design. We should not mislead the readers. We have therefore maintained the y-axis starting point and added a clarifying statement in the revised manuscript to indicate that the simulation is case-specific rather than universal.

      Results - Main Manuscript

      “These results should be interpreted as a case-specific illustration rather than a universal gain, as the magnitude of improvement depends on factors such as network structure, recording duration, noise level, and stimulation design. This simulation demonstrates that our perturbation design framework enables the step-by-step refinement of system identification even without prior knowledge of the system.”

      (14) Please provide a derivation for the factorised covariance in Equation S.2. This is the equation that underpins the main result of the paper - the error in dynamical system estimation. Currently, it appears to be taken from [Hamilton, J. D. Time Series Analysis], yet we were unable to find this result in the book. Most of the derivations in Chapters 8.1 and 8.2 assume diagonal noise covariance, and even isotropic noise (cov = sigma*2 * I), which simplifies the particular case of the derivations. Could the authors provide a reference or the derivation for the case with full covariance?

      As described in our response to the Weakness above, the relevant result is Proposition 11.1 of Hamilton (1994), not Chapters 8.1–8.2. Proposition 11.1 states the asymptotic distribution of the vectorised OLS estimator in a VAR model with i.i.d. innovations whose covariance matrix Ω is an arbitrary positive definite matrix. The factorisation in our Eq. (S.2) corresponds directly to the revised manuscript we explicitly cite “Hamilton (1994), Proposition 11.1” at Eq. (S.1) and in Hamilton’s Proposition 11.1, with and . In state the i.i.d. assumption on ξ(t) in the preceding paragraph, so that the connection to this result is unambiguous. See Theoretical Details Supplementary.

      (15) Strongly perturbing/Exciting fast-decaying modes to give them more ’runway’ and increase observed variability makes intuitive sense for a normal system. But what would happen in the case of non-normal dynamics, where stimulating one dimension only transiently amplifies it, but then excites other dimensions? Non-normality breaks the alignment between PC components and dynamic modes (see Kumar, Ankit, Loren M. Frank, and Kristofer E. Bouchard. "Identifying feedforward and feedback controllable subspaces of neural population dynamics." arXiv preprint arXiv:2408.05875 (2024)), so eigenvectors of A and Σ<sub>X</sub> won’t align for non-normal dynamics. This alignment appears to be a hidden assumption of this paper. Should the normal dynamics then be stated as an assumption/limitation of the framework? Is the example in Figure 6a highly non-normal? Does considering out-degree provide an empirical approach, an alternative to ’enlargement’, to dealing with non-normality?

      We thank the reviewer for this insightful comment regarding non-normal dynamics.

      Should normal dynamics be stated as an assumption/limitation of the framework?

      No. Our framework does not assume normal dynamics. For example, Figure 5 and the revised Figure 6 illustrate non-normal cases. In Figure 5, the row and column norms of A differ substantially (row norms ≈ [1.96, 4.05, 1.33, 4.70]; column norms ≈ [6.14, 1.94, 1.39, 0.87]). In Figure 6, Node 7 is a hub with only outgoing edges, so row 7 of A is zero while column 7 has nonzero entries. In addition, the Nodes 5–6 and Nodes 3–4 are strictly one-way, leaving the corresponding off-diagonal block upper-triangular. Both features break the symmetry required for normality.

      We additionally computed a commutator-based non-normality index which equals 0 for any normal matrix and approaches for the canonical maximally non-normal example (the 2×2 nilpotent Jordan block). The non-normality index is 1.34 for Figure 5 and 0.41 for Figure 6. These values confirm that both are clearly non-normal.

      Is the example in Figure 6a highly non-normal?

      Yes. As stated above, the network structure of Figure 6a clearly indicates non-normality.

      Does considering out-degree provide an empirical approach, an alternative to ‘enlargement’, to dealing with non-normality?

      No. In the previous manuscript, the results of out-degree and eigenvector structure were provided as supplementary intuition. However, the central contribution of our framework lies in maximizing the minimum eigenvalue µ of the observed state covariance (Equation 9), and for systems where non-normality is strong and subnetwork structure is less modular, the iterative optimization of provides a principled, assumption-free method. We clearly mentioned this point in the revised manuscript as follows. See Results in the main manuscript.

      (16) In the iterative experiment at the very end of the paper (Figure 10), what was the strategy for designing ‘u<sub>design</sub>´? How many stimulations were applied in an iteration? Are you stimulating along eigenvectors? Do you sample from components randomly, or perturb each of them individually, with the amplitude proportional to reciprocals?

      Thank you for this important question. We have clarified the design rule and stimulation protocol in the revised manuscript, and address each sub-question below.

      How many stimulations were applied in an iteration?

      One stimulation session was applied in an iteration.

      Are you stimulating along eigenvectors?

      The optimal stimulation is designed as the target node of the perturbation is determined through numerical optimization that minimizes

      Do you sample from components randomly, or perturb each of them individually, with the amplitude proportional to reciprocals?

      The optimal stimulation is designed as a composite-frequency sinusoidal input encompassing all modes of the estimated Â, and the target node of the perturbation is determined through numerical optimization that minimizes . See Results - Main Manuscript.

      (17) While it is clear that the proposed active method performs better than passive observation, some results lack a comparison with stimulating random directions/nodes with a comparable control energy (Figure 7e & Figure 10b).

      Thank you for your helpful comment. Although the optimally designed stimulation outperforms random stimulation, the previous stimulation settings were not configured to explicitly demonstrate this difference. Therefore, we modified the network structure, recording length, and stimulation intensity so that both the main messages and the difference from random stimulation can be shown simultaneously. Accordingly, we have added a comparison with random stimulation in both Figure 7e and Figure 10b.

      In Figure 7e, we included a random stimulation condition where the target node is selected randomly, and the results show that the optimized stimulation outperforms random stimulation.

      These additions strengthen the evidence for the effectiveness of our proposed method compared to non-optimized approaches.

      (18) Is it reasonable to assume full observability of the system? It would be interesting to consider biases arising from the partial observability of the system, in the spirit of Figure 9, which looked at partial controllability.

      We thank the reviewer for this suggestion. We agree that partial observability is an important consideration, however it is out of scope for the current work. Thus, we have added a future direction in the Discussion addressing partial observability. We note that Takens’ embedding theorem and Hankel DMD enable recovery of a system’s eigenvalues from partial observations, and since our framework relies on the eigenvalue structure of A, the proposed perturbation design remains applicable under partial observability.

      Discussion - Main Manuscript

      “Two directions warrant further investigation: extending the framework to partial observability, and validating it through stimulation experiments. In experimental neuroscience, recordings are often limited to a subset of neural populations, resulting in partial observability. A growing body of work has leveraged delay-embedding techniques, represented by Takens’ embedding theorem [25], to reconstruct hidden dynamics from partial observations [2, 3, 23, 11]. Applying such techniques enables the estimation of the full connectivity matrix, thereby extending our framework to settings with partial observability. The second direction concerns experimental validation. Validating a theoretical framework through experimental design is an essential in bridging the gap between theory and practice.”

      Recommendations for improving the writing and presentation.

      (1) The word ’state’ is overloaded in Figure 7. When talking about neural state classification, the ’state’ refers to a regime guided by a distinct dynamics A (should it be A<sub>i</sub>? Figure 7A bottom). However, each dynamical system also has a ’state’. A different word should be used in Figure 7A and the corresponding text.

      We agree that the terminology is potentially confusing. To avoid the confusion, we replaced the term “state” with “task condition” and “neural signal” in the main text, and revised Fig. 7.

      Results - Main Manuscript

      “We designed a neural network with clearly distinct task conditions and considered a simulation setting in which these conditions are classified using signals of a fixed duration. These distinct task conditions consist of five types, each defined by a unique linear dynamical system characterized by differing eigenvalue spectra and connectivity topologies of matrix A (Fig. 7a). These task conditions are intended to mimic different cognitive or behavioral contexts. For example, in a typical motor task experiment, such conditions could correspond to motor execution or imagery involving the left or right hand, or resting state [1, 24]. The neural signal was simulated under five different task conditions and two stimulation conditions: passive observation and external perturbation. Perturbation was applied as impulse-type inputs, such as TMS. The stimulus location was determined for each task condition by applying an impulse to each node and selecting the one that minimized . The resulting time-series data are shown in Fig. 7b. Using this data, we estimated the underlying dynamical system via a control-based identification approach presented in Eq. 5, which corresponds to an estimation of functional connectivity.”

      (2) It would be helpful to provide dimensions of matrices around Equation S.2, since vectorization makes dimensions hard to track.

      We thank the reviewer for this helpful suggestion. We agree that explicitly stating the matrix dimensions improves readability, particularly around the Kronecker product where vectorization can obscure the size of the resulting covariance matrix. We have revised the text as follows (the equation S.2 is 3 now). See Theoretical Details of the Supplementary Materials.

      (3) The background sections of the paper would benefit from referring to similar active perturbation methods: Wagenmaker, Andrew, et al. "Active learning of neural population dynamics using two-photon holographic optogenetics." Advances in Neural Information Processing Systems 37 (2024): 31659-31687. Minai, Yuki, et al. "MiSO: Optimizing brain stimulation to create neural activity states." Advances in Neural Information Processing Systems 37 (2024): 24126-24149.

      We thank the reviewer for these helpful suggestions. We have incorporated Wagenmaker et al. (2024) and Minai et al. (2024) into the Introduction. Their works focus on developing algorithmic approaches to active stimulation design for specific experimental platforms, while our work aims to establish a general theoretical framework for why and which perturbation inputs are effective has yet to be established. We cited these works and have clarified this distinction in the revised manuscript as follows.

      Introduction - Main Manuscript

      “In this paper, we propose a framework for designing the optimal perturbation input through control theory in neuroscience. We interpret neural dynamics as a control system [8, 6, 14, 12, 20, 24], and treat external perturbations as control inputs to design properties of neural stimulation (Fig. 1d). If the optimal perturbation input can be systematically designed, it becomes possible to steer the neural system toward states that are maximally informative (Fig. 1e), thereby enhancing the accuracy of the inferred connectivity (Fig. 1f). While recent studies have begun to develop algorithmic approaches to active stimulation design for specific experimental platforms [18, 28], a general theoretical framework for why and which perturbation inputs are effective has yet to be established. We first describe how to formulate neural dynamics as a control system and how to estimate the model parameters from observed data. Building upon this formulation, we derive a theoretical basis that enables us to design the optimal perturbation inputs for the neural system identification. We demonstrate the validity and utility of this theoretical basis by exploring its implications for optimizing parameters of common neurostimulation techniques and by applying it to practical examples, including neural state classification [4, 9, 1, 24] and control of neural states [8, 13, 14, 12]. In these demonstrations, we define concrete problems and apply the theory to validate its practical utility.”

      Minor corrections to the text and figures.

      (1) In Figure 3, it would be helpful to point out that the plots are in the subspace spanned by the first three principal components.

      Thank you for pointing out. We revised the caption of the Fig.3 as follows. See Results - Main Manuscript.

      (2) Figure 4 caption: "eigenvectors" –> "eigenvectors of Σ<sub>X</sub>", just to make it crystal clear (since A also has eigenvectors).

      Thank you for your suggestion. We have revised the caption of Figure 4. See Results - Main Manuscript.

      (3) Page 5, Equation (8): Capital xi should be introduced in the main text of the paper as the covariance matrix for the noise term xi. Now it can only be understood after reading the Supplementary. Or maybe call the covariance matrix Σ<sub>ξ</sub>rather than Σ<sub>ξ</sub>? That will probably make it clearer.

      Thank you for this suggestion. We have made both changes. First, we introduced the noise covariance matrix explicitly in the main text immediately after Eq. (1), defining. Second, we replaced the notation Σ<sub>ξ</sub> with Σ<sub>ξ</sub>(lowercase subscript matching the noise variable throughout the main text and supplementary.

      (4) Page 7, Equation (14): The description of the equation states "The covariance matrices of x(t) for impulse inputs can be written as:... ". We assume this is supposed to be the "state vector," not "covariance matrices".

      Thank you for catching this. We have corrected the wording to “state vector” instead of “covariance matrices". See Results - Main Manuscript

      (5) Page 10, after introducing Figures 6a-b, potentially a sentence is missing (’...’ in the first line of the last paragraph).

      Thank you for pointing this out. We have removed the placeholder along with updating the stimulation settings for Figure 6.

      (6) Page 10, Figure 6 (e): A clearer labelling would be helpful, e.g., a title for the legend (e.g. node index) and a more informative title (e.g. Eigenvector alignment with nodes), a caption (what is the take-home message?), and y-axis labels (what is the ’value’?).

      Thank you for this helpful suggestion. We have revised the Figure 6 taking your suggestion into account. The new figure includes a clearer title, axis labels, and an informative caption that highlights the key take-home message. See Results - Main Manuscript

      (7) Page 11, paragraph 1: The sentence "...(as discussed in Section)" is missing a section reference.

      Thank you for pointing this out. We have replaced the section reference placeholder with the equation reference to Equation (9), which is the relevant theoretical result.

      (8) Page 13, caption of Figure 8c: compputed -> computed.

      We have corrected this typo. We also reviewed the manuscript for any similar typographical errors and corrected them.

      (9) Page 14 Figure 9: The y-axis label for "Controlled State Process" plots is missing.

      Thank you for catching this. The y-axis was hidden by other elements in the figure. We have revised the figure layout to ensure that the y-axis label is visible.

      (10) Page 16: The sentence "As described in Section, a preliminary ... " is missing a section reference

      Thank you for pointing this out. We have corrected the missing reference. The sentence now reads “as shown in Fig. 10” rather than the incomplete “as described in Section.”

      (11) Page 16, Fig. 10c: It looks like the errors have inconsistent color ranges. A shared colorbar would help.

      Thank you for your suggestion. We have revised Figure 10c to use a shared colorbar across all subplots.

      (12) Page 17, Paragraph preceding eq. 16: "the eigenvectors of X" --> "the eigenvectors of Sigma_X".

      Thank you for catching this. We have revised the text. See Methods - Main Manuscript.

      (13) S.1: This is a GLS, not an OLS estimator, if this assumes full noise covariance.

      As stated in our response to the Weakness above, we assume the noise term ξ(t) is i.i.d. with covariance matrix Σ<sub>ξ</sub>, and the estimator we analyze is the OLS estimator. See Theoretical Details - Supplementary Material.

      (14) S.2: Unclear that⊗is the Kronecker product (not outer), as it is not defined.

      We thank the reviewer for pointing out this ambiguity. In the revised Supplementary Material, we have explicitly explained the Kronecker product with a reference to Hamilton [10, Appendix A.4, p. 732]. See Theoretical Details - Supplementary Material.

      (15) S.51: There is an accidental comma between alpha and A after "xdiff(t) ="

      Thank you for pointing this out. We have removed the accidental comma. See Theoretical Details - Supplementary Material.

      (16) Section B.2: It would be useful to have the simulation details for that section (as is given for the other section in B.1).

      We thank the reviewer for this helpful suggestion. We have added a dedicated “Simulation details” paragraph to Appendix B.5 (the section containing Fig. B.4) so that the setup is now described with the same level of specificity as the other appendix sections. See Experimental Conditions and Results - Supplementary Material.

      References

      (1) Irma N Angulo-Sherman, Marisol Rodríguez-Ugarte, Nadia Sciacca, Eduardo Iáñez, and José M Azorín. Effect of tDCS stimulation of motor cortex and cerebellum on EEG classification of motor imagery and sensorimotor band power. J. Neuroeng. Rehabil., 14(1):31, April 2017.

      (2) Hassan Arbabi and I Mezić. Computation of transient koopman spectrum using hankeldynamic mode decompoisition. APS, page G1.009, November 2017.

      (3) Steven L Brunton, Bingni W Brunton, Joshua L Proctor, Eurika Kaiser, and J Nathan Kutz. Chaos as an intermittently forced linear system. Nat. Commun., 8(1):19, May 2017.

      (4) Adenauer G Casali, Olivia Gosseries, Mario Rosanova, Mélanie Boly, Simone Sarasso, Karina R Casali, Silvia Casarotto, Marie-Aurélie Bruno, Steven Laureys, Giulio Tononi, and Marcello Massimini. A theoretically based index of consciousness independent of sensory processing and behavior. Sci. Transl. Med., 5(198), August 2013.

      (5) Karl Deisseroth. Optogenetics. Nat. Methods, 8(1):26–29, January 2011.

      (6) Shikuang Deng, Jingwei Li, B T Thomas Yeo, and Shi Gu. Control theory illustrates the energy efficiency in the dynamic reconfiguration of functional connectivity. Commun. Biol., 5(1):295, April 2022.

      (7) Shrey Grover, Renata Fayzullina, Breanna M Bullard, Victoria Levina, and Robert M G Reinhart. A meta-analysis suggests that tACS improves cognition in healthy, aging, and psychiatric populations. Sci. Transl. Med., 15(697):eabo2044, May 2023.

      (8) Shi Gu, Fabio Pasqualetti, Matthew Cieslak, Qawi K Telesford, Alfred B Yu, Ari E Kahn, John D Medaglia, Jean M Vettel, Michael B Miller, Scott T Grafton, and Danielle S Bassett. Controllability of structural brain networks. Nat. Commun., 6:8414, October 2015.

      (9) Mark Hallett, Riccardo Di Iorio, Paolo Maria Rossini, Jung E Park, Robert Chen, Pablo Celnik, Antonio P Strafella, Hideyuki Matsumoto, and Yoshikazu Ugawa. Contribution of transcranial magnetic stimulation to assessment of brain connectivity and networks. Clin. Neurophysiol., 128(11):2125–2139, November 2017.

      (10) James Douglas Hamilton. Time Series Analysis. Princeton University Press, Princeton, 1994.

      (11) Ann Huang, Mitchell Ostrow, Satpreet H Singh, Leo Kozachkov, Ila Fiete, and Kanaka Rajan. InputDSA: Demixing then comparing recurrent and externally driven dynamics. arXiv [q-bio.NC], November 2025.

      (12) Shunsuke Kamiya, Genji Kawakita, Shuntaro Sasai, Jun Kitazono, and Masafumi Oizumi. Optimal control costs of brain state transitions in linear stochastic systems. J. Neurosci., 43(2):270–281, January 2023.

      (13) Teresa M Karrer, Jason Z Kim, Jennifer Stiso, Ari E Kahn, Fabio Pasqualetti, Ute Habel, and Danielle S Bassett. A practical guide to methodological considerations in the controllability of structural brain networks. J. Neural Eng., 17(2):026031, April 2020.

      (14) Genji Kawakita, Shunsuke Kamiya, Shuntaro Sasai, Jun Kitazono, and Masafumi Oizumi. Quantifying brain state transition cost via schrödinger bridge. Netw. Neurosci., 6(1):118– 134, February 2022.

      (15) Hassan K Khalil. Nonlinear systems. Prentice-Hall, Upper Saddle River, NJ, 2002.

      (16) Paul K LaFosse, Zhishang Zhou, Jonathan F O’Rawe, Nina G Friedman, Victoria M Scott, Yanting Deng, and Mark H Histed. Single-cell optogenetics reveals attenuationby-suppression in visual cortical neurons. bioRxivorg, page 2023.09.13.557650, May 2024.

      (17) Andres M Lozano, Nir Lipsman, Hagai Bergman, Peter Brown, Stephan Chabardes, Jin Woo Chang, Keith Matthews, Cameron C McIntyre, Thomas E Schlaepfer, Michael Schulder, Yasin Temel, Jens Volkmann, and Joachim K Krauss. Deep brain stimulation: current challenges and future directions. Nat. Rev. Neurol., 15(3):148–160, March 2019.

      (18) Yuki Minai, Matthew Smith, Joana Soldado-Magraner, and Byron Yu. MiSO: Optimizing brain stimulation to create neural activity states. In A Globerson, L Mackey, D Belgrave, A Fan, U Paquet, J Tomczak, and C Zhang, editors, Advances in Neural Information Processing Systems 37, volume 37, pages 24126–24149, San Diego, California, USA, 2024. Neural Information Processing Systems Foundation, Inc. (NeurIPS).

      (19) Davide Momi, Zheng Wang, and John D Griffiths. TMS-evoked responses are driven by recurrent large-scale network dynamics. Elife, 12(e83232), April 2023.

      (20) Ali Moradi Amani, Amirhessam Tahmassebi, Andreas Stadlbauer, Uwe Meyer-Baese, Vincent Noblet, Frederic Blanc, Hagen Malberg, and Anke Meyer-Baese. Controllability of functional and structural brain networks. Complexity, 2024(1), January 2024.

      (21) Norman S Nise. Control Systems Engineering. John Wiley & Sons, 8 edition, 2020.

      (22) Katsuhiko Ogata. Modern Control Engineering. Prentice Hall, 2010.

      (23) Mitchell Ostrow, Adam Eisen, and Ila Fiete. Delay embedding theory of neural sequence models. arXiv [cs.LG], June 2024.

      (24) Yumi Shikauchi, Mitsuaki Takemi, Leo Tomasevic, Jun Kitazono, Hartwig R Siebner, and Masafumi Oizumi. Quantifying state-dependent control properties of brain dynamics from perturbation responses. J. Neurosci., page e0364252025, December 2025.

      (25) Floris Takens. Detecting strange attractors in turbulence. In David Rand and Lai-SangYoung, editors, Dynamical Systems and Turbulence, Warwick 1980, volume 898 of Lecture Notes in Mathematics, pages 366–381. Springer, Berlin, Heidelberg, 1981.

      (26) Liam C Tapsell, Matheus D Pinto, Ann-Maree Vallence, Casey Whife, Maria Luciana Perez Armendariz, Shaswat Senger, Jack Andringa-Bate, Dana Hince, and Myles C Murphy. What are the optimal transcranial direct current stimulation parameters and design elements to modulate corticospinal excitability? a systematic review and longitudinal meta-analysis. Neurol. Res. Pract., 7(1):86, November 2025.

      (27) Lei Tong, Shanshan Han, Yao Xue, Minggang Chen, Fuyi Chen, Wei Ke, Yousheng Shu, Ning Ding, Joerg Bewersdorf, Z Jimmy Zhou, Peng Yuan, and Jaime Grutzendler. Single cell in vivo optogenetic stimulation by two-photon excitation fluorescence transfer. iScience, 26(10):107857, October 2023.

      (28) Andrew Wagenmaker, Lu Mi, Marton Rozsa, Matthew S Bull, Karel Svoboda, Kayvon Daie, Matthew D Golub, and Kevin Jamieson. Active learning of neural population dynamics using two-photon holographic optogenetics. Adv. Neural Inf. Process. Syst., 37:31659–31687, 2024.

    1. eLife Assessment

      This work presents a software and hardware suite for targeted photostimulation that can be used in vivo. The package is a well-designed and documented hardware/software suite with a comprehensive build guide. This tool will likely promote important neuroscience advances through targeted real-time perturbation of the cerebral cortex. Overall, this manuscript makes a compelling case on how to design and make available power tools for the research community.

    2. Reviewer #1 (Public review):

      Lohse et al. describe an open-source system for laser scanning photostimulation (LSPS) in head-fixed animals. Although similar systems have been developed and used by different groups, Zapit provides an open-source solution requiring few custom parts and minimal coding. This tool can clearly facilitate and speed the adoption of LSPS, particularly for the increasingly used purpose of mapping the effects of focal cortical silencing during behavior. Other potential uses include mapping optogenetically evoked movements and selectively activating genetically labeled neuronal subtypes of interest in the cortex. The design is well thought through, and the presentation is mostly clear and well written.

      In general, the more modular such a system is, the better, in terms of compatibility with existing hardware and software that potential users may already have purchased - laser, galvo, and camera in particular. The system has struck a reasonable balance between allowing modularity and providing an integrated complete package, but even more flexibility would be welcome for potential users looking to cut costs, as would clearer presentation of such flexibility as already exists.

      Comments and suggestions are mostly minor, as follows.

      (1) Command signals:

      How is the relationship between analog voltage commands and laser power determined? Is this assumed (or required) to be linear (as Figure 7F implies)? Usability and modularity would be improved by an option to measure or provide a calibration curve for systems with a nonlinear mapping between command voltage and laser power.

      For the grid calibration step, how is the initial mapping from galvo voltage commands to image position determined? Presumably, some sort of initial guess or calculation based on the hardware specifications is needed for the grid calibration to be feasible. Also, how are the number of grid lines and the distance between them determined?

      Why is the mapping between analog outputs and hardware (galvos, laser, masking light) fixed? This would be trivial to make configurable and allow labs with existing setups to adopt Zapit without rewiring existing hardware.

      (2) Laser and optics:

      In Figure 1, the authors should consider explaining the scanning principle schematically, i.e., depicting how tilting of the scan mirrors translates via the scan lens into beam displacement in the specimen plane. Perhaps Zemax can be used for accurate rendering.

      Since the unexpanded beam greatly under-fills the back aperture of the lens, the z resolution is presumably terrible - which is good! That is, for the purposes of LSPS, this advantageously avoids focus-dependent effects, which might otherwise arise due to (e.g.) skull curvature. The authors should consider pointing this out, as well as providing an estimate of the z resolution.

      What is the working distance?

    3. Reviewer #2 (Public review):

      Summary:

      In this work, Lohse and colleagues develop a system for doing targeted photostimulation in mouse cortex. The system uses a camera image to target laser stimulation to stereotactically defined locations in mouse dorsal cortex.

      Strengths:

      The hardware is well designed, and the software is well documented and supported. The build guide and well-documented software package should allow for simple implementation of the technology. Without a doubt, this is a valuable community resource for the circuit neuroscience field.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    4. Reviewer #3 (Public review):

      Zappit is an open-source implementation of arbitrary-access laser-scanning optogenetics for manipulation of neuronal activity in mice. As the method requires expertise ranging from optics, hardware control and programming, the authors make the point that this powerful strategy is underutilized in the field, and put forward a well-documented modular hardware and software platform aligned to the Allen Mouse Brain Atlas aimed at enabling the larger scientific community to use this approach (democratizing) for controlling cortical activity during behavior in mice.

      The authors favor a galvanometric approach to laser targeting. The system is inexpensive, easy to build, well-documented and user friendly (Matlab based GUI and GitHub repository). The photo-stimulation laser is directed into an X-Y galvo scanner targeted to the specimen using a dichroic mirror and focused on the sample using a Plössl lens as scan lens which is also used as an objective. The scan lens/objective images the specimen onto a camera via tube lens (also a Plössl lens) in a 0.5X magnification ensuring to fit the extent of the mouse brain onto the camera sensor (USB-3 Basler acA120-40um).

      The authors report short and reproducible onsite time (~ 0.5 ms) and block (mask) the stimulation source using the laser analog control (~0.5 ms). The system is reliable, aiming at up to 20 stimulation sites per sequence considered as quasi-simultaneous (10 ms). They minimize rebound by gentle ramping down of stimulation over 250 ms.

      The system is fast to calibrate by mapping scanner positions to pixel space in the camera space and mapping stereotaxic coordinate onto the image of the exposed skull. The theoretical x-y PSF is 70 µm (measured ~90µm) while the authors make the point that due to scattering the photo-stimulation spot size (lateral extent) is about 1 mm in diameter. This is what they also observe in electrophysiological recordings using silicon probes. The effective radius of inactivation depends on laser power, but was about 1 mm for laser powers (1-2-4 mW) on which the authors observed significant behavioral perturbations - in several tasks: 1) a delayed response somatosensory discrimination, 2) a visual detection task assessing changes in temporal frequency of a drifting visual stimulus; and 3) a visual discrimination (International Brain Laboratory task) in which mice were tasked to report the location of visual stimuli by turning a wheel. As proof of principle, the authors used a photo-stimulation set composed of 52 bilateral sites positioned at 0.5 mm interval covering a large network of frontal, motor and somatosensory cortical areas. Indeed, photo-inhibition of frontal motor cortex sites produced robust increases in reaction time. In contrast, stimulation at other motor and somatosensory sites produced modest, but significant decreases in reaction times.

      While the approach is not novel, it does serve the need of better disseminating this technique in the research community. Overall, the Zappit is well-documented and easy to build and use, and will have impact in increasing robust use of site directed photo-stimulation (exciting/inhibiting ensembles of neurons at particular ~1 mm size regions of interests across the dorsal surface of the brain). The authors also note that the axial resolution is ~1.5 mm.

      Concerns & comments:

      (1) While the authors argue that it offers the best utility to affordability trade-off - faster than motorized drivers and require much less power than DMDs (100X) and less expensive/easier to use compared to SLMs, in the current form, the manuscript does not clearly list the limitations of the approach. At such, in my opinion, the authors should include side by side comparisons (perhaps as a table). For example, clear statements should be included with respect to comparisons in lateral (x-y), axial (z) spatial resolution, as well as temporal sequential aspect of Zappit and other photo-stimulation techniques involving DMDs or SLMs.

      (2) Is power really a limitation in terms of the laser sources? Or is this a disadvantage mainly because using less power has beneficial effects on the tissue health? It may be useful to provide metrics of comparisons along these lines between Zappit and DMD-based approaches.

      (3) Arbitrary-scanning vs random scanning may be more appropriate to describe to strategy.

    1. eLife Assessment

      Mechanical transduction channels of sensory hair cells possess lipid scramblase activity. Membrane lipid disruption resulting from mechanical transduction is thought to be restored by flippase activities. This fundamental study provides compelling evidence that ATP8B1, a P4-ATP flippase and its subunit TMEM30B, are key in mediating this restorative function in outer hair cells of the mammalian cochlea.

      [Editors’ note, September 9, 2026: There is a potential confound with the findings shown in Figure 7 suggesting that specific targeting of ATP8B1 and TMEM30B subunits to the hair cell stereocilia is dependent on mechanoelectrical transduction channel activity. This does not affect the central findings or conclusions, but readers should be aware that a revised version of the paper is being prepared.]

    2. Reviewer #1 (Public review):

      Sensory hair cells of the inner ear convert mechanical sound vibrations into electrical signals through mechano-electrical transduction (MET). While the protein components of the MET machinery have been studied extensively, much less is known about how the surrounding membrane lipid environment contributes to hair cell function. The recent discovery that TMC1 and TMC2 also function as lipid scramblases has brought renewed attention to the importance of membrane lipid asymmetry and the mechanisms that maintain it in sensory hair cells.

      In this study, the authors identify the P4-ATPase ATP8B1 and its partner TMEM30B as key regulators of membrane lipid asymmetry in outer hair cells. Using complementary genetic models, HA-tagged knock-in mice, localization analyses, and functional experiments, they show that ATP8B1-TMEM30B is enriched in stereocilia and the apical membrane of outer hair cells and is required to maintain phosphatidylserine asymmetry, support hair cell survival, and preserve normal hearing. The parallels between the ATP8B1/TMEM30B loss-of-function phenotypes and TMC1 deafness-associated mutants with constitutive scrambling support a model in which ATP8B1-TMEM30B flippase activity maintains membrane lipid asymmetry and homeostasis, whereas constitutive TMC1-mediated phospholipid scrambling disrupts this balance and contributes to membrane instability.

      The authors have addressed the points raised during the initial review thoroughly. The revised manuscript includes clearer methodological details, additional physiological characterization, improved presentation and quantification of several datasets, and a more balanced interpretation of the localization and mechanistic findings. These changes improve both the clarity and rigor of the study while leaving its main conclusions unchanged.

      As with any study that opens a new area of investigation, important mechanistic questions remain. In particular, it will be interesting to determine how disruption of membrane lipid asymmetry ultimately impairs MET function and triggers hair cell degeneration, how flippase and scramblase activities are coordinated in vivo, and how these pathways are integrated with the broader molecular machinery underlying mechanotransduction. These questions highlight the exciting directions that this study opens for the field.

      Overall, this work provides evidence that ATP8B1-TMEM30B is a critical regulator of stereocilia membrane lipid asymmetry and represents an important contribution to our understanding of membrane homeostasis in auditory hair cells. I have no further major concerns and support publication.

    3. Reviewer #2 (Public review):

      Summary:

      Prior work identified TMEM30B (knockout mice) as well as ATP8B1 (human genetics and mouse model), ATP8A2 (knockout mice), and ATP811A (human genetics) as relevant for hearing. The authors also reasoned that given the recent discovery of TMC1 and TMC2's dual function as mechanotransduction channels of the inner ear and as lipid scramblases, a counterpart flippase should be in the sensory hair-cell stereocilia bundle where mechanotransduction happens. They use CRISPR/CAS to modify the endogenous mouse genes and add an HA tag at the N-terminus of the ATP8B1, ATP8A1, ATP8A2, and ATP11A proteins. Their experiments with these mice unambiguously localized ATP8B1 at the base of outer hair cell stereocilia bundles. Knockout of ATP8B1 results in loss of outer hair cells, deficient auditory function (ABR), and degeneration of outer hair cell stereocilia bundles. Similarly, hair cells from genetically modified mice with endogenous HA-tagged TMEM30B proteins show localization of this protein to outer hair cell stereocilia bundles. TMEM30B knock out mice phenocopy the ATP8B1 knock out model. Interestingly, the authors show that annexing V staining precedes hair cell loss in ATP8B1 and TMEM30B knockout mice and that proper localization of these proteins is lost in mice that lack CIB2, a protein essential for hair cell mechanotransduction.

      Strengths:

      (1) Use of knock-in HA-tagged proteins to unambiguously localize ATP8B1 and TMEM30B

      (2) Systematic characterization of auditory function (ABR), hair cell loss, and hair-cell stereocilia bundle morphology.

      (3) Advances our understanding of the role played by lipid homeostasis in auditory function.

      (4) Reports on mouse models that will be helpful to further understand the mechanistic role played by ATP8B1 and TMEM30B in normal hearing and hereditary deafness.

      Weaknesses:

      (1) Are the HA tags causing any functional issues? Function and localization of tagged proteins can sometimes be compromised. This is checked for TMEM30B and ATP8B1, but not for ATP8A1, ATP8A2, and ATP11A.

      (2) Following on the point above, is it possible that ATP8B1-HA is well localized, but localization for the other three flippases (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA) is compromised by the tag? Is this potential miss-localization causing any functional phenotypes? I find surprising that there are flippases only in outer hair cells and only formed by ATP8B1. A possible explanation is that the tag is interfering with trafficking. If so, there should be a phenotype (ABRs), although this might be masked by redundancy among these flippases or caused by systemic issues (admittedly difficult to sort out).

    4. Author response:

      The following is the authors’ response to the original reviews.

      Summary of Revisions Performed:

      We have clarified the qPCR methodology in the methods section and stated the housekeeping gene GAPDH to address potential misunderstandings.

      We have assessed hearing in the generated HA-tagged mouse lines and included an adequately powered ABR measurements analysis in the revised manuscript as a supplemental figure.

      We have included powered DPOAE experiments in both ATP8B1 and TMEM30B KO mice to strengthen the findings of the ABRs.

      We have clarified the presentation of the z-stack in Figure 1F.

      We have elaborated on the analysis for Figure 7B to strengthen comprehension by readers.

      We have revised the statement to read: “No IHC stereocilia-enriched P4-ATPases were detected under the conditions examined.”

      While we appreciate the suggestion to examine TMEM30B localization on the ATP8B1 KO background, this is not feasible within a reasonable timeframe; we have clarified this limitation in the manuscript.

      We have incorporated relevant prior work (e.g., George and Ricci, 2026) demonstrating minimal Annexin V labeling prior to P6 and lack of PS externalization in TMC1/2 double knockout models.

      We have clarified that hearing thresholds for TMEM30B-HA and ATP8B1-HA lines were addressed in this study, while additional HA-tagged flippase lines (ATP8A1, ATP8A2, ATP11A) are part of ongoing work to be reported separately.

      We have softened statements regarding HA-tag insertion and clarified that, to our knowledge, localization and function are not disrupted, while acknowledging this as a potential limitation.

      We have revised the Methods section to clarify differences in fluorescence measurements across experiments.

      Public Reviews:

      Reviewer #1 (Public review):

      Figure1D.

      The authors should clarify how the qPCR data were normalized and specify the reference (housekeeping) genes used. This information is necessary to evaluate the robustness and comparability of the gene expression data.

      We thank the reviewer for this comment. qPCR data were normalized to GAPDH as the reference (housekeeping) gene. We have clarified this in the Methods section to ensure transparency and reproducibility.

      (2) Figure 1F.

      The lack of F-actin staining at the hair cell base raises the possibility that the permeabilization conditions may have limited antibody access to certain membrane regions. This is especially important given that the authors used a gentle permeabilization agent such as saponin to preserve membrane integrity. Because the authors conclude that ATP8B1 and TMEM30B are localized "almost exclusively to OHC bundles and the apical membrane, with minimal staining in the remaining plasma membrane," (line 128). Including co-labeling with a plasma membrane marker or more comprehensive F-actin visualization of lateral and basal regions would help ensure that the restricted localization is biological rather than technical. In the absence of such controls, the localization claim may be somewhat overstated and should be tempered accordingly.

      We thank the reviewer for this important point. The image shown represents a single z-slice from a larger stack, and the hair cell body lies outside the plane of this section. To clarify this, we revised the accompanying text.

      (3) Figure 7B.

      Although quantification of ATP8B1-HA intensity at the bundle appears similar between WT and Cib2 KO samples, the representative image suggests that some bundles lack detectable labeling. To better capture phenotype variability, it would be helpful to include an additional quantification showing the fraction or number of bundles with detectable ATP8B1-HA signal in Cib2 KO mice.

      We thank the reviewer for this suggestion. We have clarified the quantification of the fraction of hair cell bundles with detectable ATP8B1-HA and TMEM30B-HA signal per field of view. Although the representative images may give the impression that some hair bundles lack staining, this is due to changes in ATP8B1-HA and TMEM30B-HA distribution within the cell body. In all cases, detectable ATP8B1-HA and TMEM30B-HA signal remained present in the hair bundles.

      (4) Lines 346-349

      The manuscript suggests that IHCs lack stereocilia-enriched P4-ATPases. However, this conclusion is not directly supported by the presented data. The authors should either provide supporting localization or expression data for other P4-ATPases or soften the statement to indicate that no stereocilia-enriched P4-ATPases were detected under the conditions examined.

      We agree with the reviewer and have revised this statement to read: “No IHC stereocilia-enriched P4-ATPases were detected under the conditions examined.”

      Recommendations:

      (5) The authors convincingly demonstrate that TMEM30B loss results in ATP8B1 mislocalization. While not essential to the central conclusions, examining TMEM30B localization in ATP8B1 KO hair cells would clarify whether this interdependence is reciprocal, as described for other P4-ATPase-CDC50 complexes.

      While we agree that this experiment would provide valuable information, performing it would require generation of a compound mouse line carrying both the TMEM30B-HA allele and the ATP8B1 knockout allele. This work is beyond the scope of the current revision and cannot be completed within a reasonable timeframe.

      (6) Lines 359-374. The discussion of Annexin V labeling is careful and balanced. This paragraph would benefit from referencing other studies that showed minimal Annexin V labeling in healthy P6 organ of Corti, reinforcing that robust PS externalization in the present study is pathological rather than developmental.

      We thank the reviewer for this suggestion and have incorporated relevant prior work, including George and Ricci (2026), which demonstrates minimal Annexin V labeling prior to P6 and further supports our interpretation.

      (7) Lines 392-399.

      The proposed feedback model linking MET activity and ATP8B1-TMEM30B localization is compelling. The discussion could be strengthened by noting that in TMC1/2 double knockout hair cells, PS externalization is not observed, consistent with the idea that flippase activity becomes critical specifically when scrambling occurs. The mislocalization observed in Cib2 KO hair cells further supports the coupling between TMC-mediated scrambling and flippase-mediated membrane restoration.

      We agree and have revised the text to include that TMC1/2 double knockout hair cells do not exhibit phosphatidylserine externalization, supporting the idea that flippase activity becomes critical in the context of scrambling.

      Reviewer #2 (Public review):

      Weaknesses:

      (1) Are the HA tags causing any functional issues? Function and localization of tagged proteins can sometimes be compromised. It would be good to know, for each knock-in model (TMEM30B, ATP8B1, ATP8A1, ATP8A2, and ATP11A), whether the HA-tagged protein is causing any issues with the mice and particularly with hearing (ABRs). Are these mice normal? Can they hear? These data are missing.

      We thank the reviewer for raising this important point. In this study, we focus on TMEM30B-HA and ATP8B1-HA mouse lines, while additional HA-tagged flippase lines (ATP8A1, ATP8A2, ATP11A) are part of ongoing work to be reported separately.

      Both TMEM30B-HA and ATP8B1-HA mice are viable and exhibit normal breeding and ageing. We have included adequately powered ABR measurements of both TMEM30B-HA and ATP8B1-HA which indicate wild-type–like hearing thresholds.

      (2) Following on the point above, is it possible that ATP8B1-HA is well localized, but localization for the other three flippases (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA) is compromised by the tag? Is this potential mislocalization causing any functional phenotypes? (ABRs of point 1). I find it surprising that there are flippases only in outer hair cells and only formed by ATP8B1. A possible explanation is that the tag is interfering with trafficking. If so, there should be a phenotype (ABRs), although this might be masked by redundancy among these flippases or caused by systemic issues (admittedly difficult to sort out). Given that this manuscript will likely become foundational, and that there is evidence that at least two of the other flippases are involved in hearing loss, it would be good to provide more information about the mice and HA-tagged proteins in the other knock-ins (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA). Depending on the data available for the knock-ins, the authors may want to discuss these scenarios and soften the statement indicating that inner-hair cells may lack flippase activity altogether.

      We appreciate this concern. To our knowledge, the HA tag does not appear to disrupt localization or function of the tagged proteins. However, we agree that this cannot be fully excluded. We have therefore softened our conclusions about IHC flippases and clarified that additional flippases (ATP8A1, ATP8A2, ATP11A) are under investigation and will be described in a separate study.

      (3) Expression of ATP8B1 at P0 (Figure 1D), when there should not be protein in outer hair cells yet seems high. Does this mean that other cells in the cochlea also express ATP8B1? Is this a concern?

      We thank the reviewer for this observation. We interpret the elevated ATP8B1 transcript levels at P0 as reflecting transcription that precedes detectable protein accumulation in OHC stereocilia. While expression in other cochlear cell types cannot be excluded, we did not detect ATP8B1-HA immunolabeling outside hair cells in the knock-in model.

      (4) Fluorescence scales in Figure 6 B and D and Figure 7 B and D are very different. So are the values for WT. One would expect that the WT would be similar in all cases (at least within the same compartments), given that the methods section indicates that "All images were collected using identical acquisition parameters, including zoom and laser power, across genotypes". If WT shows such variability, how can we compare?

      We appreciate the need for clarification. Identical acquisition parameters were maintained within each experiment used for direct comparison (e.g., within a given panel). However, different panels (e.g., Figures 6B vs. 6D) were acquired on different days using different imaging settings.

      We have revised the Methods section to explicitly state this and clarify that comparisons are intended only within panels, not across experiments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 42: When discussing TMC similarity to TMEM16 scramblases, it may be helpful to mention that some TMEM16 family members (TMEM16A and B) function as ion channels, highlighting the dual ion/lipid functionality within the superfamily. The similarity to TMEM63/OSCA ion channels and lipid scramblases could also be noted. The fact that TMC, TMEM16, and TMEM63/OSCA belong to the same superfamily would provide a broader context. I also suggest referencing the work that initially suggested this relationship: (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0192851)

      We have included this citation and expanded on this discussion in the revised version.

      (2) Line 45: Consider including the recent Cryo-EM structure of CeTMC2 (PNAS, 2023), which provides structural insight into TMC-lipid interactions.

      We have included this citation in the revised version.

      (3) Line 62: For precision, consider removing "calcium-activated," as caspase-activated scramblases also disrupt membrane asymmetry.

      For precision, we have removed “calcium-activated” in the revised version.

      (4) Line 71: Clarify whether this refers to "fusion of membranes" or "cell fusion."

      We have clarified this statement to mean cell-cell fusion.

      (5) Line 85: Consider citing studies showing constitutive PS externalization in TMC1 mutant mouse models linked to deafness.

      We have added a citation to show that constitutive PS externalization is linked to deafness (Ballesteros and Swartz, 2022, and Beurg et al. 2025).

      (6) Line 93: TMEM30C is not discussed. A brief comment on its expression or relevance in hair cells would provide completeness.

      We have added a brief statement regarding TMEM30C and cited prior work describing its expression pattern (Osada et al. 2007).

      (7) Figure 1A: Use distinct colors for the P4-ATPase and CDC50 subunit rather than a rainbow scheme to improve clarity.

      We have retained the original color scheme in this panel.

      (8) Figures 3C-D and 5C-D: Increase legend symbol size for clarity. Update Y-axis labels to "Number of OHCs/100 μm" and "Number of IHCs/100 μm." Correct "um" to "μm."

      We changed the legend to improve the presentation of these panels to be more legible and changed the measurement to μm.

      (9) Figures 3F, 5F, 5H: Add scale bars.

      We have added scale bars to these figures.

      (10) Figure 7: The confocal images (A, C) show the bundle on top and cell body below, but the quantification (B, D) is in the opposite order. Reorganizing the panels for consistent orientation would improve clarity.

      We have reorganized the panels to improve clarity.

      (11) ABR measurements: Please specify the sex of the mice tested or clarify whether both sexes were included.

      We have included both male and female mice in this study as there were no differences in hearing function. We have added this clarification to the methods section under hearing tests.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 1A, the panels show CDC50. I would either change to TMEM30B or mention in the caption that TMEM30B is also known as CDC50 as labeled in the figure.

      We have changed CDC50 to TMEM30B.

      (2) Figures 1F and 1G are missing scale bars.

      We have added scale bars to these figures.

      (3) Figures 2 C, D, and 5 C, D - difficult to tell what's what in the legend. Perhaps make symbols larger in front of WT P17, KO P17, etc.?

      We changed the legend to improve the presentation of these panels.

      (4) Text under "TMEM30B is required for hearing and OHC maintenance". There is a difference in phenotype between the TMEM30B (Figure 5C) and ATP8B1 (Figure 3C) knockouts that is not discussed, as apical and middle cells seem to be okay. Should this be discussed?

      We appreciate this observation. We have elected not to expand the discussion of these regional differences because apical and middle hair cells also undergo degeneration at later ages (after P30), suggesting that the observed differences primarily reflect the timing of degeneration rather than distinct underlying mechanisms.

      (5) In the discussion text, under "Why do ATP8B1/TMEM30B-deficient OHCs die?", "Tmc1/2 or Cib2" should probably be "TMC1/2 or CIB2" or "Tmc1/2 or Cib2"

      We have changed this to read TMC1/2 or CIB2.

      (6) The methods section states "..., whereas non-significant comparisons are not shown." However, non-significant p values are shown in Figures 7B and D (bottom panels).

      We have removed the nonsignificant comparisons from Fig 7B and D to be consistent with the rest of the paper.

    1. eLife Assessment

      This study used several approaches (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial recordings) to address a significant question in epilepsy research, with additional relevance to EEG studies more broadly: Are high frequency oscillations (or "fast ripples", defined as >200 Hz by the authors) distinct from randomly occurring clustering of spikes? The results suggest fast ripples can occur by chance and how this may occur. The significance was considered important and the strength of evidence convincing, with minor limitations related to the need to address behavioral state, explaining the results in relation to epileptiform activity described by others, and discussing implications.

    2. Reviewer #1 (Public review):

      Summary:

      This is a study utilizing several types of analyses (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial multi-scale recordings) to address a highly relevant conceptual question: Are fast ripples (FRs) distinct pathological entities or largely emergent products of stochastic spike clustering? The results can potentially reshape current approaches to incorporating fast ripples into the epilepsy surgery evaluation.

      Strengths:

      The conceptualization of fast ripples as potentially arising by chance is highly novel and builds effectively on questions raised in prior studies that have never been satisfactorily resolved. Integration across biological scales and models provides a rigorous approach, now improved by addressing theoretical concerns regarding validity of the shuffling approach and state dependence. The discussion has been updated to provide a more nuanced interpretation of the study's findings.

      Weaknesses:

      The authors have satisfactorily and thoughtfully addressed the critiques provided in the first review. However, there remain two points that I would like authors to address:

      (1) Synchronized burst firing is a key feature of an epileptic site generating interictal discharges, and one that could generate either oscillatory or stochastic FRs as documented in multiple prior publications cited in the manuscript and/or in the prior review. Paroxysmal depolarization, for example, has been very well described, and consists of strong, disorganized burst firing (resulting in summated postsynaptic potentials strong enough to generate high gamma signal) in a neuronal population coinciding with a large low-frequency deflection. I would like to see the results described in this context, and to avoid blanket dismissal of stochastic FRs without a clear oscillatory component.

      (2) It would be highly useful to add a conclusion paragraph that spells out implications of the study for use of FRs as epileptic biomarkers in clinical invasive EEG recordings.

      Please address the above critiques in Discussion, or elsewhere as deemed necessary by the authors.

    3. Reviewer #2 (Public review):

      Summary:

      This paper asks an important question that has not been discussed much in the extensive literature on the High Frequency Oscillations (HFOs) that have been extensively studied in patients with epilepsy and experimental models of epilepsy. The question is whether the Fast Ripples (FRs), the HFOs in the 250-500 Hz frequency band, represent a pathological phenomenon or represent a physiological phenomenon that occurs in the healthy brain but happens to be more frequent in epileptic tissue. It is an important question that has not been systematically addressed until now. The authors conclude, from very extensive simulations, from extensive experimental animal studies (the systemic kianate model of epilepsy in rats), and from a modest amount of human data, that FRs occur in healthy brains as a result of the chance occurrence of bursts of action potentials, and that in epileptic tissue, their frequency of occurrence is approximately 30% higher than what is expected by chance. They conclude that FRs are not a separate phenomenon of epileptic tissue. This finding is reinforced by the recent findings of FRs in experimental models of Alzheimer's disease.

      Strengths:

      This is a valuable study because it asks an important and original question and because it evaluates it from several angles (simulation, tissue culture, experimental animals, and human patients). The simulations and the analyses of real data are performed very carefully and with original and solidly documented approaches, using extensive simulations and extensive data sets in the cultured cell data and in the in vivo experiments. The paper is clearly written and well-illustrated.

      Comments on revised version.

      The authors have appropriately addressed the questions I raised in the first review.

    4. Reviewer #3 (Public review):

      Summary:

      An outstanding question in the field of high frequency oscillations (HFOs) in the context of epilepsy is how these oscillations emerge, considering that they occur at such high frequencies i.e., 250Hz well above the firing ability of single neurons. One hypothesis that has been suggested in the past is that neurons that fire in an out of phase fashion or rather at random intervals may contribute to a spectrum of HFOs ranging from 250-500Hz that observed in epilepsy. However, how possible it is that random action potentials could aggregate to the extent that they could give rise to HFOs in the so-called fast ripple (FRs) frequency range (>200 according to the authors) remains unclear. To test this hypothesis, they used computational modeling to randomly insert action potentials in a signal, and they found that this approach is sufficient to generate FRs. Some of the predictors of whether FRs could occur were neuronal count, firing rate and synchronization. Besides computational modeling, they used different model systems to test whether that would be possible to be observed in neuronal cultures, in epileptic rats (intrahippocampal kainic acid model), and human data. Neuronal cultures treated with picrotoxin did not show evidence that FRs could be generated more than chance aggregation of action potentials. They then asked whether synchronization and firing rate could play a role in the emergence of FRs. They found that changes in neural firing and synchronization, such as those occurring during differences phase of the sleep-wake cycle could affect the number of FRs occurring by chance aggregation, with more FRs seen during periods of wakefulness, a result that they replicated in human data.

      The authors largely achieve their proposed aims of demonstrating that random neuronal firing can, in principle, generate FRs. Results from this study could influence current thinking around mechanisms generating FRs in epilepsy. The use of different computational approaches and model systems could offer new analytical methodologies for the study of FRs in the context of brain disease.

      Strengths:

      (1) The authors used a multi-level approach combining computational modeling with experimental datasets, including neuronal cultures, a rat model of temporal lobe epilepsy and human data.

      (2) Identification of key parameters such as neuronal count, firing rate, synchronization and brain state in observed incidence of FRs generated through random aggregation of neural firing.

      (3) Cross-species validation increases the likelihood of generalizability of the findings.

      Minor weakness:

      (1)The analyses conducted in human data lack direct comparison with sleep data due to no available data, but would encourage future investigations directly comparing HFOs during wakefulness and nocturnal sleep.

      Comments on revised version.

      The authors have addressed my comments and I have no further suggestions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a study utilizing several types of analyses (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial multi-scale recordings) to address a highly relevant conceptual question: Are fast ripples (FRs) distinct pathological entities or largely emergent products of stochastic spike clustering? The results can potentially reshape current approaches to incorporating fast ripples into the epilepsy surgery evaluation.

      Strengths:

      The conceptualization of fast ripples as potentially arising by chance is highly novel and builds effectively on questions raised in prior studies that have never been satisfactorily resolved.

      The integration across biological scales and models is a major strength. The state dependency analysis provides additional, strong support. The methodology and statistical approaches used are thoughtfully presented and rigorously applied.

      In particular, this paper provides a strong response to the findings from Gliske et al, Nat Commun 2018. This study utilized long-term data analysis to uncover low rates of FRs detected from most recording sites, suggesting spurious detections, although FRs were concentrated within seizure onset areas.

      We fully agree with this comparison. Although we had already cited this paper, we now further emphasize this observation in the Discussion:

      “Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence.”

      Weaknesses:

      The authors clearly aimed to use a statistical rather than a mechanism-based approach in this work. However, the paper's framing of true fast ripples as oscillatory events with stochastic fast ripples considered as confounders does not take prior investigations into biological mechanisms, particularly prior studies that point to an important role for stochastic fast ripples in some contexts. Incorporating recognition of these mechanisms would strengthen the manuscript and provide a more complete and nuanced characterization.

      Some examples from the literature:

      Eissa et al, eNeuro 2016, a paper that closely parallels this manuscript but took a mechanistic rather than statistical approach, showed that fast ripples can arise from population paroxysmal depolarizations - a key feature of epileptiform discharges - as temporally clustered, jittered population firing, with FRs appearing in LFP or EEG due to summated postsynaptic potentials (which are slower than action potentials and can generate signals in the high gamma range).

      Foffani et al., 2007, Neuron, and Ibarz et al., 2010, J Neurosci, argue that FRs are pseudo-oscillations created by jittered neuronal populations in the setting of altered spike timing.

      Smith et al., 2020, Sci Rep, contrasts FR characteristics in different regimes, i.e., intact inhibition early in a seizure vs. implied collapse of inhibition after recruitment. Schlingloff et al., 2025, J Neurosci, reported analogous findings in an animal model.

      We agree with the reviewer that even stochastic events may be of biological importance and an increase in stochastic events will occur when there is an increase in synchronisation and excitability, two properties of pathological cortex. We also don’t disagree that FRs can occur as distinct entities, although our work indicates that most are due to chance.

      To address this point, we have clarified our claims in the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      In addition, we expand on these points at various junctures in the Discussion. In particular, we reiterate our assertion that FRs may still be a useful biomarker, but that their interpretation should be moderated to reflect the fact that they often occur by chance:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      The computational model and subtraction approach provide a strong case for the random emergence of clustered activity in the high gamma band, given its assumptions. However, any such modeling effort needs to account for inhibitory activity, including impaired inhibitory function that is expected in epileptic brain regions, which has a strong modulating effect on excitatory firing and is thought to play a significant role in FR generation.

      We appreciate the reviewer’s concerns, but we believe that the impact of inhibitory interneuron activity on excitatory firing rates and synchronisation is incorporated indirectly into our simulations, while keeping our model as parsimonious as possible by not directly incorporating interneuron activity into our simulations. We have addressed this point in the Methods section:

      “Varying synchrony allowed us to test the impact, on the network, of inhibitory cells, which have been shown to favour synchrony (Bocchio et al., 2024; Cobb et al., 1995).”

      The shuffling procedure aims to preserve the power spectrum but randomizes high frequency phase (>200 Hz). However, this procedure removes biologically meaningful spike timing correlations, as well as structured cross-frequency coupling. The subtraction method thus likely underestimates the incidence of structured "distinct" FRs, while perhaps overestimating "chance" FRs due to biologically infeasible activity, making the statement that most FRs are due to chance correlation too strong.

      We appreciate this concern, which is especially important given that our results depend crucially on the validity of our shuffling procedure (as described in the Discussion). To address this issue, we have implemented an additional shuffling algorithm that preserves cross-frequency coupling (see last section of the Results, especially Supplementary Fig. 10f). This method showed no qualitative difference, compared with other alternative methods presented in Supplementary Fig. 10. These new results are described in the Methods section:

      “Last, we also implemented a method based on wavelet-IAAFT with preservation of cross-frequency coupling, since fast ripples are typically locked to low-frequency phase (Sheybani et al., 2019). The code detects the highest phase-amplitude coupling (PAC) in the original signal between [300-6000 Hz] for amplitude and several low-frequency bands ranging from 2-20 Hz, bandwidth of 3 Hz. PAC is computed using the modulation index (Tort et al., 2008). Then, in the shuffled signal under construction and during convergence testing of PSD (see above), the PAC between high-frequency part of the signal (300-6000 Hz) and the identified low frequency for phase is normalized to that of the highest PAC identified earlier.”

      The kainate findings underscore this point: the increase in the number of FR detections could be, as the authors state, an increase in chance clustering due to increased network excitability generally. However, the likelihood of a parallel increase in pathological FRs cannot be ruled out, given likely pro-epileptic alterations in spike timing and circuit function.

      We appreciate the reviewer’s point but wish to re-emphasise our interpretation of these findings – that the observed increase in the incidence of FRs occurs as a result of increased network excitability/synchrony, secondary to the pathological mechanisms of epilepsy. We have updated the Discussion accordingly:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-duration FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      To further emphasise this important point, we have also updated the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      Reviewer #2 (Public review):

      Summary:

      This paper asks an important question that has not been discussed much in the extensive literature on the High Frequency Oscillations (HFOs) that have been extensively studied in patients with epilepsy and experimental models of epilepsy. The question is whether the Fast Ripples (FRs), the HFOs in the 250-500 Hz frequency band, represent a pathological phenomenon or represent a physiological phenomenon that occurs in the healthy brain but happens to be more frequent in epileptic tissue. It is an important question that has not been systematically addressed until now. The authors conclude, from very extensive simulations, from extensive experimental animal studies (the systemic kianate model of epilepsy in rats), and from a modest amount of human data, that FRs occur in healthy brains as a result of the chance occurrence of bursts of action potentials, and that in epileptic tissue, their frequency of occurrence is approximately 30% higher than what is expected by chance. They conclude that FRs are not a separate phenomenon of epileptic tissue. This finding is reinforced by the recent findings of FRs in experimental models of Alzheimer's disease.

      Strengths:

      This is a valuable study because it asks an important and original question and because it evaluates it from several angles (simulation, tissue culture, experimental animals, and human patients). The simulations and the analyses of real data are performed very carefully and with original and solidly documented approaches, using extensive simulations and extensive data sets in the cultured cell data and in the in vivo experiments. The paper is clearly written and well-illustrated.

      Weaknesses:

      I found only one serious weakness in this study, but it is one that is of importance. Although the original work on FRs was done in an experimental model of epilepsy, the field really became prominent when ripples and fast ripples were found first in microelectrode recordings of epileptic patients and then in the intracerebral EEG of such patients. Numerous studies have been performed since then, with a valuable meta-analysis including 700 patients (Wang Z, Guo J, van 't Klooster M, Hoogteijling S, Jacobs J, Zijlmans M. Prognostic Value of Complete Resection of the High-Frequency Oscillation Area in Intracranial EEG: A Systematic Review and Meta-Analysis. Neurology. 2024 May 14;102(9). Although the consensus at this point is that FRs are not the ideal and totally specific marker of epileptic tissue that many thought it could be, FRs are nevertheless much more frequent in epileptic tissue than in non-epileptic tissue and are a solid biomarker.

      We agree with the reviewer, and do not intend to challenge the role of FRs as a marker of the seizure-onset zone, and potentially the epileptogenic zone. Instead, the aim of this study was to address the question of whether FRs are generated by intrinsic pathological mechanisms, or whether they arise due to the chance co-occurrence of action potentials that follow different dynamics in epileptogenic parenchyma. We have updated the Discussion accordingly:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      To further emphasise this important point, we have also updated the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      It is also well established that they are much more frequent in NREM sleep than in wakefulness, as reported in the original paper of Staba et al (Staba RJ, Wilson CL, Bragin A, Jhung D, Fried I, Engel J Jr. High-frequency oscillations recorded in human medial temporal lobe during sleep. Ann Neurol. 2004 Jul;56(1):108-15., not mentioned in this paper) and in the study of Bagshaw et al (2009). In this last paper, using SEEG in various brain regions, the average rate of FRs in NREM sleep is about 6 times that in wakefulness. In the paper by Staba, with microelectrodes in mesial temporal structures, it is about twice. As a separate issue, the paper of Fraucher et al (Frauscher B, von Ellenrieder N, Zelmann R, Rogers C, Nguyen DK, Kahane P, Dubeau F, Gotman J. High-Frequency Oscillations in the Normal Human Brain. Ann Neurol. 2018 Sep;84(3):374-385), which is not quoted, found that, in an extensive sample, non-epileptic human tissue sampled with SEEG generated extremely rare FRs (an average rate of 0.04/min/channel, i.e. 1 every 25 min).

      The results above are mentioned because they do not fit with the data provided in the present study: FRs are much more frequent in NREM sleep than in wakefulness in human epileptic patients, and they are much more frequent (not 30% more, but many hundreds of percent more) in epileptic tissue than in non-epileptic human tissue. The fundamental phenomenon of interest is, I believe, the FRs in epileptic patients. The animal experiments, tissue studies, and simulations are models to study the human phenomenon. With respect to the modulation by sleep and the differentiation between epileptic and non-epileptic tissue, it seems that the systems studied in this paper are not good models of the human condition. The human results presented in the study only reflect wakefulness recordings, which is not the condition in which most HFO studies have been done and in which most HFOs occur. The authors refer to the study of long-term fluctuations in HFO rates by Gliske et al. (2018) to say that one has to be careful with the results regarding sleep, for example, Bagshaw et al (2009), but the clear predominance in of HFOs in NREM sleep has been observed by many studies. The cautions regarding fluctuations over extended periods also apply to the awake human data analyzed in this study. The study's conclusions regarding the generation of FRs are therefore questionably applicable to the human condition. I do not dispute their validity for the models and situations in which they were studied.

      We looked at this in more detail. Our simulations were intended to test how the incidence of FRs can vary with different parameters of network activity (neuronal count, firing rate, synchronization). Indeed, since their incidence is known to vary across regions and within regions and across states, we wanted to test how FRs are controlled by different factors. As such, we do not wish to draw firm conclusions about the observed sleep-wake changes in FR incidence in rodents, and how it relates to humans – evidence shows that pathological FRs in rodents do not display state-specific preferential occurrence (Ewell et al., 2019). We have added new text to the Abstract and Discussion to emphasize this.

      Abstract:

      “Our simulations showed that chance aggregation can generate fast-ripples and that their incidence changes depending on brain state, an observation that we confirmed in our rodent data.”

      We acknowledge that previous publications have reported higher rates during sleep, although with shorter recordings than in our rodent recordings (Staba, 2004: one night; Bagshaw, 2009: 10 min; Frauscher, 2018: 20 min – only sleep recordings). We have rewritten the part of the Discussion on the effect of the sleep-wake cycle on FRs incidence:

      “In our rodent data, we were initially surprised to find a higher rate of FRs during wakefulness, which contrasts with previous reports in humans (Bagshaw et al., 2009; Staba et al., 2004). However, previous studies only indicate that physiological vs pathological FRs are more easily distinguished during NREM sleep (von Ellenrieder et al., 2016) and that their incidence varies during sleep (Von Ellenrieder et al., 2017), but in hours-long recordings, no differences in incidence have been reported in the mesial temporal lobe (Dümpelmann et al., 2015). Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence. Last, but not least, another report did not find a state-dependent expression of FRs in the kainate rat model of temporal lobe epilepsy (Ewell et al., 2019), thus indicating that the variability of FRs across sleep and wake is still an open question, at least in rodents. Hence, the main conclusion on the effect of sleep-wake transitions is that these transitions impact the likelihood of stochastic events, more than dictating the direction (increases vs decreases) of change. It also highlights that the specificity of FRs to epileptogenic parenchyma could vary across the sleep-wake cycle, which would be crucial in epileptology (Dimakopoulos et al., 2024; Roehri et al., 2018; Sheybani et al., 2019, 2018; Zijlmans et al., 2012, 2009). Hence, FRs reflect and are highly susceptible to changes in network excitability.”

      Reviewer #3 (Public review):

      Summary:

      An outstanding question in the field of high-frequency oscillations (HFOs) in the context of epilepsy is how these oscillations emerge, considering that they occur at such high frequencies, i.e., 250Hz, well above the firing ability of single neurons. One hypothesis that has been suggested in the past is that neurons that fire in an out-of-phase fashion, or rather at random intervals, may contribute to a spectrum of HFOs ranging from 250-500Hz that are observed in epilepsy. However, how possible it is that random action potentials could aggregate to the extent that they could give rise to HFOs in the so-called fast ripple (FRs) frequency range (>200 according to the authors) remains unclear. To test this hypothesis, they used computational modeling to randomly insert action potentials in a signal, and they found that this approach is sufficient to generate FRs. Some of the predictors of whether FRs could occur were neuronal count, firing rate, and synchronization. Besides computational modeling, they used different model systems to test whether that would be possible to be observed in neuronal cultures, in epileptic rats (intrahippocampal kainic acid model), and human data. Neuronal cultures treated with picrotoxin did not show evidence that FRs could be generated beyond chance aggregation of action potentials. They then asked whether synchronization and firing rate could play a role in the emergence of FRs. They found that changes in neural firing and synchronization, such as those occurring during differences phase of the sleep-wake cycle, could affect the number of FRs occurring by chance aggregation, with more FRs seen during periods of wakefulness, a result that they replicated in human data.

      The authors largely achieve their proposed aims of demonstrating that random neuronal firing can, in principle, generate FRs. Results from this study could influence current thinking around mechanisms generating FRs in epilepsy. The use of different computational approaches and model systems could offer new analytical methodologies for the study of FRs in the context of brain disease.

      Strengths:

      (1) The authors used a multi-level approach combining computational modeling with experimental datasets, including neuronal cultures, a rat model of temporal lobe epilepsy, and human data.

      (2) Identification of key parameters such as neuronal count, firing rate, synchronization, and brain state in observed incidence of FRs generated through random aggregation of neural firing.

      (3) Cross-species validation increases the likelihood of generalizability of the findings.

      Weaknesses:

      (1) Some of the simulated FRs appear short in duration and may not meet standard detection and definition criteria, potentially influencing validity.

      We thank the reviewer for raising this important concern. To address this issue, we quantified and compared the duration of FRs in original and shuffled rodent data. Consistent with the reviewer’s suspicions, we found that FRs in shuffled signals are shorter than FRs in original signals. This is important because it shows that: (i) a longer duration should be considered a core feature of genuine FRs; and (ii) depending on the basal duration of FRs, the shuffling procedure will lead to different ratios of genuine to stochastic FRs. We have updated the Results accordingly:

      “These findings demonstrate the challenge of identifying distinct FRs within a composite population of distinct and stochastic events. One parameter that could help disentangle these events is their duration. Indeed, one might expect stochastic events to be more likely to be short-lived, since the probability of consecutive APs continuing to co-occur across neurons decreases over time. Hence, we next compared the distribution of FR durations between original and shuffled rodent data and found that FRs in shuffled data are shorter than those in original data (Supplementary Fig. 9). This makes duration a key feature that could help identify distinctly generated FRs.”

      And Discussion accordingly:

      “In addition, we show that long durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      (2) The neuronal culture approach does not directly test random insertion of action potentials, limiting interpretation.

      Neither the neuronal culture approach, the rat data or the human data directly test random insertion of action potentials. The insertion of random action potentials is only performed in the simulated data to test if FRs can arise from the chance insertion of action potentials. Once this was confirmed in the simulations, we then used the shuffling procedure in biological data to test if FRs are more frequent than expected by chance.

      (3) Sleep is treated as a homogeneous state in the rat dataset, without accounting for stage-specific differences in synchronization, which may affect the results and interpretation.

      We agree with the reviewer, but our primary aim was to answer the question of whether FRs can arise by chance. Although it was interesting to see that, in our longitudinal rodent data, the incidence of FRs varies across the sleep-wake cycle, any sleep-stage-specific changes are beyond the scope of this work.

      (4) The analyses conducted in human data lack direct comparison with sleep data.

      We agree that it would have been useful to investigate variations in the incidence of FRs across the sleep-wake cycle in human microelectrode recordings. Unfortunately, however, such sleep recordings were not available. Hence, while we cannot compare variations in FR incidence across brain states between humans and animal models, our conclusions that FRs arise mostly by the chance co-occurrence of action potentials still holds.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please indicate where corrections for multiple comparisons were used.

      P-values corrected for multiple comparisons are indicated by the accompanying phrase: “adjusted p-value”. We had previously omitted to mention this once in the Results, which we have now corrected.

      (2) Delta amplitude is likely sufficient for detecting sleep-wake transitions, but the beta/delta ratio is better supported in the literature. Do the results change if beta activity is incorporated?

      We have now computed the beta (15-40 Hz) to delta (0.5-4 Hz) ratio and find that this is closely correlated with delta across time. We have updated the Results accordingly:

      “Importantly, these findings were robust to the specific method used to detect FRs (Supplementary Fig. 4c) (Padmasola et al., 2024; Sheybani et al., 2019, 2018). Also our use of delta power to identify periods of presumed wakefulness and sleep was highly (negatively) correlated with an alternative method of using the beta-to-delta ratio across time (another marker of increased vigilance; (Fraigne et al., 2023), see Supplementary Fig. 4d).”

      Methods:

      “We further verified that delta power across time displayed similar fluctuations to beta (15-40 Hz) to delta power ratio, another marker of vigilance (Fraigne et al., 2023).”

      And we updated Supplementary Fig. 4d

      “(d) Beta to delta power ratio across time is superimposed over delta power across time. There is a strong (inverse) correlation between the two time-series (inset), which is confirmed by the correlation coefficient across animals (right).”

      (3) Figure 2's axis labeling with the 3D plots is hard to read.

      We have enlarged the font size.

      (4) The scaling of the histogram in Figure 3 is unclear.

      This was on omission. The scale has now been added to Figure 3.

      (5) There is a risk of overfitting in the regression model. Was cross-validation used?

      We have now repeated this analysis with cross-validation, without any qualitative impact on the results (e.g. the model still performs well above chance). We have updated the Methods:

      “To further confirm the performance of GBT, we used a cross-validation procedure where the GBT is trained on 80% of data and then tested on the 20% remaining. The procedure is repeated 1000 times and the r<sup>2</sup> is saved at each round. We repeated the analysis with randomization of the outputs across 1000 rounds and saved this null distribution r<sup>2</sup>. We then compared the performance against original data.”

      Legend of Fig. 3:

      “(e) Performance of the GBT classifier using cross-validation (training: 80% of data; test: 20% remaining) using original (orange) and shuffled (blue) data. The difference is significant (paired t-test, p<0.0001).”

      And Results:

      “Furthermore, using a cross-validation approach with 80% of the data as training set and the remaining 20% as the test set, we obtained a significantly higher explained variance than when outputs were shuffled across the 125,000 solution points (paired t-test, p<0.0001, Fig. 3e), […]”

      Reviewer #2 (Recommendations for the authors):

      Maybe I missed it, but I did not find the length of human data analyzed or how the sections were selected.

      Apologies for this omission. The methods have been updated accordingly:

      “Microwire signals were selected based on high signal-to-noise ratio, as reflected by the detection of ≥ 1 single unit. Duration of recordings was of (median, interquartile range) 10 min and 17 s [3-13 min] and number of electrodes per patient was 4.5 [2.75-8].”

      The authors use the term "virtual simulation", which I find odd. I think the simulation is very real in the sense that it simulates reality, and I do not understand how a simulation can be virtual.

      We have updated the manuscript accordingly.

      Reviewer #3 (Recommendations for the authors):

      Major Comments:

      (1) In Figure 1, the authors suggest that random insertion of action potentials in a signal is sufficient to yield FRs. However, the observed FRs shown in panel 1b (also in supplemental Figure 5) seem pretty short in duration and may not meet the mentioned criteria in methods that require at least 4 cycles and ".whose amplitude is 3 times that of the surrounding baseline..". Moreover, in panel 1b, it seems that the FR shows a candle-like appearance, which has often been associated with filtering of sharp transients. How did the authors validate that the detected FRs were "real" FRs?

      Given the very large amount of data, it was not possible to visually verify all FRs. However, FRs were detected with published methods (Roehri et al., 2016; Roehri et al., 2017; and Sheybani et al., 2018 for confirmation of 24-hour variability in rodents) that have subsequently been used in several publications.

      Regarding the candle-like appearance of the spectrogram, the Delphos algorithm precisely looks for isolated “islands” of increased power (see Roehri et al., 2018, Ann Neurol), thus excluding any candle-like appearance. Similarly, the detector in Sheybani et al. (2018) J Neurosci first detects candidate FRs but then excludes those that are associated with a peak in lower frequencies, thus also limiting the risk of detecting candle-like events.

      Regarding duration, we have compared the duration of FRs in original and shuffled rodent data and found that FRs in original signals are indeed longer. This makes duration a key feature to identify distinct FRs. We have updated the Results accordingly:

      “These findings demonstrate the challenge of identifying distinct FRs within a composite population of distinct and stochastic events. One parameter that could help disentangle these events is their duration. Indeed, one might expect stochastic events to be more likely to be short-lived, since the probability of consecutive APs continuing to co-occur across neurons decreases over time. Hence, we next compared the distribution of FR durations between original and shuffled rodent data and found that FRs in shuffled data are shorter than those in original data (Supplementary Fig. 9). This makes duration a key feature that could help identify distinctly generated FRs.”

      (2) In the context of neuronal cultures, it is unclear how it could be deducted that the result relates to chance incidence of action potentials considering that no random action potentials were inserted, but only random shuffling of the high frequency component of the signal was attempted "Hence, neural networks with limited complexity (Kim et al., 2020; Saglam-Metiner et al., 2024; Sanchez-Vives and McCormick, 2000; Timofeev and Chauvette) fail to generate FRs beyond that expected from the chance coincidence of APs, even after increasing network excitability."

      FRs arise from series of action potentials occurring at a delay corresponding to their oscillatory frequency (250-500 Hz). Simulations demonstrated that FRs can occur by chance. When the EEG is shuffled, the only FRs that remain are those occurring by chance, because those occurring as individual entities have been broken up. Hence, if the original EEG displays more FRs than the shuffled EEG, then it means that these additional FRs were generated as individual entities. We have improved the Results section to clarify this:

      “We hypothesized that if FRs arise purely from chance firing, then temporally shuffling these recordings while conserving their spectral properties (Supplementary Fig. 3) would disrupt any oscillatory structure, leaving only FRs that occur due to chance.] Any additional FRs in the original data, compared to the number of FRs in the shuffled EEG, should thus be assumed to be individual entities.”

      (3) In the rat dataset, sleep was treated rather homogenously, without accounting for the sleep stage that is characterized by different synchronization and firing. An analysis of different sleep stages would be valuable.

      Although we agree that it would be scientifically interesting, we believe that our claim – that the ratio of genuine to stochastic FRs changes across the sleep-wake cycle – would hold. Unfortunately, lack of EMG prevents us from performing reliable sleep scoring. However, we do now include an alternative method for differentiating sleep from wake using the beta-to-delta ratio, which was highly correlated with delta activity, supporting our previous approach. Please refer to Supplementary Fig. 4d for further information.

      (4) The authors found that chance aggregation was highest during periods of wakefulness. Analyses of human data also confirmed that FRs could occur by chance aggregation during wakefulness. However, a comparison with sleep data would further strengthen this finding.

      We fully agree, but unfortunately, we do not have sleep data using microwires. Although our central claim – that FRs can occur by chance clustering of action potentials – would hold, we agree that it would have been scientifically interesting to add sleep data.

      (5) The statistics section would benefit from addressing how normality was determined and power analysis, as well as the inclusion of the exact sample size for all experiments.

      With large sample sizes, ANOVA and linear mixed models are robust to non-normality. Given the large sample sizes of our data, we thus used ANOVA and linear mixed model. For tests with small sample sizes where normality was violated, we used non-parametric tests, indicated by their name, e.g., Wilcoxon test for Supplementary Fig. 3b.

      (6) Greater discussion on the implications of this study for proposed in-phase or out-of-phase FR generation mechanisms is suggested.

      We have added further discussion on this. In the aim to keep the Discussion short and impactful, we could not elaborate too much. We have synthetized other parts of the Discussion to keep it within the right length. Here is the additional part:

      “It has been argued that the very high frequency that can be obtained during FRs are due to out-of-phase firing of excitatory neurons (Foffani et al., 2007; Ibarz et al., 2010), which is also consistent with our concept of stochastic firing. The conceptual difference is the degree to which there is any underlying organization of this firing. We argue that in the majority of cases there is no organization, although a substantial minority cannot be explained on a stochastic basis.”

      (7) More explanation around why wakefulness may drive chance aggregation and the clinical relevance of it, as often presurgical epilepsy recordings are being evaluated during sleep.

      We have profoundly rewritten the Discussion regarding the effect of the sleep-wake cycle on FRs incidence:

      “In our rodent data, we were initially surprised to find a higher rate of FRs during wakefulness, which contrasts with previous reports in humans (Bagshaw et al., 2009; Staba et al., 2004). However, previous studies only indicate that physiological vs pathological FRs are more easily distinguished during NREM sleep (von Ellenrieder et al., 2016) and that their incidence varies during sleep (Von Ellenrieder et al., 2017), but in hours-long recordings, no differences in incidence have been reported in the mesial temporal lobe (Dümpelmann et al., 2015). Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence. Last, but not least, another report did not find a state-dependent expression of FRs in the kainate rat model of temporal lobe epilepsy (Ewell et al., 2019), thus indicating that the variability of FRs across sleep and wake is still an open question, at least in rodents. Hence, the main conclusion on the effect of sleep-wake transitions is that these transitions impact the likelihood of stochastic events, more than dictating the direction (increases vs decreases) of change. It also highlights that the specificity of FRs to epileptogenic parenchyma could vary across the sleep-wake cycle, which would be crucial in epileptology (Dimakopoulos et al., 2024; Roehri et al., 2018; Sheybani et al., 2019, 2018; Zijlmans et al., 2012, 2009).”

      Minor Comments:

      (1) Abstract, please include the frequency range of fast ripples explored in this study.

      The abstract has been updated accordingly.

      (2) Abstract, consider including the exact epilepsy model system in rats instead of "a rodent model of hippocampal epilepsy".

      The abstract has been updated accordingly.

      (3) Line 87, while Ylinen uses the term "high frequency oscillations" to refer to ripples up to 200Hz, which are different from the ones discussed here, better to rephrase or use another reference.

      The reference has been changed for Bragin et al. (1999), Epilepsia

    1. eLife Assessment

      This valuable study demonstrates molecular changes associated with age related impairment in oligodendrocyte differentiation and ability to myelinate. The identification of particular genes that are associated with this decline will provide potential future targets for therapeutic interventions. The reviewers felt that the quality of the evidence was convincing while identifying some minor weaknesses that were largely addressed in the review process.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript by Ghosh and colleagues investigates the transcriptional changes within the oligodendrocyte lineage that contribute to age-related declines in oligodendrocyte differentiation and myelination. Combining bulk RNA-Seq on acutely purified oligodendrocyte lineage cells with bioinformatic approaches, the authors identify groups of genes that show different patterns of dynamic regulation during differentiation (which they term "switch" genes, or "switches"). A subset of these switch genes are differentially regulated with age. The authors identify two transcription factors, Bcl11a and Foxm1 that are downregulated during differentiation, have predicted binding site enrichment at other switch genes and are downregulated in aged OPCs. Functionally testing Bcl11a, the authors show that Bcl11a knockdown inhibits the differentiation of young OPCs in culture, whereas overexpression promotes differentiation of aged OPCs. Viral expression of Bcl11a in Sox10 expressing cells accelerates the formation of Plp1+ oligodendrocytes in aged rodents following lysolecithin induced demyelination.

      Strengths:

      The work is clearly presented and addresses an important biological problem. The bioinformatic approaches used in the manuscript are powerful, and the identification of Bcl11a as a modulator of oligodendrocyte differentiation is a novel finding. The combined in vitro and in vivo approaches to assess the function of Bcl11a in oligodendrocyte differentiation are a substantial strength of the work.

      Comment on revised version.

      In the revised version the authors now provide analysis of expression of stage-specific markers for OPCs, preOls and OLs in their isolated cells. It is slightly concerning that the OPC markers show higher expression in the isolated preOLs than in the isolated OPCs, but the authors do provide some discussion on this point in the supplementary text.

    3. Reviewer #2 (Public review):

      Ageing poses a significant challenge to the regenerative capacity of oligodendrocyte precursor cells (OPCs). Myelin abnormalities accumulate with age, while the ability of OPCs to differentiate into myelinating oligodendrocytes progressively declines. This likely contributes to inefficient replacement of damaged myelin and oligodendrocytes, impaired remyelination following injury, and reduced adaptive myelination. Identifying the molecular changes associated with this decline is therefore important for understanding and potentially treating age-related deterioration of CNS white matter.

      This study sought to identify transcriptional regulators involved in oligodendrocyte-lineage progression whose expression is altered in aged OPCs. The authors developed gSWITCH, a computational tool that identifies genes showing defined dynamic expression patterns across ordered biological states. By combining this analysis with comparisons of young and aged OPC transcriptomes and transcription-factor-binding-site enrichment, they identified Bcl11a as a candidate regulator. Bcl11a transcripts are abundant in young OPCs, decline during oligodendrocyte differentiation, and are markedly reduced in aged OPCs.

      A major strength of the study is its combination of computational candidate identification with functional experiments. Bcl11a knockdown substantially impaired the differentiation of young OPCs without measurably affecting their proliferation. Conversely, transient Bcl11a overexpression increased the differentiation of aged OPCs in vitro. Oligodendrocyte-lineage-specific expression of Bcl11a in aged mice also increased the generation of PLP1-positive oligodendrocytes following focal demyelinating injury. Together, these complementary loss- and gain-of-function experiments support the conclusion that Bcl11a expression is functionally important for OPC differentiation and that restoring its expression can improve the differentiation competence of aged OPCs.

      While the transcription-factor-binding-site enrichment analysis predicts a Bcl11a-regulated network, the current study does not establish direct binding or identify the downstream genes responsible for its effect on OPC differentiation. Similarly, the upstream mechanisms responsible for the age-associated reduction in Bcl11a expression were not investigated. Further work may help establish a more complete mechanistic framework explaining how restoration of Bcl11a expression improves OPC differentiation.

      Overall, this study offers valuable insights into the age-related loss of regenerative capacity in the central nervous system and introduces a computational framework that may be broadly useful for investigating dynamic gene regulation in other biological contexts.

      Comments on revised version.

      The authors have addressed my previous comments, and the revised manuscript has been substantially strengthened by the inclusion of additional supporting data and an expanded discussion.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Ghosh and colleagues investigates the transcriptional changes within the oligodendrocyte lineage that contribute to age-related declines in oligodendrocyte differentiation and myelination. Combining bulk RNA-Seq on acutely purified oligodendrocyte lineage cells with bioinformatic approaches, the authors identify groups of genes that show different patterns of dynamic regulation during differentiation (which they term "switch" genes, or "switches"). A subset of these switch genes is differentially regulated with age. The authors identify two transcription factors, Bcl11a and Foxm1, that are downregulated during differentiation, have predicted binding site enrichment at other switch genes, and are downregulated in aged OPCs. Functionally testing Bcl11a, the authors show that Bcl11a knockdown inhibits the differentiation of young OPCs in culture, whereas overexpression promotes the differentiation of aged OPCs. Viral expression of Bcl11a in Sox10-expressing cells accelerates the formation of Plp1+ oligodendrocytes in aged rodents following lysolecithin induced demyelination.

      Strengths:

      The work is clearly presented and addresses an important biological problem. The bioinformatic approaches used in the manuscript are powerful, and the identification of Bcl11a as a modulator of oligodendrocyte differentiation is a novel finding. The combined in vitro and in vivo approaches to assess the function of Bcl11a in oligodendrocyte differentiation are a substantial strength of the work.

      We sincerely thank the reviewer for their positive assessment and for recognising the significance of our study, as well as the bioinformatics approach and tool developed as part of this work.

      Weaknesses:

      Although the PCA plots show distinct and reproducible global gene expression differences between the different isolated cell populations, the authors do not present a figure showing expression levels of typical stage-specific markers (e.g., Pdgfra, Pcdh15, C1ql1 for OPCs, Bcas1, Enpp6, Gpr17 for preOLs, Mobp, Mog, etc. for OLs) or confirm the absence of markers of other lineages (astrocytes, neurons, microglia, etc.). This makes it difficult to evaluate the success of their cell isolation strategy at different ages without reanalyzing the raw data.

      Thank you for this suggestion. We have presented markers expression in a new figure (Supplementary Figure 1) and included a description in the new Supplementary text.

      We observed elevated expression of Hes1 in OPCs as compared to both PreOL and OL, consistent with its role as a Notch effector that maintains the OPC progenitor state and inhibits oligodendrocyte maturation (PMID: 19104146, PMID: 21167918).

      Compared with PreOLs, adult OPCs isolated from 2–3-month-old rats did not show higher RNA expression of canonical OPC markers: Pdgfra, Pcdh15, and C1ql1. However, as expected, OPCs expressed higher levels of these markers than mature OLs.

      One possible explanation is the intrinsic heterogeneity of adult OPC populations. Adult OPCs exist in multiple transcriptional states, including quiescent-like and differentiation-primed states. During early differentiation, OPC markers such as Pdgfra are not immediately extinguished, and PreOLs may transiently retain these transcripts. The PreOL population captured in our study represents intermediate states transitioning from OPC to OL, potentially still carrying residual OPC-associated RNAs from activated OPCs. Therefore, comparing PreOLs with the total heterogeneous OPC pool, which includes quiescent-like OPCs, may give the appearance of higher canonical OPC marker expression in PreOLs.

      Among the PreOL-specific markers, Gpr17 clearly distinguished the PreOL state in our data, showing higher expression compared with both OPCs and OLs. Bcas1 and Enpp6 showed higher expression in PreOLs compared with OPCs. However, when PreOLs were compared with OLs, Bcas1 appeared to be lower in PreOLs, whereas Enpp6 expression remained largely unchanged.

      The OL markers Mobp and Mog showed significantly higher expression in OLs compared with OPCs, whereas their expression was not altered between OPCs and PreOLs. However, the canonical OL maturity marker Mbp showed a progressive and significant increase during differentiation, with expression levels clearly following the expected pattern OL > PreOL > OPC.

      We did not find any difference of astrocytes marker Gfap in those cell types comparison, suggesting similar level of unavoidable contamination which will not affect determination of differential gene expression. Regarding this please also see reviewer #2 major point 1.

      We now included this in the supplementary text:

      “Please see Supplementary Figure 1. We observed elevated expression of Hes1 in OPCs compared with both PreOLs and OLs, consistent with its role as a Notch effector that maintains the OPC progenitor state and inhibits oligodendrocyte maturation (Brosnan et al, 2009; Ogata et al., 2011).

      Compared with PreOLs, adult OPCs isolated from 2–3-month-old rats did not show higher RNA expression of canonical OPC markers: Pdgfra, Pcdh15, and C1ql1. However, as expected, OPCs expressed higher levels of these markers than mature OLs. One possible explanation is the intrinsic heterogeneity of adult OPC populations. Adult OPCs exist in multiple transcriptional states, including quiescent-like and differentiation-primed states. During early differentiation, OPC markers such as Pdgfra may not be immediately extinguished, and PreOLs may transiently retain these transcripts. The PreOL population captured in our study represents intermediate states transitioning from OPCs to OLs, potentially still carrying residual OPC-associated RNAs from activated OPCs. Therefore, comparison of PreOLs with the total heterogeneous OPC pool, which includes quiescent-like OPCs, may give the appearance of higher canonical OPC marker expression in PreOLs.

      Among the PreOL-specific markers, Gpr17 clearly distinguished the PreOL state in our data, showing higher expression compared with both OPCs and OLs. Bcas1 and Enpp6 showed higher expression in PreOLs compared with OPCs. However, when PreOLs were compared with OLs, Bcas1 appeared lower in PreOLs, whereas Enpp6 expression remained largely unchanged.

      The OL markers Mobp and Mog showed significantly higher expression in OLs compared with OPCs, whereas their expression was not altered between OPCs and PreOLs. In contrast, the canonical OL maturity marker Mbp showed a progressive and significant increase during differentiation, with expression levels clearly following the expected pattern: OL > PreOL > OPC.

      We did not detect any difference in the astrocyte marker Gfap across these cell-type comparisons, suggesting a similar level of unavoidable astrocytic contamination across groups. Therefore, such contamination is unlikely to confound the interpretation of differential gene expression among OPCs, PreOLs and OLs.”

      In the main text we have added the following text:

      “The expression patterns of cell-type-specific markers were consistent with their being distinct OPC, Pre-OL, and OL populations (Supplementary Figure 1, see Supplementary text for detailed description).”

      Please note that a detailed discussion of marker expression in the main text will disrupt the flow of the manuscript in manner we feel would detract from its clarity. We have therefore provided this discussion in the Supplementary Text.

      In addition, other publicly available datasets (e.g., the Barres lab bulk RNA-Seq datasets from PMID 25186741 or the Castelo-Branco lab single cell datasets from PMID 27284195) do not show downregulation of Bcl11a during OL differentiation as is described here - this apparent discrepancy is not discussed.

      Thank you for raising this point. We have now included new data as a Supplementary Figure 4. We performed RT-qPCR (reverse transcription followed by qPCR) to quantify Bcl11a expression and found that it was significantly lower in OLs than in OPCs, and significantly lower in aged OPCs than in young OPCs. These data were presented together with stage-specific markers.

      Regarding the comparison with PMID: 25186741: we extracted Bcl11a FPKM values from their dataset (GSE52564) and plotted, as shown in Author response image 1. We found that Bcl11a expression is downregulated during differentiation. However, the dataset contains only two replicates, and the SEM between the two OL replicates is very high, which may have contributed to the apparent lack of clarity. With such high SEM and only two replicates, the statistical power is poor, making robust statistical inference difficult.

      Author response image 1.

      Plotting of FPKM values of Bcl11a (obtained from GSE52564). mean+SEM shown along with individual data points. OPC: Oligodendrocytes progenitor cells, NFO: Newly formed oligodendrocytes, MO: myelinating oligodendrocytes. Dotted red line: linear regression line.

      Regarding comparison with PMID 27284195: we contacted the Castelo-Branco laboratory, and they kindly provided us with the analysis shown below in Author response table 1. This analysis showed that Bcl11a expression is lower in myelinating oligodendrocytes (MOLs) compared with OPCs. The apparent discrepancy observed in the web interface is likely because MOLs are displayed separately by subtype in the online resource. In single-cell datasets, particularly earlier pre-10x datasets with relatively lower cell numbers and sparser transcript detection, visualisations such as violin plots or t-SNE plots can be difficult to interpret when expression is distributed across multiple subclusters. Therefore, directly examining the differential expression statistics, including fold-change and significance values, provides a clearer and more quantitative assessment of the expression change.

      Author response table 1.

      Bcl11a expression difference in MOLs vs OPCs (dataset: GSE75330)

      FC: fold change, p_val_adj: adjusted p-value.

      Therefore, our bulk RNA-seq and RT-qPCR analyses presented in this paper are consistent with the Barres laboratory bulk RNA-seq dataset (PMID: 25186741) and the Castelo-Branco laboratory scRNAseq dataset (PMID: 27284195).

      Reviewer #2 (Public review):

      Aging poses a significant challenge to the regenerative capacity of oligodendrocyte precursor cells (OPCs) to differentiate and myelinate neuronal axons. Myelin abnormalities accumulate with age, and it is likely that the ability of OPCs to differentiate into myelinating oligodendrocytes becomes progressively impaired during aging, leading to inefficient turnover of damaged myelin and oligodendrocytes, as well as reduced adaptive myelination. Understanding the molecular mechanisms underlying the compromised capacity of aged OPCs is therefore critical for addressing age-related white matter decline.

      This study aims to decipher the intrinsic molecular changes that occur in aged OPCs. By profiling differentially expressed transcription factors (TFs) between young and aged OPCs, and by employing a novel bioinformatic tool to identify key TFs that undergo dynamic changes across distinct stages of OPC differentiation, the authors identify Bcl11a as a potential regulator. Bcl11a is highly expressed in young OPCs but markedly reduced in aged cells. Functional experiments further demonstrate that while Bcl11a does not affect OPC proliferation, it significantly promotes the differentiation of aged OPCs. Importantly, this effect is also observed in vivo following demyelinating injury in aged mice.

      While the study provides compelling evidence that BCL11A represents a limiting factor for OPC differentiation during ageing, the downstream targets and molecular mechanisms through which BCL11A exerts its effects are not directly addressed. As such, the work should be interpreted primarily as identifying a key regulatory node rather than a fully defined molecular pathway.

      Overall, this study offers valuable insights into the age-related loss of regenerative capacity in the central nervous system and introduces a computational framework that may be broadly useful for investigating dynamic gene regulation in other biological contexts.

      We are grateful to the reviewer for their supportive comments and for highlighting the broader relevance of our computational framework beyond our specific subfield.

      Major Points:

      (1) MACS mouse anti-A2B5 microbeads are not OPC-specific and may also label astrocyte precursor cells or immature astrocytes. How do the authors justify this caveat? Could some of the claimed "OPCspecific" switch genes in fact be enriched in astrocyte lineage cells?

      We thank the reviewer for raising this important point. While anti-A2B5 is a well-established and widely used antibody for isolating OPCs, we nonetheless agree that no technique can isolate a specific cell type with 100% purity, and this also applies to OPC-specific isolation using a validated anti-A2B5 antibody.

      To check whether astrocyte contamination could be an issue in determining differential expression, and specifically whether the OPC population was affected by astrocyte contamination, we checked the relative expression and statistical significance of the astrocyte marker Gfap. We refer to our new Supplementary Figure 1 and Supplementary text. This suggests that no difference exists in Gfap levels when comparing OPC, PreOL and OL populations. Therefore, we contend that it is unlikely that the differential expression observed in any cell population is actually due to astrocyte contamination, or that the OPC population is selectively contaminated by astrocytes.

      We now included the following in the Supplementary text:

      “We did not detect any difference in the astrocyte marker Gfap across these cell-type comparisons, suggesting a similar level of unavoidable astrocytic contamination across groups. Therefore, such contamination is unlikely to confound the interpretation of differential gene expression among OPCs, PreOLs and OLs.”

      (2) Overall, Figures 1 and 2 are not very informative in terms of biological insight. The authors should provide more detail in the main figures regarding the enriched gene sets associated with each of the Type 1-4 switch categories. For example, summarizing the top Gene Ontology terms for each switch type would greatly enhance interpretability.

      We agree that GO analysis can add further interpretability. We have now prepared a new Supplementary Figure 3A to summarise the significant top GO-term enrichment for switch Types 1– 4, for which gSWITCH-identified patterns are presented in Figure 1C. We also prepared a Supplementary Figure 3B to summarise the top significant GO-term enrichment for the 135 Type 3 switch genes affected in ageing, presented in Figure 2C. Please note that only 8 Type 4 genes overlapped with differentially expressed genes in ageing. Due to this small number, we could not identify any significant GO-term enrichment, and therefore this was not plotted.

      (3) A similar issue applies to Figure 3. The authors should explicitly specify the transcription factors in the main figure, particularly the 27 TFs identified through theENCODE/ReMap2 analysis.

      Thank you for raising this point. We have now prepared a new Supplementary Table 3, where we list 27 TFs and highlight, with light grey shading, the 5 TFs that overlapped with Type 3 switches.

      (4) Have the authors validated Bcl11a expression across different CNS cell types and between young and aged conditions using independent methods such as qPCR, immunofluorescence, or western blotting?

      Thank you for this suggestion. We performed qPCR and presented this data in a new Supplementary Figure 4. We found that Bcl11a expression is lower in OLs than in OPCs (Supplementary Figure 4A). We also observed reduced Bcl11a expression in aged OPCs compared with young OPCs (Supplementary Figure 4B). (see also response to Reviewer 1’s recommendations).

      (5) Regarding OPC aging, an open question is whether the reduced differentiation capacity of aged OPCs is an intrinsic property of the cells themselves or whether it results from prolonged exposure to an aging environment that induces non-cell-autonomous epigenetic or genetic changes, thereby rendering OPCs less efficient at differentiating. It would be helpful if the authors could expand on this point in the Discussion, with reference to relevant previous studies and experimental evidence.

      We thank the reviewer for suggesting this important aspect be discussed. We have now included the following paragraph in the discussion section:

      “The extent to which the reduced differentiation capacity of aged OPCs is intrinsically encoded within the cells themselves or induced by prolonged exposure to an aged tissue environment is an interesting question. Based on our previous work, we favour the view that loss of OPC function is primarily determined extrinsically since various manipulations of the aged environment such as heterochronic parabiosis (Ruckh et al., 2012), fasting and calorie restriction mimetics (Neumann et al., 2019), and niche biomechanics (Segel et al., 2019) can all alter the cell-intrinsic state, reverting aged cells to a ‘youthful state’. Significantly, when aged OPCs are transplanted into the neonatal CNS they proliferate and differentiate as if they were neonatal OPCs (Segel et al., 2019). The reversion of aged OPCs to a functional state by changes in their external environment necessarily operates through changes in cell intrinsic function, suggesting that the same intrinsic mechanisms could be targeted directly to restore declining OPC function—for example through epigenetic regulation of differentiation inhibitors (Shen et al., 2008) or overexpression of transcriptional regulators such as c-Myc (Neumann et al., 2021, Dimas et al., 2025).”

      (6) Do the authors observe a change in the number or density of OPCs between young and aged mice?

      Thank you for asking this important question. In 2002 we reported that there was no difference in the OPCS density between young adult and old adult rats, at least in the deep cerebellar white matter (Sim et al. 2002 - PMID: 11923409). We also refer the reviewer to Figure S1 of another previous study, published in Cell Stem Cell in 2019 (PMID: 31585093). We did not find any difference in OPC number between young and aged brains. Quantification was performed using FACS, where freshly isolated cells were stained with A2B5 (OPC marker), CD11b (microglia marker), and MOG (oligodendrocyte marker). Thus, we do not find any evidence for an age-related decline in OPC densities.

      (7) The in vivo characterization of Bcl11a overexpression using the AAV-based approach appears incomplete. Do aged mice overexpressing Bcl11a in Sox10⁺ cells exhibit reduced age-related myelin degeneration under baseline conditions? In the LPC model, do the authors observe differences in lesion size and/or remyelination efficiency?

      Again, we thank the reviewer for raising these interesting points. To assess whether Bcl11a overexpression in Sox10+ myelinating oligodendrocytes exhibit less age-related myelin degeneration would, we suspect, require long-term experiments. For this to be the case would require a role for Bcl11a in myelin maintenance – and interesting question but one we feel (and hope the reviewer agrees) is beyond the scope of the current study. We do not see any difference in lesion size (and would not expect the expression of elevated levels of Bcl11a to protect against the membrane-solubilising effects of LPC) but do see changes in remyelination efficiency as shown in Figure 6.

      (8) Are the authors presenting gSWITCH for the first time in this manuscript? Given that the gSWITCH framework is novel and central to the study, its conceptual contribution could be emphasized more strongly. A brief comparison with existing trajectory- or pattern-based methods-ideally in the main text around Figure 1-would help readers better appreciate its novelty.

      We thank the reviewer for this important suggestion. Yes, gSWITCH is presented for the first time in this manuscript as a new computational framework and web application. We agree that its conceptual contribution should be made clearer in the main text itself, although we explained its concept in detail in ‘Materials and Methods’ and in the supplementary Figure 2 (which was Supplementary Figure 1 in first version of this manuscript).

      We now included the following paragraph in the manuscript:

      “Existing computational tools such as Monocle (Trapnell et al., 2014), tradeSeq (Van den Berge et al., 2020) and maSigPro (Nueda et al., 2014) are highly valuable for identifying genes with dynamic expression changes across pseudotime or time-course data. gSWITCH addresses a different question. It does not aim to infer trajectories. It works with a user-defined ordered series of biological states or time points and asks a more specific question — does this gene show a statistically supported "switchlike" change in expression as cells move through these states, and if so, what shape does that change take? It combines GLM-based statistical testing with criteria that capture where a gene reaches its highest or lowest expression and whether its expression changes steadily in one direction across the ordered series. To our knowledge, no existing tool combines significance testing with this type of explicit, shape-based classification into discrete, interpretable switch categories. gSWITCH sorts genes into four biologically meaningful patterns, rather than producing only a ranked list of significant genes based on pairwise comparisons between multiple conditions or states. gSWITCH also flags which of these switch genes are transcription factors, making it easier to prioritise candidates for follow-up experiments.

      This biologist-friendly tool is freely available as a web application requiring no programming, works with experimental designs containing three or more stages or time points with at least two replicates per stage (no upper limit on either), and can be applied to bulk RNA-seq or to single-cell RNA-seq data aggregated as pseudobulk.”

      (9) The evolutionary analysis also appears somewhat disconnected from the rest of the study. Could the authors leverage available public datasets to test whether a similar Bcl11a expression trajectory is observed in human oligodendrocyte lineage cells?

      We thank reviewer for mentioning this. We would like to clarify that the evolutionary analysis was included to examine whether Bcl11a sequences across vertebrates, including humans, show evidence of selective constraint, meaning that the sequence has been preserved during evolution because changes in it are likely to be disadvantageous. This analysis was therefore intended to provide broader evolutionary support for the functional importance of Bcl11a, rather than to stand as a separate or disconnected component of the study.

      For this analysis, we included Bcl11a DNA and protein sequences from 23 vertebrate species, including humans. We refer the reviewer to the Methods section of this paper, under “dN/dS analysis”, for further details. To provide further clarity regarding the different species used in this study, we have now prepared a new Supplementary Table 4, listing the 23 species together with their DNA and protein sequence accession numbers for Bcl11a.

      We also added this sentence in the main text:

      “We included twenty-three vertebrate species, including humans (Supplementary Table 4).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Given how central the isolated cells are to the subsequent analysis, the manuscript would be strengthened by a figure showing expression of stage and lineage-specific markers.

      Ideally, the authors would provide some sort of orthogonal experimental approach to confirm downregulation of Bcl11a during oligodendrocyte differentiation and loss with age (e.g., IF or RNAScope in conjunction with stage-specific markers in tissue, or western blot in culture).

      Thank you again. We have performed these. Please see the Reviewer #1 comment (above).

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1A: It should be 'anti-O4' instead of 'anti-04'.

      This is now corrected. Thank you.

      (2) Figure 1C: The authors should specify what the connecting lines indicate (e.g., gene sets or gene modules).

      Each coloured line represents one gene and connects its log2 fold-change values across the three oligodendrocyte lineage states: OPC, PreOL and OL. The connecting lines are used to visualize gene-wise patterns of expression change across these cell states. For example, in Type 1, each line shows a pattern in which gene expression increases progressively from OPC to PreOL to OL, with the highest expression change observed in OLs: OL > PreOL > OPC.

      We now included the following line in the figure legend:

      “Each coloured line represents one gene and connects its log<sub>2</sub> fold-change values across the three oligodendrocyte lineage states: OPC, PreOL and OL. The connecting lines are used to visualise gene-wise patterns of expression change across these cell states.”

      (3) Figure 2C: The authors should specify "DF genes" in the figure legend.

      Thank you for pointing this out. This was a typo: it was written as DF, but it should be DE (differentially expressed) genes. We have now corrected this in the figure and spelled out the abbreviation in the legend. Also, DE gene list is accessible through GEO accession: GSE303317. This also mentioned in the figure legend as:

      “DE: Differentially expressed. DE gene list is accessible through GEO accession: GSE303317.”

      (4) Figure 4C & Figure 5B: the title for the y-axis of the bar graph is confusing. The authors should specify what "#" indicates. Does it represent the counts? What are the thresholding criteria to judge whether an Olig2 cell is MBP-positive or not? It's unclear what the unit is here for the 0-100 scale.

      We apologise for the confusion. We used ‘#’, which is a common notation in mathematical and quantitative contexts, to denote counts, so you are correct. We now mentioned in the legend: “The symbol “#” indicates cell count.”

      We counted the number of MBP+OLIG2+ cells, divided this by the total number of OLIG2+ cells, and expressed the value as a percentage. For greater clarity, instead of writing #MBP+/#OLIG2+, we have now written #MBP+OLIG2+/#OLIG2+.

      Regarding the 0–100 scale, the unit of the Y-axis is percentage, as stated in both figure legends.

      The criterion for classifying an OLIG2+ cell as MBP+ was morphological: an OLIG2+ nucleus, shown in white, had to be surrounded by MBP+ staining, shown in red. Cells meeting this criterion were counted as MBP+OLIG2+ cells. Manual counting was performed blinded to sample identity.

      (5) Figure 6B: To discriminate from IF staining, the authors should use italic'Plp1' to indicate the RNA in situ results.

      Thank you for pointing this out; we have now corrected it.

    1. eLife Assessment

      This important work sets out to identify the neural substrates of associative fear responses in adult zebrafish. Through a compelling and innovative paradigm and analysis, the authors identify brain regions associated with individual differences in fear memory expression. While most findings are well supported, there is limited evidence that the four behavioral clusters represent discrete forms of associative fear-memory expression and the related interpretations would benefit from additional analysis or more cautious framing. Nonetheless, this study showcases the strength of zebrafish for systems-level neuroscience and will be of broad interest to the neuroscience community.

    2. Reviewer #1 (Public review):

      Summary:

      This work provides a comprehensive analysis of how adult zebrafish show fear responses to conspecific alarm substances (CAS) and retain their associative memory. It shows that freezing is a more reliable measure of fear response and memory compared to evasive swimming, and that the reactivity and the type of responses depend on the zebrafish strain. It further suggests neuronal substrates of different fear responses based on c-Fos mapping.

      Strengths:

      The behavioral part is the most comprehensive and detailed yet in the zebrafish field, providing strong support for the authors' claim. The flow from Figure 1 to Figure 4 is very smooth. They provide extremely detailed, yet complementary and necessary, analyses of how different categories of behavior emerge over time during the CAS exposure and memory retrieval. I'm convinced that neuro researchers who study fear/stress responses will always refer to this paper to plan and interpret their future experiments.

      Comments on revised version:

      The authors successfully addressed my comments, including the addition of Figure S6-2, which gives us some intuition into the relationships between c-Fos levels in individual areas and the behavioral outputs.

    3. Reviewer #2 (Public review):

      In this study, Fontana et al. develop a paradigm for associative conditioning by pairing exposure to alarm substance with a novel tank. Exposure to conspecific alarm substance (CAS) in the novel tank triggers freezing and what they characterize as evasive swimming behaviour, which are subsequently seen in a re-exposure to the novel tank without the CAS present. Importantly, these states are identified via automated processes including postural tracking and a random forest classification process, which could be very useful tools for subsequent studies.

      In their experiments they focus on the differences in behaviour among strains of zebrafish (both males and females), and among individual zebrafish. For males and females of different strains they find some differences, though the clearest message seems to be that the most robust measure of the behaviour in response to both the CAS and in the memory trials is the freezing behaviour, while evasive behaviour is more variable and not always seen. This may relate to their observation of significant "evasiveness" in vehicle control experiments (discussed further below).

      Moving on to individual variation from within this multi-strain male/female dataset, they first examine transition matrices between states, and find this is not dramatically altered by stimulus exposure. They then use clustering to identify 4 different "classes" of zebrafish that differ in their expression (or not) of two types of behaviour: freezing and/or evasive behaviour. They show that over the three exposure epochs of the experiment this classification is somewhat stable in an individual fish, though many fish change their behaviour -- e.g. evading + freezing -> only freezing.

      In the final set of experiments they move beyond behavioural analyses and perform whole-brain cFos mapping of these individual zebrafish, and perform analyses aimed at identifying correlations between individual behavioural expression and the number of cFos positive cells in different brain regions. Using partial least squares analysis they find areas associated with two types of behavioural contrasts, which differ in their weighting of different behavioural expression during the Memory trials. Covariation and network structure analysis within different classes of fish also find some differences in covariation among brain areas, providing hypotheses as to underlying network effects that may govern the expression of freezing and/or evasive behavior in the memory trial phases.

      Overall, I find this to be an interesting study that employs state of the art methods of behavioural analyses and whole-brain cFos analyses. The revision has clarified the take-home message considerably: the abstract is now more careful about which behavioural groups are memory-associated, and the causal language in the conclusions has been appropriately softened. Two of my three original main concerns have been addressed. The first is not and having looked at the data again I can now be more specific about what concerns me.

      Comments on revised version.

      (1) My first concern related to the claim that fear memory behaviour falls into four distinct groups, and specifically to the role of evasiveness in defining them. The authors give three reasons for retaining it, but I remain unconvinced.

      The first is that variable evasion in response to alarm substance is a long-standing observation (von Frisch; Suboski et al.), and that dissecting this individual variation is the purpose of the paper. I agree with the motivation, and it is a good reason to measure evasion. But it does not establish that evasion on memory day reflects fear memory, and memory day is the only day used for the clustering and neural activity mapping. The manuscript's own results point the other way: relative to pre-exposure, no strain or sex increased evasion on memory day, and relative to vehicle only female TUs did. The temporal profiles show evasion on memory day to be largely similar between vehicle and CAS-treated fish. Historical observations of variable evasion during CAS exposure do not carry over to the memory phase.

      The second is that the clustering itself reveals two kinds of freezing fish - one freezing between bouts of normal swimming, the other between bouts of evasion - demonstrating that a subset of fish increase evasion. In absolute terms, this does not match the data. In Figure 4B, evading freezers are below the population mean for absolute evasion, as are freezers. The text describes evading freezers as "high in freezing and evasive behaviors," and I do not think Figure 4B supports this.

      What actually separates the two freezing groups is the third measure, evasion as a percentage of active time. And this is where I have difficulty, because that measure is not an independent behavioural readout. The classifier assigns every window to normal, evasive or freezing, and active time is simply non-freezing time, so evasion-as-percent-of-active is fully determined once the other two are known.

      This matters for the clustering specifically. Distance-based methods weight each input dimension equally, so a variable that carries no information beyond the other two nonetheless contributes a full third of the distance between any two fish - and it contributes it in a way that counts freezing twice, once directly and once through the denominator of the derived measure. The space is nonetheless described as three-dimensional throughout, including in the Methods and the Figure 4 legend, when there are only two independent behaviours in it.

      The consequences fall hardest on exactly the animals at issue. Both freezing groups sit at 65-70% freezing, so there is very little active time to divide by, and small absolute differences in evasion - together with any noise in estimating them from a couple of minutes of non-frozen behaviour - are inflated into large differences on the rescaled measure. In terms of what the fish actually did, the two groups differ by a few percent of trial time. That is the boundary on which much of the rest of the paper rests.

      I recognise that evasion as a proportion of active time is in some respects the more biologically meaningful quantity, and the authors are right that a fish freezing 70% of the time has limited opportunity to do anything else. But that is an argument for reporting it as a descriptive measure, not for entering it into the clustering alongside the two variables from which it is computed.

      This impression is reinforced by Figure 4A itself. While the freezer group occupies a reasonably distinct region, the non-reactive, evader and evading freezer groups appear as a single continuous distribution with cluster boundaries drawn through it rather than around visible gaps. I appreciate that UMAP is a projection and that visual separation is not required for genuine structure, but this is the figure by which most readers will judge whether four discrete types exist, and it does not obviously support that reading - particularly given that the embedding is built from the same variables, including the rescaled measure, that most favour the separation.

      I would suggest that the authors re-run the clustering using only the two directly measured behaviours, percent freezing and percent evasion of total time, and report whether four groups still emerge and, in particular, whether the evading freezer / freezer split survives.

      The third is that the two groups have distinct functional networks despite equally high freezing, so the behavioural difference is real and is manifesting in the brain. This is the strongest of the three arguments, and I accept part of it: something about how a frozen fish spends its remaining active time does appear to be neurally meaningful, which is interesting in its own right. But it does not establish that these are two distinct types, nor that the difference has anything to do with the conditioning. Fish taken from either side of a cut through a continuous distribution will differ neurally if that continuum tracks brain state, so the network result is equally compatible with graded variation. More importantly, Figure 5A shows that a substantial proportion of fish are classified as evaders in the vehicle condition and at pre-exposure, before any CAS has been given. This suggests a pre-existing individual tendency toward evasive behaviour that is independent of the alarm substance, and one would expect such a tendency to persist into the memory trial. If so, the distinction the network analysis is drawing between freezers and evading freezers may simply reflect that baseline trait, and its neural correlates would be correlates of the trait rather than of fear memory. I am therefore not convinced that this distinction is related to CAS or to memory.

      (2) This concern is fully resolved. I had misread the CAS preparation: it was pooled from eight donors spanning all four strains and both sexes, so every fish received identical material and the strain and sex differences cannot be attributed to donor variability. The clarification now added to the Results will prevent other readers making the same error. The addition of FDR correction to the Figure 2 comparisons also addresses my related concern about multiple testing.

      (3) Somewhat resolved. The conclusion no longer states that behavioural variation is "driven by" activity in particular regions, and the added caveat that neural activity was not directly manipulated sets the right expectation for a mapping study. The scatterplots in Figure S6-2 are a useful addition and give a much better intuition for what the PLS contrasts represent. My remaining reservation is the one above: a great deal of the neural story rests on the evading freezer / freezer contrast, and I am not persuaded that this contrast marks a boundary relevant to fear memory.

    4. Reviewer #3 (Public review):

      This revised manuscript by Fontana et al. aims to study how animals respond to fearful stimuli, with a specific focus on brain regions involved in predicting animals that passively freeze or those that actively evade the threat. I continue to be enthusiastic about the study. The study addresses an important question regarding individual variation in fear-related behavior and links these behavioral phenotypes to whole-brain activity patterns in adult zebrafish. The combination of a contextual fear conditioning paradigm, strain/sex comparisons, behavioral clustering, and AZBA-based c-Fos mapping makes this a valuable contribution to the field, not just in answering the question posed by the authors, but also in formulating a framework for using adult zebrafish for whole brain analysis of complex behaviors. Overall, I find the authors have responded to my concerns:

      (1) I still think that separating memory acquisition and consolidation is an interesting question, and further use of the framework will need to eventually solve that; however, I also appreciate that this may be beyond the scope of the current study, and I appreciate the authors acknowledging this in the manuscript.

      (2) Regarding Figure 3, I also agree that this is difficult to present differently, and I appreciate the authors adding text to the body to clarify things. My one request is that the sentence (lines 214-215) that reads: "This increase in evasion in the vehicle group likely represents a response to the water disturbance that occurs when solution is added to the tank." Be changed to: "This increase in evasion in the vehicle group may represent a response to the water disturbance that occurs when solution is added to the tank." While it is entirely possible, there are no concrete data to support that this is "likely."

      (3) I appreciate the clarification regarding the PLS-derived contrasts in Figure 6A and in the body.

      Overall, this is a really interesting paper that will have a wide-ranging impact. All of my concerns have been addressed.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewer’s for their thoughtful comments that have significantly strengthened the paper. Below, we have outlined our responses to both the public reviews and recommendations.

      In addition to the alterations to the manuscript based on the reviews, during our review of the data analysis we uncovered some small errors that we have now corrected. In looking back over the image registration, we identified three animals whose olfactory bulbs did not register properly and one with poor cell counting in the telencephalon. To account for these issues, we imputed the missing data using an iterative soft-threshold singular value decomposition (described on lines 779-783 of the updated manuscript). This update had little impact on the results. We also identified a small error in how we determined ‘unique’ and ‘overlapping’ edges in the network analysis (Figure 8). In the previous analysis we had incorrectly noted that all ‘unique’ edges did not have an overlapping confidence interval with the two other networks (i.e., the networks for evading freezers, freezers, and non-reactive). Instead, the ‘unique’ edges in the prior version of the manuscript did not have an overlap with at least one other network. We have now updated the analysis so the reader can distinguish between edges that are truly ‘unique’ versus those with ‘1 overlapping confidence interval’ or ‘2 overlapping confidence intervals’ with other networks. As before, this update and change to the analysis does not materially affect the results or conclusions.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      The neural analysis part is very comprehensive. Figure 5 and Figure 6 are independent but complement each other very well. They together support that the cerebellar system is the key brain component for a freezing response. Their extreme focus on high-level analyses, however, came at the expense of biological intuitions. I suggest adding some figure panels and result/discussion paragraphs to help with that aspect.

      Thank you for the suggestion. We have made extensive edits to the manuscript to include additional discussion and biological intuition. Specifically:

      We added a supplemental figure (Figure S6-2) that has scatterplots showing how cfos levels vary with the different behavioral contrasts. Although the PLS analysis is multivariate, this univariate analysis should help give readers a better intuition of how the behavior relates to brain function.

      We have also rewritten the results sections for both the PLS analysis (lines 303-361) and network analysis (lines 396-437) to incorporate more of a discussion about the biological context of different regions identified. Thank you for this suggestion, we feel that this significantly strengthens the biological interpretation of the data for the reader.

      Reviewer #2 (Public review):

      (1) My first concern relates to the claim in the abstract that "We found that fear memory behavior fell into four distinct groups: non-reactive, evaders, evading freezers, and freezers".

      In my opinion, the "freezing" aspect is well supported as being both triggered by the CAS and for memory effect upon re-exposure to the tank, but I am less convinced about the "evasive" behaviour. In Figure 2, it appears that "evasiveness" is generally not increased in both the Exposure or Memory phases for many groups, and in Figure 5, it appears that "evasiveness" is expressed by nearly 50% of the fish in the pre-exposure condition before CAS addition and in all phases in the vehicle condition. Therefore, it appears that most of the expression of this behaviour is independent of any memorybased effect.

      We thank the reviewer for this suggestion and we agree that this line in the abstract was unintentionally misleading. We have now altered this line in the abstract (lines 34-36) to read:

      “We also found that that behavior fell into four distinct groups: non-reactive, evaders, evading freezers, and freezers with the evading freezer and freezer groups most clearly associated with memory formation.”

      On the larger point of the inclusion of evasion as part of the fear response, we believe this is warranted for the following reasons: (1) evasive behavior has long been acknowledged as a highly variable aspect of how fish respond to alarm substance where some fish exhibit evasion and others do not. This observation goes back to the original work from Karl von Frisch in minnows (von Frisch, 1938), and others in zebrafish (e.g., Suboski et al, 1990). One goal of our paper (and the work from the lab in general) is to try dissecting out this individual variation that can get lost when only considering population averages. (2) The unsupervised clustering also suggests that there are two distinct types of freezing clusters (Figure 4B) where some fish freeze intermittently with normal swimming and others freeze intermittently with evasive behavior. This suggests that evasion is increased in response to CAS, but only in a subset of fish. (3) The brain networks from the evading freezer and freezer groups are distinct (Figure 8A) despite having equally high levels of freezing behavior (Figures 4B and C). This means the difference we’re able to distinguish behaviorally is also manifesting in the brain, suggesting that it is not anomalous. Thus, while we agree that freezing is definitely the strongest and clearest behavioral response to CAS, we believe the analysis of this large dataset supports the interpretation that, in a subset of fish, increased evasive behavior in response to CAS is also a part of the response.

      (2) My second concern relates to the claim in the abstract that "background strain and sex influenced how fish respond to CAS, with males more likely to increase evasive behaviors than females and the TU strain more likely to be non-reactive."

      My understanding, based on the introduction and on the methods, is that it is likely important that the CAS be prepared from conspecifics of the same strain and sex, and for this reason, they prepared different CAS specific for each strain and each sex. Therefore, the "CAS" that is applied is necessarily different for each condition, and I am concerned about if the differences observed could relate more to variation in the quality, purity, concentration, etc. of the specific CAS samples for different groups, rather than their reactivity to the substance or their ability to form memories based on such experiences.

      The CAS was prepared by mixing extracts from all four strains and both sexes (so 8 fish per batch). Thus, all the fish were exposed to the same CAS mix derived from the same donors. This is described in the methods (lines 626-629). However, to ensure that this is clear to readers, we’ve now included a line indicating this in the results section (lines 123-124).

      (3) My third concern relates to the interpretation of the cFos data.

      As I mentioned above, I feel as though the behavioural analysis is perhaps more complex than is warranted via the inclusion of evasiveness, and I wonder if the conclusions from the experiments would be simpler if analyzed only from the perspective of freezing.

      We agree that the freezing response is driving the majority of the neural cfos response that we are seeing (e.g., Figure 6A-C). However, we feel that the network analysis (Figure 8) justifies the distinction between freezers and evading freezers. This is because the brain networks for these two groups (freezers and evading freezers) are quite distinct, even though these groups both have the same levels of freezing behavior (Figure 4). This stark difference in patterns of neural activity suggests the brain of a freezer and an evading freezer are engaging with the world in two distinct ways that is worth noting. We’ve updated the abstract to make this point clearer (abstract: lines 39-48) and discuss the biological interpretations of patterns of brain activity unique to evasion or evading freezers in more depth (lines 303-361; lines 396-437).

      Reviewer #3 (Public review):

      (1) The three-day contextual fear paradigm, as implemented - one CAS pairing on day 2 followed by a single recall test on day 3 - inevitably conflates acquisition and long-term memory, making it impossible to know whether strains like TU truly recall the association poorly or simply learn it more slowly. For example, given that TU fish extinguish fear faster than AB or TL strains in extended protocols, they may simply require additional or repeated CAS pairings to achieve the same asymptotic performance. To disentangle learning kinetics from recall strength, the assay could be revised to include multiple acquisition trials (e.g., conditioning on two or more consecutive days) with an immediate post-conditioning probe to assess acquisition independent of consolidation, and continuous measurement of freezing and evasive behaviors across each trial to fit learning curves for each strain. Such refinements - even if on a subset of the strains - would reveal whether "non-reactive" phenotypes reflect genuine recall deficits or merely delayed acquisition.

      We thank the reviewer for this thoughtful comment. We agree that it is difficult to disentangle acquisition from consolidation. Indeed, the TU fish do appear to have lower levels of freezing in response to the CAS (Figure 2A), supporting the idea that reduced performance at memory day could be due to some sort of deficit at acquisition. However, pursuing a detailed examination of strain dependent differences in fear memory acquisition versus consolidation is beyond the scope of the current paper where we primarily focus on individual differences in behavior. Nonetheless, we have included this important point in the discussion (lines 470-471).

      (2) My second major question is with respect to Figure 3 panel B. This is a complex figure, and I can understand the gist of what the authors are attempting to show, but it is difficult to understand as it is. Can this be represented in a way that is clearer and explained a bit more easily?

      We agree that this figure is one of the more complex in the paper. However, we’ve struggled to come up with a better way to present it. We have improved the presentation based on other reviewer comments by making the vehicle and CAS groups more easily distinguishable by using open versus closed circles. We’ve also included additional interpretations of the data in the results, which we hope will help guide readers through this figure better (lines 208-223).

      (3) The brain mapping is by far one of the most interesting aspects of this study, and the methods that the group used are interesting. The brain mapping, however, relies on generating "contrasting" groups (Figure 6A), and I was not clear as to how these two groups were formed. Could the authors elaborate a bit?

      These contrasting groups (contrast 1, contrast 2) arise analytically from the partial least squares (PLS) analysis; they are not defined by the experimenter. In brief, PLS is a multivariate technique that identifies latent variables that capture axes of maximal covariation between two datasets: behavior and brain activity. As an analogy to a more widely known technique, principal components analysis (PCA) uncovers axes of maximal variance within a single dataset. PLS, in contrast, simultaneously analyzes the covariation in two datasets. The contrast groups in Figure 6A represent the behavioral weights of the latent variables that capture the most covariance, which illustrates how the four behaviors load onto these top two contrasts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) The c-Fos analysis in Figure 5 is very comprehensive and convincing, but lacks intuitive presentations. In my understanding, the increase in c-Fos expression in red areas means increased freezing behavior for Contrast 1 for the PLS analysis? Do you have representative c-Fos expression images between different groups of fish?

      We decided not to include a representative cfos image because the data is derived from a large number of fish (N=87) and thus it can easily be cherry-picked to choose images that match the narrative. Instead, to more accurately capture the breadth of the data while providing a more intuitive presentation, we have included an additional supplemental figure that includes scatterplots of scaled cfos data against behavioral scores for each of the two contrasts (S6-2). We believe this more fully and accurately captures the relationship between behavior and brain activity. We included six different example brain regions and scatterplots for cfos activity against behavioral scores for contrasts 1 and 2, demonstrating a range of relationships. However, we should note that PLS is a multivariate technique, and so this univariate analysis does not fully capture the subtleties of the PLS analysis. Nonetheless, we think this will help give a more intuitive interpretation of the data to readers. We have also referenced this additional data in the manuscript (lines 307-309). We thank the reviewer for this excellent suggestion that improves the ability of readers to understand the paper.

      (2) Also related to Figure 5, the result section only describes the PLS statistics and does not try to describe the biological interpretation. Do the authors think the c-Fos expression directly represents lowlevel behavior, such as swimming, or a high-level behavioral state or learning? Maybe different areas mediate different aspects?

      For example, the medullary locomotor areas, which are usually highly correlated with swimming in terms of neural activity, seem to have higher c-Fos expression in freezing fish. I'm not saying this shouldn't be the case. c-Fos expression in this area was not elevated in larval fish during OMR in Shainer et al., 2023, indicating that it doesn't linearly reflect neural activity. But discussing a bit of intuition on the connection between c-Fos expression and biological process, rather than just saying "the cerebellum could regulate emotional states", would help us guide through this highly complex analysis.

      We have now added more interpretation of the data in both the PLS and network analysis sections (lines 303-361 and lines 396-437). Again, thank you for this excellent suggestion. This helps make the biological interpretation of the data clearer.

      Minor points:

      (1) Figure 2B titles: please write "memory" on the right side.

      We considered writing ‘memory’ on the right-hand side, but we thought this may add confusion because it would not apply to both graphs in the row. The left-hand graphs are the responses during ‘exposure’ and the right-hand graphs are the responses during the ‘memory’ phase. This is indicated by the titles above the left and right-hand sets of graphs.

      (2) Figure 2C: needs legend lines.

      We have now moved the legend lines from the top of the graphs to below the graph to make them more visible to readers.

      (3) Line 187: "aggregated" data.

      This has now been changed to ‘aggregated’ (now line 194).

      (4) Line 371: I'm not sure what "Beyond" means.

      We have now significantly changed this part of the paper and we no longer use the word ‘beyond’ here.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding point (1) in the Public Review:

      I would encourage the authors to consider whether this study might be better focused exclusively on the freezing behaviour, which does appear to be reliably expressed during CAS exposure and in the memory phases, and would significantly simplify the subsequent analyses of neural activity, and perhaps may lead to a more coherent conclusion.

      As noted in our response to the public review, we appreciate this suggestion, but we have decided to keep the inclusion of the evasive behavior. This is because (1) evasive behavior has long been acknowledged as a highly variable aspect of how fish respond to alarm substance where some fish exhibit evasion and others do not. This observation goes back to the original work from Karl von Frisch in minnows (von Frisch, 1938), and others in zebrafish (e.g., Suboski et al, 1990). One goal of our paper (and the work from the lab in general) is to try dissecting out this individual variation that can get lost when only considering population averages. (2) The unsupervised clustering also suggests that there are two distinct types of freezing clusters (Figure 4B) where some fish freeze intermittently with normal swimming and others freeze intermittently with evasive behavior. This suggests that evasion is increased in response to CAS, but only in a subset of fish. (3) The brain networks from the evading freezer and freezer groups are very distinct (Figure 8A) despite having equally high levels of freezing behavior (Figures 4B and C). This means the difference we’re able to distinguish behaviorally is also manifesting in the brain, suggesting that it is not anomalous. Thus, while we agree that freezing is definitely the strongest and clearest behavioral response to CAS, we believe the analysis of this large dataset supports the interpretation that, in a subset of fish, increased evasive behavior in response to CAS is also a part of the response.

      A more minor concern related to the analyses in Figure 2: in the figure legend, it is stated that "*-P < 0.05 compared to vehicle treated fish via t-tests". How are the authors dealing with the multiple comparisons problem? Would something like an ANOVA not be more appropriate?

      Thank you for bringing this point up. We did not initially correct for multiple comparisons because we considered each of these experiments across sex and strain separate since we did not compare across strains. However, the way we’ve grouped the data together in figure 2 makes it appear as if they are one large experiment. To alleviate any concern about multiple testing, we have now corrected for multiple comparisons using the false discover rate (FDR) correction. The statistics in the figure and captions have now been updated.

      (2) Regarding point (2) in the Public Review:

      If the authors agree with my concern regarding potential variability in the CAS samples, I would suggest either testing for differences among strains using the same batch of CAS, or including and explaining this caveat in the text.

      As noted in our response to the public review, the CAS was the same for all the fish. Each batch was derived from 8 donor fish, one fish from each strain and sex (described in lines 123-124 of the results and lines 626-629 of the methods).

      (3) Regarding point (3) in the Public Review:

      I feel like the standard in the field for such conclusions would be after

      (a) Direct analyses of the activity states in these areas. I was surprised not to see a direct analysis of the cFos stainings in the cerebellum relative to freezing behaviour, for example, ideally in a different animal cohort.

      The PLS analysis does relate activity in the cerebellum (and other brain regions) to specific behaviors via the the behavioral contrasts (Figure 6A). We believe this approach (instead of dividing fish into ‘high and low freezers’) is a more powerful way to leverage the data from all the animals tested (87 fish). However, we appreciate that the interpretation of the PLS analysis is not as intuitive as seeing scatterplots or bar charts comparing neural activity. For this reason (and in response to a comment from reviewer 1), we have included as a supplementary figure (Figure S6-2) scatterplots showing how standardized c-fos activity varies with the behavioral scores from the contrasts identified from the PLS analysis. Given that contrast 1 weights heavily in the positive direction on freezing, these figures can essentially be read as looking at cfos activity as a function of freezing levels. What can clearly be seen is that for regions of the cerebelleum (E.g., the LCa and CC) there is a clear positive relationship between cfos activity and the behavior scores for contrast 1.

      (b) Some kind of manipulation of the brain area resulting in the relevant behavioural modification.

      We completely agree with the reviewer. However, at the moment, we do not have the tools to do this in adult zebrafish. It is something we’re actively working on.

      Of course, I appreciate that such experiments might not be possible or feasible, and in which case I would suggest adjusting the claims accordingly and highlighting the caveats to their interpretations.

      We have incorporated the caveat that we have not directly altered neural activity into the discussion (lines 542-543) and adjusted how we discuss our findings in the abstract (lines 39-41) to more accurately represent the type of evidence we provide. Hopefully we’ll be able to do so in the near future!

      MINOR CONCERNS:

      (1) In Figure 3, how is the end of a behavioural epoch defined? I am surprised to see that you consider transitions between the same behavioural state. How does erratic swimming -> erratic swimming differ from a longer single epoch of erratic swimming? In general, I find this analysis confusing, and I am not sure if it adds significantly to the message of the paper.

      Thank you for this question as it prompted us to realize we were missing this in our methods section. We have now updated the methods to include how we calculated the behavioral transitions (lines 644-650). In short, we used a 750 ms behavioral epoch time that corresponds to the size of the sliding window we used for the random forest model.

      We have also updated the description of this analysis in the results to indicate the main finding from it (lines 207-223). In brief, the main finding is that exposure to CAS results in longer bouts of evasive behavior without increasing its frequency. Whereas CAS induced freezing arises from both longer bouts and likelihood of occuring. While we agree that this is a relatively minor finding in the paper, one of our goals is to provide as comprehensive analysis of fear behavior as possible to help guide future researchers interested in using fish for understanding different aspects of fear-related behaviors.

      (2) In the PLS analyses, two measures of evasion are used: evasion time, and evasion as a percent of active behavior. I don't understand the justification for both of these being used rather than one. Again, my overall recommendation is to reduce the focus on the analysis of evasion behaviour, but if you do not choose to do this, I think the rationale of how both measures are used and why needs explanation.

      We chose to incorporate two different measures of evasion throughout the study because the high levels of freezing in some animals results in little opportunity to express other behaviors (like evasion). Thus, to better capture what fish may be doing in the absence of freezing (i.e., when they are active) we also calculate the amount of active time spent performing evasive behaviors (instead of normal swimming). We have now included an explanation for this earlier in the results section when we first use this metric (lines 149-152).

      (3) In the methods, I don't understand this: "Animals that were assigned the wrong sex were removed from data analysis, as well as its paired fish (< 2%)".

      We determine the sex of fish when we set them up for dual housing. However, we occasionally make errors in sex determination. To ensure we properly sexed the fish, at the end of experiments, we euthanize the fish and check for the presence of eggs. If we incorrectly assigned the sex to a fish, they are removed from the experiment alongside the other fish they were dual housed with. This is because we want to ensure all fish are housed in the same way (i.e., a male fish with a female fish).

      (4) How was this determined differently from the first time, resulting in exclusion?

      After experiments, fish were euthanized and we checked for the presence of eggs (line 599-601). We’ve now added a line in this other part of the methods referring back to where we describe this (lines 676678).

      Reviewer #3 (Recommendations for the authors):

      Here are some minor concerns and errors found in the manuscript:

      (1) For Figure 2B and Figure 3B, can the group make the lines solid and dotted? The circle or triangle designation is difficult to see, and since the crux of the figure depends on comparing Veh and CAS, it would be easier to see if the lines were altered.

      Thank you for this suggestion. Instead of making the lines solid and dotted, we decided to make both the CAS and vehicle group circles and then have open and closed circles. We believe this solves the issue of being able to distinguish these groups and makes the data more readable.

      (2) Figure 2C: It appears that the line colors in the legend are missing.

      We have moved the line colors below the graphs to make them more obvious.

      (3) Figure 8A: Same thing here - could the text be enlarged? It's really difficult to make out each node, and when I zoom the text becomes pixelated. This is an important figure and one that will likely be referenced, and making it clear would be helpful.

      This one is difficult. We have made the network images as large as would fit on a page. We have now uploaded vectorized versions of the images so that they do not become pixelated when zooming in. As part of our supplemental materials we also include a cystoscope file that can be explored in greater depth as well.

      (4) The paper is really well written: I found a few typos, though:

      (a) Line 529: "Institutional Cara and Use Committee" should be "Institutional Animal Care and Use Committee" (Change cara to care and add animal).

      (b) Line 274: "hybdridization" should read hybridization.

      Thank you for catching these typos. They have now been fixed.

    1. eLife Assessment

      This study provides useful information for the Drosophila ageing community by characterising the auxin-based gene expression system (AGES) and identifying caveats associated with its use in adult flies. The authors provide solid evidence for sex-, age-, tissue- and dose-dependent variability in transgene induction, as well as effects of auxin feeding and AGES activation on stress resistance, metabolism, and lifespan. While the study would benefit from a broader characterisation of some of these limitations, the findings provide a helpful benchmark for researchers using AGES in ageing studies.

    2. Reviewer #1 (Public review):

      Summary:

      The authors set out to evaluate whether AGES, a recently developed auxin/TIR1-based conditional GAL4 expression system, is a suitable tool for Drosophila ageing research. They characterise induction efficiency across sex, transgene insertion site, auxin dose and age, then test whether AGES can replicate a well-established pro-longevity manipulation (dominant-negative insulin receptor expression).

      Strengths:

      The study is thorough and methodical. The authors use appropriate genetic controls throughout, which is required to properly interpret AGES-based experiments. They identify an important issue, in that activation of the AGES machinery itself (independent of any UAS-transgene) shortens lifespan and alters protein levels, while high-dose auxin independently affects body mass, and even a moderate dose (5 mM) impairs stress resistance across all genotypes. These findings are important for researchers when interpreting their experiments. The tissue and age mapping of induction efficiency (brain, fat body, gut) is also useful, and the inclusion of driver-only positive controls at each age (Figure 2) establishes that da-GAL4 activity itself is stable across the ages tested, ruling out declining driver activity as an explanation for the reduced induction seen in older flies (though, as noted below, reduced auxin ingestion with age remains a very plausible contributing factor alongside declining AGES efficacy).

      Weaknesses:

      Longevity and stress assays were conducted only in females, which, combined with the finding that males show weaker and less consistent induction, means the study cannot speak to whether the metabolic and survival costs of auxin/AGES activation observed here also apply to, or differ in, males. The KCl vehicle control matches the potassium cation (K⁺) content of K-NAA across conditions; therefore, chloride (Cl⁻) concentration differs between control and auxin-fed media (both minor weaknesses).

      Achievement of aims and impact:

      The authors achieve their stated aim. Rather than validating AGES as unambiguously suitable for longevity work, they set out to characterise its behaviour and limitations in this context, which they do convincingly. The data support their overall conclusion that AGES can be used to conditionally induce transgene expression at advanced ages, but that its use in longevity/healthspan studies requires caution and rigorous control genotypes. This is a useful contribution with direct practical value: it will help other researchers make informed decisions about whether and how to deploy AGES in ageing-related work, and the cautionary findings regarding auxin/AGES toxicity are likely to be of broad relevance to the growing community of AGES users beyond the ageing field specifically.

    3. Reviewer #2 (Public review):

      McGilvary et al. evaluate the recently developed auxin-based gene expression system (AGES) for use in aging studies of Drosophila melanogaster. This system is based on the widely used Gal4/UAS system that enables cell-specific expression of UAS-transgenes under Gal4 activator control. AGES uses an auxin-inducible degron-tagged Gal80 repressor that should prevent Gal4-dependent activation unless flies are fed auxin, providing a useful approach for temporal control of transgene induction - something that would be highly useful for aging studies. The authors perform a comprehensive analysis of AGES-dependent transgene induction in male and female flies at different ages with multiple controls, demonstrating some moderate induction in female flies only - albeit with some substantial background induction even in the absence of auxin.

      Overall, transgene induction appears to be both much lower with the AGES system compared to Gal4 driver controls and very leaky, with some tissue-specific differences in induction observed as well. Combined with their observations that auxin feeding has impacts on body mass, triacylglycerol and protein levels, and lifespan, these data raise some concerns regarding the interpretation of data obtained using the AGES system for aging or longevity studies in flies. This study provides well-needed validation for the recently developed AGES system and highlights critical caveats that will support future studies.

      Most conclusions of the paper are well supported by data, but additional controls and textual edits would strengthen and clarify the findings. In addition, the abstract and conclusions of this study should more accurately reflect the limitations of transgene induction using this AGES system in adult flies.

    4. Reviewer #3 (Public review):

      Summary:

      In this useful work, the authors characterize the auxin-based gene expression system (AGES) as a tool for studying ageing. They found that this system can be applied to ageing studies. In addition, they identified important drawbacks of the methods, including effects of insertion sites and sex on induction of the system, and that some auxin doses may have inadvertent effects on body mass and physiology. Overall, the study extends the AGES system for use in fly ageing studies and highlights some caveats. While the findings are solid pointers, more extensive characterization is needed to benchmark the extent of the caveats identified.

      Strengths:

      The study provides the first longitudinal evaluation of the AGES system's induction efficiency across the entire Drosophila lifespan. The authors also highlighted a number of caveats of the AGES system in ageing animals. These are all important points to be considered when using this system, and findings should be interpreted keeping these caveats in mind.

      Weaknesses:

      There were inconsistencies with auxin dosages between the figures.

      (1) The authors used a higher dose of auxin (20mM) compared to the original AGES paper (McClure, 2022) in Fig 3. The auxin dose-dependent effects are not linear for TAG and protein levels, highlighting that the genotype-dependent effects may be highly variable and may yield quite different results in other studies.

      (2) The results in Figure 4 showing the lack of induction in the brain are quite interesting; however, only 5mM auxin is tested. Characterizing dose-dependence for the variability of the induction in tissue types would be useful. At the very least, recapitulating prior results at 10mM should be done.

      (3) Food intake and hydration status were not measured alongside body mass, TAG, and protein endpoints. The changes seen could be an effect of decreased feeding or fluid balance rather than metabolic reprogramming.

    5. Author response:

      We thank the editorial team and all three reviewers for their time and attention to detail in reviewing our manuscript. We are particularly grateful for comments recognising the “direct practical value” and “comprehensive analysis” in our work (Reviewer #2) as well as its “thorough and methodical” approaches (Reviewer #1).

      We also appreciate the reviewers’ constructive comments to improve our manuscript, which we plan to address in a revised version. Specifically, we plan to more explicitly acknowledge some of the limitations of our study, including clearer highlighting of the experiments performed only in females (Reviewer #1); the inability of our experimental design to control for chloride concentration (Reviewers #1 and #2); the limitations of transgene induction using the AGES system in adult flies (Reviewer #2); and the rationale for using different concentrations of auxin across different experiments (Reviewer #3). We also note that Reviewers #1 and #2 have provided additional recommendations beyond the public reviews (largely relating to helpful ways to clarify our text and more explicitly acknowledge limitations), which we also plan to address.

      In addition, we plan to perform additional experiments to address specific points raised by reviewers in both their public reviews and additional recommendations. In the first instance, we plan to follow Reviewer #3’s suggestion to characterise induction in the brain using 10mM auxin, as well as Reviewer #2 and #3’s suggestions to explore feeding behaviour of flies fed auxin within our experimental setup.

      We look forward to submitting an improved manuscript guided by the reviews, with the aim of strengthening this “well-needed validation” study that also “highlights critical caveats that will support future studies” (Reviewer #2).

    1. eLife Assessment

      This important study integrates mouse genetics with sequencing, electrophysiological and behavioral tools to uncover the behavioral role and molecular profile of a developmentally defined subpopulation in the lateral septum. The data collected and analyzed are convincing, highlighting the role of this subpopulation in threat avoidance and stress, providing a framework by which developmental origin is linked to mature neuronal function. The work will be of interest to biologists and neuroscientists working on development and behavior.

    2. Reviewer #1 (Public review):

      This study investigates the role of a specific neuronal population in the lateral septum (LS) in balancing exploratory and defensive behaviors. The authors created a mouse model (cKO) lacking Nkx2.1-lineage neurons in the LS by deleting the Prdm16 gene. They discovered that this ablation specifically eliminated Crhr2-expressing neurons, which are normally targeted by urocortin-3 (UCN-3) inputs. Behaviorally, cKO mice did not show general changes in anxiety but displayed a significantly increased exploratory drive. In a predator odor test (using TMT), cKO mice spent more time investigating the aversive stimulus compared to controls, suggesting these LS neurons normally suppress exploration during threat. Furthermore, the study found that Nkx2.1-lineage neurons in the LS are specifically activated by acute stress (body restraint), as shown by an increased number of c-Fos-positive neurons. While the loss of these neurons caused some connectivity and electrophysiological changes, the remaining Nkx2.1-lineage neurons were more excitable. Therefore, the authors demonstrate that LS Nkx2.1-lineage/Crhr2+ neurons are a distinct population crucial for calibrating behavioral responses to stress, acting to inhibit exploration in favor of defensive strategies.

      This work provides new insights into the neural circuitry underlying anxiety and threat avoidance. However, some of the methods and data analyses require revision for greater clarity, and additional experiments and analyses are needed to further substantiate the conclusions.

      Some of my specific questions and concerns are as follows:

      (1) The authors showed a reduction in the size of LS and a specific decrease in Crhr2+ neurons in cKO mice. I would suggest examining whether the density of other types of neurons (e.g., Crhr1+ cells or other known cell types in LS) was altered in the cKO mice.

      (2) For the single-cell sequencing experiment (Figure 2), it is unclear whether tissues from the 3 male and 3 female mice within each genotype were pooled together or processed individually (i.e., as 6 separate samples). This information is not clearly stated in the manuscript. Given that male and female mice exhibited behavioral differences, it would be valuable to examine sex-dependent effects in the analysis shown in Figure 2.

      (3) Previous studies have shown that LS neurons exhibit distinct firing patterns, including regular spiking, bursting, complex-bursting, and phasic spiking. Since the authors recorded from both tdTomato-positive and -negative LS cells, it would be interesting to determine whether the positive cells display a unique firing pattern, thereby representing a distinct electrophysiological cell type within the LS.

      (4) More detailed descriptions of the electrophysiological data analysis should be provided in the Methods section. Some LS neurons display spontaneous firing without current injection; therefore, it should be clarified how the resting membrane potential was measured in these cells. The amplitude and onset latency of the first spike are presented in the figures; however, it is unclear how the first spike was selected-whether from spiking responses to rheobase current or to a specific current pulse. I would suggest defining the first spike based on responses at a certain firing frequency. The method used to determine the spike voltage threshold should also be specified.

      (5) Could the authors analyze the single-cell sequencing data to examine whether changes in ion channel expression might explain the observed alterations in spike waveforms?

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Selective loss of Nkx2.1-lineage neurons in the lateral septum alters the balance between novelty seeking and threat avoidance" is an interesting study by Miguel Turrero García and colleagues. Here, the authors report a novel mouse model allowing complete ablation of neurons pertaining to the Nkx2.1 lineage by conditionally ablating the transcriptional regulator Prdm16 from the Nkx2.1 lineage. The authors combined single-nucleus RNA sequencing, histological and electrophysiological approaches, as well as behavioral analyses to demonstrate that a large portion of LS neurons are profoundly altered by Prdm16 deletion from the Nkx2.1 lineage. This manipulation preferentially impacts Crhr2-expressing neurons, leading to electrophysiological defects. At the behavioral level, this cell population is preferentially recruited in a stressful situation, and ablation of Prdm16 from this lineage leads to enhanced exploratory behavior even in the presence of a perceived threat.

      Strengths:

      The strengths of this manuscript are (i) leveraging a transcriptional regulator within a specific cell lineage and restricted to an early stage of ontogeny (ii) disrupting the developmental trajectory of a discrete neuronal population identified with an elegant snRNAseq approach and (iii) without obvious compensation (iv) and its impact on behavior in adult mice, with a special emphasis on exploratory drive in the presence of an acute stressor. The manuscript is well written; the experiments are well conducted, organized, and presented in a logical framework. Statistical analyses are well described. Each experimental group includes a sufficient number of subjects, allowing robust statistical comparisons.

      Overall, I very much enjoyed this manuscript and the elegant mouse model bridging developmental biology with systems neuroscience. Insights generated from this line of work could illuminate how discrete perturbations in gene expression programs at early stages of ontogeny could have a profound impact on the development, organization, and function of select neural circuits and how they may impinge on behavior at later stages of ontogeny.

      Weaknesses:

      Some comments and suggestions:

      (1) General-

      Photoinhibition of LS Crhr2-expressing neurons has no effect on anxiety-like behaviors in the absence of a stressor (Anthony et al., Cell, 2014). It would thus be interesting to reappraise the behavioral experiments performed with cKO mice in response to an acute stressor. The authors duly acknowledge this important point in the discussion section.

      (2) Specific-

      (a) Figure 2J: Was the increase in Crhr2 expression observed in tdTom-cells from cKO mice in the snRNAseq as well? If so, was Crhr2 expression enhanced in a specific cluster that did not belong to the Nkx2.1 lineage, or was it randomly enhanced across distributed clusters?

      (b) Figure 3 and S3: Does the lack of UCN3+/ENK+ terminals reflect a downregulation of UCN3 and ENK, or does it reflect the absence of innervation? Restricting a retrograde viral vector in iLS that expresses a fluorophore to illuminate the UCN3+ cell bodies (and lack thereof) in PefAH of cKO mice could address this question.

      (c) Figure 3: Does immunostaining for UCN3 in the PefAH area reveal cell bodies in cKO mice? In other words, is the loss of UCN3 terminal-specific or does it reflect a general downregulation of UCN3 in the PefAH?

      (d) Figure 4: The remaining tdTomato+ neurons are more excitable in cKO mice. To what extent can alterations in the electrophysiological properties of tdTomato+ neurons lacking Prdm16 be related to their survival? Is it a general response to Prdm16 deletion that is unrelated to survival? Is it a compensation mechanism in surviving cells? Or, alternatively, is it a unique property of these specific cells that favored their survival despite Prdm16 deletion?

      (e) Figure 5 and S5: Really nice figures. Great use of MoSeq with the predator odor test.

      (f) Figure 6: Interesting that the decrease in cFos induction in NeuN+ cells of cKO mice is more prominently observed in LSd when tdTomato+ cells are prominently found in the LSi/LSv (Figure S2B). Could this be related to intra-septal connectivity?

      (g) Figure 6: If tdTomato+ cells consist of 10-30% of neurons, and these tdTomato+ cells are preferentially found in the LSi/LSv, shouldn't we expect a decrease in c-Fos+NeuN+ density in the LSi/LSv (since there are generally fewer neurons in cKO mice)? If I am not mistaken, this could suggest that another unrelated LS population that is tdTom- displays an increase in cFos expression in cKO mice compared to WT mice. Could be interesting to see if the Crhr2+ neurons that are tdTom- are preferentially recruited in cKO mice as a compensation mechanism in LSi/LSv.

    4. Reviewer #3 (Public review):

      In the current work, Turrero Garcia et al. investigate the molecular, electrophysiological, and behavioral outcomes of conditionally knocking out cells from a unique developmentally-defined subpopulation in the lateral septum (LS). The authors focused on targeting cells from the Nkx2.1 developmental lineage that also expressed the transcriptional regulator Prdm16, which is uniquely upregulated in LS postmitotic neurons. They observed that this mutant line (Nkx2.1Cre;Prdm16fl/fl;Ai14; cKO) resulted in complete ablation of neurons from the Nkx2.1-lineage exclusively in the LS and not in the medial septum, making it an ideal model to study a lineage-defined subpopulation in the LS. They performed single-nucleus RNA-sequencing and uncovered four neuronal subtypes missing in the cKO. Furthermore, the authors validated this expression loss in the subtype expressing Crhr2 and observed a reduction of UNC-3 inputs (a neuropeptide with high affinity for Crhr2), highlighting additional disruptions in circuit connectivity. Loss of Prdm16 in Nkx2.1-lineage cells only resulted in mild electrophysiological changes in the LS. Finally, the authors performed a battery of behavioral assays to study anxiety-like behaviors and threat avoidance and observed increases in exploratory behaviors in some but not all assays. It should be noted that the cKO line also results in 30% loss of cortical interneurons, which could be contributing to the behavioral phenotype and not be exclusively due to LS loss. Overall, this manuscript takes a novel perspective by providing unique insights into how embryonic origin gives rise to mature molecular identity and distinct behaviors in adults. Therefore, it elegantly links developmental origin to mature molecular identity and function in the LS, an important question understudied in the field.

      The conclusions of the work are overall supported by the data and limitations discussed, but some of the findings, in particular the histological and behavioral results, need to be extended.

      Strengths:

      (1) Utilizing developmental origin as a marker for mature neuronal identity and function is a valuable approach which remains under-utilized in the field and serves to provide a deeper understanding of how circuits are shaped to allow for appropriate behavioral responses.

      (2) The authors perform a comprehensive analysis of the cKO mutant to determine the role of the developmentally defined Prdm16 in Nkx2.1-lineage cells at the molecular, cellular, electrophysiological, and behavioral level.

      (3) The sn-RNAseq dataset in the LS of WT and cKO mice will be valuable to the neuroscience community.

      (4) The authors perform an extensive array of behavioral paradigms investigating the balance between threat avoidance and exploratory behavior, performing all experiments in both male and female mice to determine whether the same developmental origin can lead to sex-specific differences.

      Weaknesses:

      (1) It remains unknown whether the reduction of UCN3 inputs to the LS is due to loss of the Nkx2.1-lineage in the LS itself or due to reductions in the number of UCN3 cells that provide innervation to the LS in the cKO (Figure 3). The authors speculate and include anecdotal observations that the perifornical region of the hypothalamus (PeFAH) provides inputs to the LS and could be the region driving the differences in UCN3 inputs in cKO. The authors should expand on this histological data and directly test whether the UCN3 inputs are indeed originating from PeFAH and whether the loss of Prdm16 in the Nkx2.1 lineage leads to a reduction in cell numbers in these inputs. These would disentangle the authors' claims on whether it is due to loss of Prdm16 in the Nkx2.1 lineage cells in the LS or whether it is due to loss of Nkx2.1 lineage neurons in upstream regions.

      (2) The authors claim that there is an increase in exploratory drive in cKO mice, even though the dark-light test showed increases in time spent in the dark side for cKO mice in comparison to controls (Figure 5C). They discuss that this could be due to the mice being placed first in the light compartment of the chamber during the light hours, so they spend more time exploring the dark compartment of the chamber instead, which would be considered 'novel'. If this were the case, to make the results more solid, authors should place cKO mice in the dark compartment of the chamber during the dark hours and then record time spent in both chambers. If there was indeed an increase in exploratory drive in new environments, the authors should see increases in time spent in the light compartment.

      (3) The result that cKO mice spend more time than controls exploring the inlet with TMT is of interest (Figure 5E, F). It needs to be highlighted that this is primarily the case in male mice and there is a trend in females. To confirm that the increases in exploration time of the inlet are not due to overall increases in general arousal, locomotion (e.g., velocity; pixels/frame) should be assessed in cKO vs controls.

      (4) Statistical analysis correcting for repeated testing should be performed when running multiple t-tests in the same dataset, such as when analyzing histological results in Figure 1 and Figure 6 to increase confidence in the presented results.

      (5) An important consideration that the authors address in the discussion is that the loss of Prdm16 in Nkx2.1-lineage cells is, for the most part, restricted to the LS, but other regions such as the cortex also show decreases in this population. Therefore, to strengthen the authors' conclusions that Nkx2.1-lineage neurons in the LS are indeed directly responsible for balancing threat avoidance and exploratory drive, targeted manipulation experiments, or at least additional c-Fos experiments assessing activity of Nkx2.1-derived cells in the LS will need to be performed in the future.

    1. eLife Assessment

      This study examines how prefrontal population dynamics tune social behaviors, providing important findings that oxytocin receptor-expressing interneurons support sociosexual choice in female mice. The multidisciplinary evidence is convincing and could be strengthened by more comprehensive reporting of the methodology and a clearer distinction between inferential and directly tested insights from modelling. This work will be of interest to those interested in prefrontal cortex, oxytocin, and/or social behavior.

    2. Reviewer #1 (Public review):

      Summary:

      Amadei et al investigate how excitation/inhibition balance in the prefrontal cortex plays a role in social behavior. To address this question, they developed a behavioral task where adult female mice can choose between a social reward (e.g., an adult male for sociosexual choice, or an adolescent female mouse) and a non-social reward (e.g., milk). They found that optogenetic inhibition of inhibitory neurons expressing oxytocin receptors (OXTR neurons) in the prefrontal cortex (PFC) reduces choice for sociosexual interaction compared to non-social reward and to a greater extent in sexually receptive females. They also found that this manipulation increases pyramidal neuron activity. Specifically, the authors identified a neuronal ensemble which represent the male option. Inhibition of OXTR disrupts the ability of the neuronal ensemble to represent the male option during decision-making in the behavioral task. Thus, using computational modeling, the authors proposed that OXTR neurons promote male choice by letting a male-representing pyramidal ensemble outcompete other pyramidal populations in the mPFC.

      Strengths:

      The study addresses an important topic in social behaviour and reward neuroscience with a focused hypothesis. The combination of behavioral testing and circuit manipulation combined with calcium imaging is a clear strength, and the work has the potential to make a solid contribution.

      Weaknesses:

      The main weaknesses are limited methodological clarity and details.

    3. Reviewer #2 (Public review):

      The authors aim to understand how inhibitory circuitry within the medial prefrontal cortex regulates the selection of sociosexual behaviour. Rather than studying social interaction in isolation, they develop an elegant behavioural paradigm in which female mice repeatedly choose between interacting with a male and obtaining an appetitive non-social reward. This task allows the authors to examine behavioural choice under conditions that more closely resemble natural decision-making. They combine optogenetic inhibition of oxytocin receptor-expressing interneurons, large-scale calcium imaging of pyramidal neurons, slice electrophysiology, and computational modelling to investigate how inhibition shapes cortical representations that ultimately bias behavioural choice.

      The study has several notable strengths. The behavioural paradigm is novel and well-designed, allowing repeated choice measurements while controlling for general social motivation by including both male and juvenile female stimuli. The integration of multiple experimental approaches is particularly impressive. The behavioural effects of optogenetic inhibition are complemented by population imaging demonstrating elevated pyramidal activity, electrophysiological recordings confirming monosynaptic regulation of pyramidal neurons, and a computational model that provides a mechanistic interpretation of the observed circuit dynamics. The work therefore spans multiple levels of analysis, from synaptic interactions to behaviour, and the individual datasets are generally of high technical quality.

      The imaging analyses identifying a putative "MALE" ensemble are particularly interesting. The observation that a relatively small subset of pyramidal neurons preferentially represents the male option before behavioural commitment provides an attractive framework for understanding how inhibition can stabilise specific behavioural representations. The temporal analysis suggesting that disruption of this representation precedes impaired behavioural choice is especially compelling, as it moves beyond simple correlations between neural activity and behaviour.

      Several aspects of the mechanistic interpretation remain somewhat speculative. The central conclusion relies heavily on the computational competition model, which assumes an asymmetric competition between a relatively small male-selective ensemble and a much larger default pyramidal population. While the model successfully reproduces several experimental observations, many of its architectural assumptions are inferred rather than experimentally demonstrated. In particular, the designation of the remaining pyramidal neurons as a functional "OTHER" population representing the non-social alternative is not directly established experimentally. Alternative circuit architectures may be capable of producing similar behavioural and population-level effects, and the current data do not fully distinguish among these possibilities.

      Similarly, although the identification of MALE cells is thoughtfully performed, the classification depends on an operational threshold derived from ROC analysis and correlated activity. It remains uncertain whether these neurons constitute a stable functional ensemble across sessions or merely reflect one end of a continuous representational spectrum. Longitudinal analyses examining the stability of these ensembles across days or across changes in behavioural state would strengthen the claim that they represent a dedicated neuronal population.

      An additional limitation concerns the specificity of the behavioural interpretation. The reduction in male choice is interpreted primarily as impaired sociosexual decision-making. While the inclusion of juvenile female stimuli substantially improves the experimental design, it remains difficult to completely separate altered sociosexual motivation from broader changes in motivational salience, valuation, or action selection. The observed changes could reflect alterations in multiple components of the decision-making process, and this distinction deserves a somewhat more balanced discussion.

      The interaction with the oestrous state is a very interesting aspect of the work and is consistent with previous studies of oxytocin-dependent sociosexual behaviour. However, this analysis is based on relatively modest numbers of animals and sessions, making it difficult to judge the robustness of these effects. The conclusions regarding hormonal modulation would therefore benefit from a more cautious interpretation.

      Overall, the authors achieve their primary objective of demonstrating that oxytocin receptor-expressing interneuron-mediated inhibition contributes to the selection of sociosexual behaviour while regulating pyramidal population dynamics in the medial prefrontal cortex. The behavioural, imaging, and electrophysiological datasets provide convincing evidence that inhibition shapes cortical activity during decision-making. The computational model offers a plausible mechanistic framework linking these observations, although some aspects of this framework remain hypothetical and await further experimental testing.

      The work is likely to have a significant impact on the fields of cortical circuit function, social neuroscience, and decision-making. Beyond its specific findings, the study introduces a behavioural paradigm that should prove broadly useful for investigating how competing behavioural options are represented within prefrontal circuits. The combination of behavioural neuroscience, population imaging, and computational modelling represents a valuable resource for the community and provides an important foundation for future studies examining how excitation-inhibition balance shapes flexible social behaviour.

    4. Reviewer #3 (Public review):

      Summary:

      Using a combination of Miniscope imaging and optogenetic manipulation, Amadei et al. reveal how oxytocin receptor neurons in the prefrontal cortex of mice control pyramidal subpopulations and socio-sexual behavior. This work was planned and executed carefully and provides a novel and important angle to study the oxytocin system in the cortex. According to their results, oxytocin receptor neurons help discriminate between sexual and non-sexual stimuli, most likely by controlling different pyramidal subpopulations that are either most active during trials that include a sexual stimulus or that include non-sexual stimuli. I highly appreciate this article; however, I have one major concern related to the modeling part.

      Strengths:

      (1) Well-designed experiments.

      (2) Rigorous analysis.

      (3) Generates a new avenue to study socio-sexual decision making and creates a hypothesis about the connectivity of oxytocin-sensitive circuits.

      Weaknesses:

      (1) Major

      In their last figure (Figure 4), the authors generated a computational model that, according to the authors, reveals a potential network mechanism in which oxytocin receptor (OXTR) neurons are connected to both pyramidal populations with certain connectivity rules. Although the model seems to reproduce the experimental results, some assumptions of the model seem to be poorly supported. If I understood correctly, the authors simply assumed that the strength of the connections between OXTR neurons and MALE neurons is the same as the strength of the connections between OXTR neurons and OTHER neurons. The authors neither discuss literature supporting such connectivity nor provide experimental evidence for this. I also could not find information about the magnitude of the synaptic weights to each of these populations. I guess these parameters are critical for the outcome of the simulation, and it may be worth exploring the outcome of simulating the different combinations of connectivity and synaptic weights between OXTR neurons and pyramids, as well as the degree of recurrent connectivity within the pyramidal subpopulations. Further, the authors should at least discuss in depth how inhibitory OXTR neuronal subtypes (they have different properties that could potentially be implemented in the modeling) may match their computational model best. If the current model remains the most promising, the authors should clearly discuss which experimental trajectory should be taken next to actually provide proof for its correctness (e.g. whether and how it would be possible to determine the predicted connectivity experimentally).

      (2) Minor

      The authors state regarding counterbalancing in Figure 1 and Figure S4C: "The social presentation order (male or female first), as well as the locations of the social and milk options (left or right relative to start arm), were fixed over sessions within a given subject, but varied over subjects (Figure S4C)". In Figure S4C, it looks as if there are fewer animals in which the male was always presented first than animals in which the female was always presented first (~ 16 vs. 20). While the difference is not very big, it may be influential. To fully exclude a sequence effect, I would suggest adding male-first animals until both groups are the same size.

    1. eLife Assessment

      This important paper provides evidence that visual representations, as observed through drawing, preferentially preserve topological features over metric features. Across three experiments with different age groups and creative converging methods, they demonstrate that holes and T-junctions are relatively preserved as compared to topologically irrelevant L-junctions, and Euclidian features like length or precise angle. Together, these data provide convincing evidence for topology as an organizing principle for spatial representation, but the work would be strengthened by considering a broader range of topological features, or by ruling out alternative accounts, like preserved memory for complex features, motor production limitations, or noisy compression, that may produce a similar pattern without a topological prior.

    2. Reviewer #1 (Public review):

      The paper presents novel evidence that spatial representations prioritize coarse topological features (T‑junctions, holes, crosses) over precise Euclidean metrics like angle and length, using drawing-based memory tasks with adults and children. The study is interesting and well‑motivated, and the importance of topological relations is clear, but stronger and more nuanced evidence is needed before concluding that topological relations are more important than metric details, as task difficulty and the potentially distinct roles of metric and topological information in spatial representation have not yet been fully disentangled.

      Introduction<br /> (1) P.5: Please explain in more detail what you mean by "What is relevant is the relative prioritization of each of these features."

      Results<br /> (2) P.8: Please clarify how the "proportion of drawings with angles biased towards 90{degree sign}" was computed. Specify the criterion for counting a drawing as biased (e.g., a certain absolute deviation toward 90{degree sign} from the original angle), and explicitly state in the Results that absolute degrees of deviation were used, as described in Methods.

      (3) It would help to spell out whether the findings imply that obtuse angles are typically drawn smaller (closer to 90{degree sign}) and acute angles larger (closer to 90{degree sign}). Also, would angles be more biased toward 90{degree sign} or 180{degree sign} (or 0{degree sign}) depending on the angle? (e.g., 175{degree sign} is seen more as 180{degree sign} while 95 is seen more as 90{degree sign})

      (4) Figure 4B: The statement that "positive values indicate bias in the direction of 90 degrees" needs a more precise explanation. Please explain exactly how the bias metric is computed (e.g., signed difference between drawn and original angle, with the sign indicating movement toward or away from 90{degree sign}) and what the y-axis values represent. Given that the Methods refer to absolute deviations, it would be useful to reconcile where the positive/negative signs come from in this plot.

      (5) Figure 4C: The description in the Results seems to use a different metric than what is plotted. Please ensure that the measure in the text matches the measure shown in the figure, and adjust labels or wording so they align clearly.

      (6) P.11: Consider briefly justifying why the authors predicted that participants would also add L‑junctions, rather than only remove them.

      (7) P.12: The last sentence: Weren't the overall rates of feature preservation 'higher' in the adult sample?

      Methods<br /> (8) Experiment 1: Please clarify whether the angles associated with T‑ and L‑junctions were equated or differed systematically. A short description of stimulus generation (e.g., angle ranges, line lengths, junction configurations) would be helpful.

      (9) It would also be helpful to specify the statistical tests used (e.g., t‑tests, ANOVAs, mixed‑effects models), including the main factors and any random effects, so readers can clearly follow your analysis pipeline.

      Discussion<br /> (10) It may be important to note that task difficulty likely differs across feature types: junctions involve presence/absence or counting, whereas angle and length reproduction require finer metric precision. The authors' claim of "prioritization" and possible difficulty effects should be disentangled.

      (11) Furthermore, would it be possible that people retain relative order/comparison of different angles/lengths rather than computing precise values?

      (12) I agree that topological relations are extremely important. However, for above reasons, it seems like stronger/stricter evidence is needed to claim that topological relations are 'more' important than metric details. They also might serve different roles in spatial representations

      (13) The Discussion would benefit from a short paragraph on where different junction types (T, L, crosses) typically appear in everyday scenes and objects (e.g., as cues to occlusion, surface intersections, 3D structure) and what functions they serve. This would help connect your experimental findings to the ecological importance of these features for natural vision and spatial cognition.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting study that uses drawings to evaluate the extent to which visual representations of letter- and graph-like figures (preferentially) include topological features, like junctions and holes.

      The main claim is based on the observation that when participants are asked to draw presented figures from memory, they tend to (1) regularise angles towards 90deg and lengths towards the average length of the lines in the figure, while (2) preserving topological features like T-junctions more assiduously than non-topological features like L-junctions. A third experiment with 'serial reproductions' in which participants copy drawings made by other participants (like a visual version of the 'broken telephone' game) reproduce these patterns in exaggerated form. These findings were also reproduced in children (Experiment 4).

      These findings are consistent with the idea that memory representations are low-bandwidth or noisy approximations to the original figure. I would suggest that when participants are asked to reproduce the figure, it is if they combine the noisy stored representation, with generic priors about angles and the average line length. The preferential preservation of T- over L-junctions indicates that they are somehow more salient or memorable. This is not inconsistent with the authors' preferred interpretation of an explicit representation of topological structure. However, it is also not inconsistent with the idea that in order to compress the visual signals for storage, high-information (complex) components of the source are given preferential treatment. This would be compatible with optimal use of limited resources when compressing the information. Additional comparisons and control conditions would help tease these alternatives apart.

      Strengths:

      + Innovative use of drawing methods to probe internal visual representations<br /> + Experiments spanning both adults and children

      Weaknesses:

      - Failure to consider alternative hypotheses that are consistent with the findings

    4. Reviewer #3 (Public review):

      Kittur et al. ask whether human spatial memory is organized around topological relations (meaning coarse structural properties such as T-junctions, crosses, and holes) rather than around Euclidean properties such as angle and length. Across four experiments, adults and children studied letter-like figures and reproduced them from memory by drawing. The authors report two complementary patterns: metric features are systematically distorted, with angles pulled toward 90 degrees and line-length ratios compressed toward an average, while topologically critical features are comparatively well preserved. The central test contrasts T-junctions with L-junctions, which are visually similar but topologically distinct, since an L-junction reduces to a straight line, whereas a T-junction does not. A serial reproduction experiment amplifies both patterns across chains of participants, and a fourth experiment extends the findings to children aged five to eight.

      Strengths:

      The question is a good one and sits at a productive intersection of topics. It bears on debates about the representational format of cognitive maps, on proposals about the primitives of visual perception, and on a classic developmental claim from Piaget and Inhelder that has rarely been tested directly.

      The drawing paradigm is well chosen and offers something that the group's earlier forced-choice work could not. Because participants produce an open-ended response, distortion of metric detail and preservation of structure can be observed within a single response, and the relationship between them can be examined directly. The serial reproduction experiment is a particularly effective use of this affordance. The choice of the T-junction versus L-junction contrast as the primary test is well-motivated, since it holds the number of junctions constant and varies only topological relevance. It is also worth noting for readers that the central claim of a representational privilege for topologically distinct features was previously established by this group using forced-choice paradigms in both adults and children. That a similar conclusion emerges from free generation is a genuine strength, since the two methods have very different sources of error.

      The work is carefully executed. Sample sizes, dependent variables, and analyses were preregistered; stimuli were purpose-built for each question, including the deliberate exclusion of 90-degree angles so that no reference angle was available; drawings were double-coded; and the full set of raw drawings is being released publicly.

      Weaknesses:

      Drawing is treated as a transparent window onto representation, and motor limitations are not considered. Drawing is a motor act, drawing skill varies widely across individuals, and the manuscript does not discuss motor limitations at any point. As the study is designed, representational imprecision cannot be separated from difficulty of precise reproduction. The clearest way to resolve it might be asking adults to copy the figures exactly while the stimulus remains visible. If the biases persist under direct copying, then some portion of the effect is production rather than memory. Because the topological findings have already been demonstrated in keypress-only paradigms, this concern affects the metric distortion results most heavily, which are the novel contribution of the present paper.

      Three distinct claims are treated as one, and the data speak mainly to the weakest of them. The paper moves between a claim about mnemonic robustness (topological features survive degradation better than metric features), one about representational architecture (topology is a base layer with metric detail superimposed on top), and a claim about priority (topology is encoded prior to metric detail). The experiments show evidence for robustness, which is a claim about what is lost first. Robustness does not entail architecture: an encoder with a single layer, whose loss happens to spare structure, produces the same pattern with no layered format and no claim about encoding order. Earlier work does address format, because false "same" judgments to topologically matched but metrically different items show that topology plays a role in what the system treats as equivalent. Preservation counting measures robustness, rather than equivalence.

      Some alternative hypotheses to consider/address: (a) A capacity-limited memory that reconstructs from a prior produces the metric distortions with no commitment to topology. A literal absence of angle encoding, which the authors invoke, predicts noisy and unconstrained recall rather than recall pulled toward a particular value. The observed pattern reflects a structured prior. (b) The result that does discriminate might be confounded with local salience. A memory that adds uniform noise to all parts of a figure does not predict that T-junctions are preserved better than L-junctions; however, a three-way branch point is plausibly more locally distinctive than a corner, so a salience-weighted-but-topology-free account predicts the same ordering. (c) Motor simplification also predicts the same ordering, since omitting an L-junction converts a bend into a straight line, which is easier to draw, whereas omitting a T-junction requires dropping a stroke.

      The better a feature works as a topological marker, the less variance it produces and the harder it is to test, so the method is best powered where the theoretical signal is weakest. Holes are the textbook case of a topological invariant and are reported to disappear from drawings less than one percent of the time, but they are excluded from formal analysis because they are at ceiling and have no matched comparison feature. It would be useful to see bidirectional rates for holes (both how often a hole disappears and how often participants spuriously close an open figure into a loop).

      In Experiment 4, the conclusion of developmental stability rests on a nonsignificant effect of age, which is failure to detect a change rather than evidence of stability. An equivalence test or an estimate of the precision of the null is better support for developmental stability. Motor skill is confounded with age throughout. So this is an experiment that shows that the effect generalizes to childhood, but cannot adjudicate a developmental question.

      In the serial reproduction experiment, chains were intermixed so that each participant contributed one drawing to each of ten chains. The final drawings are therefore linked through shared intermediate participants, and an individual with an idiosyncratic drawing style influences ten chains at once, so the reported degrees of freedom are somewhat generous. The analysis also focuses on the final drawings and sets aside the nine hundred intermediate ones, which are the data that would show where in a chain metric detail collapses and whether structure ever breaks.

      To formalize the topology is to strengthen the argument: each figure is a one-dimensional complex (its underlying graph), treated intrinsically and up to homeomorphism. The homeomorphism type is what remains after suppressing all degree-2 vertices. Under this definition, every feature in this paper's taxonomy becomes one kind of object, namely a homeomorphism invariant of the graph: number of components, first Betti number, and the degree sequence of three or greater with its adjacency structure. Relatedly, the term "metric" needs to be unpacked. The 90-degree bias concerns angle, whereas the 4:2:1 result concerns length ratios, which are affine rather than strictly metric. The stronger statement available is that distortion appears at every level above topology in the transformation hierarchy while preservation occurs at the topological level, which connects directly to Chen's (2005) invariance hierarchy that is already cited.

      Appraisal and impact:

      The authors aimed to show that topological structure is preferentially retained in memory while metric detail is lost, and in the sense of relative preservation they succeed. The dissociation is real, replicates across two stimulus sets, amplifies under serial reproduction, and appears in young children. What the data do not establish is the stronger architectural claim that topology is a base representational layer, nor that the metric distortions specifically implicate topology rather than general properties of reconstructive memory. Separating these claims would communicate the well-supported result better. Conceptually, the work strengthens a growing case that coarse relational structure deserves a place alongside Euclidean properties in accounts of spatial representation. Practically, the public release of the full set of adult and child drawings, including excluded ones, is a resource that will support analyses well beyond those reported in this paper, and the serial reproduction design is a method that others will want to borrow.

    1. eLife Assessment

      This important study introduces a Bayesian method to determine bacterial counts that accounts for the experimental noise inherent to dilution and plating methods and distinguishes it from biological uncertainty. The evidence supporting the conclusions is compelling, combining simulated data and experimental data. The method will be of interest to microbial ecologists, and potentially to the broader community interested in inference from biological data, even more so if the domain of application and the limitations are further clarified.

    2. Reviewer #1 (Public review):

      Summary:

      The authors developed a novel theoretical/computational procedure to count bacterial populations without introducing artificial randomness effects due to dilution. Surprisingly, this very important aspect of studies of bacterial systems has been overlooked. The proposed method provides a simple and transparent approach to eliminate the randomness of bacterial accounting procedures, allowing now to fully concentrate on the intrinsic effects of the studied systems.

      Strengths:

      A very simple and clear procedure is introduced and explained in full detail. This elegant approach finds an excellent compromise between mathematical rigor and computational efficiency, which is important for practical applications. The provided examples are convincing beyond a doubt, clearly indicating the potential strong impact of the proposed framework. Various complications and possible issues are also discussed and analyzed. This seems to be a very powerful novel method that should significantly advance the analysis of complex biological systems.

      Weaknesses:

      The only minor weakness that I found is the assumption of independence of bacterial species, which is expressed as the well-stirred approximation. One could imagine that bacterial species might cooperate, leading to non-uniform distributions that are real. How to distinguish such situations?

      I believe that this method can be extended to determine if this is the case or not before the application. For example, if the bacteria species are independent of each other and one can use the binomial distributions - then the Fano factor would be proportional to the overall relative fraction of bacterial species. Maybe a simple test can be added to test it before the application of REPOP. However, I believe that this is a minor issue.

      Comments on revised version.

      I am satisfied with the correction proposed by the authors. The method is already quite impressive, and there is no need to complicate it at this stage.

    3. Reviewer #2 (Public review):

      I appreciate the thorough responses from the authors, which address my concerns. The expansion of Appendix B as well as the addition of text discussing how the REPOP method interacts with data collection efforts are very useful. These new sections show that relative error decreases with increasing samples, as expected, yet these error metrics, including KL divergence, describing the fit of the full distributions not just the modes, drop off fairly quickly with increasing number of samples showing that REPOP likely minimizes discrepancies between estimated and true distributions even at lower sampling efforts.

      Additionally, the extension of the REPOP method to the quantification of multiple bacterial species or phenotypes shows the potential utility of the method in contexts beyond basic plate counts. Between this example and the additional information on how to implement REPOP, I believe this workflow will be attractive and accessible to the audience.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (R1C1) The only minor weakness that I found is the assumption of independence of bacterial species, which is expressed as the well-stirred approximation. One could imagine that bacterial species might cooperate, leading to non-uniform distributions that are real. How to distinguish such situations?

      I believe that this method can be extended to determine if this is the case or not before the application. For example, if the bacteria species are independent of each other and one can use the binomial distributions, then the Fano factor would be proportional to the overall relative fraction of bacterial species. Maybe a simple test can be added to test it before the application of REPOP. However, I believe that this is a minor issue.

      This is an interesting point raised by the reviewer.

      First, we need to clarify an important point: we do not make a well-stirred assumption. Samples can be drawn and plated from any region of space however small and that region’s population can be quantified using our method. The stirring only occurs after we collect a sample in order to dilute the contents and pour the solution homogeneously over the plate.

      As such, learning multiple independent species is possible and not impacted by the dilution (“well-stirred” assumption). In the new first paragraph of the methods section, we made it clear that this assumption concerns the dilution process. REPOP is designed to recover the true underlying heterogeneity in species abundance (even from limited data) by leveraging a Bayesian framework that remains valid regardless of whether species are independent or correlated.

      If the method is applied to multiple species as currently implemented, REPOP can recover the marginal distribution of each species, provided that the species are either selectively cultured or produce sufficiently distinguishable colonies on the same plate. To demonstrate this, we have added a new Results subsection with a synthetic two-species example in which the species abundances are correlated across samples.

      However, in order to learn the joint distribution and capture correlations between species within samples, the method would need to be extended. At present, in Eq. 5 we sum the likelihood over all values of n, using a data-driven cutoff (twice the largest naïvely estimated count times the dilution factor). Extending this to multiple species adding up to (n<sub>1</sub>,n<sub>2</sub>), while retain the generality of the method, would require quadratically scaling memory with this cutoff in the population number. For this reason while we comment on this in the new paragraph in the conclusion, it is not implemented as part of REPOP.

      Reviewer #2 (Public review):

      (R2C1) A more thorough discussion of when and by how much estimated microbial population abundance distributions differ from the ground truth would be helpful in determining the best practices for applying this method. Not only would this allow researchers to understand the sampling effort necessary to achieve the results presented here, but it would also contextualize the experimental results presented in the paper. Particularly, there is a disconnect between the discussion of the large sample sizes necessary to achieve accurate multimodal distribution estimates and the small sample sizes used in both experiments.

      That is a great suggestion from the reviewer. To address it, we expanded Appendix B. We know report (1) the relative error in the estimated means (as already done for Fig. 4 formally 3), and (2) the Kullback-Leibler (KL) divergence between the reconstructed and ground-truth distributions. These metrics will are show as a function of the size of the dataset, for the examples in Fig 3. enabling a direct assessment of how the sampling effort affects the precision of the inference.

      That said, we now highlight in the Conclusion that, by explicitly modeling the dilution process within a Bayesian framework, REPOP extracts the maximum information available from each individual sample at a given sample size. This strategy therefore enables more accurate inference with fewer measurements, which is particularly important in applications such as plate counting, where data acquisition is labour-intensive.

      Reviewer #3 (Public review):

      (R3C1) While the study is promising, there are a few areas where the paper could be strengthened to increase its impact and usability. First, the extent to which dilution and plating introduce noise is not fully explored. Could this noise significantly affect experimental conclusions? And under what conditions does it matter most? Does it depend on experimental design or specific parameter values? Clarifying this would help readers appreciate when and why REPOP should be used.

      We agree with the reviewer that this is an important point, and we expanded Appendix B to include a quantitative analysis using simulated data (Fig. 3, formely 2), reporting both relative error and KL divergence as a function of dataset size. This complements our response to R2C1 clarifying when REPOP offers the greatest benefit.

      In addition, we will expand the discussion on how modeling dilution noise becomes essential when learning population dynamics. In particular, we emphasize? the role of Model 3, especially relevant when working with multiple plates and approaching the asymptotic regime; an aspect that was alluded to in Fig. 3 but not fully explored.

      (R3C2) Second, more practical details about the tool itself would be very helpful. Simply stating that it is available on GitHub may not be enough. Readers will want to know what programming language it uses, what the input data should look like, and ideally, see a step-by-step diagram of the workflow. Packaging the tool as an easy-to-use resource, perhaps even submitting it to CRAN or including example scripts, would go a long way, especially since microbiologists tend to favor user-friendly, recipe-like solutions.

      In the new paragraphs of the introduction, we made clear that REPOP is written in Python (PyTorch), installable via pip, and designed for ease of use. We are also expanding the tutorials to include clearer guidance on data formatting and common workflows. The new workflow figure (Fig 2) better illustrates the full process.

      (R3C3) Third, it would be great to see the method tested on existing datasets, such as those from Nic Vega and Jeff Gore (2017), which explore how colonization frequency impacts abundance fluctuation distributions. Even if the general conclusions remain unchanged, showing that REPOP can better match observed patterns would strengthen the paper’s real-world relevance.

      We thank the reviewer for this interesting suggestion. We agree that applying REPOP to additional existing datasets would make REPOP’s relevance clearer. However, the Vega and Gore datasets lack the information required. REPOP requires the plate count measurement process to be specified, including the dilution factors used for each measurement. Furthermore, we can leverage on additional information about the experimental procedure when the colony cutoffs and dilution schedules used are reported. Without the dilution factors, the likelihood connecting the observed colony counts to the underlying population size is not possible. We hope this clarification will help make future datasets made available publicly more useful for purposes of uncertainty propagation.

      (R3C4) Lastly, it would be helpful for the authors to briefly discuss the limitations of their method, as no approach is without its constraints. Acknowledging these would provide a more balanced and transparent perspective.

      We agree with the reviewer. We have added two new paragraphs to the conclusion highlighting important current constraints and future development directions of the framework. In particular, we now discuss that, in its present implementation, REPOP focuses on the population distribution that maximizes the posterior, rather than returning posterior uncertainty over the reconstructed distributions themselves. We also note the computational demands of the method, making GPU acceleration highly beneficial and more complex multi-population inference computationally challenging. This discussion synthesizes points raised throughout our response to R1C1 and the reviewers and provides a more balanced perspective on the current scope of the method.

    1. eLife Assessment

      This valuable study examines how the prelimbic cortex represents learned and generalized threat over time and identifies potentially distinct stable and dynamic subnetworks that may support these functions. The work is conceptually interesting and is strengthened by the longitudinal calcium imaging approach and the inclusion of key control groups. However, the evidence supporting the claims is incomplete, particularly because the interpretations regarding inference, time-dependent representational change, and the dissociation of neural activity from freezing behavior extend beyond what is currently established by the data.

    2. Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.

      To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

    3. Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      a. 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      b. 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      c. 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS⁺. In CS⁺15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      (8) Another example of what has been a common theme in this review :

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      (10) The following points all relate to the Discussion and reiterate many of the points above.

      a. 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      b. '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      c. 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      d. 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

    4. Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.

      To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to some neural changes across sessions. However, several aspects of our behavioral design and results argue against extinction or repeated nonreinforced retrieval as the primary drivers of the observed effects. Importantly, discrimination ratios remained stable or increased across time rather than progressively diminishing as would be expected under extinction (this new analysis will be added to the resubmission). Nevertheless, we will address this important point in the Discussion and explicitly acknowledge that repeated retrieval may contribute to some component of the observed representational changes.

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      We thank the reviewer for raising this important point, which was also noted by the other reviewers. To address this issue, we will implement Reviewer 3’s suggested Generalized Linear Model (GLM) analysis using inferred spiking activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. Because freezing behavior varies across trials whereas stimulus identity is fixed, this approach will allow us to dissociate their respective contributions to neuronal activity. If, after accounting for freezing behavior, responsive neurons continue to exhibit graded coding consistent with inferred threat value, this would strengthen the interpretation that the identified ensembles reflect generalization gradients related to aversive value rather than freezing behavior alone. Otherwise, we will adjust the conclusions according to the interpretation that freezing itself drives the generalization gradients.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      We thank the reviewer for this suggestion. We will include measures of registration quality in the resubmission.

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      We will correct correspondence between text and Figure.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      Thank you for this suggestion. We will clarify the labelling of the Figure 2a and call the graphs “activity-plots”.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      We will redo the Statistics of Figure 3 to take into account the following variables: group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30).

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term “inference” may overstate the cognitive processes engaged by the current task. Accordingly, we will revise the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

      We will correct the language and use “threat value”.

      Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      We thank the reviewer for these important points, which we address in further detail below.

      Ongoing learning during test sessions: The reviewer correctly notes that unreinforced test presentations may constitute extinction-learning trials and that some neural changes across days could therefore reflect ongoing learning rather than spontaneous ensemble reorganization. However, new analyses indicate that extinction is unlikely to be the primary driver of our findings. Discrimination ratios do not decay over time; instead, they either sharpen or remain stable across sessions (new analyses to be included in the resubmission). These results argue against robust extinction as the primary source of the neural changes observed across sessions. This interpretation is also consistent with the strength of our conditioning protocol, which used 10 CS+ shock pairings and 10 CS− no-shock pairings specifically to minimize extinction across repeated testing sessions. Nevertheless, we acknowledge that the current design cannot fully dissociate time-dependent consolidation from retrieval-induced plasticity, and we will explicitly discuss this limitation in the revised Discussion.

      Stability reflecting behavioral consistency: We agree this alternative cannot be fully excluded. However, the cluster stability analyses assess identity at the level of response profile across all four frequencies, not response magnitude alone. Tone-selective clusters, which also show consistent behavioral correlates (firing rate correlates with threat-value, Fig. S8), do not show equivalent profile stability, suggesting that the stability of graded clusters is not simply a consequence of behavioral consistency. This point will be added to the Discussion in the resubmission.

      Language of "inference" and "integration": The reviewer is correct that responses to novel tones are consistent with graded stimulus generalization. We will substantially revise the manuscript to replace "inference" and "integration" with more precise language describing graded frequency generalization gradients.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      We thank the reviewer for these specific and constructive criticisms. We will revise the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.

      (a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      The hypothesis will be rewritten as: "We hypothesized that generalization to tones acoustically similar to the CS+ and CS− depends on the emergence of stable ensembles encoding threat value, and that population-level response similarity across stimuli would correlate with the degree of behavioral fear generalization, consistent with prior work in auditory cortex [1]."

      (b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      The summary statement will be rewritten: "Our results show that stable cortical sub-ensembles preserve the emotional content of the fear memory over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity in response to tones associated with threat correlates with the degree of behavioral threat generalization."

      (c) 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      We will adopt the reviewer's suggested rewording: "In CS+15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization."

      We will systematically review the entire manuscript to ensure consistency with this revised framing.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we will follow Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using inferred spiking activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis will allow us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we will revise and appropriately adjust our conclusions.

      In addition, we will revise the section heading and surrounding text to remove the terminology of “learned and inferred valence.” Instead, the findings will be described more conservatively as: “PL population responses reflect behavioral generalization to auditory stimuli following discriminative fear conditioning.”

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      The reviewer is correct that controls do not form threat associations; however, these animals still could respond differentially to distinct frequencies, something that is not reflected in the data. We will correct the section indicating that distinct neutral frequencies do not produce graded responses: "graded responses require associative learning" will be retained but reframed simply as: "graded frequency-dependent population responses were absent in animals that did not receive fear conditioning." The concluding statement of the paragraph will be rewritten as: "PL population responses reflect behavioral generalization to acoustically similar stimuli following discriminative conditioning," in line with the reviewer's suggestion.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      We respectfully disagree with the reviewer’s interpretation. While repeated testing cannot be entirely excluded as a contributing factor, several lines of evidence suggest that it cannot fully account for our observations.

      Regarding extinction: discrimination ratios between CS+ and all other frequencies either remained stable or increased over time (new analysis included in resubmission), indicating that animals continued to discriminate threat value across the testing period rather than showing the progressive suppression expected under extinction — the opposite of what we observe.

      Regarding the recruitment of new neurons: repeated non-reinforced tone exposure would be expected to produce stimulus-specific adaptation — characterized by reduced, less discriminative neural responsiveness and flatter tuning profiles [2]— not the progressive sharpening we observe. The same would be expected if these neurons represent or are associated with new extinction learning.

      Finally, sharpening of generalization gradients during repeated within-subjects testing has been reported previously [3], suggesting that successive exposures may promote more precise discrimination in some cases. Consistent with this, discrimination learning has also been shown to narrow or sharpen fear generalization gradients rather than broaden them [4], supporting the idea that discriminative conditioning enhances stimulus specificity during testing. Although we cannot exclude the possibility that more extended training could eventually broaden the generalization gradient, under the training parameters and temporal window used in our study, the data support a progressive sharpening of the gradient over time. In the revised Discussion, we will present systems consolidation as the primary interpretive framework and further elaborate on why repeated testing is unlikely to account for the full pattern of behavioral and neural findings reported here.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS⁺. In CS⁺15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      The GLM analysis described in our response to reviewers 1 and 3 will directly address the contribution of freezing. We will report these results in the resubmission and revise the interpretive language in the manuscript accordingly.

      However, regarding the analysis of population vector similarity, we need to clarify a point of confusion. The reviewer states “Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained”. The similarity vectors were calculated by correlating activity across all tone presentations within each testing day, not only the first two presentations. In Fig. 4, “Early” and “Late” refer to the order of a tone within a trial, which we will clarify more explicitly in the resubmission. Notably, repeated-measures analyses did not reveal any effect of the time variable (Fig. 4e,f), indicating that similarity across tone presentations remained high for tones associated with high threat value. Importantly, our data showed no evidence that responses to 11 kHz or 15 kHz in the CS15 group, or to 3 kHz in the CS3 group, exhibited extinction-like patterns at either the behavioral or neural level. Therefore, the persistence of high population similarity across time provides additional evidence against extinction as the primary explanation for our findings.

      We will remove the word "determines" from the manuscript, as our data cannot conclusively establish a causal relationship.

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      We will make this correction.

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      We thank the reviewer for raising this point. The phrase "integrated learned valence across tones" refers specifically to a subpopulation of neurons that respond to all four frequencies in a graded manner, with response magnitude scaling according to threat value. This is distinct from tone-selective neurons, which respond preferentially to a single frequency. The neurons responding to all tones in a graded manner are present only in conditioned animals and not in no-shock controls, demonstrating that their graded response profile is shaped by associative learning.

      We agree, however, that the phrase "integrated learned valence" is unnecessarily opaque and we will replace it with more precise language: these neurons will be described as showing graded frequency-dependent responses whose magnitude scales with threat value. We believe this subpopulation represents a genuinely novel finding that complements the behavioral generalization literature by identifying a specific neural substrate for the generalization gradient within PL.

      (8) Another example of what has been a common theme in this review:

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      As stated in our previous responses, in the resubmission we will determine the contribution of freezing. If we find that freezing predicts graded neural responses, we will adjust the language of the manuscript.

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      We thank the reviewer for highlighting the lack of clarity in this passage and agree that the original phrasing was insufficiently precise. What we intended to convey is that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. Therefore, we suggest that neurons recruited into the active population after conditioning — likely frequency-selective neurons — contribute to the graded population responses through changes in their firing-rate activity, which is modulated by threat value (Fig. S8). We will rewrite this passage in the resubmission to make this interpretation explicit rather than leaving it to the reader to infer.

      Regarding the reviewer's suggestion that the characteristics of newly recruited neurons more likely reflect learning from non-reinforced exposures during repeated test sessions, we respectfully maintain that this interpretation is difficult to reconcile with two aspects of our data. First, graded-response neurons are absent in no-shock controls that are exposed to nonreinforced repeated testing. Second, as detailed in our responses to previous points, the progressive sharpening of population responses over time is inconsistent with what would be expected from repeated non-reinforced exposure, which would more plausibly produce broader or flatter tuning profiles.

      We agree that the phrase "shaped by associative processes" was ambiguous and will replace it with explicit language clarifying that we refer to fear conditioning as the associative process driving the emergence of graded responses, rather than any learning occurring during the test sessions themselves.

      (10) The following points all relate to the Discussion and reiterate many of the points above. 

      (a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      We will modify the language as stated in the prior points.

      (b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      We will incorporate new analysis (GLM) to better address this point and conclusions.

      (c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      We thank the reviewer for this comment and agree that both phrases require clarification or revision.

      By "graded representational axis" we intended to convey that PL population activity varies systematically as a function of stimulus similarity to the conditioned tone — that is, population responses are not categorical but scale continuously with spectral proximity to the CS+. We agree this was not clearly stated and will revise the manuscript accordingly.

      Regarding recurrent connectivity, we agree with the reviewer that nothing in our data directly measures or manipulates connectivity between neurons. This statement was intended as a speculative interpretive hypothesis in the Discussion, motivated by the established literature linking strong recurrent connectivity in prefrontal circuits to stable population-level representations [5]. However, we acknowledge that invoking it in this context, without direct evidence, risks overstating our conclusions. We will revise this sentence to make its speculative nature explicit and ground it more carefully in the cited literature rather than presenting it as an inference from our own data.

      In summary, we will ensure our conclusions will be restricted to population-level coding of learned threat value and its generalization across auditory frequencies. We will revise the relevant passages in the Discussion to ensure that speculative interpretations regarding emotional memory content are either removed or clearly flagged as speculative hypotheses.

      (d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

      We thank the reviewer for this point. Our interpretation that these neurons reflect pre-existing PL sensory-driven properties was based on the observation that tone-selective responses were present in control animals that never received conditioning, consistent with prior reports of sensory responsiveness in PL cortex ([6, 7]. Because these responses emerge from the first time we expose mice to the intermediate frequencies, they cannot be explained by repeated exposure. Moreover, we did not observe progressive refinement, emergence of discrimination-like changes, or suppression of responding to non-reinforced tones in control mice. This difference between conditioned and control animals indicates that repeated tone exposure alone is not sufficient to produce the observed dynamics — associative learning is necessary. We therefore maintain that the tone-selective responses of these neurons reflect pre-existing sensory-driven properties of PL cortex that are present independently of conditioning history.

      In summary, we thank the reviewer for suggesting clarifications to our interpretation, for raising the possibility that freezing behavior may contribute to graded neural responses, and for raising the question of whether repeated tone exposure may contribute to the properties of neurons recruited after conditioning. In the revised manuscript, we will include additional analyses to better dissociate the contributions of freezing behavior and tone identity, clarify passages that were insufficiently precise, and include a paragraph in the Discussion addressing potential alternative explanations alongside our own interpretation of the data.

      Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded-response ensembles may partially reflect freezing-related activity rather than mnemonic or salience-related representations of the conditioned stimuli themselves. In the revision, we will acknowledge that prior work has identified PL neurons that encode freezing independently of stimulus identity or associative content. Furthermore, we will implement the reviewer’s suggested generalized linear model (GLM) approach using inferred spiking activity derived from the Ca2+ signals. Specifically, we will include both stimulus identity and freezing behavior as predictors. Because freezing varies across trials whereas stimulus presentation is fixed, this analysis will allow us to dissociate the relative contributions of stimulus-related versus freezing-related activity to the graded neuronal responses. We thank the reviewer for this excellent suggestion.

      If graded stimulus coding remains significant after accounting for freezing behavior, this would strengthen the interpretation that these ensembles encode learned salience or associative properties of the stimuli rather than behavioral output alone. Conversely, if freezing explains a substantial proportion of the variance, we will revise our interpretation accordingly.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

      We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such holographic two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.

      To provide an indirect link between ensemble organization and behavior within the constraints of the current dataset, we will examine inter-individual variability in the revised manuscript. Specifically, we will test whether the proportion of neurons participating in stable graded-response ensembles versus dynamic stimulus-specific ensembles predicts individual differences in freezing behavior and fear generalization across retrieval sessions. If animals with a higher proportion of stable graded-response neurons show stronger discrimination and less generalization to non-conditioned tones, this would strengthen the association between ensemble organization and behavioral outcome, while remaining correlational in interpretation.

      We will modify the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the associative nature of our conclusions.

      References

      (1) Aschauer, D.F., et al., Learning-induced biases in the ongoing dynamics of sensory representations predict stimulus generalization. Cell Rep, 2022. 38(6): p. 110340.

      (2) Kato, H.K., S.N. Gillet, and J.S. Isaacson, Flexible Sensory Representations in Auditory Cortex Driven by Behavioral Relevance. Neuron, 2015. 88(5): p. 1027–1039.

      (3) Vervliet, B., et al., Generalization gradients in human predictive learning: Effects of discrimination training and within-subjects testing. Learning and Motivation, 2011. 42(3): p. 210–220.

      (4) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.

      (5) Mante, V., et al., Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature, 2013. 503(7474): p. 78–84.

      (6) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.

      (7) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.

    1. eLife Assessment

      This work provides a fundamental advance through a detailed, integrative analysis of how the tsetse fly feeds on blood, demonstrating that successful penetration depends on subtle structural adaptations rather than extreme forces or unusual anatomy. By combining high-resolution imaging, innovative biomechanical measurements, and experiments on artificial skin, the study offers complementary and compelling evidence, with clear data supporting a robust mechanistic interpretation. These findings have broad significance, as they clarify the biomechanics of vector feeding and have implications for the transmission of diseases such as African trypanosomiasis across diverse hosts.

    2. Reviewer #3 (Public review):

      Summary:

      Human and animal trypanosomiasis are fatal illnesses caused by African trypanosomes transmitted by tsetse flies during a bloodmeal. Thus, tsetse fly feeding is the key physical step in disease transmission to mammals. Tsetse fly feeding is not a new story, but it is revisited here through the application of sophisticated imaging techniques and novel biomechanical methods of analysis. The author's aim is to provide a high-resolution picture of the structures and forces involved in feeding to provide mechanistic insights into the process of feeding, from attachment, penetration, drinking and retraction of the feeding parts.

      Largely the authors have achieved their aims. They (i) examine the structures and forces involved in attachment; (ii) they provide detailed multi image analysis of the proboscis providing insights into its probing ability and physical mechanism of penetration; (iii) they conduct a controlled analysis of the physical forces involved in penetration and report that they are in the low nM range, not especially strong but much higher that the mosquito bite and finally they provide a first analysis of blood uptake during feeding.

      Strengths:

      The study images the tsetse fly feeding structures in unprecedented detail, with resolution to the uM scale, in 3-D, and during feeding. The resulting images are dramatic and insightful (and beautiful and frightening!) that researchers interested in trypanosomes, tsetse flies or blood feeding by flies in general will want to see.

      They conclude that flies attach strongly to smooth surfaces, because of interactions possible via the array of acanthae of the pulvillus pad at the ends of the tarsi. The estimated attachment forces are similar in male & female flies, in the low mM range (they look impressively strong in video 1). They provide a very striking analysis of the proboscis and labellum and associated tooth structures (Figs 4 & 5). I recall many years ago observing that tsetse flies are messy feeders, and these structures, especially the rasping teeth structures on the reverse folded labial tips explain why! This seems more like a chainsaw than a jigsaw in action, but the authors are probably correct that these structures and probing/retraction mechanism explain many features of tsetse fly feeding and their ability to feed on a wide range of hosts with very different skin types.

      The impressive aspect of this paper is the range of imaging techniques, (CLSM, SEM, uCT, FIB SEM), the quality of the images which attests to the obvious care taken with sample preparation. The biomechanically analysis, especially the penetration analysis is impressive. Finally, the paper is clearly written and presented, it was a very easy read and overall, a very engaging study.

      Weaknesses:

      I suppose it could be said that the paper is a descriptive study; it doesn't really test a hypothesis but that is not a prerequisite for publication. Perhaps the least convincing prats are the imaging of the flexible v rigid parts of the structures, which is based on amount of resilin (flexible) and chitin-protein (stiff) based on their autofluorescence. In seems odd that the joints would be less blue (stiffer) in Fig 1i, or what the blue structures correspond to in Fig. 6B-D.

      Comments on revised version.

      In revised version these issues have been satisfactorily addressed

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript provides a comprehensive and mechanistic analysis of how tsetse flies feed on blood across a wide range of host skin types. The authors combine detailed anatomical characterization of the feeding apparatus with quantitative measurements of mechanical properties, probing forces, and blood uptake, complemented by experiments using artificial skin. They show that tsetse flies do not rely on extreme forces or uniquely specialized structures, but instead on subtle and highly efficient structural and mechanical adaptations (such as the toothed labellum and coordinated proboscis movements) to achieve effective blood pool feeding. The study successfully moves beyond descriptive anatomy to a quantitative, functional analysis that explains how feeding is accomplished across diverse substrates.

      Strengths:

      A major strength of the work is the impressive integration of multiple complementary approaches. Advanced imaging tools provide a convincing three-dimensional view of the proboscis, labellum, and associated structures, while direct force measurements and blood intake quantification place these observations on a solid quantitative footing. The use of artificial skin with different mechanical properties is particularly powerful, as it allows structure-function relationships to be tested under controlled and reproducible conditions. Together, these datasets provide strong and coherent support for the authors' central conclusions. The quantitative treatment of feeding mechanics represents a significant advance over largely descriptive prior work by others (e.g., Gibson W et al 2017) and establishes a valuable mechanistic insight for studying blood feeding in insect vectors more broadly.

      Weaknesses:

      The study focuses almost entirely on uninfected flies and does not address how infection might alter feeding mechanics or performance. Previous work has shown that trypanosome infection can affect salivary gland function and feeding time (Van Den Abbeele et al 2010), and even cause damage to mouthparts, all of which can influence feeding behavior and efficiency. While this does not detract from the technical quality or the core findings of the study, a more explicit discussion of these biological variables would help place the results in a broader transmissionrelevant context and clarify how generalizable the conclusions are to natural infection settings.

      We thank the reviewer for this important comment. While our study focused on uninfected flies, we agree that parasite infection may influence feeding performance and should therefore be considered when assessing the broader relevance of our findings. Previous studies have shown that trypanosome infections can alter salivary gland physiology and saliva composition (Van Den Abbeele et al., 2010; Matetovici et al., 2016). In addition, transcriptomic analyses suggest that infection with T. congolense may affect the molecular and physiological state of the proboscis (Awuoche et al., 2017). However, there is currently no direct evidence that these changes translate into fundamental alterations of the mechanical properties or function of the mouthparts themselves, which were the primary focus of our study. We have now expanded our manuscript to discuss this (lines 508-520):

      "It is also important to note that our experiments were conducted using uninfected flies. Previous studies have shown that trypanosome infection can alter feeding behaviour, leading to increased probing activity and prolonged feeding times (Jenni et al., 1980; Van den Abbeele et al., 2010). These effects have primarily been attributed to infection-induced changes in saliva composition and the resulting interactions with host blood (Van den Abbeele et al., 2010). However, a different study found no significant effects of infection with either salivary gland-resident T. brucei or with proboscis-colonizing species such as T. congolense and T. vivax on Glossina feeding behaviour (Moloo, 1983). Furthermore, although infection-associated transcriptional changes in the salivary glands and proboscis have been reported (Awuoche et al., 2017; Matetovici et al., 2016), there is currently no direct evidence that trypanosome infection alters the mechanical properties or function of the mouthparts themselves."

      Overall, this is an outstanding and carefully executed study that will have a significant impact on the fields of vector biology and parasite transmission.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an impressively detailed, multidisciplinary analysis of the mechanics of blood feeding in Glossina spp. Combining SEM, CLSM, µCT, FIB-SEM, macro-videography, and quantitative force measurements, the authors characterize the structures and biomechanics of attachment, proboscis deployment, tissue penetration, and blood uptake. They also examine interactions with diverse host-type substrates, from human skin equivalents to cow, deer, and lizard skin, and integrate these with force measurements to quantify penetration and retraction dynamics.

      The work's key conclusion is that the tsetse fly does not rely on any single exceptional morphological innovation, but rather uses a suite of subtle structural features and retractive forces to feed efficiently across diverse hosts. This result is novel, insightful, and evolutionarily compelling. Overall, this is a strong manuscript that combines methodological sophistication with biological relevance. It should be of high interest to researchers studying vector biology, biomechanics, parasite transmission, and vector-host interactions.

      Strengths:

      (1) The combination of SEM, CLSM, µCT, and FIB-SEM provides an unusually comprehensive anatomical characterization of the tsetse feeding apparatus.

      (2) The direct measurement of proboscis penetration and retraction forces across diverse substrates is highly original and fills a major knowledge gap in vector-host interaction mechanics.

      (3) The study bridges morphology, mechanics, behavior, and host tissue properties, which strengthens the overall conclusions.

      (4) Imaging of trypanosomes within the hypopharynx and surrounding tissue during feeding provides new information about parasite delivery mechanisms.

      Main Comments:

      (1) The authors conclude that feeding versatility arises from the sum of subtle adaptations. This interpretation is reasonable, but it would help to sharpen which findings most robustly support this statement. For example, the relative similarity of proboscis forces across skin types is compelling evidence that the proboscis is broadly tuned rather than specialized. The observation that tsetse targets softer interscale regions on lizard skin suggests behavioural selectivity, not morphological specialisation. It would strengthen the discussion to highlight which data most directly refute the hypothesis of a unique specialization.

      We thank the reviewer for this comment. To address this point more explicitly and to sharpen the interpretation of our findings, we have expanded the final conclusion in the Discussion (lines 528544):

      "Ultimately, the objective of this study was to investigate how tsetse flies can feed on a seemingly random selection of animals with highly diverse skin structures. In our detailed anatomical studies and force measurements, we did not identify a single dominant trait that explains the fly's feeding versatility.

      Instead, our results indicate that this capability emerges from the combined effect of multiple, more subtle traits. In particular, the proboscis generates broadly similar penetration forces across a wide range of skin types, suggesting a generalised mechanical mechanism rather than hostspecific optimisation. The intricate architecture of the labellum and the strong retractile forces during probing likely contribute to efficient penetration and blood pool formation across heterogeneous substrates. Behaviourally, tsetse flies further increase feeding success by flexibly targeting mechanically favourable sites, such as the softer interscale regions on lizard skin, rather than relying on specialised morphological adaptations.

      This composite strategy likely reflects evolutionary fine-tuning that enables the broad host range of tsetse flies. By allowing efficient blood feeding across diverse vertebrate hosts, this versatility may also have facilitated the ecological success and transmission opportunities of African trypanosomes.”

      (2) A central finding is that retraction forces exceed penetration forces across substrates, implying that backward pulling is a key component of wound creation. However, the biological interpretation could be deepened. Specifically, do the authors believe retraction serves primarily to enlarge the pool-feeding site? How does this compare mechanically to mosquito fascicle oscillation or other blood-feeding arthropods (especially other flies such as those in the tabanidae family)? Could retraction forces contribute to anchoring or resisting host grooming behaviors?

      The stronger retraction forces observed during probing indeed suggest that backward pulling is not a passive withdrawal, but likely an active component of tissue disruption. As discussed in the manuscript (lines 475–482), we interpret these repeated pullback movements, together with the outward-facing prestomal teeth of the everted labellum, primarily as a mechanism to enlarge the feeding lesion and improve access to blood, consistent with the blood pool feeding strategy of tsetse flies. To make this more clear, we have added a half sentence to line 482 "..., thereby creating a larger blood pool for feeding."

      We also already compare this mechanism to mosquito feeding mechanics in the discussion (starting from line 487). In mosquitoes, high-frequency fascicle oscillations are thought to reduce insertion resistance and facilitate minimally invasive capillary feeding. Although we also observed oscillatory movements during tsetse feeding (Video 4), the underlying mechanical strategy appears fundamentally different. In contrast to the mosquito’s system optimized for delicate penetration, the tsetse proboscis appears adapted for forceful tissue disruption during pool feeding. Notably, the oscillations observed in tsetse flies seem to occur during active blood uptake rather than initial tissue penetration. Consequently, the functional role of these oscillations in tsetse flies remains unclear. We have now addressed this more specifically in the discussion (lines 490-495):

      "Oscillatory movements were also observed during tsetse probing (Video 4). Notably, these oscillations appeared predominantly during active blood uptake rather than during the initial penetration phase, suggesting that they are associated with ingestion rather than insertion. Whether they facilitate blood flow, prevent occlusion of the feeding canal, or simply reflect pump activity remains unknown."

      When looking at other species, stable flies (Stomoxys) may represent a particularly relevant comparison because they employ a similar penetration mechanism and are pool feeders with prominent prestomal teeth (Krenn and Aspöck. Function and evolution of the mouthparts of blood-feeding Arthropoda. Arthropod structure and development, 2012). In contrast, tabanids employ a different mouthpart architecture with rasping/cutting structures but without comparable prestomal teeth. Whereas mosquito mouthparts have been described as functioning like a syringe, and we compare the tsetse proboscis to a saw, tabanid mouthparts have been likened to scissors (Krenn and Aspöck. Function and evolution of the mouthparts of blood-feeding Arthropoda. Arthropod structure and development, 2012). Although tabanids are also known to inflict substantial tissue damage, it remains unclear whether their feeding movements produce retraction-dominated force patterns comparable to those we observed in tsetse flies.

      Lastly, we agree that the elevated resistance generated during retraction may contribute to withstanding host defensive behaviour such as shake-off responses. Structurally, the orientation of the prestomal teeth and the architecture of the everted labellum could provide temporary anchoring during feeding, as we have already briefly discussed in the manuscript (lines 458– 461). However, while stronger anchoring may increase feeding stability, it could also increase the risk of injury to the fly if detected by the host. Compared to other pool-feeding flies such as stable flies, tsetse flies have been reported to respond more readily to host defensive behaviour (Schofield and Torr. A comparison of the feeding behaviour of tsetse and stable flies. Medical and Veterinary Entomology, 2002). We therefore currently consider anchoring to be a possible secondary function but lack direct experimental evidence to assess its practical importance.

      (3) The study analyzes a diverse set of substrates, which is a strength. However, some caveats deserve explicit discussion. Human skin equivalents and dermal equivalents lack the full mechanical complexity of real skin (e.g., innervation, perfusion, tension). Frozen or ethanol-stored samples, particularly reptile skin, may also exhibit altered mechanical properties compared to live tissues. These limitations do not undermine the findings but should be explicitly acknowledged as they influence the interpretation of absolute force magnitudes.

      The reviewer raises a valid point regarding the interpretation of absolute force magnitudes across the measured substrates. We have therefore added a clarifying statement to the discussion (lines 469-476):

      "When interpreting absolute force magnitudes, it is important to bear in mind that our samples do not fully recapitulate physiological conditions. Skin explants and skin equivalents may behave differently to skin under active perfusion and native tissue tension, as may our fixed and frozen animal skin samples. Nevertheless, comparative force measurements revealed consistent biomechanical signatures across substrates, suggesting that the observed force patterns reflect fundamental aspects of the feeding mechanism that are likely relevant in vivo.”

      (4) The SEM and FIB-SEM images showing trypanosomes in the hypopharynx and surrounding tissue during penetration are visually striking and suggest rapid dispersal. It would be helpful to connect these observations more clearly to the kinetics of parasite deposition and whether mechanical tissue laceration is likely to increase inoculation efficiency. Without conducting additional experiments, the authors could discuss whether these findings support or modify existing models of salivary-gland-derived parasite release.

      We have now expanded the Discussion to clarify that our observations of trypanosomes in the hypopharynx are consistent with the established model of salivary-gland-derived parasite release during probing and feeding, in which infective metacyclic trypanosomes are delivered with saliva into the host tissue. Furthermore, the presence of trypanosomes beyond the immediate feeding canal supports rapid parasite dispersal following inoculation, as described in previous work (Reuter et al., 2023). In this context, the tissue laceration generated by the tsetse proboscis may facilitate local parasite distribution by creating a larger, mechanically disrupted feeding lesion. However, our data provide high-resolution structural snapshots and were not designed to quantify deposition kinetics or inoculation efficiency. We therefore refrain from concluding that mechanical laceration increases transmission efficiency and instead view this as a plausible consequence that should be tested directly in future work. Specifically, we have added this paragraph to the discussion (521-527):

      "Overall, our observations of trypanosomes within the fly's hypopharynx, labial gutter, and host tissue are consistent with the established model of salivary-gland-derived parasite release during probing and feeding. Their presence beyond the immediate feeding canal is consistent with rapid local dispersal following inoculation, as described previously (Reuter et al., 2023). This process may be facilitated by the extensive tissue disruption caused by the tsetse mouthparts, although this hypothesis will require direct experimental testing."

      (5) The authors demonstrate that tsetse attachment abilities fall within the range of generalist insects and are far lower than those of obligate ectoparasites. However, the manuscript could discuss how attachment forces relate to the tsetse's ecological context, e.g., whether their attachment is generally brief, whether host shaking strongly selects for grip strength, etc. Is there evidence that other Glossina species or tabanids with different host preferences show variation in attachment performance? This would broaden the relevance of the findings.

      Tsetse flies are obligate blood feeders, but host contact is typically brief and frequently interrupted by host defensive behaviour. As a result, selection may favour rapid and efficient feeding rather than exceptionally strong attachment. This interpretation is supported by Schofield and Torr (A comparison of the feeding behaviour of tsetse and stable flies. Medical and Veterinary Entomology, 2002), showing that tsetse flies experience more feeding interruptions than the stable fly Stomoxys calcitrans, despite completing successful blood meals in less time. These differences are consistent with life-history theory (Anderson and Roitberg. Modelling trade-offs between mortality and fitness associated with persistent blood feeding by mosquitoes. Ecology Letters, 1999), which predicts that long-lived species with low reproductive rates, such as tsetse flies, should be less willing to risk injury by persisting on a host than shorter-lived, more fecund species. Against this background, our finding that tsetse attachment forces fall within the range reported for generalist insects, appears biologically plausible. Their attachment performance needs to be functionally sufficient for brief feeding events rather than maximized for prolonged host retention.

      We are not aware of comparative biomechanical data on attachment performance across different Glossina species or tabanids. We agree that such comparative studies would be valuable to test whether differences in host preference and feeding ecology correlate with variation in attachment capacity.

      (6) In video 4, could the authors clarify whether the observed maxillary vibrations are hypothesized to reduce penetration resistance or serve another function?

      The vibrations of the maxilla specifically appear during active blood uptake rather than during initial tissue penetration, suggesting they are linked to the ingestion phase. Whether they serve a mechanical function, such as facilitating blood flow or preventing canal occlusion, or represent a passive consequence of pump activity, remains unclear. We consider this an open and interesting question that warrants dedicated investigation.

      We have therefore clarified that the functional significance of these oscillations remains unresolved to date (lines 490-495). This reads: “Oscillatory movements were also observed during tsetse probing (Video 4). Notably, these oscillations appeared predominantly during active blood uptake rather than during the initial penetration phase, suggesting that they are associated with ingestion rather than insertion. Whether they facilitate blood flow, prevent occlusion of the feeding canal, or simply reflect pump activity remains unknown.”

      Reviewer #3 (Public review):

      Summary:

      Human and animal trypanosomiasis are fatal illnesses caused by African trypanosomes transmitted by tsetse flies during a bloodmeal. Thus, tsetse fly feeding is the key physical step in disease transmission to mammals. Tsetse fly feeding is not a new story, but it is revisited here through the application of sophisticated imaging techniques and novel biomechanical methods of analysis. The authors aim to provide a high-resolution picture of the structures and forces involved in feeding to provide mechanistic insights into the process of feeding, from attachment, penetration, drinking and retraction of the feeding parts.

      Largely, the authors have achieved their aims. They (i) examine the structures and forces involved in attachment; (ii) they provide detailed multi image analysis of the proboscis providing insights into its probing ability and physical mechanism of penetration; (iii) they conduct a controlled analysis of the physical forces involved in penetration and report that they are in the low nM range, not especially strong but much higher that the mosquito bite and finally they provide a first analysis of blood uptake during feeding.

      Strengths:

      The study images the tsetse fly feeding structures in unprecedented detail, with resolution to the uM scale, in 3-D, and during feeding. The resulting images are dramatic and insightful (and beautiful and frightening!), so researchers interested in trypanosomes, tsetse flies, or blood feeding by flies in general will want to see.

      They conclude that flies attach strongly to smooth surfaces because of interactions possible via the array of acanthae of the pulvillus pad at the ends of the tarsi. The estimated attachment forces are similar in male & female flies, in the low mM range (they look impressively strong in video 1). They provide a very striking analysis of the proboscis and labellum and associated tooth structures (Figures 4 & 5). I recall many years ago observing that tsetse flies are messy feeders, and these structures, especially the rasping teeth structures on the reverse folded labial tips, explain why! This seems more like a chainsaw than a jigsaw in action, but the authors are probably correct that these structures and the probing/retraction mechanism explain many features of tsetse fly feeding and their ability to feed on a wide range of hosts with very different skin types.

      We agree that “jigsaw” may be too specific and not fully appropriate in this context. We have therefore replaced it in the manuscript with the more general term “saw.”

      The impressive aspect of this paper is the range of imaging techniques (CLSM, SEM, uCT, FIB SEM), the quality of the images, which attests to the obvious care taken with sample preparation. The biomechanical analysis, especially the penetration analysis, is impressive. Finally, the paper is clearly written and presented; it was a very easy read and, overall, a very engaging study.

      Weaknesses:

      I suppose it could be said that the paper is a descriptive study; it doesn't really test a hypothesis, but that is not a prerequisite for sharing it. Perhaps the least convincing parts are the imaging of the flexible versus rigid parts of the structures, which is based on the amount of resilin (flexible) and chitin-protein (stiff), based on their autofluorescence. It seems odd that the joints would be less blue (stiffer) in Figure 1i, or what the blue structures correspond to in Figure 6B-D.

      Our analysis is based on established CLSM approaches that use exoskeleton autofluorescence as a proxy for relative differences in cuticular composition and material properties (Michels & Gorb, 2012; Michels et al., 2016). In the tarsus, the observed differences in inferred stiffness are relatively subtle, with most regions exhibiting broadly comparable material properties. This becomes particularly evident when compared with the proboscis, where the contrasts in cuticular composition are much more pronounced (Figure 6). We also note that locally stiffer regions at joints are not unexpected, as stiffness gradients in arthropod joints can provide mechanical support and help constrain the direction of movement. Importantly, our images show a flexible, ring-like blue region directly at the articulation, surrounded by slightly stiffer material. We therefore interpret this pattern as a combination of a flexible hinge region and adjacent supporting structures that together enable controlled joint motion.

      The blue structures in Figure 6B–D correspond to flexible regions of the furca (f). Because this spring-like cuticular element undergoes substantial configuration changes during labellar eversion, the presence of highly flexible regions is consistent with its proposed mechanical function.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further experiments or analyses are suggested. However, the Discussion would benefit from briefly acknowledging how trypanosome infection can alter feeding behavior and mouthpart function, based on prior work, to place the mechanical findings in a more biologically relevant transmission context.

      We thank the reviewer for this suggestion. Previous studies have indeed shown that trypanosome infection can alter tsetse feeding behavior, primarily through changes in saliva composition. Van den Abbeele et al. (2010) demonstrated that infection with T. brucei significantly impairs the anti-haemostatic activity of tsetse saliva, resulting in prolonged prefeeding probing and therefore extended feeding times. These findings are consistent with earlier observations by Jenni (1980), who reported increased probing frequency in infected flies.

      Jenni (1980) also proposed that these behavioral changes might be linked to altered mechanoreceptor function. However, Van den Abbeele et al. (2010) argued against this interpretation for T. brucei, noting that this parasite does not colonize the mouthparts where these mechanoreceptors are located. Taken together, the available evidence suggests that the observed changes in feeding behavior are mediated primarily through altered interactions with host blood rather than through direct effects on the mouthparts themselves.

      It should be noted that this conclusion is specific to T. brucei. Other tsetse-transmitted trypanosome species, such as T. congolense, do colonize the proboscis. However, a comparative study examining flies infected with T. brucei, T. congolense, or T. vivax found no significant effects of infection on feeding behaviour relative to uninfected controls (Moloo, 1983). To our knowledge, there is currently also no direct evidence that any trypanosome species alters the physical properties or mechanical function of the mouthparts, or causes damage that would directly affect feeding performance. We have added this paragraph to the Discussion (lines 508520):

      "It is also important to note that our experiments were conducted using uninfected flies. Previous studies have shown that trypanosome infection can alter feeding behaviour, including increased probing activity and prolonged feeding times (Jenni et al., 1980; Van den Abbeele et al., 2010). These effects have primarily been attributed to infection-induced changes in saliva composition and the resulting interactions with host blood (Van den Abbeele et al., 2010). However, a different study found no significant effects of infection with either salivary gland-resident T. brucei or with proboscis-colonizing species such as T. congolense and T. vivax on Glossina feeding behaviour (Moloo, 1983). Furthermore, although infection-associated transcriptional changes in the salivary glands and proboscis have been reported (Awuoche et al., 2017; Matetovici et al., 2016), there is currently no direct evidence that trypanosome infection alters the mechanical properties or function of the mouthparts themselves."

      Reviewer #2 (Recommendations for the authors):

      Several figures (particularly SEM-based ones) contain very dense labeling. Consider providing simplified overviews or annotated "orientation guides" in figure supplements to improve navigability for readers unfamiliar with proboscis anatomy.

      We thank the reviewer for this helpful suggestion. While we agree that orientation aids can be valuable, we have decided not to include additional simplified overview figures, as we consider that introducing separate schematic summaries could potentially complicate rather than improve navigation of the structural detail. We therefore rely on consistent labelling within the existing figures and detailed captions to guide interpretation.

      The manuscript uses appropriate non-parametric tests, but could benefit from reporting effect sizes and indicating sample sizes on all plots.

      Sample sizes are reported in the figure legends, Methods section, and Supplementary material for all experiments. We agree that reporting effect sizes can be informative and will consider this in future studies. However, because the primary objective of the statistical analyses in the present work was to support comparisons between experimental conditions rather than to estimate effect magnitudes, and because the figures are already information-dense, we therefore decided not to further modify the graphical presentation in this revision.

      Reviewer #3 (Recommendations for the authors):

      (1) P5 L111. Perhaps indicate these knobs on the image Figure 1S). I assume these are the structures visible under the pointer labelled spa? Maybe highlight some of the worn areas in Figure 1G.

      The knob-like structures in Figure S1 are highlighted in green and we have now revised the figure description from:

      “…showing fine crests on the underside and surface modifications (green) on the upper side.”

      to:

      “…showing fine crests on the underside and knob-like surface modifications (green) on the upper side.”

      Regarding Figure 1G, the purpose of the panel is to illustrate the contrast between deformed spatulae (Figure 1G) and intact spatulae (Figure 1H). We therefore chose to retain the original presentation, as we feel that additional markings would not substantially improve interpretation and could obscure structural details. We hope that the direct comparison between the two panels provides sufficient visual guidance.

      (2) P9. The frictional force (and P38/39) has the units of N (kg.m/Sexp2). The safety factor is this force divided by the weight of the fly? So are there units (Kg/sexp2) or are these not shown? Perhaps this is a convention.

      The safety factor is defined as the ratio of the total frictional force to the fly’s weight force (m·g), where m is body mass and g is gravitational acceleration. Since both quantities are express in Newtons (kg·m·s<sup>-2</sup>), the safety factor is dimensionless.

      We agree that the terminology in the original manuscript may have been ambiguous, as “body weight” is sometimes used colloquially to refer to body mass. To avoid confusion, we have revised the text to explicitly refer to weight force and now define the safety factor as the total friction force divided by weight force (mg, where m is body mass and g is gravitational acceleration). We have clarified this in the main text, the Figure 2 legend, and the description of Supplementary Material 1.

      (3) P10 Figure 2G & H. It is not very clear...are these the data, the average of all readings across all surfaces in E and F? If so, why is this value useful...how does it add to what is already shown?

      The figures 2G and 2H summarize the friction forces (G) and safety factors (H) across all tested substrates, based on the values from the male (B, E) and female (C, F) datasets. The purpose of these panels is to provide an overall comparison between sexes independent of substrate type. While this information can also be inferred from the substrate-specific plots, the sex-separated presentation does not make the absence of an overall sex difference immediately obvious. Figures 2G and 2H therefore serve as concise summary plots highlighting this result.

      (4) P12. For the nonspecialist, it might be useful to draw a cartoon showing the organisation of the labium, labrum and the hypopharynx...this is visible in Figure 4i but not in the dissected proboscis and labellum ....only the labium as the labrum doesn't extend this far?

      To clarify the anatomical arrangement in the dissected specimen, we have added the following statement to the Figure 4 legend (lines 224–226):

      “In an intact fly, the labrum would be positioned within the empty groove of the labium visible in J; however, it is absent in this dissected preparation.”

      (5) P17 legend to Figure 5. Include the Lm abbreviation in the legend, and maybe a close-up of the rsp teeth?

      We have added “lm, labellum” to the Figure 5 legend (line 250), as this abbreviation was previously missing. Panel J is a close-up of the rasping teeth.

      (6) F3S and Video 3. Are the images in B and C taken from the FIB SEM video images? It is not clear. A small legend descriptor for video 3 would be helpful.

      The images in Supplementary Figure 3B and C are reconstructed from the same FIB-SEM dataset shown in Video 3, but they are displayed in a different orientation. This is indicated schematically in Supplementary Figure 3A, which illustrates the viewing plane used for the reconstruction.

      We already included the following legend for Video 3 (lines 1146–1149):

      "Video 3: FIB-SEM of the tsetse labellum. Sequential cross sections reveal internal ultrastructure progressing from near the tip of the labellum downward. Data were acquired on a Crossbeam 540 (Zeiss) with the EsB detector in continuous milling mode."

      To improve clarity, we have now added a sentence to the video legend linking the figures to the video: (lines 1149-1150)

      “Reconstructed images from this dataset are shown in Figure 5A and Supplementary Figure 3B and C.”

      In addition, we have now explicitly cross-referenced Video 3 in the legends of Figures 5 and S3 to make the connection clearer for the reader.

      (7) Figure 7. These are amazing images, especially G-I.

      Thank you for this positive feedback, we appreciate it.

      (8) P24. It is really good to see that there is a difference in force penetration for full skin v dermal...this deserves a comment.

      We agree and have revised the text accordingly. We replaced:

      "Human skin substrates required the lowest penetration forces, with 0.97 mN for full-thickness skin equivalents, 0.67 mN for dermal equivalents, and 0.85 mN for native skin explants (Figure 8C, D)."

      With this (lines 363-367):

      "Human skin substrates showed the lowest penetration forces, with dermal equivalents requiring less force (0.67 mN) than full-thickness skin equivalents (0.97 mN), reflecting the additional mechanical resistance of the epidermal layer absent in dermal-only constructs. Native skin explants fell intermediate at 0.85 mN (Figure 8C, D)."

      (9) P26 Figure S5. Panel c, there seems to be a big scatter in the drinking time. Was there an outlier?

      Indeed, the observed scatter is due to a single fly with an unusually long drinking time of 184.44 seconds, which is approximately six times the median duration. We have verified the underlying data and found no indication of a measurement error; the value therefore remains included in the analysis. The data for the plots in Supplementary Figure 5 are also available in Supplementary Material 3.

    1. eLife Assessment

      This important work addresses a very relevant biological question: what is the cellular basis of wound healing? Using the Drosophila pupal notum as a model, the paper provides an elegant, thorough, descriptive characterization of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors meticulously characterize the cell-cell fusion events during wound healing and inhibit cell fusion to show to that it is necessary to speed wound closure. This study provides convincing evidence that cell fusion allows actin resources at be partitioned to the leading edge.

    2. Reviewer #1 (Public Review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental yet poorly understood biological process.

    3. Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laser-induced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFP-positive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to syncytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

    4. Reviewer #3 (Public Review):

      In this revised manuscript, White et al. aimed to understand the wound-induced syncytia formation behavior in wound repair of Drosophila melanogaster pupal notum. For this purpose, the authors characterized two different types of adherens junctions' outcomes during syncytia formation around the wound region - border breakdown versus apical shrinking which appear to happen in different time points and for different time durations. The authors characterized cell-cell fusion events using cytoplasmic, junctional and nuclear markers. They determined that about half of the cells within 70 um radii from the wound undergo cell-cell fusion. They studied wound induction on the border between control epithelia and pnr domain suggesting that Atg1 is required for post-wound syncytia formation and wound closure. They showed that during wound closure syncytia gradually invade the wound leading edge mostly by radial fusion events. The data suggests that intercalation of cells from the leading edge slows down the wound closure process. They propose that cell fluidity of syncytial cells plays a role in wound closure speed. Finally, the authors showed that actin is concentrated to the front edge of syncytia located in the wound leading edge. The authors described some aspects of syncytia formation during wound closure using different approaches.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental yet poorly understood biological process.

      Comments on revised version.

      The manuscript overall is significantly improved and authors addressed majority of my concerns. The addition of the computational vertex model (Figure 7) as well as Atg1 RNAi (Figure 4) to inhibit cell fusion provide more mechanistic insight to their study. However, the analysis of Atg1 RNAi wound assay falls short as it does directly measure changes in syncytium frequency nor size to confirm that cell fusion is reduced. The authors should quantify the number of nuclei per syncytium over the 2hr wound healing period as performed for WT in Figure 1C. It would have been ideal if they could have also performed the Act-GFP spreading assay in WT and Atg1 RNAi strains to determine if Act-GFP movement is dependent on cell fusion as purposed. At the least, further quantification of Atg1 RNAi phenotype is warranted to support their conclusions.

      In response to the reviewer's comment, we have repeated the analysis of Fig 1C and generated a new panel, Fig. 4C, which is directly comparable to the control and shows that syncytial size is dramatically reduced in the Atg1 knockdown area. Unfortunately, we cannot perform the second analysis of actin-GFP spreading in the Atg1 knockdown cells because we need Gal4 for labeling individual cells and for knocking down Atg1, and we can't do both at the same time.

      Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laser-induced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFP-positive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to syncytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Comments on revised version:

      The authors have extended their original manuscript by adding two key parts. First, they show a role of Atg1 in mediating cell fusion (Figure 4). Second, they provide additional evidence for a contribution of radial border fusions to wound closure through its effect on tissue fluidity and through computational modelling (Figure 7).

      This new version of the manuscript is greatly improved and provides significant new insights into the role of syncytia in aiding wound repair. There are just a few minor, yet important, additions needed to back up Figure 4 which should not require new experiments.

      Minor but important points:

      The authors show a role of Atg1 in mediating syncytia formation in Figure 4. However, since the Pnr>+ side of the wound closes slower than the non-Pnr side (control side), a few additions to this figure would be important and should not require additional experiments.

      (1) The authors should show, similar to the data shown in Figure 4D of the wound radius over time for control versus Pnr>Atg1RNAi, also the same type of data for control versus Pnr>+.

      The data the reviewer requests is available in our bioRxiv manuscript, in Fig. 6B (Hua, Krystofiak, Pumford, Page-McCaw, and Hutson, https://doi.org/10.64898/2026.05.31.728998). These experiments were all done at the same time. As you can see, the difference in closure rate is quite subtle in control wounds.

      (2) Since Pnr>+ also slows down wound healing, albeit to a lesser extent than Pnr>Atg1, the authors should also show an extra graph that provides evidence that Pnr>Atg1RNAi reduces syncytia formation more than Pnr>+ does. E.g. Two graphs could be added that show individual cell size at 4 or 5h post wounding for control versus Pnr>Atg1RNAi as well as for control versus Pnr>+ and also another graph with the same data but comparing cell size between Pnr>+ and Pnr>Atg1RNAi. Otherwise, if the expected minimum cell size for a syncytium is easy to estimate, a graph could be added that shows the percentage of cells that are above this threshold (e.g. above 100 square micron) for control versus Pnr>Atg1RNAi and control versus Pnr>+ and Pnr>+ versus Pnr>Atg1RNAi.

      In response to this comment and the comment from reviewer 1, we have now added new Fig. 4C, which addresses the reviewer's question about the comparative frequency of fusion in pnr>Atg1RNAi and pnr>+. These graphs show that Atg1 knockdown significantly reduces the size of syncytia.

      Reviewer #3 (Public Review):

      In this revised manuscript, White et al. aimed to understand the wound-induced syncytia formation behavior in wound repair of Drosophila melanogaster pupal notum. For this purpose, the authors characterized two different types of adherens junctions' outcomes during syncytia formation around the wound region - border breakdown versus apical shrinking which appear to happen in different time points and for different time durations. The authors characterized cell-cell fusion events using cytoplasmic, junctional and nuclear markers. They determined that about half of the cells within 70 um radii from the wound undergo cell-cell fusion. They studied wound induction on the border between control epithelia and pnr domain suggesting that Atg1 is required for post-wound syncytia formation and wound closure. They showed that during wound closure syncytia gradually invade the wound leading edge mostly by radial fusion events. The data suggests that intercalation of cells from the leading edge slows down the wound closure process. They propose that cell fluidity of syncytial cells plays a role in wound closure speed. Finally, the authors showed that actin is concentrated to the front edge of syncytia located in the wound leading edge. The authors described some aspects of syncytia formation during wound closure using different approaches. Some clarifications are needed as described below.

      Major suggestions:

      (1) Introduction, page 4. The examples of developmental syncytia formation of invertebrates and vertebrates are confusing. The authors may want to make the examples clear and add additional examples. Currently, readers may assume that C. elegans cell fusions occur only in the hypodermis - other structures can be mentioned like the vulva, pharyngeal muscles, glia, tail. In addition, the authors may want to add injury-induced fusions like the C. elegans' PLM and PVD neurons (Ghosh-Roy et al., 2010; Newman et al., 2015; Oren-Suissa et al., 2017).

      We appreciate the suggestions and have included the additional examples of C. elegans vulva and PLM and PVD neurons. We are limiting ourselves to those because we don't want to focus too heavily on C. elegans examples, as that's not the direction this paper is heading.

      (2) In cases where it is not clear whether fusion has occurred or whether mononucleated cells were ejected from the leading edge, membrane markers can be used. Page 6. Lines 96-99. The authors may want to use a membrane marker like RFP-PH driven by the epithelial cell promoter.

      At this point in the manuscript, we are introducing syncytia and are not concerned yet with their origin. 

      (3) Pages 8-10. The authors may want to clearly explain that apical junctions shrinking is a post fusion event. That the apical shrinking is caused by the expansion of fusion pores and the migration of apical junctions towards the basolateral domain. This is something that was clearly shown during physiological epidermal cell-cell fusion in C. elegans by Mohler et al., 1998 and 2002. A cartoon showing the process of cell-cell fusion, pore expansion and apical junction dynamics would make the manuscript much clearer.

      Apical shrinking cannot be caused by the "migration of apical junctions towards the basolateral domain" because that is not what we observed -- rather, we observed labeled adherens junctions remaining at the apical surface while the area they enclose becomes smaller (shrinks). Further, despite close reading of the Mohler papers, it is not clear how similar the apical shrinking events of this manuscript are to the fusion events described there. Finally, we do not want to include a schematic describing this process because that would suggest certainty that we do not have. Unlike in C. elegans development, wound-induced cell fusion is stochastic, not stereotyped; with cells that display apical shrinking, the fusion partner of a labeled cell is difficult to identify because it is often not a neighboring cell. These factors make it difficult to describe this process in detail, but we have sufficient data to conclude that these are indeed cell fusion events.

      (4) Page 9. Line 170. "...as these cells represent fusion initiation events (fusion pore) but were unable to productively stabilize and expand the site of fusion and so returned to the diploid state." The authors may want to make clear that this is an assumption that needs to be tested. Live imaging using a membrane marker may resolve whether a reversible fusion pore was generated.

      Thank you for the suggestion; we have updated this text to make it clear that this is an interpretation.

      (5) Page 11. It is not clear whether Atg1 is directly required for cell fusion, or that autophagy is required for efficient cell fusion or both Atg1 and autophagy participate in the fusion process.

      Our data show that Atg1 is required for cell fusion. The work that inspired this experiment, Kakanj et al 2022, concluded from their more comprehensive studies that the process of autophagy was required. We have clarified the text.

      (6) Page 12. Line 235. "Indeed, we observed that several hours after wounding, the entire leading edge was occupied by syncytia." This observation is based only on the adherens junction marker. Can they test basal cell membrane marker? Is it possible that the mononucleate cell in the leading edge is under the two syncytia?

      Unfortunately, there are not good basal markers -- the recently reported basal spot markers also label adherens junctions. Nonetheless, we are confident that the mononuclear cell is not under the syncytia because we image Z-stacks and thus can detect cell overlap.

      Recommendations for the authors:

      Reviewer #3 (Recommendations For The Authors):

      Minor suggestions:

      (1) Figure 1. The authors may want to add an image immediately after laser ablation of the actual wound and the area around the wound. Add an arrow to mark the wound.

      With this wounding modality, the extent of the wound is unclear for ~30 min. As we reported in O'Connor et al, PLoS One, 2021, there is a gradient of damage emanating out from the center of the wound, and cells with greater amounts of damage die while those with less damage repair and survive. Immediately after laser ablation, very little visible damage is evident by 120ctn-RFP and Histone-GFP (the markers in Fig. 1) until the cells die and the surrounding cells respond.

      (2) Page 6. Line 86. "A mitotic tissue utilizes cell-cell fusions during wound repair." replace "during wound repair" with "after wound induction" since in this section the authors do not show that this process is part of wound repair.

      Thank you for the suggestion - we reworded this heading to remove "wound repair".

      (3) Page 6. Line 92. The authors may want to be consistent with the terms used in the text and in the figure - His2GFP in the text versus Histone GFP in the figures.

      Thank you for the suggestion, we have revised for consistency.

      (4) Figure 1 - supplement figure 1D. The "v" of Div panel moved below D.

      Thank you, we have corrected it.

      (5) Figure 1 - supplement figure 1G. add "i" to second Gii to make it Giii.

      Thank you, we have corrected it.

      (6) Page 24. Figure 1H legend. 3 or 4 wounds?

      Thank you for catching this error - 4 wounds.

      (7) Page 7. Line 124. "GFP mixing always preceded border breakdowns (n=11)" instead of "always" use "in all observed cases".

      We have made this change.

      (8) Figure 2. Switch the writing "Apical Shrinking: Nuclear Transfer" since apical shrinking represented in panel 2A and Nuclear Transfer in panel 2B. If this description applies only to panel 2B, make it clear.

      We consider this heading to apply to panels A and B together (as they show the same sample, just different channels).

      (9) Figure 2C. Is ActinGFP a cytoplasmic GFP driven by actin promoter or Actin-bound GFP? Cytoplasmic GFP versus membrane-cortex GFP?

      It is a transgene expressing an actin-GFP fusion protein, as noted in the key reagents table and discussed in Fig. 8. We corrected the manuscript to ensure it is always referred to now as Actin-GFP in the text, figures, and legends.

      (10) Video 3 - Impressive movie!

      Thank you!

      (11) Page 9. Line 155. "In both these cells, as the cell lost its basal volume, cytoplasm moved laterally to join the neighboring syncytia." It seems that the apical shrinking cells' cytoplasm joined the neighboring syncytia even before.

      Because both indicated cells (yellow and white arrows) and the neighboring syncytium are all labeled with GFP, it is not possible to determine precisely when the cells' cytoplasm joined the syncytium.

      (12) Page 9. Line 158. "...but fusions associated with apical shrinking occurred later and were more numerous." Did the fusion occur later or the apical shrinking itself as was mentioned before and shown in Figure 2F?

      We have changed the wording, as for many apical shrinking events we cannot tell exactly when the fusions were initiated.

      (13) Page 25. Figure 3A legend. What is the meaning of morphological fusion? Border breakdown and apical shrinking? The authors may want to define it.

      We have defined it now in the legend.

      (14) Page 26. Figure 3B-C legend. "Panel C shows that apical shrinking fusion and border-breakdown fusion occur at similar distances from the wound." It seems that fusion by apical shrinking mostly occurs within 60-70 um from wound center and fusion by breakdown occurs equally at all distances up to 80 um.

      We don't disagree with your comment, but we feel the dataset is too small to make such a statement. The data is presented so the interested reader can make their own conclusion.

      (15) Page 9. Line 165. "...but infrequently (n=3) with GFP mixing and no subsequent cell fusion..." Does this mean that there were GFP mixing without border breakdown or apical shrinking?

      Yes, that is correct. We assume that in this case a fusion pore opened and then closed again. We have added a phrase to clarify.

      (16) Page 9. Line 175. "...the spatial distribution of fusing cells that shrank vs. lost borders was similar (compare Figures 1G and 2E)." Even though the visual comparison suggests similar spatial distribution, the carefully quantified distribution in figure 3C suggests more fusion by shrinkage at 60-70 um from wound center of the 5 tested wounds.

      As we noted to comment 14, we feel the data set is too small to make such a statement. The data is presented so the interested reader can make their own conclusion.

      (17) Figure 3. The shown pies sum the results from 5 wounds. It would be interesting to add a graph comparing the percentage of fused and persisted cells per wound to see the variability, if exists.

      Unfortunately, the number of fused/persisting cells in each wound is greatly affected by the heat-shock conditions that generate the labeled clones; even the ratio of these fates would be heavily influenced by noise because the numbers are small in each animal. Further, the frequency of fusion is determined by the wound size as shown in Fig. 1. Because of these variables, such data could be easily misinterpreted.

      (18) Figure 3 - figure supplement 1D. Even though it was mentioned that the duration of some border breakdown is finished within minutes it is worth comparing it with shrinking duration on one graph.

      Unlike apical shrinking, it is difficult to identify exactly when border breakdown concludes, so this data is difficult to compare. We have provided several examples of border breakdown in the manuscript that give an overview of the process.

      (19) Video 1 is not mentioned in the main text.

      Thank you for catching that omission. We now refer to it in the first paragraph of the results.

      (20) Figure 4B. The difference between the treated group and the control group is unclear. Add arrows.

      We have added some arrows to Fig. 4B.

      (21) Figure 4C. For consistency use percentage for both border breakdown and shrinking cells.

      In response to the reviewer's comment, we now provide the consistent metric of number of lost borders and number of shrinking cells.

      (22) Page 11. Did the authors try other wound types (e.g. mechanical/chemical wounds)? May other wound causes besides laser ablation result in different response? This may help to answer whether there is a causation between syncytia formation and speed wound closure.

      There are reports of puncture and pinch wounds inducing cell fusion. Perhaps the reviewer is suggesting that we might be able to identify a wounding method that does not induce cell fusion and then compare the rate of wound closure. However, another type of wound would probably inflict different amounts of cell damage and so would be hard to compare. Overall, we think the half-and-half system of comparing responses on the two sides of the wound is the best, most controlled comparison.

      (23) Figure 4F-G. It was mentioned that there is less syncytia formation in Atg KD cells, however the difference in cell area between control, WT and Atg KD is not obvious. The authors may want to mark the dots that represent syncytia to distinguish them from mononucleated cells.

      The point we are trying to make (now Fig. 4G-H) is that cell area is related to distance moved, regardless of how cell area is determined. We do not have the ability to count nuclei in the control sides (nuclei are labeled only on the pnr side), and further, we have reported separately (White et al, 2024) that there is a limited amount of endocycling in these cells, which should also increase area.

      (24) Figure 5G. y axis. The authors may want to change "small cells" to "mononucleate cells".

      We changed it to "unfused cells" which is the term we used in the legend. In these wounds we were unable to visualize nuclei.

      (25) Page 12. Line 243. (Figure 5D,G) instead (Figure 5D).

      We changed it to read (Figure 5D,G).

      (26) Page 12-13. Lines 241-246. The description of "mononuclear cells removed" and "syncytia outcompete unfused cells" may be clearer if explained here as mononuclear cells joining the syncytium by cell-cell fusion.

      Here we are describing a different phenomenon - not that fusion is removing all the smaller cells but rather that the syncytia are faster/better/more effective at wound closure than the smaller cells. This is illustrated in Fig. 5Cii-Ciii.

      (27) Page 13. Line 261. "Thus, there were about five-fold more tangential borders lost to fusion than radial" Is this conclusion also true when analyzing each wound individually?

      This is a reproducible finding, that there is more fusion across tangential borders than across radial borders. The ratio of tangential-border loss: radial-border loss for each wound is as follows:

      wound 1, 63:12

      wound 2, 39: 11

      wound 3, 44:8

      wound 4, 50:8

      (28) Page 31. Figure 6 - figure supplement 1 legend, Line 652. Make "B)" bold.

      Done.

      (29) Page 14. Lines 270-275. If there is an advantage to radial fusion versus cell intercalation for wound closure speed, how do the authors explain that the percentage of radial fusion is lower than the percentage of intercalation? (Figure 6D) How does the wound affect the molecular level (fusogen expression?) of the surrounding cells? 

      We expect that radial fusion specifically reduces the need for intercalation at the leading edge, as shown in Fig. 6C. Both would speed closure, however, as any increase in cell area will allow more efficient redistribution of resources such as actin and will also reduce the total number of junctions needing to be remodeled as the wound closes. Since we don't know the fusogen, we can't say how the wound affects its distribution.

      (30) Page 14. It is not clear where the experimental data ends and the model starts. For example, in line 276, it would be clearer to describe the "tissue fluidity as measured" or is it more precise to write instead "as estimated/calculated". The fusion between observations and model is confusing and maybe this should be unfused.

      This text, referring to the analysis in Fig. 7A, B, is not a computational model but rather a quantitative analysis of tissue fluidity as measured by a pre-existing metric, the shape index. This is experimental data. The computational model begins in the next paragraph, accompanying Fig. 7C, D. We edited the language slightly in this paragraph to clarify.

      (31) Figure 8 versus Figure 2C. Actin-bound GFP versus cytoplasmic GFP? Both mentioned as Actin GFP. Make it clear.

      They are indeed the same thing, actin protein fused to GFP, as described in the text and legend, and they are labeled identically.

      (32) Figure 8Biii, Div. Nice presentation of signal distribution between the cells.

      Thank you!

      (33) Page 30. Figure 8G legend. Lines 623-628. Is the shown mean profile plot based on specific images shown in Fi and Fv or just the cells represented there? Since Fi is a single z slice and Fv is maximum intensity projection which are not comparable.

      In response to the reviewer's question, we reanalyzed the image. Fig. 8G compares Z-projections.

      (34) Page 15. Line 303-304. "Tangential border fusions allow resources from distant cells to be mobilized to the wound edge." Does not this leading-edge actin localization happen in radial fusions close to the region of the wound?

      Fusions along radial borders, as shown in the top panel of Fig. 6A, would not offer the opportunity to move actin from distant cells to cells nearer to the wound.

      (35) Page 15. Did the authors test any predictions from the simulations of the model experimentally?

      This isn’t so much a predictive model as an exploratory model that addresses one question: is it plausible that the presence of syncytia can speed closure by reducing the need for intercalations, even if the syncytia have no other special properties. The only prediction would be that inhibiting fusion would slow wound closure.

      (36) Page 17. It would be interesting to discuss the following questions: (A) Is autophagy required for fusion. (B) Is Atg1 required for epithelial cell fusion? (C) Is autophagy required for wound repair? Are any of the combinations correct (A&B, A&C, B&C, A&B&C)

      The role of autophagy in wound-induced cell fusion was thoroughly explored in the 2022 EMBO J paper from Maria Leptin's lab, "Autophagy-mediated plasma membrane removal promotes the formation of epithelial syncytia" by Kakanj et al. We merely knockdown a gene they discovered to be important for wound-induced epithelial fusion, Atg1, as one means of investigating how syncytia contribute to wound closure. Our results don't add to their findings, and the role of autophagy is not what we want to focus on in our Discussion.

      (37) Page 19. Line 381-384. "If N represents the number of cells that fused, our results suggests that syncytia can apply up to N times more actin to the leading edge; considering that we observed syncytia with dozens of nuclei, this could represent a significant enhancement of actin at the leading edge. Increased actin might explain the ability of syncytia to outcompete diploid cells at the leading edge." To enhance this suggestion, the authors may want to compare actin signal in the leading edge of different size syncytia.

      We thought a lot about this experiment because reviewer 2 asked for it in the previous round of review, but as we said then, we can imagine too many caveats to the interpretation to make it worthwhile.

      (38) Page 22. Line 447-448. "...fusion would act the fastest after wounding because there is no need for DNA replication." There may be a potential need for protein (fusogen) synthesis.

      The timing of fusion, which we report here begins within 10 minutes after wounding, suggests that if there is a fusogen, it is already present in the cells before wounding.

      (39) Page 34. Line 716. Add "C" to "29{degree sign}".

      Done.

      (40) Page 38. "Wound closure analysis" part. Can the wound closure be visualized using brightfield?

      The scar also impedes imaging through bright-field microscopy.

    1. eLife Assessment

      This valuable study addresses the effects of selection for aggression on fitness and life-history trade-offs in Drosophila melanogaster. The evidence presented is overall solid, however, the data as they are do not completely support the claims of increased survival of highly aggressive males at the expense of reproductive success. The main limitation is the choice to use males from only one aggressive Drosophila line, that do not allow disambiguation between nonaggression-related factors and aggression-related factors influencing lifespan.

    2. Reviewer #1 (Public review):

      Summary:

      This study asks how selection for male aggressiveness affects life-history and reproductive fitness traits in Drosophila melanogaster males.

      Strengths:

      Multiple comprehensive assays are used to address the question.

      Weaknesses:

      (1) The flies used for comparisons are inadequate. Behavioral assays compare Bully males mated top non-coevolved Cs females with Cs males mated to coevolved Cs females.

      (2) Lifespan analysis is done on male progeny of Cs females mated to either genetically more distant Bully or co-evolved Cs males, the longer lifespan and performance on the former is interpreted as trade-off with aggressiveness, rather than a simple explanation of hybrid vigor.

      (3) Differences in CHCs between Bully and Cs males and Cs females mated to those males are not shown to cause difference in measured behavioral outcomes.

      Comments on revised version.

      I appreciate authors responding to reviewer's comments. The inclusion of additional Bully lines in behavioral analysis, and Bully homozygous male progeny in lifespan analysis gives more strength to the authors' conclusions. It does not exclude other possible explanations for the observed results, but now authors note genetic drift as an alternative explanation for some of their results.

      I do want to point to a potential misunderstanding of male-female co-evolution by authors. The authors state that "The Bully lines used in our work were derived from Canton-S flies and thus did co-evolve with Cs". This statement is incorrect if the process of selection and line maintenance in this study was the following:

      In my understanding to create Bully lines the most aggressive males were first chosen from an ancestral Cs line and their most aggressive male progeny were mated to their sibling females, repeating the process for 37 generations. Therefore, Bully females were co-evolving with Bully males during selection process of over 37 generations, while Cs females were staying co-evolved with their own males, since they mated within the line. Moreover, after aggressive lines were created, they were kept separate from each other, and from Cs line since about the year 2010, until the experiments described in the paper were performed (which must over 10 years?). Over 10 years, a significant genetic drift can happen, that changes allele frequencies, and may results in differences in male-female co-evolved traits and in lifespan that are unrelated to selection for aggression.

      Also, decapitating females does not completely prevent female influence over mating process, but just removes central brain control over it. In Drosophila, however, the main control over copulation process for males and female is not central. Therefore, you do not completely remove the effect of coevolved or non-coevolved female traits over copulatory and post-copulatory processes.

    3. Reviewer #2 (Public review):

      Summary:

      The authors compare "Bully" lines, selected for male aggression, to Canton-S controls and find that Bully males have lower mating success, shorter mating durations, and remate sooner. Chemical analyses show Bully males have distinct cuticular hydrocarbon (CHC) signatures and transfer markedly less cVA to females, offering a plausible mechanistic link to weaker mate-guarding. Paradoxically, Bully males live longer and remain fertile at older ages when Cs males no longer mate, indicating a shift in the reproduction-survival trade-off in aggression-selected populations. Importantly, the work sheds light on proximate mechanisms, demonstrating that shifts in CHCs and pheromone transfer co-occur with changes in fitness traits.

      Strengths:

      The manuscript's strengths lie in its comprehensive and integrative approach framed within an evolutionary context. By combining behavioral assays, chemical profiling, and lifespan measurements, the authors reveal a coherent pattern linking aggression selection to life-history trade-offs. The direct quantification of cVA in the female reproductive tract after mating provides a particularly compelling mechanistic correlate, strengthening the link between behavior and chemical signaling. Findings on altered 5-T and 5-P levels further highlight how chemical communication shapes mating and mate-guarding strategies. Analytical approaches are largely rigorous, and the results provide valuable insights into the pleiotropic effects of selection on socially relevant traits.

      The revision responds directly to the main concerns raised previously. The addition of a third, independently selected line (Bully C), together with the Bully × Bully data, considerably reduces the concern that the behavioral phenotypes reflect line-specific drift or founder effects rather than a correlated response to selection. The reorganized survival figure (Figure 5) is a clear improvement over the previous version, with isolated and group-housed males in separate panels and a heterozygous Bully condition added, so the longevity claim can be evaluated more directly. The isolated-male data are especially useful here, since those flies never mate, and a longevity difference under that condition argues that the effect is not simply a consequence of Bully males mating less often. The behavioral schematic, corrected symbols, and reported sample sizes also help, as does the reinterpretation of the post-mating courtship data in terms of courtship motivation rather than a refractory-period effect once no latency difference was found.

      Weaknesses:

      Most of the remaining weaknesses are ones I raised in the first round, and the revision has narrowed them. The links between the altered CHC profiles, the reduced cVA transfer, and the behavioral outcomes remain correlative. The causal experiments that would establish them (for example, perfuming or cVA-equalization) are acknowledged by the authors as future directions, which is reasonable, but it means the mechanistic claims should be read as candidate explanations rather than demonstrated ones. It is also worth noting that the CHC differences and the behavioral differences may both be downstream of a common selection target (for instance, genes affecting oenocyte function or CHC biosynthesis) rather than one causing the other; the Discussion would be more balanced if this alternative were stated explicitly.

      My main remaining concern is with the lifespan data. The behavioral phenotypes are replicated across Bully A, B, and C, but the survival assays were done on Bully A only, so a line-specific contribution to the longevity result, including drift, cannot be excluded, even though this has been addressed for the behavioral traits. This matters because the title and abstract present the survival-reproduction trade-off as a general consequence of selection for aggression, whereas the survival evidence rests on a single line. The authors can either run the lifespan assays on a second line, or calibrate the text, title, and abstract so that the strength of the survival claim matches the single-line evidence behind it, with second-line lifespan data noted as a future step.

      The Bully C line is currently underused. Its intermediate aggression, together with the absence of a significant reduction in mating duration, points to a graded rather than binary relationship between aggression intensity and mating duration. This is one of the more interesting features of the expanded dataset, and it deserves more than its present role as a justification for focusing on Bully A.

      The authors have appropriately softened causal language in the title, subheadings, and much of the Discussion. A few residual passages still imply causation or directional transfer and would benefit from the same treatment.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This valuable study addresses the effects of selection on aggression on fitness and life-history trade-offs in Drosophila melanogaster. However, the evidence presented is incomplete and does not support the claims proposed in the study of increased survival of highly aggressive males at the expense of reproductive success and shorter mating duration. The main limitation of the study is the choice to use males from only one aggressive Drosophila line in combination with Canton-S females, that do not allow disambiguation between nonaggression-related factors, such as hybrid vigor and aggression-related factors influencing mating and lifespan.

      We would like to clarify the points raised in the eLife assessment.

      The report states that we relied on a single line of hyper-aggressive males tested with Canton-S females, and implies that Bully and Cs have not co-evolved. This is a misunderstanding: Bully flies were derived from Cs population. Thus, Bully and Cs have co-evolved. In addition to the Bully A line presented in the main figures of the manuscript, we replicated several of our findings with a second independent selected line, Bully B. Results from courtship assays involving both Bully A and Bully B couples males and females were presented in Figure Supp1. We apologies for not having made this more explicit in the original manuscript, which we will correct. These experiments should alleviate the concerns from the reviewers; they demonstrate that our conclusions are supported by two independent hyper-aggressive lines, and these include assays with selected male and female flies.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study asks how selection for male aggressiveness affects life-history and reproductive fitness traits in Drosophila melanogaster males.

      Strengths:

      Multiple comprehensive assays are used to address the question.

      We thank the reviewer for recognizing these strengths.

      Weaknesses:

      (1) The flies used for comparisons are inadequate. Behavioral assays compare Bully males mated to non-coevolved Cs females with Cs males mated to coevolved Cs females.

      We thank the reviewer for this comment, which made us realize that we had not sufficiently highlighted some of our experiments. The Bully lines used in our work were derived from Canton-S flies and thus did co-evolve with Cs. As originally described by Penn et al. (2010), highly aggressive “Bully” lines were generated through selective breeding from Canton-S males that consistently won aggressive encounters. After 34–37 generations, stable Bully lines were established. Thus, 1) Bully and Cs flies have co-evolved and 2) the selection applied was male-specific. Independent selection replicates produced distinct lines, including Bully A and Bully B. Previous studies only characterized Bully A (Penn et al., 2010; Chowdhury et al., 2017), but our work includes both Bully A and Bully B (Fig. S1).

      The rationale for pairing Bully or Cs males with Cs females (with which both male types co-evolved) follows the approach used by Dierick et al. (2006), who investigated how the male-specific selection for aggression affected courtship and mating behaviors by testing them with standard Canton-S females. This design allows to isolate the effects of male genotype and behavior on courtship and mating outcomes, avoiding confounding effects from female behavioral changes.

      We initially compared selected Bully pairs (Bully males × Bully females) (Fig. S1) with Cs pairs and observed similarly shortened mating durations in both Bully × Bully and Bully × Cs matings (Fig. S1, Fig. 1F and G). Thus, the reduction in mating duration arises specifically from Bully males. We therefore chose to use Cs females as a standard background to assess the consequences of male-specific selection for aggression on reproductive behaviors.

      (2) Lifespan analysis is done on male progeny of Cs females mated to either genetically more distant Bully or co-evolved Cs males; the longer lifespan and performance on the former is interpreted as a trade-off with aggressiveness, rather than a simple explanation of hybrid vigor.

      We appreciate this comment, which again stems from a poor explanation from our part about the origin of the Bully line in the original manuscript. The Bully flies were derived from the same original population as the Cs line. Hybrid vigor typically arises when crossing individuals from distinct populations, which is not the case here as both Bully and CS come from the same population.

      To further support our conclusions, we conducted additional experiments using progeny from within-line crosses (Bully males × Bully females) and results revealed the same phenotype: the progeny of these flies also exhibited significantly longer lifespans than Cs males x Cs females progeny. This finding argues against hybrid vigor as the main explanation for the observed phenotype, since both the Bully and Cs crosses result in inbreeding, yet give longer lifespan in Bully. We will include these additional longevity data (currently not included in the manuscript) to strengthen our results and reinforce our interpretation.

      (3) Differences in CHCs between Bully and Cs males and Cs females mated to those males are not shown to cause differences in measured behavioral outcomes.

      We thank the reviewer for raising this important point regarding causality. One way to establish a causal link between differences in CHCs observed in Bully and Cs flies and the corresponding behavioral outcomes would be to experimentally manipulate CHC profiles. For instance, one could perfume oenocyte-less males with the compounds found in higher abundance in Bully flies, then perform behavioral assays to assess causality. We agree that such experiments would be highly informative in determining the functional roles of specific CHCs elevated in Bully males. However, this approach is technically challenging, as the perfuming technique must be optimized to transfer precise amounts of each compound. For example, this method can be used to gradually perfume flies to assess dose–response behavioral effects, whereas matching exactly the natural concentrations found in individuals, especially given inter-individual variability, remains difficult.

      We considered conducting such experiments during our study but did not pursue them for these technical reasons. Nevertheless, we can include a statement in the Discussion acknowledging this as an important future direction to test the causal relationship between CHC variation and behavior.

      Reviewer #2 (Public review):

      Summary:

      The authors compare "Bully" lines, selected for male aggression, to Canton-S controls and find that Bully males have lower mating success, shorter mating durations, and remate sooner. Chemical analyses show Bully males have distinct cuticular hydrocarbons (CHC) signatures and transfer markedly less cVA to females, offering a plausible mechanistic link to weaker mate-guarding.

      Paradoxically, Bully males live longer and remain fertile at older ages when CS males no longer mate, indicating a shift in the reproduction-survival trade-off in aggression-selected populations.

      Importantly, the work sheds light on proximate mechanisms, demonstrating that shifts in CHCs and pheromone transfer co-occur with changes in fitness traits, thus offering new entry points for understanding life-history evolution.

      We thank the reviewer for this positive summary of our work.

      Strengths:

      The manuscript's strengths lie in its comprehensive and integrative approach framed within an evolutionary context. By combining behavioral assays, chemical profiling, and lifespan measurements, the authors reveal a coherent pattern linking aggression selection to life-history trade-offs. The direct quantification of cVA in female reproductive tracts after mating provides a particularly compelling mechanistic correlate, strengthening the link between behavior and chemical signaling. Findings on altered 5-T and 5-P levels further highlight how chemical communication shapes mating and mate-guarding strategies. Analytical approaches are largely rigorous, and the results provide valuable insights into the pleiotropic effects of selection on socially relevant traits. The study will be of interest to Drosophila biologists working on sexual selection, behavioral evolution, and aging.

      We thank the reviewer for recognizing the integrative design and mechanistic contributions of our study.

      Weaknesses:

      The weaknesses are primarily conceptual rather than procedural. The generality of the findings is uncertain, as selection appears to be represented by only one (and a second closely related) Bully line, limiting conclusions about selection responses versus line-specific drift or founder effects. The causal link between aggression selection and increased longevity is not established: the data show a correlated shift but do not identify mechanisms underlying lifespan extension. In several places, the manuscript uses causal language (e.g., that selection 'influences' longevity or mating strategy) where association would be more accurate; this should be toned down to avoid overstatement. Ecological relevance is also not addressed, since laboratory conditions may bias the balance between costs and benefits of aggression compared with variable natural environments. Addressing these points would strengthen both the impact and clarity of the study.

      (1) Generality of findings and potential line effects

      We agree that our results presented in the main figures of the manuscript relied mainly on one Bully line (Bully A). To address potential line-specific effects, we replicated key courtship experiments with another independent line, Bully B, selected in parallel from the same Canton-S stock but through distinct selection replicates. The results obtained from Bully B closely matched those from Bully A, suggesting that the observed phenotypes are consistent consequences of aggression selection rather than random drift or founder effects.

      (2) Causality versus correlation

      We concur that some sentences in the manuscript could overstate causal interpretations. We will revise the text to clearly distinguish correlation from causation and to avoid implying direct causal relationships where data only support association.

      (3) Ecological relevance

      We appreciate this point. Our experiments were performed under controlled laboratory conditions, which may not fully capture the ecological contexts shaping the costs and benefits of aggression. We will acknowledge this limitation and expand the Discussion to consider how environmental variability could modulate the fitness trade-offs associated with aggression in natural populations.

      We thank both reviewers for their constructive feedback, which will help us strengthen the rigor and clarity of the manuscript. We believe that the additional results and revisions will satisfactorily address their concerns.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The major weaknesses raised by the reviewers, namely the flies used in the study (CsxBully compared to CsxCs) where the effect on lifespan could be explained by hybrid vigor and the use of only one Bully and one Cs line that does not allow to link unambiguously the observed effect to the selection for aggression, should be addressed by a different experimental design and additional lines to exclude the effect of non-aggression related factors.

      We thank the Reviewing Editor for these comments.

      (i) Experimental design and hybrid vigor:

      Hybrid vigor typically arises from crosses between genetically divergent populations. In our study, Bully lines were derived from Canton-S background and are not therefore not genetically distant from controls. To directly address this concern, we included new data from Bully × Bully pairs (Figure 1), using independently selected Bully lines. These experiments reproduce the key aggression and courtship phenotypes observed in Cs × Bully assays, indicating that the effects are not attributable to hybrid vigor.

      (ii) Use of additional selected lines:

      We now include data from two independently selected lines (Bully A and Bully B), both derived from Cs, which show consistent behavioral phenotypes. This supports the conclusion that the observed effects are associated with selection for aggression rather than line-specific artifacts. We note that generating such lines is time- and labor-intensive, and only a few laboratories have established aggression-selected lines in Drosophila melanogaster (e.g., Penn et al., 2010; Dierick et al., 2006; Edwards et al., 2006). Accordingly, we have revised the manuscript to explicitly acknowledge this limitation and to frame our conclusions in terms of association rather than causation.

      Reviewer #1 (Recommendations for the authors):

      I can't see any way to interpret the data using CsxCs vs CsxBully comparisons.

      We thank the reviewer for this important point. This concern appears to arise from the assumption that Bully and Cs represent genetically distinct or non-coevolved populations. However, Bully lines were directly derived from Cs and therefore share a common genetic background. We have clarified this point in the Introduction (lines 104-106) and Results (lines 132-137).

      Importantly, we now include additional data showing that key phenotypes, including reduced mating duration, are also observed in Bully × Bully pairings and across independently selected Bully lines (new Figure 1). These results demonstrate that the observed effects are driven by the male genotype and do not depend on the female background.

      Because selection for aggression was applied specifically to males, we used Cs females as a standardized background to isolate male-specific effects while minimizing variability arising from female genotype or behavior. This rationale is now explicitly stated in the Results (lines 163-166) and at the beginning of the Discussion (lines 326-330). This experimental design allows interpretation of male-specific effects, and the observed differences cannot be attributed to cross design artifacts or hybrid vigor.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) Several passages currently imply causality, whereas the data support correlations between selection and trait differences rather than direct causation. This overstatement also appears in section subheadings within the Results, such as "Hyper-aggressive males display reduced mate-guarding efficiency, without compromising female fertility." Please consider toning down the wording by replacing active causal verbs with more neutral phrasing. Additionally, it would be important to include a clear, explicit sentence in the Discussion acknowledging this caveat, as the existing phrase "is associated with changes in reproductive traits" does not fully convey this nuance.

      We thank the reviewer for this important comment. We have revised the manuscript throughout, including Results subheadings, to replace causal language with association-based phrasing. We also rephrase the first sentence of the Discussion to clarify that our conclusions are correlational (see line 320).

      (2) Figure 1 would benefit from a simple schematic of the behavioral paradigm and the arena, since the authors' arena design minimizes manual handling; a cartoon would help readers quickly grasp the assay flow and the conditions under which interactions occur.

      We thank the reviewer for this helpful suggestion. We have added a schematic to Figure 1 illustrating the behavioral paradigm and arena design. Additional details are provided in the Materials and Methods (Trannoy et al., 2015). This improves clarity and accessibility of the experimental design.

      (3) In Figures 2A-B and A'-B', the higher post-mating UWE in Bully males is intriguing, but these panels do not actually measure the refractory period. It would be helpful to include 'latency' in the first UWE after mating in these swapped-female conditions. This could also be repeated with pheromone-standardized (cVA/CHC-equalized) decapitated females to disentangle effects of female pheromone load from male sensory perception. In addition, a baseline courtship control (naive males with decapitated virgins) is necessary to test whether Bully males simply have a lower threshold for initiating courtship.

      We thank the reviewer for this suggestion. The referenced panels are now shown in Figure 3. We quantified post-mating courtship latency; however, latencies were very short across conditions, and no differences were observed between genotypes. We therefore revised the text to interpret these results in terms of post-mating courtship motivation rather than refractory period. The baseline courtship control with decapitated virgins is provided in Fig 2G. These changes clarify the interpretation of post-mating behavior and address the reviewer’s concerns.

      (4) Related to my above point, the results in Figure 2B-B′ raise the possibility that Bully males have reduced perception or neural sensitivity to anti-aphrodisiac pheromones deposited by CS males, which could account for their elevated post-mating courtship; the authors might consider experiments that directly test male sensory responsiveness to these cues or mention this possibility in the Discussion.

      We thank the reviewer for this point. We performed additional assays to test males’ sensory responsiveness using binary choice assays and measured the time spent performing UWE towards decapitated females versus males. These results were added in Figure 3-Figure Supp 1, and indicate that both Cs and Bully males displayed courtship preferentially towards females, providing a control for sensory perception.

      (5) In multiple figure panels, virgin and mated females are depicted with the same symbols, which makes interpretation confusing.

      Thank you for pointing this out. We have updated the figure panels to use distinct symbols for virgin and mated females to improve clarity.

      (6) For Figure 4, it would be helpful to provide standalone KM curves for Bully versus CS males, including a separate panel for isolated (never-mated) males, and present mating counts in a separate panel while reporting survival models that incorporate mating frequency (or use it as a time-dependent covariate). Although Figures 4C-D report median survivals, full KM plots and an isolated-male curve are important since mating itself elevates mortality and can otherwise confound intrinsic lifespan differences.

      Thank you for this important point to improve clarity of this figure. We have reorganized this figure (now Figure 5) to now, present survival curves first (isolated and group-housed males), followed by lifetime mating counts. This reorganization separates survival from mating activity and addresses the potential confounding effect of mating on lifespan. For the lifetime mating data, we used bar plots rather than curves to better visualize individual mating events.

      (7) The following sentences overstate the results and imply causality; consider toning down: "These findings suggest that 5-P and 5-T might contribute to promoting remating in females that have previously mated with Bully males (Figure 2F' and G'). Given that Bully males also showed higher levels of both 5-P and 5-T compared to naïve Cs males (Figure 3B), it is likely that the elevated levels of these compounds observed in females result from their transfer during mating."

      We thank the reviewer for this important point. We have rephrased these sentences to remove causal language and instead describe associations between CHC profiles and behavioral outcomes. In particular, statements implying that 5-P and 5-T promote remating or are directly transferred during mating have been revised to reflect correlational evidence only (see lines 253-256).

      (8) It is not entirely clear how aggression was quantified in each generation, what proportion of males were selected to breed, and whether the findings generalize beyond a single Bully line (Figure Supplement 1 shows data from a closely related Bully line). Without independent replicate lines or sham-selected controls, it remains difficult to rule out drift or line-specific artifacts, and this limitation should be explicitly acknowledged.

      We thank the reviewer for this important point. We have clarified the aggression selection procedure by adding methodological details from Penn et al., including how aggression was quantified and how breeders were selected (lines 104-106 and 131-137). Briefly, independent selection replicates were initiated from the same Canton-S population, generating three lines (Bully A, B, and C), which were maintained separately.

      To address generality, we now include data from multiple lines. In particular, a new Figure 1 presents aggression and courtship phenotypes across Bully A, B, and C, and key behavioral results are consistent across independent lines.

      We acknowledge that additional independent lines would further strengthen generality; this limitation is now explicitly stated in the Discussion (lines 325-326).

      These additions clarify the selection procedure and support that the observed phenotypes are associated with aggression selection rather than line-specific artifacts.

      Minor Comments:

      (1) Exact sample sizes for every experiment should be included in the main figure legends.

      Thank you. We have added the number of replicates in each figure legends.

      (2) In Figure 1-Supplement 1, the orientation for depicting mating success is reversed compared to Figure 1, which is a bit jarring; it would be clearer to keep the orientation consistent with the main figure.

      Thank you. We have incorporated the results initially presented in Figure 1-Sup 1 into a new Figure 1 with additional results, and have taken into account reviewers’ comment.

      (3) For multivariate analyses, I suggest including important details such as group sample sizes, p-value, and the percent variance, etc., in the figure legend rather than keeping this only in Supplementary Table S1.

      We have inserted these details directly into the figure legends for clarity.

      (4) Why was the food cup used for arenas where decapitated virgins were used in mating assays?

      Thank you for pointing this. We now have clarified the experimental procedure in the M&M of the revised manuscript (lines 465-469).

      (5) For cartoons in Figure 2, the current yellow background makes it very difficult to distinguish flies drawn in yellow or green. Please adjust to a higher-contrast background or add darker outlines so that the cartoons are clearly legible.

      Thank you. We have increased the contrast of the female bodies to ensure the cartoons are clearly distinguishable (now figure 3).

      (6) Addition of line numbers in the manuscript would be helpful during the review process.

      Line numbers have been added throughout the manuscript.

      (7) I noticed a few typos in the manuscript. For example, in the Introduction, "seminal fuids" should be corrected to "seminal fluid." In the Discussion, the phrase "CHCs profiles compared those" requires a "to" before "those." Please carefully review the manuscript for similar errors.

      Thank you for pointing this out. We carefully reviewed the manuscript for typos and corrected all identified errors.

    1. eLife Assessment

      This potentially valuable study investigates the anti-senescence effects of red light exposure, proposing that reduced SIRT4 levels enhance fatty acid metabolism and H3K9ac, thereby attenuating ageing-related phenotypes. The authors use multiple approaches, including cultured cells, animal models, and molecular analyses, to support their conclusions. Following revision, many of the concerns previously raised by the reviewers have been addressed. The evidence is solid, whereas additional controls and stronger mechanistic data are still needed to fully substantiate the proposed pathway, particularly the mechanism by which red light exposure leads to reduced SIRT4 levels.

    2. Reviewer #1 (Public review):

      Summary:

      Deng and colleagues pursue the possibility that red light exposure can provide some benefits and anti-senescence effects in aged mouse models. In addition, they show how red light influence metabolism in cultured keratinocytes. The authors provide a long dissection of the potential paths involved in the changes promoted by red light exposure, identifying CytC oxidase, SIRT4, PPARa and MCD as key players.

      Strengths:

      The authors did a thorough exploration of the multiple potential avenues by which red light exposure influence metabolism. The in vitro and in vivo evidence nicely complement each other.

      Weaknesses:

      This is a challenging hypothesis that would require some additional experimental controls. The pathway dissection, while extensive, sometimes is approach in unconvincing ways and the results are not always evident to judge or interpret. Technically, the western blots and transcriptomic analyses require notable improvements.

      Comments on revised version.

      The revised version of the manuscript provides some improvements. However, I feel that many aspects remain poorly addressed. In the authors' favour, many of these limitations are now acknowledged in their rebuttal, as well as in the discussion section.

    3. Reviewer #2 (Public review):

      Summary:

      This work identifies a previously unknown way that red light can slow ageing. The authors show that red light lowers the level of a protein called SIRT4 in skin cells. Reducing SIRT4 boosts fatty acid use and increases a type of histone modification that keeps genes active. These changes help cells clear away signs of ageing, reduce inflammation, and restore normal metabolism. The findings open the possibility of developing new treatments that target SIRT4 to reverse age‑related decline.

      Strengths:

      The evidence is solid because the authors use several complementary methods. They test red light in both cultured cells and naturally aged mice, and they confirm the key role of SIRT4 by silencing its gene. Measurements of metabolism, protein changes, and ageing markers all point in the same direction. However, the exact way red light lowers SIRT4 levels is not fully explained, which leaves a minor gap. Overall, the conclusions are well supported and convincing.

      Weaknesses:

      The paper does not evolve to use the mechanistic discoveries of the manuscript to help our community to identify the mechanism of photobiomodulation, which is not known so far.

      I would like to draw your attention to a recently published paper by Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), which shows that red light (660 nm) stimulates mitochondrial fatty acid oxidation in keratinocytes via AMPK‑dependent phosphorylation of ACC, without altering expression of electron transport chain complexes. I believe this paper is highly complementary to current study.

      Herrera et al. demonstrate that red light increases basal, ATP‑linked, and maximal oxygen consumption rates in keratinocytes specifically through enhanced fatty acid oxidation (inhibited by etomoxir). This independently validates the central finding of the current manuscript ,i.e., red light boosts lipid metabolism, strengthening the robustness of this concept.

      While the current manuscript focusses on the SIRT4‑MCD axis, Herrera et al. identify AMPK phosphorylation and ACC inhibition as key effectors. Authors can integrate and expand their discussion, since SIRT4 downregulation may converge on AMPK activation, or they may represent parallel, reinforcing mechanisms. This would enrich the mechanistic model and open new hypotheses.

      The mechanism of photobiomodulation: Herrera et al. explicitly challenge the prevailing paradigm that red light acts solely via cytochrome c oxidase (by showing long‑lasting effects, unchanged OXPHOS protein levels, and no difference in permeabilized cells). The current finding (red light acts through SIRT4 downregulation, i.e., not direct enzymatic activation, aligns perfectly with Herrera´s critique.

      Long‑term metabolic effects - Herrera et al. show that a single red light exposure elevates oxygen consumption for up to 2 days. The current study focuses on changes at 12‑24 h. Their data extend the time window and suggest that the metabolic reprogramming you describe may persist longer than currently discussed, which is clinically relevant.

      Discussing Herrera et al. results would not only acknowledge independent, corroborating evidence but also allow the authors to position your SIRT4‑centric mechanism within a broader, emerging understanding of red‑light photobiomodulation.

      Comments on the latest version:

      The authors have made a terrific work in answering the reviewers and modifying the manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the editors and reviewers for your careful evaluation of our manuscript and for the constructive recommendations that have helped us improve the rigor, clarity, and balance of the study. We are pleased that the reviewers recognized the potential value of linking red light exposure to SIRT4 downregulation, fatty acid metabolism, H3K9 acetylation, and attenuation of ageing-related phenotypes. We have revised the manuscript extensively in response to the reviewers’ comments.

      In particular, we have clarified the wavelength specificity of the red-light response, reanalyzed and more cautiously interpreted the omics data, improved the presentation and quantification of semi-quantitative experiments, revised statistical reporting, corrected gene/pathway annotations, toned down mechanistic claims where direct evidence was insufficient, and expanded the Discussion to integrate recent evidence on red-light-induced fatty acid oxidation and AMPK/ACC signaling. We also added a dedicated limitations paragraph addressing the use of female mice, the absence of a complete in vivo wavelength-control and source-blocked sham cohort, and the need for future direct metabolic flux and isolated mitochondria studies.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      This is a challenging hypothesis that would require some additional experimental controls. The pathway dissection, while extensive, is sometimes approached in unconvincing ways, and the results are not always evident to judge or interpret. Technically, the western blots and transcriptomic analyses require notable improvements.

      We would like to thank the reviewer for the careful and patient examination of the issues identified in our manuscript. The poor quality of some of the Western blot bands in Figure 4 may have been caused by inappropriate electrophoresis conditions during the Western blot experiments. In the revised manuscript, we will optimize the electrophoresis conditions to obtain higher-quality protein bands and update the quantitative data. Regarding the quantification format, we believe that heatmaps provide a more intuitive representation of trends in protein expression across different treatment groups. This approach more accurately reflects the results of our biological replicates than simply analyzing the significance of differences in the grayscale values of protein bands. For the analysis of transcriptomic data, we will conduct a more detailed analysis of signal pathway enrichment and the identified differentially expressed genes to ensure that predicted genes are excluded from our current results and redundant data presentation is removed.

      Regarding additional experimental controls, such as incorporating experimental data under blue light treatment conditions as a control for red light. While exploring the optimal red light irradiation dose at the cellular level, we simultaneously conducted experiments on the effects of blue light irradiation at the same dose on keratinocyte activity. The results indicated that as the blue light irradiation dose increased (0–160 J/cm<sup>2</sup>), the keratinocyte activity exhibited a dose-dependent decline. This indicates that blue light is phototoxic to keratinocytes. The relevant experimental results have already been published in our previous study (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). Taken together with the data from our study, this demonstrates that the anti-ageing effects of red light reported in the current manuscript are indeed driven by red light.

      Reviewer #2 (Public review):

      Weaknesses:

      The paper does not evolve to use the mechanistic discoveries of the manuscript to help our community to identify the mechanism of photobiomodulation, which is not known so far.

      I would like to draw attention to a recently published paper by Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), which shows that red light (660 nm) stimulates mitochondrial fatty acid oxidation in keratinocytes via AMPK‑dependent phosphorylation of ACC, without altering expression of electron transport chain complexes. I believe this paper is highly complementary to the current study.

      Herrera et al. demonstrate that red light increases basal, ATP-linked, and maximal oxygen consumption rates in keratinocytes specifically through enhanced fatty acid oxidation (inhibited by etomoxir). This independently validates the central finding of the current manuscript, i.e., red light boosts lipid metabolism, strengthening the robustness of this concept.

      While the current manuscript focuses on the SIRT4-MCD axis, Herrera et al. identify AMPK phosphorylation and ACC inhibition as key effectors. The authors can integrate and expand their discussion, since SIRT4 downregulation may converge on AMPK activation, or they may represent parallel, reinforcing mechanisms. This would enrich the mechanistic model and open new hypotheses.

      The mechanism of photobiomodulation: Herrera et al. explicitly challenge the prevailing paradigm that red light acts solely via cytochrome c oxidase (by showing long-lasting effects, unchanged OXPHOS protein levels, and no difference in permeabilised cells). The current finding (red light acts through SIRT4 downregulation, i.e., not direct enzymatic activation) aligns perfectly with Herrera´s critique.

      Long-term metabolic effects-Herrera et al. show that a single red light exposure elevates oxygen consumption for up to 2 days. The current study focuses on changes at 12-24 h. Their data extend the time window and suggest that the metabolic reprogramming you describe may persist longer than currently discussed, which is clinically relevant.

      Discussing Herrera et al.'s results would not only acknowledge independent, corroborating evidence but would also allow the authors to position their SIRT4-centric mechanism within a broader, emerging understanding of red-light photobiomodulation.

      We would like to thank the reviewer for providing us with constructive suggestions for discussion. Our results showed that under red light conditions, both glycolipid and lipid metabolism were activated in keratinocytes, and cellular metabolic flux increased. The activation of lipid metabolism directly led to an increase in metabolism-associated H3K9ac and drove the upregulation of anti-ageing-related genes; we believe this is key to the anti-ageing effects of red light. Mechanistic analysis combining proteomics and acetylation proteomics revealed that red light significantly downregulated SIRT4 expression and increased the acetylation of MCD, a protein regulated by SIRT4 that governs cellular fatty acid oxidation rates. Through validation using cell-level knockdown and inhibitors, we confirmed that SIRT4 inhibition exerts anti-ageing effects in vitro and that inhibiting MCD function under red light conditions suppresses H3K9ac. These results establish the role of the SIRT4-MCD signalling axis in mediating the anti-ageing effects of red light.

      The study by Herrera et al. included a substantial body of validation data confirming the role of red light in promoting fatty acid oxidation, providing robust empirical support for our research. Furthermore, Herrera et al. revealed that red light-induced fatty acid oxidation depends on AMPK and ACC phosphorylation. This mechanism of red-light photobiomodulation may refute the notion that its bio-regulatory effects rely solely on the action of mitochondrial cytochrome c oxidase. Furthermore, together with our study revealing that red light exerts anti-ageing photobiomodulatory effects via the SIRT4-MCD signalling axis, these findings independently confirm that red light regulates cellular fatty acid oxidation, thereby demonstrating the pivotal role of activated fatty acid oxidation in the bio-regulatory effects of red light. In the revised manuscript, we will include a discussion on the potential link between the red light-driven downregulation of SIRT4 and the phosphorylation of AMPK/ACC. This will be of positive value in elucidating how SIRT4 exerts its anti-ageing effects by regulating lipid metabolism, as well as in explaining the possible mechanisms by which red light downregulates SIRT4.

      Recommendations for the authors:

      Summary of Major Revisions

      Changes made in the revised manuscript:

      (1) Added a clearer explanation of why the 625-635 nm red-light regimen was considered the active intervention and how the available blue-light data from our previous work support wavelength-dependent effects on keratinocytes.

      (2) Revised the language describing inflammatory regulation. We now avoid presenting red light as producing a uniform anti-inflammatory effect and instead describe selective remodeling of ageing-associated inflammatory and SASP signatures.

      (3) Improved figure presentation and quantification for immunofluorescence, metabolite, and western blot assays; clarified image-analysis regions, replicate numbers, and normalization procedures.

      (4) Reanalyzed transcriptomic, proteomic, and acetyl-proteomic datasets with appropriate multiple-testing correction and corrected erroneous pathway/gene annotations in metabolic gene panels.

      (5) Replaced overly strong causal wording with more conservative language, especially regarding PI3K/Akt/mTOR, cytochrome c oxidase, SIRT4 localization, PPARα immunofluorescence, and direct fatty acid oxidation flux.

      (6) Expanded the Discussion to incorporate Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), highlighting convergence between the SIRT4-MCD model and AMPK/ACC-dependent fatty acid oxidation.

      (7) Corrected typographical, nomenclature, and figure-legend inconsistencies throughout the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) Wavelength specificity and need for a non-red-light control

      As a reader, one is left wondering whether the effects are due to red light specifically. An important control would have been to irradiate mice and cells with another light wavelength, such as blue light.

      We agree that wavelength specificity is a critical issue for interpreting photobiomodulation studies. In the revised manuscript, we have clarified that the anti-ageing and metabolic effects described here apply specifically to our 625-635 nm red-light regimen, rather than to visible light in general. We have also added a discussion of our previously published blue-light experiments, in which keratinocyte viability decreased in a dose-dependent manner across the same 0-160 J/cm<sup>2</sup> dose range (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). These data indicate that blue light and red light produce distinct biological outcomes in keratinocytes. Because high-dose blue light was cytotoxic under comparable cellular conditions and because the present study was designed to investigate the long-term effects of red light in aged mice, we did not perform prolonged in vivo blue-light irradiation as an ageing intervention.

      Changes made in the revised manuscript:

      Clarified in the revised Introduction and Discussion that the conclusions are specific to 625-635 nm red light under the irradiation parameters used in this study. (Lines 92 to 94, Lines 1146-1149)

      Added text summarizing the published blue-light comparison data from our previous study (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1), including the dose-dependent decline in keratinocyte activity after blue-light irradiation. (Lines 90 to 92, Lines 1146-1149)

      Added data on the wavelength range of the red light used in this study. (Lines 529 to 531, Fig S1a)

      (2) Complexity of inflammatory effects

      The manuscript repeatedly emphasizes anti-inflammatory effects, yet some cytokines such as IL-18, Ccl2, TNF-α, Ccl2, and IL-8 appear increased. This suggests that the effects may be more complex than presented and may require additional readouts or stronger statistical power.

      We unanimously agree that the inflammatory response to red light should not be described as a simple, uniform suppression of all cytokines. We have demonstrated that changes in the levels of the senescence-associated secretory phenotype (SASP) at the cellular level and in skin tissue following red light treatment not only indicate that red light-induced metabolic activation can reduce the age-related inflammatory baseline, but also reveal red light-driven short-term reparative effects or stress-related cytokine responses. We consider this to be consistent with the findings, and the downregulation of NF-κB-related signalling observed in skin tissue following periodic red light irradiation of aged mice further supports the conclusion that red light alleviates the age-related inflammatory baseline. In fact, in our previous study, we did observe that red light treatment promoted increased levels of the cytokine Ccl2, which plays an important positive role in rapid wound healing (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). In the revised manuscript, we have reworded the relevant Results and Discussion sections to indicate that red light remodels ageing-associated inflammatory signalling rather than globally reducing every inflammatory mediator. We have also toned down statements suggesting that red light ‘reverses’ or ‘suppresses’ inflammation where the underlying data support a more selective effect.

      Changes made in the revised manuscript:

      Replaced broad terms such as “anti-inflammatory effects” with more precise wording such as “remodeling of ageing-associated inflammatory signaling” where appropriate. (Lines 541 to 543, Lines 573 to 574, Lines 893 to 895)

      Expanded the Discussion to explain that red-light-induced metabolic activation may simultaneously reduce senescence-associated inflammatory tone while allowing transient reparative or stress-related cytokine responses. (Lines 1135 to 1142)

      (3) Figure clarity, semi-quantitative methods, western blot quality, and inconsistent band patterns

      Many differences are difficult to see or require orthogonal validation. Some tissue-specific signals and western blots are difficult to judge. Several western blots are of poor quality, and multiple markers show inconsistent band profiles across experiments, including SIRT4 in Figure 5.

      We thank the reviewer for highlighting these technical and presentation issues. We have reviewed the semi-quantitative data and revised the presentation of the figures to improve their interpretability. For Western blot experiments, we optimised the electrophoresis and transfer conditions, replaced low-quality representative images where possible, and updated the semi-quantitative results. Furthermore, regarding the lack of clarity in the SIRT4 protein band, we have conducted repeat experiments and updated the main text to include a clearer image of the band. The issue with the annotation of the protein location was in fact due to an oversight during the data analysis process; we have carried out a detailed review and provided the uncropped full-length Western blot images for all experiments in the Supplementary Materials for the reviewers’ scrutiny. Finally, we would also like to point out that factors such as sample origin, protein extraction, electrophoresis conditions, antibody exposure time, and potential non-specific detection may all contribute to differences in band patterns. At present, the core conclusions regarding SIRT4 are supported by multiple lines of evidence, including mRNA analysis, immunofluorescence, Western blotting of bands at the expected sizes, and SIRT4 knockdown experiments, rather than being based solely on any single semi-quantitative Western blot result. We therefore believe that the conclusions drawn from the data presented in the revised manuscript are equally convincing.

      Changes made in the revised manuscript:

      Replaced or improved low-quality Western blot panels and updated quantitative analyses in revised Figures 3-5 and associated supplementary material. (Fig 3m, Fig 4, Fig 5f)

      Clarified the normalization approach for H3K9ac/H3 and target/loading-control comparisons, and the use of Actin or H3 as appropriate loading controls. (Lines 283 to 293)

      Bands with nonspecific profiles were excluded from quantitative conclusions and the manuscript conclusions no longer depend on those ambiguous signals.

      Revised the Results (Repeat the experiment to update the low-quality Bands) to avoid overstating changes that are not clearly visible or not supported by statistical analysis. (Fig 4n and r)

      (4) Choice of pharmacological agents and need for genetic strategies

      The choice of drugs in Figure 4 is puzzling. More specific and widely used inhibitors could be used to block PI3K/Akt or mTOR, and natural agonists such as insulin or EGF could be used. Genetic strategies should complement these observations.

      We agree that pharmacological perturbation experiments should be interpreted with caution. In this study, our criteria for selecting inhibitors were based on transcriptomic and proteomic analyses; we sought to determine how the most direct inhibition of red light-activated signalling pathways would affect H3K9ac levels. In the revised manuscript, we have clarified the rationale for the compounds used and have reduced the causal weight assigned to these inhibitor/agonist experiments. These data are now presented as supportive evidence that red light is associated with metabolism-related signalling changes, rather than as definitive proof that PI3K/Akt/mTOR is the primary upstream mechanism. We have also emphasised the genetic SIRT4 knockdown experiments as a more direct mechanistic test for the SIRT4-centred part of the model. We acknowledge that additional experiments using more selective inhibitors, physiological agonists such as insulin or EGF, and genetic perturbation of PI3K/Akt/mTOR components would be valuable for future studies.

      Changes made in the revised manuscript:

      Revised the text describing pharmacological experiments to distinguish supportive pathway modulation from direct causal evidence. (Lines 787 to 789)

      Added a limitation and future direction noting that genetic perturbation of PI3K/Akt/mTOR and physiological pathway activation with insulin or EGF would strengthen the model. (Lines 1190 to 1194)

      (5) Serum NADH measurement

      In Figure 1t, the authors measure serum NADH. NADH is poorly detectable in serum or plasma, and changes may reflect blood-cell lysis during collection rather than circulating NADH.

      We appreciate this technical concern. We have revised the manuscript so that serum NADH is no longer used as a central mechanistic readout. We now treat this measurement only as an exploratory indicator of systemic redox-related changes and explicitly acknowledge that serum or plasma NADH is vulnerable to artifacts from blood-cell disruption during sampling. The mechanistic interpretation has been shifted toward cellular and tissue measurements, including intracellular NADH/NADPH/GSH, ATP, acetyl-CoA, fatty acid uptake, and H3K9ac, which are more directly relevant to keratinocyte metabolic remodeling.

      Changes made in the revised manuscript:

      Removed serum NADH from the main causal argument linking red light to metabolic flux and H3K9ac.

      Placed greater emphasis on cell-based metabolite assays, tissue acetyl-CoA, and H3K9ac measurements as the main metabolic-epigenetic evidence. (Lines 582 to 587, Fig 1s)

      (6) Direct assessment of glycolysis and fatty acid oxidation

      The authors propose that red light increases glycolysis and fatty acid oxidation, but this could be assessed directly rather than through surrogate measures.

      We agree. Our current data include multiple metabolic readouts, including glucose and fatty acid uptake, ATP, NADH/NADPH/GSH, triglycerides, fatty acids, pyruvate, lactate, acetyl-CoA, and MCD-dependent changes; however, these assays are not equivalent to direct flux measurements such as Seahorse extracellular flux analysis, isotope tracing, or etomoxir-sensitive respiration. We have therefore revised the wording throughout the manuscript to distinguish metabolic remodeling and fatty-acid-oxidation-related signatures from direct measurements of fatty acid oxidation flux. We also incorporated the recent independent work by Herrera et al., which directly measured oxygen consumption and demonstrated red-light-induced fatty acid oxidation in keratinocytes. This external evidence supports the biological plausibility of our SIRT4-MCD model while making clear which aspects are directly measured in our study and which are inferred.

      Changes made in the revised manuscript:

      Replaced overstrong language such as “red light increases fatty acid oxidation” with “red light promotes PPAR-α-related fatty acid metabolism pathway” where direct flux data were not measured in our experiments. (Lines 882 to 883)

      Expanded the Discussion to integrate direct FAO evidence from Herrera et al. and to place the SIRT4-MCD axis within a broader red-light metabolic framework. (Lines 1169 to 1189)

      (7) Incorrect annotation of metabolic genes in Figure 3e

      Acss2, Aldh3b1 and Aldh3a1 are not glycolytic enzymes, Aldh3a3 does not appear to exist, and several enzymes classified as FAO are fatty acid synthesis enzymes. This questions the interpretation of the data.

      We thank the reviewer for identifying these annotation errors. We have rechecked the gene names and pathway assignments in the transcriptomic analysis and corrected the metabolic gene panels. We have confirmed that Acss2 is an acetyl-CoA synthase involved in the metabolism of acetate to acetyl-CoA. The Aldh family genes, meanwhile, are associated with aldehyde metabolism and detoxification. We have made the corresponding adjustments in the manuscript. The incorrectly listed Aldh3a3 entry has been removed. Furthermore, we have categorised genes involved in fatty acid metabolism as ‘fatty acid metabolism-related genes’, rather than grouping them all under the FAO category. These revisions have significantly improved the accuracy of the metabolic interpretation.

      Changes made in the revised manuscript:

      Reannotated Figure 3e and the corresponding Results text to correct glycolysis, TCA cycle, pentose phosphate pathway, and fatty acid metabolism categories. (Fig 3d)

      Removed the erroneous Acss2 and Aldh family genes. (Fig 3d)

      Revised the metabolic model to avoid using incorrectly grouped genes as evidence for direct fatty acid oxidation. (Lines 707 to 709)

      (8) Transcriptomic analysis and implausible volcano-plot p-values

      The transcriptomic analysis raises concerns. For example, the volcano plot in Figure 3d appears incorrect, with -log<sub>10</sub>(P-value) around 300 despite n=3 biological replicates.

      We thank the reviewer for pointing out this important issue. We have reopened the transcriptomic data and found that the extremely high -log<sub>10</sub>(P-value) in the original volcano plot were caused by the automatic replacement of very small P-values—generated during the differential expression analysis—with zero in the tabular data. To avoid misleading visualisations, we have regenerated the volcano plot using Q-values in place of the original P-values. Differentially expressed genes were defined as those with a Q-value < 0.05 and |log<sub>2</sub> fold change| > 1. Furthermore, for visualisation purposes only, the upper limit for q-values was set to 1 × 10<sup>-50</sup> for values below 1 × 10<sup>-50</sup>. This adjustment does not affect the statistical classification of differentially expressed genes but prevents over-interpretation of extremely small values. The revised volcano plots and legends have been updated accordingly. To avoid any potential misinterpretation arising from these updates, the updated volcano plots are presented in the supplementary materials.

      Changes made in the revised manuscript:

      Reanalyzed transcriptomic data using appropriate multiple-testing correction and revised the volcano plot. (Fig S3a)

      Corrected the y-axis transformation and removed implausible -log<sub>10</sub>(P-value) presentation. (Fig S3a)

      Updated Methods to specify the statistical workflow for transcriptomic differential expression and pathway enrichment. (Supplementary materials Lines 46 to 52)

      Moved the analysis of metabolic pathways based on transcriptomic data to the supplementary material, thereby reducing the reliance of the conclusions on transcriptomic data (Fig S3b).

      (9) Need to tone down mechanistic claims regarding PI3K/Akt/mTOR, cytochrome c oxidase, SIRT4, and PPARα

      The mechanisms proposed must be toned down. PI3K/Akt/mTOR should not be called glycolytic pathways, the link to red light or cytochrome c oxidase is vague, SIRT4 reduction requires mitochondrial counterstaining, and PPARα appears cytosolic after SIRT4 knockdown.

      We agree and have substantially revised the mechanistic language. PI3K/Akt/mTOR is no longer referred to as a ‘glycolytic pathway’; instead, it is described as a metabolism-related signalling axis that may influence glucose uptake, growth and nutrient-responsive metabolism. We have also toned down statements attributing red-light effects directly to cytochrome c oxidase, as our study primarily examines downstream metabolic and epigenetic remodelling rather than direct photoreceptor activation. With regard to SIRT4, we have revised the text to avoid interpreting changes in SIRT4 immunofluorescence alone as evidence of altered mitochondrial abundance or mitochondrial localisation. The conclusion is now based on a combination of SIRT4 mRNA levels, western blot bands of the expected size, immunofluorescence trends, and SIRT4 knockdown phenotypes. With regard to PPARα, we have re-examined the PPARα antibody used for the cellular immunofluorescence experiments. In the original Figure 5p, we mistakenly used a PPARα antibody (PPARα, Abclonal, A25296) that is only suitable for Western blot (WB) experiments; we believe this was the cause of the mislocalisation of the fluorescent signal; Consequently, we conducted new experiments using a PPARα antibody (PPARα, Abclonal, A22887) specifically designed for cellular immunofluorescence. The relevant experimental data have been corrected in the manuscript.

      Changes made in the revised manuscript:

      Replaced “PI3K/Akt/mTOR glycolytic pathway” with “The PI3K-AKT signalling pathway is involved in the regulation of glucose metabolism” throughout the revised manuscript. (Lines 701 to 705, Lines 809 to 810, Lines 813, Lines 1193)

      Reduced mechanistic certainty around cytochrome c oxidase and framed it as a possible upstream photoreceptor rather than an experimentally proven mechanism in this study. (Lines 827 to 830, Lines 842 to 848)

      Repeat the PPARα immunofluorescence staining experiment. (Fig 5p)

      (10) Need for isolated mitochondria experiments and red/blue light comparison of mitochondrial respiration

      If the effect of red light relies on mitochondrial cytochromes, additional proof would be needed, potentially using isolated mitochondria and comparing how red and blue light influence respiration capacity.

      We agree that isolated mitochondria experiments would be an important way to test direct mitochondrial photoreception. Because the present study was designed around cellular and in vivo metabolic-epigenetic remodeling, we did not perform isolated mitochondria irradiation experiments. To address this concern, we have toned down statements implying direct cytochrome activation and revised the Discussion to distinguish between direct mitochondrial photoreceptor models and downstream metabolic reprogramming. We also added a future direction proposing isolated mitochondria or permeabilized-cell experiments comparing red and blue light effects on respiration, ATP-linked OCR, maximal respiration, and FAO-dependent respiration. The revised manuscript now emphasizes that our data support a downstream SIRT4-MCD-H3K9ac mechanism after red-light exposure, while the proximal photophysical event remains to be fully defined.

      Changes made in the revised manuscript:

      Added discussion of the need for isolated mitochondria, permeabilized-cell, and wavelength-comparison respiration experiments. (Lines 1194 to 1197)

      Reviewer #2 (Recommendations for the authors):

      (1) Statistical reporting, post-hoc tests, normality/equal-variance tests, exact p-values, and FDR control

      The manuscript states that one-way ANOVA followed by Tukey or Dunnett tests was used, but it does not consistently specify the post-hoc correction for each figure. Normality and equal-variance tests are not reported, p-values are shown only as asterisks, and FDR control is not mentioned for transcriptomics and proteomics.

      We agree that the statistical reporting needed to be more complete. We have revised the Statistics and reproducibility section and the figure legends to specify the statistical test used for each experiment, the post-hoc correction applied after ANOVA, the number of independent biological replicates, and the definition of error bars. Where multiple comparisons were performed, we now state whether Tukey’s or Dunnett’s correction was used. Regarding P-value presentation, we have retained the use of asterisks in the figures as visual indicators of statistical significance to maintain figure readability. For transcriptomic, proteomic, and acetyl-proteomic analyses, we have revised the Methods section to state that multiple-testing correction was performed using the Benjamini–Hochberg false-discovery-rate procedure. Adjusted P values or Q values were used for differential-expression and pathway-enrichment analyses. These revisions clarify the statistical workflow and strengthen the reproducibility of the study.

      Changes made in the revised manuscript:

      Revised the Statistics and reproducibility section to define statistical tests, post-hoc corrections, assumption checks, and multiple-testing correction. (Lines 509 to 522)

      Updated relevant figure legends to include n values, statistical tests, post-hoc corrections, and definitions of significance symbols. (Lines 515 to 516)

      Added FDR control details for RNA-seq, proteomics, acetyl-proteomics, and pathway-enrichment analyses. (Lines 450 to 455)

      (2) Figure clarity and quantitative analysis of fluorescence, JC-1, metabolite, and western blot data

      Several figures lack clarity or appropriate quantification. Figure 1i-j H3K9ac quantification should be based on whole-image or multiple fields; Figure 2e JC-1 should include red/green ratio quantification; Figure 2k-p metabolite data should include absolute concentrations; Figure 3j needs appropriate loading controls.

      We appreciate these specific suggestions and have revised the figure presentation accordingly. For H3K9ac immunofluorescence in skin sections, we have clarified the anatomical region quantified and performed a more objective quantification using multiple fields/regions per section rather than relying on a visually selected dashed area. The dashed regions in the representative images were made clearer and the quantification criteria were added to the Methods and legend. For JC-1 staining, the bar chart on the right-hand side of the mitochondrial membrane potential fluorescence image in Figure 2e shows the quantitative data for the red/green fluorescence ratio obtained from independent experiments; compared with providing only a representative image, these data offer a more easily interpretable quantitative measure of mitochondrial membrane potential. To avoid any potential misunderstanding, we have corrected the vertical axis. For metabolite assays, we clarified normalization to cell number or protein content and revised the data presentation to include absolute or normalized concentrations where available, rather than relying solely on fold changes with variable y-axis scaling.

      Changes made in the revised manuscript:

      Revised Figure 1i-j quantification using multiple fields/regions per mouse section and improved dashed-region visibility. (Lines 431 to 440, Fig 1i and j)

      Corrected the vertical axis of the quantitative data for the JC-1 red/green fluorescence ratio. (Fig 2e and Fig S2a)

      Updated metabolite panels and/or source data to include absolute or protein-normalized values where available, and standardized y-axis interpretation. (Fig 2k-p and Fig 4s and v, Given the diversity of intracellular fatty acid and triglyceride species, absolute quantification based solely on absorbance measurements would be technically challenging and may not accurately reflect the content of each molecular component. Therefore, we presented the changes in fatty acid and glycerol levels as percentage-normalized relative absorbance values, which allowed consistent comparison among the experimental groups.)

      (3) Experimental design limitations: sex of mice and sham control

      Only female C57BL/6 mice were used, although aging and metabolic responses can be sex-dependent. The thermal-control argument lacks a true sham control in which mice are placed in the same apparatus with the light blocked at the source.

      We agree with these points. We have added a section on limitations stating that all aged mice used in this study were female, and that sex-dependent responses to red light, SIRT4 regulation, metabolism and skin ageing should be investigated in future studies using both male and female cohorts. Furthermore, regarding the design of the non-irradiated control group: although the control mice underwent the same depilation and routine procedures, they did not receive red light irradiation. However, we also acknowledge that establishing a sham-irradiated control group with light shielding would allow for stricter control of factors such as restraint, contact with equipment and procedural stress. However, given that this experiment involved a continuous cyclic photoperiodic treatment lasting two years, we were unable to supplement the study with a control experiment involving only red light shielding. Nevertheless, based on the fact that we observed only minimal changes in the mice’s skin temperature following red light irradiation, we believe that the primary factor driving the alleviation of the skin ageing phenotype in the mice remains red light-induced.

      Changes made in the revised manuscript:

      Added a limitation noting that the study used female C57BL/6 mice only and that sex as a biological variable should be addressed in future studies. (Lines 1198 to 1201)

      (4) Textual errors, nomenclature inconsistencies, and ChIP-qPCR normalization

      Several textual errors and inconsistencies should be corrected, including Pparg1a/Ppargc1a, Ricotr/Rictor, Pi3k/PI3K, Sirt4/SIRT4 protein nomenclature, and the use of RPL30 normalization in ChIP-qPCR without showing that RPL30 is unchanged.

      We thank the reviewer for their careful reading. We have corrected the typographical errors and standardised gene and protein nomenclature throughout the manuscript and figure legends. Specifically, Ppargc1α has been corrected to Ppargc1a, Ricotr to Rictor, and the capitalisation of PI3K has been standardised. We now use Sirt4 for the mouse gene and SIRT4 for the protein, applying the same convention to other genes and proteins. For ChIP-qPCR, we have revised the Methods and Results sections to describe normalisation against input and IgG controls more clearly, and to specify the role of the RPL30 locus as an internal control. We have also included data in the Supplementary Materials showing relative enrichment of H3K9ac in the RPL30 promoter region in PAM212 cells before and after red light irradiation; the results indicate that H3K9ac enrichment at the RPL30 locus remained stable across treatment groups after normalization to input DNA and correction against IgG background. This result indicates that the use of RPL30 as an internal control in ChIP-qPCR experiments is feasible.

      Changes made in the revised manuscript:

      Corrected Ppargc1a, Rictor, PI3K, Sirt4/SIRT4, and related nomenclature throughout the manuscript.

      Supplement the experimental results on the effect of red-light irradiation on the level of H3K9ac enrichment at the RPL30 locus in keratinocytes. (Lines 339-352, Lines 539 to 541, Fig S1d)

      (5) Additional Revision Addressing the Public Review and Herrera et al.

      The reviewer suggested integrating the recent study by Herrera et al. showing that 660 nm red light stimulates mitochondrial fatty acid oxidation in keratinocytes through AMPK-dependent phosphorylation of ACC, without changing electron transport chain complex expression. The reviewer also noted that these findings may complement the SIRT4-MCD axis and challenge a cytochrome-c-oxidase-only model of photobiomodulation.

      We are grateful for this constructive suggestion. We have expanded the Discussion to incorporate Herrera et al. and to place our SIRT4-MCD-centered mechanism within the broader emerging model of red-light-driven metabolic remodeling. Herrera et al. provide direct oxygen-consumption evidence that red light enhances fatty acid oxidation in keratinocytes and that this effect involves AMPK/ACC signaling. This is highly complementary to our data, in which red light decreases SIRT4, increases acetylation of MCD, promotes fatty-acid-metabolism-related signatures, elevates acetyl-CoA, and increases H3K9ac. In the revised Discussion, we propose two nonexclusive models: red light-induced SIRT4 downregulation may converge with AMPK/ACC-dependent relief of fatty acid oxidation, or the two pathways may represent parallel reinforcing mechanisms that together enhance lipid metabolic flux.

      Changes made in the revised manuscript:

      Added a paragraph discussing Herrera et al. in the revised Discussion. (Lines A1169 to 1189, Lines 1206 and 1208)

      Revised the conceptual model of red-light photobiomodulation to emphasize downstream metabolic reprogramming rather than direct cytochrome c oxidase activation alone. (Lines 827 to 830, Lines 843 to 844)

      Added future directions to test whether red-light-induced SIRT4 downregulation causally affects AMPK/ACC phosphorylation and FAO-dependent respiration. (Lines 1169 to 1189)

      We again thank the editors and reviewers for their thoughtful and constructive comments. The revised manuscript now provides a more rigorous and balanced presentation of the evidence, distinguishes direct measurements from inferred metabolic flux, corrects pathway annotations, improves figure quantification and statistical transparency, and places the SIRT4-MCD-H3K9ac mechanism within a broader framework of red-light-induced fatty acid metabolic remodeling. We believe these revisions substantially strengthen the manuscript and clarify both the significance and the limitations of our findings.

    1. eLife Assessment

      This study presents valuable findings on the role of specific dopamine neurons for aversive learning and modulation of innate behavior in Drosophila larvae. The authors present convincing evidence backed up by detailed behavioral quantification and rigorous testing. Their data confirms previous findings and will be of interest to the learning and memory community.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate the role of different specific dopaminergic neurons in the mushroom body of Drosophila larvae for learning and innate behavior. All the tested neurons are thought to be involved in punishment learning. The authors discover that artificial activation of single DANs in training leads to safety learning, but not punishment learning. Furthermore, activation of single DANs can lead to changes in locomotion behavior, which can affect light preference. The authors provide a deeper understanding of the functional diversity of single dopamine neurons; however, it is unclear how translatable these findings are to learning experiments with real punishment stimuli.

      The authors provide a detailed behavioral analysis of locomotion in response to activation of various dopamine neurons. This analysis allows them to exclude that the locomotion defects affect memory recall behavior.

      Strengths:

      The authors disentangle which kind of memories are formed with the activation of different dopamine neurons - safety learning and/or punishment learning. They further investigate whether the US is required in the test for recall. They do indeed find differences, and the results will be of interest to the learning and memory community.

      Interestingly, optogenetic activation of a single DAN during training leads to safety memory, but not punishment memory. Furthermore, DAN activation also affects innate locomotion, and the authors show that optogenetic activation of different DANs affects locomotion differently.

      Weaknesses:

      All experiments in the manuscript use optogenetic activation of DANs, thus it is not clear what kind of memories are formed. Several stimuli can be used as punishment, such as electric shock, salt, bitter, and light - it is not clear what kind of memory the authors investigate here. The findings could be discussed in the context of what DANs respond to. Furthermore, studies in adults and larvae showed that most DANs can code for both valences - etc., aversive DANs can be activated by punishment, and inhibited by reward. Thus, safety learning might be a result of a decrease in activity in DANs during odor presentation. The authors also do not discuss possible feedback loops from MBONs to DANs across compartments. Could such connections allow for safety learning in larvae?

      The authors show that artificial activation with different light intensities can form different memories and that increasing the light intensity sometimes leads to no memories. Also, using different optogenetic tools reveals different results. This again raises the question of how applicable the results will be for learning with real stimuli. Is there a natural stimulus that only induces safety learning, but no punishment learning? The authors discuss these limitations.

    3. Reviewer #2 (Public review):

      Summary:

      This study provides valuable context for ongoing research on the role of dopamine in memory and locomotion. DANs have been a fascinating area of study due to their complexity, and this work dissects specific DANs, exploring their roles in different memory-related behaviors while offering some explanations. The discussions provided by the authors effectively situates the study in the broader field of learning, memory, DAN circuitry and behavioral computation in insect brains. The study achieves what it sets out to and it does so unequivocally. The experiments were elegantly designed, leaving little room for doubt in the study's claims. However, the study lacks context regarding the molecular pathways underlying these results. While it strengthens current knowledge by providing robust evidence, it does little to explore the molecular mechanisms behind these effects.

      Strengths:

      (1) Experiment design is one of the strengths of this study. The experiments are thorough and cover the length and breadth of the core findings of the study. Although a lot of work has already been done in studying the role of dopamine in memory and locomotion, the dissection of the functions of distinct DANs in larvae has been done meticulously with well-structured experiments.

      (2) This study fits quite nicely into the puzzle of memory, especially in the context of Dopamine. Previous studies in *Drosophila* adults have shown the opposing roles of DANs in locomotion depending on the context of DAN activation. This study drives that point home for larvae, providing conclusive evidence in that regard.

      (3) The use of clear figures and simple language is one of the strengths of this paper. The figures are comprehensive, complete and manage to narrate the story by themselves. The flow of information is smooth. The simple and effective language used maintains scientific rigor while remaining accessible to those new to the field. A pleasant read.

      Weaknesses:

      (1) The authors have done a great job at structuring the figures. But some main figures would benefit from including the controls instead of placing them in supplementary.

      (2) The paper would benefit from a deeper discussion regarding molecular mechanisms underlying their results. It would be interesting to see what the authors think about different Dopamine receptors and how they relate to the findings of this paper.

      (3) Throughout the paper, the authors have been clear and comprehensive, but in some cases, further explanation of their choices were missing. For example, the choice to compare bending and tail velocity over other parameters within the same clusters is unclear.

      Comments on revised version.

      Most of the comments have been addressed.

    4. Reviewer #3 (Public review):

      Summary

      Across species, dopamine release serves seemingly diverse functions, such as reinforcing memories and regulating locomotion and flight. However, whether distinct dopaminergic neurons (DANs) are allocated to each function is unclear. In this study, Toshima et al. have used the numerically simple organization of the Drosophila larval brain to answer this question. They use optogenetic activation to systematically stimulate a small set of DANs, individually and collectively, and study the effect on diverse functions such as memory formation, retrieval, and locomotion. The reproducibility of optogenetic activation is a strength of this approach. At the same time, this is a caveat, as optogenetic activation may not recapitulate natural modes of activation and may lead to outcomes not observed under natural conditions. They find that singly or collectively, DL1 DANs can induce punishment and/or safety memory formation and retrieval. DANs can even gate the expression of memory. Finally, the same DANs also modulate locomotion in the larvae. The authors speculate that dopaminergic neurons in other species may also share such overlapping functions. Their findings are nicely summarised in Figure 9.

      Strengths

      The study systematically activates neurons in the DL1 cluster. Individual and collective stimulation of the Dl1 DANs has been conducted to assess the induction and gating of aversive punishment memory, safety memory, and acute locomotion.

      Specific adult Drosophila DANs are known to induce dual behaviors and functions. The same MP1/y1pedc DANs are recognized for gating appetitive memory expression and representing aversive teaching signals downstream of sensory stimuli such as electric shocks, bitter tastes, and heat. Neurons in the PPL1 cluster regulate adult flight and food-seeking behavior. The authors deserve credit for conducting an organized examination of dopaminergic neuronal functions in larvae, thereby making their findings more comparable and facilitating the proposal of a holistic model.

      They have provided substantial evidence for their findings and have frequently presented replicated behavioral datasets. They have been transparent about the results that were difficult to explain. Additionally, they have provided an impressive body of supporting data to strengthen their main findings.

      Weaknesses

      As mentioned above, optogenetic activation may not recreate natural neuronal activation in response to external stimuli. This could have led to outcomes that will not occur under other natural circumstances.

      Comments on revised version.

      I appreciate the author's responses, and I do not have any comments or suggestions at this point.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      All experiments in the manuscript use optogenetic activation of DANs, thus it is not clear what kind of memories are formed. Several stimuli can be used as punishment, such as electric shock, salt, bitter, and light - it is not clear what kind of memory the authors investigate here. The findings could be discussed in the context of what DANs respond to.

      This is indeed a caveat of our study and we discuss this issue now in lines 557-566. We refrained from testing necessity to specific US on purpose as we knew that another research group was focussing on this question in parallel (Weber et al., 2023, also published in eLife) and therefore rather focussed on complementary experiments. That study also includes a rather deep discussion about the inputs to individual DANs. We briefly refer to this discussion (lines 442-444) but decided to not go into detail to avoid too much overlap.

      Furthermore, studies in adults and larvae showed that most DANs can code for both valences - etc., aversive DANs can be activated by punishment, and inhibited by reward. Thus, safety learning might be a result of a decrease in activity in DANs during odor presentation. The authors also do not discuss possible feedback loops from MBONs to DANs across compartments. Could such connections allow for safety learning in larvae?

      We thank the reviewer for raising these points and included a brief discussion of both scenarios in lines 469-474.

      The authors show that artificial activation with different light intensities can form different memories and that increasing the light intensity sometimes leads to no memories. Also, using different optogenetic tools reveals different results. This again raises the question of how applicable the results will be for learning with real stimuli. Is there a natural stimulus that only induces safety learning, but no punishment learning?

      We do not know of such a stimulus. Based on our data, a US that only activates a single DAN should only make safety memory – however, the available data of which US activates which DAN is very limited in larvae and currently no such US is known. We discuss this point briefly in lines 557-563 and 572-574.

      The authors provide a detailed behavioral analysis of locomotion behavior; however, the detailed analysis seems unnecessary for that dataset. Modulation of speed and bending rate has been described before with simpler methods (specifically for MBONs). The revealed locomotion phenotypes probably affect larval locomotion during memory recall with light activation, thus the authors should show that larvae are potentially able to move during light-on memory tests.

      We expanded our locomotion analysis of the innate and learned odor preference experiments (new Fig. 6) and show that in these experiments, even with TH-DANs being activated, larvae indeed can move relatively normal and the existing locomotion phenotypes are not correlated to their olfactory choices.

      We do not agree that the locomotion analysis is unnecessary. Modulations of speed and bending have been described for MBONs but to our knowledge not for DANs. It is not trivial at all that DANs and MBONs cause the same behavioral modulations (see, for example, this adult study: Mohammad et al., 2024 Plos Biol). There is extremely limited knowledge about the motoric effects of dopaminergic neurons in larvae - we therefore find it important to describe our results in detail. We added some further rationale of why we think it is crucial to explore the functions of DANs for learning and movements together (lines 81-86).

      Reviewer #2 (Public review):

      Weaknesses:

      (1) The authors have done a great job at structuring the figures. But some main figures would benefit from including the controls instead of placing them in supplementary.

      We had decided to put the controls into the supplement in some cases to prevent the main figures to be overcrowded. We revised this decision upon the reviewer’s comment for Fig. 8 (previously Fig. 7) but decided to keep other figures unchanged as we feel that the current design best fits the purpose of each figure. We provide a figure-for-figure rationale in our response to the recommendations for the authors.

      (2) The paper would benefit from a deeper discussion regarding molecular mechanisms underlying their results. It would be interesting to see what the authors think about different Dopamine receptors and how they relate to the findings of this paper.

      We thank the reviewer for the suggestion. Although we agree that such a discussion would be interesting, we hesitate to expand on this topic, as the discussion is already quite long and our study does not contribute any new data to clarify the molecular dopaminergic mechanism.

      (3) Throughout the paper, the authors have been clear and comprehensive, but in some cases, further explanation of their choices were missing. For example, the choice to compare bending and tail velocity over other parameters within the same clusters is unclear.

      We understand that this choice was not clearly explained and expanded on our rationale in lines 244-251.

      Reviewer #3 (Public review):

      Weaknesses:

      The larvae exhibit directed locomotory action to express punishment or safety memory. If the larvae did not move, we would not be able to assess memory function. Hence, functional activation of DANs could result in one action, which seems like two different functions of memory expression and locomotion. It can also be argued that activation of DANs represents a teaching signal to the KCs, and then eventually, downstream of the MBONs, it results in locomotion modulation. Hence, the seeming functional diversity could be a function of different downstream neuronal pathways and not molecular context-dependent diversity inside dopaminergic neurons. The authors should address this possibility or point out the fallacy in the above argument.

      We thank the reviewer for raising this issue. To the first point, we expanded our locomotion analysis of the innate and learned odour preference experiments (new Fig. 6) and show that in these experiments, even with TH-DANs being activated, larvae indeed can move relatively normal. In addition, the existing locomotion phenotypes in these experiments were not correlated with the animals’ olfactory choice. This makes it unlikely that the changed locomotion directly determines our observation during the olfactory experiments.

      We do agree that it is possible that both the preference after learning and the changed locomotion could come through the same dopaminergic mechanism via diverse downstream pathways. We cover this hypothesis in Fig. 10G and address this question briefly in lines 580-582.

      The finding that activation of TH-GAL4 conveys aversive valence and R58E02-GAL4 conveys appetitive valence seems redundant (Figure 6). I understand they say this in the context of locomotion. However, they may not have mentioned similar findings in adults. In adults, artificial activation of DANs covered by the same GAL4 lines acts as aversive and appetitive teaching signals for memory formation. These references should be cited appropriately in the results and discussion if not currently included.

      We thank the reviewer for this comment and tried to include the relevant adult literature (see e.g. lines 351-356 and 583-604). In particular, we added a quite detailed discussion about a paper published after our initial submission that performed similar experiments for the adult PAM-DANs Lozada-Perdomo et al., 2025, iScience).

      We do not agree, however, that the experiments in Fig. 7 (previously Fig. 6) are redundant. Recent studies in adults found no correlation between the rewarding/punishing effects and the innate valence a given dopaminergic neuron induces (Rohrsen et al., 2021, bioRxiv; Mohammad et al., 2024, PLOS Biol; Lozada-Perdomo et al., 2025, iScience). To our knowledge, no such studies have been carried out in larvae so far. Therefore, we think that it is not only important to test it but that the respective results compared to the results in adults are of relevance for the readership.

      The evidence for the role of dopamine (Figure 7) can be bolstered by using other available RNAi lines against TH. A valium20 vector-based shRNA line is recommended. The current evidence is based mainly on non-specific pharmacological intervention with 3IY.

      We agree to this caveat and made it transparent now in lines 387-389 and 401-403. We nevertheless chose, for the time being, to not include further experiments to address this point in the current study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Activation of specific or multiple DANs seems to increase naïve odor preference (f1 or TH). Is this due to locomotion defects in TH - how does the odor preference develop over time? Can this increased odor preference explain the safety learning - where they also approach the odor stimulus?

      The reviewer is right that in presence of light, we see increased odor preference both innate and after unpaired training – theoretically, that could be the same effect. However, this would not explain why we see the same increased preferences after unpaired learning with all driver strains but increased innate preferences only for some of them. Moreover, when activating TH, we also see increased odor preference in absence of light after unpaired training but not innately. We therefore think that these are independent effects.

      We also include a new Fig. 6 providing additional information, including the development of preference over time, and addressing the question whether the modulations in locomotion can explain differences in odor preference.

      Locomotion behavior was assessed in 30s light-on periods - was the behavior different from the memory test or naïve preference test which had light on for 3 (or 2.5) minutes - which light intensity was used for the data shown in Figure 5.2? How did the larvae move in the high light concentration/ or with ATR - did they not show memory due to impaired locomotion?

      Light intensity for all odour preference and learning experiments with ChR2-XXL was 100 µW/cm<sup>2</sup> (except Fig. 2 – S2C), i.e. equivalent to the experiments in Fig. 5 – S1 and Fig. 9 – S2 (weak light). We replaced Fig. 5 – S2 with a new expanded Fig. 6 analysing the locomotion of our experiments shown in Fig. 2F, 3F and Fig. 2 – S1F. We show in this figure that the larvae can move relatively normal and that the locomotion effects are weaker than in our experiments with 30s light periods. Unfortunately, we do not have videos available for all experiments and therefore cannot make a similar analysis for Fig. 2 – S2C and D when we used strong light or ATR feeding. The experimenters did not notice impaired locomotion during the experiment and the animals did show normal odour preferences similar to those shown in Fig. 3 – but the preferences were the same after paired and unpaired training, resulting in zero Memory Scores. Therefore, we do not think that the locomotion prevented the memory expression.

      In several experiments, even genetic controls seem to show learning with blue light activation. Thus, the light itself seems to activate DANs. The authors should discuss these effects and explain what this could mean for the findings. The light stimulus might not just activate the specific DAN that expresses the optogenetics, but also additionally other DANs which respond to light.

      We thank the reviewer to point this out and point out this caveat in lines 159-165.

      The authors speculate about the function of potential MB circuits - the DAN-MBON circuit is not well described so far and might be required for the US in test memory recall. A straightforward experiment to investigate the involvement of this circuit in punishment or safety memory recall would be to block dopamine receptors in the MBON.

      We very much agree to this suggestion, but believe these experiments are beyond the scope of the current study. We therefore decided to not perform these experiments for the current paper.

      Reviewer #2 (Recommendations for the authors):

      (1) As self-explanatory as the figures are, it would be interesting to also see controls in some of them. For example, in Figure 5, the effect size graph (Figure 5C) clarifies to an extent the difference between control genotypes and the experimental genotypes. It would be nice to see the results of genetic controls in Figures 5A and 5B instead of in Figure 4 - supplement 2.

      We originally decided to put the controls into the supplement to prevent the main figures to be overcrowded. We revised this decision upon the reviewer’s comment for Fig. 8 (originally 7). For Fig. 4 and 5 specifically, we decided to keep the current layout because each serves a different purpose: Fig. 4D-L, Fig. 5 – S1 and S2 present the actual data with all genotypes that were made in parallel and therefore can be compared directly. Fig. 4 – S2 aims to visualize the effect of the light by comparing all controls across all experiments, normalized to the same starting value. Fig. 5A and B aim to compare the shape and effect size of activating DANs on top of the effect of the light - therefore, we subtracted the controls in each experiment from the experimental group. We think that adding the controls’ behaviour to Fig. 5A and B would undermine the aim of this figure.

      (2) It is a bit unclear why bending and tail velocities were the parameters chosen to compare between groups while in most cases they were of lower relative importance according to Figure 4 - supplement 1. Elaborating on this would strengthen the differences in behavior and also the claims of this study.

      We thank the reviewer for the suggestion and tried to make our choice clearer. Please see our answer to the respective part of the public review.

      (3) In adults, it has been shown that the same DAN can encode opposing valence depending on whether it was activated before or after odor presentation. Discussing the importance of temporal order of stimulus processing would bolster the results regarding paired and unpaired training in Figure 3.

      We thank the reviewer for this very good suggestion – also in larvae, this temporal function has been described. We discuss these observations in relation to our results in lines in 481-494.

      Reviewer #3 (Recommendations for the authors):

      Toshima et al., as stated in the public reviews, have done an admirable job with this manuscript. Below are specific suggestions that could improve the manuscript. It is, of course, up to the authors to decide which ones to attend to.

      Treat controls consistently. In Figure 2 and others, parental controls are not pooled, but in Figure 3, for odor preference, controls are pooled.

      We agree that the same things should be treated in the same way throughout a study and normally adhere to this principle. We nevertheless made an exception for Fig. 3 only because its goal is to provide a post-hoc analysis across several replications of experiments, some of which included genetic controls, others not (from Fig. 2, Fig. 2-S2 and S3). Due to relatively small effect sizes and high variability in odor preferences, to answer the question of paired and unpaired learning, we need higher sample sizes than each individual experiment provided. We therefore decided to pool all “equivalent” data across all these experiments. We do agree that this is a suboptimal approach but hope the reviewer can understand the rationale behind it. We explained our rationale clearer now (lines 180-183).

      I prefer to see all data points in a graph. It is more transparent than the box plots. Also, could you note why the data median is preferable to show over the mean?

      Although we in principle agree to the notion that presenting all data points is more transparent, we opted against it as it makes some graphs harder to read in particular with high sample sizes – in some of our figures, we have hundreds of data points per group. We explain our choice, including for using the median, in the method section (lines 835-839).

      Please undertake another round of language editing to handle spelling errors, etc. Use consistent British/American English.

      We thank the reviewer for their suggestion and tried our best to fix any spelling and grammar errors.

      I urge the authors to move beyond the false dichotomy of 'p' value statistics to using the statistical framework of estimation statistics for data analysis. I understand switching from familiar statistical analysis in such a late manuscript stage is very difficult. However, the authors can consider the estimation statistics framework in subsequent studies. https://www.estimationstats.com is a good starting point for biologists to get to know a framework that has been extensively worked on and is arguably a more 'honest' way of analyzing data. Disclaimer: I am not associated with the above website.

      We agree that the p-value has problems and are aware of the estimation statistics framework. We had considered applying it here, but we decided against switching to a completely different statistical framework for a research project that was ongoing since several years. However, we are sincerely considering it for our current research projects.

    1. eLife Assessment

      In this solid work, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role for two conserved acidic residues rather than a single one. This valuable study used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript the applicants study two residues in the GHKL ATPase active site of Aq MutL and GyrB, and argue that the catalytic base function is shared between two conserved acidic residues that are 3 residues apart.

      In the manuscript, they generated mutant versions in MutL and GyrB (both ala and the appropriate Asn/Gln version) and performed ATPase analysis. They also generated high resolution crystal structures of the GyrB NTD with AMPPnP for WT and mutants of the two acidic residues. The data show that mutation in either of these residues does not fully kill activity (with the exception of the Alanine mutation of the first of the two, that interferes with ATP (or AMPPnP) binding). When the acidic residues are mutated to Asn/Gln, the catalytic water can still be positioned, and hence these mutants are more active than the Ala mutants. In both cases the double mutation is catalytic dead.<br /> The authors then perform phylogenetic analysis and ancestral gene reconstruction and based on this they argue that HSP90 forms a different class of GHKL ATPases, and lost rather than gained this separate status.

      Strengths:

      The biochemical analysis seems solid.

      Weaknesses:

      - A major question that remains, is why the mutations have so much more detrimental effect in MutL (100-fold lower kcat/KM) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      The authors need to discuss this issue explicitly to make it clear that conservation of the mechanism is not complete and that other interpretations are possible.

      - The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the the acidic residues and their mutants are not shown.

      This has been addressed.

      There are some issues with figure S2B and S5.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role of two conserved acidic residues rather than one. The authors have used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism.

      Through a detailed re-analysis of their previously published structure of the aqMutL NTD (ATPase domain) in complex with AMPPCP, they identified Glu29 and Glu32 as interacting with nucleophilic water for the catalysis. The authors carefully dissected the respective roles of these two acidic residues with a series of site-directed mutations. Mutations at Glu29 impaired ATPase activity without affecting protein secondary structure or ATP binding in the case of the E29Q mutant. Moreover, mutations at Glu32 did not affect secondary structure (except for E32G) but reduce ATPase activity. Activity was abolished when both residues (E29Q/E32Q) are mutated.

      The authors extended their study to another GHKL ATPase, aqGyrB. Their findings further supported the cooperative function of the corresponding acidic residues in aqGyrB (Glu48 and Asp51) during ATP hydrolysis. Mutation of these residues partially impaired ATP hydrolysis without affecting protein secondary structure. ATPase activity was completely lost in the double mutant E48Q/D51M. While the E48Q mutant retained the ability to bind ATP, the E48A mutant did not. High-resolution structures of the WT and E48A, E48Q, D51A and D51N mutants of the aqGyrB NTD demonstrated that nucleophilic water positioning depended on these residues. E48 played a dominant role in water positioning and is critical for stabilising ATP lid formation and associated conformational changes, whereas D51 contributed cooperatively to catalysis.

      The authors investigated the functional impact of mutating the corresponding residues in the human MutL homologs PMS2 and MLH1. Clinical variants consistently exhibited reduced or abolished ATPase activity, providing a potential molecular basis for Lynch syndrome, through impaired DNA mismatch repair.

      Lastly, through evolutionary analysis, the authors inferred that the second acidic residue was likely present in the common ancestor of MutL, GyrB, and MORC proteins, but was lost in the case of Hsp90.

      Strengths:

      (1) This study contains a detailed structural and biochemical analysis of a biologically important set of GHKL ATPases. The authors identify a second acidic residue that is conserved and contributes to catalysis in a large subset of GHKL ATPases. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed, which involves cooperative and partially overlapping roles for the catalytic residue pair. This revised mechanistic model is invaluable for the interpretation of clinical variants of GHKL ATPases such as PMS2 and MLH1.

      (2) The work described was performed to an excellent and rigorous technical standard. The structural and biochemical data are sound. The evidence supporting the claims is compelling.

      Weaknesses:

      (1) The identification in this study of a second acidic residue contributing to catalysis but not absolutely essential for catalysis is a useful finding. However, given that many structures of GHLK ATPases have been determined with different nucleotide analogs bound and that the essential role of the first acidic residue is well established, the importance and scope of the advances described here remain focused within the field of study of GHKL ATPases.

      (2) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of any effect of disease-linked mutations in GHKL ATPases would have strengthened this study.

      (3) The effect of other aqMutL NTD E32 mutants, particularly, the E32K mutant on ATP binding remains unclear, although experimental assessment of nucleotide binding would be challenging due to the high protein concentrations required for the equilibrium dialysis assay.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) A major question that remains is why the mutations have so much more detrimental effect in MutL (100-fold lower k<sub>cat</sub>/K<sub>M</sub>) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      We agree that the quantitative effects of the mutations differ between MutL and GyrB. However, we do not think that this difference argues against conservation of the catalytic mechanism. The trends of the mutational effects are highly consistent between the two enzymes. In both proteins, replacement of the conserved catalytic glutamate with Ala (E29A in aqMutL and E48A in aqGyrB) abolished ATPase activity and ATP binding, whereas replacement with the isosteric amide residue (E29Q and E48Q), which preserves hydrogen-bonding capability but lacks proton-accepting capacity, retained ATP binding and measurable ATPase activity. Likewise, substitutions of the second acidic residue (E32Q in aqMutL and D51N in aqGyrB) also retained substantial ATPase activity despite the loss of proton-accepting capability. Most importantly, simultaneous substitution of both acidic residues (E29Q/E32Q in aqMutL and E48Q/D51N in aqGyrB) completely abolished ATPase activity in both enzymes. Therefore, although the magnitude of the activity reduction caused by the individual mutations differs between MutL and GyrB, the qualitative pattern is essentially identical. We therefore think that the proposed catalytic mechanism is conserved, while the quantitative differences likely reflect differences in the local catalytic environment rather than differences in the underlying mechanism.

      (2) The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the acidic residues and their mutants is not shown.

      We have added Supplementary Fig. S2, which shows the 2F<sub>o</sub>–F<sub>c</sub> electron density maps around residues Glu48/Asp51 (or their substituted residues) in the wildtype and all mutant structures. The following sentences have been added in the revised manuscript:

      “To show that the introduced substitutions were unambiguously supported by the crystallographic data, electron density maps around residues 48 and 51 are shown in Supplementary Fig. S2.” (p. 4 line 198-200 in the revised manuscript)

      Reviewer #2 (Public review):

      (1) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of the effect (or likely effect) of disease-linked mutations in GHKL ATPases would have strengthened this study.

      We agree that extending the analysis to additional disease-associated variants in other human GHKL ATPases would further strengthen our understanding of the conserved catalytic mechanism and its clinical relevance. However, we believe that such a comprehensive analysis is beyond the scope of the present study, which focuses on establishing the fundamental catalytic mechanism shared between MutL and GyrB. We consider systematic functional and structural analyses of disease-associated variants across the GHKL ATPase family to be an important direction for future research. We have now added a statement to the Results and Discussion section to acknowledge this limitation and highlight this future perspective:

      “Although we focused here on pathogenic variants in the MutL homologs MLH1 and PMS2, extending similar structural and biochemical analyses to disease-associated variants in other human GHKL ATPases will be important for evaluating the generality and clinical relevance of the conserved catalytic mechanism proposed in this study.” (p. 6 line 303-306 in the revised manuscript)

      (2) In MLH1, the E37K mutation completely abolishes ATPase activity, but the corresponding mutations in aqMutL, aqGyrB, and PMS2 do not. It remains unclear why E37K in MLH1 leads to complete loss of activity, as the authors propose that water molecule positioning via the first acidic residue, as well as ATP lid stabilisation and associated conformational changes, should still be possible.

      We agree that the complete loss of ATPase activity caused by the MLH1 E37K variant cannot be explained solely by loss of the catalytic carboxylate. However, we note that the corresponding aqMutL E32K variant analyzed in this study also exhibited essentially no detectable ATPase activity, indicating that this phenotype is not unique to MLH1. It can be thought that the severe defect of the lysine variants arises not merely from loss of the acidic side chain but from charge reversal. We have clarified this point in the Results and Discussion sections:

      “In contrast, the E37K mutation in the MLH1 NTD completely abolished the ATPase activity under our assay conditions (Fig. 5B and Table 1) unlike the corresponding glutamine substitutions, which retained substantial residual ATPase activity in aqMutL, aqGyrB, and PMS2 NTDs. A similar complete loss of ATPase activity was also observed for the E32K mutant form of the aqMutL NTD. These observations suggest that the severe defect caused by the lysine substitution cannot be attributed simply to loss of the catalytic carboxylate. Instead, introduction of a positively charged side chain (charge reversal) is likely to perturb the local electrostatic environment. Structural characterization of the MLH1 E37K and aqMutL E32K mutant forms will be required to clarify the molecular basis of this severe functional defect.” (p. 6 line 281-289 in the revised manuscript)

      (3) The authors do not examine ATP binding in the E32 mutants of aqMutL NTD and the D51 mutants of aqGyrB, or AMPPNP binding of the NLH1 and PMS2 mutants. Hence, the relative contributions of the acidic residues to ATP binding and hydrolysis remain partially unclear.

      We performed additional ATP-binding experiments using the aqMutL NTD E32A and aqGyrB NTD D51A mutant forms. Both mutant forms exhibited ATP-binding activities comparable to those of the corresponding wildtype forms. These results support our conclusion that the second acidic residue primarily contributes to ATP hydrolysis rather than ATP binding, whereas the first acidic residue plays dual roles in ATP binding and catalysis.

      Although we agree that nucleotide-binding analyses of the MLH1 and PMS2 variants would be informative, these experiments were not feasible because the equilibrium dialysis assay requires high protein concentrations, which we were unable to obtain for the recombinant human MLH1 and PMS2 N-terminal domains.

      We have incorporated these new data into the Results and Discussion section:

      “In contrast to the E29A mutant form of the aqMutL NTD, the E32A mutant form exhibited ATP binding ability comparable to that of the wildtype form (Supplementary Fig. S1A), indicating that Glu32 does not contribute to ATP binding.” (p. 3 line 143-145 in the revised manuscript)

      “The D51A mutant form of the aqGyrB NTD retained ATP binding ability comparable to that of the wildtype form, indicating that Asp51 is not required for nucleotide binding (Supplementary Fig. S1B).” (p. 4 line 183-185 in the revised manuscript)

      (4) The ATPase assays for PMS2 and MLH1 (Figure 7 and Table 1) were performed with purification/solubility tags still present. Hence, it cannot be ruled out that these tags influence the measured activities.

      We thank the reviewer for raising this important point. We agree that the possible influence of the purification/solubility tags on the absolute ATPase activities of the PMS2 and MLH1 NTDs cannot be completely excluded. However, the wild-type and mutant forms for each homolog were analyzed using identical constructs under the same experimental conditions. Therefore, the affinity/solubility tags are unlikely to affect the relative comparisons of the mutational effects. Furthermore, because the affinity tags are located at the N terminus and are distant from the ATPase active site, they are unlikely to directly perturb the catalytic center.

      (5) The authors suggest that the two-acidic-residue mechanism proposed in this study could be shared among several GHKL ATPase families, yet they also state that the hydrogen-bonding network was not observed in MutL and MORC family proteins. This raises doubt about how conserved the mechanism is, e.g., in MutL and MORC proteins.

      We thank the reviewer for this insightful comment. Our proposed mechanism is based on the cooperative catalytic roles of the two conserved acidic residues, namely the involvement of the first acidic residue in ATP binding and nucleophilic water positioning and the role of the second acidic residue in proton abstraction. In contrast, the Glu48–Gln340 hydrogen-bonding interaction described in aqGyrB was proposed only as a structural feature that may modulate the contribution of the first acidic residue to ATP binding. It is not an essential component of the catalytic mechanism proposed in this study. Therefore, the absence of this particular hydrogen-bonding network in the currently available structures of MutL and MORC proteins does not argue against conservation of the catalytic mechanism itself.

      Recommendations for the authors:

      Reviewing Editor Comments:

      One of the structures (Crystal Structure of the E48A variant) has relatively poor statistics in the PDB validation report. Please improve this structure.

      We performed additional refinement of the E48A crystal structure. This resulted in a clear improvement in the overall model quality, with the Ramachandran favored residues increasing from 93.4% to 95.4%, the percentage of side-chain outliers decreasing from 6.1% to 1.4%. The refined structural model has been used throughout the revised manuscript, and the updated refinement statistics are provided in Table 2.

      Reviewer #1 (Recommendations for the authors):

      Please show conventional density maps (e.g., sigmaA weighted 2fo-fc maps).

      This comment is closely related to Comment (2) in the Public Review by the Reviewer #1. In response, we have added Supplementary Fig. S2, which presents conventional σA-weighted 2F<sub>o</sub>–F<sub>c</sub> electron density maps around the catalytic acidic residues in the wild-type and mutant aqGyrB structures.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding the analysis of clinical variants, it would be informative to note that the second allele is lost before tumor growth in Lynch syndrome.

      “Therefore, these variants might contribute to the development of Lynch syndrome by weakening the ATPase-driven regulatory functions of MutL.” (p. 6 line 280-281 in the original manuscript) has been changed to:

      “In individuals carrying these germline variants, subsequent loss or inactivation of the remaining wildtype allele would leave only the ATPase-defective MutL protein, thereby compromising mismatch repair and promoting tumorigenesis.” (p. 6 line 296-298 in the revised manuscript)

      (2) P. 4, in the paragraph "Conserved roles of two acidic residues of aqGyrB in ATP hydrolysis", the E48Q mutant retains approximately one third of the WT activity, not one quarter as stated in the text (Table 1). Additionally, later in the article, the D51 mutant is reported to retain approximately one-sixth (~17%) of the WT activity, rather than ~25% as written.

      We thank the reviewer for carefully identifying these inconsistencies. The text has been corrected to accurately reflect the data presented in Table 1: “…one third of the wildtype activity” (p. 4 line 180) and “…retaining ~16%...” (p. 4 line 186 in the revised manuscript)

      (3) P. 6, lines 275-276, this sentence should be rephrased for clarity, as the authors note at the end of page 5 that not all members of the GHKL ATPase family possess this second acidic residue.

      “…this second acidic residue plays a conserved and functionally significant role in ATP hydrolysis across the GHKL ATPase family.” in the original manuscript has been changed to:

      “…this second acidic residue plays a conserved and functionally significant role in ATP hydrolysis among some members of the GHKL ATPase family.” (p. 6 line 292 in the revised manuscript)

      (4) P. 9, in the "Data Accessibility Statement", the PDB code 23UY is missing. This entry corresponds to the crystal structure of the D51A mutant of aqGyrB NTD and should be included.

      The Data Accessibility Statement has been revised to include the code 23UY. (p. 9 line 454 in the revised manuscript)

      (5) P. 14, the table should be labelled "Table 2. Data collection and refinement statistics for the aqGyrB NTDs", rather than "Supplementary Table 2", to ensure consistency with how it is cited in the main text.

      The table title has been corrected from "Supplementary Table 2" to "Table 2”. (p. 4 line 198 in the revised manuscript)

      (6) It is difficult to determine from the figures whether the magnesium ion is positioned equivalently in aqMutL and aqGyrB. Did the authors observe any differences in ion positioning?

      To facilitate direct comparison of the catalytic Mg<sup>2+</sup> ion between the aqMutL and aqGyrB NTDs, we have added Supplementary Fig. S3, which shows a structural superimposition of the ATPase active sites of the two proteins:

      “Structural superposition of the aqGyrB NTD and aqMutL NTD revealed that the catalytic Mg<sup>2+</sup> ion occupies essentially the same position in the two ATPase active sites (Supplementary Fig. S3), indicating that the metal-binding geometry is highly conserved, where the Mg<sup>2+</sup> ion is coordinated by the side chain of the conserved Asn, AMPPNP, and surrounding water molecules. Neither Glu48 of aqMutL nor Asp51 of aqGyrB directly coordinated the Mg<sup>2+</sup> ion.” (p. 5 line 201-205 in the revised manuscript)

      (7) The authors should discuss the interaction between the aqGyrB NTD, Mg<sup>2+</sup>, and ATP during the binding step. In the case of the E48A mutant, where ATP binding is lost, does E48 directly establish contacts with Mg<sup>2+</sup>, or is another residue involved (with conformational changes preventing this interaction)?

      Our structural analyses indicate that Glu48 does not directly coordinate the catalytic Mg<sup>2+</sup> ion. Instead, as shown in Supplementary Fig. S3, the Mg<sup>2+</sup> ion is coordinated by the side chain of Asn52, AMPPNP, and surrounding water molecules. We have clarified this point in the Results and Discussion sections:

      “Structural superposition of the aqGyrB NTD and aqMutL NTD revealed that the catalytic Mg<sup>2+</sup> ion occupies essentially the same position in the two ATPase active sites (Supplementary Fig. S3), indicating that the metal-binding geometry is highly conserved, where the Mg<sup>2+</sup> ion is coordinated by the side chain of Asn52, AMPPNP, and surrounding water molecules. Neither Glu48 nor Asp51 directly coordinated the Mg<sup>2+</sup> ion.” (p. 5 line 201-205 in the revised manuscript)

      (8) Figures 1 and 3: Use ribbon representation and no shadows, at least for the inset panels, to enhance clarity and interpretability.

      We have revised Figures 1 and 3 by displaying the protein structures in ribbon representation and removing shadows from the inset panels.

      (9) Combine Figures 1 and 2, and combine Figures 3 and 4.

      Following the reviewer's recommendation, we have combined the original Figures 1 and 2 into a single figure and the original Figures 3 and 4 into another single figure.

      (10) Figure 5: Zoom in further and remove shadows. The current panels are not very effective in highlighting how ATP is bound by the different protein variants.

      Figure 5 has been revised by increasing the magnification of the ATP-binding sites and removing shadows from the structural renderings.

      (11) Figure 8. Add a scale bar to show evolutionary distance.

      We thank the reviewer for this helpful suggestion. To provide information on evolutionary distances while preserving the clarity of the main figure, we have added a new Supplementary Fig. S5 showing the same phylogenetic tree with branch lengths proportional to the inferred evolutionary distances and an evolutionary distance scale bar. Figure 6 has been retained in its simplified form with equal branch lengths to facilitate visualization of the ancestral-state reconstruction, and we have clarified this distinction in the Materials and Methods section:

      “For visualization purposes, branch lengths were not scaled and were displayed with equal lengths in Fig. 6. The corresponding phylogeny with branch lengths proportional to the inferred evolutionary distances is provided in Supplementary Fig. S5.” (p. 9 line 429-432 in the revised manuscript)

    1. eLife Assessment

      This study offers valuable insights into brain responses to somatosensory and visual temporal and spatial tasks in the auditory cortex of Deaf and hearing individuals. The evidence for a sensorily-bound representation in the deprived cortex -- rather than abstract -- is solid; however, the study will benefit from some additional analyses to better clarify and contextualise the results. This work will be of broad interest to neuroscientists investigating brain plasticity and development.

    2. Reviewer #1 (Public review):

      Summary:

      The authors conducted a carefully constructed experiment to test reorganization in the auditory cortex in deafness in response to task (spatial vs. temporal working memory) and modality (visual vs. somatosensory). They found a complex pattern of results, which included changes to univariate response strength in deafness that differed between the primary and association auditory cortex. HG showed a preference for the somatosensory working memory task, whereas STG/S responded more in both modalities for the temporal task. They further showed multivariate similarity of their results to models representing both task and modality in both groups, which were increased in deafness. Curiously, the task effect was for a sensorily-bound model, which shows mid-level representation, and not high-level task ("metamodal") code.

      Strengths:

      I appreciated the matched design, rigorous analysis, and careful interpretation of the nuanced results.

      Weaknesses:

      Only minor weaknesses: behavior and residual hearing can be better controlled.

    3. Reviewer #2 (Public review):

      Manini and colleagues present an interesting study on the consequences of early deafness on the organization of temporal regions chiefly engaged in audition in hearing people. Mainly relying on representational similarity analyses, they show that the auditory cortex in deaf individuals represents information about task, sensory modality, and somatosensory frequency. Critically, task and modality representations were also found in the auditory cortex of hearing individuals. There were significant differences between groups, implying that these representations are enhanced as a consequence of deafness.

      Overall, I feel that the paper could gain in clarity and impact if the hypothesis space tested in the introduction and discussion was made clearer, if some new analyses were provided to support some claims, and if the authors better matched their conclusions to the observed results.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript examines functional plasticity in auditory cortices of people born deaf (or early deaf). Deaf participants (N=13, all native BSL signers) and hearing controls (N=18) performed a delay-match-to-sample working memory task in visual and somatosensory modalities, attending either to frequency (temporal task) or spatial pattern (spatial task). Using fMRI and univariate and representational similarity analysis (RSA), the authors test what type of information is represented in auditory areas of hearing controls and deaf participants. Different types of tasks and stimulus features are examined, including low-level sensory features (vibration/movement frequency and spatial position of stimuli on screen/hand) as well as task features (attending to space vs. frequency) and stimulus modality (somatosensory vs. visual).

      There are a number of interesting findings. In early auditory cortices (right Heschl's gyrus), only Deaf participants show above baseline responses and only in the somatosensory task. In secondary auditory/multisensory STS, only deaf participants show above-rest responses to both visual and somatosensory stimuli. In RSA analysis, d/Deaf, but not hearing participants, show sensitivity to the frequency of somatosensory vibration when finger position is held constant (SFRm model). RSA in auditory cortices finds sensitivity to sensory modality (visual vs. somatosensory) information in both hearing and deaf participants, with an enhanced effect in the d/Deaf group. A similar d/Deafness enhancement effect was observed for task (frequency vs. spatial) but only within modality, not across, suggesting a less abstract representation.

      Strengths:

      The paper has a number of strengths. Running tasks in multiple modalities in the same study is a technical challenge and adds valuable information, since, as it turns out, early auditory areas of deaf people are sensitive to somatosensory but not visual frequency. Moreover, it makes it possible to look for modality coding.

      The RSA analysis finds sensitivity to task and modality features not detectable with univariate analyses.

      Overall, the findings of differences across modalities (visual and somatosensory) in auditory cortices are very interesting. Having parallel tasks in the two modalities makes it all the more important that only somatosensory stimuli activate HG. The discussion of this finding and ideas about alternative possible paths of somatosensory and visual information to the auditory cortices in d/Deaf individuals is also interesting.

      The authors conclude that there is both evidence for shared and different function across hearing and deaf groups, and this makes sense. It's refreshing that the authors acknowledge that the simplicity dichotomy of preservation vs. change present in the literature and presented in the introduction turns out not to explain the findings.

      Weaknesses:

      The participant sample is not very large, but this is a difficult-to-recruit population, and it is within an acceptable range since the authors have taken care to run a robust study design and collect a large amount of data from each person.

      Some weakness of the paper includes incomplete presentation of the results and conclusions that do not follow from the data.

      The results are presented in a way that makes it hard to track what is significant and how the auditory ROIs are different or similar to the control regions. It is also difficult to connect the written text results with the figures. In some cases, the figures look like there is no significant effect, but the results report that there is one.

      It is not clear that three-way interactions were tested for (e.g., group, by task, by modality). This complicates the interpretation of significant two-way interactions, e.g., task by group. For example, in STG a task by group interaction is reported, but the plot suggests that this is driven primarily by the somatosensory modality.

      The paper suggests that frequency/temporal information is coded in auditory areas of d/Deaf participants but not location/spatial information. This is stated in the Results and in the Discussion as a major point. But this is not quite true in the somatosensory case and not true in the visual case at all, as far as I can tell. In the somatosensory case, since frequency is only coded in a finger-dependent manner, location is in fact coded in this regard. This pattern differs from what is observed in the somatosensory cortex, where frequency is coded in a finger-dependent and independent manner as well as location as such. In the visual case, it seems like temporal, i.e., frequency information, is not coded at all in the auditory ROIs. Although it is also puzzling that visual frequency is not coded in the visual 'control' ROI. An incomplete presentation of results motivates some conclusions that are not warranted, e.g., coding in the auditory cortex reflects a preservation of its function - i.e., frequency but not location coding.

      A second related issue is that some dimensions (e.g., visual frequency, modality-independent task) show no neural response anywhere in the brain, including in the canonical visual, somatosensory, and amodal networks. For these dimensions, the current experiment does not offer a good test. This is fine but should be clearly stated in the results and the Discussion so there is no confusion about which hypotheses are really tested. Right now, the results say things like "the control ROI shows the expected pattern", but in some cases it's more complicated and failures to observe effects constrain what can ultimately be expected in auditory areas. This is okay, but needs to be made clear in the results and discussion, and auditory results need to be interpreted in this context.

      Relatedly, the paper has no whole cortex searchlight analyses, and it is not clear why. If nothing comes out in the small sample of n=13, that is okay, but at least the collapsed hearing and d/Deaf sample data should be shown for each dimension. This will give the reader a sense of what is to be expected and contextualize the results.

      Some claims are made which are not supported by the data: "However, it is likely that the representation of somatosensory frequency in deaf individuals relies on the same mechanisms used to represent auditory temporal frequency in hearing individuals." Likewise, the paper goes on to say, 'the underlying computations might be the same'. True, they might be, but they also might be different. No evidence is presented to support one or the other hypothesis, so both should be stated, and it should be stated that these cannot be distinguished based on the presented data.

      Conclusions-wise, the paper sometimes makes sweeping claims that go well beyond what the evidence supports and fails to provide caveats. The first paragraph of the Discussion states: "Overall, these findings suggest that crossmodal plasticity relies on representational and functional configurations that are present across individuals and modulated by sensory experience." Such a sweeping conclusion about cross-modal plasticity in general is not supported or refuted by the present data. There is no evidence that 'representations' or 'functional configurations' are the same across groups. What are functional configurations? What representations are shared? Second, it is far too general to make claims about 'cross-modal plasticity' based on one study with one population.

    5. Author response:

      We would like to thank the editor and reviewers for their thoughtful and constructive feedback. We appreciate the time and care devoted to reviewing our manuscript, as well as the recognition of rigorous experimental design, the technical challenges involved in conducting an fMRI study with two sensory modalities and two tasks in both deaf and hearing participants, and the value of the findings for understanding crossmodal plasticity and cortical organisation in deafness. We are encouraged by the overall assessment of the study, and appreciate the suggestions for strengthening the manuscript. Below, we provide a summary of how we plan to address the reviewers’ comments in our formal revision of the manuscript:

      (1) Additional analyses

      (a) We will incorporate behavioural performance measures into the relevant analyses to disentangle potential behavioural contributions to the observed effects.

      (b) We will calculate the noise ceiling value for each of the RSA analyses. 

      (c) We will conduct a correlation analysis between RDMs of auditory and control regions, to investigate the similarity between these computations and whether this is influenced by sensory experience.

      (d) Regarding the suggestion to conduct whole-brain searchlight analyses, we respectfully do not believe that this approach would address the primary research question of the study, namely whether and how representations within auditory cortex differ between deaf and hearing individuals. Our central hypotheses specifically concern representational content within predefined auditory cortical regions, making the ROI-based approach the most appropriate and sensitive method for testing these questions.

      Furthermore, the searchlight approach would require adequately powered group comparisons at the whole-brain level. Given the challenges associated with recruiting deaf native signers participants and the resulting sample size, we do not believe the study is sufficiently powered to draw reliable conclusions from this analysis. We will further clarify this rationale in the revised manuscript.

      (2) Presentation of the results

      We will revise the presentation of the findings to better guide the reader through the analyses and facilitate interpretation of the figures. In particular, we will ensure that significant effects, interactions, and their relationship to the corresponding figures are described more explicitly throughout the manuscript.

      (3) Revision of the discussion

      Following the reviewers’ feedback, we will revise the Discussion to more clearly distinguish between results that directly support a conclusion and hypotheses that remain speculative.

    1. eLife Assessment

      This study provides a useful investigation of machine learning approaches that can lessen potential gaps in the prediction of behaviour from brain imaging data across majority and minority samples. The authors provide incomplete evidence to suggest that domain adaptation methods can mitigate these gaps. The analysis would benefit from further testing of model generalisability and the inclusion of recommended workflows for how the tested approaches can be used in future research. This work will be of interest to scientists using machine learning in brain imaging.

    2. Reviewer #1 (Public review):

      Summary:

      The present report describes an investigation into the use of machine learning techniques to improve cross-racial/ethnic performance of brain models of cognitive function. The authors tested several approaches to boost prediction of NIH cognitive toolbox scores using brain imaging data (function, structure) for minoritized (Black) participants in the ABCD Study sample compared to white (majority) participants. Structural (e.g., volume) measures showed the greatest performance gap, and a balanced weighting method showed the greatest performance gain across features. The authors conclude that supervised domain adaptive methods can improve models for cognitive prediction and mitigate cross-racial/ethnic performance disparities.

      Strengths:

      This investigation makes some headway into issues by identifying computational methods that may help to improve some models for limited outcome variables (i.e., general cognitive performance). Addressing racial/ethnic disparities in brain imaging research has significant implications for generalizability of findings and for the practical utility of imaging findings in the wider population. A comparative approach to evaluate the improvements in a "prediction gap" across various methods could have benefits for neuroimaging beyond racial/ethnic disparities. The use of the ABCD Study, given its deep phenotyping of individuals, is also a benefit.

      Weaknesses:

      Despite its strengths, there are several large conceptual and related methodological issues that impact its conclusions and the overall utility of the approach. The sample selection approach limits insight into likely drivers of the performance gap (e.g., socioenvironmental disparities known to exist between groups and associated with neurodevelopment), and in so doing ignores a critical component of understanding racial brain differences, particularly in relation to cognitive functions. Further, while the relative gaps in performance of a single cognitive score across features are well described, the actual performance (and therefore relative benefit to these techniques) is unclear. Specific examples include the following.

      (1) The overarching conceptual issue with the manuscript is a lack of engagement with a substantial and growing evidence base on the drivers of racial disparities in brain imaging which impact model performance. Racial/ethnic groups in the US (and other regions of the world) are not equivalent in terms of developmental environments that shape brain function and structure (see Harnett et al., 2023, Neuropsychopharmacology; Ricard et al., 2023, Nature Neuroscience; Cardenas-Iniguez & Gonzalez, 2024, Nature Neuroscience for some overview here). The socioenvironmental disparities inherent to race in the US further shape cognitive development and brain associations with cognitive performance (e.g., Marek et al., 2025, Science). The framing of the manuscript focuses almost exclusively on broad sampling issues, and in doing so treats racial/ethnic variability as if it reflects statistical abnormality rather than a critical component of understanding human brain development. This lack of contextualizing racial disparities significantly impacts the overall utility of the proposed approach and the conclusions of the manuscript.

      (2) In relation to the above, another conceptual issue in this approach of using a majority to inform minority brain associations with cognitive variables is an assumption that minority brain patterns should match the majority, rather than developmental stressors inducing alternative brain-weighting to predict outcomes. This framework does not assess this possibility and may in fact obscure such an outcome, limiting our inferences into neurodevelopment.

      (3) Another conceptual/methodological issue here is the use of "matched groups" for analysis. The specifics of matching are fairly vague, but given the description one would assume the w/B groups are matched on a number of behavioral/socioenvironmental variables, which is a significant issue for interpretability and applicability. As noted, w/B groups in the US (and the ABCD Study) differ substantially across variables; matching has the likely consequence of creating a highly non-generalizable sample, particularly when the minority group is restricted to N = 10.

    3. Reviewer #2 (Public review):

      In the manuscript "Supervised domain adaptation mitigates cross-ethnicity prediction errors in neuroimaging-based cognitive prediction", the authors investigated the efficacy of data adaptation techniques to reduce ethnicity-related prediction bias in neuroimaging-based cognitive prediction. They found that data adaptation algorithms, particularly balanced weighting, contributed to mitigating ethnicity-related performance disparities. Furthermore, these bias mitigations could be achieved without requiring a large set of data from the underrepresented ethnic group. This study addressed an important concern in the field of neuroimaging-based behaviour prediction, providing many intriguing results. Nevertheless, the manuscript also suffers from a lack of coherent methods design, the unorganised presentation of information, and the lack of in-depth discussion of results.

      The conclusions claimed by the authors are sometimes over-generalised and not fully supported by the study outcomes. Overall, this study demonstrated strong technical designs and convincing statistical analysis for the main outcomes, although clearer presentation would be needed to convey the messages in the manuscript.

      The central investigation of this study is whether domain adaptation techniques improve ethnicity-related performance disparities. However, these improvements were only measured against a very weak baseline model, where a small set of African American (AA) subjects were added to the training sample consisting purely of White American (WA) subjects. While the authors recognised that balancing the training sample could already mitigate the ethnicity-related disparities, they considered that such approaches are unfeasible in their experimental scenario, where only a small amount of AA data were available. However, as Li et al. (2022) showed, a balanced sample of around 90-150 AA subjects could already reduce the ethnicity-related bias. Even from a practical standpoint, this balanced sample approach would be a more valid baseline for domain adaptation models to compare against.

      The authors made two main conclusions: that domain adaptation methods reduced ethnicity-related bias, and that balanced weighting performed the best and the most stably. Both claims were over-generalised to some extent. First, the adaptation benefit claimed in the first conclusion is not seen in the functional connectivity (FC) modality, which is the most popular modality for neuroimaging-based prediction of behaviour. This difference in adaptation benefit across modalities is an important finding that is meaningful for future studies, the omission of which also removes interesting insights that the audience could take away from this article.

      Second, the judgement of prediction performance is based on the area under the improvement curve (AUIC) metric, which summarises a model's performance across different availability of labelled AA data. As a result, the analysis of prediction performance naturally favours algorithms that could perform well with a small amount of added AA data. On the one hand, this provides an easy decision point for users to pick an algorithm to use without being concerned about data availability. On the other hand, important insights could be overlooked with the oversimplified recommendation of balanced weighting. As the authors have also observed, in some cases, domain adaptation strategies do not improve ethnicity-related bias more than the non-adaptation baseline. If the message is to recommend simple, low-cost strategies to reduce ethnicity-related prediction bias, it would be misleading not to note that the simplest and lowest-cost strategy could also be non-adaptation methods sometimes.

      Regardless, for the general audience, the underlying assumptions when interpreting the AUIC metric are not immediately clear, which could cause the conclusions to be misleading. Apart from aggregating over different amounts of available AA data, the statistical comparison of AUIC gain across data adaptation algorithms also did not account for the impact of brain phenotype modalities. Even though the upstream analyses have confirmed that adaptation benefits vary greatly across brain modalities, this major observation was not followed in the final analysis where conclusions were made about which algorithm performed the best. Based on visual inspection of Figure 3b, it may be suspected that PRED performed better than or comparably to balanced weighting when task contrasts based on the Destrieux atlas were used.

      Finally, the findings from this study align with the common hypothesis that ethnicity-related prediction bias originates from disparities already manifested during data collection and preprocessing. As the authors have noted, the modalities with the most tendency for ethnicity-related bias are the anatomical ones, including all three volume-based modalities (cortical volume, T1 and T2 subcortical volume) in the top ten phenotypes with the largest performance gap. Most prominently, brain features in the occipital pole, frontal pole, and a range of subcortical areas were found to contribute highly to adaptation gain. Subcortical areas are often reported to show noisier measurements compared to cortical areas, whereas the poles of the brain are likely more strongly warped/distorted during alignment to a standard template. From a data quality perspective, these results support the interpretation that ethnicity-related prediction bias may stem from loss of data quality during data collection or preprocessing. In the prediction models based on anatomical brain features, data adaptation methods may have helped to address these disparities in the data, without the more resource-intensive need to improve the bias in preprocessing pipelines.

      Li, J., Bzdok, D., Chen, J., ... Genon, S. (2022). Cross-ethnicity/race generalization failure of behavioral prediction from resting-state functional connectivity. Science Advances, 8(11), eabj1812.

    4. Reviewer #3 (Public review):

      The manuscript frames its work in fairness and disparities but does not show or directly test that its approach decreases differences between White and African Americans. While it is stated that the objective is not to equalize performance across groups, large parts of the paper repeatedly claim that the methods mitigate cross-ethnicity disparities and improve fairness. Improving prediction in African American participants relative to a non-adapted model is not necessarily the same as reducing the disparity between African American and White American participants. The adapted model should be evaluated in both groups, and the post-adaptation performance gap should be reported directly.

      Additional prediction performance measures are needed. For example, in Li et al, different results and conclusions are made with MSE and the correlation between observed and predicted variables. In that paper particularly, aggression measures showed better correlation in African Americans but better MSE in White Americans. Such differences are important to note as they likely suggest different mechanisms.

      Similarly, characteristics of the cognitive outcome need to be understood. For example, differences in MSE or MAE may reflect a difference in variance between the groups. The group with a larger variance will have a larger MSE. Correlation or other performance measures that are invariant to different mean or variance shifts can be helpful here.

      While the authors note that for the paper they treat racial and ethnic backgrounds interchangeably, I do not think that is the best given the differences between them and the impact and history they have in American culture. Overall, the authors likely need to do a better job conceptualizing their results in the history of minoritized populations in the United States. It is immensely important not to treat them as biological domains without considerable qualification and to avoid language implying that observed domain differences are intrinsic properties of racial groups.

      Changes in feature weight are not a proper way to identify the mechanisms of improved performance. At most, these analyses characterize how model coefficients change when target-group data are incorporated or upweighted.

      The cross-validation strategy is suboptimal. First, the use of the matched splits of the ABCD data introduces data leakage. To match a validation set to the training set in such a manner requires that each split knows about the other split's characteristics. That is data leakage. Though the impact could be small. Second, African American breakdowns are not balanced across sites and scanners. Domain adaption methods may be learning a shortcut or proxy for African American like site, scanner, or something else. A likely better approach would be some sort of leave X sites out approach, where a model is trained on White Americans from a set of sites, adapted with African Americans from those sites, and applied (with and without adaptation) to the White and African Americans from the left-out sites.

      Baseline models for comparisons to the domain adaptation are missing. Some simpler ones include a target-only model trained on the same 10-100 African American participants and a pooled model with a group indicator and group-by-feature interaction. Without these comparisons, it is difficult to know whether balanced weighting is learning target-specific neurobiological information or merely recalibrating the prediction distribution.

      There are a few statistical issues:

      (1) The repeated MAE estimates are therefore not independent observations. Paired t-tests cannot be applied across repetitions. Subject-level bootstrap or permutation procedures that repeat the complete training and testing process are needed

      (2) The caption describes approximate 95% confidence intervals as {plus minus}1.96 × SD/n. Conventionally, the standard error would involve SD/sqrt(n). However, even if corrected, there would still be issues about the dependence among the overlapping resamples.

      (3) Ten repetitions are likely insufficient, especially in the case of ten target participants. Results at n = 10 may be extremely sensitive to which children are selected. The authors should use substantially more repetitions and report the full distribution of results.

      (4) The Friedman and Wilcoxon comparisons treat the 80 imaging phenotypes as the observational units. These phenotypes are highly dependent because they are derived from the same participants, many use overlapping images, and numerous task contrasts and structural measures are strongly correlated. This non-independence can make the comparison among adaptation methods look much more precise than it is. A hierarchical analysis by modality or a resampling strategy that preserves dependence among phenotypes would be more appropriate.

      (5) The gap metric and AUIC are difficult to interpret. Gap is the absolute relative difference between target-group and source-group MAE, normalized by source-group MAE. It is sensitive to the denominator and may produce large values whenever source-group. MAE is relatively small. Reporting signed raw MAE differences and MAE ratios alongside this derived score would help. Similarly, the AUIC combines errors with an arbitrary sequence of target-sample sizes. IStatistically significant differences in AUIC do not necessarily indicate practically meaningful differences among methods.

      (6) The correlation between baseline gap and adaptation gain is partly tautological. Those with the widest gaps have the most room for improvement and likely thus show the greatest improvement. While still of value, the authors may want to tone down their interpretation of the correlation and describe its limitation.

      (7) Given that the sample sizes vary from approximately 4,000 to more than 11,000 depending on modality, the authors may want to consider a reduced sample matched in size across modalities. It is hard to fully know if the conclusion that connectivity is more robust given the wide-scale differences in sample size and feature dimensionality.

      (8) PLS are sensitive to many factors like scaling and collinearity. Many recent papers have been written about their limitations when used for subtyping. Some of these hold for prediction too. I think showing the results are consistent with different prediction algorithms is needed. SVR and ridge regression are two common methods for regression prediction with neuroimaging data.

      (9) The feature interpretation is partly circular. The method with the largest performance gain is selected, and its coefficient changes are then used to explain that gain. A method designed to give target observations greater influence will unsurprisingly change its coefficients more than naïve inclusion.

      (10) The practical and ethical deployment scenario is underspecified. Supervised adaptation requires labelled cognitive outcomes from the target population and, as currently framed, may require choosing a model based on an individual's racial category. What are the implications of deploying race-specific models that need to be considered? It is not self-evident that this approach is preferable to developing a broadly representative model or directly modeling the social and technical sources of distribution shift.

      (11) The paper is worded and interpreted much too strongly. The current study supports the conclusion that, within ABCD, giving a small labelled target-group sample greater influence can sometimes improve held-out target-group MAE relative to naïvely adding the same participants. It does not yet establish that the method improves fairness or identifies mechanisms of racial bias. Likewise, the abstract and conclusion overstate the results. The abstract states that all adaptation methods reduced target-group prediction error, while the Results show near-zero or negative benefits for several functional-connectivity phenotypes and instability of PRED and interpolation below 30 target participants. Similarly, "substantially reduce disparities," "improve equity," "consistently," and "practical path forward" are stronger than the analyses support. Finally, the limitations section is incomplete and omits the more consequential limitations.

      (11) That only ten labelled participants are needed to change the results is troubling. This is a shockingly low number. Giving ten target observations disproportionate influence can move the fitted model, particularly when the balanced-weighting ratio is high. A measurable MAE change is therefore possible, but it may reflect a shift in intercept or slope rather than learning a stable target-group brain-cognition relationship. Further, the manuscript does not report the numerical improvement for the n = 10 condition in the text or a table. Visual inspection of Figure 4 suggests reductions of roughly 0.10-0.25 standardized MAE units for some high-gap structural phenotypes, approximately 10-20%, while low-gap connectivity phenotypes show little or no gain.

      (12) The study lacks genuine external validation, which may be needed to fully convince readers that such a low number of subjects is needed to reduce biases.

    1. eLife Assessment

      This study addresses a fundamental question about how large-scale brain networks interact, and specifically how the default mode network exchanges information with sensory cortex. The analyses provide solid evidence for the claims made in the paper. The findings should be of broad interest to researchers studying brain network organization and dynamics.

    2. Reviewer #1 (Public review):

      Summary:

      This paper leverages 7T fMRI data from the Natural Scenes Dataset to investigate whether retinotopic coding the position-selective organization of visual responses structures spontaneous resting-state interactions between the Default Network (DN) and the Dorsal Attention Network (dATN). Using individualized network parcellations and population receptive field (pRF) modeling, the authors show that DN voxels can be split into two subpopulations based on their response to visual stimulation: those with position-specific positive BOLD responses (+pRFs) and those with position-specific negative BOLD responses (-pRFs). Critically, these subpopulations relate differently to the dATN during rest: -pRFs are anticorrelated with the dATN, +pRFs are positively correlated, and non-retinotopic DN voxels show no coupling. The anticorrelation (and positive correlation) is enhanced when DN and dATN voxels share visual field preferences. An event-triggered analysis suggests that retinotopic coding shapes both "top-down" (DN-initiated) and "bottom-up" (dATN-initiated) spontaneous activity transients, supporting the claim that the retinotopic scaffold is intrinsic to the DN. These findings challenge the prevailing view of global DN-dATN antagonism and suggest retinotopic coding as an organizing principle for cross-network communication.

      Strengths:

      The central finding that what looks like network-level independence between DN and dATN decomposes into structured, bivalent interactions organized by voxel-level visual field preferences is a compelling demonstration that macro-scale network descriptions can hide meaningful substructure. The logic of the analysis is clean: pRF properties are estimated from retinotopic mapping data and then used to predict resting-state coupling in completely independent scanning sessions. This cross-session, cross-modality design rules out many circularity concerns.

      The use of individualized multi-session hierarchical Bayesian parcellation (Kong et al.) to define DN and dATN boundaries within each subject is the right methodological choice for this question. Network boundaries in posterior cortex, where DN and dATN interdigitate most closely, vary considerably across individuals, and group-average approaches would introduce exactly the kind of misassignment that would most confound the result.

      The matched-vs-random pRF analysis is well-controlled. The authors demonstrate that cortical distance between matched and randomly matched dATN pRFs does not differ, effectively ruling out spatial proximity on the cortical surface as a confound. tSNR controls further show that signal quality differences do not drive the effect.

      The event-triggered analysis (Figure 3) is creative and adds genuine value. Showing that retinotopically-specific coupling persists during DN-initiated activity transients not only dATN-initiated ones is the key piece of evidence for the claim that the code is intrinsic to the DN rather than passively inherited through bottom-up visual drive.

      The result is observed consistently across all individual participants, which provides strong evidence for the robustness of the qualitative pattern despite the small sample size inherent to densely sampled designs.

      Comments on revised version:

      I'm content with the additional analyses and alterations to the writing that the authors have performed. I'm convinced that this work will spawn a very productive thread in the literature.

    3. Reviewer #2 (Public review):

      Summary:

      Using a public dataset of retinotopic mapping and resting-state data, the authors find that the default mode network has voxels that respond (positively or negatively) to visual stimulation at specific retinotopic positions, and that resting-state activity in these voxels is correlated with activity in more traditional sensory voxels with the same visual-location preference. The retinotopic specificity is bidirectional, such that high activity in default mode voxels drives activity only in voxels with matching receptive fields in sensory cortex, and vice versa. These findings are at odds with traditional views of the default mode network as having abstract (non-retinotopic) representations and competing (rather than cooperating) with external sensory representations.

      Strengths:

      This study continues an intriguing line of research about how default mode regions interact with sensory cortex. Demonstrating that there are structured interactions between these regions at rest, and that these interactions are in fact organized according to retinotopic location (as opposed to traditional views of representational format in the default mode network), provides a new framework for thinking about large-scale internal and external brain networks. The authors make use of a well-powered public dataset that allows for precise estimates of pRFs and individual-specific resting-state networks and develop a number of interesting analyses that characterize the relationships between DN and dATN voxels. The findings are exciting and could have a major impact on future studies in cognitive neuroimaging.

      The authors mention that these findings could shed light on internal/external interactions such as "anticipatory saccades or memory-guided attention," which is true, though I would argue that constructing DN representations of external stimuli is in fact even more fundamental than these specific cases (e.g. see Barnett and Bellana, 2025, "Situation models and the default mode network"). The "highways" identified in this study could play a vital role in real-world perceptual processes that are constantly translating external input into internal mental models.

      Weaknesses:

      (1) The criterion used for defining voxels as retinotopic seems very liberal. The authors show that only 5% of voxels have R^2>0.14 in a null analysis and therefore define voxels with R^2>0.14 as retinotopic. Although all the networks in Fig 1C show voxel distributions that differ from the null, the number of false positives above R^2>0.14 seems problematic, especially for the DN positive pRFs (red distribution) and to a lesser extent the DN negative pRFs (blue distribution). From visual inspection of the plot, the false discovery rate (fraction of voxels labeled as retinotopic that are false positives) looks like it would be greater than 50% for the DN positive pRFs. The authors do show that the positive pRF voxels have above-chance consistency across runs and also show in a supplementary analysis (Fig S5) that applying a stricter R^2 criterion yields similar results. These help to mitigate this concern, providing evidence that there are true positive voxels in this set which are driving the effects.

      (2) The claim that "voxel-level visual response profiles shape DN-dATN coupling during spontaneous resting-state activity" is well-supported for specific sub-groups of DN voxels, though it is unclear whether the overall DN-dATN correlation at rest is primarily driven by the pRF-tuned voxels investigated in this study.

      (3) The event-triggered analysis is effective at testing the bidirectional relationship between DN and dATN, with high activity in either network triggering a response in the other network. However, it would be helpful to show more validation that these "events" are meaningful windows of time to study, and that 13 TRs a typical length of time that activity is elevated during one of these events.

      (4) The framing of this paper relative to the authors past work, such as Steel et al. 2024 ("A retinotopic code structures the interaction between perception and memory systems") could be improved. The primary novelty here is that this paper examines resting-state data and individually defined whole-brain networks, showing that there are widespread spontaneous interactions between broad internal and external networks, but this distinction is not made explicit in the Introduction.

    4. Reviewer #3 (Public review):

      Summary:

      This paper addresses an important question (relationship between DN and dATN, and the role of retinotopic coding) and uses a set of novel analyses.

      Strengths:

      Important question, novel analytical approaches (pRF-informed functional connectivity analysis).

      Weaknesses:

      Some of the analyses are not described with sufficient clarity, especially the final analysis related to Fig. 3.

      Comments on revised version.

      Related to my previous comment 3), the removal of the labels "bottom-up" and "top-down" in the final analysis is a big improvement. However, I still don't fully understand how the 10 most aligned pRFs and the 10 most anti-matched pRFs are selected. The methods section on this has some ambiguity: "the 10 with the smallest Euclidean distance in RF center (x,y)". Does this mean that these are the pRFs closest to fovea? If not, what is the Euclidean distance referring to? Likewise, I don't understand how the anti-matched voxels are selected. This makes the interpretation of Fig 3 difficult.

      My previous comment about baseline activation was to compare the matched voxels with randomly selected voxels, instead of with anti-matched voxels. The authors responded that there was a technical difficulty with this.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper leverages 7T fMRI data from the Natural Scenes Dataset to investigate whether retinotopic coding, the position-selective organization of visual response structures, spontaneous resting-state interactions between the Default Network (DN) and the Dorsal Attention Network (dATN). Using individualized network parcellations and population receptive field (pRF) modeling, the authors show that DN voxels can be split into two subpopulations based on their response to visual stimulation: those with position-specific positive BOLD responses (+pRFs) and those with position-specific negative BOLD responses (-pRFs). Critically, these subpopulations relate differently to the dATN during rest: -pRFs are anticorrelated with the dATN, +pRFs are positively correlated, and non-retinotopic DN voxels show no coupling. The anticorrelation (and positive correlation) is enhanced when DN and dATN voxels share visual field preferences. An eventtriggered analysis suggests that retinotopic coding shapes both "top-down" (DNinitiated) and "bottom-up" (dATN-initiated) spontaneous activity transients, supporting the claim that the retinotopic scaffold is intrinsic to the DN. These findings challenge the prevailing view of global DN-dATN antagonism and suggest retinotopic coding as an organizing principle for cross-network communication.

      Strengths:

      The central finding that what looks like network-level independence between DN and dATN decomposes into structured, bivalent interactions organized by voxellevel visual field preferences is a compelling demonstration that macro-scale network descriptions can hide meaningful substructure. The logic of the analysis is clean: pRF properties are estimated from retinotopic mapping data and then used to predict resting-state coupling in completely independent scanning sessions. This cross-session, cross-modality design rules out many circularity concerns.

      The use of individualized multi-session hierarchical Bayesian parcellation (Kong et al.) to define DN and dATN boundaries within each subject is the right methodological choice for this question. Network boundaries in posterior cortex, where DN and dATN interdigitate most closely, vary considerably across individuals, and group-average approaches would introduce exactly the kind of misassignment that would most confound the result.

      The matched-vs-random pRF analysis is well-controlled. The authors demonstrate that cortical distance between matched and randomly-matched dATN pRFs does not differ, effectively ruling out spatial proximity on the cortical surface as a confound. tSNR controls further show that signal quality differences do not drive the effect.

      The event-triggered analysis (Figure 3) is creative and adds genuine value. Showing that retinotopically-specific coupling persists during DN-initiated activity transients, not only dATN-initiated ones, is the key piece of evidence for the claim that the code is intrinsic to the DN rather than passively inherited through bottom-up visual drive.

      The result is observed consistently across all individual participants, which provides strong evidence for the robustness of the qualitative pattern despite the small sample size inherent to densely-sampled designs.

      Weaknesses

      (1) The nature of negative pRFs requires more scrutiny

      The entire interpretive framework depends on treating negative pRFs in the DN as genuine position-selective neural responses (suppression). However, negative BOLD signals are well known to arise from non-neural sources, specifically, vascular stealing (where activation in nearby tissue diverts blood from adjacent voxels) and macrovascular draining vein effects that produce spatially displaced signal inversions. These concerns are amplified at 7T, where T2*-weighted GEEPI carries substantial macrovascular weighting. The DN and dATN interdigitate extensively in the posterior cortex, often within millimeters. A negative pRF in a DN voxel adjacent to a positive dATN voxel could, in principle, reflect the hemodynamic shadow of its neighbor rather than an independent neural response.

      The spatial dispersion control (matched vs. random pRFs have similar cortical distribution) is valuable but addresses long-range confounds, not local hemodynamic crosstalk. The reliability of sign and center position across runs is reassuring but does not exclude a vascular origin, as vascular architecture is itself stable across sessions. I would encourage the authors to test whether the matched-vs-random effect survives exclusion of voxels near large pial vessels (identifiable from T2* contrast or the venograms available in the NSD). These analyses would not be dispositive, but they would meaningfully strengthen the neural interpretation.

      The reviewer raises an important concern about the interpretation of negative pRFs in the DN, namely that spatially specific negative BOLD responses could, in principle, reflect local vascular effects rather than genuine position-selective suppression. The reviewer suggests excluding voxels near large vessels to address this issue.

      Based on the reviewer’s suggestion, we repeated the pRF matching analysis excluding any voxels within 3mm of a major vein, as identified using the time-of-flight (TOF) MR venography included in the NSD. This analysis therefore tests whether the retinotopically specific DN–dATN coupling persists after removing voxels most likely to be affected by vascular signal.

      Excluding these voxels did not impact our results: we found preferential coupling according to response valence and center position, with stronger correlation between matched +DN and +dATN voxels (t(6) = 6.054, p < 0.001), and a more pronounced negative correlation between matched -DN and -dATN voxels (t(6) = -5.0448, p < 0.01). We have added these results to the supplemental figures (Fig. S7), and also added to the text (Pg. 8). Together with the run-wise reliability of pRF sign and position, and the persistence of the matched-versus-random effect after vessel exclusion, this analysis supports the interpretation that negative DN pRFs reflect structured, spatially specific responses rather than a vascular artifact.

      “Finally, to rule out any possible influences from vascular stealing (i.e. the shunting of blood into active tissue from nearby regions), we repeated the matching analysis after excluding any voxels within a 3mm radius of a major vessel (Fig. S7; see Methods). Both matching effects remained after excluding vascularly susceptible voxels (+DN x dATN: t(6) = 6.054; p < 0.001; -DN x dATN: t(6) = -5.0448; p < 0.01).”

      (2) Amount of retinotopic mapping data and choice of pRF pipeline

      The NSD includes 6 runs of retinotopic mapping (~5 minutes each; 3 baraperture, 3 wedge/ring). The authors use only the 3 bar-aperture runs (~15 minutes total per subject) and fit their own pRFs using AFNI's 3dNLfim procedure, rather than using the pRF estimates provided as part of the NSD release (which were fitted using the analyzePRF toolbox with all 6 runs).

      Fifteen minutes of bar data is quite limited for reliable voxel-wise pRF estimation, especially in regions far from the early visual cortex, where signal-to-noise is inherently lower. Standard recommendations for robust pRF mapping in higherorder regions generally suggest substantially more data. The variance-explained threshold is close to the noise floor by design, meaning that a non-trivial number of the "retinotopic" DN voxels may be poorly estimated. Given that the core analyses depend on both the sign and the center position of these pRFs, the limited data is a significant concern.

      The authors do not explain why they chose to re-fit pRFs rather than use the NSD-provided estimates. If the motivation was methodological (e.g., the NSD pRF pipeline does not readily yield signed amplitude, or the bar-only fits were judged more appropriate for detecting negative responses), this should be made explicit. If the NSD-provided pRFs can reproduce the key findings, this would substantially increase confidence in the results. If they cannot, that divergence itself would be important to understand. I would ask the authors to address this choice and, if feasible, to report whether the core results replicate using the NSDprovided pRF estimates and/or whether using all 6 runs of retinotopy data changes the findings.

      The reviewer raises two related concerns: first, that the amount of retinotopic mapping data available in the NSD may be limited for estimating voxel-wise pRFs in higher-order cortical regions; and second, that we re-fit the pRF model using AFNI rather than relying on the pRF estimates provided with the NSD release. We appreciate the opportunity to clarify both points. We agree with the reviewer that more travelling bar data would be preferable and would likely yield more robust model fits, particularly in higher-order regions with lower SNR. This is a limitation of our paper that we now acknowledge in the discussion section. However, we do not think that more data would fundamentally change the pattern of our results for the following reasons.

      First, we implemented a novel data-driven approach to derive a threshold for thresholding significant pRF fits (a noise floor). Importantly, our noise floor estimation yields a conservative threshold (R<sup>2</sup> > 0.14), which is greater than both our previous work characterizing cortical pRFs (Steel et al. 2024: R<sup>2</sup> > 0.08) and other work exploring visual responses in the default network (Klink et al. 2021: R<sup>2</sup> > 0.05; no threshold: Szinte and Knapen 2020; Knapen 2021).

      Second, the key pRF features used in our analyses – response sign and centre position – were reliable across retinotopic mapping runs. This reliability is important because our central matching analysis depends on voxel-wise estimates of both response valence and visual-field position.

      Third, the matching analysis asks whether pRF parameters estimated from the retinotopic mapping task predict functional coupling measured during independent resting-state scans. Noisy or unstable pRF estimates should weaken this relationship, because they would degrade the accuracy of voxel-wise matching. Thus, parameter instability would be expected to obscure retinotopically specific coupling rather than systematically produce the observed matched-versus-random effects.

      To the reviewer’s question about our decision to re-fit the pRF model using AFNI, the reviewer is correct that this was motivated by the requirements of our analysis: we re-fit the pRF estimates using AFNI because it allows for both positive and negative signed amplitudes. The pRF model fits provided with the NSD do not allow bivalent amplitude estimates. We have made this decision clearer in the text, reproduced below (Pg. 4-5; Pg. 14-15).

      “We chose to re-fit the data using a simple Gaussian approach as implemented in AFNI to allow for both positive and negative signed amplitudes.”

      “The limited amount of pRF mapping task data included in the NSD posed a challenge for establishing reliable visual response estimates. Here, we addressed this issue by developing a novel thresholding method to establish robust voxel-wise model fits. Among voxels that passed this empirical threshold, we observed a significant correlation in voxel-wise estimates of centre position and visual response amplitude. In addition, our pRF matching results were based on the relationship between the voxel-wise estimates of centre position and response amplitude with resting-state fMRI – a completely independent measure. Crucially, noisy estimates of pRF parameters would obscure this relationship and make our results less likely. Therefore, despite the relatively limited pRF mapping data available, unstable pRF estimates are unlikely to drive our results.”

      (3) pRF model adequacy for the Default Network

      The isotropic Gaussian pRF model was developed for and validated in early and mid-level visual cortex, where it captures the dominant spatial selectivity of neuronal populations. In DN voxels where the model explains comparatively little variance, it is less clear that the model is capturing the right quantity.

      Specifically, the negative pRFs could conceivably be described by a model with a dominant suppressive surround (e.g., a difference-of-Gaussians model), in which what appears as a "negative pRF" in the standard model is actually the surround component of a center-surround mechanism whose center is poorly resolved. This distinction matters: a genuine inverted code (negative center response) implies a qualitatively different computation than inherited surround suppression from nearby visual cortex.

      The authors should consider discussing why the standard model is sufficient for the questions asked, or ideally, testing whether the sign distinction survives under alternative pRF model specifications.

      We appreciate the reviewer’s comment about the limitations of a single gaussian pRF model. We chose the single gaussian model as a direct extension of prior work from our lab and others (Steel et al., 2024, Klink et al., 2022, Szinte and Knapen, 2021). We agree that a negative response in this model could, in principle, reflect a more complex spatial profile, such as a dominant suppressive surround. However, adjudicating among alternative pRF models would require more retinotopic mapping data than are available in the NSD, particularly for higher-order cortex. Thus, we feel that it is outside the scope of the current work. We now address this limitation in our discussion (Pg. 15).

      “Relatedly, here we used a single gaussian model, consistent with prior work on negative visual responses in memory systems (31, 33, 34). However, other models of visual receptive fields might offer further insight into the DN’s visual responsiveness, such as double gaussian models of surround suppression (65) or compressive summation (66). Future studies might directly compare different visual models to further refine the computations underpinning visual responses in the DN.”

      (4) Interpreting resting-state transients as top-down vs. bottom-up The event-triggered analysis labels high-amplitude DN pRF activations as "topdown events" and dATN activations as "bottom-up events." This is a reasonable inference given experience-sampling studies showing that rest involves alternation between internal and external attention, but it remains an inference. Without concurrent experience sampling, eye-tracking, or physiological monitoring, we cannot establish that a spontaneous DN transient reflects memory retrieval or internally-directed thought rather than a global arousal fluctuation. Similarly, dATN transients during rest could reflect covert shifts of spatial attention to remembered or imagined locations rather than bottom-up processing per se. I would ask the authors to soften this framing or to discuss what additional data would be needed to validate the top-down/bottom-up attribution.

      The reviewer raises an important concern about the strong interpretation of elevated BOLD activity detected in the DN and dATN as top-down and bottom-up events. We agree that the limitations of fMRI in our current data prevent these strong claims about the origin of these signals. We have therefore softened this framing throughout the manuscript, and we now refer to these events as DN-driven and dATN-driven. We think that this more directly describes the analysis: events were defined by transient high-amplitude activity in DN or dATN pRFs, respectively.  

      (5) The "retinotopic code" vs. "visual field bias" distinction The paper uses the language of a "retinotopic code" throughout and correctly distinguishes this from a "retinotopic map," noting that DN voxels do not form a continuous topographic representation on the cortical surface. This distinction deserves greater emphasis. In vision science, retinotopic maps carry computational significance through their topographic continuity and relationship to cortical wiring. A distributed collection of voxels with coarse visual field preferences but no cortical topography is a fundamentally different organizational feature. Recent reviews have drawn an explicit distinction between retinotopic maps and visual field biases (Groen, Dekker, Knapen & Silson, TiCS 2022), and the present findings may be more accurately characterized as the latter. Perhaps the authors think that the distinction is merely a signal-to-noise distinction, in which case I would invite them to clearly speak to this interpretation. In any case, this is not a criticism of the findings themselves, but clarity on this point would prevent conflation of two different organizational principles and would help position the work for both the vision and network neuroscience communities.

      The reviewer raises a valuable point about the distinction between a retinotopic code, a retinotopic map, and a visual field bias, and we are happy to add discussion of this topic to our manuscript.

      Our results show that the DN does not exhibit a continuous retinotopic map in the sense used in early visual cortex. Rather, our results suggest a distributed voxel-level code for visual-field position: individual DN voxels show reliable spatial preferences, and these preferences predict retinotopically specific functional coupling with dATN voxels. This voxel-level organization is analogous to other distributed spatial codes, such as head-direction coding in retrosplenial cortex, where spatial variables are represented by population activity without requiring a topographic map on the cortical surface. This differs from a coarse visual-field bias, including preferential responses to the contralateral visual field, although we do also observe such biases. We have added text unpacking this important distinction to the Discussion (Pg. 15-16):

      “Prior work has emphasized the visual response bias in regions where voxel-wise retinotopic responses lack a map-like organization(35); overall, the DN does exhibit this kind of bias. However, our results show that the voxel-scale activity underpinning this bias reflects the latent connectivity of those voxels. Thus, we adopt the term “retinotopic coding”, because this voxel-scale coding scheme exists without a map-like organization on the cortical surface. For example, rodent and bat head direction cells are not laid out in a literal ring, but the population code of these neurons forms a ring manifold(68, 69).”

      Reviewer #2 (Public review):

      Summary:

      Using a public dataset of retinotopic mapping and resting-state data, the authors find that the default mode network has voxels that respond (positively or negatively) to visual stimulation at specific retinotopic positions, and that restingstate activity in these voxels is correlated with activity in more traditional sensory voxels with the same visual-location preference. The retinotopic specificity is bidirectional, such that high activity in default mode voxels drives activity only in voxels with matching receptive fields in sensory cortex, and vice versa. These findings are at odds with traditional views of the default mode network as having abstract (non-retinotopic) representations and competing (rather than cooperating) with external sensory representations.

      Strengths:

      This study continues an intriguing line of research about how default mode regions interact with the sensory cortex. Demonstrating that there are structured interactions between these regions at rest, and that these interactions are in fact organized according to retinotopic location (as opposed to traditional views of representational format in the default mode network), provides a new framework for thinking about large-scale internal and external brain networks. The authors make use of a well-powered public dataset that allows for precise estimates of pRFs and individual-specific resting-state networks, and develop a number of interesting analyses that characterize the relationships between DN and dATN voxels. The findings are exciting and could have a major impact on future studies in cognitive neuroimaging.

      The authors mention that these findings could shed light on internal/external interactions such as "anticipatory saccades or memory-guided attention," which is true, though I would argue that constructing DN representations of external stimuli is in fact even more fundamental than these specific cases (e.g., see Barnett and Bellana, 2025, "Situation models and the default mode network"). The "highways" identified in this study could play a vital role in real-world perceptual processes that are constantly translating external input into internal mental models.

      Weaknesses:

      (1) The criterion used for defining voxels as retinotopic seems very liberal. The authors show that only 5% of voxels have R^2>0.14 in a null analysis, and therefore define voxels with R^2>0.14 as retinotopic. Although all the networks in 1C show voxel distributions that differ from the null, the number of false positives above R^2>0.14 seems problematic, especially for the DN positive pRFs (red distribution) and to a lesser extent the DN negative pRFs (blue distribution). From visual inspection of the plot, the false discovery rate (fraction of voxels labeled as retinotopic that are false positives) looks like it would be greater than 50% for the DN-positive pRFs. The authors do show that the positive pRF voxels have abovechance consistency across runs, again providing evidence that there are true positive voxels in this set, but perhaps a stricter criterion (such as having consistent negative fits across runs) would provide more targeted identification of the DN voxels with true retinotopic sensitivity.

      We thank the reviewer for giving us the opportunity to discuss this important decision. We agree with the reviewer that a stricter R<sup>2</sup> criterion could result in more targeted pRF identification. Motivated by the reviewer’s suggestion, we repeated the cross-region pRF matching analysis across multiple R2 thresholds.

      The retinotopic matching effects were not dependent on the original threshold. In fact, we found that the pRF matching effects are enhanced as the R<sup>2</sup> value increases (Fig. S5). This pattern suggests that any false-positive voxels admitted near the original threshold would dilute, rather than drive, the observed matched-versus-random effects. We have added text to the results highlighting this finding (Pg. 7):

      “In contrast, DN voxels that responded positively to visual stimulation (DN positive pRFs, +pRFs) had a positive correlation with the dATN (mean correlation = 0.284±0.152, t(6) = 4.96, p = 0.0025), while DN voxels with systematic negative responses to visual stimulation (DN negative pRFs, -pRFs) were anti-correlated with the dATN (mean correlation = -0.21±0.149, t(6) = -3.75, p = 0.0094). This relationship was further strengthened by adopting more conservative R<sup>2</sup> thresholds up to 0.30 despite the overall number of included voxels decreasing, suggesting that this effect is not driven by false-positive voxels at the edge of our threshold criteria (Fig. S5).”

      (2) The claim that "opponency at rest between the DN and dATN appears to be driven by the subset of DN voxels with negative retinotopic tuning" is not well supported. The fraction of DN voxels with negative pRFs is small: 9.42% of DN voxels have pRFs, and 58.77% are negative, so about 6% of DN voxels have negative pRFs. The fact that any DN voxels have negative pRFs is notable, but the authors do not provide evidence that these 6% are driving the overall behavior of the DN. They do show (e.g., in Figure 2B) that negative and positive pRFs have opposing influences, but the overall correlation with dATN does not look similar to the negative pRF connectivity. I'm also unsure whether "opponency" is a reasonable description for two networks that are "independent (i.e., not correlated)" in this analysis.

      The reviewer raises an important point about whether negative DN pRFs should be described as driving the overall DN–dATN relationship. We agree that this language was too strong. Negative pRFs constitute a small subset of DN voxels, and our analyses show that this subset has a distinct pattern of functional coupling with the dATN, not that it explains the global relationship between the DN and dATN as a whole.

      We have therefore revised the manuscript to avoid implying that negative DN pRFs drive overall DN–dATN opponency. Instead, we now frame these voxels as an important retinotopically tuned subpopulation nested within broader network dynamics. Specifically, our results show that visually responsive DN voxels are not homogeneous: positive and negative DN pRFs show opposing patterns of coupling with dATN pRFs, and these interactions are strengthened when voxels share visual-field preferences. This suggests that a small but structured subset of DN voxels may provide a route for retinotopically specific communication between internally and externally oriented networks, without implying that this subset determines the mean activity pattern of the entire DN:

      “Spontaneous DN and dATN activity during rest is uncorrelated at the network level. However, voxel-scale functional coupling across networks is shaped by the latent visual field preferences of individual voxels in each network, as measured during independent retinotopic mapping.” Abstract (Pg. 2)

      “This result shows that voxel-level visual response profiles shape DN-dATN coupling during spontaneous resting-state activity. Specifically, the DN and dATN activation is independent during rest. However, at the voxel-level, specific sub-groups of DN voxels have distinct coupling patterns with the dATN that depends on the valence of voxels’ visual responses. DN and dATN voxels with positive visual responses show a positive relationship during rest, and a notable subset of DN voxels with negative visual responses display the canonical opponency with dATN voxels. This suggests that retinotopic coding may be a mechanism that enables visual information to be exchanged between these large-scale brain systems. Specifically, opponency at rest between the DN and dATN appears to be driven by the subset of DN voxels with negative retinotopic tuning.” Results (Pg. 7)

      These findings offer a multi-scale account of neural communication, in which interactions among sub-populations of voxels with shared tuning preferences are nested within macro-scale network dynamics. Nesting multiple neural codes might enable ongoing computations within a larger brain system (e.g., attending to internal mental states within the DN during memory recall), while simultaneously allowing for the sharing of fine-grained representations across brain systems (34). Discussion (Pg. 14)

      (3) The event-triggered analysis is effective at testing the bidirectional relationship between DN and dATN, with high activity in either network triggering a response in the other network. However, it would be helpful to show more validation that these "events" are meaningful windows of time to study. First, is 13 TRs a typical length of time that activity is elevated during one of these events? Second, the top-down and bottom-up terminology is perhaps too loaded and not well-justified; if the negative pRFs in the DN reflect a meaningful coding system, then couldn't low (rather than high) activity indicate a top-down event?

      We thank the reviewer for these helpful suggestions. To the best of our knowledge, there is not currently a widely agreed-upon time window for performing event-based fMRI analyses. We chose a 13 TR time window to balance between sufficiently capturing BOLD signal related to the chosen event while also minimizing influence from other signal fluctuations, based on the procedure adopted in Gordon et al. (Nature, 2023) and Mitra et al. (J. Neuro Phys, 2014), which considered temporal relationships among brain regions over comparable timescales. In our analysis, this window considered 6 TRs (9.6s) on either side of the detected event, which we felt comfortably captures the peak BOLD signal that would result from an impulse at the event time, and responses that may reflect upstream activity leading into it.

      The reviewer has raised an additional comment about the terms “top-down” and “bottom-up.” These concerns were shared by Reviewer 1. Based on these comments, we have adopted the terms “DN-driven” and “dATN-driven”, which we think aligns more closely with our analysis approach.

      (4) The framing of this paper relative to the authors' past work, such as Steel et al. 2024 ("A retinotopic code structures the interaction between perception and memory systems"), could be improved. The existence of negative pRFs in the DN and a functional relationship between these pRFs and the sensory pRFs have already been described in prior work. My understanding of the primary novelty here is that this paper examines resting-state data, showing that there are widespread spontaneous interactions between broad internal and external networks, but this distinction is not made explicit in the Introduction.

      We appreciate the opportunity to clarify the novel aspects of our paper. The reviewer correctly identifies the extent of prior work, which identified -pRFs in regions of the canonical default network (Szinte and Knapen, 2021; Klink et al., 2022) and characterized the local interactions between adjacent perceptual and mnemonic regions (Steel et al., 2024). Our current work builds upon these findings in two key ways.

      First, we explore the effect across individually-defined whole brain networks. While the DN and dATN are often adjacent, these networks are spatially discontinuous and are comprised of distinct sub-regions (e.g. in prefrontal cortex). Whether retinotopic patterning of activity would persist in distributed networks could not have been extrapolated from our prior work. We think that finding will be of broad interest to the community studying perception and memory systems, because it offers a mechanistic account of how information is read in/out of memory.

      Second, here we considered whether spontaneous activity across networks would be structured by a retinotopic code. Our previous work characterized activity during tasks that depended on visual information: either scene perception or mental imagery. While the prior work was an important first step, it left open the possibility that retinotopic coding may only be relevant in visual tasks. By demonstrating that the retinotopic coding structures voxel-specific coactivation during rest, which entails no overt visual demands, we provide evidence that retinotopic features are a general, mode-agnostic code between regions.

      (5) The definition of the default mode (DN) in this study aligns with past research, but the definition of the dorsal attention network (dATN) seems at odds with standard terminology. For example, the authors cite Fox et al. 2006, which depicts the dATN as including regions such as IPS, FEF, SMA, and MT+. Here, however, the "dATN" seems to be primarily lateral and ventral visual cortex (e.g., Figure S5). The exact location of these sensory pRFs is not critical to the authors' claims, but this labeling seems incorrect, and the motivation for defining/selecting the sensory network in this way is not described.

      We thank the reviewer for this insightful comment and their careful consideration of our network definition.

      Our method for network identification, and the topography of the resulting networks, are broadly consistent with more recent conceptualizations of the DN and dATN (e.g. Du et al. 2024, Gordon et al. 2017, Braga and Buckner 2017). Relatedly, because we defined brain networks based on the unique connectivity patterns of each individual participant, we expect them to differ from previous group-level network descriptions. The increased resolution of the 7T data in the NSD may also result in greater departure from prior definitions compared to previous work done at 3T.

      Further study into dATN differences between group-level 3T, individualized 3T, and individualized 7T networks could be a valuable future direction, but this is outside the scope of this work.

      Reviewer #3 (Public review):

      Summary:

      This paper addresses an important question (the relationship between DN and dATN, and the role of retinotopic coding) and uses a set of novel analyses.

      Strengths:

      Important question, novel analytical approaches (pRF-informed functional connectivity analysis).

      Weaknesses:

      Some of the key claims are not fully supported by the data presented. There is also a concern about over-interpretation of the results. Key issues:

      (1)  The authors claim that retinotopic coding scaffolds the interaction between DMN and dATN. However, retinotopically tuned voxels account for a mere 9% of DMN voxels. So this appears to be a major overstatement. For instance, the statement that "these findings would position retinotopy as a unifying framework for brain-wide information processing" is not justified given the presented data.

      We appreciate the reviewer’s concern about the framing of our conclusion, which was shared by reviewers 1 and 2. In response to these comments, we have revised our paper to more accurately reflect the observed data. Specifically, we focus on the specific sub-populations of voxels within the DN and dATN that show retinotopic responses, and we have removed references to explaining the overall pattern of activity across networks.

      (2) Given that positive pRF voxels in DMN positively correlate with dATN voxels and negative pRF voxels in DMN negatively correlate with dATN voxels, there is a concern that these results could be contributed to by imprecise brain network parcellations. E.g., could some of the positive pRF voxels in DMN be erroneously assigned to DMN and actually belong to one of the other task-positive networks? There is insufficient validation of network parcellation to put this worry to rest, especially since it depends on ICA, which has a degree of arbitrariness built in.

      We thank the reviewer for the opportunity to clarify our method for network definition.

      Precision functional mapping is a growing field with many methods for defining personalized functional networks for each individual. Because the NSD resting-state data is relatively high resolution, we chose an approach designed to improve the stability of voxel-wise network assignment: Multi-Session Hierarchical Bayesian Modeling approach (Kong et al. 2019; Du et al. 2024). This approach enhances stability of network assignment by including a group-based prior and accounting for both within- and across-subject variability. This approach is more stable than ICA, and, because this approach leverages a prior, there is less concern about arbitrary or idiosyncratic network definitions.

      However, it is still common for network assignments to have lower confidence around the borders between networks. Yet, we also do not think border misassignment is likely to explain the present results for two reasons: first, while DN and dATN nodes are sometimes adjacent, there are many regions where they are spatially distant, such as the IPS for dATN and the lateral temporal lobe for DN. Second, the DN pRFs do not appear to cluster selectively along DN–dATN borders, suggesting that they are not simply misassigned dATN voxels.(Fig. S3) Therefore, we think voxels on the edge of these networks are unlikely to drive the effects observed here (see Fig. S3).

      (3) The claim that retinotopic coding is intrinsic to the DN network is not supported by rigorous analysis and results. The analysis here has many arbitrary factors, including: the threshold of the 99th percentile of resting-state distribution; the designation of DN as "top-down" and dATN as "bottom-up"; the definition of "anti-matched" voxels instead of using randomly selected voxels; and the statistics being paired between matched and anti-matched voxels instead of using comparisons to baseline. Overall, I do not think that the result supports the conclusion that retinotopic coding in DN is intrinsic instead of being bottomup-driven, given the very high threshold (99%) used and the fact that many other networks could also send bottom-up input to DN. Furthermore, the idea that bottom-up inputs only occur when the dATN (or any other RSN)'s spontaneous BOLD activity is above a certain threshold is a huge and unvalidated assumption.

      The reviewer raises several interesting concerns about decisions in our event-detection analysis. Here, we clarify the rationale for several analytic choices:

      (1) The 99th-percentile threshold was chosen to identify sparse, high-amplitude events while minimizing contamination from smaller ongoing fluctuations.

      (2) The other reviewers also noted a concern with the top-down/bottom-up terminology. We have revised these terms to DN-driven and dATN-driven, which we think reflect our approach more accurately.

      (3) We used anti-matched rather than randomly selected voxels because the full event-by-voxel randomization procedure was computationally intractable at the network level.

      (4) We did not understand the reviewer’s contention about activation baseline, but we would welcome clarification.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor points

      (1) The reliability analysis (Figure 1D) notes that dATN negative pRF amplitude was not reliable above chance in 2 of 7 participants. This could be discussed more prominently as it suggests that negative pRFs may not be stable features in all networks, which tempers the generality of the sign distinction as a fundamental organizational property.

      We thank the reviewer for raising this point of clarification. It is true that 2/7 participants did not show reliable negative pRFs in the dATN. However, the majority of participants show stable negative pRFs, and even in these 2 participants, the negative result does not indicate that negative pRFs would not be stable in those individuals with additional data. 

      Based on the reviewer’s comment, we have added emphasis to this point, but we do not feel that this warrants greater discussion in the paper. 

      Both positive and negative pRF amplitude was reliable in the DN for all subjects. In the dATN, positive amplitude pRFs were reliable in all participants, and negative amplitude pRFs (which constituted a small proportion of the overall pRFs in this network) were reliable in 5/7 participants. For the remainder of the paper, we only consider positive pRFs in the dATN. Importantly, pRF center position was highly reproducible across runs of pRF data in the dATN and DN in all subjects (Fig. 1D).  Pg. 5

      (2) The paper would benefit from situating the findings more explicitly within the cortical gradient framework (Margulies et al., 2016), which predicts that DN regions have maximally abstract, transmodal codes. The present findings complicate this view productively and deserve to be "situated" within that ongoing debate.

      We agree that the gradient framework is interesting, and we have added discussion of Margulies to our paper. (Pg. 16)

      Relatedly, the DN is considered a transmodal hub for cortical processing, where disparate sensory and motor processes converge (59, 75) The DN’s position at the cortical apex implies connections with and influence over unimodal cortical areas. However, the mechanism for liking unimodal and transmodal networks had been unknown. Prior work posited that sensory coding in transmodal areas might serve this function (31, 35), and our data provide direct empirical support for this account: specific visually-responsive voxels provide an input/output interface linking perceptual and memory systems. This complements work delineating specific affective and effective subregions within the DN that link the DN to other brain areas (76). Thus, while the DN may be “distant from input” (28), these data suggest that it is not disengaged from sensory processing.

      (3) It would be informative to know whether the *proportion* of negative vs. positive pRFs differs between DN-A and DN-B, given their distinct functional roles.

      Despite the functional specialization of DN-A and DN-B, and the slightly higher mean proportion of negative pRFs in DN-A (61% vs 56%), we found no statistically significant difference in the proportion of negative pRFs across the two networks (t(6) = 0.888, p = 0.409).

      (4) Low N is inherent to the densely-sampled NSD design, and the within-subject consistency is a strength. Nevertheless, with 6 degrees of freedom, the precision of specific quantitative estimates (e.g., that 58.77% of DN pRFs are negative) is uncertain, and the authors should be cautious about the generalizability of these point estimates.

      The reviewer raises a concern about the inclusion of specific levels of decimal place in our statistical reporting. We do not think that this is a major issue with the paper, but we are willing to change if the reviewer feels strongly.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1C could use an explicit legend (I believe it is following the color convention from the bar plots in 1F?). Also, for consistency, it would be helpful to make all the colormaps in 1F correspond to the bars (i.e., change the dATN colormap to go white->green).

      We thank the reviewer for this suggestion, and we have added explicit labels to Fig. 1C

      (2) Providing a scatter plot, in which each dot is a voxel and the x and y axes are the pRF amplitude estimates in different runs, could help provide evidence that there are voxels with pRFs that have consistently negative amplitudes across runs. This would also go beyond the binary consistency analysis in Figure 1D to show that the magnitudes of the amplitude estimates are also consistent.

      We thank the reviewer for this suggestion. We feel that the binary consistency conveys sufficient information. Because the analysis is done using pairwise correlation, how the scatter plot would reflect the three-way consistency is not clear. 

      (3) For understanding how the overall correlation between DN and dATN could be driven by voxel populations with opposing effects (e.g., Figure 2B), it would be useful to show how the +pRF and -pRF voxels compare to other voxels within the DN. For example, are these the voxels with the strongest negative and positive correlations with dATN, or are there many other DN voxels (among the 90% that do not have pRFs) that also have similarly-strong dATN correlations?

      The reviewer offers a very interesting suggestion. Based on the reviewer’s suggestion, we have refocused our paper on the particular subpopulations of +/- pRFs in the DN, rather than on an explanation for the overall pattern of correlation between the DN and dATN. Because our revised framing focuses on the properties of these retinotopically defined voxel populations, rather than on explaining whole-network DN–dATN coupling, we have not added this additional analysis. We have revised the relevant text to avoid implying that these pRF subpopulations drive the overall network-level relationship.

      (4) Initially, the baseline comparison pRFs for the matched pRFs are labeled "random" pRFs, which seems misleading; these are closer to "mismatched"/"anti-matched" pRFs since they are selected from the 1/3 that are farthest away. Then the comparison switched to using the anti-matched pRFs that are the 10 very farthest away, though I didn't understand the rationale that "the large number of pRFs made the random matching procedure impractical" - in what way is the number of pRFs larger in this analysis? Having a more consistent baseline (e.g., just using the 10 anti-matched pRFs the whole time) would be easier to interpret.

      We thank the reviewer for this suggestion. We have compared the results between the randomly-sampled bottom ⅓ matched versus the 10 worst matched, and the pattern of results is identical (the effect is strongest in the 10 worst matched). Therefore, we include the bottom ⅓ matched in the main text as a more conservative test of this effect. We are happy to include this as a supplemental figure if the reviewer feels it is essential. 

      (5) In the past, I have only seen the terminology "bootstrapped" to refer to sampling with replacement from the data sample, producing samples/statistics that are centered on the observed data. Here (lines 704-708), the sampling is coming from the null distribution of randomly-chosen voxels, and therefore the term "bootstrapped" would not apply (and could just be replaced with "null").

      We have revised this terminology in the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract and Discussion should be significantly toned down. E.g., the claim that "These findings challenge the prevailing view of global DN-dATN antagonism" is not really supported by the data provided. The claim that "retinotopic coding underpins the dynamic coordination of perception and thought" is also unsupported by the presented data.

      We have revised the manuscript in light of this comment.

      (2) Line 233-235: The null statistical result cannot support the claim reached here. Correlation analysis or Bayesian statistics should be used.

      We have revised the manuscript in light of this comment.

      (3) Line 250-254: Comparison to baseline should be used, in addition to comparing matched and random voxels.

      We agree that baseline comparisons can be useful in event-triggered analyses. However, for the pRF-matching analysis discussed here, the critical question is whether shared visual-field preference influences resting-state functional coupling between DN and dATN voxels. For this question, we believe that the appropriate baseline is the coupling observed for pRFs that do not share visual-field preferences. We therefore compare retinotopically matched pRFs to randomly matched pRFs drawn from the same networks. 

      (4) Line 271: "not" is missing.

      We have revised the manuscript in light of this comment.

    1. eLife Assessment

      This important study combines chromatin accessibility and genomic DNA sequence conservation data from low-coverage genome sequencing of related species (without assembly), for the in-silico identification of cis-regulatory elements in large genomes. The approach and results are compelling and well supported by the experimental validations. The work will be of interest to researchers working in the field of gene regulation and evolution, particularly because the methodology proposed can be applied to a large variety of experimental organisms.

    2. Reviewer #1 (Public review):

      Summary:

      Forbes et al. developed an integrated approach to identify cis-regulatory elements (CREs) in the large (3.6 Gbp) genome of the crustacean Parhyale hawaiensis, addressing the challenge of pinpointing these regions among large regions of non-coding sequences. They combined ATAC-seq chromatin accessibility profiling (both bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis). Without assembling congener genomes, they mapped reads with low stringency to the P. hawaiensis reference, identifying about 55k conserved islands that overlap ATAC peaks more than expected by chance. This dual filter was used to select CRE candidates for transgenic reporter validation, yielding 6 functional elements (out of 11 tested) driving ubiquitous, neuronal, or muscle-specific expression, a major advance for non-model systems with large genomes.

      Strengths:

      Forbes et al. generated high-quality ATAC data across multiple scales. Using bulk ATAC-seq (from whole embryos, developing and adult legs) they identified tens of thousands of open chromatin peaks across the assembled P. hawaiensis large genome. Moreover, using single-nucleus ATAC-seq from adult legs, they could resolve differentially accessible chromatin profiles across more than 15 cell types previously identified by scRNA-seq, enabling cell-type-specific candidate selection.

      Furthermore, their innovative low-coverage comparative genomics method mapped 0.46-6.4% of congener reads to P. hawaiensis without genome assembly, revealing hundreds of thousands of conserved non-coding islands, including about 55k showing conservation in all four species, far exceeding random expectation.

      Using the developed approach, the authors could validate 6 (out of 11 candidates) reporter constructs, driving robust ubiquitous and tissue-specific expression, succeeding where prior promoter-only screening failed and providing immediately useful genetic tools for the Parhyale community.

      Weaknesses:

      The primary limitation is that functional CRE testing was performed only in P. hawaiensis. While the conservation maps provide a valuable resource for comparative analyses, functional validation in congener species was not performed, so the extent to which the identified CREs or the prioritization strategy can be functionally generalized across related species remains to be established.

      The approach did not successfully identify developmental CREs among the candidates tested. None of the candidates selected using the combined ATAC-seq and conservation filtering drove reporter expression matching the expected endogenous patterns. The authors appropriately discuss possible technical and biological explanations.

      Overall Assessment:

      Forbes et al. fully succeed with their integrated approach to (1) generate an ATAC-seq atlas plus functional CRE discovery and (2) innovative low-coverage sequencing for conservation mapping in the large 3.6 Gbp genome of Parhyale hawaiensis. Their combination of ATAC-seq chromatin accessibility profiling (bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis), without congener genome assembly, drastically shrank the CRE search space. Using this approach, the authors could validate six out of 11 candidate transgenic reporters (ubiquitous, neuronal, and muscle-specific) where prior promoter-only screening failed.

      The low-coverage mapping innovation cuts cost and labour while snATAC-seq provides cell-type resolution, making these resources valuable for building new genetic and imaging tools in Parhyale.

      This compelling method also has the potential to enable labs with limited resources to identify and characterize regulatory elements in more non-model organisms, advancing our understanding of their evolution while establishing a scalable pipeline for large-genome systems.

      Comments on revised version.

      The authors have adequately addressed all my previous comments. I have no further specific suggestions or requests.

    3. Reviewer #2 (Public review):

      The manuscript by Forbes, Skafida, Karapidaki et al. concerns the in-silico identification of cis-regulatory elements (CREs) in large genomes using chromatin accessibility (ATAC-seq) and sequence conservation (genomic DNA sequencing) data. They exemplify this method by applying it to identify novel CREs in Parhyale hawaiensis, which they validated using reporter constructs.

      The results are convincing and are well supported by the data and validations. Identified CREs are valuable for researchers interested in the regulation of the expression of genes they control.

      The methodology on the whole is also valid, as suggested by the results and previous publications on various taxa. Sequence conservation, as stated by the authors, was long used as a method to identify regions of non-coding DNA with functional and evolutionary constraints. The same applies to ATAC-seq data, which has also been used as a proxy for functional regions in different animals such as sea urchins and amphioxus. The methodology proposed is likely to be successfully used by researchers working on a variety of experimental organisms.

      The authors do not use existing genome assemblies and use short-read sequencing to identify conserved regions, and while it is not conceptually novel, such an approach is becoming more and more viable and useful considering the recent advances in next generation sequencing technology and the decrease in price of short-read sequencing.

      The authors have addressed and discussed the limitations and weaknesses of the approach as well as explicitly indicated the advantages.

      All in all, the authors provide a valid method to strengthen CRE identification via sequence conservation without the need of multiple complete close species genome assemblies, making it a compelling option for non-model organism research.

    4. Reviewer #3 (Public review):

      Summary:

      Forbes et al. present a new approach for identifying cis-regulatory elements in large genomes. Using Parhyale hawaiensis, a crustacean with a large genome (~3.6 Gb, comparable in size to the human genome), the authors show that current methods for identifying cis-regulatory elements, effective in smaller genomes, are markedly inefficient in organisms with large genomes. To address this limitation, they combine bulk ATAC-seq and single-cell (sc) ATAC-seq to identify chromatin regions that are either ubiquitously accessible or specifically accessible in particular cell types. They further integrate comparative genomics across multiple Parhyale species (P. hawaiensis, P. aquilina, and P. darvishi), selected at appropriate phylogenetic distances (20-95 million years divergence), to pinpoint conserved open chromatin regions likely under functional constraint.

      Using this strategy, the authors predict a set of ubiquitous and cell-type-specific cis-regulatory elements. Importantly, they validate these predictions using rigorous transgenic reporter assays, convincingly demonstrating that their approach can successfully identify functional regulatory elements where previous methods had failed.

      Strengths:

      The approach introduced by Forbes et al. is conceptually straightforward, efficient, and readily transferable to other organisms. The validation experiments show not only that a substantial proportion of the predicted elements are functional, but also that the method is capable of identifying both ubiquitous and cell-type-specific regulatory elements. Given that the identification of regulatory regions remains a major bottleneck in understanding the molecular mechanisms underlying processes of development and regeneration, this work has the potential to make a significant impact in developmental and regeneration biology, particularly for studies involving non-model organisms with large genomes.

      An additional strength is the demonstration that only the genome of the focal species requires high-quality sequencing and assembly. In contrast, species used solely for comparative analysis can be sequenced at low coverage without assembly, substantially reducing costs and increasing the accessibility of the approach.

      Weaknesses:

      While the method is effective in identifying regulatory elements that are active ubiquitously or in differentiated cell types, it failed in detecting elements associated with developmentally regulated genes. This may be due to trivial reasons, such as very low level of expression of the selected genes. However, as acknowledged by the authors, it may also indicate inherent challenges in identifying regulatory elements associated with developmentally dynamic gene regulation, compared to those associated with genes expressed in differentiated cell types.

      A second limitation, also acknowledged by the authors, is the absence of chromatin conformation capture data, which would help link distal regulatory elements to their target genes. This limitation may be particularly relevant for developmentally regulated genes, where long-range regulatory interactions may be critical.

      Addressing these limitations will be an important direction for future work. Nonetheless, the approach as presented in this manuscript represents a key contribution that sets the stage for further methodological advances in the identification of cis-regulatory elements in large genomes.

      Comments on revised version.

      I am fully satisfied with the current version of the manuscript.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      Forbes et al. developed an integrated approach to identify cis-regulatory elements (CREs) in the large (3.6 Gbp) genome of the crustacean Parhyale hawaiensis, addressing the challenge of pinpointing these regions among large regions of non-coding sequences. They combined ATAC-seq chromatin accessibility profiling (both bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis). Without assembling congener genomes, they mapped reads with low stringency to the P. hawaiensis reference, identifying about 55k conserved islands that overlap ATAC peaks more than expected by chance. This dual filter was used to select CRE candidates for transgenic reporter validation, yielding 6 functional elements (out of 11 tested) driving ubiquitous, neuronal, or muscle-specific expression, a major advance for non-model systems with large genomes.

      Strengths:

      Forbes et al. generated high-quality ATAC data across multiple scales. Using bulk ATAC-seq (from whole embryos, developing and adult legs), they identified tens of thousands of open chromatin peaks across the assembled P. hawaiensis large genome. Moreover, using single-nucleus ATAC-seq from adult legs, they could resolve differentially accessible chromatin profiles across over 15 cell types previously identified by scRNA-seq, enabling cell-type-specific candidate selection.

      Furthermore, their innovative low-coverage comparative genomics method mapped 0.46-6.4% of congener reads to P. hawaiensis without genome assembly, revealing hundreds of thousands of conserved non-coding islands, including about 55k showing conservation in all four species, far exceeding random expectation.

      Using the developed approach, the authors could validate 6 (out of 11 candidates) reporter constructs, driving robust ubiquitous and tissue-specific expression, succeeding where prior promoter-only screening failed and providing immediately useful genetic tools for the Parhyale community.

      Weaknesses:

      The primary limitation is that functional CRE testing was performed only in P. hawaiensis. While conservation maps are valuable resources, the manuscript lacks functional validation in congener species, limiting claims about broad applicability across related genomes/species.

      The approach also failed to validate developmental CREs. None of the candidates from combined ATAC and conservation filtering drove reporter expression matching endogenous patterns. The authors appropriately hypothesize technical limits (low expression) or biological factors (long-range enhancers, shadow enhancers).

      Overall Assessment:

      Forbes et al. fully succeed with their integrated approach to (1) generate an ATAC-seq atlas plus functional CRE discovery and (2) innovative low-coverage sequencing for conservation mapping in the large 3.6 Gbp genome of Parhyale hawaiensis. Their combination of ATAC-seq chromatin accessibility profiling (bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis), without congener genome assembly, drastically shrank the CRE search space. Using this approach, the authors could validate six out of 11 candidate transgenic reporters (ubiquitous, neuronal, and muscle-specific), where prior promoter-only screening failed.

      The low-coverage mapping innovation cuts cost and labour while snATAC-seq provides cell-type resolution, making these resources valuable for building new genetic and imaging tools in Parhyale.

      This compelling method also has the potential to enable labs with limited resources to identify and characterize regulatory elements in more non-model organisms, advancing our understanding of their evolution while establishing a scalable pipeline for large-genome systems.

      We thank the reviewer for their comments and valuable feedback.

      Reviewer #1 (Recommendations for the authors):

      (1) Standardize terminology and introduce acronyms properly:

      I suggest standardizing technique names throughout the manuscript (e.g., ATAC-seq rather than ATACseq, RNA-seq rather than RNAseq, ChIP-seq rather than ChipSeq). Please also introduce technical terms with their full name at first use, such as 'Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq)' rather than just the acronym.

      Corrected.

      (2) Correct author name:

      On page 15, "Lo brutto" appears to be misspelt. Please review and fix throughout the text and the corresponding reference in the bibliography if this is the case.

      Corrected.

      (3) Revise the title to reflect dual contributions:

      The paper delivers two key advances: (a) ATAC-seq atlas plus functional CRE discovery in P. hawaiensis, and (b) innovative low-coverage sequencing for conservation mapping across congenerics. Current title highlights only the first. Consider modifying the title to better reflect both.

      The title reflects our overall objective, without highlighting one of the two approaches (advances) in particular. We would like to keep this concise title and invite readers to read about the two approaches in the abstract.

      (4) Emphasize the combined power of the approach in the Discussion and Conclusions:

      A significant innovation of this manuscript is integrating low-coverage comparative genomics with ATAC-seq to prioritize functional CREs in a large non-model genome. The abstract highlights this well, but the Discussion and Conclusions could better emphasize the power of this pipeline over ATAC-seq alone. The authors could also add 2-3 sentences quantifying cost savings versus traditional assemblies and reiterating how this complements ATAC-seq for efficient CRE prioritization in non-model species.

      We have added the following text in the Discussion to describe complementary contributions of ATAC-seq and sequence conservation to CRE discovery: "Previous efforts to identify cis-regulatory elements in Parhyale relied on reporter constructs carrying a few kb of sequences upstream of selected target genes, an approach that has worked well in animals and plants with relatively small genomes. As presented earlier, however, this approach was often unsuccessful in Parhyale: from tens of reporters tested, only four robust native drivers had so far been identified (refs). The present work adds 5 new drivers to that collection, including ones with ubiquitous, neuron- and muscle-specific activities. This result comes from combining information on genome-wide chromatin accessibility and evolutionary conservation profiles.

      We cannot at this point distinguish the relative contributions of chromatin profiling and sequence conservation to CRE discovery, because these sources of information were not tested separately. At minimum, we can state that (1) ATAC-seq profiles serve to identify robustly the promoters of candidate genes that are active in a particular cellular context (cell type and stage) and (2) coupling this information with sequence conservation narrows down candidate promoters and distant CREs by a factor of 4 to 10, since only a fraction of ATAC-seq peaks show sequence conservation (Figure 3D). This represents a great improvement in our ability to select candidates to test by transgenesis, the most labour-intensive step in the process."

      Further, we have added this text to explain the advantages and cost savings of our low coverage sequencing strategy: "Our strategy of mapping regions of sequence conservation by direct mapping of short sequence reads across species is much more accessible than conventional strategies that rely on genome assembly. The latter require much higher sequence coverage (> 50x) from multiple libraries, long-read sequencing or other scaffolding methods, and complex bioinformatic pipelines to assemble large genomes. Moreover, these approaches are often compromised by high levels of polymorphism found in natural populations. We estimate that our approach is 5- to 10-fold cheaper than assembly-based methods, even excluding labor costs."

      (5) Improve figure readability:

      The authors could improve figure readability by introducing a schematic representation of the specimens and more references in Figures 4, 5 and Supplementary Figure 5. A schematic representation would help non-experts in Parhyale understand what they are looking at. Some figures might also benefit from improved color contrast (e.g., Figure 3 has very similar orange/red colors; black dots on a dark grey background are hard to distinguish).

      We have added additional labels and explanations in the legends of Figures 4, 5 and Supplementary Figure 5, which we think will make the images more intelligible to the readers. In Figure 3 we modified the colouring in panels B and C to improve contrast.

      (6) Quantify reporter validation efficiencies:

      The authors should add a summary table/plot (e.g. n surviving, n fluorescent) or label the figures (n of specimens showing that pattern/n of specimens that do not show the pattern) with exact numbers to explicitly illustrate the observation. For example, Supplementary Table 3 contains excellent data for the putative developmental CREs tested, but the main text lacks equivalent quantification for successful reporters.

      This information is already provided in Table 3.

      (7) Discuss the rapid evolution of developmental CREs:

      The failure to validate developmental CREs using conserved candidates may also reflect the rapid evolution and turnover of developmental enhancers, which can erode detectable sequence conservation over these phylogenetic distances. As a result, functionally relevant elements may have been excluded during candidate selection. It may be worth discussing this possibility alongside the proposed long-range and shadow enhancers hypothesis.

      We added the phrase “or the rapid evolution of these enhancers leading to low sequence conservation" in the relevant part of the Discussion.

      Reviewer #2 (Public review):

      The manuscript by Forbes, Skafida, Karapidaki et al. concerns the in silico identification of cis-regulatory elements (CREs) in large genomes using chromatin accessibility (ATAC-seq) and sequence conservation (genomic DNA sequencing) data. They exemplify this method by applying it to identify novel CREs in Parhyale hawaiensis, which they validated using reporter constructs.

      The results are convincing and are well supported by the data and validations. Identified CREs are valuable for researchers interested in the regulation of the expression of genes they control.

      The methodology on the whole is also valid, as suggested by the results and previous publications on various taxa. Sequence conservation, as stated by the authors, was long used as a method to identify regions of non-coding DNA with functional and evolutionary constraints. The same applies to ATAC-seq data, which has also been used as a proxy for functional regions in different animals such as sea urchins and amphioxus. The methodology proposed is likely to be successfully used by researchers working on a variety of experimental organisms.

      The authors do not use existing genome assemblies and use short-read sequencing to identify conserved regions, and while it is not conceptually novel, such an approach is becoming more and more viable and useful considering the recent advances in next-generation sequencing technology and the decrease in price of short-read sequencing.

      We thank the reviewer for their comments and valuable feedback.

      Two major weaknesses are:

      (1) The novelty of the approach and its advantages should be more explicitly stated.

      (2) The authors do not discuss in depth the strength of using a combination of two methods rather than either of the two, especially considering that previously known CREs do not overlap with conserved sequences.

      We have added two paragraphs at the start of the Discussion to address the reviewer's comments 1 and 2 more explicitly (see response to reviewer 1, comment 4).

      Previously known CREs do include some conserved sequences, see Suppl. Figure 7.

      Reviewer #2 (Recommendations for the authors):

      In addition to addressing the two above-mentioned weaknesses, the authors should address the following minor issues:

      (1) It is difficult to refer to particular regions of text without line numbers.

      Spelling of ChIPseq is inconsistent in the Introduction.

      Spelling corrected. (Sorry for not including line numbering, we'll try to remember next time.)

      (2) "6.4% of reads from P. aquilina, 4.1% of reads from P. darvishi, and 0.46 % of reads from P. plumicornis could be mapped unambiguously to the P. hawaiensis genome" seems quite low for closely related species. Do the authors expect such low rates?

      Neutral nucleotide substitution rates in multicellular animals are in the order of 1 per site per 100 million years (e.g. https://pubmed.ncbi.nlm.nih.gov/12949132/) or a little lower (e.g. https://pubmed.ncbi.nlm.nih.gov/11792858/, https://pubmed.ncbi.nlm.nih.gov/34049492/). With the evolutionary times separating P. hawaiensis from P. aquilina/darvishi and P. plumicornis estimated at roughly 50 and 180 million years (2x25 and 2x90 million years, respectively), we expect a large fraction of neutrally evolving nucleotides in these genomes to have changed. We performed the read mapping using bowtie2, which requires a ~20 nt long perfect match with the reference sequence. We were therefore not surprised to obtain such low rates of read mapping. In fact, these low mapping rates (long divergence times) are important for islands of sequence conservation to stand out.

      (3) "Of these, 37% are found in introns, 54% in intergenic regions, and 1% overlap with promoters (TSS), marking regions that evolve at a lower rate than surrounding non-coding sequences". The authors explain in the Methods why they omit exons, but in the Results and Discussion, it is not stated. In addition, discussing the conservation with exons would be helpful, and the % in exons should be compared to non-coding regions.

      We have added "Of these, 8% are found in exons, likely reflecting conservation in protein-coding sequences".

      (4) "Two of the 7 reporters we tested, named neuro5 and neuro6," if I understood correctly, neuro5 and neuro6 are CREs, however, they are named quite ambiguously, and their names can be mistaken for gene names.

      Indeed, neuro5 and neur6 are the names of the CRE reporters. We have now added the names of the corresponding genes ("carrying CREs associated with the genes αTub and Cdk5α, respectively"). The gene names are also given in Table 3.

      (5) Why was single-end sequencing done for E24?

      We now explain this in the Methods: "Sequencing was carried out on an Illumina NextSeq 500 sequencer; we carried out single-end 76 bp sequencing for the first sample we generated (E24), and then switched to paired-end 76 bp sequencing for the other samples, because this leads to more specific read mapping."

      (6) Syntax related to in-line references should be double checked as the following sentences are broken by parentheses, e.g., "updated in (Almazán et al. 2022))".

      Corrected.

      (7) Could the authors discuss the P. aquilina genome size, which was estimated to be 3-times less than P. hawaiensis? Considering that in their phylogeny these two species are closest, it is quite surprising that they have such differing genome sizes. Do you expect it to be true? If yes, what could be the reason?

      As we explain in the manuscript, our estimates of genome size were obtained by dividing the total number of nucleotides sequenced by the estimated genome coverage, for each species. This method could overestimate genome sizes if there was a significant fraction of contaminating DNA in the preps, or a high degree of sequence variation that would prevent efficient mapping to BUSCO genes (both would underestimate genome coverage), but we find no evidence of this when we estimate the genome size of P. hawaiensis (see manuscript). We used the same method to estimate genome size in all four Parhyale species and have no reason to think that the method would be biased in one species and not in others. We therefore think that we have comparable estimates of genome size for the 4 species and the size difference is real.

      Variations in genome size can be driven by changes in the fraction of repetitive sequences found in a genome. We therefore checked the proportion of repetitive elements in each Parhyale genome using dnaPipeTE (https://github.com/clemgoub/dnaPipeTE). Based on this method (which likely underestimates the repetitive genome content) we find that the genomes of P. hawaiensis, P. aquiline, P. darvishi and P. plumicornis contain 31%, 22%, 18% and 39% of repetitive sequences, respectively. These figures do not fully account for the differences in genome size (particularly since P. darvishi appears to have even fewer repetitive sequences than P. aquiline). We therefore hesitate to add this very preliminary analysis to the manuscript.

      Of note, such rapid change in genome size is not unprecedented: in fruit flies genome size can vary more than 3-fold in species that have diverged over about 30 million years (https://elifesciences.org/articles/66405).

      (8) Wording "and found a genome coverage of 5.8x, corresponding to a genome size of 3.0 Gbp instead of 3.6 Gbp" is confusing and unclear as to what the authors exactly did here.

      We modified the sentence: "As a control, we followed the same procedure for P. hawaiensis, for which genome size is known (ref), and found a genome size of 3.0 Gbp instead of 3.6 Gbp (with a genome coverage of 5.8x)."

      (9) In the figures and supplementary figures, the genome browser screenshots should also include tracks of macs2 called peaks (those in narrowPeak format).

      Each ATACseq and sequence conservation track has its own set of peaks; we think that adding more tracks would overcrowd the figures. All the tracks (including called peaks) are provided as genome-browser-readable files in Supplementary Data files 1-3, so readers should be able to explore the data and reconstruct the figure panels without much effort.

      Reviewer #3 (Public review):

      Summary:

      Forbes et al. present a new approach for identifying cis-regulatory elements in large genomes. Using Parhyale hawaiensis, a crustacean with a large genome (~3.6 Gb, comparable in size to the human genome), the authors show that current methods for identifying cis-regulatory elements, effective in smaller genomes, are markedly inefficient in organisms with large genomes. To address this limitation, they combine bulk ATAC-seq and single-cell (sc) ATAC-seq to identify chromatin regions that are either ubiquitously accessible or specifically accessible in particular cell types. They further integrate comparative genomics across multiple Parhyale species (P. hawaiensis, P. aquilina, and P. darvishi), selected at appropriate phylogenetic distances (20-95 million years divergence), to pinpoint conserved open chromatin regions likely under functional constraint.

      Using this strategy, the authors predict a set of ubiquitous and cell-type-specific cis-regulatory elements. Importantly, they validate these predictions using rigorous transgenic reporter assays, convincingly demonstrating that their approach can successfully identify functional regulatory elements where previous methods had failed.

      Strengths:

      The approach introduced by Forbes et al. is conceptually straightforward, efficient, and readily transferable to other organisms. The validation experiments show not only that a substantial proportion of the predicted elements are functional, but also that the method is capable of identifying both ubiquitous and cell-type-specific regulatory elements. Given that the identification of regulatory regions remains a major bottleneck in understanding the molecular mechanisms underlying processes of development and regeneration, this work has the potential to make a significant impact in developmental and regeneration biology, particularly for studies involving non-model organisms with large genomes.

      An additional strength is the demonstration that only the genome of the focal species requires high-quality sequencing and assembly. In contrast, species used solely for comparative analysis can be sequenced at low coverage without assembly, substantially reducing costs and increasing the accessibility of the approach.

      Weaknesses:

      While the method is effective in identifying regulatory elements that are active ubiquitously or in differentiated cell types, it failed in detecting elements associated with developmentally regulated genes. This may be due to trivial reasons, such as a very low level of expression of the selected genes. However, as acknowledged by the authors, it may also indicate inherent challenges in identifying regulatory elements associated with developmentally dynamic gene regulation, compared to those associated with genes expressed in differentiated cell types.

      A second limitation, also acknowledged by the authors, is the absence of chromatin conformation capture data, which would help link distal regulatory elements to their target genes. This limitation may be particularly relevant for developmentally regulated genes, where long-range regulatory interactions may be critical.

      Addressing these limitations will be an important direction for future work. Nonetheless, the approach as presented in this manuscript represents a key contribution that sets the stage for further methodological advances in the identification of cis-regulatory elements in large genomes.

      Reviewer #3 (Recommendations for the authors):

      I have no specific comment for the authors. While in my opinion the study has two limitations (as described in the public review), these are clearly acknowledged and properly discussed in the manuscript.

      The manuscript is extremely well written. It has been a great pleasure to read it. Excellent job!

      Thank you!

    1. eLife Assessment

      The report by Liu and colleagues reports on an analysis of environmental adaptation across diverse lineages of the grass Phragmites australis differing by their level of ploidy. These results represent important findings for understanding the environmental adaptation of species complexes with mixed ploidy and the analysis reports solid evidence that lineages with distinct levels of ploidy occupy different climate niches. The use of regional survey in tandem with common garden experiment represents a convincing approach to suggest a correlation between ploidy and climate adaptation. This manuscript will be of interest to a broad community of ecological genomicists interested in how structural variation in gene dosage potentially affects the pattern of adaptation.

    2. Reviewer #1 (Public review):

      Summary:

      The article is testing the relative advantages of plant lineages with differing ploidy and admixture across environmental gradients. The results show that intraspecific variation in ploidy and admixture between lineages impacts plant traits that may enable persistence and range expansion.

      Strengths:

      Suitable marker panel size and strong results that include attempts to analyse mixed ploidy level data which is a challenge.

      Weaknesses:

      The sample sizes of the common garden experiments are very low making it difficult to draw robust conclusions.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Overview of Revisions

      We thank the editors and reviewers for their constructive and insightful comments, which have substantially improved the manuscript. We have carefully addressed every point raised. The major revisions include:

      (1) Methods 2.1: completely reorganized to clarify the allotetraploid genome structure of Phragmites australis, the rationale for single-chromosome-anchored microsatellite markers, the maximum distinguishable alleles per ploidy level, and the conservative Ploidies(mydata) <- 4 setting in polysat.

      (2) Methods 2.3 / Results 3.3 / Discussion 4.4: clarified common garden sample sizes, added Cohen's d effect sizes, and acknowledged the correlational nature of the lineage-level comparisons.

      (3) Introduction: added a new paragraph on the eco-evolutionary significance of gene flow in mixed-ploidy systems, and another paragraph emphasizing the novelty of integrating SDMs with physiological and common garden experiments.

      (4) Discussion 4.1 and 4.4: reframed all "polyploidy-driven" language to "polyploidy-associated", explicitly acknowledging that ploidy is confounded with genetic background and was not experimentally manipulated.

      Public Reviews:

      Reviewer #1 (Public review):

      (R1-P1) Inadequate explanation of allele dosage for ploidy levels

      Inadequate explanation of allele dosage for ploidy levels, some of which do not match the allele counts expected for genome copy number.

      We appreciate this comment and have substantially revised Methods 2.1. The key clarifications are:

      (A) Marker specificity: All 42 microsatellite markers were aligned to the P. australis reference genome, and each marker mapped to a single unique chromosome (Table S2). This confirms that each marker amplifies a locus specific to one subgenome only. Therefore, in tetraploids each marker detects at most two alleles (the two homologous copies of that chromosome from one subgenome), while in octoploids (autopolyploid derivatives with four copies of the same chromosome) each marker detects up to four alleles.

      (B) Conservative ploidy setting: We set Ploidies(mydata) <- 4 for all samples in polysat because the exact ploidy of many samples could not be confidently assigned a priori. This uniform treatment is conservative: for actual tetraploids, the two unobserved "copies" are scored as null; for actual octoploids, all four detected alleles are accommodated.

      (C) Dosage estimation: Allele dosage was estimated from high-coverage sequencing read counts (mean >5,000× per locus per sample) using the SSRSeq V1.1 pipeline (Cui et al., 2022), which applies stutter correction, amplification bias correction, and ploidy-optimized dosage calling, not inferred from allele presence/absence alone.

      Methods 2.1, second paragraph onward

      Phragmites australis has a base allotetraploid genome. All 42 microsatellite markers used in this study were aligned to the P. australis reference genome, and each marker mapped to a single unique chromosome (Table S2), confirming that each marker amplifies from one subgenome only. Therefore, in tetraploids each marker detects at most two alleles (the two homologous copies of that chromosome from the target subgenome) while the homologous region from the other subgenome is not amplified (Saltonstall, 2003). In Asia, the prevalent octoploids are most likely autopolyploid derivatives of tetraploids, carrying four homologous copies of the same chromosome and thus capable of up to four distinguishable alleles per locus (Liu et al., 2022; Wang et al., 2024). Hexaploid individuals are rare and occur primarily in contact zones, likely originating from inter‑lineage hybridization (Wang et al., 2024).

      In practice, the 42 selected markers very rarely produced more than four alleles in any single individual (Table S2), consistent with a ploidy ceiling of octoploid. Because the exact ploidy of many samples could not be confidently assigned a priori (ploidy was inferred from a combination of chloroplast haplotype, geographic origin, and flow cytometry from prior studies; Lambertini et al., 2020; Liu et al., 2022), we consistently set Ploidies(mydata) <- 4 in the polysat R package (Clark & Jasieniuk, 2011), treating every individual as having four homologous copies. This uniform treatment is conservative: for an actual tetraploid (two copies per locus), the two unobserved "copies" are simply scored as null (missing data) in the dosage matrix; for an actual octoploid, all four detected alleles are accommodated. The allele dosage itself was estimated from high-coverage sequencing read counts (mean >5,000× per locus per sample) using the SSRSeq V1.1 pipeline (Cui et al., 2022), not inferred from allele counts alone.”

      (R1-P2) Common garden setup and sample sizes unclear

      The setup and sample sizes of the common garden experiments are very unclear. The numbers implied are extremely low to draw robust conclusions.

      We agree that the original description was insufficiently detailed. We have made three modifications:

      (A) Methods 2.3: clarified that each population contributed one rhizome segment planted in one pot (one biological replicate per population per site), and emphasized that the key inference rests on the direction and consistency of differences across four climatically distinct sites rather than on significance at any single site.

      (B) Results 3.3: added Cohen's d effect sizes (verified from the raw data), showing that where lineage differences are present, they are biologically substantial (d = 1.10–1.43 at three of four sites).

      (C) Discussion 4.4: added a paragraph acknowledging the limited number of populations per lineage and the need for future confirmation with a larger panel.

      Methods 2.3

      “(1) A previously published common garden experiment (Song et al., 2021) conducted in 2017 across Jinan (36.43°N, 117.45°E) and Panjin (41.20°N, 122.02°E), using CN (n = 11) and FEAU (n = 9) lineages, for which we determined the haplotype information of all samples; (2) A new common garden experiment established in 2021 across Qingdao (36.36°N, 120.69°E) and Shanghai (30.20°N, 121.29°E), with CN (n = 9) and FEAU (n = 8) lineages (Table S3). Each rhizome segment (2–3 buds per segment, one segment per population) was transplanted into an individual 20 L pot, yielding one biological replicate per population per site. Although the number of populations per lineage is modest, the key inference rests on the direction and consistency of lineage differences across four climatically distinct sites rather than on the statistical significance at any single site.”

      Results 3.3

      “Similarly, plant height was significantly greater in the FEAU lineage than in the CN lineage in Jinan, but not in the other common gardens (Figure 3C). Effect sizes for total biomass were large in Jinan (Cohen's d = 1.10), Panjin (d = 1.12), and Qingdao (d = 1.43), but negligible in Shanghai (d = 0.37), confirming that the lineage differences, where present, are biologically substantial.”

      Discussion 4.4

      “We also acknowledge that the common garden experiments, while replicated across four climatically distinct sites, involved a limited number of populations per lineage (9–11 CN and 8–9 FEAU), which constrains our ability to fully separate lineage-level effects from population-level variation. The consistent direction of biomass differences across three of four sites, supported by large effect sizes, nonetheless provides robust evidence for a lineage-level performance advantage that merits further confirmation with a larger, more geographically representative panel of populations.”

      (R1-P3) How allele dosage is determined

      Unclear how allele dosage is determined. Given it's so central to many analyses, it would be useful to see how this is done rather than use a citation.

      We agree that a self-contained description is warranted. We have rewritten the relevant paragraph in Methods 2.1 to describe the three core steps of the SSRSeq V1.1 pipeline (Cui et al., 2022): (i) stutter correction based on empirically estimated slip ratios; (ii) amplification bias correction across alleles of different repeat lengths; and (iii) ploidy-adjusted allele dosage calling. The pipeline's source code and full documentation are available at https://github.com/ccoo22/SSRseq_count.

      Methods 2.1

      “Microsatellite genotyping was performed using the SSRSeq V1.1 pipeline (Cui et al., 2022; https://github.com/ccoo22/SSRseq_count). Briefly, the pipeline takes the per-locus per-sample read count table generated from high-throughput sequencing and processes it through three core steps fully described in Cui et al. (2022): (i) stutter correction, which reallocates a fraction of reads from each allele to its adjacent repeat class based on empirically estimated slip ratios; (ii) amplification bias correction, which normalizes read counts across alleles of different repeat lengths using locus-specific bias coefficients; and (iii) allele dosage calling, which selects the maximum number of alleles consistent with the specified ploidy (four in this study) and assigns integer dosages (0–4) by comparing corrected read ratios to a ploidy-adjusted threshold optimized to minimize both allelic dropout and false positives. The final output is a genotype matrix with integer allele dosages for all samples and loci, which was used directly as input to the polysat R package for subsequent population genetic analyses.”

      Reviewer #2 (Public review):

      (R2-P1) Polyploidy has no causal evidence; confounded with genetic background

      First, no data support the claims that polyploidy has any causal effect. The ploidy levels are, in fact, completely confounded with other genetic differences, so it is not possible to eliminate genetic variation, independent of ploidy, as the causative factor. As the authors note, ploidy was not manipulated in the reported experiments. Thus, the focus on polyploidy in the introduction and elsewhere distracts from the novel and informative experiments that were conducted.

      We fully acknowledge this critical limitation and thank the reviewer for this important critique. We have revised the manuscript at four locations to reframe all claims from "polyploidy-driven" to "polyploidy-associated" and to explicitly state that ploidy is confounded with lineage identity and was not experimentally manipulated.

      Abstract (last sentence)

      “These results demonstrate that climate change interacts with intraspecific variation among polyploidy-associated lineages, manifested through differences in thermal tolerance, biomass production, and asymmetric gene flow, to drive potential lineage replacement within a native range”

      Introduction (polyploidy paragraph)

      “The octoploid FEAU lineage is distinguished from its tetraploid relatives not only by ploidy level but also by its distinct evolutionary history, genomic background, and geographic origin. Polyploidy has been shown in other systems to generate genetic novelty, alter gene expression, and enhance physiological stress tolerance (Bureš et al., 2024; Cheng et al., 2021; Kolář et al., 2017; Van de Peer et al., 2017), potentially pre-equipping polyploid lineages to occupy new geographical ranges and endure environmental shifts (Cheng et al., 2021; López-Jurado et al., 2019). The FEAU lineage's superior thermal tolerance and biomass are consistent with such polyploidy-associated effects, although ploidy is correlated with, rather than experimentally separable from, the broader genetic identity of each lineage.”

      Discussion 4.1 (title and opening paragraph)

      “Our findings demonstrate that the octoploid FEAU lineage of P. australis possesses greater heat tolerance and biomass production than the tetraploid CN lineage. Under a high emission scenario (SSP5-8.5), the projected suitable habitat for the FEAU lineage expands by 18.6%, while the CN lineage exhibits a much smaller relative increase. Several non-mutually-exclusive mechanisms could explain these lineage-level differences, including increased gene dosage from whole-genome duplication, divergent selection histories, and/or standing genetic variation in thermal tolerance loci unlinked to ploidy (Bures et al., 2024; Cheng et al., 2021; Van de Peer et al., 2017). Our data cannot fully partition these factors, but the strong association between lineage identity and both physiological performance and projected range dynamics highlights the importance of incorporating intraspecific lineage information into ecological forecasts, regardless of the ultimate causal mechanism.”

      Discussion 4.4 (limitations paragraph)

      “Crucially, ploidy was not experimentally manipulated in this study; it is inherently confounded with the distinct evolutionary history and genomic background of each lineage. While the observed thermal tolerance and biomass differences are consistently associated with the octoploid FEAU lineage, we cannot formally exclude the possibility that these traits are driven by genetic factors independent of ploidy per se. Future studies using experimental approaches that can partition ploidy effects from lineage-specific genetic effects, such as common gardens with synthetic polyploids or transcriptomic analyses comparing gene expression dosage responses, are needed to strengthen causal inference (Wei et al., 2020). Similarly, the common garden results should be interpreted as lineage-associated rather than ploidy-causal performance differences. The potential role of admixture in facilitating the adaptive introgression of heat tolerance alleles also warrants deeper investigation (Suarez-Gonzalez et al., 2018).”

      (R2-P2) SDMs treat lineages as homogeneous entities

      Second, the manuscript indicates that intraspecific variation is critical for the evolutionary potential of a species to respond to environmental change, but intraspecific variation is seldom considered in species distribution models... the manuscript performs species distribution modeling on a small number of sub-specific lineages, essentially treating them as homogeneous "species"—thus the analysis commits the same oversimplification that the manuscript highlights, but does so at a finer evolutionary scale than species.

      We acknowledge this important limitation and agree that it deserves explicit discussion. While disaggregating the species into three major genetic lineages is a step forward from species-as-monolith approaches, within-lineage variation in thermal tolerance, growth, and dispersal capacity is plausible given the broad geographic ranges of the CN and FEAU lineages. We have added a new paragraph in Discussion 4.4 to address this point.

      Discussion 4.4 (new paragraph)

      “We also recognize that our SDM approach, while disaggregating the species into three major genetic lineages, still treats each lineage as a homogeneous entity. This simplification parallels—albeit at a finer scale—the species-as-monolith assumption that we critique in the Introduction. Within-lineage variation in thermal tolerance, growth, and dispersal capacity is plausible, particularly given the broad geographic ranges of the CN and FEAU lineages. By modelling each lineage as a uniform group, our projections may overestimate the precision of range forecasts and underestimate the evolutionary potential of standing variation within lineages (Chardon et al., 2020). Future frameworks that incorporate trait variation at multiple hierarchical levels (population, lineage, ploidy) will be necessary to capture both the adaptive potential and the ecological constraints that shape species' responses to climate change.”

      (R2-P3) Asymmetric introgression and thermal tolerance lack context in Introduction and Discussion

      The title suggests that asymmetric introgression and thermal tolerance are the most important findings of the work. However, the introduction contains no explanation of the potential importance of gene flow (other than to say that asymmetric gene flow was suggested by some preliminary analyses), and the discussion offers only a limited explanation of either the potential mechanisms underlying the asymmetric gene flow or its importance for the long-term evolution of the species.

      We agree that the evolutionary significance of asymmetric gene flow was underdeveloped. We have added two substantial new passages:

      (A) Introduction: a new paragraph explaining the dual role of gene flow in climate adaptation: introgression of adaptive alleles vs. asymmetric introgression as a mechanism of gradual lineage replacement. This paragraph explicitly connects genome dosage differences (octoploid vs. tetraploid) to the natural directionality of backcrossing.

      (B) Discussion 4.2: a new paragraph extending the discussion of asymmetric introgression into its long-term evolutionary consequences, including the potential erosion of the CN lineage's genetic distinctiveness and the risk of losing cold-adapted alleles under future climate volatility.

      Introduction (new paragraph)

      “Gene flow between lineages of differing ploidy can play a dual role in climate adaptation. Introgression may introduce adaptive alleles (e.g., heat tolerance loci) into a recipient lineage, facilitating its persistence under warming (Suarez-Gonzalez et al., 2018). Conversely, if introgression is asymmetric, such that one lineage's genome is disproportionately represented in admixed populations, it can drive a gradual but systematic shift in genetic composition within the contact zone—effectively functioning as a mechanism of lineage replacement without requiring complete competitive exclusion (Bartolić et al., 2024; Zohren et al., 2016). In mixed-ploidy systems, genome dosage differences create a natural directionality in backcrossing: hybrids tend to backcross more frequently with the high-ploidy parent (Bartolić et al., 2024). In the present study, we test whether such a bias exists between the octoploid FEAU and tetraploid CN lineages and examine its consequences for future distribution under climate warming.”

      Discussion 4.2 (new paragraph)

      “From an evolutionary standpoint, asymmetric introgression can erode the genetic distinctiveness of the minority lineage (CN) while enriching the majority lineage (FEAU) with alleles that may have been locally adapted in the CN genomic background. This could reduce the species' overall evolutionary potential, even if the FEAU lineage itself thrives—because cold-adapted alleles from the CN lineage, which may be valuable under future climate volatility (including extreme cold events), risk being diluted or lost (Exposito-Alonso et al., 2022). The directionality of introgression is also not fixed; it could shift if environmental conditions alter hybrid fitness or if the demographic balance between lineages changes. Long-term genomic monitoring of the CN–FEAU contact zone will be essential to determine whether the asymmetric gene flow documented here represents a transient phase or a persistent trajectory toward genomic homogenization.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      (RE-1) Explain how ploidy level is inferred

      We invite the authors to clearly explain how the level of ploidy is being inferred (Reviewer #1).

      We agree that the rationale for ploidy assignment and its relationship to allele counts needed greater clarity. This has been addressed by the comprehensive revision of Methods 2.1 described in response to R1-P1 (Part 1, Reviewer #1 Public Reviews). The revised text now presents a complete step-by-step logical chain: (i) P. australis has an allotetraploid base genome; (ii) all 42 markers map to a single unique chromosome in the reference genome, confirming that each marker amplifies from only one subgenome; (iii) tetraploids therefore show at most two distinguishable alleles per locus, while octoploids (autopolyploid derivatives of tetraploids) show up to four; (iv) because many samples lacked independent ploidy confirmation, we uniformly set Ploidies(mydata) <- 4 in polysat as a conservative treatment that accommodates both tetraploids (two observed copies + two null) and octoploids (four observed copies).

      See the full revised text under R1-P1 (Part 1) above.

      (RE-2) Common garden results are correlational

      We note that the findings of the common garden experiment, although interesting, are mostly correlational (not causative) and rely on a relatively small sample size and confound lineage isolation and adaptive differentiation (both Reviewers).

      We fully acknowledge this limitation. Because all octoploids belong to the FEAU lineage and all tetraploids to CN, ploidy and lineage identity are inherently confounded. This has been addressed by the four-part revision described in response to R2-P1 (Part 1, Reviewer #2 Public Reviews). Specifically:

      The Abstract now frames the findings as "polyploidy-associated" rather than "rooted in polyploidy."

      The Introduction now explicitly states that ploidy is correlated with—but not experimentally separable from—the broader genetic identity of each lineage.

      The Discussion 4.1 title was changed to "Polyploidy-associated thermal tolerance" and the opening paragraph now presents multiple non-mutually-exclusive mechanisms rather than asserting a causal role for polyploidy.

      The Discussion 4.4 now includes an expanded limitations paragraph acknowledging that ploidy was not experimentally manipulated and that common garden results should be interpreted as lineage-associated rather than ploidy-causal.

      In addition, the Methods 2.3 and Discussion 4.4 revisions described in response to R1-P2 (Part 1) address the sample size concern by clarifying the experimental design and adding a dedicated acknowledgement of the limited population replication.

      See the full revised text under R2-P1 and R1-P2 (Part 1) above.

      Reviewer #1 (Recommendations for the authors):

      (R1-R1) "Large morphological traits"

      Line 85. Large morphological traits. Does this mean physically large? Or higher values of some trait.

      We agree the original phrasing was ambiguous. We have replaced "large morphological traits" with explicit trait descriptions.

      Introduction

      “The octoploid FEAU lineage exhibits greater shoot height, larger leaf size, and thicker stems (K. Chen et al., 1993; Guo et al., 2025; Liu et al., 2021b, 2026; Yin et al., 2024), along with stronger salt tolerance and higher thermal tolerance”

      (R1-R2) Why 2 allele copies expected for a tetraploid

      Line 119. Not clear why this allele copy number is expected. A tetraploid can have up to 4 unique alleles (e.g., ABCD), not two.

      This comment arises from the same conceptual gap addressed in R1-P1. The key point is that P. australis is an allotetraploid with two subgenomes, and our 42 markers each map to a single unique chromosome (one subgenome). Therefore, the marker only amplifies the two homologous copies from that subgenome, giving at most two distinguishable alleles. A true autotetraploid would indeed show up to four alleles—but that is not the genomic architecture of P. australis. The revised Methods 2.1 (see R1-P1 in Part 1) now explicitly explains this logic.

      Fully addressed by the Methods 2.1 revision in R1-P1.

      (R1-R3) Theoretical expectation of allele number vs. ploidy

      Line 127-129. As above, this is unclear and not what we expect theoretically. If there is a reasonable number of alleles, there should be a maximum of 4 for tets, 6 for hex, and 8 for octs.

      Same point as R1-R2. The reviewer's expectation (4 for tetraploids, 6 for hexaploids, 8 for octoploids) is correct for autopolyploids with markers that amplify all homologous copies. The discrepancy arises because P. australis is an allotetraploid and our markers are single-chromosome-anchored (each amplifying from only one subgenome). The revised Methods 2.1 now clarifies this distinction explicitly.

      Fully addressed by the Methods 2.1 revision in R1-P1.

      (R1-R4) Typo "makers"

      Line 136. Should be 'markers' not 'makers'

      We have performed a full-text search and corrected all instances of "makers" to "markers" in the manuscript.

      Full-text search and replace.

      (R1-R5) Reason for removing markers with >4 alleles

      Line 142. The reason for the removal of more than 4 alleles is not clear. What about hexaploids and octoploids? They can carry 6 or 8 alleles, respectively.

      We agree the original text did not adequately justify this quality-control step. In our study, the maximum expected distinguishable alleles (given the allotetraploid genome and single-chromosome-anchored markers) is two for tetraploids and four for octoploids. The observation of five or more alleles in multiple individuals is therefore diagnostic of multi-locus amplification (the marker amplifying more than one genomic locus), not of high ploidy. This is a quality-control filter, not a ploidy assignment criterion.

      Methods 2.1

      “During genotyping, eleven markers (including four of the five multi-mapping markers) were removed because more than ten samples exhibited more than four alleles per sample at these loci. Because the maximum number of distinguishable alleles expected under our ploidy model is two (tetraploid) to four (octoploid), the observation of five or more alleles in multiple individuals indicates that these markers amplify more than one genomic locus, rendering them unsuitable for dosage-based genotyping. This filtration is a quality-control step, not a ploidy assignment criterion.”

      (R1-R6) Unclear sample sizes in common garden

      Line 240. Unclear sample sizes. If these are the numbers, they are a very low level of replication expected for a common garden experiment.

      Same point as R1-P2 (Public Review). Please see the full response under R1-P2 in Part 1, where we have (A) clarified the experimental design in Methods 2.3, (B) added Cohen's d effect sizes in Results 3.3, and (C) acknowledged the sample size limitation in Discussion 4.4.

      Fully addressed by the three-part revision in R1-P2.

      Reviewer #2 (Recommendations for the authors):

      (R2-A1) De-emphasize polyploidy

      De-emphasize polyploidy, as it's not manipulated in the study and is entirely confounded with the genotypes of the distinct lineages, and the putative links between polyploidy and heat tolerance are circumstantial and lacking in a clear mechanism.

      We agree fully. This has been addressed comprehensively across four locations in the manuscript (Abstract, Introduction, Discussion 4.1, Discussion 4.4). See the full response under R2-P1 (Part 1) for the revised text at each location.

      Fully addressed by the four-part revision in R2-P1.

      (R2-A2) Emphasize the novelty of combining SDMs with experiments

      Emphasize the novelty of combining SDMs with experiments (or, if I'm not up on the literature and they are more common, explain how they have been used to make new insights).

      We appreciate this suggestion and agree that explicitly stating the novelty of our integrative approach strengthens the manuscript. We have added a new paragraph at the end of the Introduction.

      Introduction (end, before "Here, we integrate population genomics…")

      “Studies that combine species distribution models with physiological or common garden experiments remain surprisingly uncommon (but see López-Jurado et al., 2019). Such integration is essential for transforming correlative SDM projections into mechanistically grounded predictions. In the present study, we adopt this integrative approach: common garden experiments directly test growth performance under controlled conditions, heat-tolerance measurements identify the specific physiological thresholds (T<sub>crit</sub>, T<sub>50</sub>) underlying lineage-specific climate responses, and SDMs project how these experimentally documented differences translate into spatial dynamics under future warming. By linking experimental data with spatial forecasting, we move beyond correlative climate matching toward a trait-based understanding of how intraspecific variation shapes species' future distributions.”

      (R2-A3) Elaborate on the importance of gene flow

      Elaborate on the importance of gene flow and the potential connections between gene flow and evolving species (or lineage) geographical limits.

      This has been addressed by the two new paragraphs described under R2-P3 (Part 1)—one in the Introduction on the dual role of gene flow in climate adaptation (introgression of adaptive alleles vs. asymmetric introgression as a mechanism of lineage replacement), and one in Discussion 4.2 on the long-term evolutionary consequences of asymmetric introgression.

      Fully addressed by the two-part revision in R2-P3.

    1. eLife Assessment

      This study provides valuable insights into the cellular dynamics underlying accelerated tooth regeneration in a vertebrate model. Using single-nucleus RNA sequencing across multiple time points, the authors present a well-structured analysis of cell populations, trajectories, and intercellular signaling events associated with this process. The strength of evidence is solid but only partially supported, as the conclusions are primarily supported by computational inference, without experimental validation of key findings.

    2. Reviewer #1 (Public review):

      Summary:

      The authors used single-nucleus RNA sequencing (snRNA-seq) to investigate accelerated tooth replacement following tooth plucking in cichlid fish. They analyzed four stages of regeneration using elegant and well-designed approaches to characterize cellular trajectories and interactions within the dental epithelium and mesenchyme during the accelerated replacement process. Their analyses identified cell type-specific gene expression profiles and intercellular signaling interactions associated with whole-tooth regeneration.

      Strengths:

      This is a highly interesting and thoughtfully executed study that provides compelling and convincing insights into the mechanisms underlying accelerated tooth regeneration.

      Comments on revised version.

      I noted in my initial review that "the manuscript currently lacks experimental validation of the single-nucleus RNA-seq data." In response, the authors have added a statement indicating that their cell-type annotations and pathway interpretations are supported by extensive prior experimental work in the cichlid tooth model, including histology, in situ hybridization, immunohistochemistry, and pharmacological perturbation of major developmental pathways. They have also acknowledged this limitation in the Study Limitations and Future Directions section, stating that direct experimental validation of the single-nucleus RNA-seq findings will be the focus of future studies.

      The authors have carefully addressed my comments, particularly the Major Points (2), (3), and (4), as well as all of the Minor Points. I appreciate their efforts to further characterize the mesenchymal landscape surrounding the putative successional lamina and to provide additional evidence supporting the presence of a specialized stromal microenvironment associated with tooth regeneration. Overall, the revisions have substantially strengthened the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      Mubeen and colleagues study the cellular basis of tooth regeneration in cichlid fish. Using an elegant tooth plunking strategy followed by single nucleus RNA-sequencing, the authors were hoping to achieve an atlas of cellular and transcriptional changes that occur within and between cells during whole tooth replacement.

      Strengths:

      The major strengths of the methods and results are high novelty in the approach in a vertebrate with continuous tooth replacement, the temporal analysis of analyzing at plucking and three later time points, the thorough and sophisticated analysis of the snRNA-seq data including the inferring of trajectories and signaling events, and the robust signal of transcriptional differences induced by tooth plucking.

      Weaknesses:

      The major weaknesses of the methods and results are no validation of any of the inferred cell types, no functional tests of whether any of the changes in signaling pathways affect the plucking-induced tooth replacement process, and perhaps no clear take-away message for biologists not necessarily interested in tooth replacement.

      Conclusions:

      The authors achieved their aims of identifying the changes in gene expression and cellular composition that occur during whole tooth replacement accelerated by plucking. Overall, the results support their conclusions, although some slight semantic qualifiers should probably be added (e.g. referring to "cell types" as "putative cell types").

      The work should have high impact in the field of tooth and organ regeneration, and the novel methodological paradigm established here of accelerating tooth replacement three-fold by plucking has great promise for future follow up studies to further study this process. The work also could have strong impact by the computational methods used here to infer trajectories and signaling interactions. Specific pathways, genes, and cell types could be tested in other fish such as zebrafish to test function during tooth replacement.

      The work is unique and interdisciplinary and also has significance by establishing that robust phenotypically plastic accelerations in regeneration rates occur upon tooth removal. There are very few studies like this one that combine genetic x environmental studies of regeneration. The result that three different species of cichlid fish that normally have very different tooth patterns all accelerate tooth replacement threefold upon tooth plucking also has significance in revealing a highly conserved plucking response.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Many thanks to the three reviewers and the editors for their thoughtful comments and careful evaluation of our manuscript. These are fair, consistent and largely expected comments. On behalf of my co-authors, we provide this response to the public reviews to summarize the main issues raised and the corresponding revisions we have made in the revised manuscript.

      (1) The main consistent comment from all three referees was that our single-nucleus RNA-seq data should be further validated. The reviewers differ in the detail of exactly what they think should be validated, but collectively these comments referred to validation of: (1) the identified cell types, (2) pathways inferred from trajectory analysis, (3) differentially expressed genes between plucked and control conditions across the four sampled time points, and/or (4) inferred ligand–receptor pairs from the cell–cell communication analysis.

      We believe that we are on strong footing for some of these points because of extensive work we’ve done in the past in the cichlid fish model.

      In the references cited in the manuscript and highlighted below (References 1, 10, 11, 29, 30, 31), we tally 29 figures with 273 individual figure panels presenting histology, in situ hybridization, and immunohistochemistry featuring genes expressed in cichlid (replacement) teeth. Most of these genes are markers of dental competency and/or indicative of regenerative potential.

      In addition, in multiple of these papers, we use pharmacology to manipulate the role of key pathways (Hh, BMP, Wnt, Notch) in cichlid tooth development and replacement. Validation of cell types in the present study therefore draws on these published data in cichlids (and other vertebrates), as well as on an unbiased comparative approach, SAMap, which identifies homology between cichlid and mouse dental cell types based on shared gene expression.

      In short, experiments to validate cell types and pathways active in cichlid teeth have been published and are referenced herein. We recognized, however, that these references (some of which include Gareth Fraser as an author, when he was a postdoc in my group; for Reviewer 2) were cited primarily in the Introduction, rather than in the Rationale/Methods or Results sections. We have therefore clarified these connections in the revised manuscript (line 173-74).

      We have not validated nor analyzed functionally the ligand-receptor pairs we inferred from cell-cell communication analysis. This work is beyond the scope of the current paper, and we now state more clearly that these inferences represent hypotheses to be tested in future studies, although many of these ligand–receptor pairs have been noted in other tooth-related publications cited in the manuscript.

      (2) The biggest weakness of our manuscript, noted by referees, is that we do not provide serial histology to accompany our snRNA-seq time course after plucking. We previously described this as a limitation in the “Study limitations and future direction” section of the Discussion, but we have now strengthened this discussion. In particular, we more explicitly acknowledge that we do not directly document the histological progression of tissue responses across the plucking time course or the degree of tissue damage caused by the plucking paradigm at each sampled time point.

      In the “study limitations” section, we note both issues 1 and 2 and suggest that a spatial transcriptomics experiment across the timespan of plucking<>recovery would address simultaneously the desire to understand cellular context of plucking and cellular/spatial differences in plucked vs control cell-type gene expression.

      (3) Reviewers also asked about the presence and interpretation of stromal cells in our snRNA-seq data. In response, we re-examined the mesenchymal compartment and added additional analyses to better characterize stromal/mesenchymal populations and their inferred trajectories in the revised manuscript. This includes a revised Figure 4, revised text around Figure 4 and revised Supplementary Figures.

      (4) Multiple (minor) suggestions for clarification in text and figures have been adopted throughout the revised manuscript, figure legends, and supplemental materials.

      Overall, we do not anticipate that further reviewer engagement will be necessary, and we believe that editorial review of the revised manuscript should be sufficient.

      References cited in the manuscript, highlighted here:

      (1) Fraser, G. J. et al. An Ancient Gene Network Is Co-opted for Teeth on Old and New Jaws. PLoS Biol. 7, e1000031 (2009).

      (10) Fraser, G. J., Bloomquist, R. F. & Streelman, J. T. Common developmental pathways link tooth shape to regeneration. Dev. Biol. 377, 399–414 (2013).

      (11) Bloomquist, R. F. et al. Developmental plasticity of epithelial stem cells in tooth and taste bud renewal. Proc. Natl. Acad. Sci. 116, 17858–17866 (2019).

      (29) Streelman, J. T., Webb, J. F., Albertson, R. C. & Kocher, T. D. The cusp of evolution and development: a model of cichlid tooth shape diversity. Evol. Dev. 5, 600–608 (2003).

      (30) Fraser, G. J., Bloomquist, R. F. & Streelman, J. T. A periodic pattern generator for dental diversity. BMC Biol. 6, 32 (2008).

      (31) Bloomquist, R. F. et al. Coevolutionary patterning of teeth and taste buds. Proc. Natl. Acad. Sci. 112, (2015).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors used single-nucleus RNA sequencing (snRNA-seq) to investigate accelerated tooth replacement following tooth plucking in cichlid fish. They analyzed four stages of regeneration using elegant and well-designed approaches to characterize cellular trajectories and interactions within the dental epithelium and mesenchyme during the accelerated replacement process. Their analyses identified cell-type-specific gene expression profiles and intercellular signaling interactions associated with whole-tooth regeneration.

      Strengths:

      This is a highly interesting and thoughtfully executed study that provides compelling and convincing insights into the mechanisms underlying accelerated tooth regeneration.

      Weaknesses:

      The manuscript currently lacks experimental validation of the single-nucleus RNA-seq data.

      We thank Reviewer #1 for the thoughtful and positive assessment of our study, including the recognition that our snRNA-seq time course provides insight into cellular trajectories, cell-type-specific gene expression, and inferred intercellular signaling during accelerated tooth replacement in cichlid fish. We also appreciate the reviewer’s central concern that the manuscript would be strengthened by additional experimental validation of the single-nucleus RNA-seq data.

      As summarized above and discussed in more detail in our point-by-point responses below, we have clarified how the present cell-type annotations and pathway interpretations are supported by extensive prior experimental work in the cichlid tooth model, including histology, in situ hybridization, immunohistochemistry, and pharmacological perturbation of major developmental pathways. We have also added analyses demonstrating reproducibility across biological test subjects and consistency of representative differentially expressed genes between paired plucked and control samples. Finally, we have strengthened the Study Limitations section to more clearly state that future spatial transcriptomic, histological, and functional validation experiments will be important next steps.

      Reviewer #2 (Public review):

      Summary:

      Mubeen and colleagues studied the cellular basis of tooth regeneration in cichlid fish. Using an elegant tooth plunking strategy followed by single-nucleus RNA-sequencing, the authors were hoping to achieve an atlas of cellular and transcriptional changes that occur within and between cells during whole tooth replacement.

      Strengths:

      The major strengths of the methods and results are high novelty in the approach in a vertebrate with continuous tooth replacement, the temporal analysis of analyzing at plucking and three later time points, the thorough and sophisticated analysis of the snRNA-seq data, including the inference of trajectories and signaling events, and the robust signal of transcriptional differences induced by tooth plucking.

      Weaknesses:

      The major weaknesses of the methods and results are no validation of any of the inferred cell types, no functional tests of whether any of the changes in signaling pathways affect the plucking-induced tooth replacement process, and perhaps no clear takeaway message for biologists not necessarily interested in tooth replacement.

      Conclusion:

      The authors achieved their aims of identifying the changes in gene expression and cellular composition that occur during whole tooth replacement accelerated by plucking. Overall, the results support their conclusions, although some slight semantic qualifiers should probably be added (e.g., referring to "cell types" as "putative cell types").

      The work should have a high impact in the field of tooth and organ regeneration, and the novel methodological paradigm established here of accelerating tooth replacement three-fold by plucking has great promise for future follow-up studies to further study this process. The work could also have a strong impact through the computational methods used here to infer trajectories and signaling interactions. Specific pathways, genes, and cell types could be tested in other fish, such as zebrafish, to test function during tooth replacement.

      The work is unique and interdisciplinary, and also has significance by establishing that robust phenotypically plastic accelerations in regeneration rates occur upon tooth removal. There are very few studies like this one that combine genetic and environmental studies of regeneration. The result that three different species of cichlid fish that normally have very different tooth patterns all accelerate tooth replacement threefold upon tooth plucking also has significance in revealing a highly conserved plucking response.

      We thank Reviewer #2 for the careful and constructive evaluation of our manuscript and for highlighting the novelty of the cichlid tooth-plucking paradigm, the temporal design of the snRNA-seq experiment, and the computational analyses used to infer cellular trajectories and signaling interactions during accelerated tooth replacement. We also appreciate the reviewer’s comments regarding validation of inferred cell types and interpretation of signaling pathways.

      In response, we have revised the manuscript to clarify that our cell-type annotations are supported by marker-gene expression, previously published cichlid tooth studies, and an unbiased comparative approach, SAMap, which relates cichlid and mouse dental cell types based on shared gene-expression structure. We have also clarified that inferred ligand-receptor interactions represent computational hypotheses rather than functionally validated mechanisms. In addition, we revised Figure 6 and Figure 7A to improve the readability and interpretation of inferred signaling results, and we edited the relevant text and figure legends to make these results easier to follow. These points are addressed in greater detail in the point-by-point responses below.

      Reviewer #3 (Public review):

      Summary:

      This is an interesting paper. The process of tooth exfoliation and replacement in vertebrates remains an intriguing and fascinating subject of inquiry. As the scientists noted, there are no mammalian models that can be used to examine signaling pathways in real time.

      Strengths:

      This work integrates in vivo and high-resolution transcriptomics. The study confirms previous findings and emphasizes the need for additional research into the processes that drive the restoration of missing teeth for future therapeutic uses.

      Weaknesses:

      I disagree with the use of the phrase "plucking". Instead, the authors use tooth extraction or tooth removal, which is clinically more correct for the procedure they are doing.

      The inspiration for our ‘plucking’ experiment is work done in the hair follicle model (lines 73-74). Because cichlid teeth are so numerous, are very small, and lack dental roots, this is an accurate description of the procedure. We opt to retain the phrasing.

      The title is rather broad and appears to be more appropriate for a review than an original research work. I would advise specifying the species under research and/or the sort of damage model used in the transcriptome analysis.

      We opt to retain the title.

      It's uncertain whether the findings are exclusively based on regeneration. The presence of tooth remnants, as well as unintended harm to surrounding tissues, may have triggered repair mechanisms, thereby biasing the current data. How did the authors handle this issue? The oral cavity was under severe manipulation, increasing the inflammatory stimuli, a situation that does not take place in physiological exfoliation.

      In the revised manuscript, we have more clearly acknowledged that our plucking paradigm may induce tissue damage and repair-associated responses in addition to accelerated tooth replacement. We have strengthened the Study Limitations section to state that we do not directly document the histological progression of tissue responses across the time course or the degree of damage caused by plucking at each sampled time point. One caveat, however, is that bone remodeling and immune response is likely triggered on the ‘control’ side of the jaw also, just not to the same degree as after plucking.

      The authors indicated the use of microCT analysis; however, no such information appears in the main text. In fact, this manuscript lacks anatomical information. It is required to conduct histological examinations of the regenerated teeth at various time points.

      microCT data were included as a Supplemental Figure to demonstrate the dental formulae of our chosen species; but we did not characterize post-plucking recovery using this technique (see above summary and below point-by-point comments).

      Although the current findings confirm previously found and verified signaling pathways, the absence of functional data lends uniqueness to this work.

      In the revised manuscript, we also clarify that, while our transcriptomic analyses identify candidate cell states, pathways, and signaling interactions associated with accelerated replacement, the functional roles of these inferred pathways remain to be verified in future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Points:

      (1) Figure 1 should include representative H&E staining images comparing the left control side and the regenerated side at 7 days post-plucking. This would provide important histological context for the regeneration process and help readers better interpret the molecular findings.

      This would indeed be valuable information, but we did not carry out histology of paired control vs plucked jaws to accompany our pulse-chase and dissections for single-nucleus isolation. This comment is similar to that below about validation of what is happening on plucked vs. control jaw halves and is the first limitation we discuss in the “study limitations and future directions” section (from line 538).

      (2) Each tooth position consists of a functional tooth, a replacement tooth, and the dental (successional) lamina. On the control side, the successional lamina contains teeth at different developmental stages, analogous to the mammalian bud, cap, and bell stages. Can the snRNA-seq analysis distinguish among tooth families at different developmental stages, as well as the individual components within a single tooth unit? Clarification of this point would enhance the developmental interpretation of the dataset.

      No, our approach does not distinguish among teeth at different stages, nor among teeth in even vs odd positions that tend to be synchronized in replacement cycles. Theoretically, one could do this by dissecting individual teeth and pooling by tooth stage, but we did not.

      (3) Identifying successional lamina cells is critical, and the authors report putative SL cells within the VEE cluster. However, the stromal cells surrounding the successional lamina are also known to play important roles in tooth regeneration. Can the authors further annotate and characterize stromal cell populations in the snRNA-seq dataset? Additional analysis of these supporting cells would strengthen the conclusions regarding epithelial-mesenchymal interactions.

      We thank the reviewer for this insightful suggestion. In response, we further characterized mesenchymal subpopulations and included these new analyses in the revised manuscript (updated Figure 4 and Supplemental Figure 8). Specifically, pseudotime and CellRank analyses identified a mesenchymal subpopulation enriched for Twist1, Dnmt1, and Runx2, which we interpret as putative dental ectomesenchyme (DEM) based on the established roles of these genes in odontogenic mesenchymal development and differentiation, as well as their reported expression in mouse and human tooth single-cell transcriptomic studies. Notably, this putative dental ectomesenchymal population resides within the broader dental follicle compartment identified in our dataset (Figure 4A-C).

      To further assess supporting stromal populations, we examined the expression of established stromal marker genes, including Lum, Col6a3, Aspn, and Vegfc. These markers were broadly restricted to mesenchymal populations and showed strong enrichment overlapping the newly identified putative dental ectomesenchymal region (Supplemental Figure 8B). Consistent with these observations, differential expression analysis identified additional DEM-enriched genes that substantially overlap canonical stromal markers, including extracellular matrix-associated genes, supporting a close transcriptional relationship between the putative dental ectomesenchyme and the surrounding stromal mesenchymal compartment (Supplemental Figure 8C). Together, these findings refine the mesenchymal landscape surrounding the putative successional lamina and support the presence of a specialized stromal microenvironment associated with tooth regeneration.

      Consistent with this interpretation, our CellChat analysis identified significantly increased interactions between the dental ectomesenchyme and cycling ameloblast populations on the plucked side at Day 0 (Supplemental Figure 8D). These interactions were enriched for signaling pathways including SEMA4, EPHB, SLIT and SPP1, all of which have established roles in tissue remodeling, extracellular matrix organization, and regenerative processes. (Supplemental Figure 8E). Because Day 0 contained sufficient biological replicates and cell numbers for robust statistical comparison, we focused our interaction analyses on this time point. Collectively, these additional analyses provide a more comprehensive characterization of the stromal compartment and further support the conclusion that a specialized dental ectomesenchymal population actively participates in epithelial-mesenchymal communication during the earliest stages of tooth regeneration. So, in total, Figure 4 was revised, the text on lines 281-309 was revised, and Supplemental Figure 8 was added.

      (4) The manuscript currently lacks experimental validation of the single-nucleus RNA-seq data. The authors should validate the expression of major signature genes using RNAscope or immunostaining, ideally comparing regenerated samples with the left-side control. Such validation would significantly enhance the robustness of the conclusions.

      We did not validate up- or down-regulation of differentially expressed genes in intact tissue, owing in part to (1) the complexity of this experiment, (2) the fact that the majority of DEGs, or ‘major signature genes’ have been observed to be expressed in dentitions generally, and often by us in previous work on cichlid teeth, and the fact that (3) independent biological replicates were strongly consistent in the direction of effects (see below). In the “study limitations” section, we note this issue and suggest that a spatial transcriptomics experiment across the timespan of plucking<>recovery would address simultaneously the desire to understand cellular context of plucking and cellular/spatial differences in plucked vs control cell-type gene expression.

      Minor Points

      (1) In Figure 1, the color scheme used in the schematic drawing (Figure 1A) should match the corresponding structures shown in Figure 1B to improve clarity and consistency.

      We appreciate the reviewer’s thoughtful suggestion regarding the color consistency between the schematic (Figure 1A) and the fluorescence images (Figure 1B). However, the color scheme in the schematic (Figure 1A) was intentionally selected to maximize accessibility, particularly for readers with color vision deficiencies, and therefore differs from the magenta and green fluorescence channels used in Figure 1B. In the fluorescence images, the magenta and green colors reflect the native display colors used for the Alizarin Red and Calcein labeling channels in the pulse-chase experiment. Directly matching the schematic colors to the fluorescence images could reduce the visual contrast between key anatomical structures and compromise accessibility for some readers. We have therefore retained the current color scheme in Figure 1A while ensuring that the corresponding structures are clearly identified through consistent labels and annotations across both panels.

      (2) The abbreviation for successional lamina (SL) should be defined upon first use in the Introduction.

      We thank the reviewer for catching this omission. We have now defined the abbreviation “successional lamina (SL)” upon its first appearance in the Introduction.

      (3) Regarding biological replicates, the authors should provide data demonstrating the consistency and reproducibility across replicated samples.

      We thank the reviewer for this suggestion. To demonstrate the consistency and reproducibility across biological test subjects, we have added analyses summarizing sequencing quality metrics, test subject contributions, integrated clustering, and representative differential gene expression across individual samples (see Figure S4, panels C, D & E and Author response image 1). Panel A shows that nuclei from different biological test subjects are well integrated across clusters rather than segregating by sample origin. Finally, Panel B presents representative differentially expressed genes from multiple cell populations, demonstrating consistent expression differences between paired plucked and control samples across biological test subjects.

      Author response image 1.

      (A) UMAP embedding of dental nuclei. Each point represents a single nucleus, colored by test subject. (B) Representative differentially expressed genes show consistent expression differences between plucked and control samples across biological replicates. Paired boxplots of average gene expression for representative differentially expressed genes from multiple cell populations at Days 0, 1, 3, and 7. Each point represents one biological replicate (test subject), with paired plucked and control samples connected by dashed lines. The y-axis shows average gene expression, and the x-axis indicates the experimental condition. These representative examples illustrate the consistent direction of differential expression across biological replicates, supporting the reproducibility of the single-nucleus RNA-seq dataset.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1: Can the panels to the right of panel B be labeled? It's not clear what these six images are showing, so giving them letters and explaining briefly in the legend what the point of each panel is would clarify. "Right, example of individually classified teeth" - can the authors elaborate on what each tooth is an example of (i.e., how each tooth shown was classified"?) For clarity, the graphs in panels C and D should have the y-axes labeled

      We thank the reviewer for this helpful suggestion. In response, we revised the Figure 1B legend to clarify the classification criteria used for dye incorporation analyses and to better describe the representative fluorescence images. Specifically, teeth positive for both Alizarin and Calcein were classified as pre-existing old teeth, whereas teeth positive only for Calcein were classified as newly formed teeth. We additionally clarified that the images to the right of panel B show representative individually classified teeth, with the top row representing pre-existing old teeth and the bottom row representing newly formed teeth. We also added y-axis labels to panels C and D to improve figure clarity and readability.

      (2) Figure 2 legend: should "the cell type" instead be "the putative cell type"? Without validation for all cell types, it seems adding some sort of qualifier is in order here. Can the authors comment further on examples of validation from other studies? For example, Gareth Fraser has published numerous studies that show Pitx2 expression marking dental epithelium in different fish, yet none of these older papers are cited.

      Identification and validation of cell types make use of multiple published datasets in cichlids (for markers matched to mouse), as well as an unbiased computational approach (SAMap) that draws homology between cichlid and mouse dental cell types, based on shared global patterns of gene expression. There is perhaps a philosophical debate to be had about the validity of ‘cell types,’ generally, but our data are validated using two methods. We edited the text in lines 167-177 to clarify, including citing references to our own work (these studies include Gareth Fraser as an author, when he was a postdoc with Streelman).

      (3) Figure 6 is extremely complicated. Can any portions of rows or columns in these tables be highlighted in the figure to help the reader follow the proposed signaling interactions highlighted in the text?

      We thank the reviewer for this helpful suggestion. To improve the readability of Figure 6 and better guide readers through the dynamic signaling patterns described in the text, we revised the figure by visually highlighting the key sender-receiver interaction regions discussed in the Results. Specifically, we annotated the interactions involving mesenchymal subpopulations and alveolar bone (OST) signaling toward CYC-AMB at Days 0 and 7, mesenchymal signaling toward NK/T cells at Day 1, and epithelial cross-talk centred around ES-2 at Day 3. These visual annotations allow readers to more readily identify the signaling interactions highlighted in the text and relate them to the corresponding regions of the interaction heatmaps.

      (4) In Figure 7A, what does the black font indicate (if grey is up in control and red is up in plucked)? I'd guess not up in either, which then makes it unclear whether the sets in black are different or why they are being presented.

      We thank the reviewer for pointing out this ambiguity. In Figure 7A, blue and red labels indicate signaling pathways identified by CellChat as condition-specific, with blue representing pathways detected only in the control condition and red representing pathways detected only in the plucked condition. In contrast, pathways shown in black represent signaling pathways detected in both conditions but exhibiting significant differences in inferred communication probability between conditions. Thus, the black labels denote shared signaling pathways whose activity differs significantly between control and plucked samples, rather than pathways unique to either condition. We have revised the figure legend to clarify this distinction and improve interpretability.

      Reviewer #3 (Recommendations for the authors):

      (1) I encourage the authors to offer information on the histological differences between teeth during physiological and accelerated replacement. I'm curious if the eruption's accelerated rate has any effect on the mineralization of those teeth.

      We did not examine the histology of individual teeth, and so can’t comment on differences in mineralization.

      (2) The findings section contains multiple sentences that should be moved under material and techniques.

      We expect the reviewer is referring to paragraph lines 104-114, which was a tricky paragraph to place in the manuscript. In the end, we believe it represents important context necessary to interpret findings (which could be missed if moved to ‘methods’) and so we’ve chosen to keep this paragraph in its place.

      (3) It would be useful to include a table showing sample distribution by experimental design.

      We thank the reviewer for this suggestion. Sample distributions across experimental conditions, time points, biological test subjects, and identified cell populations are already provided in Supplementary Table 1. To improve clarity and accessibility, we have revised the table legend to more explicitly describe the experimental design and sample annotations represented in the table.

      (4) The writers did a nice job with the graphics in Figure 8; however, the schematics in Figure C are difficult to follow and are not adequately discussed anywhere. Please note that this text may be of great interest to the dentistry community, including clinicians, and that a clear and succinct explanation of the schemes at the end would be quite beneficial.

      We thank the reviewer for this helpful suggestion. We have revised the Figure 8 legend to more clearly explain Panel C as a summary schematic of inferred cell–cell communication events associated with accelerated tooth replacement after plucking. The updated legend clarifies that the pathway labels in Figure 8C summarize results directly from Figure 7A: red pathway labels indicate plucked-only signaling events, corresponding to pathways shown as full red bars in Figure 7A, while black pathway labels indicate signaling interactions detected in both plucked and control conditions but showing significant differences in interaction probability between conditions. Panel C also includes a cell-type legend at the bottom to identify the relevant cell populations.

    1. eLife Assessment

      This manuscript provides an important contribution to the field of platelet biogenesis, and the convincing evidence will advance our understanding of signal transduction driving the development of late megakaryopoiesis and platelet reactivity that results in bleeding diathesis. The paper is noteworthy for analyzing two related tyrosine phosphatases, using single or combined conditional gene knockouts at different developmental stages. Because SHP1 is a negative regulator and SHP2 is an activator, the synergistic effects found in the double knockout were surprising.

    2. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Barré et al utilize the Gp1ba-Cre transgenic mouse model to build upon previous findings in a Pf4-Cre system to investigate the effects of individual and combined Shp1 and Shp2 deletion in megakaryocytes and platelets. They report decreased megakaryocyte maturation, macrothrombocytopenia, and increased blood loss primarily in association with the Shp1/Shp2 double-knockout condition. The authors further show that this phenotype appears to be driven primarily by Shp2 and implicate dysregulation of Tpo signaling and downstream Ras/MAPK pathways, including ERK1/2. They propose that Shp1 may be functioning through a distinct pathway that has yet to be identified, opening up areas for future study.

      Strengths:

      Overall, the experiments combine in vitro, in vivo, and ex vivo approaches and appear to have been carefully designed and carried out, with multiple technical and biological replicates where relevant. The authors make a compelling argument for using the Gp1ba-Cre as opposed to the Pf4-Cre system and demonstrate both the dose- and stage-dependent effects of Shp1 and Shp2 on megakaryopoiesis and thrombopoiesis. They find that Shp1 and Shp2 are required in late-stage megakaryocyte maturation and that even low levels of expression compared to baseline are likely sufficient to yield generally normal megakaryocytes. Their findings also lead to specific future directions, such as the mechanism by which Shp1 regulates megakaryopoiesis and thrombopoiesis that is distinct from Tpo-mediated signaling. Figure 8 is particularly effective in summarizing the different models and pathways presented.

      Weaknesses:

      The effects of Shp1 and Shp2 knockouts are described as "synergistic," but it is not always clear that the effects are synergistic vs. additive, especially as the specific mechanism by which Shp1 functions in megakaryocyte development has yet to be identified. On a more minor point, although a significant part of the introduction focuses on the role of Mpl signaling in human disease, there is ultimately limited reference to Mpl (although there is of course a strong focus on Tpo) and the potential clinical implications of the findings presented here.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This manuscript provides an important contribution to the field of platelet biogenesis, and the convincing evidence will advance our understanding of signal transduction driving the development of late megakaryopoiesis and platelet reactivity that results in bleeding diathesis. The paper is noteworthy for analyzing two related, either singly or in combination, tyrosine phosphatases in this conditional, stage development gene knockout. Because SHP1 is a negative regulator and SHP2 is an activator, the synergistic effects found in the double knockout were surprising.

      We thank the reviewer for acknowledging the importance and novelty of our findings.

      Public Reviews:

      Reviewer #1 (Public review):

      Barré et al. investigated the role of Shp1 and Shp2 in megakaryocytes (MKs) and platelets by conditional knock-out of Shp1, Shp2, or both under the control of the Gp1ba promoter. Deletion of Shp1 and Shp2 in MKs and platelets was almost complete. The Shp1/Shp2 double knock-out mice displayed macrothrombocytopenia and increased bleeding, whereas the single knock-outs did not show significant defects. Platelet function was aberrant in DKOs, but not in single knock-outs, and so was ligand-induced signaling, particularly Syk phosphorylation.

      Megakaryocyte maturation was impaired in Shp1/Shp2 DKO mice. Ligand-induced signaling was impaired in Shp2 knock-out and DKO. Ex vivo formation of platelets and in vivo maturation of MKs were impaired in DKO mice. Pharmacological inhibitors of Shp1 and Shp2 had largely similar effects as observed in the single knock-outs. The authors conclude that Shp1 and Shp2 have synergistic functions in the MK/platelet lineage, and that Shp2 may be a potential therapeutic target in myeloproliferative neoplasms.

      Strengths:

      The data clearly show effects of the Shp1/Shp2 double knock-out on MKs and platelets.

      Weaknesses:

      There appears to be a discrepancy between the results with the Shp2 single knock-out and the Shp2 inhibitor: the Shp2 knock-out does not affect MKs and platelets, except Erk1/2 signaling, whereas the Shp2 inhibitors appear to affect MK function.

      This work is interesting and may have potential from a therapeutic point of view.

      Pharmacological effects do not always correlate with congenital anomalies arising for genetic defects. The Shp2 allosteric inhibitors used in our study only inhibit catalytically inactive Shp2, whereas targeted deletion of Ptpn11 results in a loss of total Shp2 expression, including catalytic and non-catalytic related functions, with developmental consequences. Further, Gp1ba-Cre+; Shp2fl/fl megakaryocytes express approximately 22% of normal Shp2 level, which likely also contributes to differences observed between pharmacological inhibition and genetic ablation of Shp2.

      We thank the reviewer for recognizing the therapeutic potential of our findings.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Barré et al. investigate the roles of the phosphatases Shp1 and Shp2 in the megakaryocyte and platelet lineage using genetic depletion in mice. By employing Gp1ba-Cre-based models, the study builds on the authors' previous work and addresses some limitations associated with earlier Pf4-Cre approaches. The authors report relatively mild alterations in megakaryocyte and platelet parameters in mice lacking either Shp1 or Shp2 alone, whereas combined deletion of both phosphatases results in macrothrombocytopenia, mild bleeding, and impaired GPVI-dependent platelet aggregation accompanied by reduced Syk phosphorylation. The functional platelet defects are linked to reduced expression of GPVI and integrin α2, while thrombocytopenia is associated with impaired megakaryocyte maturation, reduced ploidy, defective proplatelet formation, and altered TPO-dependent Ras/MAPK signaling. Similar effects on megakaryopoiesis are also observed in vitro following treatment with newly developed Shp2 inhibitors.

      Strengths and Weaknesses:

      The study addresses an important biological question and presents a substantial dataset that could contribute to a better understanding of Shp1 and Shp2 function in platelet biology. However, several aspects of data presentation and interpretation would benefit from additional clarification. In particular, while the authors conclude that single genetic deletion or pharmacological inhibition of Shp1 has a limited impact and that the major phenotypes are specific to combined Shp1/2 deletion or Shp2 inhibition, some of the data suggest more nuanced effects that may warrant further discussion.

      We thank the reviewer for raising this point. The manuscript is being revised accordingly, including highlighting the potential role of Shp1 in megakaryopoiesis and thrombopoiesis under steady-state and stressed conditions, requiring more detailed investigation.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Barré et al utilize the Gp1ba-Cre transgenic mouse model to build upon previous findings in a Pf4-Cre system to investigate the effects of individual and combined Shp1 and Shp2 deletion in megakaryocytes and platelets. They report decreased megakaryocyte maturation, macrothrombocytopenia, and increased bleeding primarily in association with the Shp1/Shp2 double-knockout condition. The authors further show that this phenotype appears to be driven primarily by Shp2 and implicate dysregulation of Mpl signaling and downstream Ras/MAPK pathways, including ERK1/2. Given the key role of these pathways in human diseases such as myeloproliferative neoplasms and the challenges associated with modulating such a central pathway, identification of a specific regulator of Mpl signaling poses intriguing questions for future studies on clinical applicability.

      We thank the reviewer for acknowledging the importance and novelty of our findings.

      Strengths:

      Overall, the experiments combine in vitro, in vivo, and ex vivo approaches and appear to have been carefully designed and carried out, with multiple technical and biological replicates where relevant. The authors make a compelling argument for using the Gp1baCre as opposed to the Pf4-Cre system and demonstrate both the dose- and stagedependent effects of Shp1 and Shp2 on megakaryopoiesis and thrombopoiesis. They find that Shp1 and Shp2 are required in late-stage megakaryocyte maturation and that even low levels of expression compared to baseline are likely sufficient to yield generally normal megakaryocytes. Their findings also lead to specific future directions, such as the mechanism by which Shp1 regulates megakaryopoiesis and thrombopoiesis that is distinct from TPO-mediated signaling.

      Weaknesses:

      While the experiments have been thoughtfully designed and carried out, there is limited background explanation on relatively complex or niche pathways/mechanisms, such as the relationship between P-selectin, CRP, and PAR4p; the interactions between SFK, Syk, GPVI, and CLEC-2; and TPO, MPL, ERK1/2, AKT, and STAT3, which, while likely intuitive to experts in their respective fields, may be less obvious to a reader approaching this manuscript with a global interest in megakaryopoiesis/thrombopoiesis and thus detract from the impact of the findings.

      We thank the reviewer for raising this point. The manuscript is being revised to better explain the rationale and molecular mechanisms linking these pathways and functions.

      With regard to the science itself, some of the conclusions feel premature based on the available data.

      (1) The section "Aberrant ITAM signaling in Shp1- and Shp2-deficient platelets" is challenging to follow for those not well-versed in ITAM signaling and associated pathways, and may take additional outside reading to follow the conclusion that Syk-dependent signaling is modulated downstream of GPVI and CLEC-2 based on lack of change in Src p-Tyr418, especially considering that Src p-Tyr418 was previously introduced as a measure of SFK rather than Syk. In the introduction, Shp1 is specifically mentioned as a negative regulator of the ITAM/Syk/phospholipase pathway. However, in Figure 4Ai and Bi, Syk phosphorylation/activation in Shp1 knockout cells did not appear to be different from Shp2 knockout cells, and is lower than the control, which is surprising for a negative regulator. It is also not clear why, in the section (Figure 4A-B), there is reduced Syk activation in Shp1 and Shp2 single knockout cells upon CLEC2 stimulation (but apparently not with CRP) when there was no difference in response to CLEC2 (but a difference in response to CRP) in the previous section (Figure 3A, C).

      We thank the reviewer for raising these important points. The manuscript is being revised accordingly, including clarifying the roles of SFKs, Shp1 and Shp2 in the ITAM-Syk-PLCγ2 signaling pathway.

      Briefly, SFKs are essential for phosphorylating ITAMs, allowing SH2-dependent docking of Syk. Reduced reactivity of Shp1/2 DKO platelets to CRP and collagen is likely due to downregulation of the ITAM-containing GPVI-FcR γ-chain complex and integrin α2 subunit, and concomitant reduction in Syk phosphorylation.

      However, the marginal albeit significant reduction in Syk phosphorylation downstream of CLEC-2 in Shp1 and Shp2 KO platelets was not determined and was insufficient to impact CLEC-2-mediated platelet aggregation under the conditions tested.

      Differences in the stoichiometry and docking of Syk to phosphorylated GPVI-FcR γ-chain and CLEC-2 likely contribute to the differences in platelet reactivity and Syk phosphorylation downstream of the two receptors in the absence of Shp1 and Shp2.

      (2) In the section "Reduced Tpo signaling in Shp1/2-deficient MKs," only Western blot data for (p)ERK1/2, AKT, and STAT3 are presented before concluding that decreased ERK1/2 activity is a mechanistic explanation for thrombocytopenia seen in the Shp1/2 doubleknockout condition. Such a statement would benefit from additional experiments, such as protein or transcriptional levels of ERK1/2 targets specifically relevant to megakaryopoiesis, such as ETS, FOS, and JUN, to assess the consequences of decreased phosphorylated ERK1/2.

      We thank the reviewers for these constructive comments. Further experiments are being planned to determine the biological and transcriptional consequences of reduced ERK1/2 phosphorylation during megakaryopoiesis and thrombopoiesis.

      (3) Suggesting that "inhibiting Shp2 will not have any bleeding consequence in patients" and that Shp2 may be a therapeutic target in myeloproliferative neoplasms when none of these studies have been carried out in a human model is a bold conclusion. There are no data presented on, for example, whether Shp2 inhibition can help reverse the MPL/JAK/STAT pathway in the setting of gain-of-function mutations specifically associated with myeloproliferative neoplasms.

      This conclusion is being tempered in the revised manuscript. Genetic- and pharmacological-based approaches will be used to establish the therapeutic potential of inhibiting Shp1 and Shp2 in mouse models of MPN, including Jak2 gain-of-function mice. Bleeding and thrombotic complications of inhibiting Shp1 and Shp2 will be explored as part of these studies.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Altogether, we feel that this is an important study for those in the fields of hematology or signal transduction. Your important study characterizes the roles in late megakaryopoiesis and platelet biogenesis of single or combined conditional deletion of two tyrosine phosphatases, Shp1 and Shp2. Strengths include technical advances in single and combined deletions, the somewhat surprising results of synergy between the two phosphatases, focusing on the critical stage of late megakaryopoiesis, and clinical implications in bleeding diathesis.

      Weaknesses are mostly minor, but the numerous points raised by reviewer 3 need to be addressed and typographical errors corrected. Further discussion should include the relevance or dissimilarity in megakaryopoiesis and platelet biogenesis between murine and human blood health and disease. Since SHP1 is a negative regulator and SHP2 is a positive activator, additional discussion about how they coordinate and fine-tune ("nuanced") signal transduction in TPO- or GPVI-induced signaling in an explicitly stated pathway.

      We invite you to respond to the critiques and submit a revised manuscript.

      Sincerely,

      Seth Corey, MD MPH

      We thank the editor for the positive evaluation of our study and for highlighting its relevance to the fields of haematology and signal transduction. We have carefully addressed all comments raised by Reviewer 3 and corrected typographical errors throughout the manuscript.

      As suggested, we expanded the Discussion to better address the relevance of our murine findings to human megakaryopoiesis and platelet biogenesis. While our study relies on mouse models, key components of TPO/MPL signaling and platelet production are conserved between mice and humans, although differences in megakaryocyte maturation dynamics and platelet biology are acknowledged and now discussed.

      We also clarified the coordinated roles of Shp1 and Shp2 in signaling. Although Shp1 generally acts as a negative regulator and Shp2 as a positive mediator of signal transduction, our results suggest that they function in a complementary manner to optimize signaling downstream of TPO/MPL and GPVI pathways, thereby ensuring appropriate regulation of late megakaryopoiesis, platelet production and activation.

      These additional considerations have been incorporated into the revised manuscript to provide a clearer conceptual framework for how Shp1 and Shp2 cooperate to regulate platelet biogenesis.

      Reviewer #1 (Recommendations for the authors):

      (1) The effects of the Shp1/Shp2 DKO are clear, but the effect of the Shp2 single knock-out is less clear on all parameters that were tested. The exception is ERK1/2 phosphorylation, which was reduced in the Shp2 knock-out as well as the Shp1/Shp2 DKO. Why do the authors conclude that Shp2 may be a potential therapeutic target, while the data show that knock-out of Shp1 and Shp2 is required for the observed effects?

      We agree that the most pronounced phenotypes were observed in the Shp1/Shp2 DKO. However, Shp2 single knock-out consistently reduced ERK1/2 phosphorylation, indicating that Shp2 contributes to MPL downstream signaling in megakaryocytes. The absence of a strong phenotype in Shp2 single knock-out may be due to residual Shp2 protein. However, given the established role of the Shp2–ERK pathway in megakaryopoiesis and the observation that pharmacological Shp2 inhibition significantly affected MK ploidy, proplatelet formation, and ERK1/2 phosphorylation, our data support a contribution of Shp2 to these processes and suggest it as a potential therapeutic target.

      (2) Inhibitors of Shp1 and Shp2 had largely similar effects as Shp1 and Shp2 single knock-outs, respectively. The effect of Shp2 knock-out on MK ploidy is not clear, cf. Figure 5Ai (no effect) and Figure 5Aii (reduction, which is not significant), whereas a clear and significant effect was reported for the Shp1/Shp2 DKO. In contrast, in Figure 7Ciii, the Shp2 inhibitors SHP099 and RMC-4550 clearly affect MK ploidy and the percentage of MKs forming proplatelets. The discrepancy between the effect of Shp2 knock-out and Shp2 inhibitors suggests that the inhibitors may affect other targets. The authors should consider using the Shp2 inhibitors on the Shp2 knock-out to prove or disprove that the effects of the Shp2 inhibitors are mediated exclusively by Shp2.

      Pharmacological inhibition does not necessarily phenocopy genetic deletion. The allosteric Shp2 inhibitors used in our study (SHP099 and RMC-4550) stabilize Shp2 in an inactive conformation and inhibit its catalytic activity, whereas Ptpn11 deletion results in complete loss of the Shp2 protein, including both catalytic and scaffolding functions. These mechanistic differences may lead to distinct biological outcomes and could explain the discrepancy observed between Shp2 knockout and inhibitor treatments.

      (3) Since the most profound effects were found in the Shp1/Shp2 DKO, it would be interesting to use combinations of the Shp1 and Shp2 pharmacological inhibitors to mimic the effect of the Shp1/Shp2 DKO.

      We thank the reviewers for these constructive comments. Further experiments are indeed being planned to use combinations of the Shp1 and Shp2 pharmacological inhibitors to mimic the effect of the Shp1/2 DKO.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Additional details on the strategy used to isolate megakaryocyte progenitors from mouse bone marrow would improve clarity, including sorting approach, gating strategy, and assessment of population purity.

      We thank the reviewer for this suggestion. We have now expanded the Methods section to provide a more detailed description of the strategy used to isolate megakaryocyte progenitors from mouse bone marrow.

      Briefly, bone marrow cells were first enriched for hematopoietic progenitors and stained with antibodies against lineage markers and megakaryocyte-associated markers. Megakaryocyte progenitors were then isolated by flow cytometric sorting based on established surface marker combinations, including c-Kit and CD41 expression. The gating strategy excluded lineage-positive cells and debris before selecting the progenitor population of interest.

      (2) Platelet GPVI expression appears reduced not only in Shp1/2 double-knockout mice but also, to some extent, in single Shp1- or Shp2-deficient models. A more detailed quantitative comparison and discussion would be helpful.

      We thank the reviewer for this observation. Although the most pronounced reduction in GPVI surface expression was observed in Shp1/Shp2 double knock-out platelets, minor variations may appear in the single knock-out models. To address this, we performed additional statistical analyses comparing WT platelets with each single knock-out genotype. These analyses did not reveal any significant statistical differences in GPVI expression between WT and either Shp1- or Shp2-deficient platelets, indicating that the apparent variations fall within the range of biological variability.

      (3) The aggregation traces shown in Figures 3A and 3B would benefit from clarification regarding their representativeness relative to the corresponding quantitative analyses.

      We thank the reviewer for this comment. The aggregation traces in Figures 3A and 3B represent experiments selected from independent replicates included in the quantitative analysis. The figure legends have been revised to clarify that these traces are representative of the experiments summarized in the quantification panels, which include data from multiple independent mice.

      (4) In several experiments, statistical significance may be influenced by differences in sample size across genotypes (e.g., Figures 2Ci, 3Ai, 3Di, and 6Ai). Using comparable numbers of replicates would strengthen the interpretation.

      We appreciate the reviewer’s attention to statistical rigour. The differences in sample size between genotypes reflect the availability of animals from the different breeding cohorts. Importantly, all statistical analyses were performed using appropriate tests that account for unequal sample sizes. The observed differences remain consistent across independent experiments.

      (5) The rationale for assessing only P-selectin exposure following CRP and PAR4p stimulation is not fully explained. Including integrin αIIbβ3 activation, or clarifying its exclusion, would provide a more complete assessment of platelet activation.

      We thank the reviewer for this suggestion. P-selectin exposure was used as a primary readout because it provides a robust measure of α-granule secretion downstream of GPVI and PAR signaling. Integrin αIIbβ3 activation was not assessed in these experiments because platelet aggregation assays were performed in parallel, which already provide a functional readout of integrin activation, as aggregation requires αIIbβ3 engagement. Nonetheless, we agree with the reviewer that direct measurement of integrin activation (e.g., fibrinogen binding) would provide complementary information and will be considered in future studies.

      (6) Figure 3Dii is described as an aggregation assay, although it appears to report P-selectin exposure; this distinction should be clarified.

      We thank the reviewer for identifying this inconsistency. Figure 3Dii reports indeed P-selectin exposure measured by flow cytometry, rather than platelet aggregation. We have corrected the description in the Results section.

      (7) The suggestion of compensatory extramedullary hematopoiesis based on splenomegaly would be strengthened by immunophenotypic analysis of splenic hematopoietic progenitor populations.

      We appreciate this important suggestion. In the current study, the evidence for possible compensatory extramedullary hematopoiesis is mainly based on the splenomegaly observed in Shp1/2 DKO mice. We agree that detailed immunophenotypic analysis of splenic hematopoietic progenitors would provide additional mechanistic insight; however, this was beyond the scope of the present study, which focuses on the intrinsic role of Shp1 and Shp2 in the megakaryocyte and platelet lineage. We have therefore revised the Discussion to present this interpretation more cautiously and to indicate that further studies will be required to determine whether splenic hematopoiesis contributes to compensatory platelet production in this model.

      (8) In Figure S3, differences in platelet recovery kinetics among genotypes appear evident. Clarification of the statistical tests used to assess these differences would be useful.

      We thank the reviewer for this comment. Platelet recovery kinetics were analyzed using two-way ANOVA with appropriate post hoc tests. No statistically significant differences between genotypes were observed. These details have been added to the Methods and figure legend for clarity.

      Reviewer #3 (Recommendations for the authors):

      Overall, the manuscript suffers from multiple typographical and grammatical errors that distract from the data being presented.

      We have carefully revised the manuscript to correct typographical and grammatical errors throughout, improving clarity and readability.

      (1) Figure S1: I believe this should be referenced in the first paragraph of the results section.

      We have now referenced the Supplemental Figure S1 in the first paragraph of the results section as suggested.

      (2) Figure 2A: Although the individual points for the replicates are informative, they do make it difficult to appreciate the SEM, and to my eye it appears that, for example, there may not be a difference between Shp2 and Shp1/2 or that there may be a difference between Shp1 and Shp1/2 in (ii), as Table S2 suggests. In other words, it seems that the increased MPV (as well as the leukocyte phenotype) may be driven by the knockout of Shp2; are there statistical analyses that could be performed to show that the increased MPV is specific to the double knockout?

      We thank the reviewer for this comment. Despite the slightly higher MPV observed in Shp2 single knockouts, statistical analysis using one-way ANOVA, which is appropriate for comparing means across multiple independent groups, and taking all individual data points into account, revealed no significant differences between Shp2 or Shp1 single KO and the Shp1/2 DKO.

      (3) Figure 2Bi: Is this missing a statistical significance bar, or was there no significant difference in cumulative bleeding time between the conditions? If the latter, this should be clarified in the main text (although the specific sentence regarding bleeding time only claims "mildly prolonged," the preceding sentence indicates "significant increase in bleeding").

      Thank you for this comment. There was no statistically significant difference in cumulative bleeding time between the groups. We have now modified the text accordingly to clarify this point and to indicate that, while bleeding time was not significantly different, blood loss was significantly increased in Shp1/2 DKO mice.

      (4) Figure 2Ci: What was the extent (statistically) of GPVI reduction in the Shp1 and Shp2 single knockout mice compared to the control? It seems that although there was no change in alpha2 expression in the single-knockout conditions, the contributions of Shp1 and Shp2 loss may be additive on GPVI (although I acknowledge that this is not necessarily borne out in Figure 3Ai).

      Thank you for this comment. After reanalyzing the data using an appropriate statistical test (one-way ANOVA followed by Tukey’s post hoc test), we found that GPVI expression is significantly reduced in both Shp1 and Shp2 single knockout platelets compared with controls. However, this reduction did not result in detectable functional consequences on platelet aggregation, as shown in Figure 3Ai.

      (5) Figure 3Ai: It seems that the individual replicates for the Shp1/2 double knockout cluster in two populations, extreme non-responders and arguably normal responders to CRP. Are there any biological or technical explanations for this?

      We thank the reviewer for this observation. We agree that the distribution of individual replicates in the Shp1/2 DKO group suggests the presence of two subpopulations, with some samples showing markedly impaired aggregation and others retaining near-normal responsiveness to CRP. While all experiments were performed under standardized conditions, subtle differences in platelet preparation, agonist sensitivity, or assay timing could also contribute to dispersion within this group. Importantly, despite this variability, the overall trend indicates a significant reduction in aggregation in the Shp1/2 DKO condition compared to controls, supporting a critical and partially redundant role for Shp1 and Shp2 in GPVI-mediated platelet activation.

      (6) "Aberrant functional responses of Shp1/2-deficient platelets": It may be helpful, in the last paragraph of this section, to briefly explain the relationship between P-selectin, CRP, and PAR4p. If short on space/words, the introduction likely does not need an explanation of platelet function and definitions of megakaryopoiesis and thrombopoiesis.

      We thank the reviewer for this suggestion. We have revised the last paragraph to clarify that P-selectin surface expression reflects α-granule secretion following platelet activation. We now specify that CRP activates platelets via GPVI signaling, whereas PAR-4 peptide signals through thrombin receptors, providing context for the differential responses observed in Shp1/2-deficient platelets.

      (7) "Aberrant ITAM signaling in Shp1- and Shp2-deficient platelets": Is there a cartoon figure panel that could be added to clarify how SFK (which, as an aside, is not defined as an acronym), Syk, GPVI, CLEC-2 receptor, Shp1, and Shp2 are interrelated? In addition to the comments left in the public review, I was perplexed by Figure 4Bi, as the band for the Shp1/2 double knockout condition appears to be stronger than the other 3 conditions, but this is not what is depicted in the bar graph on the right.

      We thank the reviewer for this helpful comment. We have now added a schematic cartoon (new Figure 8) to clarify the relationships between SFKs, Syk, GPVI, and the regulatory roles of Shp1 and Shp2. All acronyms, including SFK, are now defined at first mention to improve accessibility.

      Regarding Figure 4Bi, we appreciate this observation. The apparent discrepancy between the representative blot and the quantification reflects variability across experiments. The bar graph represents the average of independent replicates.

      (8) I would also recommend considering reshuffling the panels in Figure 4 so that the 2 assays measuring Syk phosphorylation and the 2 assays measuring Src phosphorylation are next to each other, as opposed to grouped by agonist. They should also be presented in the order of the text, which states that SFK activation was measured via Src before mentioning Syk (but the data are presented in reverse).

      We thank the reviewer for this suggestion. We have reorganized Figure 4 so that the panels measuring Src and Syk phosphorylation are presented together and, in the order, described in the text. The manuscript text has also been updated accordingly to match the revised figure layout.

      (9) GPVI overexpression experiments in these megakaryocytes or, conversely, Syk inhibition in control cells, to reverse or recapitulate the phenotype, respectively, may be additionally informative.

      We thank the reviewer for this suggestion. We agree that modulating GPVI or Syk activity could provide additional mechanistic insight. While these experiments were beyond the scope of the current study, we plan to explore GPVI overexpression and Syk inhibition in follow-up studies to further validate the pathway’s role in the observed phenotype.

      (10) "Reduced Tpo signaling in Shp1/2-deficient MKs": In addition to the comments left in the public review, I would suggest moving this section to after "Defective proplatelet formation and MK maturation in Shp1/2-deficient mice" so that the 2 sets of proplatelet and ploidy data are consecutively presented.

      We thank the reviewer for this helpful suggestion. We have now revised the manuscript accordingly by reorganizing both the text and figures. The ploidy and proplatelet formation data are now presented together in Figure 5, followed by the Tpo signaling data in Figure 6, improving the overall flow and clarity of the results section.

      (11) Figure 6Cii: Why does Shp1 add up to >100%?

      The reason the Shp1 bar exceeds 100% is due to how the data were quantified and normalized. Each segment represents the mean from separate experiments. Stacking these means can exceed 100% because the sum of averages is not equal to the average of the total.

      (12) Figure 7D: How do you reconcile these findings of impaired AKT phosphorylation with the addition of a Shp2 inhibitor but no change with Shp2 knockout (Figure 5C)? Would you attribute it to the residual Shp1 and Shp2 in the Cre-Lox MKs?

      Pharmacological effects do not always correlate with congenital anomalies arising for genetic defects. The Shp2 allosteric inhibitors used in our study only inhibit catalytically inactive Shp2, whereas targeted deletion of Ptpn11 results in a loss of total Shp2 expression, including catalytic and non-catalytic related functions, with developmental consequences. Further, Gp1ba-Cre+; Shp2fl/fl megakaryocytes express approximately 22% of normal Shp2 level, which likely also contributes to differences observed between pharmacological inhibition and genetic ablation of Shp2.

    1. eLife Assessment

      This important study demonstrates how ablation or silencing of hilar mossy cells in the mouse influences the primary location where the mossy cells project, the inner molecular layer of the dentate gyrus. The anatomical findings are convincing and include altered adult-born granule cells and the shrinkage of the inner molecular layer following mossy cell ablation. However, the mechanisms and their functional significance are unclear, so more of these types of experiments/analyses would strengthen the study, especially the support for the broader conclusions.