10,000 Matching Annotations
  1. Jul 2026
    1. eLife Assessment

      In this important study, the authors have performed a zebrafish drug screen to identify suppressors of atherogenic lipoproteins. They utilize a well-established LipoGlo assay to find molecules that modulate these lipoproteins, identifying 49 potential hits and they perform validation experiments, including studies linking enoxolone to its likely inhibitory effect on a specific transcription factor, HNF4alpha. Overall, the results are convincing and robust, and will open up new areas of exploration for those investigators interested in in vivo lipid biology.

    2. Reviewer #1 (Public review):

      [Editors' note: The authors addressed reviewer comments well, further strengthening the conclusions of the study.]

      Summary:

      A whole-organism drug screen was performed to identify molecules that decrease Apolipoprotein B (ApoB) as a target for agents to reduce atherosclerosis. Kelpsch et al. used a zebrafish reporter line, LipoGlo, which is a fusion of the Nano-luciferase protein to the ApoB protein as a proxy for the presence of ApoB-containing lipoproteins (B-lps) in larval stages. The LipoGlo line was screened against a well-characterized drug library and identified 49 hits from their primary screen. Follow-up studies further refined this list to 19 molecules that reproducibly reduced B-lps significantly. The authors focused their studies on enoxolone, a licorice root extract, and showed that larvae treated with this agent can reduce the production of B-lps. As enoxolone has been reported to suppress Hepatocyte Nuclear factor 4a (HNF4a), the authors investigated whether loss-of-hnf4a or pharmacological inhibition of hnf4a in zebrafish also produced similar phenotypes as enoxolone treatment. Their studies showed that this was the case. Transcriptomic studies after enoxolone treatment resulted in altered expression of genes involved in cholesterol biosynthesis and in glucose/insulin signaling pathways. This study highlights the utility of a zebrafish whole-organism chemical screen for modifiers of B-lps production and/or its clearance. A significant finding is that enoxolone inhibits hnf4a in zebrafish to reduce B-lps production and supports targeting HNF4a as a therapeutic means to reduce the emergence of atherosclerosis.

      Strengths:

      The authors performed a whole-organism chemical screen with over 3000 agents. Such screens are challenging, and the authors used strict criteria for determining hits. The conclusions of this study are well supported by the presented data.

      Comment on revised version:

      The authors have addressed all my comments.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aimed to develop a large-scale drug screen to identify B-lp modulators in a vertebrate whole-animal system. Using the zebrafish LipoGlo system that the authors had previously published and validated, the authors screened 2762 drug candidates to generate 49 hits and ultimately validated 19 drugs as genuine ApoB-lowering drugs. Using LipoGlo-Electrophoresis, the authors are able to obtain insights into the ApoB-lipoprotein size/subclass distribution. The authors further validate and study the mechanism of a strong hit, Enoxolone, known as also known as 18β-Glycyrrhetinic acid, which has previously been reported to modulate lipid metabolism. The authors also show that Enoxolone effects are mediated through HNF4⍺, which has been previously shown in the mouse system, but this is the first time it has been shown in the zebrafish.

      Strengths:

      The study was methodical and robust, using a published and well-validated zebrafish LipoGlo model. The authors validated the hits from the screen independently and considered the possibility that some drugs may have been detected as false positive results due to effects on the enzymatic activity of NanoLuciferase; only one hit, verteporfin, was shown to be a false positive. Using LipoGlo-Electrophoresis, the authors are able to obtain extra insights into the ApoB-lipoprotein size/subclass distribution. They showed that while enoxolone treatment reduces total B-lps, there are no overt changes in B-lp size distribution compared to vehicle-treated animals, other than a slight increase in the zero mobility (ZM) fraction, which contains very large particles and/or tissue aggregates. In contrast, the positive control, lomitapide, does show a change in B-lp size distribution compared to vehicle-treated animals - an increase in frequency of LDLs (low-density lipoprotein), but a decrease in VLDLs (very low-density lipoprotein). This study also assesses the LipoGlo-Electrophoresis profile of HNF4⍺ inhibitors. Work in the zebrafish larvae means that the effect on overall development and an entire vertebrate organism can also be assessed. Finally, the authors applied a thorough statistical measure to define a hit, using the Strictly Standardized Mean Difference (SSMD) method.

    4. Reviewer #3 (Public review):

      Summary:

      In "A‬‭ whole-animal‬‭ phenotypic‬‭ drug‬‭ screen‬‭ identifies‬‭ suppressors‬‭ of‬‭ atherogenic‬ lipoproteins", Kelpsch et al seek to identify new, chemically targetable pathways that regulate ApoB function and could ultimately serve as treatments for elevated lipid disorders and/or cardiovascular disease. Given the interconnected nature of lipid regulation in the whole organism with interdependent organs and secreted components (i.e. lipoproteins), they use the vertebrate model zebrafish to screen a large library of ~3000 compounds for their ability to lower the important ApoB-containing lipoproteins. They find 49 hits with 19 compounds passing a higher level of scrutiny, and focus on the role of enoxolone in modulating B-Ip levels at least partly through the HNF4alpha transcription factor and, putatively, through downstream cholesterol/lipid biosynthetic pathways.

      Strengths:

      The study uses a well-validated in vivo stain (LipoGlo) for measuring lipoproteins in the context of a developing whole organism with a quantitative read-out on a high-throughput platform, allowing for screening of thousands of compounds altering the complex metabolic/physiologic functions necessary for lipoprotein production.

      The use of genetic mutant HNF4alpha to assign the mechanism of action to the prime candidate compound studied (enoxolone) is a powerful approach for this challenging aspect of chemical genetics studies.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this important study, the authors have performed a zebrafish drug screen to identify suppressors of atherogenic lipoproteins. They utilize a well-established LipoGlo assay to find molecules that modulate these lipoproteins, identifying 49 potential hits. They perform some validation experiments, including studies linking enoxolone to its likely inhibitory effect on a specific transcription factor, HNF4alpha. Overall, the results are convincing and robust, and will open up new areas of exploration for those investigators interested in in vivo lipid biology.

      We appreciate the manuscript assessment and provide new data (see below) that significantly increases the “strength of the evidence”.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The authors performed a whole-organism chemical screen with over 3000 agents. Such screens are challenging, and the authors used strict criteria for determining hits. The conclusions of this study are well supported by the presented data.

      Weaknesses:

      There are areas within the study and writing that can be improved and extended, specifically within the gene expression studies.

      We appreciate the Reviewer recognizing the strength of the data and the challenging nature of performing a whole animal small molecule screen. With regard the “Weaknesses”, as the Reviewer suggested we improved and clarified the text and “extended” the study by adding analysis and discussion of six additional hits (see Figure 3 and Sup. Fig.1) and modified the results section with the addition of over a page of text describing the additional phenotyping results:

      “Secondary characterization of selected validated B-lp lowering hits.

      To further evaluate the biological relevance and potential mechanisms of selected validated hits, we performed secondary analyses assessing total B-lp levels, larval morphology, and lipoprotein size distribution. Our lab previously defined the method to measure total B-lp levels from whole animals by using the homogenate of a single zebrafish larva [35]. We confirmed that treatment of animals with 4 µM pomiferin significantly reduced total B-lp levels measured from homogenates collected from whole animals after treatment (p = 8.6x10<sup>-4</sup>; Figure 3A). However, pomiferin treatment (4 µM) produced animals with reduced body length and lethality at higher doses, suggesting developmental toxicity may confound interpretation of its B-lp-lowering effect.

      Treatment of animals with riboflavin tetrabutyrate (Figure 3B) and calcipotriene (Figure 3C) reduced total B-lp levels (p < 2x10<sup>-16</sup>) but did not affect larval morphology. A key feature of B-lps is their size, often a proxy for the total amount of lipid in the particle [35]. Particle size can impact the particle's lifetime (e.g. in metabolically healthy humans, small particles are cleared rapidly by the liver) [54–56]. Thus, we also assessed whether these compounds alter B-lp size distribution. Animals were treated for 48 h with vehicle, 5 µM lomitapide, or a drug of interest, and whole-animal homogenates were prepared and subjected to native polyacrylamide gel electrophoresis followed by chemiluminescent imaging. B-lps were classified into four classes based on gel migration: zero mobility (ZM), very low-density lipoproteins (VLDL), intermediate-density lipoproteins (IDL), and low-density lipoproteins (LDL) as previously described [35]. Lomitapide treatment effectively reduces VLDL particles and increases LDL particles [35] (Figure 3E), whereas riboflavin and calcipotriene did not affect lipoprotein classes. Thus, riboflavin tetrabutyrate and calcipotriene reduce total B-lp levels without overt developmental toxicity or changes in lipoprotein subclass distribution, suggesting they may act through mechanisms that decrease overall particle abundance rather than altering lipoprotein turnover or catabolism.

      Although doxycycline treatment lowered B-lp levels in whole fixed animals in the primary screen and validation studies, we did not observe a reduction in total B-lps in whole-animal homogenates (Figure 3D). However, we detected a slight increase in VLDL levels (p < 2x10<sup>-16</sup>; Figure 3E), suggesting that doxycycline may alter lipoprotein composition or distribution rather than total particle abundance.

      Alternatively, two structurally related compounds, thiethylperazine and prochlorperazine, at 4 µM significantly reduced (p < 1.4x10<sup>-10</sup> and p < 1.2x10<sup>-6</sup> respectively), B-lp levels measured from whole-animal homogenates (Figure 3F and 3H). Furthermore, both 8 µM thiethylperazine and 8 µM prochlorperazine increased relative LDL (p = 1.2x10<sup>-3</sup> and p = 2.3x10<sup>-3</sup>, respectively) and decreased relative VLDL levels (p = 8.4x10<sup>-4</sup> and p = 9.1x10<sup>-4</sup>, respectively; Figure 3G and 3I) suggesting a shift toward smaller lipoprotein particles and a potential alteration in lipid processing or clearance pathways. Together, these results highlight the diversity of mechanisms among validated hits, ranging from compounds that reduce total B-lp abundance without affecting B-lp class composition to those that shift B-lp class distribution, while also underscoring the importance of secondary assays to distinguish true B-lp modulators from those that likely produce a B-lp effect through generalized toxicity.

      Enoxolone significantly reduces B-lps in the larval zebrafish.

      Hit compounds were prioritized for follow-up studies based on reproducible dose-dependent responses, minimal toxicity as indicated by normal morphology over development, lack of direct NanoLuciferase inhibition, and the presence of literature suggesting potential links to lipid metabolism. One compound meeting these criteria was enoxolone, also known as 18β-Glycyrrhetinic acid, (Figure 2 Drug 20, Supplemental Table 1, Supplemental Figure 1T, Supplemental Figure 2L).”

      The Discussion now has the following additional text:

      “Further validation of these hits demonstrated a wide range of potential mechanisms of lipoprotein regulation. We identified hits that affected larval development, some hits that reduced total B-lp levels, and several structurally related compounds that directly reduced B-lp particle size (Figure 3).”

      Reviewer #2 (Public review):

      Strengths:

      The study was methodical and robust, using a published and well-validated zebrafish LipoGlo model. The authors validated the hits from the screen independently and considered the possibility that some drugs may have been detected as false positive results due to effects on the enzymatic activity of NanoLuciferase; only one hit, verteporfin, was shown to be a false positive. Using LipoGlo-Electrophoresis, the authors are able to obtain extra insights into the ApoB-lipoprotein size/subclass distribution. They showed that while enoxolone treatment reduces total B-lps, there are no overt changes in B-lp size distribution compared to vehicle-treated animals, other than a slight increase in the zero mobility (ZM) fraction, which contains very large particles and/or tissue aggregates. In contrast, the positive control, lomitapide, does show a change in B-lp size distribution compared to vehicle-treated animals - an increase in frequency of LDLs (low-density lipoprotein), but a decrease in VLDLs (very low density lipoprotein). This study also assesses the LipoGlo-Electrophoresis profile of HNF4⍺ inhibitors. Work in the zebrafish larvae means that the effect on overall development and an entire vertebrate organism can also be assessed. Finally, the authors applied a thorough statistical measure to define a hit, using the Strictly Standardized Mean Difference (SSMD) method.

      We appreciate that the Reviewer valued the rigour and robustness of our approach.

      Weaknesses:

      While the screen was thorough and well-validated, the authors missed a chance to provide a lot of extra significance to a wide range of readership. While the hits were thoroughly validated and displayed, the authors could have also presented the LipoGlo-Electrophoresis for all validated hits or at least a number of them. This would hugely increase the insights into these compounds. Also, the authors chose to validate and follow up a mechanism for Enoxolone, yet this hit was already known to modulate lipid metabolism through HNF4⍺, therefore, hugely limiting the impact of the paper. So what the authors have shown that is novel is only subtly added to this - consistent in vertebrate models, RNA sequencing of pathways, further validation of the HNF4⍺ pathway, and a profile of resulting B-lp size distribution. It seemed an easy way out to pick such a candidate, and they could have followed up by validating more thoroughly a completely novel drug. Also, the authors' prior paper showing the methodology also depicted complementary EM and LipoGlo-microscopy approaches. The microscopy especially, would have been an easy complementary add-on to the screen to really give extra insights into B-lp metabolism in a whole organism for all candidates. This felt like a missed opportunity.

      Here we agree and added Fig. 3 describing the phenotyping of 6 additional compounds including some LipoGlo-Electrophoresis analyses as suggested by the Reviewer. The text of the Results section was modified as described for Reviewer 1 (see above).

      Reviewer #3 (Public review):

      Strengths:

      The study uses a well-validated in vivo stain (LipoGlo) for measuring lipoproteins in the context of a developing whole organism with a quantitative read-out on a high-throughput platform, allowing for screening of thousands of compounds altering the complex metabolic/physiologic functions necessary for lipoprotein production.

      The use of genetic mutant HNF4alpha to assign the mechanism of action to the prime candidate compound studied (enoxolone) is a powerful approach for this challenging aspect of chemical genetics studies.

      We appreciate that the Reviewer understands how challenging it can be to assign a mechanism to any small molecule and thereby recognizes the power of the zebrafish model combined with our unique lipoprotein phenotyping tools.

      Weaknesses:

      As shown in Figure 5A, the HNF4alpha mutant homozygous -/- already lowers lipoproteins. Is it just that the mutant level is already at a minimum in this homozygous mutant (and thus enoxolone cannot induce even lower lipoprotein levels), or is it true that the enoxolone molecule is primarily acting through this TF (i.e. HNF4alpha homozygous mutant is truly epistatic to enoxolone function) as favored in the text.

      While it is definitely interesting to study enoxolone effects during whole embryo development, the link to HNF4alpha had previously been described in the literature, as pointed out by the authors. The generalizability of the approach to identify truly novel pathways remains to be fully realized, but sharing this available screen data to date will invite further inquiry and be very valuable to the community.

      Here too we agree that a link between enoxolone was proposed in the literature. However, we added quite a lot of additional insight regarding the transcriptional targets shared by HNF4alpha and enoxolone. The goal of identifying the mechanism(s) of action of other novel small molecule hits from the screen is important and that work is ongoing.

      Figure 5 - The same allele of HNF4alpha loss of function/hypomorph (rdu14) is used in both 5A and 5B, but labeled differently in each subpanel. This is explained in the figure legend, but could be updated to use the same nomenclature in both panels to clarify the Figure presentation.

      We thank the Reviewer for catching this and have modified the Figure (now Fig. 6) so the subpanels are labeled identically to avoid any confusion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors describe statistical methods to improve the calling of hits. In the results section, they discuss the use of fold change and of strictly standardized mean difference as criteria to account for the variability that occurs within in vivo screens. The combination of both criteria resulted in a significant reduction in the total number of hits, from 487 (16%) to 49 (1.6%), which is a large decrease in total hits. The authors should comment on whether some of the 487 compounds were randomly tested individually to confirm that they were true negatives and compared to the true hits of 49. For example, did the authors independently test calcipotriene and diphenylboric acid to confirm that these are truly negative?

      Initially, we did not directly retest any of the 487 compounds outside of the top 49 hits with the intention to compare their effect size to the top compound. While we do suspect that there is a chance that some of these compounds may still have a significant effect on lipoprotein levels, we prioritized compounds with the strongest overall biological effect size and largest statistically significant effect. We agree with the Reviewer that it would be interesting to continue to test more of these compounds and examine whether the lipoprotein reduction phenotype correlated with the primary screen effect, this would be a large experimental undertaking, though a randomized subset could be tested. Validation studies of Calcipotriene (Supplemental Figure 2E) showed a small, but significant, lipoprotein reduction at only the 0.25 µM dose across 3 independent experiments. Because the effect is small and not a dose-response, it is not a highly prioritized hit. We did not validate diphenlylboric acid in this study but will in the future.

      (2) The transcriptomics profile studies are not well described. The authors do not provide a detailed list of differential genes for each of the conditions in their dataset. This should be included.

      Supplemental Table 2 includes the fold change and p-values of all genes measured in differential expression analysis. We agree with the Reviewer and added new supplementary tables (Supplemental Table 5 and 7) that contain the differentially expressed genes at each treatment duration.

      (3) The Gene Ontology analysis results are not accompanied by a list of genes that match the GO terms. The authors should include these results. This part was confusing as it was not clear if the cholesterol pathway affected was due to an abundance of genes that were up- or down-regulated.

      We agree with Reviewer 1 and added Supplemental Table 6 that contains each gene ontology term with the associated DE genes that contribute to the significance of that term as well as explicit directionality information.

      (4) Figure 6A shows heat maps of differential gene expression, but there is no key provided in the figure or legend. Are they color-coded for fold change or log2 fold change?

      We agree that Figure 6A can be approved and added a key to the legend as suggested by the Reviewer.

      (5) The overlap of genes that are changed in hnf4a mutants and with enoxolone is provided as percentages, but the actual genes are not listed. Do these genes represent cholesterol biosynthesis pathways? There are other bioinformatic tools that the authors could apply to their dataset for further analyses. For example, enrichr (https://maayanlab.cloud/Enrichr/) is one tool to query for GO Biological processes, cell/tissue type, and even overlap of genes with genetic models and other disease states. An extension of these bioinformatic studies will be useful to determine if other pathways are relevant.

      We agree with the Reviewer and added a table of the genes that overlap between drug treatments and hnf4a mutants (Supplemental Table 7). We also appreciated the Reviewer’s suggestion to deploy Enrichr, which we have done and modified the Results section to now read:

      “Of the 439 differentially expressed genes from 12, 16, and 24 hours post-treatment, 34 differentially expressed genes are shared between all three treatment durations and are associated with gene ontology terms related to carbohydrate metabolism and signaling pathways (Figure 7D). We expanded this analysis using the bioinformatic tool Enrichr [68–70], which largely recapitulated gene ontology results described above. However, Enrichr analysis revealed significant overlap between several late enoxolone-responsive gene sets and transcriptional signatures associated with prochlorperazine, another compound identified in our screen (Figure 2, Figure 3, Supplemental Figure 1AO, Supplemental Figure 2Z). Enrichment of prochlorperazine-associated signatures was observed at 12 (prochlorperazine MCF7 up, 4/58 genes [INSIG1;IRF7;ISG15;ATF3], adjusted p-value = 0.006), 16 (prochlorperazine MCF7 up, 6/58 genes [INSIG1;DDIT4;IRF7;PMAIP1;ISG15;ATF3], adjusted p-value = 0.00008; prochlorperazine PC3 up, 3/29 genes [INSIG1;DDIT4;ATF3], adjusted p-value = 0.006), and 24 hours post-treatment (prochlorperazine PC3 up, 6/29 genes [DUSP5;INSIG1;DDIT3;TRIB3;SQSTM1;ATF3], adjusted p-value = 0.0001; prochlorperazine MCF7 up, 7/58 genes [DDIT3;INSIG1;IRF7;PMAIP1;ISG15;SAT1;ATF3], adjusted p-value = 0.0005). These results suggest that enoxolone and prochlorperazine may perturb overlapping molecular pathways, an observation that warrants further investigation. Ultimately, these data demonstrate distinct early and late responses to enoxolone treatment, and the early response modulates key lipid metabolism pathways.”

      (6) Figures 1E and 1F would benefit if the exact spots/points where enoxolone and other hits mentioned in the text were labelled.

      Great idea, we modified Figure 1F as suggested.

      (7) Figure 3 and Figure 4 graphs should state Fold change on the Y-axis title.

      We are thankful the Reviewer noticed this typo and we have fixed both Figures.

      Reviewer #2 (Recommendations for the authors):

      (1) To boost the impact for more readers, the authors should include the LipoGloElectrophoresis and LipoGlo-Microscopy results from a few more of the validated hits, especially ones that are completely novel (unlike Enoxolone, which already had a known role in lipid metabolism). Results on enoxolone are useful as a validation of the assay, mostly with some minor additional insights.

      We agreed with Reviewer 2 and added more validation testing of for a few hits (see response to Reviewer 1, Strengths and Reviewer 2, Weaknesses). We did not perform these experiments on each drug for technical reasons, mainly because these experiments are low-throughput (especially the Microscopy). Nonetheless, we performed many additional experiments to add phenotyping data for 6 new drugs that included multiple LipoGlo-Electrophoresis panels to an entirely new Figure.

      (2) The authors should include raw data from the screen from all drugs tested in the supplementary and then for which SSMD was calculated for, providing in an excel sheet or similar the values and how these were calculated, i.e. the 487 unique drugs that lower B-lp levels with an SSMD cutoff of < -1.0.

      This information was provided in the supplemental file as separate .csv files with the associated R script, which can be run locally and contains the SSMD functions. Considering the Reviewer comments, we ensured this information in provided in Supplemental Tables 1 and 2.

      (3) Page 3, lines 23-24: What does the 2 to 4 fold chance mean? Perhaps rewrite: Genetic mutations in Lipoprotein(a) increase the chance of heart attack or stroke 2-4 fold greater than without the mutation.

      We agree and the sentence now reads: Patients with genetic mutations in the Lipoprotein(a) encoding gene have a 2 to 4 fold increased risk of sudden heart attack or stroke.

      (4) Figure 1, for C and D, label some of the most significant hits and definitely show where exonolone lies.

      We agree see response to Reviewer 1 Pt6

      (5) Page 6, lines 1-9: I'm a bit confused why this is here if you do not present the data.

      We thought this was relevant information to share for researchers that running drug screens with positive controls and defining hit cutoffs. In light of the Reviewer’s comment, we removed the last sentence from this paragraph.

      (6) Page 6, line 8: This needs better justification of why you are validating enoxolone rather than other hits; otherwise, it could seem like cherry picking. Especially as enoxolone is known to affect lipid metabolism. Otherwise present more details of a couple of validated candidates.

      We agree. As the Reviewer requested, we validated more compounds (described above) and modified the text of the results to elaborate on our justification for selecting enoxolone for further study. The text of the results now reads: Hit compounds were prioritized for follow-up studies based on reproducible dose-dependent responses, minimal toxicity as indicated by normal morphology over development, lack of direct NanoLuciferase inhibition, and prior reports the presence of literature suggesting potential links to lipid metabolism. One compound meeting these criteria was enoxolone, also known as 18β-Glycyrrhetinic acid, (Figure 2 Drug 20, Supplemental Table 1, Supplemental Figure 1T, Supplemental Figure 2L).

      (7) Supplementary Figures 1 and 2: The resolution is too low, and the reader cannot even see the charts or the text.

      We agree and now have uploaded higher resolution images

      (8) Page 7, lines 31-35: Needs a higher resolution and magnified image to merit this 'offhand' statement. Also cite reference [59] here.

      We agree with the Reviewer and added magnified insets of the heart. As far as the suggestion of adding Ref 59, we do not see the connection to that paper (A point mutation decouples the lipid transfer activities of microsomal triglyceride transfer protein PLOS Genetics 16:e1008941)

      (9) Figure 3E: In addition to the proportions graph would be useful to also have an absolute amount of lipoprotein in each class graph.

      While there may be changes in total luminescence values from lane to lane in these gels, we have not fully validated the absolute quantitation of a full lane. We typically use plate-based whole-animal assays to determine total absolute lipoprotein levels and calculate the proportion of the whole lane for each lipoprotein class, as described in our prior publication detailing the assay. We do expect that there is some additional variation incorporated into the native PAGE assay due to sample freeze/thaw, dilution, and loading.

      (1) Page 10 lines 27-28: "Continual statin use for more than 1 year reduced circulating Blps and all-cause mortality by ~30% in individuals with high B-lp levels." This sentence doesn't seem right, intimates continual statin use causes death - I don't think that's right.

      We thank the Reviewer for catching this and have corrected the sentence. It now reads: “Continual statin use for more than 1 year in individuals with high B-lp levels reduced circulating Blps and lowered all-cause mortality by ~30% [12,13].”

      (11) Page 12, line 25: "canlikely" is a typo, should be can likely.

      We fixed that sentence and now reads: “Further, the drug screening paradigm we developed using the LipoGlo system is highly scalable and can be deployed to screen large novel drug libraries to identify many additional B-lp-lowering compounds.”

      (12) Figure 1 legend: "An ordered plot of each SSMD score measured from 5 μM lomitapide treated animals from each 96-well plate (n = 1381) relative to respective vehicle treatment." This comes across as though it's 5uM Iomitapide/vehicle. But it's the SSMD score of each drug compared to Iomitapide and relative to the respective vehicle (I think) - make it clearer.

      We agree and clarified the legend so it now reads: “…(D) An ordered plot of each SSMD score measured from positive control (5 µM lomitapide) treated animals from each 96-well plate (n = 1381) relative to respective vehicle treatment. The solid black line at y = 0 represents the divide in increased and decreased SSMD score, the solid blue line at y = -1.41 represents the curve's inflection point, and the dashed black line at y = -1 represents the SSMD (open circles) cutoff used to define a hit. “

      (13) Figure 1E: What is the x axis?

      Each data point on the x-axis represents each drug at every dose tested, we will clarify the test. The legend now reads: “(E) A plot of SSMD scores measured from each drug at each dose tested, each open circle represents the SSMD score of an individual drug at an individual dose.” In addition, “Compound (each dose tested)” was added to the x-axis of the figure panel.

      (14) Figure 2: Would you not have space to put the drug names in the figures? Where, for example, is enoxolone?

      We agree and have updated the figure accordingly.

      (15) Figure 3A-C: label enoxolone on the x axis.

      We agree and added this text to what is now Figure 4.

      (16) Figure 3D: Looks like delayed development with enoxolone, if left to grow, would the embryos develop normally?

      We did not examine if animals treated from 3-5 dpf develop normally beyond 5 dpf.

      (17) Figure E. Are stars all compared to vehicle control? Perhaps useful to have lines to indicate what are the significantly different relationships.

      Comparison in these experiments are always to the vehicle (negative control) and clarified the legend considering the Reviewer’s comment we modified the legend to now read,”… * <0.05 as compared to vehicle.”

      (18) Figure 3E: As well as the proportion of total lipoprotein, it would also be beneficial to see absolute lipoprotein levels.

      See above response to Reviewer 2 Pt9.

      (19) Figure 4: Does the overall health or size of the animal correlate with the luminescence score?

      While we do know that, in untreated animals, lipoprotein levels vary with age (and, thus, size), we have not examined this more granularly than in 24-hour time points after treatment. Further, we have not examined this in the context of a drug treatment.

      (20) Figure 4D: Again the absolute in each fraction would be meaningful, also the 5078 looks brighter?

      See above response to Reviewer 2 Pt9.

      (21) Figure 4A, 5A: Would the traces (line plots) not be useful here to see the overall dynamics over time?

      We considered presenting Figure 5A this way but decided to keep the plots as is because we wanted to show the individual points which make the line blots very difficult to read. Further, our analysis evaluates individual animals at each time point as it is not possible to follow the same animal over time.

      Reviewer #3 (Recommendations for the authors):

      Figure 4: Consider using standard scientific notation for the very small p values in some of the figure legends.

      We agree and adjusted the p-value notation as suggested.

    1. eLife Assessment

      This study provides important insights into how researchers can use perceptual metamers to formally explore the limits of visual representations at different processing stages. The framework is compelling and the data support the claims.

    2. Reviewer #1 (Public review):

      This is an interesting study on the nature of representations across the visual field. The question of how peripheral vision differs from foveal vision is a fascinating and important one. The majority of our visual field is extra-foveal, yet our sensory and perceptual capabilities decline in pronounced and well-documented ways away from the fovea. Part of the decline is thought to be due to spatial averaging ('pooling') of features. Here, the authors contrast two models of such feature pooling with human judgments of image content. They use much larger visual stimuli than in most previous studies, and some sophisticated image synthesis methods to tease apart the prediction of the distinct models.

      More importantly, in so doing, the researchers thoroughly explore the general approach of probing visual representations through metamers-stimuli that are physically distinct but perceptually indistinguishable. The work is embedded within a rigorous and general mathematical framework for expressing equivalence classes of images and how visual representations influence these. They describe how image-computable models can be used to make predictions about metamers, which can then be compared to make inferences about the underlying sensory representations. The main merit of the work lies in providing a formal framework for reasoning about metamers and their implications, for comparing models of sensory processing in terms of the metamers that they predict, and for mapping such models onto physiology. Importantly, they also consider the limits of what can be inferred about sensory processing from metamers derived from different models.

      Overall, the work is of a very high standard and represents a significant advance over our current understanding of perceptual representations of image structure at different locations across the visual field. The authors do a good job of capturing the limits of their approach I particularly appreciated the detailed and thoughtful Discussion section and the suggestion to extend the metamer-based approach described in the MS with observer models. The work will have an impact on researchers studying many different aspects of visual function including texture perception, crowding, natural image statistics and the physiology of low- and mid-level vision.

      The main weaknesses of the original submission relate to the writing. A clearer motivation could have been provided for the specific models that they consider, and the text could have been written in a more didactic and easy to follow manner. The authors could also have been more explicit about the assumptions that they make.

      Comments on revised version.

      The authors have now fully addressed my concerns and I think the paper is a valuable contribution. In future studies within the same research program I would appreciate seeing further consideration of how metamerism at different stages of visual processing interact to determine behaviour in tasks. For example, there are presumably interesting impacts of feedback that may modify feature spaces, thereby rendering aspects of appearance that were previously metameric perceptually discriminable.

    3. Reviewer #2 (Public review):

      Summary:

      The authors have improved clarity overall and have spoken to most of the issues raised by the reviewers. There are still two outstanding problems however, where issues raised during the review were inappropriately dismissed in the manuscript. These should be explicitly addressed as limitations to the results presented (no eye tracking), and early pilot experiments that informed the experiments as presented (pink noise) rather than brushed off as 'unnecessary' and 'would be uninformative'.

      Eye tracking:<br /> It is generally accepted that experiments testing stimuli presented at specific locations in peripheral vision require eye tracking to ensure that the stimulus is presented as expected, in particular, in the correct location. As I stated in the previous round of review, while a stimulus presentation time of 200ms does help eliminate some saccades, it does not eliminate the possibility that subjects were not fixating well during stimulus onset. I am also unclear what the authors mean by 'trained observer' in this context, though the authors state that an author subject in a different portion of the paper is an 'expert observer'. Does this mean the 'trained observers' are non-expert recruited subjects? Given the conditions tested differ from previous work (Freeman & Simoncelli, 2011) *these differences are a main contribution of the paper!* which DID include eye tracking in a subset of subjects, it is entirely possible to get similar results to this work in the context of non eye-tracking controlled stimulus presentation. The reasons now in the manuscript are not reasons that make eye tracking 'considered unnecessary'.

      I appreciate that the authors now state the lack of eye tracking explicitly, but believe the paper needs to at least state that this is a limitation of the results reported, and eyetracking being 'considered unnecessary' is unreasonable, nor a norm in this subfield.

      N=1:<br /> The authors now state clearly the limitations of a single subject in the manuscript, and state the expertise level of this subject.

      Large number of trials:<br /> The authors now address this, and include an enumeration of the large number of trials.

      Simple Models / Physiology comparison:<br /> I support the choice to reduce claims regarding tight connections to physiology, and appreciate the explanation of the luminance model.

      Previous Work:<br /> I appreciate the author's changes to the introduction, both in discussing previous work and citation fixes.

      Blurred White, Pink Noise:<br /> While the authors now address pink noise, the explanation for such stimuli being expected to be uninformative is confusing to me. The manuscript now first states that pink noise is a natural choice, then claims it would be uninformative, while also stating in the rebuttal (not the manuscript) that they tried it and it indeed reduced the artifacts they note. The logic of the experiments indeed relies on finding the smallest critical scaling value, which is measured by subjects determining if a synthesis is similar or different to a target or second synth. A synthesis free from artifacts would surely affect the subjects' responses and the smallest critical scaling measured.

      The statement that the authors experimented with pink noise early on and found this able to address the artifacts should be stated in the manuscript itself, not just in the rebuttal, and the blanket statement that this experiment would be 'uninformative' is incorrect. Surely this early pilot the authors mention in the rebuttal was informative to designing the experiments that appear in the final paper and would be an informative experiment to include.

      Comments on revised version.

      The authors have addressed my outstanding concerns, adding discussion about the limitations of not having eye tracking in the study, details about the subject pool, limitations of a subset of the study which contains a single subject, and experiments with pink noise seeds, and this relationship to largest vs smallest critical scaling. In addition, they have added clarity around internal noise vs metamerism in the context of this study as raised by the other reviewer.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments following re-submission:

      Overall, I think the authors have done a satisfactory job of addressing most of the points I raised.

      There’s one final issue which I think still needs better discussion.

      I think reviewer 2 articulated better than I have the point I was concerned about: the relationship between JNDs and metamers as depicted in the schematics and indeed in the whole conceptualization.

      I think the issue here is that there seems to be a conflating of two concepts- ’subthreshold’ and ’metamer’-and I’m not convinced it is entirely unproblematic. It’s true that two stimuli that cannot be discriminated from one another due to the physical differences being too small to detect reliably by the visual system are a form of metamer in the strict definition ’physically different, but perceptually the same’.

      However, I don’t think this is the scientifically substantial notion of metamer that enabled insights into trichromacy. That form of metamerism is due to the principle of univariance in feature encoding, and involves conditions in which physically very different stimuli are mapped to one and the same point in sensory encoding space whether or not there is any noise in the system. When I say ’physically very different’ I mean different by a large enough amount that they would be far above threshold, potentially orders of magnitude larger than a JND if the system’s noise properties were identical but the system used a different sensory basis set to measure them. This seems to be a very different kind of ’physically different, but perceptually the same’.

      We are in full agreement with this. Typically, the notion of metamers is about deterministic information loss, which can be modeled as a projection from a high-dimensional physical space to a lower-dimensional perceptual space. This is the topic of the paper, and is analogous to the work on color matching in the 19th century.

      In contrast, sensitivity to small differences within the perceptual space due to internal noise is usually addressed by other methods, such as signal detection theory, a topic which is not the focus of this paper. It is analogous to the work on discriminability of colors such as MacAdam ellipses (MacAdam, 1942). It is instructive to look at progress in the color field. The color matching experiment and the question of metamerism is quite well worked out, whereas the question of how to quantify discriminability within that space has been an ongoing topic of investigation for over a century.

      Here, we aim to develop and test a model of metamers, analogous to the color matching experiments, but we do not attempt to develop a model of discriminability. Nevertheless, while the two types of information loss are conceptually distinct, they are both present in the nervous system of the observer, and both are reflected in the performance vs. scaling plots in our paper. We have added clarifications about this point in the introduction on page 2, starting on line 45, and the discussion, starting on page 16, line 397.

      Finally, regarding physical differences between stimuli: the differences between target images and synthesized metamers are quite large (high mean squared error), as shown in Appendix 5. In no condition did we present subjects with stimulus pairs that were physically similar.

      I do think the notion of metamerism can obviously be very usefully extended beyond photoreceptors and photon absorptions. In the interesting case of texture metamers, what I think is meant is that stimuli would be discriminable if scrutinised in the fovea, but because they have the same statistics they are treated as equivalent.

      The notion of “texture metamers” is perhaps a reference to the work by Freeman and Simoncelli (2011), whose stimuli are similar to ours: when synthesized using a model with sufficiently small scaling, they are indiscriminable, and therefore metamers. The reviewer is of course correct that the stimulus pairs are not metamers when the observers move their eyes due to differences in spatial encoding as a function of eccentricity. That is, they are only metameric under a specific set of viewing conditions, and they are not metameric when those conditions are violated. The same is true for color metamers, as the spectral sensitivity of the cones also differ with eccentricity (Stockman and Sharpe, 2000).

      I think the discussion of this could still be clearly articulated in the manuscript. It would benefit from a more thorough discussion of the difference between metamerism and subthreshold, especially in the context of the Voronoi diagrams at the beginning.

      We agree that a more thorough discussion of the diagrams could help clarify the issues to the reader. We have modified the caption of figure 1 with the goal of clarifying interpretation of the diagrams, and see also our discussion earlier in this note about noise and discriminability.

      It needs to be made clear to the reader why it is that two stimuli that are physically similar (e.g., just spanning one of the edges in the diagram) can be discriminable, while at the same time, two stimuli that are very different (e.g., at opposite ends of a cell) can’t.

      Do the cells include BOTH those sets of stimuli that cannot be discriminated just because of internal noise AND those that can’t be discriminated because they are projected to literally the same point in the sensory encoding space? What are the strengths and limits of models that involve the strict binarization of sensory representations, and how can they be integrated with models dealing with continuous differences? These seem like important background concepts that ought to be included in either the introduction of discussion sections. In this context it might also be helpful to refer to the notion of ’visual equivalence’ as described by:

      This is an important point and we appreciate the reviewer raising it. In brief, as one traverses a region in one of the Voronoi diagrams, the images are changing physically but subject to the constraint that they all project to the same single point in the reduced perceptual space. When one crosses from one region to another, the images now project to a different point in the perceptual space. Whether or not that the two locations in the perceptual space are distant enough to be distinguishable given the internal noise is a question pertaining to the topic of JNDs in the perceptual space, rather than the mapping from physical space to the perceptual space. We do not address that question in detail in this paper, though we do now reference it in the caption of figure 1, as well as in the new sections in the introduction and discussion mentioned earlier in this response.

      We do note that the perceptual space is not discrete: the model outputs are real-valued. The apparent discretization is a limitation of the simplified 2-D schematics.

      Ramanarayanan, G., Ferwerda, J., Walter, B., & Bala, K. (2007). Visual equivalence: towards a new standard for image fidelity.ACM Transactions on Graphics (TOG), 26(3), 76-es.

      Other than that, I congratulate the authors on a very interesting study, and look forward to reading the final version.

      Reviewer #2 (Public review):

      Summary:

      The authors have improved clarity overall and have spoken to most of the issues raised by the reviewers. There are still two outstanding problems however, where issues raised during the review were inappropriately dismissed in the manuscript. These should be explicitly addressed as limitations to the results presented (no eye tracking), and early pilot experiments that informed the experiments as presented (pink noise) rather than brushed off as ’unnecessary’ and ’would be uninformative’.

      Eye tracking:

      It is generally accepted that experiments testing stimuli presented at specific locations in peripheral vision require eye tracking to ensure that the stimulus is presented as expected, in particular, in the correct location. As I stated in the previous round of review, while a stimulus presentation time of 200ms does help eliminate some saccades, it does not eliminate the possibility that subjects were not fixating well during stimulus onset. I am also unclear what the authors mean by ’trained observer’ in this context, though the authors state that an author subject in a different portion of the paper is an ’expert observer’. Does this mean the ’trained observers’ are non-expert recruited subjects?

      Given the conditions tested differ from previous work (Freeman & Simoncelli, 2011) ‘these differences are a main contribution of the paper!’ which DID include eye tracking in a subset of subjects, it is entirely possible to get similar results to this work in the context of non eye-tracking controlled stimulus presentation. The reasons now in the manuscript are not reasons that make eye tracking ’considered unnecessary’.

      I appreciate that the authors now state the lack of eye tracking explicitly, but believe the paper needs to at least state that this is a limitation of the results reported, and eyetracking being ’considered unnecessary’ is unreasonable, nor a norm in this subfield.

      By “trained” observers, we mean people who were recruited from the community of vision science labs at NYU and who have participated in many visual psychophysics experiments. All of the participants are “trained” in this sense, and are thus used to maintaining fixation while performing peripheral tasks. One of these participants, an author, was also an expert in the specific content area of the paper. By “expert”, we mean high familiarity with the stimulus types and models employed in the paper.

      We have now further clarified this in the text in the subsection of the methods on Observers, on page 22.

      We also discuss the issue at greater length in the methods subsection Apparatus, on page 26. We removed the word "unnecessary" and make it clear that while we don’t think our results are undermined, the lack of eye tracking is nonetheless a limitation.

      N=1: The authors now state clearly the limitations of a single subject in the manuscript, and state the expertise level of this subject.

      Large number of trials: The authors now address this and include an enumeration of the large number of trials.

      Simple Models / Physiology comparison: I support the choice to reduce claims regarding tight connections to physiology, and appreciate the explanation of the luminance model.

      Previous Work: I appreciate the author’s changes to the introduction, both in discussing previous work and citation fixes.

      Blurred White, Pink Noise: While the authors now address pink noise, the explanation for such stimuli being expected to be uninformative is confusing to me. The manuscript now first states that pink noise is a natural choice, then claims it would be uninformative, while also stating in the rebuttal (not the manuscript) that they tried it and it indeed reduced the artifacts they note. The logic of the experiments indeed relies on finding the smallest critical scaling value, which is measured by subjects determining if a synthesis is similar or different to a target or second synth. A synthesis free from artifacts would surely affect the subjects responses and the smallest critical scaling measured.

      The statement that the authors experimented with pink noise early on and found this able to address the artifacts should be stated in the manuscript itself, not just in the rebuttal, and the blanket statement that this experiment would be ’uninformative’ is incorrect. Surely this early pilot the authors mention in the rebuttal was informative to designing the experiments that appear in the final paper, and would be an informative experiment to include.

      First, we did render some test stimuli with pink noise seeds, but we did not collect psychophysical data, hence there are no results we could add. Visual inspection of these stimuli was indeed clarifying in the following sense. The pink noise stimuli had fewer high-frequency “artifacts”. If our goal was to synthesize stimuli that are indistinguishable from the original stimulus, as one might do to save compute power when in a device that for foveated rendering, then starting with pink noise would be better than starting with white noise. Our purpose was just the opposite. For our experiments, the artifacts were just what we wanted: the more artifacts, the better. The reason is that a strongest test of a metamer model is whether two stimuli that are as physically different from one another as possible, are nonetheless indistinguishable when their model representations are the same. Stimuli synthesized from pink noise seeds are harder to discriminate from the target stimulus, not easier. Thus using them in an experiment would result in a larger estimate of critical scaling. Since our explicit goal was to estimate the smallest critical scaling window, these stimuli would not bring us closer to our goal. As the reviewer points out, these metamers were “informative” in the sense that they informed our experimental design, but they are “uninformative” (relative to white noise seeds) for estimating the critical scaling.

      We have updated our description in the discussion starting on page 19, line 449, and included a new appendix to demonstrate this point (appendix 2 on page 35).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Typo: p. 19, l. 439: ’Why does asymptotic performance, but not critical scaling, depends on image content?’": remove ’s’ from ’depends’.

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      Recommendations: State that the lack of eye tracking to control stimulus presentation is a limitation of the results presented.

      Remove the claim that pink noise or filtered white noise seeds would be uninformative, and mention the fact that the authors in fact experimented with pink noise seeds in an early version of the experiments (which was surely informative to the experimental setup as presented here).

      Addressed as described above.

      References

      Freeman J, Simoncelli EP. Metamers of the ventral stream. Nature Neuroscience. 2011 aug; 14(9):1195–1201. doi: 10.1038/nn.2889.

      MacAdam DL. Visual Sensitivities To Color Differences in Daylight*. Journal of the Optical Society of America. 1942 may; 32(5):247. http://dx.doi.org/10.1364/josa.32.000247, doi: 10.1364/josa.32.000247.

      Stockman A, Sharpe LT. The Spectral Sensitivities of the Middle- and Long-Wavelength-Sensitive Cones Derived From Measurements in Observers of Known Genotype. Vision Research. 2000 jun; 40(13):1711–1737. doi: 10.1016/s0042-6989(00)00021-3.

    1. eLife Assessment

      This is an important study establishing the mechanistic principles how NUP98-KDM5A phase separates together with H3K4me3 chromatin in leukemia. Methods are sound and results are convincing, providing a framework to understand gene expression changes in leukemia patients.

    2. Reviewer #1 (Public review):

      Leukemia-driving NUP98 oncofusion proteins form chromatin-associated biomolecular condensates in the nucleus, and these structures are important for oncogenic transformation. Most NUP98 fusions do not contain domains that mediate the recognition of specific DNA elements. Instead, they entail domains that are important for chromatin regulation. For instance, the NUP98::KDM5A fusion features a fusion of the NUP98 N-terminus with the third PHD domain of the histone demethylase KDM5A. As PHD domains are critical for the recognition of methylated histones without any sequence specificity, it is not clear what controls the condensation and chromatin binding of NUP98::KDM5A, leading to the induction of oncogenic transcriptional programs.

      In this work, the authors use a combination of cellular and in vitro studies to show that biomolecular condensation of NUP98::KDM5A is dependent on H3K4me3 binding. Their model proposes that concentration-dependent chromatin-associated condensation of NUP98::KDM5A depends on local densities of H3K4me3 on chromatin and the levels of the fusion oncoprotein. In line with this, the analysis of gene expression data from NUP98::KDM5A-positive AML cells shows a positive correlation between differentially expressed genes and H3K4me3 levels.

      This is an interesting manuscript that aims to dissect the molecular mechanisms underlying biomolecular condensation of the NUP98::KDM5A oncoprotein. The work is solid, and the results are well explained and presented in a logical order. However, the study suffers from several weaknesses that if addressed would improve the study.

      Major points:

      (1) All cellular experiments are performed in settings of transient transfection of NUP98::KDM5A in non-hematopoietic cell types. These conditions are not physiologically relevant, as these cells do not depend on the fusion oncogene. Therefore, any claims about concentration-dependent effects on condensation need to be validated in AML cells that are driven by NUP98::KDM5A. While this may not be possible in primary patient-derived cells, several groups have published AML models of NUP98::KDM5A-driven AML that could be used.

      (2) The results presented in Figure 4 are not entirely supportive of the mechanism. It is known that active gene expression correlates with high H3K4me3 levels; therefore, the correlations shown by the authors are expected. Yet, the authors do not discuss the fact that many H3K4me3-positive genomic regions do not show NUP98::KDM5A binding. This should be elaborated on in the discussion section.

      (3) While the focus of the manuscript is on NUP98::KDM5A, this oncofusion is part of a family of >30 fusions that join the NUP98 N-terminus to a variety of factors with roles in epigenetic control and transcription. While the repertoire of NUP98 fusion partners is diverse with regard to functional domains, they all induce a conserved set of target genes that is characteristic of this leukemia subtype. How can this be achieved in the context of NUP98 fusion proteins that do not contain a PHD domain, such as NUP98::NSD1 or NUP98::HOXA9? Please discuss this.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors investigate how the oncogenic fusion protein NUP98-KDM5A alters gene expression in leukemia, using a combination of cellular experiments with model and patient cell lines, as well as in vitro studies. Upon transfection of U2OS cells with mEGFP-tagged NUP98-KDM5A, the authors show that the fusion proteins form sub-micrometer puncta, whereas KDM5A alone does not. These foci are also observed at expected native expression levels (using OpenCell data). The tag has an effect here, as switching to an mCherry tag raises the apparent saturation concentration for phase separation. Finally, the authors show via super-resolution imaging that the foci correlate with H3K4me3 distribution.

      In vitro, the fusion protein forms amorphous, gel-like condensates at double-digit nanomolar concentrations. Truncation analysis identifies PHD3 of KDM5A as required for maximal phase separation, consistent with the ability of the protein to bind H3K4me3 peptides. Addition of polynucleosomes increases the amount of fusion protein partitioning into the condensate in an H3K4me3-binding-dependent manner. Condensates are gel-like with slow internal dynamics in vitro; in cells, however, the dynamics depend on the position of the EGFP tag (no corresponding experiments with mCherry are shown). Reconstitution with H3K4me3- and H3K4me0-modified arrays shows colocalization with both wild-type NUP98-KDM5A and the binding mutant. Here, H3K4me3 arrays recruit ~20% more protein and yield gel-like structures in a manner dependent on the PTM and on the PHD finger.

      In cells, the fusion protein colocalizes with H3K4me3-marked loci, including the HOX clusters, as confirmed by FISH. Finally, re-analysis of published expression datasets from patient cells shows that genes are predominantly upregulated and that the upregulated genes are H3K4me3-marked.

      This is a well-executed mechanistic study. The data convincingly establish that NUP98-KDM5A forms sub-micrometer foci at realistic expression levels, that these foci correlate with H3K4me3-marked sites, that the PHD3-H3K4me3 interaction mediates chromatin binding while the NUP98 moiety drives phase separation in vitro, that foci in cells overlap genes heavily decorated with H3K4me3, and that H3K4me3-marked genes are those found to be upregulated in patient datasets. These are important mechanistic findings and of interest to the community.

      Still, the functional/causal link is a bit more tentative, as the data is mostly correlative, since it is not directly established that there is feedback between H3K4 methylation, NUP98-KDM5A recruitment, phase separation and target gene overexpression. An experiment that could further bolster this claim would be a direct test of whether NUP98-KDM5A expression drives overexpression of bound genes, e.g. expression of the fusion protein vs PHD- and NUP98-mutant variants, followed by qPCR of target genes, such as the HOX cluster, and possibly H3K4me3 ChIP at the same loci. As all the constructs and cell lines exist, this could be feasible and would substantially strengthen the manuscript.

    4. Author response:

      We greatly appreciate the positive and constructive comments from the reviewers, which recognized the intellectual contributions of our manuscript and also captured its limitations. Below is our response to the three major points from Reviewer #1 and the comments from Reviewer #2.

      Response to Reviewer #1

      (1) Use of transient transfection in non-hematopoietic cells. We appreciate the reviewer’s point regarding physiological relevance. Our goal in these cellular microscopy experiments was to dissect the biophysical principles of NUP98::KDM5A (including its mutants) condensate formation under controlled expression levels (concentration). While AML model systems driven by NUP98::KDM5A are available, they do not offer this possibility because of pre-existing NUP98::KDM5A expression. The suspension culture of hematopoietic cells also brings practical challenges for correlative FISH+IF and high-resolution microscopy analysis. We agree that validating our observed behaviors in AML models would be valuable, e.g. by creating HSPC lines with inducible expression of tagged NUP98::KDM5A and its mutants, but such experiments fall outside the scope of the current study. We will add text acknowledging this limitation and clarifying that our mechanistic conclusions are grounded in biophysical principles that should generalize across cell types.

      (2) Correlation between H3K4me3 and gene activation. We agree that active transcription correlates with H3K4me3, and that this baseline relationship must be considered. Our analysis explicitly uses fold-change between patient cells and healthy controls as the readout. This comparison inherently accounts for the activating effect of H3K4me3 itself. Regarding the reviewer’s comment that “many H3K4me3-positive genomic regions do not show NUP98::KDM5A binding”, we would like to clarify that this point is exactly what our manuscript aims to explain. Our cell line studies demonstrate that, at a patient-relevant expression level, NUP98::KDM5A condensates form preferentially at H3K4me3 locus with high local mark density. Considering that NUP98::KDM5A concentration in the nucleus is lower than the K_D between KDM5A PHD3 and H3K4me3, this means that H3K4me3 loci with lower mark density will not see NUP98::KDM5A binding without the high local concentration of the fusion protein (as a result of condensate formation). This is consistent with our observation that genes with the highest local density show disproportionately stronger upregulation. We will further clarify this point in the revised manuscript.

      (3) Generalization to other NUP98 fusions lacking PHD domains. We appreciate this important conceptual question. We will expand the discussion to note that many NUP98 fusions, despite diverse partner domains, produce similar transcriptional programs. Our current hypothesis is that the highly active status of the HOX cluster genes in HSPC attracts NUP98 oncofusions targeting H3K4me3 (e.g. KDM5A and PHF23), while NUP98::HOXA9 and other transcription factor fusions directly target the HOX cluster via DNA sequence recognition. This convergence and the broad targets of the dysregulated HOX transcription factors explain the transcriptional program similarity, sustained by the previously described enrichment of transcriptional co-activators by NUP98 oncofusion condensates. We will explicitly discuss this hypothesis in our revised manuscript.

      Response to Reviewer #2

      We thank the reviewer for the thoughtful evaluation and agree that the causal link between H3K4me3 recognition, condensate formation, and gene activation remains partly correlative. Experiments such as qPCR or ChIP following expression of WT versus mutant constructs would indeed strengthen the causal chain. However, performing these assays across multiple constructs and loci in a physiologically relevant system is not feasible within the current revision cycle. We have added text acknowledging this limitation and clarifying that our study focuses on establishing the biophysical mechanism of targeting, while functional consequences are inferred from patient datasets rather than new perturbation experiments.

    1. eLife Assessment

      In this important study, the authors describe the mechanisms by which CDK4/6 overexpression mediates resistance to Osimertinib in EGFR mutant lung cancer models. Their data show that CD4/6 overexpressing cells had increased replicative stress and genomic instability. The evidence is solid, but inclusion of some additional experiments would make this stronger.

    2. Reviewer #1 (Public review):

      Summary:

      The authors used a panel of cell models to determine whether CDK4/6 overexpression resulted in resistance to the EGFR inhibitor Osimertinib, and the mechanisms underlying the resistance.

      (1) Major Concerns (highest priority):

      There is a lack of detail about the methodology in the results section/figure legends, which makes it difficult to interpret the data. Sometimes, adequate information is also not included in the methods themselves. For example, Figure 1A: how many doses did each mouse receive? How long after dosing were animals sacrificed? Figures 1E and 2A: is this RNA-seq analysis?

      Using a second EGFR inhibitor for some of the key experiments would increase the rigor of the studies shown.

      (2) Nice to have experiments:

      Using CRISPR KO of CDK4 in the CDK4-amplified HCC827 and testing response to Osi and presence of replication stress would also increase the rigor of the studies.

      The authors show that in their patient data, some cell cycle regulators which are amplified in NSCLC at similar rates to CDK4/6, such as CCNE1, had no increase in FGA. Overexpressing CCNE1 and testing Osi response in their cell models would be a nice test of their proposed mechanism that it is the genomic instability and FGA that are driving resistance. This wouldn't need to be done in vivo, but could be done using cell culture-based methods.

      Similarly, testing the overexpression of some of the proposed target genes, such as STEAP1 and AGR2, on the therapeutic response to Osi in cell culture would also be a nice test of the mechanism proposed.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gini et al. investigate the mechanisms by which CDK4 and CDK6 upregulation drives resistance to EGFR tyrosine kinase inhibitors (TKIs) in EGFR-mutant lung adenocarcinoma (LUAD). The study utilizes preclinical models, including cell line-derived xenografts (CDXs), patient-derived xenografts (PDXs), and primary organoids, alongside large-scale clinical genomic datasets. The authors demonstrate that CDK4 or CDK6 overexpression allows cancer cells to bypass EGFR TKI-induced G1/S arrest, leading to continuous cell cycle progression. This sustained proliferation during EGFR inhibition induces DNA replication stress, activates DNA damage response pathways (such as ATM and TPX2), and ultimately causes genomic instability. The authors also show that this leads in turn to the upregulation of tumor-promoting genes (e.g., AGR2, ASNS, STEAP1) and an epithelial-mesenchymal transition (EMT) phenotype. Moreover, the authors show that combinatorial treatment utilizing TKIs such as osimertinib alongside CDK4/6 inhibitors effectively suppresses proliferation, mitigates DNA damage, and restores TKI sensitivity in preclinical models.

      Overall, this is a highly translational study that provides a strong mechanistic rationale for biomarker-driven clinical trials combining EGFR and CDK4/6 inhibitors. However, there are a few experimental and analytical areas that require clarification or additional data to fully support the authors' conclusions.

      Major Comments:

      (1) Reliance on overexpression models over loss-of-function

      The mechanistic studies mainly rely on overexpression of CDK4 and CDK6 to simulate the amplified state. Although the authors argued that the level of overexpression mimics that observed in resistant tumors, a complementary study in which CDK4/CDK6 were suppressed in a model where CDK4/CDK6 is amplified (such as HCC827 or TH116), and replication stress and osimertinib sensitivity tested would greatly strengthen their observations. Indeed, there is mention of CDK4 constructs to perform knockdown studies in the methods, but those studies are not included in this submission.

      (2) Mechanistic link between genomic instability and specific gene amplifications

      The authors highlight that CDK4/6 activation leads to recurrent copy number gains and transcriptional upregulation of specific pro-tumor genes including AGR2, ASNS, and STEAP1. While the paper establishes that CDK4/6 overexpression causes general genomic instability (increased FGA), it does not mechanistically explain why these specific genes are consistently amplified. The authors should investigate or discuss whether these specific loci are inherently fragile under replication stress, if they are direct downstream targets of the E2F transcriptional program, or if this is a result of random genomic instability followed by strong positive selection under osimertinib pressure.

      (3) Discrepancies in tumor mutational burden (TMB) reporting

      There is a slight contradiction regarding the TMB data that needs to be clarified for readers. The manuscript states that in the clinical datasets, "EGFR-mutant LUAD harboring cell cycle gene alterations exhibited significantly elevated FGA and TMB relative to cell cycle-negative tumors" (Line 265-266). However, in the next section, the authors say, "Notably, no corresponding increase in TMB was observed with CDK4 or CDK6 CNA, similar to our findings in preclinical models" (Line 271-273). The authors should clarify or discuss why broad cell cycle alterations correlate with high TMB, while CDK4/6-specific alterations drive structural instability (FGA) without increasing TMB.

    4. Author response:

      Reviewer 1:

      (1) Major Concerns (highest priority):

      There is a lack of detail about the methodology in the results section/figure legends, which makes it difficult to interpret the data. Sometimes, adequate information is also not included in the methods themselves. For example, Figure 1A: how many doses did each mouse receive? How long after dosing were animals sacrificed? Figures 1E and 2A: is this RNA-seq analysis?

      Thank you for pointing this out. We will provide additional details about the methodology in the results, figure legends, and methods to clarify our findings.

      Using a second EGFR inhibitor for some of the key experiments would increase the rigor of the studies shown.

      Thank you for this suggestion. We agree that a second EGFR inhibitor may increase the scientific rigor of the findings. However, at the time the experiments were completed, osimertinib was the clear standard of care first-line therapy for EGFR-mutated lung cancer and the clinical relevance of using other inhibitors in these experiments is not clear.

      (2) Nice to have experiments:

      Using CRISPR KO of CDK4 in the CDK4-amplified HCC827 and testing response to Osi and presence of replication stress would also increase the rigor of the studies.

      Thank you for this suggestion. We agree that this would increase the rigor of the studies, but these experiments are currently beyond the scope of this manuscript.

      The authors show that in their patient data, some cell cycle regulators which are amplified in NSCLC at similar rates to CDK4/6, such as CCNE1, had no increase in FGA. Overexpressing CCNE1 and testing Osi response in their cell models would be a nice test of their proposed mechanism that it is the genomic instability and FGA that are driving resistance. This wouldn't need to be done in vivo, but could be done using cell culture-based methods.

      Thank you for this suggestion. For this manuscript, we have chosen to focus on the role of CDK4 and CDK6 in osimertinib resistance as they have shown the clearest correlation with decreased responsiveness to EGFR TKI treatment in other studies. We agree that assessing the role of CCNE1 in osimertinib resistance is important and will address this in future studies.

      Similarly, testing the overexpression of some of the proposed target genes, such as STEAP1 and AGR2, on the therapeutic response to Osi in cell culture would also be a nice test of the mechanism proposed.

      Thank you for this suggestion. We agree that these experiments are important, but they are currently beyond the scope of this study.

      Reviewer 2:

      Major Comments:

      (1) Reliance on overexpression models over loss-of-function

      The mechanistic studies mainly rely on overexpression of CDK4 and CDK6 to simulate the amplified state. Although the authors argued that the level of overexpression mimics that observed in resistant tumors, a complementary study in which CDK4/CDK6 were suppressed in a model where CDK4/CDK6 is amplified (such as HCC827 or TH116), and replication stress and osimertinib sensitivity tested would greatly strengthen their observations. Indeed, there is mention of CDK4 constructs to perform knockdown studies in the methods, but those studies are not included in this submission.

      Thank you for this suggestion. We agree that loss of function studies are an important complement to over expression studies. However, we have chosen to focus on the use of pharmacologic inhibitors of CDK4/6, which are more clinically relevant than knockdown studies. We demonstrate that the CDK4/6 inhibitor palbociclib is able to prevent replication stress and restore osimertinib sensitivity in CDK6 amplified TH116 patient-derived xenografts (Figure 5 and Supplementary Figure 10).

      (2) Mechanistic link between genomic instability and specific gene amplifications

      The authors highlight that CDK4/6 activation leads to recurrent copy number gains and transcriptional upregulation of specific pro-tumor genes including AGR2, ASNS, and STEAP1. While the paper establishes that CDK4/6 overexpression causes general genomic instability (increased FGA), it does not mechanistically explain why these specific genes are consistently amplified. The authors should investigate or discuss whether these specific loci are inherently fragile under replication stress, if they are direct downstream targets of the E2F transcriptional program, or if this is a result of random genomic instability followed by strong positive selection under osimertinib pressure.

      Thank you for this suggestion. We will provide additional discussion about whether these specific loci are likely to be inherently fragile under replication stress, if they are direct downstream targets of the E2F transcriptional program, or if it is more likely the result of random genomic instability followed by strong positive selection under osimertinib pressure.

      (3) Discrepancies in tumor mutational burden (TMB) reporting

      There is a slight contradiction regarding the TMB data that needs to be clarified for readers. The manuscript states that in the clinical datasets, "EGFR-mutant LUAD harboring cell cycle gene alterations exhibited significantly elevated FGA and TMB relative to cell cycle-negative tumors" (Line 265-266). However, in the next section, the authors say, "Notably, no corresponding increase in TMB was observed with CDK4 or CDK6 CNA, similar to our findings in preclinical models" (Line 271-273). The authors should clarify or discuss why broad cell cycle alterations correlate with high TMB, while CDK4/6-specific alterations drive structural instability (FGA) without increasing TMB.

      Thank you for this suggestion. We will clarify our findings and provide further discussion about why there may be a difference between the effects of broad cell cycle alterations and CDK4/6-specific alterations on TMB and structural genomic instability.

    1. eLife Assessment

      This study presents a potentially important structure-interaction model describing recognition of polymyxin antibiotics by the human transporter hPepT2 and applies these insights to guide the rational design of potent polymyxin B analogues with reduced nephrotoxicity. The conclusions are supported by compelling evidence, integrating molecular dynamics simulations, transporter mutagenesis, uptake assays, protein expression analyses, antibacterial testing, and mouse nephrotoxicity studies, although some mechanistic interpretations remain limited by altered transporter expression and the absence of direct structural validation.

    2. Reviewer #1 (Public review):

      Summary:

      Polymyxins are the last line of drugs to treat gram-negative bacteria-induced multi-drug resistance; however, they cause nephrotoxicity in 60% of patients. In this work, the authors have studied the structure-interaction relationship (SIR) of polymyxins with hPepT2 using computational and experimental methods. Moreover, it is observed that the electrostatic interactions coordinate the hPepT2-Polymyxin interactions; hence, an alanine scanning strategy is used to understand the interactions and derive the polymyxin variants.

      Computational methods such as molecular modeling, coarse-grained and all-atom MD simulations, and interaction studies are performed, while the results are validated in the mouse model, which is a great strategy to prove the hypothesis.

      Strengths:

      A clear understanding of the hPepT2-Polymyxin interactions and the role of electrostatic interactions is one of the very important strengths of the paper. In addition, this work proposes a great pipeline for using computational approaches and experimental validation methods to guide the development of newer antibiotics.

      Overall, the study proposes novel polymyxin analogues with reduced or no nephrotoxicity, thereby providing a promising foundation for the rational development of safer lipopeptide antibiotics.

      Weaknesses:

      This work is very well executed and presented; however, addressing the following concerns might improve the presentation of the work:

      (1) The introduction is well articulated; however, including a paragraph on the known inhibitors might be helpful in understanding the current status. In addition, it might also help to introduce Dabs, FADDI variants, Gly-sar and MIPS.

      (2) The following details of modeling with AlphaFold2 should be included: how the final structure was selected, what the RMSD and structure alignment of the template are, and the final selected structure. A section on modeling with all the parameter details might be useful for reproducing the structure. In addition, specify how the alanine scanning was performed alongside the structure prediction of polymyxins.

      (3) In the all-atom MD simulation method, detailing several parameters might help in reproducing the results: simulation time for each system, water model, system composition, protonation state, box type and dimensions, salt ions and concentration, membrane parameters and ligand parameterization methods. Also, the following details on energy minimization might be useful: minimization algorithm, number of steps for minimization and structure restraints in place.

      (4) On page 6, line 210, the MIC is used for the first time; although MIC is given in the abbreviation list, the first occurrence should have a complete name. A one-line explanation of MIC in the introduction or wherever suitable might be better but is not mandatory.

      (5) Similarly, Gly-sar is first mentioned on page 8, line 301, but its complete name is only mentioned later on page 10, line 368. This can be addressed if a short description is included in the introduction section.

      (6) For coarse-grained MD simulation, why were 2 replicates performed? Most studies perform 3 replicates, which are also good in terms of statistics and error bar calculations. In addition, the authors should specify whether an independent minimization is done for each of the two replicates or whether the minimization step is common for both.

      (7) For MD simulation results, giving simulation movies in supplementary results might be a better way to show how the trajectories behaved.

      (8) The description of visualisation software such as VMD or PyMol is missing. The authors should specify if any visualization tool is used.

      (9) For the mouse model study, the authors claim that FADDI-795 has no observable nephrotoxicity; however, the n=3 shows that a very small number of mouse models were used to make the assumption. In addition, the number of mice used in each experiment is not explicitly mentioned in the methods section.

      (10) In Table 2, the column 8 header is not visible.

    3. Reviewer #2 (Public review):

      Summary:

      Jiang et al. sought to elucidate the molecular basis of polymyxin antibiotic interaction with the renal transporter hPepT2, a transporter previously implicated in polymyxin-induced nephrotoxicity. They combined molecular dynamics simulations with transporter mutagenesis, functional uptake assays, kinetic analyses, protein expression studies, antibacterial susceptibility testing, and mouse nephrotoxicity experiments to develop a structure-interaction relationship (SIR) model and apply this model to the rational design of polymyxin analogues.

      Overall, the study represents a substantial multidisciplinary effort that integrates computational and experimental approaches. The identification of transporter residues involved in polymyxin recognition and the subsequent design of analogues with reduced hPepT2-mediated uptake provide a valuable framework for developing safer polymyxin antibiotics. In particular, the identification of FADDI-795 as an analogue that retains antibacterial activity while exhibiting reduced nephrotoxicity represents an encouraging proof of concept.

      Strengths:

      The computational predictions are strengthened by extensive experimental validation, including site-directed mutagenesis, transport kinetics, fluorescence uptake assays, membrane expression analyses, and in vivo toxicity studies. The consistency between multiple independent experimental approaches increases confidence in many of the authors' conclusions.

      Weaknesses:

      Several conclusions would benefit from a more cautious interpretation. A major limitation is that several transporter mutations substantially altered total or membrane protein expression, making it difficult to distinguish effects on substrate binding from indirect effects caused by impaired transporter stability or trafficking. The authors acknowledge this limitation in the Discussion, but some mechanistic conclusions remain stronger than the available evidence supports.

      Similarly, while the proposed binding model is biologically plausible and supported by mutagenesis, it remains an inferred model derived from molecular simulations rather than a direct structural determination. Statements describing the model as "validated" should therefore be moderated to indicate that the experimental data provide support rather than definitive structural confirmation.

      The translational implications are promising but remain preliminary. Although FADDI-795 demonstrated reduced nephrotoxicity in the mouse model while maintaining antibacterial activity, no pharmacokinetic studies were presented to demonstrate reduced renal accumulation or altered tissue distribution, and additional efficacy studies in infection models would further strengthen the therapeutic claims.

    4. Reviewer #3 (Public review):

      Summary:

      Jiang et al. described findings aimed at interrogating the interactions of the antibiotic polymyxin B with human kidney proteins that mediate nephrotoxicity. Their findings using both computational molecular dynamics simulations and experimental approaches illustrate the importance of aspartic acid residues (D215) in mediating the antibiotic uptake into the cells, and upon mutagenesis with Alanine, the effects are less pronounced. Further, they could modify the antibiotic units interacting with proteins into less toxic peptides with retained antibacterial properties.

      Strengths:

      I was impressed by this text, which advances the knowledge of how the antibiotic causes human nephrotoxicity and how this could be exploited into less problematic antibiotic peptides.

      Weaknesses:

      Interactions of Polymyxin B with kidney proteins were not demonstrable in vivo, and with reliable technologies such as X-ray or NMR.

    1. eLife Assessment

      This valuable study provides solid evidence that N6 methylation of A74 by METTL3 is required for efficient translation directed by the SARS-CoV-2 5′UTR in uninfected cells, likely by limiting access of protein factors to the 5′UTR that are necessary for efficient translation. These findings advance our understanding of the role of RNA modification in coronavirus replication and m6A-mediated regulation of gene expression more broadly, while further supporting METTL3 as a potential anticoronaviral therapeutic target. However, the evidence that m6A acts by destabilizing the third stem-loop (SL3) of the 5′UTR remains incomplete. The impact of the study would be strengthened by biochemical analyses directly assessing structural changes in the 5′UTR induced by A74 methylation.

    2. Reviewer #1 (Public review):

      Summary:

      A prevailing view is that translation of 5' capped mRNAs, i.e. mRNAs that are translated via ribosome scanning, is inhibited by highly structured 5' untranslated regions (5' UTRs). Despite having a common, structured 5' UTR, the mRNAs produced by the SARS-CoV-2 virus are efficiently translated. In this study, the authors identified a DRACH motif in stem-loop 3 (SL3), suggesting a potential site of m6A methylation of A74 by the enzyme METTL3. Given that such m6A modifications are known to disrupt RNA structure formation, the authors tested the hypothesis that this may be the basis underlying the efficient translation of these mRNAs. Mutational approaches complemented by METTL3 siRNA knockdown were employed to support this hypothesis. Additional experiments showed that this is required for efficient association of a reporter mRNA with polysomes (indicative of active translation), and suggest that the 5' UTR is more highly structured when methylation is abrogated.

      Strengths:

      The data clearly indicate that N6 methylation of A74 is required for efficient translation of SARS-CoV-2 mRNAs.

      Weaknesses:

      While the evidence supports the authors' central hypothesis, there are two issues that should be addressed. The first is that all of the approaches are indirect. All of the evidence for the presence of mRNA structural elements is based on computational and genetic analyses. We now know that there is something there, but we still do not know what it is. The authors need to use a biochemical approach to actually map the structural elements of the 5' UTR and determine how such structure(s) are changed by loss of methylation. The second hinges on the assumption that these mRNAs are translated via canonical ribosome scanning. RNA viruses are well-known to use a variety of other mechanisms, e.g. internal ribosome entry signals and ribosome tethering, to promote efficient translation. Alternatives to ribosome scanning should be considered.

    3. Reviewer #2 (Public review):

      The study addresses the conundrum of how the mRNAs of SARS-CoV-2 are efficiently translated since the 5' leader, which has common elements for all the viral genes, is highly structured. The authors test the hypothesis that m6A modification at position 74 is key to this translation. First, the authors show convincingly that this site is modified. Then, with extensive transfected reporter experiments using luciferase assays as well as sucrose gradient sedimentation, this modification is shown to be key for efficient translation. While the mechanism of this effect is not entirely clear (see comments/suggestions below), the authors show that it is independent of YTH "reader" proteins and likely involves altered interactions between SL3 (which contains A74) and downstream elements in the UTR. These results are important because they both offer insight into the function of m6A in gene expression and suggest how they may be important for the translation of viral mRNA in particular.

      While the data on their own make the overall case that the m6A modification in the 5'UTR of the viral genes is important to their expression, there are several things worth considering that could refine the model and make it more convincing.

      It is not clear whether putative uORF translation, particularly translation of the uORF that begins with a CUG codon at position 59 in the 5'UTR (as shown in Finkel et al., Nature 2020), would be impacted by this modification (as it includes the putative m6A site at position 74). It is also worth considering whether SL3 melting by translation of this uORF would alter the proposed mechanism.

      It is a bit unclear why the A74T mutant was put in the longer construct while the C75G mutant was put in a shorter construct. While not essential, the mechanistic arguments would be stronger if the same construct had been used to compare the mutations.

      While the authors show that "global depletion of m6A modification does not grossly alter translation efficiency" in a general sense (page 11), it would be of interest to know whether any host mRNAs with 5'UTR m6A (i.e., ACTA2 and COX8A, mentioned in this study) are affected by the mechanism here (i.e. run them in the luciferase assay).

      The authors show that the YTH "reader" proteins have a very small inhibitory effect (1.6-fold) on the translation of the viral mRNA with 5'UTR m6A. However, it remains unclear how important this is or whether it is generally true for host mRNAs with this modification.

      It is reassuring to see controls for changes in RNA levels in the supplemental material. The RNAs were generally stable under the experimental parameters explored, which would rule out RNA-decay-based mechanisms of m6A regulation. However, it should be noted that mRNA level experiments appear to have been done at 24 h while luciferase measurements were done at 48 h (as noted on p. 23, gene expression vs luciferase activity). It is not clear whether any RNA decay phenotypes would be apparent at 24 h.

      The authors use the term "ribosome profiling" (for example, on page 10), but it would appear the experiment performed is actually "polysome profiling" or "sucrose gradient sedimentation" since it did not involve ribosome footprinting.

    4. Reviewer #3 (Public review):

      Aly et al investigate the potential for a single N6-methyladenosine RNA modification in the context of the 5' UTR sequence of SARS-CoV-2 to regulate translation of a downstream luciferase reporter transfected into cells. They show using meRIP (m6A RNA IP) that this site is methylated in the plasmid-driven transcript, and convincingly show it mediates reporter translational efficiency using knockdown of the m6A methyltransferase METTL3 and mutation of the modified UTR site together with analysis of the transcript's association with polyribosomes. They suggest that the benefit to translation conferred by the modification is through its effect on the secondary structure of the 5' UTR, based on an RT-PCR-based assay in control and METTL3 knockdown cells linking RT processivity to translation (luciferase) output. They also extend their conclusions to two cellular mRNA 5' UTRs, also reported to contain a single m6A modification, and show METTL3-dependent changes in RNA structure stability, hinting at a broader significance of this mechanism of m6A control of gene expression.

      The conclusions of the paper are mostly well supported by the data presented, though validation of knockdown of METTL3 (and reader proteins) is absent.

      A major limitation of the work is the exclusive use of the reductionist artificial reporter system in uninfected cells. Though the 5' UTR site they identify is methylated in the context of a transcript generated in the nucleus (where the m6A installing complex is mainly localized, and believed to act exclusively in uninfected cells), how frequently this site is modified, if at all, on viral RNAs generated within cytoplasmic membrane-bound replication organelles. Similarly, whether the translation regulation by a single m6A modification identified here occurs within the context of an infected cell, in which there are many changes to the RNA and translational regulatory landscape, also remains to be tested.

      How this work can be reconciled with others that have concluded either little potential for translational regulation by 5' UTR modification (Guca et al 2024; PMID: 38244546) or that an eIF3-mediated mechanism is responsible (Meyer et al, 2015 PMID: 26593424) is not addressed in the discussion.

    1. eLife assessment

      In this manuscript, Rademacher and colleagues examined the effect of a chemogenetic approach on the integrity of the dopamine system in mice with chronically stimulating dopamine neurons. These findings are important: (1) This approach led to an axon-first degeneration over a time course (2–4 weeks) that is suitable for experimental investigation; (2) The finding that direct excitation of dopaminergic neurons causes differential degeneration sheds light on dopaminergic neuron selective vulnerability mechanisms. Overall, the strength of the evidence is solid, but the behavior experiments that do not include a CNO control provide incomplete support for the findings.

    2. Reviewer #1 (Public Review):

      Summary:

      In this manuscript, the authors investigated the effect of chronic activation of dopamine neurons using chemogenetics. Using Gq-DREADDs, the authors chronically activated midbrain dopamine neurons and observed that these neurons, particularly their axons, exhibit increased vulnerability and degeneration, resembling the pathological symptoms of Parkinson's disease. Baseline calcium levels in midbrain dopamine neurons were also significantly elevated following the chronic activation. Lastly, to identify cellular and circuit-level changes in response to dopaminergic neuronal degeneration caused by chronic activation, the authors employed spatial genomics (Visium) and revealed comprehensive changes in gene expression in the mouse model subjected to chronic activation. In conclusion, this study presents novel data on the consequences of chronic hyperactivation of midbrain dopamine neurons.

      Strengths:

      This study provides direct evidence that the chronic activation of dopamine neurons is toxic and gives rise to neurodegeneration. In addition, the authors achieved the chronic activation of dopamine neurons using water application of clozapine-N-oxide (CNO), a method not commonly employed by researchers. This approach may offer new insights into pathophysiological alterations of dopamine neurons in Parkinson's disease. The authors also utilized state-of-the-art spatial gene expression analysis, which can provide valuable information for other researchers studying dopamine neurons. Although the authors did not elucidate the mechanisms underlying dopaminergic neuronal and axonal death, they presented a substantial number of intriguing ideas in their discussion, which are worth further investigation.

      Weaknesses:

      Many claims raised in this paper are only partially supported by the experimental results. So, additional data are necessary to strengthen the claims. The effects of chronic activation of dopamine neurons are intriguing; however, this paper does not go beyond reporting phenomena. It lacks a comprehensive explanation for the degeneration of dopamine neurons and their axons. While the authors proposed possible mechanisms for the degeneration in their discussion, such as differentially expressed genes, these remain experimentally unexplored.

    3. Reviewer #2 (Public Review):<br /> <br /> Summary:

      Rademacher et al. present a paper showing that chronic chemogenetic excitation of dopaminergic neurons in the mouse midbrain results in differential degeneration of axons and somas across distinct regions (SNc vs VTA). These findings are important. This mouse model also has the advantage of showing a axon-first degeneration over an experimentally-useful time course (2-4 weeks). 2. The findings that direct excitation of dopaminergic neurons causes differential degeneration sheds light on the mechanisms of dopaminergic neuron selective vulnerability. The evidence that activation of dopaminergic neurons causes degeneration and alters mRNA expression is convincing, as the authors use both vehicle and CNO control groups, but the evidence that chronic dopaminergic activation alters circadian rhythm and motor behavior is incomplete as the authors did not run a CNO-control condition in these experiments.

      Strengths:<br /> This is an exciting and important paper.<br /> The paper compares mouse transcriptomics with human patient data.<br /> It shows that selective degeneration can occur across the midbrain dopaminergic neurons even in the absence of a genetic, prion, or toxin neurodegeneration mechanism.

      Weaknesses:

      Major concerns:

      (1) The lack of a CNO-positive, DREADD-negative control group in the behavioral experiments is the main limitation in interpreting the behavioral data. Without knowing whether CNO on its own has an impact on circadian rhythm or motor activity, the certainty that dopaminergic hyperactivity is causing these effects is lacking.

      (2) One of the most exciting things about this paper is that the SNc degenerates more strongly than the VTA when both regions are, in theory, excited to the same extent. However, it is not perfectly clear that both regions respond to CNO to the same extent. The electrophysiological data showing CNO responsiveness is only conducted in the SNc. If the VTA response is significantly reduced vs the SNc response, then the selectivity of the SNc degeneration could just be because the SNc was more hyperactive than the VTA. Electrophysiology experiments comparing the VTA and SNc response to CNO could support the idea that the SNc has substantial intrinsic vulnerability factors compared to the VTA.

      (3) The mice have access to a running wheel for the circadian rhythm experiments. Running has been shown to alter the dopaminergic system (Bastioli et al., 2022) and so the authors should clarify whether the histology, electrophysiology, fiber photometry, and transcriptomics data are conducted on mice that have been running or sedentary.

    4. Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Rademacher and colleagues examined the effect on the integrity of the dopamine system in mice of chronically stimulating dopamine neurons using a chemogenetic approach. They find that one to two weeks of constant exposure to the chemogenetic activator CNO leads to a decrease in the density of tyrosine hydroxylase staining in striatal brain sections and to a small reduction of the global population of tyrosine hydroxylase positive neurons in the ventral midbrain. They also report alterations in gene expression in both regions using a spatial transcriptomics approach. Globally, the work is well done and valuable and some of the conclusions are interesting. However, the conceptual advance is perhaps a bit limited in the sense that there is extensive previous work in the literature showing that excessive depolarization of multiple types of neurons associated with intracellular calcium elevations promotes neuronal degeneration. The present work adds to this by showing evidence of a similar phenomenon in dopamine neurons. In terms of the mechanisms explaining the neuronal loss observed after 2 to 4 weeks of chemogenetic activation, it would be important to consider that dopamine neurons are known from a lot of previous literature to undergo a decrease in firing through a depolarization-block mechanism when chronically depolarized. Is it possible that such a phenomenon explains much of the results observed in the present study? It would be important to consider this in the manuscript. The relevance to Parkinson's disease (PD) is also not totally clear because there is not a lot of previous solid evidence showing that the firing of dopamine neurons is increased in PD, either in human subjects or in mouse models of the disease. As such, it is not clear if the present work is really modelling something that could happen in PD in humans.

      Comments on the introduction:

      The introduction cites a 1990 paper from the lab of Anthony Grace as support of the fact that DA neurons increase their firing rate in PD models. However, in this 1990 paper, the authors stated that: "With respect to DA cell activity, depletions of up to 96% of striatal DA did not result in substantial alterations in the proportion of DA neurons active, their mean firing rate, or their firing pattern. Increases in these parameters only occurred when striatal DA depletions exceeded 96%." Such results argue that an increase in firing rate is most likely to be a consequence of the almost complete loss of dopamine neurons rather than an initial driver of neuronal loss. The present introduction would thus benefit from being revised to clarify the overriding hypothesis and rationale in relation to PD and better represent the findings of the paper by Hollerman and Grace.

      It would be good that the introduction refers to some of the literature on the links between excessive neuronal activity, calcium, and neurodegeneration. There is a large literature on this and referring to it would help frame the work and its novelty in a broader context.

      Comments on the results section:

      The running wheel results of Figure 1 suggest that the CNO treatment caused a brief increase in running on the first day after which there was a strong decrease during the subsequent days in the active phase. This observation is also in line with the appearance of a depolarization block.

      The authors examined many basic electrophysiological parameters of recorded dopamine neurons in acute brain slices. However, it is surprising that they did not report the resting membrane potential, or the input resistance. It would be important that this be added because these two parameters provide key information on the basal excitability of the recorded neurons. They would also allow us to obtain insight into the possibility that the neurons are chronically depolarized and thus in depolarization block.

      It is great that the authors quantified not only TH levels but also the levels of mCherry, co-expressed with the chemogenetic receptor. This could in principle help to distinguish between TH downregulation and true loss of dopamine neuron cell bodies. However, the approach used here has a major caveat in that the number of mCherry-positive dopamine neurons depends on the proportion of dopamine neurons that were infected and expressed the DREADD and this could very well vary between different mice. It is very unlikely that the virus injection allowed to infect 100% of the neurons in the VTA and SNc. This could for example explain in part the mismatch between the number of VTA dopamine neurons counted in panel 2G when comparing TH and mCherry counts. Also, I see that the mCherry counts were not provided at the 2-week time point. If the mCherry had been expressed genetically by crossing the DAT-Cre mice with a floxed fluorescent reported mice, the interpretation would have been simpler. In this context, I am not convinced of the benefit of the mCherry quantifications. The authors should consider either removing these results from the final manuscript or discussing this important limitation.

      Although the authors conclude that there is a global decrease in the number of dopamine neurons after 4 weeks of CNO treatment, the post-hoc tests failed to confirm that the decrease in dopamine number was significant in the SNc, the region most relevant to Parkinson's. This could be due to the fact that only a small number of mice were tested. A "n" of just 4 or 5 mice is very small for a stereological counting experiment. As such, this experiment was clearly underpowered at the statistical level. Also, the choice of the image used to illustrate this in panel 2G should be reconsidered: the image suggests that a very large loss of dopamine neurons occurred in the SNc and this is not what the numbers show. A more representative image should be used.

      In Figure 3, the authors attempt to compare intracellular calcium levels in dopamine neurons using GCaMP6 fluorescence. Because this calcium indicator is not quantitative (unlike ratiometric sensors such as Fura2), it is usually used to quantify relative changes in intracellular calcium. The present use of this probe to compare absolute values is unusual and the validity of this approach is unclear. This limitation needs to be discussed. The authors also need to refer in the text to the difference between panels D and E of this figure. It is surprising that the fluctuations in calcium levels were not quantified. I guess the hypothesis was that there should be more or larger fluctuations in the mice treated with CNO if the CNO treatment led to increased firing. This needs to be clarified.

      Although the spatial transcriptomic results are intriguing and certainly a great way to start thinking about how the CNO treatment could lead to the loss of dopamine neurons, the presented results, the focussing of some broad classes of differentially expressed genes and on some specific examples, do not really suggest any clear mechanism of neurodegeneration. It would perhaps be useful for the authors to use the obtained data to validate that a state of chronic depolarization was indeed induced by the chronic CNO treatment. Were genes classically linked to increased activity like cfos or bdnf elevated in the SNc or VTA dopamine neurons? In the striatum, the authors report that the levels of DARP32, a gene whose levels are linked to dopamine levels, are unchanged. Does this mean that there were no major changes in dopamine levels in the striatum of these mice?

      The usefulness of comparing the transcriptome of human PD SNc or VTA sections to that of the present mouse model should be better explained. In the human tissues, the transcriptome reflects the state of the tissue many years after extensive loss of dopamine neurons. It is expected that there will be few if any SNc neurons left in such sections. In comparison, the mice after 7 days of CNO treatment do not appear to have lost any dopamine neurons. As such, how can the two extremely different conditions be reasonably compared?

      Comments on the discussion:

      In the discussion, the authors state that their calcium photometry results support a central role of calcium in activity-induced neurodegeneration. This conclusion, although plausible because of the very broad pre-existing literature linking calcium elevation (such as in excitotoxicity) to neuronal loss, should be toned down a bit as no causal relationship was established in the experiments that were carried out in the present study.

      In the discussion, the authors discuss some of the parallel changes in gene expression detected in the mouse model and in the human tissues. Because few if any dopamine neurons are expected to remain in the SNc of the human tissues used, this sort of comparison has important conceptual limitations and these need to be clearly addressed.

      A major limitation of the present discussion is that it does not discuss the possibility that the observed phenotypes are caused by the induction of a chronic state of depolarization block by the chronic CNO treatment. I encourage the authors to consider and discuss this hypothesis. Also, the authors need to discuss the fact that previous work was only able to detect an increase in the firing rate of dopamine neurons after more than 95% loss of dopamine neurons. As such, the authors need to clearly discuss the relevance of the present model to PD. Are changes in firing rate a driver of neuronal loss in PD, as the authors try to make the case here, or are such changes only a secondary consequence of extensive neuronal loss (for example because a major loss of dopamine would lead to reduced D2 autoreceptor activation in the remaining neurons, and to reduced autoreceptor-mediated negative feedback on firing). This needs to be discussed.

      There is a very large, multi-decade literature on calcium elevation and its effects on neuronal loss in many different types of neurons. The authors should discuss their findings in this context and refer to some of this previous work. In a nutshell, the observations of the present manuscript could be summarized by stating that the chronic membrane depolarization induced by the CNO treatment is likely to induce a chronic elevation of intracellular calcium and this is then likely to activate some of the well-known calcium-dependent cell death mechanisms. Whether such cell death is linked in any way to PD is not really demonstrated by the present results.

      The authors are encouraged to perform a thorough revision of the discussion to address all of these issues, discuss the major limitations of the present model, and refer to the broad pre-existing literature linking membrane depolarization, calcium, and neuronal loss in many neuronal cell types.

    5. Author response:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, the authors investigated the effect of chronic activation of dopamine neurons using chemogenetics. Using Gq-DREADDs, the authors chronically activated midbrain dopamine neurons and observed that these neurons, particularly their axons, exhibit increased vulnerability and degeneration, resembling the pathological symptoms of Parkinson's disease. Baseline calcium levels in midbrain dopamine neurons were also significantly elevated following the chronic activation. Lastly, to identify cellular and circuit-level changes in response to dopaminergic neuronal degeneration caused by chronic activation, the authors employed spatial genomics (Visium) and revealed comprehensive changes in gene expression in the mouse model subjected to chronic activation. In conclusion, this study presents novel data on the consequences of chronic hyperactivation of midbrain dopamine neurons.

      Strengths:

      This study provides direct evidence that the chronic activation of dopamine neurons is toxic and gives rise to neurodegeneration. In addition, the authors achieved the chronic activation of dopamine neurons using water application of clozapine-N-oxide (CNO), a method not commonly employed by researchers. This approach may offer new insights into pathophysiological alterations of dopamine neurons in Parkinson's disease. The authors also utilized state-of-the-art spatial gene expression analysis, which can provide valuable information for other researchers studying dopamine neurons. Although the authors did not elucidate the mechanisms underlying dopaminergic neuronal and axonal death, they presented a substantial number of intriguing ideas in their discussion, which are worth further investigation.

      We thank the reviewer for these positive comments.

      Weaknesses:

      Many claims raised in this paper are only partially supported by the experimental results. So, additional data are necessary to strengthen the claims. The effects of chronic activation of dopamine neurons are intriguing; however, this paper does not go beyond reporting phenomena. It lacks a comprehensive explanation for the degeneration of dopamine neurons and their axons. While the authors proposed possible mechanisms for the degeneration in their discussion, such as differentially expressed genes, these remain experimentally unexplored.

      We thank the reviewer for this review. We do believe that the manuscript has a mechanistic component, as the central experiments involve direct manipulation of neuronal activity, and we show an increase in calcium levels and gene expression changes in dopamine neurons that coincide with the degeneration. However, we agree that deeper mechanistic investigation would strengthen the conclusions of the paper. We have planned several important revisions, including the addition of CNO behavioral controls, manipulation of intracellular calcium using isradipine, additional transcriptomics experiments and further validation of findings. We anticipate that these additions will significantly bolster the conclusions of the paper.

      Reviewer #2 (Public Review):

      Summary:

      Rademacher et al. present a paper showing that chronic chemogenetic excitation of dopaminergic neurons in the mouse midbrain results in differential degeneration of axons and somas across distinct regions (SNc vs VTA). These findings are important. This mouse model also has the advantage of showing a axon-first degeneration over an experimentally-useful time course (2-4 weeks). 2. The findings that direct excitation of dopaminergic neurons causes differential degeneration sheds light on the mechanisms of dopaminergic neuron selective vulnerability. The evidence that activation of dopaminergic neurons causes degeneration and alters mRNA expression is convincing, as the authors use both vehicle and CNO control groups, but the evidence that chronic dopaminergic activation alters circadian rhythm and motor behavior is incomplete as the authors did not run a CNO-control condition in these experiments.

      Strengths:

      This is an exciting and important paper.

      The paper compares mouse transcriptomics with human patient data.

      It shows that selective degeneration can occur across the midbrain dopaminergic neurons even in the absence of a genetic, prion, or toxin neurodegeneration mechanism.

      We thank the reviewer for these insightful comments.

      Weaknesses:

      Major concerns:

      (1) The lack of a CNO-positive, DREADD-negative control group in the behavioral experiments is the main limitation in interpreting the behavioral data. Without knowing whether CNO on its own has an impact on circadian rhythm or motor activity, the certainty that dopaminergic hyperactivity is causing these effects is lacking.

      This is an important point. Although we show that CNO does not produce degeneration of DA neuron terminals, we do not exclude a contribution to the behavioral changes. We agree that this behavioral control is necessary, and will address it in revision with a CNO-only running wheel cohort.

      (2) One of the most exciting things about this paper is that the SNc degenerates more strongly than the VTA when both regions are, in theory, excited to the same extent. However, it is not perfectly clear that both regions respond to CNO to the same extent. The electrophysiological data showing CNO responsiveness is only conducted in the SNc. If the VTA response is significantly reduced vs the SNc response, then the selectivity of the SNc degeneration could just be because the SNc was more hyperactive than the VTA. Electrophysiology experiments comparing the VTA and SNc response to CNO could support the idea that the SNc has substantial intrinsic vulnerability factors compared to the VTA.

      We agree that additional electrophysiology conducted in the VTA dopamine neurons would meaningfully add to our understanding of the selective vulnerability in this model, and will complete these experiments in revision.

      (3) The mice have access to a running wheel for the circadian rhythm experiments. Running has been shown to alter the dopaminergic system (Bastioli et al., 2022) and so the authors should clarify whether the histology, electrophysiology, fiber photometry, and transcriptomics data are conducted on mice that have been running or sedentary.

      We will explicitly clarify which mice had access to a running wheel in our revision. Briefly, mice for histology, electrophysiology, and transcriptomics all had access to a running wheel during their treatment. The mice used for photometry underwent about 7 days of running wheel access approximately 3 weeks prior to the beginning of the experiment. The photometry headcaps sterically prevented mice from having access to a running wheel in their home cage.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Rademacher and colleagues examined the effect on the integrity of the dopamine system in mice of chronically stimulating dopamine neurons using a chemogenetic approach. They find that one to two weeks of constant exposure to the chemogenetic activator CNO leads to a decrease in the density of tyrosine hydroxylase staining in striatal brain sections and to a small reduction of the global population of tyrosine hydroxylase positive neurons in the ventral midbrain. They also report alterations in gene expression in both regions using a spatial transcriptomics approach. Globally, the work is well done and valuable and some of the conclusions are interesting. However, the conceptual advance is perhaps a bit limited in the sense that there is extensive previous work in the literature showing that excessive depolarization of multiple types of neurons associated with intracellular calcium elevations promotes neuronal degeneration. The present work adds to this by showing evidence of a similar phenomenon in dopamine neurons.

      We thank the reviewer for the careful and thoughtful review of our manuscript.

      While extensive depolarization and associated intracellular calcium elevations promotes degeneration generally, we emphasize that the process we describe is novel. Indeed, prior studies delivering chronic DREADDs to vulnerable neurons in models of Alzheimer’s disease did not report an increase in neurodegeneration, despite seeing changes in protein aggregation (e.g. Yuan and Grutzendler, J Neurosci 2016, PMID: 26758850; Hussaini et al., PLOS Bio 2020, PMID: 32822389). Further, a critical finding from our study is that in our paradigm, this stressor does not impact all dopamine neurons equally, as the SNc DA neurons are more vulnerable than the VTA, mirroring selective vulnerability characteristic of Parkinson’s disease. This is consistent with a large body of literature that SNc dopamine neurons are less capable of handling large energetic and calcium loads compared to neighboring VTA neurons, and the finding that chronically altered activity is sufficient to drive this preferential loss is novel.

      In addition, we are not aware of prior studies that have chronically activated DREADDs to produce neurodegeneration. Other studies have shown that acute excitotoxic stressors can produce neuronal degeneration, but the chronic increase in activity is central to our approach.

      In terms of the mechanisms explaining the neuronal loss observed after 2 to 4 weeks of chemogenetic activation, it would be important to consider that dopamine neurons are known from a lot of previous literature to undergo a decrease in firing through a depolarization-block mechanism when chronically depolarized. Is it possible that such a phenomenon explains much of the results observed in the present study? It would be important to consider this in the manuscript.

      As discussed in greater detail in the results section below, our data suggests this may not be a prominent feature in our model. However, we cannot rule out a contribution of depolarization block, and will expand on the discussion of this possibility in the revised manuscript.

      The relevance to Parkinson's disease (PD) is also not totally clear because there is not a lot of previous solid evidence showing that the firing of dopamine neurons is increased in PD, either in human subjects or in mouse models of the disease. As such, it is not clear if the present work is really modelling something that could happen in PD in humans.

      We completely agree that evidence of increased dopamine neuron activity from human PD patients is lacking and the existing data are difficult to interpret without human controls. However, as we outline in the manuscript, multiple lines of evidence suggest that the activity level of dopamine neurons almost certainly does change in PD. Therefore, it is very important that we understand how changes in the level of neural activity influence the degeneration of DA neurons. In this paper we examine the impact of increased activity. Increased activity may be compensatory after initial dopamine neuron loss, or may be an initial driver of death (Rademacher & Nakamura, Exp Neurol 2024, PMID: 38092187). Beyond what is already discussed in the manuscript, additional support for increased activity in PD models include:

      - Elevated firing rates in asymptomatic MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488)

      - Increased frequency of spontaneous firing in patient-derived iPSC dopamine neurons and primary mouse dopamine neurons that overexpress synuclein (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060)

      - Increased spontaneous firing in dopamine neurons of rats injected with synuclein preformed fibrils compared to sham (Tozzi et al., Brain 2021, PMID: 34297092)

      We will include and further discuss these important examples in our revision.

      Similarly, in future studies, it will also be important to study the impact of decreasing DA neuron activity. There will be additional levels of complexity to accurately model changes in PD, which may differ between subtypes of the disease, the disease stage, and the subtype of dopamine neuron. Our study models the possibility of chronically increased pacemaking, and interpretation of our results will be informed as we learn more about how the activity of DA neurons changes in humans in PD. We will discuss and elaborate on these important points in the revision.

      Comments on the introduction:

      The introduction cites a 1990 paper from the lab of Anthony Grace as support of the fact that DA neurons increase their firing rate in PD models. However, in this 1990 paper, the authors stated that: "With respect to DA cell activity, depletions of up to 96% of striatal DA did not result in substantial alterations in the proportion of DA neurons active, their mean firing rate, or their firing pattern. Increases in these parameters only occurred when striatal DA depletions exceeded 96%." Such results argue that an increase in firing rate is most likely to be a consequence of the almost complete loss of dopamine neurons rather than an initial driver of neuronal loss. The present introduction would thus benefit from being revised to clarify the overriding hypothesis and rationale in relation to PD and better represent the findings of the paper by Hollerman and Grace.

      We agree that the findings of Hollerman and Grace support compensatory changes in dopamine neuron activity in response to loss of dopamine neurons, rather than informing whether dopamine neuron loss can also be an initial driver of activity. We will clarify this point in our revision. In addition, the results of other studies on this point are mixed: a 50% reduction in dopamine neurons didn’t alter firing rate or bursting (Harden and Grace, J Neurosci 1995, PMID: 7666198; Bilbao et al, Brain Res 2006, PMID: 16574080), while a 40% loss was found to increase firing rate and bursting (Chen et al, Brain Res 2009. PMID: 19545547) and larger reductions alter burst firing (Hollerman & Grace, Brain Res 1990, PMID: 2126975; Stachowiak et al, J Neurosci 1987, PMID: 3110381). Importantly, even if compensatory, such late-stage increases in dopamine neuron activity may contribute to disease progression and drive a vicious cycle of degeneration in surviving neurons. In addition, we also don’t know how the threshold of dopamine neuron loss and altered activity may differ between mice and humans, and PD patients do not present with clinical symptoms until ~30-60% of nigral neurons are lost (Burke & O’Malley, Exp Neurol 2013, PMID: 22285449; Shulman et al, Annu Rev Pathol 2011, PMID: 21034221).

      Other lines of evidence support the potential role of hyperactivity in disease initiation, including increased activity before dopamine neuron loss in MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488), increased spontaneous firing in patient-derived iPSC dopamine neurons (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060), and increased activity observed in genetic models of PD (Bishop et al., J Neurophysiol 2010, PMID: 20926611; Regoni et al., Cell Death Dis 2020,  PMID: 33173027).

      It would be good that the introduction refers to some of the literature on the links between excessive neuronal activity, calcium, and neurodegeneration. There is a large literature on this and referring to it would help frame the work and its novelty in a broader context.

      We agree that a discussion of hyperactivity, calcium, and neurodegeneration would benefit the introduction. While we briefly discuss calcium and neurodegeneration in the discussion, we will expand on this literature in both the introduction and discussion sections. We will carefully review and contextualize our work within existing frameworks of calcium and neurodegeneration (e.g. Surmeier & Schumacker, J Biol Chem 2013, PMID: 23086948; Verma et al., Transl Neurodegener 2022, PMID: 35078537). We believe that the novelty of our study lies in 1) a chronic chemogenetic activation paradigm via drinking water, 2) demonstrating selective vulnerability of dopamine neurons as a result of altering their activity/excitability alone, and 3) comparing mouse and human spatial transcriptomics.

      Comments on the results section:

      The running wheel results of Figure 1 suggest that the CNO treatment caused a brief increase in running on the first day after which there was a strong decrease during the subsequent days in the active phase. This observation is also in line with the appearance of a depolarization block.

      The authors examined many basic electrophysiological parameters of recorded dopamine neurons in acute brain slices. However, it is surprising that they did not report the resting membrane potential, or the input resistance. It would be important that this be added because these two parameters provide key information on the basal excitability of the recorded neurons. They would also allow us to obtain insight into the possibility that the neurons are chronically depolarized and thus in depolarization block.

      We do report the input resistance in Supplemental Figure 1C, which was unchanged in CNO-treated animals compared to controls. We did not report the resting membrane potential because many of the DA neurons were spontaneously firing. However, we will report the initial membrane potential on first breaking into the cell for the whole cell recordings in the revision, which did not vary between groups. This is still influenced by action potential activity, but is the timepoint in the recording least impacted by dialyzing of the neuron by the internal solution. We observed increased spontaneous action potential activity ex vivo in slices from CNO-treated mice (Figure 1D), thus at least under these conditions these dopamine neurons are not in depolarization block. We also did not see strong evidence of changes in other intrinsic properties of the neurons with whole cell recordings (e.g. Figure S1C). Overall, our electrophysiology experiments are not consistent with the depolarization block model, at least not due to changes in the intrinsic properties of the neurons. Although our ex vivo findings cannot exclude a contribution of depolarization block in vivo, we do show that CNO-treated mice removed from their cages for open field testing continue to have a strong trend for increased activity for approximately 10 days (S1E).  This finding is also consistent with increased activity of the DA neurons. We will add discussion of these important considerations in the revision.

      It is great that the authors quantified not only TH levels but also the levels of mCherry, co-expressed with the chemogenetic receptor. This could in principle help to distinguish between TH downregulation and true loss of dopamine neuron cell bodies. However, the approach used here has a major caveat in that the number of mCherry-positive dopamine neurons depends on the proportion of dopamine neurons that were infected and expressed the DREADD and this could very well vary between different mice. It is very unlikely that the virus injection allowed to infect 100% of the neurons in the VTA and SNc. This could for example explain in part the mismatch between the number of VTA dopamine neurons counted in panel 2G when comparing TH and mCherry counts. Also, I see that the mCherry counts were not provided at the 2-week time point. If the mCherry had been expressed genetically by crossing the DAT-Cre mice with a floxed fluorescent reported mice, the interpretation would have been simpler. In this context, I am not convinced of the benefit of the mCherry quantifications. The authors should consider either removing these results from the final manuscript or discussing this important limitation.

      We thank the reviewer for this insightful comment, and we agree that this is a caveat of our mCherry quantification. Quantitation of the number of mCherry+ DA neurons specifically informs the impact on transduced DA neurons, and mCherry appears to be less susceptible to downregulation versus TH. As the reviewer points out, it carries the caveat that there is some variability between injections. Nonetheless, we believe that it conveys useful complementary data. As suggested, we will discuss this caveat in our revision. Note that mCherry was not quantified at the two-week timepoint because there is no loss of TH+ cells at that time.

      Although the authors conclude that there is a global decrease in the number of dopamine neurons after 4 weeks of CNO treatment, the post-hoc tests failed to confirm that the decrease in dopamine number was significant in the SNc, the region most relevant to Parkinson's. This could be due to the fact that only a small number of mice were tested. A "n" of just 4 or 5 mice is very small for a stereological counting experiment. As such, this experiment was clearly underpowered at the statistical level. Also, the choice of the image used to illustrate this in panel 2G should be reconsidered: the image suggests that a very large loss of dopamine neurons occurred in the SNc and this is not what the numbers show. A more representative image should be used.

      We agree that the stereology experiments were performed on relatively small numbers of animals. Combined with the small effect size, this may have contributed to the post-hoc tests showing a trend of p=0.1 for both the TH and mCherry dopamine cell counts in the SN at 4 weeks. As part of the planned experiments for our revision, we will perform an additional stereologic analysis to further assess the loss of SNc dopamine neurons. We will also review and ensure the images are representative.

      In Figure 3, the authors attempt to compare intracellular calcium levels in dopamine neurons using GCaMP6 fluorescence. Because this calcium indicator is not quantitative (unlike ratiometric sensors such as Fura2), it is usually used to quantify relative changes in intracellular calcium. The present use of this probe to compare absolute values is unusual and the validity of this approach is unclear. This limitation needs to be discussed. The authors also need to refer in the text to the difference between panels D and E of this figure. It is surprising that the fluctuations in calcium levels were not quantified. I guess the hypothesis was that there should be more or larger fluctuations in the mice treated with CNO if the CNO treatment led to increased firing. This needs to be clarified.

      We thank the reviewer for this comment. We understand that this method of comparing absolute values is unconventional. However, these animals were tested concurrently on the same system, and a clear effect on the absolute baseline was observed. We will include a caveat of this in our discussion. Panel D of this figure shows the raw, uncorrected photometry traces, whereas panel E shows the isosbestic corrected traces for the same recording. In panel E, the traces follow time in ascending order. We will also include frequency and amplitude data for these recordings.   

      Although the spatial transcriptomic results are intriguing and certainly a great way to start thinking about how the CNO treatment could lead to the loss of dopamine neurons, the presented results, the focusing of some broad classes of differentially expressed genes and on some specific examples, do not really suggest any clear mechanism of neurodegeneration. It would perhaps be useful for the authors to use the obtained data to validate that a state of chronic depolarization was indeed induced by the chronic CNO treatment. Were genes classically linked to increased activity like cfos or bdnf elevated in the SNc or VTA dopamine neurons? In the striatum, the authors report that the levels of DARP32, a gene whose levels are linked to dopamine levels, are unchanged. Does this mean that there were no major changes in dopamine levels in the striatum of these mice?

      We will review the expression of activity-related genes in our dataset, although we must keep in mind that these genes may behave differently in the context of chronic activation as opposed to acutely increased activity. We will also include experiments assessing striatal dopamine levels by HPLC in the revision.

      The usefulness of comparing the transcriptome of human PD SNc or VTA sections to that of the present mouse model should be better explained. In the human tissues, the transcriptome reflects the state of the tissue many years after extensive loss of dopamine neurons. It is expected that there will be few if any SNc neurons left in such sections. In comparison, the mice after 7 days of CNO treatment do not appear to have lost any dopamine neurons. As such, how can the two extremely different conditions be reasonably compared?

      Our mouse model and human PD progress over distinct timescales, as is the case with essentially all mouse models of neurodegenerative diseases. Nonetheless, in our view there is still great value in comparing gene expression changes in mouse models with those in human disease. It seems very likely that the same pathologic processes that drive degeneration early in the disease continue to drive degeneration later in the disease. Note that we have tried to address the discrepancy in time scales in part by comparing to early PD samples when there is more limited SNc DA neuron loss. Please note the numbers of DA neurons within the areas we have selected for sampling (Figure at right). Therefore, we can indeed use spatial transcriptomics to compare dopamine neurons from mice with initial degeneration and patients where degeneration is ongoing during their disease.

      Author response image 1.

      Violin plot of DA neuron proportions sampled within the vulnerable SNV (deconvoluted RCTD method used in unmasked tissue sections of the SNV).

      Control and early PD subjects.

      Comments on the discussion:

      In the discussion, the authors state that their calcium photometry results support a central role of calcium in activity-induced neurodegeneration. This conclusion, although plausible because of the very broad pre-existing literature linking calcium elevation (such as in excitotoxicity) to neuronal loss, should be toned down a bit as no causal relationship was established in the experiments that were carried out in the present study.

      Our model utilizes hM3Dq-DREADDs that function by increasing intracellular calcium to increase neuronal excitability, and our results show increased Ca2+ by fiber photometry and changes to Ca2+-related genes, strongly suggesting a causal relation and crucial role of calcium in the mechanism of degeneration. However, we agree that we have not experimentally proven this point, as we acknowledged in the text. Additionally, we have planned revision experiments involving chronic isradipine treatment to further test the role of calcium in the mechanism of degeneration in this model.

      In the discussion, the authors discuss some of the parallel changes in gene expression detected in the mouse model and in the human tissues. Because few if any dopamine neurons are expected to remain in the SNc of the human tissues used, this sort of comparison has important conceptual limitations and these need to be clearly addressed.

      As discussed, we can sample SN DA neurons in early PD (see figure above), and in our view there is great value for such comparisons. We agree that discussion of appropriate caveats is warranted and this will be clearly addressed in the revision.

      A major limitation of the present discussion is that it does not discuss the possibility that the observed phenotypes are caused by the induction of a chronic state of depolarization block by the chronic CNO treatment. I encourage the authors to consider and discuss this hypothesis.

      As discussed above, our analyses of DA neuron firing in slices and open field testing to date do not support a prominent contribution of depolarization block with chronic CNO treatment. However, we cannot rule out this hypothesis, therefore we will include additional electrophysiology experiments and add discussion of this important consideration.  

      Also, the authors need to discuss the fact that previous work was only able to detect an increase in the firing rate of dopamine neurons after more than 95% loss of dopamine neurons. As such, the authors need to clearly discuss the relevance of the present model to PD. Are changes in firing rate a driver of neuronal loss in PD, as the authors try to make the case here, or are such changes only a secondary consequence of extensive neuronal loss (for example because a major loss of dopamine would lead to reduced D2 autoreceptor activation in the remaining neurons, and to reduced autoreceptor-mediated negative feedback on firing). This needs to be discussed.

      As discussed above, while increases in dopamine neuron activity may be compensatory after loss of neurons, the precise percentage required to induce such compensatory changes is not defined in mice and varies between paradigms, and the threshold level is not known in humans. We also reiterate that a compensatory increase in activity could still promote the degeneration of critical surviving DA neurons, whose loss underlies the substantial decline in motor function that typically occurs over the course of PD. Moreover, there are also multiple lines of evidence to suggest that changes in activity can initiate and drive dopamine neuron degeneration (Rademacher & Nakamura, Exp Neurol 2024). For example, overexpression of synuclein can increase firing in cultured dopamine neurons (Dagra et al., NPJ Parkinsons Dis 2021, PMID: 34408150) while mice expressing mutant Parkin have higher mean firing rates (Regoni et al., Cell Death Dis 2020,  PMID: 33173027). Similarly, an increased firing rate has been reported in the MitoPark mouse model of PD at a time preceding DA neuron degeneration (Good et al., FASEB J 2011, PMID: 21233488). We also acknowledge that alterations to dopamine neuron activity are likely complex in PD, and that dopamine neuron health and function can be impacted not just by simple increases in activity, but also by changes in activity patterns and regularity. We will amend our discussion to include the important caveat of changes in activity occurring as compensation, as well as further evidence of changes in activity preceding dopamine neuron death.

      There is a very large, multi-decade literature on calcium elevation and its effects on neuronal loss in many different types of neurons. The authors should discuss their findings in this context and refer to some of this previous work. In a nutshell, the observations of the present manuscript could be summarized by stating that the chronic membrane depolarization induced by the CNO treatment is likely to induce a chronic elevation of intracellular calcium and this is then likely to activate some of the well-known calcium-dependent cell death mechanisms. Whether such cell death is linked in any way to PD is not really demonstrated by the present results. The authors are encouraged to perform a thorough revision of the discussion to address all of these issues, discuss the major limitations of the present model, and refer to the broad pre-existing literature linking membrane depolarization, calcium, and neuronal loss in many neuronal cell types.

      While our model demonstrates classic excitotoxic cell death pathways, we would like to emphasize both the chronic nature of our manipulation and the progressive changes observed, with increasing degeneration seen at 1, 2, and 4 weeks of hyperactivity in an axon-first manner. This is a unique aspect of our study, in contrast to much of the previous literature which has focused on shorter timescales. Thus, while we will revise the discussion to more comprehensively acknowledge previous studies of calcium-dependent neuron cell death, we believe we have made several new contributions that are not predicted by existing literature. We have shown that this chronic manipulation is specifically toxic to nigral dopamine neurons, and the data that VTA dopamine neurons continue to be resilient even at 4 weeks is interesting and disease-relevant. We therefore do not want to use findings from other neuron types to draw assumptions about DA neurons, which are a unique and very diverse population. We acknowledge that as with all preclinical models of PD, we cannot draw definitive conclusions about PD with this data. However, we reiterate that we strongly believe that drawing connections to human disease is important, as dopamine neuron activity is very likely altered in PD and a clearer understanding of how dopamine neuron survival is impacted by activity will provide insight into the mechanisms of PD.

    1. eLife Assessment

      This study addresses an important gap in drug discovery by delivering a rigorous, large-scale evaluation of widely used co-folding methods for predicting ligand-bound protein complexes and virtual screening. A key strength is the comprehensive benchmarking framework, which leverages structures and chemical compounds that were absent from the AI models training set, thereby providing particularly compelling and unbiased evidence of co-folding performance. The findings clearly delineate the complementary roles of deep learning-based co-folding and physics-based docking, offering practical guidance for their rational integration into drug discovery workflows. Overall, the conclusions are well supported by thorough analyses across a representative set of cases.

    2. Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      Strengths:

      Overall, this is a scientifically solid paper.

      The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Comments on revised version:

      The authors have adequately addressed my concerns.

    3. Reviewer #3 (Public review):

      Summary:

      Core conclusions are well-supported by data: co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods, provides an unbiased and rigorous benchmark dataset, which contains structures and compounds absent from the co-folding models training sets. Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment. The revised results clarify an intriguing finding: co-folding can predict correct ligand poses even when protein formations are mispredicted. The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      Weaknesses:

      The study identifies a major limitation of co-folding-failure to capture rare protein conformational changes, which deserve future investigation. The authors include uncalibrated Boltz-2 affinity data (addressing a prior comment) but note that large-scale free energy perturbation (FEP) comparisons are beyond their capabilities.

      Appraisal of Aims Achieved:

      The authors successfully achieved their primary aims and the results provide strong, well-supported evidence for their core conclusions. Key conclusions are grounded in the study's unbiased, training-set independent data, ensures the conclusions are not confounded by model memorization and are broadly applicable to the field's use of these co-folding models.

      Field Impact:

      This study provides a critical reality check for the field: co-folding models are powerful tools for pose prediction but are not yet standalone solutions for virtual screening, a key distinction that will prevent over-reliance on these models and guide more rational tool selection.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      We thank Reviewer 1 for this thoughtful summary of our work.

      Strengths:

      Overall, this is a scientifically solid paper. The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Weaknesses:

      My main concern is that the study's aim is a bit unclear. Modern benchmarking studies comparing physics-based docking with deep learning-based co-folding approaches (e.g., AF3, Boltz-2, Chai-1, and others) are increasingly expected to go beyond aggregate performance metrics.

      Indeed, we have gone into several examples of failures and successes for each of these methods. As we are not developing these methods ourselves, we also think this dataset will be a valuable contribution for improving them further.

      In addition to rigorous dataset construction, transparent methodology, and appropriate statistical evaluation, high-impact benchmarks typically provide actionable guidance on when each method class is most appropriate, reflecting their distinct inductive biases and practical constraints. Failure-mode analyses that link performance differences to protein flexibility, ligand chemistry, or binding-site characteristics are particularly valuable, as they move comparisons beyond "scoreboard" assessments toward mechanistic understanding.

      Right now, we do not observe meaningful trends that separate the failure modes for any individual method. This is covered in Supplementary Figures 6 and 7.

      While full biological validation is not expected, qualitative interpretation grounded in physical and biological principles strengthens conclusions. Providing reproducible workflows or reference pipelines is not mandatory, but it is increasingly viewed as a best practice because it facilitates adoption and helps contextualize results for practitioners.

      We note that our code is available (https://github.com/jongbin99/Cofolding/) and all structural data will be publicly accessible in the PDB alongside publication (we only held it back only for “blinding” during peer review to avoid contamination with any new deep learning methods).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Kim et al. evaluates the performance of three modern AI-based methods in predicting complex structures and binding affinities between proteins and chemical compounds. An honest 'prospective' evaluation is achieved by studying benchmark structures and chemical compounds that did not exist in the PDB at the time the AI structure prediction models (AlphaFold3, Chai-1, Boltz-2) were trained.

      Strengths:

      (1) The study addresses an important question in modern computational biology and drug discovery, and establishes the strengths and limitations of the three tools in solving various computational chemistry tasks, including compound pose prediction, active-inactive discrimination, and potency ranking.

      (2) The conclusions are based on examination of four separate targets and respective compound datasets, where for one of the targets, the authors also obtained numerous X-ray structures to serve as experimental answers for the binding pose prediction task.

      (3) The study reports relationships between structure prediction confidence, predicted energies (DOCK3.7), and affinity predictions (Boltz-2) with the geometric accuracy of compound pose prediction as well as the experimentally measured potency.

      (4) One of the key findings is the limited ability of co-folding methods to predict conformational rearrangements, which does not correlate with their ability to predict binding poses of the compounds inducing these rearrangements.

      (5) The findings could serve as useful guidelines for computational chemists in selecting appropriate software and scoring schemes for each task.

      We appreciate Reviewer 2’s summary of the novelty of the dataset and analysis.

      Weaknesses:

      While I consider this a solid study, several aspects would need to be addressed to make it really strong:

      (1) DOCK3.7 docking and scoring experiments were performed using one experimental structure of Mac1, selected from dozens of structures based on a criterion that is not sufficiently well justified. For sigma2 receptor, dopamine D4 receptor, and AmpC β-lactamase, it is not clear which structures or models were selected for docking at all. It is well known that geometry predictions, scoring, and active-inactive ROC AUCs are all strongly influenced by the selected structure. It would be important to attempt Mac1 docking using all available experimental Mac1 structures, or at least against representative structures in various conformations; it would also be quite insightful to compare results to docking of the same compound sets to AF3, Boltz-2 and Chai-1 predicted structures of Mac1. Same goes for the docking studies of sigma2, D4, and AmpC β-lactamase.

      In any program, a decision has to be made as to which template will be used for docking, we justified the choice in the methods:

      “We used this structure because the inhibitor (Z5014193706) was the most potent molecule with a structure determined around the same time as the ligands in this dataset were tested.”

      We stand by this as a reasonable assumption. Similarly, for sigma2, D4, and AmpC β-lactamase, the template was chosen in the respective papers:

      a) The σ2 receptor bound to cholesterol (PDB ID: 7MFI) was used in the docking calculations.

      - This structure was determined in the paper, the first structure of sigma2 and therefore a worthy template

      b) The D4 receptor campaign used PDB 5WIU

      - This was one of two D4 structures available and chosen because it was not bound to sodium

      c) For AmpC, the campaign used the structure in the Protein Data Bank (PDB) 1L2S

      - This maximizes comparisons to other docking studies that used the same receptor template.

      The major goal of this study is to compare different methods under reasonable (but perhaps as the reviewer points out, not optimal) conditions, not to optimize docking score.

      (2) For binding affinity predictions, as a control, authors should consider compound co-folding with an unrelated protein, or even with a pseudo-peptide that consists of a few random single amino acids - this would provide an honest baseline for such predictions.

      This suggestion would be valuable for understanding the performance for these methods from the perspective of ligand specificity (a valuable, but separate, goal). Surely this will generate some number or some prediction - but what would this baseline mean and how would it be relevant for drug discovery? Therefore, we do not think this suggestion is relevant for the issues being investigated in this manuscript.

      (3) ROC curves Figure 3 and elsewhere should be shown, and AUCs quantified/reported on a log or square-root scaled x-axis, to emphasize early enrichment, which is the area of practical significance for these predictions. For example, Figure 3A currently suggests that the pose prediction performance of AF3 exceeds that of Boltz-2 whereas the early enrichment is clearly better for Boltz-2.

      We agree with this, and added a semi-logAUC plot for Figure 3A. For Figure 5, we also generated a semi-logAUC plot to see early ligand enrichment clearly, added as Supplementary Figure 11. We added the text:

      “Considering its early enrichment performance, Boltz-2 Ligand ipTM was the strongest predictor of pose accuracy based on normalized logAUC (20.5% above random, Fig. 3a). In contrast, although Boltz-2 pIC50 showed poor overall discrimination, it overestimated its ability to enrich true positive poses at low false positive rates, despite having a weak early enrichment behavior”

      (4) 'Trained set' in figures and text should probably be 'training set'? Or otherwise explain this new term the first time it is introduced.

      Thank you for pointing out this for clarification. ‘Training set’ is the correct word, and we made changes appropriately across all figures and texts.

      (5) Figure 1 illustrates a projection onto the first two principal components of a space that apparently had only one (scalar) metric for each compound pair (% maximum common substructure or Tanimoto coefficient); the authors need to better explain the principle behind this analysis and visualization.

      This suggestion is valuable, since we often use PCA to reduce dimensionality for more complex features. For clarification, we actually have a full pairwise similarity matrix for all tested Mac1 compounds based on each of Tc and MCS%. PCA for each MCS% and Tc is a representation of each pairwise similarity matrix. We also made a change in Figure 1 caption to make this point clearer:

      “projection of compounds represented by their full pairwise similarity vectors (by ECFP-4 Tc and MCS%)”

      Reviewer #3 (Public review):

      Summary:

      This study's core conclusions are well-supported by data. It is shown that co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false-positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      (1) Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods.

      (2) Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment.

      (3) The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      We thank Reviewer 3 for pointing out the unprecedented and comprehensive nature of our study

      Weaknesses:

      (1) Limited generalization to diverse protein families (e.g., no ion channels/transporters).

      We agree - we have not explored the entire proteome and these are important target classes that will surely be investigated by future studies. We focused on targets here where we had large number of X-ray crystal structures (Mac1) and affinity/inhibition measurements from docking (the other three targets).

      (2) Ambiguity in the mechanism underlying co-folding's failure to predict rare conformational changes.

      Again, we agree. We are not the developers of these methods. We observe that these methods do not predict conformational changes with high fidelity and this weakness is an area that co-folding methods will surely prioritize in the future.

      (3) Virtual screen comparison is unbalanced (docking-prioritized hit lists bias results).

      We acknowledge this in the results: “An important caveat is that the hit-lists were composed of molecules prioritized by docking in the first place, giving it an advantage on these particular sets.” and discussion: “Finally, comparing co-folding to docking based on hit-lists themselves selected by docking is arguably unfair to co-folding. Counter-balancing this is the inclusion, in each of the three hit lists, of molecules that had mediocre and poor docking scores intentionally selected to test the correlation between docking score and hit-rate. Here too, the correlation between co-folding score and likelihood to bind, what we sometimes call a “dock-response-curve” was no better than docking’s, often worse (SFig.11).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are suggestions for revisions:

      (1) The writing is at times obtuse and hard to follow.

      This happens sometimes when multiple authors are writing together. We apologize and are happy to respond to specific areas that can be streamlined to be easier to follow.

      (2) In the Results section, "A set of 557 previously unreported Mac1 ligand complexes", the authors have compared the ligand poses across different metrics such as Tc - a standard, highly effective method in chemo-informatics and MCS (maximum common substructures); these are standard metrics for quantifying the structural similarity between pairs of small molecules. This part of the analysis checks whether this is memorization; it is critical to compare the two metrics, but it is not sufficient to draw a conclusion.

      Thank you for pointing out about the structural similarity of molecules co-folded to those present in the training set (resolved as Mac1 complexes and deposited in PDB before training dates). We have conducted an analysis where we do a pairwise similarity comparison for all ligands present in the PDB (regardless of the target), by both Tc and MCS, and overlay the cluster of ligands we tested (Mac1, AmpC, sigma2, D4). This should show where our tested benchmark datasets lie in the chemical space covered in the entire PDB. Each cluster (around 500 to 1300 compounds per target system) is overlaid on the cluster of all ligands deposited in PDB (over 50,000 compounds), and each cluster was relatively diverse by both Tc and MCS.

      (3) In the "Co folding can accurately reproduce poses of ligands dissimilar to those trained." Subsection under Results, the authors' conclusions are hard to follow; they state that the co-folding models often mispredict or miss the alternative conformation, but they also predict poses that are distinct from the training set. What does that imply?

      Our interpretation is actually a somewhat unsettling one: co-folding gets the ligand pose right even when it gets the protein wrong, and even when the ligand is novel. This suggests the models may be anchoring on conserved pharmacophoric interactions (like the adenosine-mimicking purine scaffold) rather than truly modeling the physics of the full complex. We added to the results section:

      This result suggests that co-folding reliably recapitulates dominant ligand-binding interactions even in the absence of accurate protein conformational modeling, providing further support to the idea that they are learning specific interaction patterns rather than a deeper physics-based representation (Masters et al. 2025).

      (4) The Discussion section connects the results and conclusions, but it can be challenging to grasp the study's overall message.

      We think the final paragraph hits on three major points:

      - Co-folding accurately predicts ligand poses for known binders, but fails to capture conformational changes

      - Co-folding does not reliably distinguish true binders from false positives in virtual screening hit lists

      - Docking and co-folding are complementary rather than competing tools

      (5) The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment. The value of the paper would be further enhanced by explaining how it differs from seemingly similar results reported in other studies, including the one cited in this manuscript (see https://www.biorxiv.org/content/10.64898/2025.12.04.692352v1).

      The Mac1 results are completely unique. However, the docking datasets are exactly the same as those analyzed in the Menon et al manuscript. We don’t think our results differs from conclusions of the Menon et al manuscript as we wrote: These observations are supported by a fascinating study on some of the same ligand sets as investigated here, using AlphaFold3, reaching similar conclusions (Menon et al. 2025).

      Reviewer #3 (Recommendations for the authors):

      (1) Expand target diversity to include ion channels, transporters, etc., beyond enzymes and GPCRs.

      (2) Investigate the cause of co-folding's failure in predicting rare conformational changes (e.g., adjust sampling, MSA inputs, or add experimental constraints).

      (3) Mitigate docking bias in virtual screens (e.g., re-analyze unbiased compound libraries).

      We addressed these three points in the public review above

      (4) Test Boltz-2's affinity predictions without linear calibration and compare with FEP.

      The data without linear calibration are included in the manuscript. Comparing such a large number of compounds with FEP is currently beyond our capabilities.

      (5) Conduct proof-of-concept to test co-folding-docking integration for better hit rates.

      We think this is well beyond the scope of this manuscript - but look forward to testing this idea in the future.

      We also got one community review that we respond to below:

      Summary

      This manuscript evaluates the performance of co-folding models when tasked with 1) the recapitulation of a large number of experimentally determined co-crystal structures of Mac1 with a series of Mac1 ligands and 2) the rescoring of hits to identify false positives originally derived from a set of large docking-based virtual screens. The evaluation leverages a dataset of crystal structures and affinity data from high-throughput crystallographic and biophysical screens, respectively. These data uniquely enable this report to focus on the ability of co-folding models to handle ligands, resulting in an analysis that is particularly timely given the wide adoption of co-folding models and the relative scarcity of such ligand-focused benchmarks among existing evaluations, which have primarily focused on protein structure prediction or binder design.

      Thank you for this thoughtful summary of our work

      Feedback

      The experiments and analyses in the manuscript are well thought-out and do not have any significant issues. There are a few high-level points that may improve the clarity and completeness of the results. Importantly, none of the suggested additional experiments will affect the conclusions of the paper, but rather help provide additional context for the results:

      The first section presents an exciting opportunity to frame the Mac1 ligands against ligands in the PDB more broadly. It would be informative to assess whether chemotypes that are easier or harder to predict accurately and confidently are over- or under-represented in the PDB as a whole. Note that this is not a recommendation that new scaffold similarity metrics be incorporated into the analysis, but rather that analyses similar to those already performed in the manuscript are performed using all ligands in the PDB. For example, PCA-based analyses similar to those in Fig. 1c could be used to examine Mac1 ligands in the context of all PDB ligands enabling questions such as whether similarity to a nearest PDB neighbor, cluster size in a Tc/MCS PCA space, or other frequency-based measures show any relationship with prediction vs. crystal structure RMSD. Such analyses could provide additional insight into how effectively models leverage ligand information present in the PDB overall, as opposed to biases arising specifically from scaffolds represented in Mac1 structures in the PDB, which are already well covered in the manuscript. The conclusion that Tc/MCS do not correlate with the ligand RMSDs for the ligands already associated with the Mac1 is well supported, and presumably suggests that a correlation would not exist against the backdrop of the PDB, but it would be interesting to see the data using analyses similar to those already done in the manuscript nonetheless.

      We are adding new figures in SFig.1 that consider how different clusters of ligands tested for our co-folding analysis are distributed across the chemical space in PDB. This is done by making a similarity comparison between every ligand in PDB and those tested in our analysis by Tc and MCS%, then plotting in PCA space for each metric. We are excited to see that each dataset covers a wide scope in PCA space, but at the same time, there are unexplored areas in the chemical space of PDB by co-folding.

      Similarly, even though the four proteins used in this manuscript are not themselves the primary focus of the analysis, it would be valuable to perform a high-level assessment of the precedent for each protein in the PDB (beyond the count of liganded structures in Table S6), either in protein sequence space (e.g., MSAs) or structural space (e.g., FoldSeek). An analysis like this would provide important context about whether any of the proteins in the study have close homologs with liganded structures in the PDB, or are generally overrepresented in the PDB. The fact that the AUC for L-pLDDT for AmpC is higher than σ2 and D4, for example, is notable given the relative abundance of liganded AmpC structures in the PDB (this raises potentially interesting questions related to where DOCK3.7 and AF3 actually place the ligands, given the orthosteric β-lactam binding pocket in AmpC, although this is outside of the scope of this manuscript).

      High-level assessment of the precedent for each protein in the PDB will definitely help to understand if proteins we used have close homologs with liganded structures in the PDB. Our Supplementary Table 6 covers the extent to which these liganded structures were available by cutoff dates for AF3, Chai-1 and Boltz-2. AmpC had more homologs than sigma2 and D4, and this may explain a better AUC for AF3 L-pLDDT specifically for this target.

      A discussion of the affinity probability results (`affinity_probability_binary`) from Boltz-2 is likely warranted in the second section in addition to the pIC50s that are already reported (`affinity_pred_value`). The former seems like it would be more applicable for section 2 of the manuscript, but both warrant inclusion—they should both be calculated by default when the affinity pipeline in Boltz-2 is turned on, so it wouldn't involve any more inference.

      As boltz-2 affinity module outputs both affinity probability binary output and affinity predicted value, we kept track of both metrics. So we tried re-ranking hit lists using both metrics. Where boltz-2 performed better (Sigma2, D4), binary probability values were more representative as a metric to differentiate true actives from non-binders. This was more clear in semi-logarithmic ROC plots. However, in AmpC, both Boltz-2 scoring metrics performed similarly. Such inconsistency in trend made it difficult to draw conclusions.

      Minor points

      A more detailed description of the experimental methods used to generate the ground-truth data in the introduction (even though these have been explained in prior works) would help orient the reader early on, and ground the benchmarking aspect of the story. In general, the abstract and introduction would benefit from a more cohesive through-line to tie the two complementary but orthogonal sections of the paper together.

      We will include a more thorough description alongside the PDB depositions. As for the two sections, we have tried to tie them together from the perspective of drug discovery workflows…

      The cutoffs in the "Co-folding can accurately reproduce..." section shift between 2.5 Å (from the ligand center of mass) and 2.0 Å. Is there a reason for this? Along similar lines, mentioning cutoffs for true positives/negatives when introducing the ROC analyses later on in the Mac1 section seems unnecessary since no cutoff should be necessary here.

      We used 2.5A distance to COM to just get at “broadly the correct binding site” for fast filtering and 2.0A RMSD because that is the broadly accepted standard in the field for “relatively correct binding pose”.

    1. eLife Assessment

      This study identifies apoptotic retinal ganglion cells as a potential source of ATP-mediated activation of PANX1 channels that initiates developmental retinal Ca²⁺ waves and coordinates microglial activation and vascular outgrowth during postnatal maturation. The work is important because it proposes an integrative framework linking programmed cell death, spontaneous neural activity, immune responses, and angiogenesis into a self-regulating developmental loop. Although the mechanistic conclusions would benefit from complementary genetic validation, the study provides a convincing foundation for future investigations into the coordination of neural circuit development and tissue remodeling.

    2. Reviewer #1 (Public review):

      Summary:

      This study presents a potentially important integrative model linking spontaneous retinal waves, apoptosis, microglial activity, and vascular development during postnatal retinal maturation. Its significance lies in proposing a mechanistic framework that could reshape understanding of how neural activity and tissue remodeling are coordinated in the developing central nervous system. The evidence is strengthened by the use of multiple complementary techniques, including Ca++ imaging, high-throughput electrophysiology, transcriptomics, histology and pharmacology.

      Strengths:

      (1) Multimodal Validation: The authors correlate large-scale functional imaging (calcium imaging and MEA) with high-resolution structural and molecular data (scRNA-seq and IHC), providing strong topographical evidence for the "centrifugal expansion" pattern.

      (2) The primary significance lies in identifying apoptotic Retinal Ganglion Cells (RGCs) as the physiological "pacemakers" for stage II retinal waves. By linking programmed cell death directly to neural activity and subsequent angiogenesis, the authors propose a self-regulating developmental loop.

      Weaknesses:

      (1) While the PANX1 pharmacological data provides compelling functional support, extending these conclusions to the broader CNS may be premature. Additional direct mechanistic validation would further strengthen the claim of causality.

      (2) While the manuscript beautifully illustrates the co-occurrence of events during retinal development, strengthening the distinction between correlation and direct causation would enhance the impact of the findings.

      Appraisal of Aims and Conclusions:

      The authors successfully achieve their aim of presenting a cohesive, multi-layered framework for postnatal retinal maturation, aligning functional physiological data with structural and transcriptomic timelines. The data robustly supports the correlation between retinal waves, microglial activity, and vascular remodeling and also identifies apoptotic RGCs as the potential "pacemakers" of Stage II waves.

      Impact, Utility, and Community Asset:

      This work will significantly impact developmental neurobiology by reframing programmed cell death as an active, instructive driver of neural network patterning and angiogenesis, rather than a passive clearance process. Methodologically, the integration of large-scale MEA recordings and live calcium imaging with scRNA-seq sets an excellent benchmark for multimodal developmental studies. Furthermore, the transcriptomic datasets mapping microglial phenotypes and vascular remodeling will serve as a highly valuable reference repository for the broader visual neuroscience community.

      Additional Context for Readers:

      To fully appreciate this study, readers should view it through the lens of neurovascular unit assembly. While Stage II cholinergic waves are traditionally studied purely in the context of visual circuit refinement, this work adds vital context by showing they also regulate the surrounding metabolic ecosystem. It effectively demonstrates that early electrical activity, programmed cell death, and vascular scaffolding do not occur in isolation, but are deeply interdependent processes.

    3. Reviewer #2 (Public review):

      Summary:

      Savage et al. investigates the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Using scRNAseq, the authors classify autofluorescence cell clusters (ACCs) at the leading edge of vasculature outgrowth as Hmox-1+ microglia. From here they show microglia engulfment of apoptotic RGCs and the potential release of ATP may contribute to Ca2+ wave generation. The authors demonstrate these mechanisms through the use of two pharmacological to agents to either block the ATP release from Panx-1 or by blocking receptor binding to ATP. Furthermore, while previous studies have described the site of initiation of retinal Ca2+ waves as random, this study shows the initiation of Ca2+ waves are biased to the leading edge of vascular growth in the developing retina. To do this, the authors use a combination of wide-field Ca2+ imaging and multi-electrode arrays to pinpoint the sites of Ca2+ wave initiation in the developing retina.

      Strengths:

      Savage et al. uses a several techniques to interrogate these mechanisms, including single cell RNAseq, wide-field Ca2+ imaging, and multi-electrode arrays. With these experiments, this manuscript proposes several novel ideas, such as ATP as the Ca2+ wave initiating cue, and the localization the Ca2+ wave initiation to the leading edge of vascular growth.

      Weaknesses:

      The main limitation of this study is the reliance on only two pharmacological agents to test their central hypotheses. In future studies, these conclusions could be strengthened if they used genetic knockout models to perturb programmed cell death and/or ATP release (i.e. BAX-KO, Panx-1 KO).

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents a potentially important integrative model linking spontaneous retinal waves, apoptosis, microglial activity, and vascular development during postnatal retinal maturation. Its significance lies in proposing a mechanistic framework that could reshape understanding of how neural activity and tissue remodeling are coordinated in the developing central nervous system. The evidence is strengthened by the use of multiple complementary techniques, including Ca++ imaging, high-throughput electrophysiology, transcriptomics, histology, and pharmacology.

      Strengths:

      (1) Multimodal Validation: The authors correlate large-scale functional imaging (calcium imaging and MEA) with high-resolution structural and molecular data (scRNA-seq and IHC), providing strong topographical evidence for the "centrifugal expansion" pattern.

      (2) The primary significance lies in identifying apoptotic Retinal Ganglion Cells (RGCs) as the physiological "pacemakers" for stage II retinal waves. By linking programmed cell death directly to neural activity and subsequent angiogenesis, the authors propose a self-regulating developmental loop.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      (1) While the PANX1 pharmacological data provide compelling functional support, extending these conclusions to the broader CNS may be premature. Additional direct mechanistic validation would further strengthen the claim of causality.

      We agree with the reviewer that the conclusions would be greatly solidified with more direct mechanistic validation. However, we are unable to conduct more experimentation as the grant is finished and the Sernagor lab is in the process of being shutdown, after the unexpected passing of the PI.

      In order to make clearer that this mechanism was found in retinal tissue, not CNS, we have moved any mention of the implications of our work to a broader CNS mechanism to the discussion section. We have also added text into the discussion highlighting the need for more mechanistic investigation to uncover the full extent of the developmental processes described herein, see Line 413.

      (2) While the manuscript beautifully illustrates the co-occurrence of events during retinal development, strengthening the distinction between correlation and direct causation would enhance the impact of the findings.

      We have been clear to only present our findings as correlational as we were unable to fully explore the causational nature within the mechanisms presented. In the discussion, we have used published evidence and experimental papers to bolster our understanding of the causal aspects of this research. We have also included sections of text to address what experimentation is be required to examine the causal interactions more directly, see Line 413.

      Reviewer #2 (Public review):

      Summary:

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Using scRNAseq, the authors classify autofluorescence cell clusters (ACCs) at the leading edge of vasculature outgrowth as Hmox-1+ microglia. From here, they show microglia engulfment of apoptotic RGCs, and the potential release of ATP may contribute to Ca2+ wave generation. The authors demonstrate these mechanisms through the use of two pharmacological agents to either block the ATP release from Panx-1 or block receptor binding to ATP. Furthermore, while previous studies have described the site of initiation of retinal Ca2+ waves as random, this study shows that the initiation of Ca2+ waves is biased to the leading edge of vascular growth in the developing retina. To do this, the authors use a combination of wide-field Ca2+ imaging and multi-electrode arrays to pinpoint the sites of Ca2+ wave initiation in the developing retina.

      Strengths:

      The authors use several techniques to interrogate these mechanisms, including single-cell RNAseq, wide-field Ca2+ imaging, and multi-electrode arrays. With these experiments, this manuscript proposes several novel ideas, such as ATP as the Ca2+ wave-initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      The main weakness of the manuscript is the overreliance on only two pharmacological agents to test the central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP release (i.e., BAX-KO, Panx-1 KO).

      We thank the reviewer for their insightful suggestions for further experimentation to bolster the research. Initially, we utilised pharmacological interventions as they provided acute and quick answering of the research question. At the outset of the research, we were not certain that purinergic release through PANX-1 channels was the mediator for the developmental mechanisms described. We tested a wide variety of specific agonists and blockers before seeing any profound effects on wave generation. These agonists and antagonists have been used before and are proven to deliver reliable results. In addition, since the ACCs had never been reported before we were unsure if a knockout animal would display the same anatomical phenotype. Furthermore, it is known that knockout mouse lines, especially connexin and hemichannel pores, do not lose function but rather have other isoforms or compensation mechanisms which can substitute the original function. For the retina, for example, it was shown that Cx36 can functionally replace Cx45 after Cx45 KO (Frank et al, 2010).

      We agree that while direct mechanistic validation would significantly reinforce the arguments, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      In order to address the omission of mechanistic validation in the paper we have added text into the discussion highlighting the need deeper investigation in the causality of the developmental processes described herein, see Line 413.

      M. Frank et al., Neuronal connexin-36 can functionally replace connexin-45 in mouse retina but not in the developing heart, J. Cell Sci. 123, 3605 (2010).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      General and major comments

      (A) Introduction

      (1) The introduction is currently quite extensive. I recommend streamlining the background information to more directly frame the study's core objectives.

      We have reduced the background information contained in the introduction to better align with the direct outputs of the study. However, as this research paper examines the interactions of multiple complex developmental processes a relatively in-depth introduction is needed to inform the reader of the salient points.

      (2) To improve clarity, it would be highly beneficial to conclude the introduction with a sequential summary of key observations in the order they are presented in the study. This would provide a clearer roadmap for the reader.

      We have reworked the end of the introduction to better align with the key observations in the order they are presented through the figures.

      (B) Results

      (3) The characterization of ACCs (Apoptotic Cell Clusters) would be more effective if separated from the description of their spatiotemporal occurrence with SVPs. Consider moving the microscopic description to a dedicated section or merging it with the subsequent paragraph on molecular identification.

      We thank the reviewer for their suggestion to reorganise the results section. However, we believe that highlighting the integrated nature of the ACC positioning and development of the vascular plexus is an important stepping off point for the reader. It allows us to highlight the original serendipitous discovery of the ACCs and their highly cohesive role in bridging multiple developmental processes.

      (4) Please explicitly state the specific retinal developmental stages (e.g., P3-P6) within the section regarding RNA sequencing, as this timing is critical for interpreting the transcriptomic data.

      We have amended the text to specify the postnatal days used for RNA sequencing, see Line 123

      (5) Regarding the scRNA-Seq data: were ACC-negative samples (isolated via FACS) also processed? A direct comparison between ACC+ and ACC- sampled microglia would significantly strengthen the claim that microglia are specifically attracted to ACCs. If these data are available, they would make an elegant and compelling addition to the manuscript.

      We thank the reviewer for this important suggestion and agree that direct comparison of ACC+ and ACC− microglia would further strengthen the study. Unfortunately, ACC− populations were not processed for scRNA-seq in the current study because of the prioritisation of the rare ACC+ population. Nevertheless, several independent observations support the conclusion that microglia are preferentially associated with ACCs, including: (i) the enrichment of microglia within ACC-containing regions observed histologically, (ii) the spatial proximity analyses shown in Figure 3, and (iii) the distinct transcriptional profile of ACC-associated microglia identified by scRNA-seq.

      We have also added a section to the discussion to highlight the need for a direct comparison of the ACC+ and ACC- transcriptomics profiles in future work, see Line 344.

      (6) For the Ca2+ -imaging experiments, please briefly describe the staining protocol and specify which cell types (e.g., RGCs) were labeled within the Results text to assist the reader's immediate understanding.

      We have added a short description of the labelling technique in the results, see Line 228

      (7) The manuscript notes that MEA waves are evident in the graphs of Figure 7, but the raw wave data or representative traces are not shown. Including these (similar to those of the imaging waves) would provide necessary visual verification of the physiological phenomena described.

      We have added a supplementary figure 3 which details stage 2 retinal waves recorded using MEAs.

      (C) Interpretations and Logic

      (8) The finding ' ...Wholemount staining revealed a broad centro-peripheral gradient of apoptosis; however, this apoptotic annulus was positioned more peripherally than the ACCs, SVP, and Hmox1-positive microglia (Figure 4B)' seems to contradict the HMOX1/Yo-Pro-1 stained microglia. It is not clear whether the authors aim to prove that this particular set of microglia phagocytise RGCs, or another set that lines up better with dying cells and does not show up on the HMOX-1 label. I believe the authors intend to show the time difference of the two events - cells dying and HMOX-1 microglia appear at the site later. I believe the logic is good; it may need a sentence pointing this out at the end of this paragraph.

      We have added a statement in Line 182 which clarifies our intent to show that the Hmox1 microglia phagocytose the dying RGCs after they initiate apoptotic mechanisms.

      (9) If the authors intend to demonstrate a temporal lag between cell death and the appearance of HMOX1+ microglia as evidence of causality, a concluding sentence to this effect would greatly clarify the logic of this paragraph.

      We have added a concluding sentence to the paragraph in Line 190 which indicates a causal link between RGC cell death and appearance of hmox1 positive microglia.

      (D) Figures and Presentation

      (10) The blood vessel staining in the final panel of Figure 1B is currently quite faint. Increasing the brightness/contrast for this panel would allow the reader to better appreciate the underlying architecture.

      We have updated the panel in Figure 1B to match the brightness of the others of that series.

      (11) Given that the peripherality of events in Figure 7 suggests a specific sequence, the authors should consider adding a summary timeline (P3-P6). A plot using curves (mean or median values), color-coded to match the corresponding events, would provide a much-needed visual synthesis of the data.

      We agree with reviewer that the D1/2 metrics would benefit from more clarification to show the timeline of development more clearly. We have added another panel to Figure 7, which shows mean/standard deviation plots for each developmental measure using D1/2 as timelines. This allows the reader to better compare the progression of centrifugal spread more clearly.

      (12) Please ensure that graph labels and axis titles are uniform in size across all figures. e.g., the labels in Figure 6 and several other graphs are currently too small to be legible in the PDF; these should be enlarged for better accessibility.

      We have fixed the labels and axis titles to maintain readability across the paper

      Minor comments

      (1) In line 125, there is a missing closing parenthesis after the reference to Figure 1C.

      We have fixed this error

      (2) The specific algorithm used for the unsupervised cluster analysis has not been identified in the text. Please specify whether k-means, Louvain, or another method was employed to ensure reproducibility.

      We have reworked the section detailing the cluster analysis to make clear we used Louvain-based clustering, see line 475.

      (3) While it is appropriate to leave comprehensive technical details for the Methods section, a brief conceptual explanation of the D1, D2, and D3 metrics should be included in the Results text to aid general comprehension.

      We have added a brief description of the D1/2 and D1/3 metrics when they are first mentioned in the results section, see line 243.

      (4) In the Figure 7 schematic, the representation of the starburst amacrine cell should be revised to more accurately reflect its well-characterized morphology (e.g., thin primary and gradually thickening higher-order dendrites).

      We have changed the SAC representation to better match the characteristics of that cell type.

      (5) Throughout the manuscript (e.g., in lines 293-294), it is claimed that RGC apoptosis promotes the expression of PANX-1 hemichannels. While the data effectively demonstrate the release of purinergic molecules (e.g., ATP) via PANX-1 from dying cells, the evidence for an actual upregulation or increase in PANX-1 protein/mRNA levels is not explicitly shown. Please clarify whether the findings suggest increased activity of existing channels or a true increase in expression. If the latter is not empirically supported, the phrasing should be adjusted to reflect functional activation rather than de novo expression.

      We agree with the author that our research shows a functional increase in PANX-1 and we have adjusted the language to match. In the introduction and discussion, we provide published evidence and experimental papers which describe the upregulation of the PANX-1 molecule in dying RGCs.

      Reviewer #2 (Recommendations for the authors):

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Furthermore, the authors demonstrate the initiation of Ca2+ waves occurs at the leading edge of vascular growth in the developing retina. This manuscript proposes several novel ideas, such as ATP as the Ca2+ wave initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth. The main weakness of the manuscript is the overreliance on only two pharmacological agents to test their central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP releases (i.e., BAX-KO, Panx-1 KO). In addition, the following comments should also be addressed:

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Specific comments:

      (1) Line 128: Why would ACCs be involved in SVP guidance if they are trailing the leading edge of the vasculature? Would it be the other way around, where the leading edge of the vasculature would be trailing the ACCs?

      At this point in the paper we are suggesting that the highly stereotyped position of the ACCs under the leading edge of the SVP indicates that they have a mechanistic involvement in SVP growth. Not that they are the direct cause of the expansion. As the paper progresses, we make clear that contrary to our original hypotheses which state the ACCs may cause or control the integrated development of the retina, they are a hallmark of the apoptotic RGCs in the periphery which are the chemogenic beacons for vascular growth being ‘decommissioned’ by the microglia which fine-tune vascular growth and create the ACCs.

      (2) Line 137: Please state in the text and figure legend, at what age ACCs were isolated from the retina.

      We have added the relevant information to Line 123

      (3) Line 154: Why didn't RGCs form their own cluster? Why do the RBPMS+ cells appear across the entire dataset (Figure 2G)? Have previous investigations also shown engulfed cell transcriptomes appearing in the microglia clusters using scRNAseq?

      Previous transcriptomic studies have demonstrated that phagocytic microglia can encapsulate transcripts originating from neurons and other neural cell types, which are detectable by RNA‑seq despite not belonging to a common microglial genetic signature. For example, Solga et al. showed that CNS microglia contain neuronal and oligodendrocyte‑specific mRNAs that localise within microglia but are not translated. This research group interpreted that this RNA is acquired through phagocytosis or macropinocytosis of surrounding neural cells (Solga et al., 2015). Similarly, in zebrafish, synapse‑engulfing microglia identified in situ display neuronal and synaptic gene expression in single‑cell RNA‑seq profiles, consistent with engulfed neuronal material contributing to the detected transcriptome (Sliva et al., 2021). In line with these observations, and given our FACS strategy enriching autofluorescent ACCs rather than intact RGCs, we interpret the widespread Rbpms expression across ACC‑associated clusters as an expected consequence of microglial engulfment of apoptotic RGCs, rather than evidence for a distinct population of viable RGCs that failed to form a separate cluster.

      Solga, A.C., Pong, W.W., Walker, J., Wylie, T., Magrini, V., Apicelli, A.J., Griffith, M., Griffith, O.L., Kohsaka, S., Wu, G.F. and Brody, D.L., 2015. RNA‐sequencing reveals oligodendrocyte and neuronal transcripts in microglia relevant to central nervous system disease. Glia, 63(4), pp.531-548.

      Silva, N.J., Dorman, L.C., Vainchtein, I.D., Horneck, N.C. and Molofsky, A.V., 2021. In situ and transcriptomic identification of microglia in synapse-rich regions of the developing zebrafish brain. Nature communications, 12(1), p.5916.

      We have incorporated this information into the discussion in Line 329

      (4) Figure 4C-H: The data in the figure would be strengthened if the authors added quantification for their co-localization images.

      We thank the reviewer for this important suggestion and agree that quantification of the co-localisation would further strengthen the study. Unfortunately, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      (5) Figure 5A-I: The error bars are quite large. The figure legend says they represent SEM, but are you sure they don't represent standard deviation (SD)?

      The large SEM bars reflect substantial biological variability across retinas and across individual microglia, which is expected for morphometric measures such as circularity, perimeter, branch number and total skeleton length, as well as for counts of rare cell populations (Hmox1+ microglia, YO‑PRO‑1+ cells and double‑positive cells). Importantly, despite this variability, the effects of probenecid and PSB‑0739 on microglial morphology and on the frequencies of apoptotic and double‑positive cells remain statistically robust in our non‑parametric ANOVA and post‑hoc tests, as indicated by the reported P‑values.”

      For panels where the distributions were clearly non‑Gaussian, we used non‑parametric statistics (Kruskal‑Wallis ANOVA), reporting medians and 95% confidence intervals, and we retained SEM in Figure 5 for consistency with the original plotting routine while clarifying this choice in the legend and Methods.

      (6) Figure 5: What are these measurements made at? Please add the age to the results section and the figure legend.

      We have added the postnatal day of the animals used.

      (7) Line 236-239: There is no mention of the use of Probenecid in the text and in Figure 6B. This treatment of Ca2+ waves should be mentioned before line 245.

      We have added a brief description of probenecid application in Line 227.

      (8) Figure 6: The order of this figure and the results section may be clearer if panels C-F were switched with panels G-J.

      We thank the reviewer for their suggests to improve the flow of the results section and figure 6. We have swapped the panels as suggested and amended the text to fit the new flow.

      (9) For the discussion section: Why does probenecid only affect the wave initiation at the P3 timepoint, but not later timepoints, P4-P6.

      We have added some text in the discussion to explain our findings of differential effects of PANX-1 blockade across the P3-6 timeline. See Line 404.

      Editorial revisions:

      (1) Line 92-96: Citation needed.

      We have added appropriate citations for this section.

      (2) Figure 1F: Please consider changing the color scheme from green/red to green/magenta for colorblind readers.

      We have changed the image LUT

      (3) Figure 3A: This figure may benefit from separating out each channel separately (Iba1 and ACC) and then having a "merge" panel.

      We have separated out the panels in Figure 3A

      (4) Figure 4E, H: There is no label for the immunomarkers in Figure 4H. There is no inset in Figure 4E as mentioned in the figure legend. However, it appears that Figure 4H is a magnified image of Figure 4G and not Figure 4E.

      We have fixed the error in the figure legend and included an inset indicator in Figure 4G. We have added immunomarkers to Figure 4H

      (5) Line 198: "SAC" should be "SACs".

      We have fixed this error

    1. eLife Assessment

      This convincing contribution addresses a question of practical importance: when collecting tilt-series data, what is the optimal angular step size between successive tilt images? The work provides valuable practical insights into cryo-ET data acquisition by demonstrating that balancing two competing demands - sufficient dose per individual tilt image and fine angular sampling - is essential to achieve high-quality tomographic reconstructions. They demonstrate that tilt-series acquired with finer increments (1-3 degrees) yield superior alignment accuracy and improved template-matching performance.

    2. Reviewer #1 (Public review):

      This work addresses a question of practical importance that had never been systematically analysed in the cryo-ET field: when collecting tilt-series data, what is the optimal angular step size between successive tilt images? Due to the upper limit in electron exposure (100 - 150 e⁻/Ų), this question is important, since finer angular sampling improves attainable reconstruction resolution (Crowther criterion) but reduces the signal-to-noise ratio of each individual image, potentially compromising both image quality and the ability to computationally align successive frames. To address this, the authors designed a thorough benchmarking study comparing five tilt increments (1{degree sign}, 2{degree sign}, 3{degree sign}, 5{degree sign}, and 10{degree sign}) while keeping the total dose and tilt range constant. They evaluated the consequences at every stage of the cryo-ET workflow - from raw image quality and tilt-series alignment, through template matching for ribosome detection, to high-resolution subtomogram averaging - with the goal of providing the community with an evidence-based recommendation for data acquisition.

      The manuscript is well written, and the experimental design is carefully thought out. The work provides valuable practical insights into cryo-ET data acquisition by demonstrating that balancing two competing demands - sufficient dose per individual tilt image and fine angular sampling - is essential to achieve high-quality tomographic reconstructions. The identification of a practical optimum at 3{degree sign} tilt increment is the key contribution of the work. It will be interesting to see in the future whether this optimum shifts for smaller molecular targets, and how emerging tilt interpolation strategies such as cryoTIGER may interact with the choice of experimental angular increment.

      Comments on revised version.

      Well done! I really like the manuscript and from my point of view it's an excellent piece of work and super useful for the community. Thank you so much for the meticulous work!

    3. Reviewer #2 (Public review):

      The determination of macromolecular structures directly within their native cellular environment is becoming increasingly routine, making standardized data collection strategies essential. In this manuscript, Tuijtel et al. provide a timely and valuable contribution by benchmarking key acquisition parameters and establishing practical guidelines for in situ cryo-electron tomography (cryo-ET). Critically, the authors present a systematic framework for optimizing data collection to achieve the highest attainable resolution.

      Using Dictyostelium cells as a model system, the authors generate multiple datasets at a constant total dose while varying the tilt increment. They demonstrate that tilt-series acquired with finer increments (1-3 degrees) yield superior alignment accuracy and improved template-matching performance, resulting in higher-quality reconstructions than those collected with coarser increments (5 degrees or above). Furthermore, the authors show that for subtomogram averaging, a 3-degree tilt increment outperforms all other conditions tested, particularly after per-particle refinement as implemented in M.

      Comments on revised version.

      The authors have addressed all my concerns, and I have no further issues.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work addresses a question of practical importance that had never been systematically analysed in the cryo-ET field: when collecting tilt-series data, what is the optimal angular step size between successive tilt images? Due to the upper limit in electron exposure (100 - 150 e<sup>-</sup>/Å<sup>2</sup>), this question is important, since finer angular sampling improves attainable reconstruction resolution (Crowther criterion) but reduces the signal-to-noise ratio of each individual image, potentially compromising both image quality and the ability to computationally align successive frames. To address this, the authors designed a thorough benchmarking study comparing five tilt increments (1°, 2°, 3°, 5°, and 10°) while keeping the total dose and tilt range constant. They evaluated the consequences at every stage of the cryo-ET workflow - from raw image quality and tilt-series alignment, through template matching for ribosome detection, to high-resolution subtomogram averaging - with the goal of providing the community with an evidence-based recommendation for data acquisition.

      The manuscript is well written, and the experimental design is carefully thought out. The work provides valuable practical insights into cryo-ET data acquisition by demonstrating that balancing two competing demands - sufficient dose per individual tilt image and fine angular sampling - is essential to achieve high-quality tomographic reconstructions. The identification of a practical optimum at 3° tilt increment is the key contribution of the work. It will be interesting to see in the future whether this optimum shifts for smaller molecular targets, and how emerging tilt interpolation strategies such as cryoTIGER may interact with the choice of experimental angular increment.

      The conclusions of this paper are mostly well supported by data, but some aspects of data analysis need to be clarified and/or extended, including:

      (1) Line 109: The authors state that the tilt range was kept at ± 60° relative to the lamella plane. Assuming a typical lamella pre-tilt of ~10°, the absolute stage tilt would approach its mechanical limit. Two clarifications would be appreciated: (a) What was the average pre-tilt across all lamellae? (b) How many dark tilt images, if any, were excluded during tomogram reconstruction?

      We thank the reviewer for asking for further clarification. For all our datasets, the pre-tilt of the stage was +8° with the lamella untilted under the e-beam, resulting in a tilt range of -52° to + 68°, thereby not reaching the mechanical limit, which is 70° for our microscope stage.

      Regarding “dark tilt images”, for most datasets, we did not need to remove many tilt images. However, we now noticed notably more absence of images from higher tilt values for the 1° dataset (see SFig 1). When analysing further, we noticed that for this dataset, we did not actively remove many images prior to tomogram reconstruction, but rather that they were not acquired in the first place by SerialEM. During acquisition, SerialEM performs various safeguarding checks that can abort the acquisition of a tilt series (or of a single branch). As this seems predominantly a problem for the 1-degree tilt-increment dataset, we have decided to add this to the manuscript as follows, including the figure as new SFig 1.

      In the main text:

      “For most datasets, image acquisition was largely complete, with the exception of the 1-degree dataset, which showed a markedly higher proportion of missing images at high tilt angles (SFig. 1). Closer inspection revealed that many of these images were not acquired, as SerialEM applies built-in safeguards (e.g. autofocus inconsistency or insufficient image counts) that can abort a tilt-series branch before completion.”

      (2) Line 148: "When analysing tomographic volumes, we found that tomograms from data with a smaller increment displayed higher SNR values (see Fig. 2B)." It would be helpful to specify which comparisons are statistically meaningful (e.g. Mann-Whitney U test?). While the difference between 1° and 2° appears pronounced, the differences between 2°, 3°, and 5° seem minimal. From my point of view, reporting the mean SNR values +/- standard deviations for each condition would already indicate some significance. Furthermore, since SNR is expected to depend on lamella thickness, it should be clarified whether the average lamella thickness is comparable across the five datasets.

      We have now calculated the mean and standard deviation of the tomogram SNR, as follows:

      Author response table 1.

      Furthermore, we performed a statistical significance test. Kruskal-Wallis test confirmed significant differences in SNR across tilt increments (H=270.97, p<0.001). Pairwise Mann-Whitney U tests with Bonferroni correction revealed significant differences between all pairs except 2° and 3° (p=0.093), suggesting these two conditions indeed yield comparable SNR.

      Lastly, we have now added data regarding local lamella thickness for all tilt-series used in the study, as displayed in SFig. 4.

      We incorporated this in the manuscript as follows.

      In the main text:

      “When analysing tomographic volumes, we found that tomograms from data with a smaller increment displayed higher SNR values (see Fig. 2B and Supplementary Note), whilst showing a similar lamella thickness distribution (see SFig. 4A).”

      and:

      “Firstly, we selected ca. 20 tomograms per condition, based on tomogram content and local lamella thickness [31] (for more details, see Methods and SFig. 4B).”

      As a supplementary note:

      “As the tomogram SNR distribution of particularly the 2° and 3° dataset showed similar SNR distributions, we performed formal significance testing for the data in this panel (see Fig. 2B). Kruskal-Wallis test confirmed significant differences across conditions (H=270.97, p<0.001); pairwise Mann-Whitney U tests with Bonferroni correction revealed all pairs were significantly different except 2° vs. 3° (p=0.093), indicating comparable SNR for these two tilt increments.”

      And, adding the test in the Methods:

      “To quantify differences in signal-to-noise ratio (SNR) across tilt increment conditions, a non-parametric Kruskal-Wallis test was performed as an omnibus test of the null hypothesis that all groups are drawn from the same distribution. Because SNR distributions were not assumed to be normal, and sample sizes differed across conditions, non-parametric tests were used throughout. Following the omnibus test, all 10 pairwise comparisons between conditions were assessed using two-sided Mann-Whitney U tests. To control for multiple comparisons, raw p-values were adjusted using the Bonferroni correction (multiplied by the number of comparisons, n=10, capped at 1.0). Statistical significance was defined as a Bonferroni-corrected p-value below 0.05. All analyses were performed in Python using the scipy.stats module.”

      (3) Line 167: "Indeed, the variation in maximum resolution correlates with lamella thickness across all datasets (see Fig. 2F)." The reported R<sup>2</sup> values of 0.30 (1°), 0.38 (2°), 0.66 (3°), 0.61 (5°), and 0.60 (10°) reveal a notably weak linear relationship for the finer tilt increments. It is also difficult to assess whether the lamella thickness distributions are comparable across conditions from the current figures - visually, the 1° dataset appears to be based on thinner lamellae, while the 10° dataset appears to include thicker samples. A histogram of lamella thickness distributions for each condition, provided as supplementary material, would greatly aid interpretation. Given this thickness dependency, reporting mean +/- standard deviation of lamella thickness per condition is highly appreciated.

      We have added the full lamella thickness distribution per dataset now in SFig. 4.

      The apparent weaker relationship between resolution fit and local lamella thickness for the 1 dataset seems to be largely apparent to the few very thin data points in this data (for more clarity, see the same data plotted separately in Author response image 1). We speculate that this is due to even less signal in these very thin and very low-dose images.

      Author response image 1.

      (4) Figure 4: It should be specified which tomogram subsets were used for the Rosenthal-Henderson analysis, whether lamella thickness was taken into account in the subset selection, and whether ribosomes too close to the lamella edges were excluded. Finally, linear fits should be displayed across the full x-axis range for all tilt increments to facilitate direct visual comparison.

      We have described the process of tomogram subset selection in detail in the Methods section Template matching and 3D classification. To further add clarity, we have incorporated the local lamella distribution plots for the full data, as well as specifically for the tomograms subjected to TM and STA in SFig. 4.

      Regarding the linear fits, we respectfully disagree with this suggestion. Displaying the linear fits only over the range used for their calculation avoids implying that the linear relationship extends beyond the measured data, and in our view produces a clearer figure.

      (5) General: Were ribosomes located at the lamella edges excluded from the analysis? As demonstrated in the authors' own prior work (Tuijtel et al., Science Advances, 2024), Ga-FIB milling induces structural damage at the lamella surfaces. To exclude the influence on the STA results, particles near the lamella edges should be removed prior to analysis, and the criteria for this exclusion should be stated explicitly.

      We have not excluded any ribosomes from close to the surface. As we still treated all data the same for each condition, we anticipate that the results of the comparison reported here still hold true.

      The aim of the authors was to provide the cryo-ET community with an evidence-based recommendation for the choice of tilt increment, and they largely succeeded in this goal. The identification of 3° as a practical optimum - balancing sufficient dose per tilt image for effective per-particle refinement with fine enough angular sampling for accurate tilt-series alignment - is well supported by the data and consistent across the multiple quality metrics employed. The conclusion that coarser increments (5° and 10°) compromise tomogram quality, template matching accuracy, and STA resolution is robust and clearly demonstrated. However, the conclusion rests entirely on a single biological system using ribosomes as the sole molecular target, which are exceptionally favourable due to their abundance, size, and electron contrast. Whether the identified optimum holds for smaller, lower-abundance, or lower-contrast targets remains an open question.

      In future, it would be particularly interesting to test whether emerging tilt interpolation strategies, such as cryoTIGER, which is particularly intriguing, can effectively compensate for coarser experimental angular sampling in post-processing. Here, the optimal experimental increment may shift, and the interaction between these two approaches represents a promising direction for future work. More broadly, as cryo-ET datasets grow larger and public repositories expand, the practical tradeoffs between acquisition time, data storage, and structural quality identified here will become increasingly relevant to the field.

      We agree with the reviewer and thank them for this positive assessment. An interesting note to the use of cryoTIGER in particular is that it uses already aligned tilt-series as an input, and it therefore is unlikely to overcome severe alignment issues associated with large tilt-increments.

      Reviewer #2 (Public review):

      The determination of macromolecular structures directly within their native cellular environment is becoming increasingly routine, making standardized data collection strategies essential. In this manuscript, Tuijtel et al. provide a timely and valuable contribution by benchmarking key acquisition parameters and establishing practical guidelines for in situ cryo-electron tomography (cryo-ET). Critically, the authors present a systematic framework for optimizing data collection to achieve the highest attainable resolution.

      Using Dictyostelium cells as a model system, the authors generate multiple datasets at a constant total dose while varying the tilt increment. They demonstrate that tilt-series acquired with finer increments (1-3 degrees) yield superior alignment accuracy and improved template-matching performance, resulting in higher-quality reconstructions than those collected with coarser increments (5 degrees or above). Furthermore, the authors show that for subtomogram averaging, a 3-degree tilt increment outperforms all other conditions tested, particularly after per-particle refinement as implemented in M.

      Overall, the manuscript is clearly written, and the conclusions are well supported by the data presented. I have no major concerns. There are some minor points that the authors should address, including:

      (1) The phrase "electron optical density distribution" (line 31, Introduction) should be revised to "electrostatic potential" or "Coulomb potential distribution," which more accurately reflects what is measured in cryo-EM/ET.

      We thank the reviewer for this correction and have adjusted it in the text:

      “It captures the 3-dimensional (3D) electrostatic potential of the specimen under scrutiny and enables the structural analysis of macromolecular complexes within their native context.”

      (2) The authors state that the maximum tolerable electron dose is approximately 100-150 e<sup>-</sup>/Å<sup>2</sup> (line 34, Introduction). This is an oversimplification, as bacterial specimens, for example, have been shown to tolerate doses of 200 e<sup>-</sup>/Å<sup>2</sup> or higher (see Breigel et al., PNAS, 2009; https://www.pnas.org/doi/10.1073/pnas.0905181106#T1). The statement should be revised to reflect this variability.

      We adjusted this statement to now read:

      “One of these is the maximum electron dose that can be applied to biological specimens before irreversible damage occurs, which is about 100-150 e-/Å2 for most eukaryotic cells.”

      (3) Lines 56-57: The authors do not cite their own prior work benchmarking tilt-series acquisition strategies on in vitro samples. This earlier study provides important context and should be referenced and briefly discussed.

      We assume the reviewer meant this study: Turonova et al., Nat. Comms. (2020). We have now added and discussed this reference as follows:

      “Since accumulated radiation dose progressively degrades high-resolution information, this motivated the development of the dose-symmetric tilt scheme, which prioritizes acquisition of low-tilt images early to better preserve high-resolution information [11, 16].”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 159: "Surprisingly though, the resolution to which the CTF was fitted was similar for all conditions, despite an 8-fold increase in dose (see Fig. 2B, E)." The reference to Figure 2B at this point is unclear.

      We thank the reviewer for pointing this out, we have removed the reference to panel B.

      (2) Line 162: "As the data shown in Fig. 2D-F pertains to images of the untilted specimen, ..." For clarity, this should also be stated explicitly in the figure caption.

      We have added this to the figure legend; “both estimated with Gctf on projection images of the sample at effective zero-tilt position.”

      (3) Figure 3: These are compelling results, and the 3D classification outcomes provide an excellent visual representation of the quantitative data shown in Figure 3. Including (some or all) of the initial five classes in the figure would further strengthen this already convincing presentation. Additionally, applying a uniform extraction threshold (e.g., z-score of 3.75 or 5) across all tilt increments would facilitate a more direct comparison. But this is really a minor remark, the authors and the editors may judge if the current presentation is already sufficient.

      We have adjusted Figure 3 according to the reviewer’s recommendation:

      We referenced this in the main text as:

      “To ensure similar data processing strategies for all conditions, extraction thresholds were also lowered for the 1-, 2- and 3-degree conditions, and 3D classification was performed to filter out the junk particles (see Fig. 3 A and SFig. 8). “

      In order to directly compare the particle extraction, we have already carried out such a uniform extraction threshold, with a z-score threshold of 5 (apart for the 10-degree data, where this was not possible). This extraction was then used for the 3D classification that led to the TM analysis and further STA investigations.

      Reviewer #2 (Recommendations for the authors):

      (1) Supplementary Figure 1: To improve accessibility for a broader readership, the authors should annotate or highlight the key organelles and protein complexes visible in the tomographic slices.

      We thank the reviewer for this suggestion, but adding arrowheads made this figure too crowded in our opinion. Furthermore, for most of the panels, the mentioned features of interest is centred in the image panel, which should make identification straightforward.

      (2) 'In situ' should be in italics throughout the text.

      We have changed this.

    1. eLife Assessment

      This is a valuable study on the metabolic adaptations upon succinate dehydrogenase loss in cancer cells. If confirmed, this study will offer some therapeutic vulnerabilities in treating SDH-deficient cancer. However, the evidence supporting the authors' claim is incomplete and would benefit from additional experimental evidence. This study will be of broad interest for cancer biologists focusing on metabolism.

    2. Reviewer #1 (Public review):

      Different studies have proposed distinct mechanisms by which succinate dehydrogenase (SDH)-deficient cells escape aspartate limitation, highlighting metabolic heterogeneity across experimental systems. In this study, the authors address these previously conflicting observations by longitudinally tracking the adaptation of multiple SDHB-knockout clones derived from the same parental cell line.

      The authors identify two distinct adaptive mechanisms: complex I suppression with predominantly GOT1-dependent aspartate synthesis, and preservation of complex I activity with increased PC-GOT2-dependent aspartate synthesis. They further define shared and unique dependencies associated with these adaptive states, providing a rationale for potential therapeutic targeting strategies.

      Overall, this is a strong study in cancer metabolism, integrating complementary longitudinal and mechanistic approaches, including long-term adaptation, isotope tracing, genetic perturbation, metabolomics, and functional cell growth assays. Although the study provides substantial mechanistic insight, several limitations remain.

      (1) MPC is proposed as a shared dependency of both adaptive states. Testing whether MPC inhibition suppresses SDH-deficient tumor growth in vivo would substantially strengthen the therapeutic relevance.

      (2) The distinction between complex I-intact and complex I-suppressed states is based mainly on the expression of two complex I subunits and the oxygen consumption. More direct assays of complex I activity or assembly are needed. Early-passage SDHB-knockout cells should also be included as controls in the OCR experiments.

      (3) The two adaptive states appear to rely differentially on glucose- versus glutamine-derived aspartate synthesis. Testing the sensitivity of EP and LP clones to glucose or glutamine deprivation would further support this metabolic distinction.

      (4) Since SDH is described as a tumor suppressor, the authors should clarify why SDHB loss initially inhibits hPheo1 cell proliferation.

      (5) The study focuses on SDHB loss, and it remains unclear whether similar adaptive mechanisms arise following loss of other SDH subunits, including SDHA, SDHC, or SDHD, across different biological contexts. This limitation should be discussed explicitly.

    3. Reviewer #2 (Public review):

      In the manuscript entitled "Adaptive plasticity of aspartate metabolism in succinate dehydrogenase-deficient cancer cells," Sokolov et al. delineate the metabolic adaptations that succinate dehydrogenase (SDH)-deficient cancer cells undergo over time to overcome the initial aspartate limitation. To do so, the authors generated five clonal osteosarcoma SDH subunit B (SDHB) knockout cell lines using the CRISPR/Cas9 system and compared the proliferation rates of early- and late-passage cells, revealing that the latter rewired central carbon metabolism to increase aspartate levels and therefore replicate faster than their early-passage counterparts. Using a series of pharmacological and/or genetic interventions, the authors show that this rewiring can occur via two different routes: either through reduced Complex I (CI) activity, whereby glutamine is channelled towards aspartate synthesis via reductive carboxylation, or through a metabolic rewiring in which aspartate is produced from glucose via the PC-GOT2 pathway while CI activity is preserved. The CI-suppression-independent route depends on PC expression, as evidenced by an analysis of DepMap cell-line data, in which higher PC expression is associated with decreased SDH dependency. Moreover, they find that other consequences of aspartate deprivation observed in SDH-deficient cells, including impaired pyrimidine synthesis, replication stress, and DNA damage, are ameliorated in late-passage cells.

      Overall, this study is interesting because it disentangles the different metabolic rewiring routes that SDH-deficient cells can undergo to reverse aspartate limitation and sheds light on previously reported, seemingly contradictory results in the field. However, the study's major premise requires further validation, and important controls are missing, diminishing the overall strength of the conclusions.

      Major points:

      (1) The main conclusion that two separate routes allow SDH-deficient cells to overcome aspartate limitation, defined by their CI-activity status, is not convincingly proven. Indeed, to show this dichotomous behaviour, the authors performed Western blots for two CI subunits and determined the basal oxygen consumption rate. However, these assays are insufficient to demonstrate that LP clones 2 and 3 maintain functional CI, in contrast to LP clone 1. Moreover, it is not ruled out that these clones show dysfunction in ETC complexes other than CI. To assess these points, the activities of all individual ETC complexes should be carefully measured, for instance, by Seahorse assay after permeabilization. Furthermore, given the complex nature of CI, a reduction in two subunits does not necessarily reflect a reduction in its assembly. Therefore, CI assembly should be assessed directly by BN-PAGE analysis of isolated mitochondria.

      (2) It is difficult to reconcile why the authors used an NDUFA8 KO in clone 2 EP to mimic the physiological long-term CI-suppression-dependent adaptation. Indeed, this approach seems to represent an extreme scenario of Complex I loss that may induce non-physiological adaptations that override the effects of SDH KO. To assess the distinct metabolic fluxes between the two proposed routes, it would be advisable to use a more physiological model and instead compare the tracing data from LP clone 2 with those from LP clone 1, which exhibits a "natural" CI-suppressed state. Does clone 1 LP show similar metabolic changes to A8KO, including increased reductive carboxylation?

      (3) It is unclear whether the loss of Complex I at late passage is an intrinsic progression of osteosarcoma cells rather than a feature specific to SDH-deficient cells. A proper comparison between SDHB-deficient cells and WT cells, both at early and late passage, should be carried out. This is essential to fully understand the adaptive trajectories of SDH-deficient cells. This comparison is essential to identify the baseline metabolic hardware of the osteosarcoma cells. Indeed, the authors state that "While wild-type 143B cells synthesize most aspartate from glutamine via oxidative TCA cycling and GOT2 activity, ..." (Page 7, third paragraph), but these data are not included in the manuscript and would represent an important control for assessing the observed metabolic changes in comparison with the wild-type context.

      (4) The data showing that the PC-GOT2 pathway is mainly driven by enhanced PC activity are not fully convincing, as PC activity seems to be equally important for maintaining aspartate levels in the NDUFA8 KO compared with clone 2 LP. Moreover, the extracted expression data from DepMap suggest that increased PC expression might not be transcriptionally regulated, as only a slight association between PC mRNA levels and SDH dependency was observed. Are PC mRNA levels increased in clone 2 LP? If not, PC might be regulated post-transcriptionally. To test this, the nascent translation of PC could be assessed.

    4. Author response:

      We sincerely thank the reviewers for their time, helpful critiques and overall positive evaluation of our work. We plan to investigate the points raised by the reviewers and look forward to submitting a revised manuscript that addresses concerns about therapeutic relevance, functional distinctions between the CI-intact and CI-suppressed adapted states, and PC regulation in this system.

    1. eLife Assessment

      This valuable study documents, for the first time, the degree of preference for contralateral versus ipsilateral forelimb movements in the motor cortex of mice, revealing subtly distinct profiles across cortical areas, with the orofacial region showing comparatively little limb selectivity relative to the forelimb motor areas. The experiments and analyses are, in general, solid and well-designed, and the data support the paper's conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      This study describes motor cortical activity patterns during food handling in mice, investigating whether the hand/s used is reflected in distinct neural activity. The experiments focus on forelimb M1 and M2 (fM1, fM2) and an oral-manual region LOM. The main findings are that fM1 and fM2 have largely similar relationships with forelimb control, and LOM neurons are more broadly tuned. These conclusions are reached using a variety of analyses spanning straightforward firing rate analyses, selectivity metrics, PCA, and GLM decoding methods to assess tuning generalizability. The study's significance is strengthened by including analyses of bimanual control, and in this sphere, there are descriptive data and analyses that aficionados of cortical control of dexterous behaviors will find useful. The use of unimanual control is useful as a point of comparison here, but less novel overall. There are a number of places where the descriptions of what is being analyzed, what is being concluded, and data reporting should be strengthened and clarified. Additionally, the study could be greatly improved by consolidating figures and the analyses shown, since many are redundant. Many of the analyses need clearer reporting of means and effect sizes in the text, rather than just statistical outcomes. Overall, at this juncture, the study presents analyses of a unique dataset that may seed future investigations of mechanisms of bimanual coordination.

      Strengths:

      There are relatively few studies that compare neural activity across bimanual and unimanual control. This study uses a naturalistic food handling task to explore neural relationships to forelimb kinematics under these conditions. The uniqueness of the task and analysis target is a strength of the study.

      The authors remain fairly conservative and make few strong claims in the study, which may be warranted given the diversity of tuning profiles they observed.

      Weaknesses:

      There are a number of statistical tests that were accompanied by too little information to critically evaluate. Means and effect sizes needed to be better reported; some details of analyses were difficult to parse, making the strength of the conclusions difficult to evaluate.

    3. Reviewer #2 (Public review):

      Summary:

      Barrett et al. examine how neural activity in the mouse motor cortex varies when a movement is performed with the ipsilateral or contralateral forelimb. First, they train animals to grasp and manipulate a pellet of food with either the left forepaw, the right forepaw, or both. Next, they measure activity in the primary and secondary forelimb motor areas (fl-M1 and fl-M2) and in the classical tongue-jaw area (tj-M1 / LOM). While responses in the forelimb areas are diverse, with some neurons preferring ipsilateral or bilateral movements, a plurality of cells prefer the contralateral limb. In LOM, by contrast, little limb selectivity is observed. At the neural population level, structure is preserved across conditions in LOM, but not in the forelimb areas. Finally, paw position can be decoded from activity in all three areas, and the LOM decoder generalized across limbs.

      Strengths:

      While previous studies in macaques have compared motor cortical activity during movement (and perturbation) of the contralateral and ipsilateral arms, no analogous work has been undertaken in rodents. This paper closes this knowledge gap by showing, for the first time, moderate-to-strong lateralization in the forelimb motor cortical areas of mice transporting grasped food pellets to the mouth, and a relative absence of lateralization in the classical tongue-jaw area. On the whole, I think this is a solid paper that reports novel observations of interest to the motor systems community.

      Weaknesses:

      The central question posed is whether cortical activity depends on the effector(s) used (ipsi forelimb, contra forelimb, or both). The corresponding hypotheses (Figure 1) are somewhat coarse-grained and are not mutually exclusive. One might expect to see condition-independent, limb-selective, and uni-/bimanual-selective signals in motor cortex (though their magnitudes could differ substantially), and to find these signals intermingled at the level of single neurons. The authors may wish to consider setting up a more focused question. For example, can bimanual responses be explained as a sum of the unimanual responses from the left and right limbs?

      In the area usually identified as tongue-jaw motor cortex (here referred to as LOM), unit and population activity look quite similar for ipsilateral, contralateral, and bilateral forelimb reaches. The most parsimonious explanation is that the activity is related mostly to mouth and tongue movements, rather than limb movements. Systematic mapping studies with microstimulation in the rat (Neafsy et al., Brain Res. Rev. 1986) and optogenetic stimulation in the mouse (Mayrhofer et al., Neuron 2019) tend to support the idea that tjM1/LOM is specialized for control of the tongue and mouth. Thus, I'm not entirely convinced that it "encodes ingestion-related forelimb parameters necessary for oromanual coordination." The authors could say more about this issue: what specific limb-related parameters do they think are encoded, why would these parameters be effector-independent, what evidence for this encoding is presented here, and how can limb- and mouth-related components be distinguished? The problem could potentially be addressed experimentally, as well, by delivering food pellets directly to the mouth while preventing manipulation with the paws, but this experiment isn't strictly necessary.

      Because the corticospinal tract is strongly lateralized, cortical activity presumably has a smaller effect on ipsilateral than contralateral motor output. Somatosensory feedback should also be relatively lateralized for the forelimb areas. The authors could say a bit more about this issue and how it relates to their data and conclusions in the Discussion.

      An important limitation of the behavioral task is that it involves only a single stereotyped movement for each limb, instead of multiple directions, speeds, or loads. This issue and its consequences for the analyses (especially those in Figures 7-10) and conclusions could be discussed.

    4. Reviewer #3 (Public review):

      Summary:

      Barrett et al. compare the responses of different parts of the mouse primary and secondary motor cortex in the context of a task where the animals manipulate and eat food using either or both hands. They find that roughly half the activity is conserved when reaching with one hand vs. the other hand, or with both. Similarity of activity was somewhat higher in the "lateral oral and manual" (LOM) part of the motor cortex, consistent with notions of a more generalized oromanual function there.

      Strengths:

      This work aims at addressing two worthwhile questions in a mouse model of motor control: (1) what specializations do we have for controlling feeding movements, and (2) how are the arms and hands coordinated with one another? The authors develop a simple but innovative apparatus to block either hand during food handling, track the behavior at high temporal fidelity, and record a sizable neural dataset. The analyses come from numerous angles to take good advantage of the data, and succeed in showing multiple lines of evidence for greater invariance in LOM than in the forelimb parts of M1 and M2.

      Weaknesses:

      There are several limitations of the current study. Most importantly, the behavior presents an inherent challenge: there is only one type of movement for each of the three conditions (contra hand, ipsi hand, and bimanual). This is entirely reasonable from the perspective that this is the ethological behavior when feeding, but it limits what analyses are possible. In particular, it precludes disentangling the neural relationship with many correlated aspects of behavior, and limits identifying population-level features of the neural activity meaningfully. This means that there are a number of alternative possible sources of the neuron-level area differences found here, and the population-level features may not be reliable. Second, the behavior tracking was used at a relatively coarse level, and thus the relationships to various behavioral variables were left less distinguishable than they might have been. Finally, there may be an issue with the coordinates of what is being called forelimb M1 here, which may include some hindlimb M1.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      This study describes motor cortical activity patterns during food handling in mice, investigating whether the hand/s used is reflected in distinct neural activity. The experiments focus on forelimb M1 and M2 (fM1, fM2) and an oral-manual region LOM. The main findings are that fM1 and fM2 have largely similar relationships with forelimb control, and LOM neurons are more broadly tuned. These conclusions are reached using a variety of analyses spanning straightforward firing rate analyses, selectivity metrics, PCA, and GLM decoding methods to assess tuning generalizability. The study's significance is strengthened by including analyses of bimanual control, and in this sphere, there are descriptive data and analyses that aficionados of cortical control of dexterous behaviors will find useful. The use of unimanual control is useful as a point of comparison here, but less novel overall. There are a number of places where the descriptions of what is being analyzed, what is being concluded, and data reporting should be strengthened and clarified. Additionally, the study could be greatly improved by consolidating figures and the analyses shown, since many are redundant. Many of the analyses need clearer reporting of means and effect sizes in the text, rather than just statistical outcomes. Overall, at this juncture, the study presents analyses of a unique dataset that may seed future investigations of mechanisms of bimanual coordination.

      Strengths:

      There are relatively few studies that compare neural activity across bimanual and unimanual control. This study uses a naturalistic food handling task to explore neural relationships to forelimb kinematics under these conditions. The uniqueness of the task and analysis target is a strength of the study.

      The authors remain fairly conservative and make few strong claims in the study, which may be warranted given the diversity of tuning profiles they observed.

      Weaknesses:

      There are a number of statistical tests that were accompanied by too little information to critically evaluate. Means and effect sizes needed to be better reported; some details of analyses were difficult to parse, making the strength of the conclusions difficult to evaluate.

      We will improve these aspects of the presentation and reporting of statistical analyses in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Barrett et al. examine how neural activity in the mouse motor cortex varies when a movement is performed with the ipsilateral or contralateral forelimb. First, they train animals to grasp and manipulate a pellet of food with either the left forepaw, the right forepaw, or both. Next, they measure activity in the primary and secondary forelimb motor areas (fl-M1 and fl-M2) and in the classical tongue-jaw area (tj-M1 / LOM). While responses in the forelimb areas are diverse, with some neurons preferring ipsilateral or bilateral movements, a plurality of cells prefer the contralateral limb. In LOM, by contrast, little limb selectivity is observed. At the neural population level, structure is preserved across conditions in LOM, but not in the forelimb areas. Finally, paw position can be decoded from activity in all three areas, and the LOM decoder generalized across limbs.

      Strengths:

      While previous studies in macaques have compared motor cortical activity during movement (and perturbation) of the contralateral and ipsilateral arms, no analogous work has been undertaken in rodents. This paper closes this knowledge gap by showing, for the first time, moderate-to-strong lateralization in the forelimb motor cortical areas of mice transporting grasped food pellets to the mouth, and a relative absence of lateralization in the classical tongue-jaw area. On the whole, I think this is a solid paper that reports novel observations of interest to the motor systems community.

      Weaknesses:

      The central question posed is whether cortical activity depends on the effector(s) used (ipsi forelimb, contra forelimb, or both). The corresponding hypotheses (Figure 1) are somewhat coarse-grained and are not mutually exclusive. One might expect to see condition-independent, limb-selective, and uni-/bimanual-selective signals in motor cortex (though their magnitudes could differ substantially), and to find these signals intermingled at the level of single neurons. The authors may wish to consider setting up a more focused question. For example, can bimanual responses be explained as a sum of the unimanual responses from the left and right limbs?

      In the revised manuscript, we will clarify that the possibilities illustrated in figure 1 are not intended as mutually exclusive. We indeed find all of these signals intermingled. In terms of a more focused question amenable to hypothesis testing, this can be expressed as: for each dimension (laterality vs manuality) are the activity patterns in each area closer to those predicted by invariance or dependence, as compared to the other areas? The various statistical analyses in the paper all essentially boil down to testing this question. Broadly, the answer is yes: we see activity closer to the invariant prediction LOM, and activity closer to the dependent prediction in fl-M1 and fl-M2. Testing whether bimanual activity can be explained as a simple linear sum of the left and right unimanual might provide additional insight into this question, and this analysis will be presented in the revised manuscript.

      In the area usually identified as tongue-jaw motor cortex (here referred to as LOM), unit and population activity look quite similar for ipsilateral, contralateral, and bilateral forelimb reaches. The most parsimonious explanation is that the activity is related mostly to mouth and tongue movements, rather than limb movements. Systematic mapping studies with microstimulation in the rat (Neafsy et al., Brain Res. Rev. 1986) and optogenetic stimulation in the mouse (Mayrhofer et al., Neuron 2019) tend to support the idea that tjM1/LOM is specialized for control of the tongue and mouth. Thus, I'm not entirely convinced that it "encodes ingestion-related forelimb parameters necessary for oromanual coordination." The authors could say more about this issue: what specific limb-related parameters do they think are encoded, why would these parameters be effector-independent, what evidence for this encoding is presented here, and how can limb- and mouth-related components be distinguished? The problem could potentially be addressed experimentally, as well, by delivering food pellets directly to the mouth while preventing manipulation with the paws, but this experiment isn't strictly necessary.

      The issue of orofacial movement confounds is an important one that we made a point of addressing in the discussion. There are three main points that we believe cast doubt on this as the most likely explanation for the effector-invariant representation in LOM.

      First, while we do not disagree that LOM has an important role in tongue and jaw control, there is plenty of evidence from mapping and behavioural studies (which we cite in the introduction) that it also plays a role in forelimb motor control as well.

      Second, while we cannot observe all orofacial movements, we have previously shown that the jaw is less active when the hands and LOM are most active (Barrett et al., 2024). Conversely, LOM firing is much lower during chewing, when the tongue and jaw are very active.

      Finally, the correlation between LOM firing and forelimb movements is not merely a coarse-grained one on the timescale of active manipulation vs passive holding phases. LOM firing closely tracks the position of the forelimb(s) on fast timescales and with near-zero lag, giving better decoding than from fl-M1 or fl-M2, as we have shown here and previously (Barrett et al., 2022). If we assume that LOM only encodes orofacial movements, then this result implies that orofacial movements correlate with forelimb position better than fl-M1 or fl-M2 firing correlates with forelimb position.

      The revised manuscript will include an expanded discussion to clarify these and related points.

      Because the corticospinal tract is strongly lateralized, cortical activity presumably has a smaller effect on ipsilateral than contralateral motor output. Somatosensory feedback should also be relatively lateralized for the forelimb areas. The authors could say a bit more about this issue and how it relates to their data and conclusions in the Discussion.

      We will discuss this in the revised manuscript.

      An important limitation of the behavioral task is that it involves only a single stereotyped movement for each limb, instead of multiple directions, speeds, or loads. This issue and its consequences for the analyses (especially those in Figures 7-10) and conclusions could be discussed.

      This limitation applies to the analyses relating to the transport-to-mouth movement (Figures 3-7). The population correlation structure and decoding analyses (Figures 8-10) consider activity throughout the full duration of food handling, which involves a much greater variety of movements (Barrett et al., 2020). Indeed, this was a major motivation for including these analyses. The revised manuscript will clarify this point.

      Reviewer #3 (Public review):

      Summary:

      Barrett et al. compare the responses of different parts of the mouse primary and secondary motor cortex in the context of a task where the animals manipulate and eat food using either or both hands. They find that roughly half the activity is conserved when reaching with one hand vs. the other hand, or with both. Similarity of activity was somewhat higher in the "lateral oral and manual" (LOM) part of the motor cortex, consistent with notions of a more generalized oromanual function there.

      Strengths:

      This work aims at addressing two worthwhile questions in a mouse model of motor control: (1) what specializations do we have for controlling feeding movements, and (2) how are the arms and hands coordinated with one another? The authors develop a simple but innovative apparatus to block either hand during food handling, track the behavior at high temporal fidelity, and record a sizable neural dataset. The analyses come from numerous angles to take good advantage of the data, and succeed in showing multiple lines of evidence for greater invariance in LOM than in the forelimb parts of M1 and M2.

      Weaknesses:

      There are several limitations of the current study. Most importantly, the behavior presents an inherent challenge: there is only one type of movement for each of the three conditions (contra hand, ipsi hand, and bimanual).

      See our response to reviewer #2 above regarding the variety of movements. We agree that this a limitation for the unit-level and PCA analyses, but one that is alleviated by the population correlation and decoding analyses, which relate to complex ongoing movements.

      This is entirely reasonable from the perspective that this is the ethological behavior when feeding, but it limits what analyses are possible. In particular, it precludes disentangling the neural relationship with many correlated aspects of behavior, and limits identifying population-level features of the neural activity meaningfully. This means that there are a number of alternative possible sources of the neuron-level area differences found here, and the population-level features may not be reliable.

      Important behavioural confounds include orofacial movements, non-specific movement initiation signals, and arousal. Orofacial movements we have discussed above in our response to Reviewer 2. Movement initiation signals would likely be transient and well-timed to movement onset, hence this may explain some of the effector-independent activity in fl-M1 and fl-M2 (consistent with e.g. (Kaufman et al., 2016)). However, we do not believe this to be the case in LOM as its activity is delayed and sustained relative to movement initiation. Regarding arousal, the mouse is actively engaged in consuming the food even when the hands are stationary, so there is no a priori reason to believe that arousal varies rapidly during the behaviour. Consistent with this, measurements of noradrenergic activity in the locus coeruleus during consumption suggest that arousal varies on slow timescales, on the order of seconds to tens of seconds (Sciolino et al., 2022). Such slow variation in arousal would not explain the rapid but condition-invariant changes in firing in any of the cortical areas studied here. The updated manuscript will include more detailed discussion of these points.

      Second, the behavior tracking was used at a relatively coarse level, and thus the relationships to various behavioral variables were left less distinguishable than they might have been.

      Behavior tracking was performed with kilohertz temporal resolution and submillimeter spatial resolution.

      Finally, there may be an issue with the coordinates of what is being called forelimb M1 here, which may include some hindlimb M1.

      Our recording coordinates are based on the territory of corticospinal neurons retrogradely labeled from C6 spinal cord, medial to any layer 4 labelling, as reported in our previous study (Yamawaki et al., 2021). Thus we are confident in calling this area forelimb M1.

      References:

      Barrett, J. M., Martin, M. E., Gao, M., Druzinsky, R. E., Miri, A., & Shepherd, G. M. G. (2024). Hand-jaw coordination as mice handle food is organized around intrinsic structure-function relationships. The Journal of Neuroscience, 44(42), e0856242024.

      Barrett, J. M., Martin, M. E., & Shepherd, G. M. G. (2022). Manipulation-specific cortical activity as mice handle food. Current Biology, 32(22), 4842-4853.e6.

      Barrett, J. M., Tapies, M. G. R., & Shepherd, G. M. G. (2020). Manual dexterity of mice during food-handling involves the thumb and a set of fast basic movements. PLOS ONE, 15(1), e0226774.

      Kaufman, M. T., Seely, J. S., Sussillo, D., Ryu, S. I., Shenoy, K. V., & Churchland, M. M. (2016). The Largest Response Component in the Motor Cortex Reflects Movement Timing but Not Movement Type. eNeuro, 3(4).

      Sciolino, N. R., Hsiang, M., Mazzone, C. M., Wilson, L. R., Plummer, N. W., Amin, J., Smith, K. G., McGee, C. A., Fry, S. A., Yang, C. X., Powell, J. M., Bruchas, M. R., Kravitz, A. V., Cushman, J. D., Krashes, M. J., Cui, G., & Jensen, P. (2022). Natural locus coeruleus dynamics during feeding. Science Advances, 8(33), eabn9134.

      Yamawaki, N., Raineri Tapies, M. G., Stults, A., Smith, G. A., & Shepherd, G. M. (2021). Circuit organization of the excitatory sensorimotor loop through hand/forelimb S1 and M1. eLife, 10, e66836.

    1. eLife Assessment

      This manuscript employs cryo-EM, mutational analysis, and biochemical assays to explore the molecular basis by which glutamine promotes filamentation and regulates the activity of human glutamine synthetase (hGS) by stabilizing interactions between hGS decamers. The combined structural and biochemical data supporting the proposed mechanism are solid, although the evidence for a higher-order filamentous architecture and the precise assignment of the interface density would benefit from additional support. This work will be of particular interest and useful to groups interested in understanding the molecular basis of nutrient sensing, cellular metabolism, and structural regulation of enzyme activity.

    2. Reviewer #1 (Public review):

      Summary:

      The study is methodologically solid and introduces a compelling regulatory model. However, several mechanistic aspects and interpretations require clarification or additional experimental support to strengthen the conclusions.

      Strengths:

      (1) The manuscript presents a compelling structural and biochemical analysis of human glutamine synthetase, offering novel insights into product-induced filamentation.

      (2) The combination of cryo-EM, mutational analysis, and molecular dynamics provides a multifaceted view of filament assembly and enzyme regulation.

      (3) The contrast between human and E. coli GS filamentation mechanisms highlights a potentially unique mode of metabolic feedback in higher organisms.

      Comment on revised version.

      The authors have addressed all of my comments and concerns. The revisions have substantially improved the quality of the manuscript. I have no further questions or concerns.

    3. Reviewer #2 (Public review):

      Major concern 1: The manuscript does not clearly establish a bona fide GS filament state.

      The authors repeatedly refer to GS "filaments," but the data presented appear to support primarily a di-decameric assembly rather than a well-defined filamentous polymer.

      A two-decamer reconstruction can define a putative inter-decamer interface, but it cannot by itself demonstrate propagation of a repeating filament geometry. To establish a bona fide filament, the authors should provide evidence for a reproducible one-dimensional assembly, such as at least three consecutive repeating units or equivalent quantitative evidence that the same inter-decamer transform propagates along an assembly axis.

      In the current manuscript, many of the supporting 2D classifications appear to contain at most two adjacent GS decamers. This is particularly evident in the time-resolved cryo-EM datasets shown in Supplementary Figures 9-10, where I do not see convincing 2D classes corresponding to filaments. The same concern applies to other datasets, including Supplementary Figures 2, 6, and 11, where the apparent assemblies are primarily two-decamer particles.

      Moreover, many of the selected "filament" classes show only one well-resolved GS decamer, while the neighboring decamer density is blurred. This suggests substantial variability in the relative position and/or orientation of adjacent decamers. Such heterogeneity is difficult to reconcile with a stable repeating filament geometry.

      Therefore, the authors should explicitly define what they mean by "filament." If their evidence supports only a di-decameric or short oligomeric assembly, the terminology should be changed accordingly throughout the manuscript.

      Symmetry concern

      Given the low quality and heterogeneity of the 2D classifications for the putative "filament" classes, the use of D5 symmetry requires stronger justification. The current reconstructions primarily show the result after applying D5 symmetry to a two-decamer assembly. The authors should show reconstructions of the same particle sets processed under C1, C5, and D5 symmetry, and explain why D5 symmetry is justified.

      This is particularly important because the claimed interface density and ligand interpretation are sensitive to symmetry averaging. Without showing how the reconstruction behaves under less restrictive symmetry assumptions, it is difficult to determine whether the final D5 map reflects a true biological assembly or a symmetry-imposed interpretation.

      Filament abundance and physiological relevance

      Even under the authors' broad classification criteria, the filament-like population appears to be a minor species. In some datasets, especially Supplementary Figure 10, the apparent filament fraction is very low, approximately 2-10%. This raises a major concern: if GS filaments are rare even under high-concentration cryo-EM conditions, are they expected to form to a meaningful extent under physiological conditions?

      The authors propose a concentration-dependent assembly mechanism. If so, the relevance of GS filamentation in the lower-concentration cellular environment becomes even less clear. The authors should quantify filament abundance as a function of GS concentration and glutamine concentration, ideally under conditions closer to physiological ranges.

      K52/C53 interface mutations

      The authors use K52 and C53 as filament-interface residues, but the mechanistic contribution of these residues to filament assembly remains insufficiently explained. Why should K52A or C53A disrupt filament formation? Is the effect due to loss of a specific side-chain contact, altered local electrostatics, reduced crosslinker accessibility/reactivity, local structural destabilization, or nonspecific disruption of the interface?

      The manuscript states that the interface is "concentration dependent and driven primarily by electrostatic interactions," but the data presented before that statement do not clearly establish this. The authors should explicitly identify the interacting electrostatic partners and provide structural or biochemical evidence supporting this interpretation.

      Functional linkage between filamentation and kinetics is weak.

      The authors should establish the oligomeric state of GS under the actual assay conditions. In particular, what is the filament fraction during the Figure 2F / Supplementary Figure 15 kinetic assays? Is the change in KM ammonia quantitatively correlated with filament abundance?

      This is currently unclear. The direct comparison between decamer and 2-decamer fractions does not robustly show a functional difference, and the later glutamine-addition assays are interpreted as filament-mediated without directly demonstrating the filament fraction under the same assay conditions.

      In Figure 2F, WT, K52A, and C53A already show different ammonia-dependent kinetic parameters in the absence of added glutamine. K52A and C53A appear to have lower basal kcat/KM ammonia and higher KM ammonia than WT even without glutamine. The authors should explain why these interface mutants already alter basal ammonia kinetics. Without such an explanation, K52A and C53A cannot be treated as clean controls that selectively disrupt glutamine-stabilized filamentation.

      Major concern 2: The interface density is not convincingly assigned to glutamine.

      The second foundational issue is the assignment of the interface density to glutamine. At present, the evidence is not sufficient to support the conclusion that glutamine is the ligand at this interface.

      The local density at the interface appears weak and likely has lower local resolution than the reported global resolution. The current density could represent a low-occupancy or symmetry-averaged amino-acid-like density rather than a confidently assigned glutamine molecule.

      Ligand pose and hydrogen bonding.

      The proposed glutamine pose also requires more rigorous validation. The authors state that glutamine forms hydrogen bonds with interface residues, including K52, C53, and E55. These hydrogen bonds should be shown explicitly in a figure, with distances listed.

      The proposed interaction involving C53 appears unusual and should be justified chemically and geometrically.

      Glutamate has not been excluded.

      The largest problem is that the authors do not adequately consider glutamate as an alternative ligand. They compare the density with phosphate and ATP/ADP, but this is not sufficient. Glutamate is present at high concentration during turnover, and it is chemically and structurally very similar to glutamine. Given the limited local density and possible orientational averaging, distinguishing glutamine from glutamate from the current cryo-EM density alone is not justified.

      The authors should report or estimate the concentrations of glutamate and glutamine at the vitrification time point used for the high-resolution turnover-filament reconstruction. If glutamate is present at a much higher concentration than glutamine, the authors must explain why the interface density should be assigned to glutamine rather than glutamate.

      The authors should fit both glutamine and glutamate into the interface density using the same validation criteria and compare the results. Stronger support would come from direct structural experiments, such as cryo-EM structures of GS incubated separately with glutamate and glutamine under controlled conditions.

      Unless stronger evidence is provided, the claim that "glutamine binds to the filament interface" cannot be made.

      Specific comments

      Interface assembly statement:<br /> "These data suggest that the formation of the interface is concentration dependent and driven primarily by electrostatic interactions."

      What specific data support "concentration dependent" at this point in the manuscript? Which residues or chemical groups are proposed to form the electrostatic interactions? The authors should provide a more explicit explanation.

      Line 149-150:<br /> "In both scenarios, any signal is likely to be averaged out and experiments with symmetry expansion and focused classification did not yield any convincing density."

      Please show these analyses. Negative results are important here because they bear directly on the reliability of the interface interpretation.

      "Glutamine stabilizes larger GS filaments":<br /> What does "larger" mean? Longer filaments, more decamers per filament, or larger diameter? The authors should define this quantitatively, preferably by reporting filament-length distributions or the number of decamers per assembly.

      Filament classification:<br /> The criteria used to classify particles or 2D classes as "filament" are not sufficiently clear. The authors should provide the full 2D classification results for each time-resolved dataset, including selected and discarded classes, particle numbers, and objective selection criteria. Some selected and discarded classes appear visually similar, especially in Supplementary Figures 9-10.

      R298A decamer:<br /> The R298A mutant is presented as a turnover-decamer structure, not a filament structure. The authors should clarify whether R298A forms filament-like particles under comparable turnover conditions. If R298A does not form filaments, this should be reported and explained. If filament-like particles were present but excluded during processing, the authors should provide their abundance and justify why only the decameric form was analyzed. This point matters because R298A is used to connect E305-loop disorder with the proposed filament-associated mechanism, although R298A is a loop-stabilization mutant rather than a filament-interface mutant.

      Line 231-233:<br /> "a reaction time that should yield a high concentration of product due to the higher enzyme concentration than previous experiments"

      What is the estimated product concentration at vitrification? What concentration range qualifies as "high"? The authors should provide a quantitative estimate.

      Glutamine hydrogen bonds:<br /> The proposed hydrogen bonds linking glutamine to K52, C53, and E55 should be shown explicitly with atom identities and distances.

      Glutamate comparison:

      What is the glutamate concentration in the same sample? Given that glutamate is chemically similar to glutamine and likely present at high concentration, why is the interface density not glutamate? The authors should compare glutamine and glutamate fitting using the same validation criteria.

      Line 248-253:<br /> The speculation that apo filaments may arise from high GS concentration or residual glutamine should be moved to the Discussion. In the Results, this reads as an ad hoc explanation rather than a result directly supported by data.

      Actual assay-state oligomeric distribution:<br /> What is the filament fraction under the actual kinetic assay conditions? Is the KM ammonia change quantitatively correlated with filament abundance?

      Figure 2F:<br /> Why do WT, K52A, and C53A differ in basal ammonia-dependent activity even without added glutamine? The authors should explain whether these mutations alter intrinsic ammonia kinetics independent of filamentation.

      Supplementary Figure 15 / Figure 2F:<br /> Please clarify the relationship between Figure 2F and Supplementary Figure 15. The kinetic constants in Figure 2F appear to depend on global fitting of progress curves shown in Supplementary Figure 15. The authors should provide replicate-level raw progress curves, between-replicate variability, fitting residuals, and individual fitted parameters.

      Supplementary Figure 19:<br /> Supplementary Figure 19 should be presented consistently with Supplementary Figure 18, including the corresponding 2D classification results.

      In summary, although the revised manuscript improves the presentation of cryo-EM map processing, the two foundational claims remain unresolved. The current data establish, at most, a di-decameric or filament-like GS assembly, but not a rigorously defined filamentous polymer. In addition, the interface density is not convincingly assigned to glutamine, particularly because glutamate has not been excluded as the most relevant alternative ligand. Since the proposed negative-feedback mechanism depends directly on these two points, the current evidence does not support the strength of the title, abstract, or mechanistic conclusions.

    4. Reviewer #3 (Public review):

      In this manuscript, the authors propose a product-dependent negative-feedback mechanism of human glutamine synthetase, whereby the product glutamine facilitates filament formation, leading to reduced catalytic specificity for ammonia. Using time-resolved cryo-EM, the authors demonstrate filament formation under product-rich conditions. Multiple high-quality structures, including decameric and di-decameric assemblies, were resolved under different biochemical states and combined with MD simulations, revealing that the conformational space of the active site loop is critical for the GS catalysis. The study also includes extensive steady-state kinetic assays, supporting the view that glutamine regulates GS assembly and its catalytic activity. Overall, this is a detailed and comprehensive study. However, I would advise that a few points be addressed and clarified.

      Comments on revised version.

      The revision addresses several reviewer concerns: the authors add sharpened maps, ligand-density panels, symmetry expansion/focused classification, biochemical blank/substrate/TCEP controls, and E305-loop focused classification. The E305-loop part is stronger now: turnover decamer recovers partial E-flap density in few classes, while turnover filament does not.

      My only remaining comment is that - as also the authors agree on the need to integrate density with biochemical data and that local resolution/averaging complicates modeling - I would advise softening the claim regarding glutamine from "glutamine binds" to "density consistent with glutamine/product-associated density". In general, it would be best to avoid overstating atomic certainty at the filament interface and the safest framing is the observed interface density is compatible with glutamine but not independently conclusive.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study is methodologically solid and introduces a compelling regulatory model. However, several mechanistic aspects and interpretations require clarification or additional experimental support to strengthen the conclusions.

      Strengths:

      (1) The manuscript presents a compelling structural and biochemical analysis of human glutamine synthetase, offering novel insights into product-induced filamentation.

      (2) The combination of cryo-EM, mutational analysis, and molecular dynamics provides a multifaceted view of filament assembly and enzyme regulation.

      (3) The contrast between human and E. coli GS filamentation mechanisms highlights a potentially unique mode of metabolic feedback in higher organisms.

      Weaknesses:

      (1) The mechanism underlying spontaneous di-decamer formation in the absence of glutamine is insufficiently explored and lacks quantitative biophysical validation.

      (2) Claims of decamer-only behavior in mutants rely solely on negative-stain EM and are not supported by orthogonal solution-based methods.

      We thank the reviewer for the summary and noting of the strengths. We agree that the evolutionary divergence of metabolic feedback in GS homologs is a fruitful avenue for future studies. With regard to the weaknesses, the di-decamer in the absence of glutamine only forms under high (higher than physiological) concentrations of enzyme. Our primary evidence for the mutant behavior was the lack of crosslinking (Figure 1E), with supplementary support from the negative stain. In the revised version we will soften the language to say “reduced” rather than “did not support” filament formation.

      Reviewer #2 (Public review):

      The authors set out to resolve the high-resolution structure of a glutamine synthetase (GS) decamer using cryo-EM, investigate glutamine binding at the decamer interface, and validate structural observations through biochemical assays of ATP hydrolysis linked to enzyme activity. Their work sits at the intersection of structural and functional biology, aiming to bridge atomic-level details with biological mechanisms - a goal with clear relevance to researchers studying enzyme catalysis and metabolic regulation.

      Strengths and weaknesses of methods and results:

      A key strength of the study lies in its use of cryo-EM, a technique well-suited for resolving large, dynamic macromolecular complexes like the GS decamer. The reported resolutions (down to 2.15 Å) initially suggest the potential for detailed structural insights, such as side-chain interactions and ligand density. However, several methodological limitations significantly undermine the reliability of the results:

      (1) Cryo-EM data processing: The absence of critical details about B-factor sharpening - a standard step to enhance map interpretability - is a major concern. For high-resolution maps (<3 Å), sharpening is typically applied to resolve side-chain features, yet the submitted maps (e.g., those in Figures 1D, 2D, and supplementary figures) appear unprocessed, with density quality inconsistent with the claimed resolutions. This makes it difficult to evaluate whether observed features (e.g., glutamine binding) are genuine or artifacts of unsharpened data.

      (2) Modeling and density consistency: The structural models, particularly for glutamine binding at the decamer interface, do not align with the reported resolution. The maps shown in Figure 2D and Supplementary Figure S7 lack sufficient density to confidently place glutamine or even surrounding residues, conflicting with claims of 2.15 Å resolution. Additionally, fitting a non-symmetric ligand (glutamine) into a symmetry-refined map requires justification, as symmetry constraints may distort ligand placement.

      (3) Biochemical assay controls: While the enzyme activity assays aim to link structure to function, they lack essential controls (e.g., blank reactions without GS or substrates, substrate omission tests) to confirm that ATP hydrolysis is GS-dependent. The use of TCEP, a reducing agent, is also not paired with experiments to rule out unintended effects on the PK/LDH system, further limiting confidence in activity measurements.

      Achievement of aims and support for conclusions:

      The study falls short of convincingly achieving its goals. The claimed high-resolution structural details (e.g., side-chain densities, ligand binding) are not supported by the provided maps, which lack sharpening and show inconsistencies in density quality. Similarly, the biochemical data do not robustly validate the structural claims due to missing controls. As a result, the evidence is insufficient to confirm glutamine binding at the decamer interface or the functional relevance of the observed structural features.

      Likely impact and utility:

      If these methodological gaps are addressed, the work could make a meaningful contribution to the field. A well-resolved GS decamer structure would advance understanding of enzyme assembly and ligand recognition, while validated biochemical assays would strengthen the link between structure and function. Improved data processing and clearer reporting of validation steps would also make the structural data more reliable for the community, providing a resource for future studies on GS or related enzymes.

      We disagree with the reviewer’s overall assessment.

      With regard to sharpening and resolution: we examined sharpened maps and in a revised version will present additional supplementary figures showing these maps side by side. We note that the resolutions reported are global and that the most interesting features are, of course, in the periphery and subject to conformational and compositional heterogeneity. We will include supplementary figures of core side chain densities that are more like what are expected by the reviewer in the revision. With regard to modeling: the apo filament and turnover filament datasets were handled nearly identically. The additional density is therefore likely not artefactual to the symmetry operator - however, the lower resolution in this region noted by the reviewer is worthy of further exploration. The maps are public and we think this is the most plausible interpretation of the density, which we based primarily on the biochemical data and will include more speculation in the version.

      With regard to the biochemical controls: we point the reviewer to Figure S1, which shows that omission of ammonia or glutamate in the wild-type (tagless) system removes any coupling of the reactions. We will perform the additional controls to publication quality in the revised version along with the TCEP control. We note that the reducing agent is present across all experiments, ruling out an effect on any specific result. The inclusion of TCEP is also very standard in other published uses of the Coupled ATPase assay (e.g. PMID: 31778111 and PMID: 32483380 by our first author)

      Additional context:

      Cryo-EM has transformed structural biology by enabling high-resolution analysis of large complexes, but its success hinges on rigorous data processing and validation steps that are critical to ensuring reproducibility. The challenges highlighted here are not unique to this study; they reflect broader issues in the field where incomplete reporting of methods can obscure the reliability of results. By addressing these points, the authors would not only strengthen their current work but also set a positive example for transparent and rigorous structural biology research.

      All the data is public and the reviewer or anyone is free to reinterpret the maps and models - and we encourage that rather than just an interpretation of our static figures. In addition, we will upload the raw micrograph data for the apo filament and turnover filament datasets to EMPIAR prior to submitting the revision.

      Reviewer #3 (Public review):

      In this manuscript, the authors propose a product-dependent negative-feedback mechanism of human glutamine synthetase, whereby the product glutamine facilitates filament formation, leading to reduced catalytic specificity for ammonia. Using time-resolved cryo-EM, the authors demonstrate filament formation under product-rich conditions. Multiple high-quality structures, including decameric and di-decameric assemblies, were resolved under different biochemical states and combined with MD simulations, revealing that the conformational space of the active site loop is critical for the GS catalysis. The study also includes extensive steady-state kinetic assays, supporting the view that glutamine regulates GS assembly and its catalytic activity. Overall, this is a detailed and comprehensive study. However, I would advise that a few points be addressed and clarified.

      (1) In Figure 2D and Supplementary Figure 7, the extra density observed between the two decamers does not appear to have the defining features of a glutamine. A less defined density may be expected given the nature of the complex, but even though mutagenesis assays were performed to support this assignment, none of these results constitutes direct and conclusive evidence for glutamine binding at this site. I would thus suggest showing the density maps at multiple contour thresholds to allow readers to also better evaluate the various small molecules under turnover conditions that cannot be well fitted based on this density map, helping to provide a more balanced interpretation of the results.

      (2) On the same point regarding the density for the enzyme under turnover conditions, more details should be provided about the symmetry expansion and classification performed, and also show the approximate ratio of reconstructions that include this density. Did you try symmetry expansion followed by focused classification, especially on the interface region?

      (3) The interface between the two decamers of the model needs to be double-checked and reassigned, especially for the residues surrounding the fitted glutamine. For example, the side chain of the Lys residue shown in the attached figure is most likely modeled incorrectly.

      We thank the reviewer for the feedback. As noted above, we will include supplemental figures that show maps at multiple thresholds and sharpening schemes. We noted in the manuscript and above that our interpretation here is based on integrating biochemical evidence alongside the density and will make that even more clear in the revised manuscript. The filaments +/- the putative glutamine density were processed nearly identically, but we will attempt various schemes of focused classification/symmetry expansion in the revision as well. However, we point out that there is extensive averaging there that makes modeling a bit trickier than expected given the global resolution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Limitation to Di-decamer Formation:

      Could the authors clarify why hGS, when visualized by cryo-EM, predominantly forms di-decamers rather than extended filaments? Since glutamine bridges the two decameric rings, one would expect this to promote further polymerization. It remains unclear why longer filaments are not observed under the conditions used. Moreover, data in Figure 1B-1D are not mixed with Gln, and the mechanism of the apo-form filament formation is not clearly discussed. If K52 and C53 are in charge of filamentation, it is supposed to form a long filament, not a stack of 2 decamers. The kinetics of wild type and mutations, including K52A and C53A, are different. This confused me as K52 and C53 don't participate in the reactions of glutamine synthesis. Does data reduce catalytic efficiency with K52A or C53A mutation suggest that di-decameric GS exhibits a greater catalytic turnover rate than the pure decameric GS?

      We thank the reviewer for pointing this out and it was indeed filament length that was a point of curiosity during the study. Prior to preprint, we repeated the freezing conditions under identical turnover conditions and with high protein concentration as reported and indeed found much longer filaments - pointing to the capacity of the system to form larger complexes. However, we only captured a couple of ‘screening’ images and did not collect a full second dataset. Therefore, we hypothesize that filament length may be stochastic based on the subtle differences of individual grid vitrification based on the assumption that decamers within a filament can freely and quickly exchange. However, as shown in the time-resolved cryoEM experiment, the fraction of particles that are characterized as participating in a filament form (length-agnostic metric), does not appear to be subject to individual grid vitrification conditions but rather by experimental conditions (concentration of reaction-derived glutamine).

      The filament formation in apo state was not further explored because the concentrations required to achieve filament formation in this case were supraphysiological.

      We have edited the discussion to emphasize this point more clearly:

      “While the enzyme concentrations to achieve robust filamentation in the absence of glutamine are much higher than observed in cells, the protein concentrations used in our time-resolved cryoEM experiments where filamentation is correlated with accumulation of glutamine are within the range of intracellular GS concentrations in S. cerevisiae (Engel et al. 2025) and human cell lines (Wiśniewski et al. 2014).”

      Regarding the ability of apo-GS to form filaments - we identified that these residues were important based on their structural location at the decamer: decamer interface (Figure 1D; away from the active site as pointed out) and because their individual mutation to alanine attenuated the ability to form higher order filaments (Figure 1E). Therefore, these mutants were crucial controls in the steady-state kinetic experiments reported (Figure 2E). Here, we used exogenous glutamine to seed/stabilize filaments because we identified glutamine as serving this function and, crucially, because we did not observe glutamine occupancy in the active site under turnover conditions (which would suggest an orthosteric feedback inhibition mechanism; Supplementary Figure 10). Under these conditions, the wild-type, filament-competent protein displayed a ~3-fold K<sub>M, ammonia</sub> increase with glutamine addition compared to no glutamine, which when considered in the context of the greater E-305 flap conformational heterogeneity speaks to a model of allosteric feedback inhibition. Importantly, K52A and C53A show no difference in K<sub>M, ammonia</sub> between the glutamine and no glutamine condition, suggesting that attenuation of filament formation at this interface via mutation, eliminates the kinetic deficit. Therefore, K52A and C53A are not product-inhibited in the same manner as wild-type GS.

      We have clarified the discussion to emphasize this result:

      “Importantly, point mutations of the interfacial residues do not show a K<sub>M, ammonia</sub> defect in the presence of glutamine, indicative of the importance of the filament form for product feedback.”

      (2) Origin of Di-decamer Formation in the Absence of Glutamine:

      While the manuscript demonstrates glutamine-stabilized filamentation, the spontaneous formation of di-decamers under apo conditions is not mechanistically explained. The observation of 10-mer, 20-mer, and 40-mer species in Figure 1B should be validated against molecular weight standards or through SEC-MALS. The inference of higher-order oligomers based solely on migration is insufficient. Additional characterization (e.g., SEC-MALS, AUC, or mass photometry) would clarify whether these assemblies are biologically relevant or incidental.

      We thank the reviewer for pointing out the low precision of preparatory size exclusion chromatography assignments of GS molecular weight filament depicted in Figure 1. We have included calibration standards and assignment in Supplementary Figure 1 and updated Figure 1 to include the ambiguity of these assignments in panel B. It is important to note that GS has historically been underestimated in size via these methods and was originally assigned as an octamer for which there was previous consensus (PMID: 10708854). The low concentration requirements of mass photometry preclude its use for this purpose and we are not in a position to do SEC-MALS or AUC for this. Hopefully, the negative stain, crosslinking, and cryo-EM results are sufficient to indicate that we have correlated signals with the correct species!

      (3) Validation of Interface Mutants as Decamer-only Species:

      K52A and C53A mutants are used to disrupt di-decamer formation and are shown by negative-stain EM to exist as decamers. While supportive, this is qualitative. The inclusion of quantitative biophysical data (e.g., SEC-MALS or mass photometry) would more convincingly demonstrate that these mutants do not transiently assemble into higher-order oligomers. Furthermore, molecular measurements describing the spatial relationship of interface residues - such as the distance between K52 and E55 or between C53 residues of opposing decamers - would aid interpretation. The use of the term "adjacent" (line 530) is vague and should be made more precise.

      We thank the reviewer for the thoughtful comments. Beyond negative-stain EM we also performed a biochemical validation of the filament interface through bi-functional crosslinking based on the premise that the new filament interface, as defined by the apo-filament structure, presented new/unique pairs of nucleophilic amino acid R-groups in close proximity. We used Bis-sulfosuccinimidyl glutarate (BSG) or bis-maleimoethane (BMOE) to covalently link adjacent primary amines and sulfhydryls respectively (Figure 1E). This experiment defines two key principles of GS filament formation in the absence of glutamine:

      (1) It is concentration dependent. In the wild-type case there is a protein-dependent increase on crosslinking efficiency for both crosslinkers.

      (2) It is dependent on C53 and K52. Mutation of C53 or K52 significantly attenuated crosslinking efficiency.

      To make sure that these results are more prominent, we have now included a table of these intersubunit distances between epsilon amine groups of lysines and gamma sulfhydroxyl of cysteines groups based on the apo-filament structure and labeled this as either participating in the filament interface or not. Furthermore, in line with multiple reviewers comments, we have updated Figure 1D to include a sharpened representation of the map that shows strong side chain density for the amino acid side chains to further support these reported side chain distance measurements.

      We thank the reviewer for pointing out the low precision of the SEC chromatogram interpretation of Figure 1B and the figure has been amended to show filaments of variable length, instead of defined length. We also included a calibration curve to Supplemental Figure 1B and estimated molecular weights. While these estimates of size are lower than ground truth it is important to note that GS has historically displayed smaller than predicted molecular weights via size exclusion chromatography and analytical ultracentrifugation where initial characterization papers defined the oligomeric state as an octamer rather than decamer (PMID: 10708854). These studies and the present indicate potential adherence to resin and/or other factors about the shape of GS that lead to longer retention. Lastly, these SEC procedures were performed as a preparative step rather than for analytical purposes, so resolution was not the ultimate goal.

      (4) Terminology: "Scarless" hGS:

      The term "scarless human glutamine synthetase" is unconventional and potentially confusing. If it refers to the wild-type sequence lacking N- or C-terminal tags or mutations, I recommend using the term "native hGS" for clarity.

      We usually reserve “native” for proteins isolated from the original species and not recombinantly expressed (as here). So we will leave the term scarless in the document.

      (5) Helical Parameters of Filament Assembly:

      The manuscript states a ~26{degree sign} rotation (clockwise or counterclockwise?) between decamers in the filament, yet does not describe how this value was derived. Given that hGS filaments form helices, this parameter could be assessed via helical reconstruction. Is it possible the actual helical twist is ~30{degree sign}, implying 12 stacked decamers per full turn? Please elaborate on how the rotational angle was determined.

      Rotation was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was measured. Helical reconstruction was not pursued in this work owing to the typically short filaments observed in micrographs and the relative ease by which a 20-mer species could be selected via traditional 2D and 3D classification/reconstruction methods.

      We have added to the methods the following to better illustrate this measurement:

      “Decamer: Decamer rotation across the filament interface was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was depicted in Figure 1D.”

      (6) Time-Resolved Cryo-EM and Filament Growth:

      The use of time-resolved cryo-EM is innovative; however, the accessible timescales are relatively short. I suggest complementing this approach with techniques such as dynamic light scattering (DLS) or mass photometry, which allow extended real-time monitoring of filament assembly over longer durations (e.g., hours). These methods can also provide higher temporal resolution and particle size distributions.

      We thank the reviewer for this suggestion and agree that understanding the kinetics of filament formation is critical. While DLS and mass photometry are excellent for monitoring assembly over hours, our data indicates that GS filament formation occurs on a much faster timescale.

      As shown in Figure 2C, when we added ATP and Glutamine directly to GS and vitrified the sample after only 5 minutes, the majority of particles had already formed filaments, indicating that the interaction had reached saturation. This contrasts with our time-resolved experiment, where the kinetics of filament formation were likely rate-limited by the enzymatic generation of glutamine rather than the assembly process itself.

      Consequently, we anticipate that filament assembly occurs on the order of seconds or less—a timescale we interpret as a necessary prerequisite for a rapid and effective cellular feedback mechanism. Therefore, we believe the current cryo-EM data accurately captures the biologically relevant window of assembly.

      (7) Cryo-EM Symmetry Imposition and Loop Flexibility:

      The use of D5 symmetry in cryo-EM reconstructions may obscure conformational heterogeneity in flexible elements, such as the E305 loop. Since the authors used MD simulations to characterize loop dynamics, it would strengthen the study to also perform symmetry expansion followed by non-uniform refinement and alignment-free 3D classification of individual subunits. This could provide experimental validation of the proposed conformational variability.

      We agree with the reviewer that symmetry enforcement can mask conformational heterogeneity, particularly for flexible elements like the E305 loop (the E-flap). To address this, we followed the reviewer’s suggestion and performed symmetry expansion on both the turnover decamer and filament consensus maps. This was followed by focused, alignment-free 3D classification on the asymmetric unit containing the E-flap.

      Our analysis revealed a clear distinction: while 4 out of 12 turnover decamer classes showed partial density for the E-flap (class 3, 7, 8, and 12)—consistent with the flexibility observed in our MD simulations—none of the turnover filament classes demonstrated similar density. To ensure a direct comparison, we utilized a C5-expanded turnover decamer map to maintain an identical asymmetric unit to the D5-expanded turnover filament map. We note here that the turnover decamer consensus volume is different from the deposited map for which no symmetry was applied.

      Despite the different initial symmetries (D5 for filaments vs. C5 for decamers), we utilized a C5-expanded decamer map to maintain an identical asymmetric unit. For transparency, we have uploaded this C5-refined consensus map and all resulting 3D classification maps to Zenodo. We agree that the text is now strengthened given this result and we have added the following to the main text:

      “The differential loop density between turnover-decamer and turnover-filament species was further supported by 3D classification of symmetry-expanded particles, which recovered partial E305-loop density in 4/12 turnover-decamer classes (C5; Supplemental Figure 18) compared to 0/12 classes for the turnover-filament (D5; Supplemental Figure 19).”

      (8) Crosslinking Gel Analysis (Figure 1E):

      The SDS-PAGE gels shown in Figure 1E have molecular weight ladders cropped. For proper interpretation, please include full ladders with size markers and labels. In addition, clarify whether the crosslinked samples were denatured in reducing buffer. Crosslinking efficiency and specificity using BMOE or BSG require verification under reducing conditions (e.g., DTT, β-mercaptoethanol, or TCEP) to confirm covalent linkage between decamers.

      We have added in Figure 1E molecular weight markers estimates to aid in gel interpretation and have included the uncropped gels in Supplementary Figure 1D that contain the full MW ladder. The methods were clarified to indicate that reducing reagent was used in both the crosslinking reaction and all SDS-PAGE samples.

      “Protein samples were diluted to concentrations noted in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.1 mM TCEP) and, reacted with crosslinker to a final concentration of 0.5 mM for 10 mins at room temperature followed by quench in 5X SDS-PAGE sample buffer (225 mM Tris pH 6.8, 50% glycerol, 0.05 % SDS, 0.2 mg/mL bromophenol blue, 1M DTT) supplemented with 100 mM of either NH<sub>4</sub>Cl (to quench BSG reactions only) or DTT (to quench BMOE). Protein concentrations were normalized after quench prior to SDS-PAGE analysis.”

      (9) Missing Reference for NADH-Coupled Assay:

      Line 678-679 refers to an NADH-coupled assay described "previously" without citing a source. Please provide a proper reference to ensure reproducibility.

      The original paper describing the implementation of a coupled-assay to measure ADP production from glutamine synthetase was written by Bennett Shapiro and Eric Stadtman in 1970 and has been included. We will note that the conditions of this assay have been much improved since this time with better buffers, commercially available reagents of combined lactate dehydrogenase and pyruvate kinase, and modern plate readers. We added the following reference:

      “Shapiro, B.M. and Stadtman, E.R., 1970. [130] Glutamine synthetase (Escherichia coli). In Methods in enzymology (Vol. 17, pp. 910-922). Academic Press.”

      (10) Unclear Description of the NADH Assay:

      The stability of NADH is influenced by pH and light exposure. Please specify the pH range used in the assay and whether precautions (e.g., light shielding) were taken. NADH autoxidation at high pH or degradation at low pH could impact assay reliability and should be addressed in the Methods section.

      For clarity and transparency the following text was added to the Methods section.

      “Stocks of ATP, NADH, and phosphoenolpyruvate were made in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.5 mM TCEP) and the pH was adjusted until it reached 7.5 on ice prior to aliquoting, flash freezing, and storage at -80°C in the dark. NADH was only exposed to light upon thawing and assay set-up and no appreciable change in absorbance of control experiments were noted.”

      (11) Ligand Density in Figure 2 and Supplementary Figure 7:

      The density attributed to glutamine, ADP, and phosphate appears broader than expected. Please include cross-correlation (CC) values, estimated occupancies, and Q-factors for ligand fitting. Varying the contour level to assess density consistency would clarify whether the observed volume represents multiple conformations, partial occupancy, or overfitting. A similar concern applies to the cysteine sidechain density.

      We have updated Supplemental Figures to include:

      (1) Globally refined map in comparison to locally refined map where both are sharpened per previous feedback.

      (2) Ligand placement now also include Q-scores and CC values

      We did not include multiple contour levels because these are included in the resolution representative Supplemental Figure and because alternative contours do not influence Q-scores. From this analysis it is apparent that phosphate and ADP are both worse fits to the density.

      Moreover, we have now included Supplementary Table 2 that includes all ligand validation statistics for the reader to evaluate the range of B-factor, CC values, and Q-scores for all ligands in all models.

      (12) Missing Ligand B-factors in Supplementary Table 1:

      The ligand refinement statistics in Supplementary Table 1 are incomplete. Please include B-factors and occupancy values for all ligands.

      We have updated the PDB depositions to include B-factors in .cif files that are now available. We have also included Supplementary Table 2 in the manuscript detailing the ligand statistics for all models including cross-correlation, Qscore, and Bfactor.

      (13) Style and Formatting Issues: format consistently throughout.

      (a) Kinetic Parameters: Please follow the IUPAC and IUBMB-recommended formatting:

      kcat should be italic with subscript.

      KM should be italic K with upright M.

      Use lowercase s<sup>-1</sup>, not uppercase S<sup>-1</sup>.

      Refer to:

      IUBMB enzyme nomenclature guidelines https://iubmb.org/wp-content/uploads/2021/01/Current_IUBMB_recommendations_on_enzyme_nome nclature.pdf

      IUPAC Green Book https://publications.iupac.org/books/gbook/green_book_2ed.pdf

      We have corrected the abbreviations according to the reviewers recommendations.

      (b) Inconsistent Terminology and Typography:

      cryo-EM vs. cryoEM are used inconsistently - standardize throughout.

      FSC 0.143 appears with and without subscript formatting-please unify.

      Line 123: "X-ray" should be capitalized.

      Line 266: CryoEM should cryoEM, lowercase "c"

      Line 571: "100 μg ml<sup>-1</sup>"-use superscript minus; ensure consistency with "mg ml<sup>-1</sup>".

      Lines 605, 606, 626, 628: MgCl<sub>2</sub>-ensure the <sub>2</sub> is subscripted throughout.

      Temperature units (lines 607, 630, 640): Write as "4 {degree sign}C" instead of "4C".

      Microliters (lines 660, 681, 697): Replace "uL" with "μL".

      Line 797: Use superscripts: K<sup>+</sup>, Cl<sup>-</sup>.

      Line 862: CO<sub>2</sub> should appear with subscript.

      We have made all terminology and typography consistent throughout based on these suggestions.

      Reviewer #2 (Recommendations for the authors):

      To strengthen the manuscript and address the methodological and interpretational gaps identified, we recommend the following revisions and additions:

      (1) Data processing and cryo-EM map quality

      (a) B-factor sharpening: Reprocess all cryo-EM maps using standardized B-factor sharpening workflows (e.g., the autoSharpen tool in cryoSPARC or similar methods) to enhance side-chain and ligand density visibility.

      We have updated main and supplementary figures to include sharpened maps. All maps were sharpened using the Autosharpen feature of Phenix, specifically, by half-maps. We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      (b) Document the specific parameters used (e.g., B-factor values, solvent content estimates) in the Methods section to improve transparency.

      See above regarding the additions made to the methods section.

      (c) Map replacement and reanalysis: Replace all figures and supplementary panels displaying raw (unsharpened) maps (e.g., Figures 1D, 2D, 3A/B, 4A, 5B, and Supplementary Figs. 2D, S7B, S10) with the newly sharpened versions. Reanalyze density features (e.g., glutamine binding sites, ATP triphosphate groups) using these revised maps and update results to reflect any changes in interpretation.

      As requested, we have updated figures with sharpened maps and found our original analyses to hold. In particular, we have included multiple metrics of ligand model scoring in Supplementary Figure 7B including Q-score and CC. Additionally, we have included all ligand model statistics in Supplementary Table 2.

      (2) Structural modeling and validation

      (a) Ligand fitting justification: Provide high-resolution ({less than or equal to}3 Å) density slices or side-chain density close-ups (e.g., for phenylalanine rings or glutamine-binding regions) to validate claims of atomic-level detail. For non-symmetric ligands (e.g., glutamine) fitted into symmetry-refined maps, explicitly describe how symmetry constraints were adjusted or applied during fitting (e.g., local symmetry refinement, manual adjustment of ligand orientation) and include validation metrics (e.g., cross-correlation scores, density fit plots) to support the placement.

      We have supplied 5 new supplementary figures to demonstrate the resolution of our sharpened cryo-EM maps (most notably Supplementary Figures 3, 4, 8, 12, 22 and panels in others) . Of particular note is the sharpened map features of R298A decamer under turnover conditions which demonstrates multiple instances of a ring density for aromatic residues.

      See discussion above regarding the placement of glutamine in the interface density and updated handling of symmetry during refinement.

      (b) Ligand density supplements: Include supplementary figures showing representative ligand-density fits (e.g., ATP, glutamine) with clear side-chain or functional group annotations, as is standard in structural biology publications.

      In our revision we have included the following updated figures and figure panels demonstrating ligand density into sharpened maps:

      Turnover Filament Glutamine Ligand: Figure 2D-E (updated representation) and Supplementary Figure 7 (new and updated representations).

      Turnover Filament ATP and Mg(II): Supplementary Figure 10 (updated representation)

      Turnover Decamer ADP and Mg(II): Supplementary Figure 5 (new figure panel)

      Turnover R298A ADP and Mg(II): Supplementary Figure 14 (new figure panel)

      (3) Biochemical assay rigor

      (a) Control experiments: Perform and report the following controls to strengthen enzyme activity claims:

      - A blank control (reaction mixture without GS, ammonia, or glutamate) to quantify background ATP hydrolysis.

      - Substrate omission controls (reactions lacking ammonia or glutamate) to confirm that ATP hydrolysis depends on both substrates and GS catalysis.

      - A TCEP effect control (compare ATP hydrolysis rates with and without TCEP) to rule out reducing agent interference with the PK/LDH coupled assay.

      We have provided blank, substrate omission, and TCEP controls in Supplementary Figure 1. These results demonstrate negligible ATP hydrolysis without complete substrate inclusion and do not indicate any impact from TCEP inclusion.

      (b) Direct activity validation: Consider supplementing the coupled assay with a more direct measure of GS activity (e.g., quantifying inorganic phosphate release via malachite green assay) to cross-validate results.

      On the merits of the PK/LDH coupled assay being used for >55 years to measure steady-state activity of glutamine synthetases and that it is a robust assay as supported by the additional control experiments presented above in Supplementary Figure 1 we have elected not to pursue tedious cross-validation with a non-continuous assay and believe our interpretation of the enzyme kinetic results hold.

      (4) Writing and presentation clarity

      (a) Methods detail: Expand the Methods section to explicitly describe:

      - Cryo-EM data processing steps, including B-factor sharpening parameters, map reconstruction workflows, and any post-processing (e.g., filtering, masking).

      - Criteria used to validate ligand fitting (e.g., density threshold values, manual vs. automated docking).

      See above the revisions made in response to critique from review #1 which we will briefly summarize here:

      We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      Focused masks are represented in Figure 2E, Supplementary Figure 18, and Supplementary Figure 19. The details around focused mask utilization are included in the revised figure captions and the following was included in the Methods.

      “Focused masks were generated in ChimeraX (v.1.7 and above). Focused refinement and 3D classification (3 Å filter resolution, PCA initialization) were performed in cryoSPARC. Strategy of class picking and refinement are noted in Supplementary Figures 18 and 19.”

      Map reconstruction workflows are present in the relevant Supplementary Figures. No post-processing steps beyond map sharpening in Phenix were carried out. In general, human GS represents a straightforward protein to reconstruction via cryo-EM.

      Ligand identification criteria was described throughout the results section. Supplementary Figure 7 was revised to show sharpened density for either globally refined or locally refined maps fit with all three products of the glutamine synthetase reaction (ADP, Pi, and glutamine) individually showing the best CC and Qscore for glutamine. Beyond Supplementary Figure 7 we also combined both biochemical experiments and cryoEM to make this ligand assignment supported by:

      (1) Time-resolved cryo-EM experiments that show increasing filament particles over reaction time (Figure 2B and Supplementary Figures 9 and 10)

      (2) Glutamine+ATP cryo-EM screening (Figure 2C) showing long filaments

      (3) Supplementary Table 2 showing reasonable ligand statistics for glutamine

      To clarify this in the Methods sections we include the following statement:

      “Ligands were placed with ISOLDE (v1.7) and those with >0.5 Qscore and supporting biochemical and/or literature precedent were built.”

      (b) Results framing: In the Results, clearly distinguish between observations supported by sharpened maps and preliminary/unvalidated features. Avoid over interpreting density in unprocessed maps (e.g., referring to "glutamine binding" in Figure 2D without noting current density limitations).

      We have updated our discussion of Figure 2D (and now also Figure 2E) to include discussion of only sharpened maps and noted current density limitations to the interpretation.

      (5) Data and material availability

      (a) Ensure all supporting data are publicly accessible:

      - Upload raw cryo-EM movies, particle stacks, and processed maps to the Electron Microscopy Data Bank (EMDB) with appropriate accession codes.

      - Deposit final atomic models in the Protein Data Bank (PDB) and reference these accession codes in the manuscript.

      We deposited maps and models with accession codes in advance of review. The PDB and EMDB codes are available in Supplementary Table 1.

      Furthermore, for the focused maps and focused classifications that were generated during the review, and for the benefit of not cluttering the PDB/EMDB, we have included these more specific analyses in Zenodo: 10.5281/zenodo.20298855.

      - Provide detailed protocols for biochemical assays (e.g., TCEP handling, enzyme purification) in the Methods or as supplementary information to enable reproducibility.

      See updates above to reviewer #1

      (b) By implementing these revisions, the manuscript will better align with eLife's standards for methodological rigor, transparency, and reproducibility, allowing readers to confidently evaluate the study's contributions to structural and functional biology.

      We agree!

      Reviewer #3 (Recommendations for the authors):

      (1) In line 252, it would be helpful to show negative-stain EM images for each SEC peak, further probing whether any peaks correspond to partially aggregated, as this could affect the measured Kcat and Km.

      We aren’t in a position to do this experiment. We routinely check for aggregation by noting Absorbance at 340nm for non-specific scattering indicative of aggregation and observed no evidence of aggregation in our fractions.

      (2) In Supplementary Figure 6, many of the classes in the "Selected Filament Classes" inset appear to be averages of closely spaced particles, which may bias the calculation and should be excluded. In the "Selected Decamer Classes", I would suggest removing the top-view particle classes, as these particles not only have significantly different ice penetration rates, but are also more difficult to distinguish in 2D classification.

      We agree and have provided an additional, more strenuous cutoff, analysis of the tr-cryo-EM data wherein only classes that show clearly aligned decamers are included and all top views are omitted (Additional Supplementary Figure 6). We are happy to say that even with the more strenuous cutoffs that our main conclusions that filaments increase with forward reaction time holds.

      (3) In line 445, "Figure 5A" should be corrected to "Figure 5B".

      We thank the reviewer for pointing this out and have made the correction.

    1. eLife Assessment

      This work presents important findings on quantifying gene coexpression from spatial omics. These quantification methods have been applied to gastruloid to describe how genes are spatialised. The description of the quantifying tools is characterized by exceptional evidence after a thorough revision.

    2. Reviewer #3 (Public review):

      Summary:

      Triandafillou and colleagues report a single-cell resolved spatial atlas of gene expression of 26 gastruloids. While previous work had analyzed either single-cell gene expression or spatially coarse-grained patterns of gene expression (van den Brink et al, 2020) the authors here use multiplexed sequential RNA FISH (seqFISH) to create the first gastruloid atlas which is simultaneously spatially and cellularly resolved. This atlas adds to a growing list of resources cataloging gastruloid development (see also Suppinger et al 2023).

      To analyze this dataset, the authors also describe a novel analytical framework. Their analysis centers around the 'L-score', which measures the degree to which pairs of genes are either coexpressed or mutually exclusive. While this metric is similar to calculating correlations in gene expressions, it has important differences (including that it can in principle be asymmetric, although the authors symmetrize much of their analysis). In addition to the gene-centric L-metric analysis, the authors also analyze cells in their dataset according to the cell type entropy (an information-theoretical measure of confidence in cell type assignment) and the 'exposure index' (a measure of the similarity of nearest cellular neighbors).

      Using this framework, the authors focus analysis of two major features of development. The first is the differentiation of the bipotent neuromesodermal progenitor (NMP) cells in the posterior of the gastruloid into either presomitic mesoderm (PSM) or spinal cord SC lineages. They use L-metric analysis to compare overlap in marker genes used to separate NMP, PSM, and SC fates. They highlight that L-metric analysis can recover spatial patterns of gene expression (without explicit spatial information) and discern subtle features of marker genes beyond simple binning of cell types (e.g. that Epha5 expression in anterior NMPs may predict future SC differentiation).

      The second is the formation of endothelial (spatial) clusters within the gastruloid. The authors highlight two subtypes of endothelial clusters: (1) smaller clusters within the somitic anterior region, and (2) larger clusters associated with endoderm. While the authors discern some subtle differences in gene expression between these two clusters, their different spatial patterns suggest a potential physiological difference that would not be captured in traditional droplet microfluidic-based scRNAseq pipelines.

      Overall, this manuscript is a sophisticated and technically sound study that will provide a valuable beachhead for future studies of developmental patterning in gastruloids and organoids.

      Strengths:

      The major strengths of this study are the overall technical sophistication of the data set and analysis, as well as its potential generalizability to other developmental systems (both in vitro and in vivo). The data are extensively analyzed and reasonably interpreted, and this atlas makes good use of the variability in gastruloid development to extract statistical structure of developmental processes. The L-score offers a parameter-free tool to analyze transcriptomic datasets that could overcome pitfalls of other approaches.

      Weaknesses:

      The major limitations of this study are the depth and novelty of the developmental processes studied. The authors provide very convincing proof-of-concept that their data set can recover known features of gastruloid development, including NMP differentiation and endothelial development. However, further analysis and/or investigation would be required to discover new principles of gastruloid development and patterning.

      Comments on revised manuscript:

      In their revised manuscript, Triandafillou et al have made substantial updates including analysis of variability with their 26 gastruloid datasets; formalization of the L-score (formerly L-metric) and clarification on its interpretation; and validation of their gastruloid samples (e.g. Hox gene colinearity). They have also clarified and sharpened language throughout the manuscript. With these additions further bolster the usefulness of this study as a resource for the gastruloid field, they do not provide major advances in understanding gastruloid development.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors performed seqFISH in 26 gastruloids and performed a variety of computational analyses on these novel spatial data sets. Whilst the data is valuable and the computational concepts useful (exposure index, L-metric, ...), the article falls short on novelty and is written using a very clunky language, often with contradictory conclusions.

      We thank the reviewer for their comments about the value of our data and computational concepts. We agree with the reviewer’s critical comments and have endeavored to address them all. We believe the resulting manuscript is greatly clarified and improved.

      Major issues:

      (1) The authors did well in explaining and detailing the provenance of data and the individual experiments performed. However, their 26 gastruloid data still constitute a very limited sampling from their total organoids: one experiment pooled 4 plates at an 80-94% success rate; 6 different aggregation experiments were done, making a total of 1843 gastruloids, sampled 26 (~1-2%). A simple IF stain of 2-3 markers in a bigger sample could have given a more accurate picture of specific domains of interest and their proximity. Regardless, more information should be given about the existing samples: variation across experimental batches, differences between 300-cell vs 100-cell gastruloids that were used.

      This omission was an oversight on our part and we thank the reviewer for catching it. We did the following to address this point:

      (1) Added date labels to Figure S1.2d (now S1.1a) so that the proportion correct for each separate experiment is clear:

      (2) We added the raw images of the gastruloids used in the study taken before fixation.

      (3) We segmented these images and quantified metrics of the masks to address differences across samples in morphology. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation. Elongation was measured as 1-(the ratio of the width and the height of the segmented gastruloid area); the code can be found here:

      https://github.com/arjunrajlaboratory/ImageAnalysisProject/blob/1b2f2119f77083c27f58f8c36b14 c48d97ea706c/workers/properties/blobs/blob_metrics_worker/entrypoint.py

      Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300).

      The literature also supports our assumption that combining experiments with different starting numbers of cells would not dramatically affect the results (Bennabi et al. 2024). In this paper they show that only 35 genes were differentially expressed between gastruloids formed from 100 cells and those formed from 300 cells (compared to 319 for those formed from 1200 cells and 336 for those formed from 50 cells, both compared to 300 cells). The same paper also demonstrates that the positioning of gene expression (as measured by IF staining) for several representative genes (Bra and Foxc1) does not significantly differ between gastruloids formed from 100 and 300 cells when normalized for overall size and AP axis length (as we’ve done in this paper as well).

      We have updated the text and revised Figure S1.1 to reflect these changes.

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. The experiment was performed 3 times on different days, so to ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].

      To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (see also Revised Figure S1.1a-c).

      (2) Language in the manuscript should be revised. Overall the manuscript is very long, descriptive and written "impressions and beliefs" are often not adequately justified and indeed can be contradictory, e.g. in Section 1: the title states "cell types' locations ...are consistent", a few sentences down we find "there was substantial variation" and "within range of what would be considered a 'morphologically normal' gastruloid". "quite consistent", "compelling patterning", "we don't believe"... these types of expressions are best avoided and replaced with data or used and bolstered with quantitative numbers such as percentages when a given cutoff is used. Another example: "location of each cell type relative to gastruloid morphology was quite consistent the posterior region ... mainly consisted in NMPs." Given T expression in the posterior, this result phrased as such appears quite inflated, in fact, looking at cell types in Figures S1, 2a/b/c, this reviewer would state they are all but consistent and indeed it takes sophisticated analyses to find a pattern (of sorts) beyond the coarse domains expected!

      We thank the reviewer for their careful reading of our paper and appreciate that the work would be strengthened by increasing the degree to which quantitative measures are used to justify the statement we make. We have made the following changes to the manuscript to address this criticism:

      (1) We more clearly delineate where we are making qualitative descriptions and have removed summary language (like ‘consistent’, ‘normal’, ‘variable’ etc.) from these sections. For example, the section the reviewer refers to originally read:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we characterized the organization of each by mapping where each cell type was found relative to other types and overall morphology. The approximate location of each cell type relative to gastruloid morphology was quite consistent: the posterior region, although highly variable in size (Figure S1.2a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]....”

      And now reads:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we first qualitatively examined where each cell type was found relative to other types and overall morphology. The posterior region, although variable in size (Figure S1.3a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]...”

      (2) We follow this qualitative description with a quantitative analysis of cell type proportion where we clearly state which variable aspects are statistically significant:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d).”

      (3) We added a summary paragraph at the conclusion of the results from the first two figures which clearly states which aspects of gastruloid organization we find to be variable and which are consistent, with statistical testing:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (3) Figure 6 is one of the most valuable parts of the work, as the authors use the battery of analyses developed to investigate the variable and not-so-robust endothelial clusters in gastruloids. However, this investigation is still very preliminary, and it should be further linked with known biology. It is still unclear what the unique organization of this cell type is (circularity isn't convincing) and whether any signalling cues of adjacent cells could explain it. Is there any evidence that more mature endodermal cell types are generated (like the suggested "liver") to give rise to endothelial cells? It would certainly be interesting to perform IF for this cell type together with mesodermal and endodermal markers to validate seqFISH predictions on a bigger sample.

      We appreciate the reviewer pointing out that the comparisons between different endothelial cell types was interesting, and agree that the clustering methods were insufficiently justified and that a more explicit consideration of the signaling context of the gastruloid could strengthen our findings.

      We have re-evaluated how we calculate differentially expressed genes. We restricted our analysis to only consider genes that are expressed at > 2 counts/cell in at least 50% of the subsets considered. The results are in shown in the revised Figure 6.

      We find that, as the reviewer suggested, some signaling genes are significantly differentially expressed. Specifically, Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in notch signaling in the posterior of the gastruloid. Tek, on the other hand, is more expressed in the somite-associated endothelial cells, and Tek has been annotated to be involved in retinoic acid signaling. These findings align with the reviewer’s observation that signaling from adjacent cells could explain or relate to differentially expressed genes.

      We also did a more thorough review of the literature, and found several papers that reported unique subsets of endothelial precursors, albeit in related systems. In [Rossi 2022] and [Rossi 2021] the authors find a population of endoderm-associated endothelial cells in gastruloids grown with a different protocol that involves Matrigel embedding, treatment with factors that promote blood development, and growth for 168 hours. In [Veenlveit 2020] the authors find a unique somite-associated population of endothelial cells in Trunk-Like Structures, which are similar to gastruloids but model later in development and have more physical organization with discrete somites. To address the reviewer’s request that we further link with known biology we have added the following to the text:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which show more tissue-like organization than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      Finally, we tested several methods of clustering and calculating circularity, and determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and any conclusions drawn from the text.

      (4) Figures 1c and 6b need statistical significance assessments.

      We thank the reviewer for pointing out that without significance testing these plots are difficult to interpret. We have removed plot 6b (see response above about removing the circularity assessments). For plot 1c we appreciate that it is difficult to interpret which cell types vary more than others in their occurrence without significance testing. To address this we did two things: we first calculated the coefficient of variation for the proportion of each cell type across samples:

      Author response image 1.

      To calculate significance, we first considered that since these values are proportions, they must sum to 1 and changes in one cell type will affect at least one other cell type within the same sample. To properly account for this when applying statistical tests, we calculated the CLR-transformed proportion and tested all pairs of cell types for significant variation. The results are shown in the Author response image 2:

      Author response image 2.

      We added a plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We also address said variation in the text:

      “We sought to quantify variability in cell type composition between gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centred log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.3a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.4).

      The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.4d).”

      (5) The article should include an analysis of Hox colinearity expression in these gastruloids as a validation of the system.

      We thank the reviewer for pointing out the importance of these genes in validating the gastruloid system and agree that assessing their expression specifically would help readers assess data quality.

      We analyzed the center of mass of expression along the AP axis for the Hox genes included in our panel (Hoxb6, Hoxc10, Hoxd1, Hoxb9, Hoxc8, Hoxc6, Hoxaas3, and Hoxb1). We highlighted these genes in Figure S1.3: their expression along the AP axis is consistent with previously reported expression in the tomoseq dataset from [van den Brink 2020]. A summary of the correlation coefficients for each individual gastruloid for all genes (blue) and the Hox genes (orange) is shown in the Figure S1.2. The Hox genes have similar correlation coefficients overall, although their variation is higher. This is likely due to differences in gastruloid pseudo-age; in future experiments we plan to include more Hox genes and use their expression to further classify gastruloids (see updated Figure S1.2).

      We have updated the text with these new results:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an ambitious and technically challenging spatial-transcriptomic atlas of 26 gastruloids using seqFISH. The authors introduce quantitative metrics (mixing score, exposure index, L-metric / scL-metric, spatial L-metric, triplets) to characterize spatial organization at multiple scales. The dataset is valuable, and several analyses are original, particularly the rank-based L-metric family for mutual exclusivity.

      Strengths:

      The authors generate one of the most detailed spatial transcriptomic datasets of gastruloids to date. They propose creative computational metrics (L-metric/scL-metric) to quantify mutual exclusivity of gene expression without predefined thresholds, and they explore organizational principles from single-cell topology to cluster-level structure. Many observations align well with known gastruloid biology, such as posterior robustness and anterior variability. The writing is generally clear, and the figures are rich.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      Several central claims rely on metrics whose computation and justification are insufficiently explained, making it difficult to assess how robust or interpretable the results are. Many choices in the analysis appear arbitrary or are insufficiently motivated (normalization schemes, choice of parameters such as the number of neighbors, the distance cutoffs, hierarchical clustering setup, and so on). The interpretations of spatial consistency, gene-program inference, and endothelial heterogeneity are plausible but might be stronger than the evidence currently supports.

      The manuscript would benefit from stronger benchmarking, quantification of uncertainty, and explicit controls for known artifacts in spatial transcriptomics (e.g., spillover, 2D slicing, cell type assignment entropy). The biological insights are promising, but since several depend on methodological assumptions that have not yet been demonstrated to be stable, they would benefit from clearer methodological explanation.

      We thank the reviewer for spending time to give constructive and actionable comments, and we believe the manuscript is greatly strengthened and more consistent and clear as a result of the changes suggested.

      The work is rich and could become a reference dataset. Then, clarifying and validating the quantitative methods will considerably strengthen the impact and reliability of the conclusions.

      Reviewer #3 (Public review):

      Summary:

      Triandafillou and colleagues report a single-cell resolved spatial atlas of gene expression of 26 gastruloids. While previous work had analyzed either single-cell gene expression or spatially coarse-grained patterns of gene expression (van den Brink et al, 2020), the authors here use multiplexed sequential RNA FISH (seqFISH) to create the first gastruloid atlas, which is simultaneously spatially and cellularly resolved. This atlas adds to a growing list of resources cataloging gastruloid development (see also Suppinger et al 2023).

      To analyze this dataset, the authors also describe a novel analytical framework. Their analysis centers around the 'L-metric', which measures the degree to which pairs of genes are either coexpressed or mutually exclusive. While this metric is similar to calculating correlations in gene expressions, it has important differences (including that it can, in principle, be asymmetric; although the authors symmetrize much of their analysis). In addition to the gene-centric L-metric analysis, the authors also analyze cells in their dataset according to the cell type entropy (an information-theoretical measure of confidence in cell type assignment) and the 'exposure index' (a measure of the similarity of nearest cellular neighbors).

      Using this framework, the authors focus their analysis on two major features of development. The first is the differentiation of the bipotent neuromesodermal progenitor (NMP) cells in the posterior of the gastruloid into either presomitic mesoderm (PSM) or spinal cord SC lineages. They use L-metric analysis to compare overlap in marker genes used to separate NMP, PSM, and SC fates. They highlight that L-metric analysis can recover spatial patterns of gene expression (without explicit spatial information) and discern subtle features of marker genes beyond simple binning of cell types (e.g., that Epha5 expression in anterior NMPs may predict future SC differentiation).

      The second is the formation of endothelial (spatial) clusters within the gastruloid. The authors highlight two subtypes of endothelial clusters: (1) smaller clusters within the somitic anterior region, and (2) larger clusters associated with endoderm. While the authors discern some subtle differences in gene expression between these two clusters, their different spatial patterns suggest a potential physiological difference that would not be captured in traditional droplet microfluidic-based scRNAseq pipelines.

      Overall, this manuscript is a sophisticated and technically sound study that will provide a valuable beachhead for future studies of developmental patterning in gastruloids and organoids.

      Strengths:

      The major strengths of this study are the overall technical sophistication of the data set and analysis, as well as its potential generalizability to other developmental systems (both in vitro and in vivo). The data are extensively analyzed and reasonably interpreted, and this atlas makes good use of the variability in gastruloid development to extract the statistical structure of developmental processes. The L-metric offers a parameter-free tool to analyze transcriptomic datasets that could overcome the pitfalls of other approaches.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      The major limitations of this study are the depth and novelty of the developmental processes studied. The authors provide very convincing proof-of-concept that their dataset can recover known features of gastruloid development, including NMP differentiation and endothelial development. However, further analysis and/or investigation would be required to discover new principles of gastruloid development and patterning.

      We agree that the developmental processes studied here are not inherently novel, and we hope that by showing sufficient overlap with different, less highly resolved methods we have created a convincing document that highlights the potential for this technique to be used to analyze other systems. We appreciate the reviewer’s comments and that the manuscript is improved after making the suggested changes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) S1.2 plates shown individually, but unclear from which experiment.

      The reviewer was right to point out this oversight — we have updated the figure (now S1.1a) with labels for the individual experiments:

      (2) Figure 2 could include a clearer indication of the types of triples/ doublets to make it even more informative.

      We thank the reviewer for pointing out that the types of triples were not clear — we’ve added a color key and more explanatory text to this figure in order which explicitly explains the type of triplets considered.

      (3) Figures should be presented in order. Figure 3c is before 3b, etc.

      We appreciate the reviewer’s attention to detail and have swapped these two panels so that their order in the figure reflects the order they are referenced in the text.

      (4) Figure 3 is interesting, and the L-metric appears useful to pinpoint crucial genes that, when expressed, indicate a type transition has occurred. It would be great to test this with another set of cell types besides NMPs/presomitic/spinal cord.

      We thank the reviewer for their interest in this biological transition, and agree that testing on another transition would be really interesting. There isn’t another set of cell types expected at this stage of gastruloid development that are predicted to have the same type of bifurcating differentiation. However, in an effort to address the spirit of this comment (that looking at other sets of cell types with the L-score would be interesting), we have used our analytical framework on a non-spatial, single-cell dataset from van den Brink et al 2020.

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (5) Figures 4 and 5 felt exploratory, and I would recommend combining them into a single figure highlighting the usefulness of L-metric and its spatial version.

      We thank the reviewer for the suggestion to merge the content of Figures 4 and 5. While we agree that they are thematically related, given the size of the heatmaps generated in the analyses, we were unable to combine them in a way that preserved the readability of the figure and stayed within the space constraints of the page size; thus, we have chosen to keep these as separate figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Quantification methods require clearer formalization and justification

      A key limitation is that the manuscript relies on several spatial metrics whose definitions are not sufficiently formalized.

      (a) To evaluate biological interpretations, the reader needs a precise description of:

      - How each metric is calculated (mixing score, exposure index, scL-metric, spatial L-metric).

      - Why specific choices were made (normalizations, parameter values, distance thresholds).

      - What the expected ranges and interpretations are.

      We agree that these elements are crucial to interpreting quantitative metrics and thank the reviewer for their close read of the work. The reviewer had many comments on the exposure/mixing values calculated for Figures 1 and 2, and for the L-metric (now called L-score) values calculated for Figures 3, 4, and 5. To address the reviewers concerns we have done the following:

      (1) Created a new, unified framework for calculating cell type exposure and mixing.

      (2) Rewritten the methods section for this section with a particular emphasis on including elements the reviewer suggested, including specifically outlining normalizations and what they account for, distances chosen and the biological rationale behind them, and the range of values expected for each measure:

      “Quantification of Cell Type Spatial Relationships: Exposure index

      To characterize the spatial organization of cell types, we computed two related metrics: a pairwise exposure index matrix capturing type-specific spatial relationships, and a scalar mixing index summarizing overall spatial integration. For each cell, we identified neighbours as all cells whose centroids fell within a specified radius r of the focal cell's centroid. We chose a value of r of 16 μm, which gave an average of 5-6 neighbors per cell. We chose this value as it captures local interactions, which was the primary goal of these analyses. For each ordered pair of cell types (s, t), we calculated the exposure index as the proportion of type s cells' neighbors that are type t:

      where Ns→t denotes the count of neighbour pairs in which the focal cell is type ‘s’ and the neighbour is type ‘t’, and Ns denotes the total number of neighbours across all type ‘s’ cells. Each row of the resulting exposure matrix sums to unity and represents a probability distribution over neighbour types for a given source type. To account for differences in cell type abundance, we normalized exposure indices relative to the expectation under random spatial arrangement:

      Where pt is the proportion of cells that are type ‘t’. Normalized values of zero indicate exposure consistent with random mixing, positive values indicate spatial attraction (co-localization), and negative values indicate spatial avoidance, with a minimum of −1 representing complete exclusion.”

      “Quantification of Cell Type Spatial Relationships: Mixing index

      To summarize overall spatial integration across all cell types, we computed a mixing index defined as the fraction of neighbour pairs involving different cell types:

      where N_cross is the number of neighbour pairs involving cells of different types and N_total is the total number of neighbour pairs. For normalization, we compared the observed mixing to the expectation under random spatial arrangement:

      where

      is the expected cross-type interaction rate given cell type proportions. Normalized values of zero indicate random spatial mixing, positive values indicate greater integration than expected (hyper-mixing), and negative values indicate spatial segregation.

      Exclusion of untyped cells. When computing the mixing index, cells lacking confident type assignments were optionally excluded from both the numerator and denominator, ensuring the metric reflects only spatial relationships among typed cells. These cells were retained in the exposure matrix to quantify how typed cells interact with unclassified cells.

      Statistical analysis of variance. To identify cell type pairs whose spatial relationships varied significantly across samples, we computed the variance in exposure indices across samples for each type pair. To account for the expected relationship between mean exposure and variance, we regressed log-variance against log-absolute-mean across all type pairs and computed residuals. Type pairs with residual variance exceeding the 97.5th percentile (two-tailed α = 0.05) were considered significantly variable, indicating spatial relationships that differ across samples beyond what is expected from sampling variation and composition differences.”

      (3) We have also rewritten the methods for how the L-score is calculated, adding emphasis to where we normalize, what ranges of values are expected, and what the interpretation of these values are:

      “Calculating the single-cell L-score

      The single-cell L-score (scL-score) was computed for each ordered gene pair (gene A, gene B) within a single gastruloid. The cell-by-gene expression matrix was filtered to retain only cells with at least 2 detected transcripts for both genes. Cells were sorted in descending order by gene A's expression values; gene A thus serves as the reference distribution, and the score is asymmetric with respect to gene order.

      Three reference distributions were constructed for gene B: (1) perfect coexpression, in which gene B's values were sorted in the same descending order as gene A; (2) perfect mutual exclusivity, in which gene B's values were sorted in ascending order; and (3) independence, in which every cell was assigned the mean expression value of gene B.

      Cumulative sums of expression values were computed for gene A, for the observed expression of gene B, and for each reference distribution. The cumulative sum of each gene B distribution (observed and reference) was then plotted against the cumulative sum of gene A. This cumulative-sum-versus-cumulative-sum representation captures how gene B's expression accumulates relative to gene A's: if gene B's expression is concentrated in the same high-expressing cells as gene A, gene B's cumulative curve rises steeply at first; if concentrated in opposite cells, the curve rises steeply at the end. The area under each curve was calculated using trapezoidal integration and normalized by the product of gene A's and gene B's total expression, yielding four normalized areas: A_observed (observed relationship), A_positive (perfect coexpression), A_negative (perfect mutual exclusivity), and A_uniform (independence). This normalization ensures that scores are comparable across gene pairs with different overall expression levels (Figure S3.2a,b). The scL-score was then defined as follows:

      If A_observed > A_uniform: scL-score = (A_observed − A_uniform) / (A_positive − A_uniform), yielding values in (0, 1].

      If A_observed = A_uniform: scL-score = 0.

      If A_observed < A_uniform: scL-score = −(A_observed − A_uniform) / (A_negative − A_uniform), yielding values in [−1, 0).

      A score of +1 indicates perfect coexpression, −1 indicates perfect mutual exclusivity, and 0 indicates independence.

      The scL-score was computed for all gene pairs in each gastruloid from the 05/07/2025 dataset (n = 18 gastruloids). The other two datasets (n = 8 gastruloids) were excluded due to lower transcript detection quality. To generate an average scL-score matrix, the analysis was restricted to a common set of 202 genes well-detected across all 18 gastruloids, and per-gastruloid matrices were averaged. Unless otherwise noted, a symmetrized scL-score was used: scL-score_sym(A, B) = [scL-score(A, B) + scL-score(B, A)] / 2.

      Calculating the spatial L-score

      The spatial L-score extends the scL-score to spatial regions. For each gene, a kernel density estimate (KDE) was fitted over all detected transcript spots and evaluated on a regular square grid spanning the gastruloid. Bin side length was set to twice the median nearest-neighbour distance between detected spots, calculated separately for each gastruloid. Spatial bins were ranked by KDE-derived density and processed identically to the scL-score calculation. Low-density bins were not filtered, as KDE smoothing produced non-zero density values throughout the imaging area. The spatial L-score was symmetrized as for the scL-score, except when displaying asymmetric heatmaps.

      Hierarchical clustering of L-score matrices

      Gene-gene distances were defined as Euclidean distances between L-score vectors. Agglomerative hierarchical clustering was performed using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean'). This approach operates on L-score vector differences rather than on pairwise L-score values directly, and therefore does not require the L-score itself to satisfy the properties of a mathematical distance metric; the Euclidean distance between L-score vectors is non-negative and symmetric by construction, satisfying the requirements of Ward's method. Heatmaps display pairwise L-score values, not vector distances.

      We applied this clustering procedure to the following gene sets:

      (1) A subset of NMP, presomitic mesoderm, and spinal cord marker genes in one gastruloid (n = 36 genes; Figure 3g).

      (2) All well-detected genes excluding cell cycle genes, averaged across all gastruloids (n = 166 genes; Figure 4 and Figure S4.3a).

      (3) All well-detected genes common to all gastruloids, averaged across gastruloids (n = 202 genes; Figures S4.1a, S4.3b, S5.1a).

      (4) All well-detected genes excluding cell cycle genes in one example gastruloid (n = 171 genes; Figures 5b, S5.2a).

      (5) All well-detected genes common between our seqFISH panel and those that were detected in > 3 cells in scRNA-seq data from [XXX] (n=207 genes; Figure S4.5a).

      (6) All genes detected in > 3 cells in scRNA-seq data from [XXX] (n=19075 genes, Figure S4.5c-f).

      In Figure S4.2a,b a transformation of the L-metric values was used to cluster genes. The scL-scores were averaged across gastruloids as described above, and then each pairwise scL-score was transformed to a distance-like value with(1 - scL)/2. The matrix was then symmetrized as described previously. Agglomerative hierarchical clustering was performed directly on this transformed gene-gene distance matrix using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean').”

      (4) We have added an illustrative figure about how the L-score is calculated which defines expected behaviour for several cases, gives a visual explanation of the process, and shows several extreme behaviours and their biological interpretation (see Revised Figure S3.2).

      (b) For example:

      - Mixing score: The normalization is unclear. Why only 1-nearest neighbor instead of k-NN? Why not consider existing spatial-autocorrelation metrics such as Moran's I, which would also apply to gene-level mixing?

      - Exposure index: The normalization makes the metric unbounded (e.g., exposure > 1 when local frequency > global frequency). It is unclear whether this behavior is intended. Since exposure to self is meaningful, the same metric could replace the mixing score and simplify the framework. The choice of k = 5 is not justified; parameter-free approaches like Delaunay triangulation could avoid arbitrary cutoffs. If k-nn is preferred, then the robustness of the score to change the k value should be studied.

      These issues make it difficult to interpret the magnitude of reported effects or compare them across studies.

      The reviewer makes an excellent point — we have completely overhauled this analysis in the following way to address the issues raised:

      (1) Created one unified metric (see points 1 and 2 above) which considers for every cell, the identity of its neighbors in a 16 um radius (on average 5 or 6 neighbors for each cell in each gastruloid). This value was chosen so that in most cases, the cells in the immediate vicinity of a cell were considered, but not those further out (i.e. the measure is sensitive to close interactions rather than far ones). We made a matrix of all interaction pairs for a given gastruloid, with diagonal elements representing within-type interactions and off-diagonal elements representing cross-type interactions. We have replaced the previous description with the following:

      “Several studies of gene expression in gastruloids have used pooled measurements to infer the AP axis-location of genes and cell types [3,7,13,24,27] and our data are largely consistent with these lower-resolution findings (Figure S1.3a). Yet it is obvious from individual gene staining [1,4,8,28] and our detailed 2D maps of cell identity and location that gastruloid organization is much more complex than the average order of cells along the AP axis. We thus needed an analytical method for quantifying spatial organization beyond distributions along the AP axis. To further characterize spatial organization, we sought to quantify the degree to which cells were mixed in each gastruloid, and how that mixing might vary between gastruloids. For each cell in each gastruloid, we counted the interactions between that cell and all its neighbours within a 16 μm radius (on average 5-6 neighbors per cell), and summarized all these interactions for all cells in the gastruloid in a matrix, normalizing each element by the frequency of the cell type considered to be the ‘neighbour’ in the interaction.”

      (2) To quantify overall mixing, we calculate the sum across types of the frequency of self interactions (normalized to the total interactions) and then take the inverse (1-M). We then normalize this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of that cell type).

      Because the density of cells is fairly consistent across the gastruloids, w is very close to p (the proportion of that type).

      is the expectation of cross-type rate with random mixing.

      Mixing ranges from -1 (totally segregated) to +1 (more mixed than random, i.e. there is attraction between unlike types). 0 is completely random, and negative values indicate that cell types within that gastruloid tend to cluster. We added the following to the text to explain this:

      “To quantify overall mixing, we calculated the sum (across types) of the frequency of self interactions, normalized to the total interactions) and then took the inverse. We normalized this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of the neighbouring cell type). This gave us, for each gastruloid, a value that we call the mixing index that ranged from -1 (totally segregated) to +1 (totally mixed with less frequent self-interactions than expected from chance). A mixing index of 0 indicates a random distribution, i.e., neighbour frequency is exactly what would be predicted by that cell type’s frequency alone.”

      We also edited the following description of the overall distribution of mixing indices:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). All gastruloids had a negative mixing index, indicating that they all, on average, had more like-cell type interactions than would be expected given random mixing of types. However, we note that there is a ~14% difference in the mixing index across gastruloids, meaning some variation in overall mixing is present.”

      (3) The exposure index for a given pair can be found from the off-diagonal elements of the interaction matrix, and the normalization means it represents relative overexposure/clustering (positive values) or underexposure/avoidance (negative values). The minimum value is -1 and the maximum is (1-pt)/pt. We changed the description of how the exposure index is calculated to reflect this unified method of quantification:

      We noted that the off-diagonal elements of the matrix we used to calculate the mixing index were informative about cell type-cell type interactions. Specifically they quantify the degree to which each cell type (source) is exposed to another cell type (neighbours). To assess the overall frequency of cell type-cell type interactions, we first pooled the data from all gastruloids together into one interaction matrix (Figure 2b). The measure can range from -1 (no interactions at all), with higher values indicating a greater frequency of being found in close proximity. It is asymmetric in that the exposure of cell type A to B may not be the same as the exposure of cell type B to A.

      We have updated all of the quantification in Figures 1 and 2 with these new measures.

      (2) L-metric: unclear justification and interoperability

      (a) First of all, even if it is not a major issue, the L-metric is not a "metric" at least in the mathematical sense of a metric since a metric is always positive. The L-metric is central to several major conclusions (gene exclusivity, modules, spatial organization), but its conceptual basis and computational steps need more justification.

      We thank the reviewer for pointing this out and have changed “L-metric” to “L-score” throughout. We have also endeavoured to clarify the conceptual basis and have fleshed out the various computational steps as outlined in more detail in our responses below.

      (b) Several steps (ranking, cumulative curves, area under the curve) are difficult to interpret biologically It is unclear why each transformation is required and how it responds to common scenarios (highly expressed genes, correlated vs mutually exclusive patterns).

      Since the metric is rank-based, two genes that are both highly expressed in all cells may show low L-metric despite being truly correlated.

      We appreciate the reviewer’s comments about both the interpretation of the scL-score calculation and how it behaves in common expression scenarios, particularly for genes that are broadly expressed across many cells. To address these points, we generated a set of simulated examples spanning five scenarios: ubiquitously expressed genes with similarly high average expression, ubiquitously expressed genes with differing average expression, ubiquitously expressed genes with similarly low average expression, genes coexpressed across a subset of cells rather than all cells, and genes generally expressed in opposite subsets of cells. For the first three simulations, we independently sampled two genes across 50 cells using Poisson distributions with mean expression set to 100 or 50 (to simulate a gene with high or low average expression, respectively), without any expression bias towards any subsets of cells. For the latter two simulations of dependent expression relationships, we first sampled gene 1 across 50 cells using a Poisson distribution with mean expression set to 4, then generated gene 2 from gene 1 by sampling from cell-specific Poisson distributions fitted to either generally match or oppose the transcript count obtained for gene 1 in that cell. These simulations show that genes can independently appear broadly coexpressed at the population level (regardless of average expression) simply by being ubiquitously expressed, yet still receive low scL-score values. The simulations of dependent coexpression or mutually exclusive expression relationships receive scL-score values near +1 and -1, respectively. These results align with the reviewer’s prediction, but they reflect why we designed the L-metric to follow a rank-based methodology since they preserve the specificity of the metric’s upper bound (+1) for detecting non-random coexpression relationships rather than chance coexpression relationships resulting from independently ubiquitous expression. The results of these simulations are depicted (See Revised Figure 3.3).

      We have made the following edits to the text to specifically address the case the reviewer raised about highly expressed genes:

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3.”

      (c) Interpretation of L-metric values is ambiguous

      What does 0 represent? Randoms? Co-expression? Is 1 the strongest exclusivity? The manuscript currently mixes "co-expression" and "mutual exclusivity" scales.

      We agree with the reviewer that clearly defining what values of the L-score mean is critical to understanding the text. We have added a more explicit discussion of this in the text (excerpted from the response to 2c):

      “A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes.”

      And made a figure representing visually how the L-score is calculated which shows the behaviour and biological interpretation of several extreme values and an example of how the L-score is calculated. See new Figure S3.2:

      We also edited figure captions where we referred to plots as ‘co-expression’ plots, since in some cases the plots showed genes that were mutually exclusive or not expressed together in most cells. Figure S3.1:

      “c. Spatial distribution of the expression of Nkx1-2 and Rfx4 in an example gastruloid.”

      In all other cases we checked, we used the term “co-expression” to mean the opposite of mutually exclusive, as outlined in the definition above.

      (d) Additional issues also limit interpretability - Benchmarking is missing.

      - No tests on synthetic datasets, negative controls, or curated examples.

      - Prior exclusivity methods (EEI, COTAN) routinely benchmark against ground truth; this is now standard.

      - The code for the L-metric seems to be missing in the repository.

      We appreciate this suggestion offered by the reviewer as benchmarking against prior exclusivity-oriented methods provides an important comparison for clarifying both where the scL-score agrees with existing approaches and where it offers distinct advantages. To address this, we explicitly compared the scL-score to the Exclusively Expressed Index (EEI), which is bounded below by 0 and increases with mutual exclusivity, and to the signed coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework, in which positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. We performed this comparison using seven simulated scenarios as well as four representative gene pairs from one gastruloid sample (2025-05-07_roi2). In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (high_high, high_low, low_low), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained low, and COEX was also positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.

      We then applied the same comparison to four gene pairs from one gastruloid sample (2025-05-07_roi2). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types. Pax6-Eogt, Rfx4-Eogt, and Nkx1-2-Rfx4 all showed opposing expression by scL-score, with values of -0.572 and -0.526, -0.970 and -0.955, and -0.537 and -0.615, respectively. EEI detected exclusivity most strongly for Rfx4-Eogt (0.171), but gave values of 0 or approximately 0 for the other two pairs, while COEX was negative for all three pairs (Pax6-Eogt: -0.235, Rfx4-Eogt: -0.350, and Nkx1-2-Rfx4: -0.106), consistent with opposing expression, with strongest signal for Rfx4-Eogt. By contrast, Cdx4-Cdx2 showed moderate levels of coexpression by scL-score (0.402 and 0.424) and EEI (0), but COEX indicated that they were not coexpressed (-0.331).

      These comparisons also clarify the practical advantage of the L-metric over the EEI and COTAN frameworks. EEI is based on binary zero/non-zero quantification and is therefore designed specifically to measure exclusivity rather than coexpression. COEX provides a signed value and, in our simulations, tracked both exclusivity and coexpression; however, on representative gene pairs from one gastruloid sample, scL and COEX diverged in magnitude for highly exclusive expression relationships (Rfx4-Eogt) and sign for a coexpression relationship (Cdx4-Cdx2), motivating our introduction of a signed measure based on the mutual exclusivity of expression with the scL-score (see New Figures S3.3 and S3.4).

      We have updated the text to address the reviewer’s concerns: we benchmark using simulations of commonly occurring scenarios (such as varying expression levels, degree of mutual exclusivity, and amount of noise present in the relationship between the two genes in question) as the reviewer suggested. We also provided a direct comparison to an earlier exclusivity measure (EEI):

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3 and Figure 3.4.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We compared the scL-score to two existing measures of exclusivity. The Exclusively Expressed Index (EEI) (Nakajima et al. 2021) is bounded below by 0 and increases with mutual exclusivity. EEI is computed from binary zero/non-zero quantification and is designed specifically to measure exclusivity but not coexpression. The coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework (Galfrè et al. 2021) can also be used to quantify relationships between genes: positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (but with varying relative expression levels), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained close to 0, and COEX was positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We calculated the scL-score, EEI, and COEX for four gene pairs from one gastruloid sample (2025-05-07_roi2). The scL-score consistently delineated gene pairs possessing opposing expression profiles, while EEI was not always able to measure those exclusivity patterns (Pax6-Eogt: scL-score=-0.572 and -0.526, EEI=0; Rfx4-Eogt: scL-score=-0.970 and -0.955, EEI=0.171; Nkx1-2-Rfx4: scL-score=-0.537 and -0.615, EEI~0). Only the scL-score was able to detect the coexpression pattern present between the positively associated expression profiles of Cdx4 and Cdx2 (scL-score=0.413, EEI=0, COEX -0.331) (shown visually in Figure S3.4a). Thus, while EEI was informative for measuring gene expression relationships characterized by mutual exclusivity, the scL-score more clearly separated positive, random, and mutually exclusive relationships on a single signed bounded scale. The COEX value trended in the opposite direction than expected, but this may be due to the fact that it cannot be calculated on single-gene pairs and necessarily uses information from the entire count table, which here only consisted of 6 genes. These comparisons combined with the simulations in Figure S3.3, clarify a conceptual advantage of the L-metric over the EEI and COTAN frameworks. In contrast, by leveraging the ranked structure of transcript counts across cells, the L-metric framework does not binarize expression and does not require fitting a parametric distribution. It can be calculated on single gene pairs, and is more sensitive to mutual exclusivity.”

      We also amended the Data and Code Availability section to include a specific reference to the L-metric package that was previously missing:

      “All code used to process the raw data and generate figures, as well as the processed data and figures can be found at the following link :

      https://www.dropbox.com/scl/fo/bchkqlbcjb8ub9m606did/AIudcWZaC566toXzb2L-jXc?rlkey=u0wgtkq8j erxoqb5ump599oip&dl=0

      Additional custom scripts used to process the raw seqFISH data can be found on GitHub:

      https://github.com/arjunrajlaboratory/NimbusImage/

      The code for calculating the L-score can be found on GitHub:

      https://github.com/arjunrajlaboratory/l-metric

      Images of all gastruloids generated for this study, as well as single-channel seqFISH images with segmentation and annotations are available here:

      https://app.nimbusimage.com/#/project/69d3fa8f1be4701f5fab6359 Raw seqFISH images are available upon request.”

      (e) Clustering using L-metric vectors

      In this study, the authors use hierarchical clustering to group genes according to the L-metric. This choice is reasonable: hierarchical clustering provides a natural representation of similarity relationships across multiple scales, and the L-metric captures a form of signed dissimilarity between genes. However, this approach raises an important issue. Standard hierarchical clustering methods typically assume a non-negative metric, whereas, as noted earlier, the L-metric can take negative values, meaning it does not strictly satisfy the requirements of a metric in the mathematical sense.

      To the best of our understanding from both the text and the source code, the authors address this issue by defining the distance between two genes A and B as the Euclidean distance between two vectors: the L-metric values from A to all other genes, and from B to all other genes. Although this procedure is mentioned in the manuscript, it is neither justified nor accompanied by any discussion of how such a distance should be interpreted. It is not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes. These two gene distances are not without overlap, but they are not the same.

      This choice magnifies the interpretability issue: readers must understand two layers of transformations. If the [-1,1] range poses problems for hierarchical clustering, simple transformations (e.g., 1 − L) or alternative clustering methods could avoid these issues.

      Given that the L-metric underlies major biological inferences (novel gene modules, spatial subclusters, endothelial states), clearer justification and benchmarking are essential. We think this can lead to more consistency in the spatial metrics.

      We really appreciate that the reviewer took the time to understand our proposed method thoroughly, and apologize for any confusion resulting from a lack of clarity in how it is calculated, and the language used to describe it. The reviewer is absolutely correct that it is not a metric in the mathematical sense; we have replaced the word ‘metric’ with the word ‘score’ throughout the text.

      The reviewer also raised concern about how the hierarchical clustering was performed, and they were absolutely correct about what the vectors represent—they are, as the reviewer states, “not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes”. The heatmaps in figures 3, 4, and 5 are intended to cluster genes that have similar expression patterns, i.e. similar L-score values with all other genes. The reviewer pointed out that this transformation wasn’t clear, so we have added an explicit explanation of what the vectors represent, as well as an explanation of why we were interested in how these vectors, which represent a ‘fingerprint’ of how the gene interacts with all other genes, clustered (because this is in the section discussing NMP differentiation we focus on a specific subset of genes here, but later apply to the entire panel):

      “For each gene annotated as belonging to any of the three cell types (NMP, PSM, or spinal cord), we calculated a vector of scL-score values with all other genes. Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      We also appreciate the reviewer’s suggestion that alternative transformations of the scL-score may improve clustering interpretability. We tried using the reviewer’s suggestion of doing a simple transform: we averaged the scL-score matrices across the 18 gastruloids using the shared gene panels, transformed each scL-score from the original [-1,1] scale to a [0,1] scale using (1-scL)/2, symmetrized the resulting matrix so that each gene pair was represented by a single value, and then performed hierarchical clustering directly on this gene-by-gene distance matrix using average linkage. We used this analysis to test whether a more direct distance-based approach would change the gene groupings recovered by our original clustering method.

      This alternative approach largely recovered the same cell type-associated groupings, but the separation between branches in the dendrogram was smaller, making fine-scale ordering harder to interpret. We measured this by evaluating the average cell type dispersion, measured in terms of additive branch length, which was 0.682 and 0.671 for the 166-gene and 202-gene panels, respectively. Since the branch separation on our original dendrograms was greater and thus representative of more robust groupings, we chose to continue using hierarchical clustering based on Euclidean distances between scL-score vectors.

      We comment on this alternative transformation we tried in the next section, when considering the clustering of the entire gene panel. We feel this is appropriate as the motivation for looking at the difference vector was derived from expected behaviour of a smaller set of genes, and as the reviewer pointed out it is not clear that that expectation should or would hold for the entire panel.

      “To generate the heatmap shown in Figure 4a and S4.1a, we used the same clustering method as described earlier with the Euclidean distance between scL-score vectors. However, we also tried clustering directly on the scL-scores themselves, by transforming each scL-score from the original [-1,1] scale to a [0,1] distance-like scale using a (1-scL)/2 mapping (Figure S4.3a,b). The results were largely consistent, however the cophenetic distance scale was relatively compressed when the transformed values were used (Figure S4.3a,b). This shallow structure implies that many branches are separated by only modest distances, so fine-scale ordering within the dendrogram should be interpreted more cautiously than the larger-scale cell type block structure. We chose to continue using Euclidean distance-based hierarchical clustering of scL-score profiles, where cell type grouping is observed alongside larger cophenetic separations between clusters” (See New Figure S4.3).

      (3) Claims of "remarkably consistent" spatial organization are stronger than the data currently support

      (a) The manuscript emphasizes reproducible organization across gastruloids, but several factors complicate this interpretation.

      We agree that the distinction between what is reproducible/invariant between gastruloids and what varies was not clear in the original manuscript. To address this, we have updated the text to more explicitly distinguish between the two categories. The other suggestions made by the reviewer to sharpen the quantitative measures to strengthen these claims was very helpful and we appreciate the thought put into them, and have used that framework (emphasizing the statistically significant variations and consistencies) in summarizing our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrate several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior:posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (b) Possible selection bias. Only elongated, QC-passing gastruloids were retained; 18/26 datasets remain. seqFISH runs with uneven housekeeping signals were excluded.

      We agree with the reviewer that our data are elongated, QC-passing gastruloids, although these represent two sources of variation (biological and technical respectively). Our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category, and we have attempted to signal this to readers by consistently including language like “morphologically normal” and “elongated”. We have further updated the language in the manuscript to emphasize this point:

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. To ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].”

      In regards to the exclusion of datasets, the only time 18 out of 26 were used was when calculating the averaged L scores for all genes in Figures 4 and 5. In this case we used all 18 gastruloids from the seqFISH run performed on 4/4/2025; this dataset had the highest spot counts due to protocol improvement between runs, and integrating the datasets with very different spot counts was problematic because a minimum expression level is needed to calculate L scores. We used all 26 samples for the spatial metrics calculated in Figures 1 and 2 (Figures 3 and 6 focus on specific gastruloids). We have added additional labels in Figures 1, 2, 3, and 6 to make clear when all 26 datasets are used and when only a subset is used.

      (c) Gastruloids are known to be variable; restricting to morphologically "normal" samples could inflate apparent regularity.

      We thank the reviewer for this observation and agree that the degree of variability among gastruloids is an important consideration. As stated in response to b), our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category. The rate of occurrence of ‘normal’ gastruloids in our hands is ~80% (Figure S1.1a). We agree that it would be interesting to consider how variations from this baseline affect cell type composition and arrangement, and while we make no claims about it in this paper, we have updated the introduction to highlight this point:

      “To address these gaps, and to create a systematic, high-resolution dataset of gene expression in gastruloids considered to be morphologically normal, we developed a spatially resolved, single-cell molecular map of the location, identity, and gene expression of cells within 26 individual gastruloids with normal morphologies. We found that despite some morphological variability within the qualitative category of elongated and polarized, “normal” gastruloids had largely reproducible cell type composition.”

      (d) Partial lack of statistical validation. The manuscript shows descriptive consistency but no formal tests across runs or batches (e.g., mixed-effects models, ICCs, leave-one-run-out validation).

      We agree with the reviewer that a quantitative comparison between batches is important. We have added the following to the text to address this point:

      “To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (See Revised Figure S1.1a-c).

      (e) 2D sampling limitations. Spatial metrics rely on a single imaging plane chosen as the "midplane," but z-position varies between gastruloids. AP projections, mixing, and triplet analyses could all be sensitive to z-plane choice. Prior work shows that 2D slices can underestimate distances and contacts by large margins (https://pmc.ncbi.nlm.nih.gov/articles/PMC5522766). These limitations should be acknowledged explicitly.

      We agree that sampling in 2D can limit the interpretation of our findings and we thank the reviewer for bringing up this important point. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      “Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.”

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2017].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (4) Uncertainty in cell-type assignment is not incorporated into spatial metrics

      Many spatial measurements depend directly on cell-type calls (exposure, triplets, mixing). However:

      (a) Anterior cell types have higher entropy in their marker-based scores (Figure S1.1a).

      We thank the reviewer for pointing out that several of the cell types in the anterior have high entropy — specifically cardiac mesoderm and paraxial mesoderm. However, we think there is nuance to this point; two of the other prominent anterior cell types (somite and endothelial) have low entropy scores overall, and that spinal cord/neural precursor cells, which are mostly posterior, have somewhat higher entropy; higher entropy scores are not exclusive to the anterior, nor is low entropy exclusive to the posterior. Inspired by the reviewer’s comments, we have re-analyzed our data to include this nuance (see response to point d) below.

      (b) Uncertain labels inflate apparent "mixing" or "disorder," because misclassifications randomly create mixed neighbors and triplets.

      We agree that uncertainty in labels could affect the interpretation of mixing. We appreciate these comments and the reviewer’s suggestions, and we have followed them in our response to point d) below.

      (c) Posterior cell types have low entropy, so comparisons between anterior vs posterior mixing may partly reflect label uncertainty, not biology.

      We thank the reviewer for bringing up this important caveat to our findings. We incorporated discussion of this in our text edits (see point d) below.

      (d) The authors should incorporate confidence measures (e.g., probability-weighted neighbors, entropy filtering, bootstrapping) to confirm that patterns hold independently of classification noise.

      We appreciate these suggestions and have chosen to use entropy filtering to assess whether the spatial organization we observe is highly sensitive to what values are considered ‘low’ entropy. We have updated the text (see below), and added Figure S2.2 to address the reviewer’s comments:

      “The contrast between organized posterior clustering and disorganized anterior mixing matches expectations based on literature that shows that self-organization mechanisms in gastruloids in the anterior vs. posterior are differentially sensitive to culture conditions, with somitic patterning requiring external matrix support [1,4,8,24], distinguishing it from the seemingly more autonomous organization observed in posterior cell types.”

      “One potential caveat to this finding is that differences in uncertainty in cell typing could be the primary driver of mixing and cell type interaction differences, both between the anterior and the posterior within an individual gastruloid, or overall between gastruloid. To control for this, we applied an entropy filter to our dataset. We filtered out cells that had entropy > 1.5 (see plot below for cutoff), which was chosen based on the distribution of entropy values for ‘none’ type cells, which effectively describe the upper limit of random transcript assignment (99.7% of ‘none’ type cells are removed with this filter, and about 50% of cardiac mesoderm cells and paraxial mesoderm cells, see Figure S2.2a). We first examined overall mixing; there was strong correlation between the per-gastruloid mixing indices before and after entropy filtering (Pearson r = 0.809, Figure S2.2b). Globally, mixing indices decreased with filtering, meaning that overall the cell types were more clustered. When we compared the absolute value of the change in mixing index pre and post-filtering to the proportion of each cell type, the only significant correlation was with cardiac mesoderm (Figure S2.2c). Exposure indices were overall quite similar after filtering, although the strength of somite-somite and somite-paraxial mesoderm interactions increased (Figure S2.2d).”

      “We also examined how entropy filtering might affect the exposure index, given that the mixing index is calculated from the exposure index of across cell types. In general, the magnitude of the exposure index values increased when more uncertain cells were excluded, but the directionality and relative ordering was not affected. Although the magnitude of change in the posterior cells types was less than the anterior cell types, the cross-cell type exposure values, particularly between paraxial mesoderm/endothelium and somites, doubled. From these results we conclude that mixing in the posterior is driven mainly by NMP/presomitic mesoderm interactions, and is overall lower than mixing in the anterior, which is driven by rarer cell types like cardiac mesoderm, paraxial mesoderm, and endothelium, being interspersed within somite cells” (See Figure S2.2).

      (e) This leads to reviewing the claim on endothelial heterogeneity, which strongly depend on spatial adjacency and gene exclusivity metrics.

      We have extensively considered claims of endothelial cell heterogeneity, and these are discussed in detail in response to the reviewer’s next point. We have also copied them here for the reviewer’s convenience:

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in the Author response image 3:

      Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although many of the somite-associated genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0). See Author response image 4.

      Author response image 4.

      (5) Endothelial "spatially dependent" gene expression may reflect spillover rather than intrinsic state

      The comparison between anterior-associated and posterior-associated endothelial nuclei suggests two transcriptional states. However, spatial adjacency confounds the interpretation:

      (a) seqFISH assigns transcripts to nuclei in dense tissue; partial-volume effects can mix RNA from neighboring endodermal or somitic cells.

      We thank the reviewer for their attention to detail and agree that a careful consideration of these points is important. We also note that given that some of our differentially expressed genes in endothelial cells are endoderm or somite genes, there indeed may be some transcript misassignment.

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript misassignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for none-typed cells at any dilation considered.

      (b) Endoderm and endothelium are closely intermixed (Figure S6.1d), and their gene expression co-localizes in KDE maps (Figure 5b-c).

      We agree with the reviewer and thank them for their close reading of the manuscript. We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      (c) Without explicitly quantifying spillover, differential expression between these two endothelial subsets cannot be confidently attributed to cell-intrinsic differences.

      (d) You may control for spatial proximity with any of the following:

      - Include adjacency index as a covariate in DE models.

      - Use scL-metric to test the mutual exclusivity of endothelial vs endoderm genes within the same nucleus.

      - Apply local permutation nulls: shuffle transcripts within local windows and recompute DE.

      - Restrict analysis to gastruloids that contain both endothelial subsets.

      We thank the reviewer for bringing up this important point, which we were eager to address. We agree with the reviewer’s point that by only considering gastruloids that contain both subsets of endothelial cells is the correct way to do the analysis. We were already only considering this case (specifically the values calculated in Figure 6 and for 1 gastruloid pictured in Figure 6a). We added an n=1 label to revise Figure 6d to emphasize this point.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although some genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0) (See Author response image 4).

      We have also included in the supplement larger images of some of the top differentially expressed genes, which more intuitively show the differential expression results (Endoderm enriched and Somite enriched).

      Finally, we have referenced several previously-reported instances in the literature where distinct subsets of endothelial precursors with unique gene expression programs were identified. Although in these cases 1) the embryo models were different (in [Rossi 2021, Rossi 2022] gastruloids made with a different protocol and treated with factors designed to promote blood development, and in [Veenlveit 2020] trunk-like structures) and 2) the methods were different (IF and 10x single-cell sequencing) this at least establishes a precedent for the observation of multiple types of endothelial precursors. In the case of [Veenlveit 2020] the authors specifically note that one subset is associated with somites, and we have updated the text to reflect these new results:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b,c). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Pecam1 also clustered uniquely in our L-metric analysis (Figure 4c), suggesting this differential expression is conserved across gastruloids. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behaviour is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6)

      (6) Interpretation of gene-program modules may be overstated

      Claims that the L-metric reveals "novel gene programs" should be softened:

      (a) The seqFISH panel is an approx. 200-gene marker-enriched panel, already biased toward known cell-type markers.

      This is true and we appreciate that this came through in the text since it’s important for the reader to understand the approach we took in this study.

      (b) Strong blocks in Figure 4a and S4.1a may reflect panel design rather than newly discovered programs.

      We agree with the reviewer that the panel design was not sufficiently highlighted, so we have made the following changes to the text to emphasize which patterns would be expected due to the genes we are probing for, and which findings were surprising given the known functional role of the gene.

      At the end of the section titled ‘The L-metric captures the spatial distribution of gene expression despite being calculated without spatial information’:

      “These analyses demonstrate that information contained within the hierarchical relationships between genes, determined by scL-score can reveal novel information about cell states within cell types, although we acknowledge that since cell type is determined by a limited panel of marker genes, results should be further functionally verified. scL-score analysis can also identify distinct spatial locations of cells in this cell state, all without explicit encoding of spatial information, but rather quantifying and clustering the degree to which genes are mutually exclusively expressed with one another.”

      At the end of the section titled ‘Clustering scL-metric vectors clearly resolves cell types and reveals novel genetic interactions’:

      “Finally, although the strong blocks we find in the heatmaps in Figures 4a and S4.1a largely reflect cell types, as is consistent with our panel design, we discovered some novel functions of genes in the panel through their location in the scL-score tree: although Tgfβ was initially included in our panel to generally detect inflammatory and growth signaling, clustering by expression patterns revealed its unique association with endothelial precursors.”

      To further address the concern that the generality of clustering is due to gene selection, we performed random gene drop-out and assessed how well cell types clustered as a function of the number of genes removed:

      Author response image 5.

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      Finally, we performed scL-score analysis on an unbiased, scRNA-seq dataset without any panel selection, and were able to show similar groupings of cell type markers:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (c) Cluster robustness is not assessed (bootstrap, stability).

      We appreciate the reviewer’s suggestion that cluster robustness should be assessed. To address this point, we performed a clustering-stability analysis on the common 202-gene panel that included cell cycle genes by asking to what extent hierarchical clustering could recapitulate cell type-based groupings of genes as the number of genes used for clustering was varied across progressively larger, randomly sampled panel subsets. For each resulting tree, we averaged cell types’ dispersion of genes across the tree using the framework described in our response to suggestion 6m and compared the observed value to a permutation-based distribution generated on the same tree.

      This analysis showed that the observed cell-type dispersion remained consistently lower than the corresponding permutation distribution across all subset sizes examined. In other words, genes assigned to the same annotated cell type remained closer together in the dendrogram than expected by chance even when clustering was performed on reduced random subsets of the panel. We also observed that dispersion values increased as larger gene subsets were included, which is expected as the clustering problem becomes more complex with increasing panel size; however, the separation between the permuted distribution of average cell type dispersion and the observed dispersion value increased as the panel subset size increased. Taken together, these results indicate that the cell type-resolved organization captured by the scL-score derived hierarchy is not dependent on one particular subset of genes, but is instead a stable property of the broader gene panel. New Figure S4.2 addressing cluster stability:

      We have added the following explanation in the text:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      (d) Agreement with cNMF (claimed in text) is not quantified (ARI, Jaccard, hypergeometric overlap).

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See Revised Figure S4.2 (now S4.4)).

      We have updated the text to reflect these quantitative comparisons:

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (e) Testing the scL-metric on larger, unbiased scRNA-seq datasets would help demonstrate generality.

      We agree with the reviewer that this would demonstrate generality, so we applied scL-score analysis to the scRNA-seq dataset from van den Brink 2020 — several gastruloids at the same stage of development as those used in this paper were pooled and sequenced. We added a figure, new Figure S4.5 with the results of this analysis, and a new section in the paper describing them:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (7) Minor Comments

      (a) We suggest including representative raw seqFISH images. The manuscript does not show raw images, which makes it difficult to evaluate the quality of the underlying data that all spatial analyses depend on. A figure showing raw fluorescence channels, detected spots, and nuclei segmentation masks for at least one anterior region, one posterior region, and one dense interface (e.g., endoderm-endothelial) would allow readers to assess spot intensity and background levels, signal-to-noise ratio, segmentation accuracy, and potential over/under-segmentation, channel cross-talk, and transcript crowding or dropouts in dense tissues. A small panel of raw images would improve the transferability to the spatial metrics.

      This is an excellent suggestion and we have included examples of the raw images for a representative gene for all hybridizations, the spots as determined by the spot-finding algorithm distributed with the seqFISH instrument, deconvolved spots, and nuclear segmentation. These images can be found in Revised Figure S1.1d.

      (b) The Introduction mentions 3D organization, which can confuse readers into thinking the seqFISH dataset is volumetric. The data shown and analyzed come from a single 2D plane per gastruloid, not from full 3D z-stacks. Since all spatial metrics rely on true adjacency, the manuscript should explicitly state early on that the dataset is 2D and briefly justify why a single plane is sufficient for the analyses.

      We appreciate the reviewer bringing up this point and we have removed references to 3D so as not to confuse readers.

      We also have included an extensive discussion of the 2D nature of the data. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2018].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (c) Results, first paragraph: "good agreement" along the AP axis should be quantified or defined.

      In the first paragraph we compare the peak in gene expression along the (length-normalized) AP axis of each gene with a similar but orthogonally measured dataset from another group (van den Brink 2020). The text specifically reads:

      “When we compared how gene expression varies along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section.”

      We have revised Figure S1.2 with a summary plot showing the distribution of correlation coefficients for all genes and for the Hox genes in our panel (which are known to be expressed sequentially along the AP axis).

      We have updated the text as follows:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      (d) Clarify what "greater cell type distinction" means when using marker panels, and what metric demonstrates improvement?

      By “greater cell type distinction” we meant that when using traditional clustering methods we were not able to individually resolve some cell types: there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together. Cluster labeling was performed by considering which genes showed up as being differentially expressed in each cluster using the same associations as were used to perform the cell type scoring with the marker gene panel. While we acknowledge that these results are subjective to clustering parameters, this is a general problem with clustering and not specific to this study. We have added additional clarification in the text and removed the phrase “greater cell type resolution” since it wasn’t clear what we were comparing to:

      “To profile the spatial organization of individual gastruloids, we assigned a cell type to each nucleus using a cell type scoring method with known marker genes. We compared these results to those obtained with unsupervised clustering. We found that although clustering did produce clusters, they were not strongly separated and we were not able to individually resolve some cell types we expected to find: specifically, there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together (Figure S3.6a). Therefore, we proceeded with the scoring-based method; see Methods for additional details. A representative gallery of typed gastruloids is shown in Figure 1b; the full dataset is in Figure S1.3. Because each cell received a cell type score for each type, we could use the entropy of the cell type score probability distribution to assess confidence of our cell type assignment: a cell that received a similar score for multiple cell types would have a high entropy distribution, while one which scored highly for one type and low for the rest would have low entropy. On average, the cell type entropies for most cells in a given cell type were low, with the exception of paraxial mesoderm and cardiac mesoderm cells, which had intermediate values (see additional discussion below). The generally low entropies indicate that most of our cell type assignments were high-confidence (Figure S1.4a).”

      (e) The variation in "none-typed" cells across gastruloids should be expanded and shown quantitatively.

      We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      (f) Figure S1.1b: explain the criteria used to decide which tissues "varied significantly." Showing the proportion of "None" per gastruloid would help.

      We thank the reviewer for pointing this out. We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      To justify the use of the phrase ‘varied significantly’ we have calculated the degree to which cell type pairs significantly co-vary (adjusted p value < 0.05) in their proportions and added the plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We address said variation in the text:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids profiled. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.32a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.41dc).”

      “The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.41dc).”

      (g) In Figure 1c, showing the variability of "None" cells is important.

      We made this adjustment and added the plot to Figure S1.3c

      (h) For Figure 1e, consider normalizing to T-expression proportion, not only AP length. This could clarify multimodality in the endoderm and spinal cord.

      We thank the reviewer for this suggestion to normalize by molecular as well as physical features. We agree this could be a useful projection of the data, since expression of T is used in many contexts to define the posterior of gastruloids.

      We incorporated T expression into the length normalization in the following way: we generated a cumulative distribution of all T spots along the AP axis, and when this value exceeded a threshold (specifically 90% of all spots) we defined this as the midpoint of the gastruloid, and linearly normalized space before and after it from 0-0.5 and 0.5-1 respectively. We then re-projected nuclei for each gastruloid onto this new coordinate system, and visualized in the same way as the main text figure (See Figure 1f).

      To assess whether this increased or decreased variability in position, we calculated how much the location of the peak of each cell type in each gastruloid differed from the peak position of that cell type in all samples pooled together. We compared this difference to a bootstrapped null drawn from the pooled distribution. Interestingly, we found that nearly all cell types had statistically significant variation (meaning the average distance from the mean fell outside the bootstrapped distribution in the positive direction), however the effect size for most cell types was very small.

      Notably, as the reviewer suggested, normalizing to T expression decreased the effect size of this variability in all cases but one, and particularly decreased spinal cord variability (despite not being a marker for spinal cord):

      We interpret these findings in the following way — that most of the variation observed in cell type arrangement along the AP axis is due to morphological and molecular variability (likely due to stochasticity in initial cell number and differences in developmental timing between gastruloids), and that once these factors are taken into account, for most cell types the effect size of variability is very small. However for some cell types, notably the location of the differentiation front, the position varies among gastruloids. We hypothesize that this may be due to the rhythmic nature of somite development. We have updated the text to reflect these changes.

      “Given that the proportions of cell types within each gastruloid were fairly consistent, to what degree did their spatial organization vary? We first projected each cell’s expression onto the AP axis and looked at the distribution of AP axis locations across gastruloids (Figure 1f, left-hand side). The most posterior cell types (NMP and presomitic mesoderm) showed wide distributions from 0-30% of the axis. Centred around 30%, spinal cord precursors and endoderm had distinctive peaks (clusters), the exact location of which varied between gastruloids. Using a threshold of T expression to define the midpoint of each gastruloid and uniformly length-normalizing each half caused these cell types to collapse into a single peak (Figure 1f, right-hand side). The differentiation front was similarly located in one peak (between 30-40% of the AP axis), and the variation in the location of this peak decreased with T-expression normalization. To quantify this variability, we bootstrapped a null distribution of cell type locations by pooling each cell type together across samples and using the resulting distribution to define a reference mean. When we compared how much peak variation there was among samples randomly drawn from this null to our observed data, we found that although nearly every cell type (excepting paraxial mesoderm) had more variability than expected by chance, the effect size of this variation was small (Figure S1.5a). Normalizing location to T-expression decreased variability in most cell types, most notably for differentiation front and spinal cord (Figure S1.5b).

      We also asked how the order of cell types along the AP axis varied between gastruloids. When we ranked the peaks shown in Figure 1f per gastruloid, we found that Kendall’s W, an overall measure of rank coherence across independent samples that spans from 0 (no agreement) to 1 (complete agreement), was 0.834 (Figure S1.5c). We found that the cell types most likely to swap rank order were spinal cord and endoderm, and paraxial mesoderm and endothelium (Figure S1.5d).”

      (i) Clarify "mixing coefficient values range from 0.29-0.58" (incomplete sentence). Section "Cell types' arrangement is consistent across gastruloids, but varies by type" second paragraph.

      We appreciate that there was some confusion here and we thank the reviewer for pointing it out. We have updated the text with the new values and a more specific interpretation to aid the reader and address the reviewer’s comment:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). While all values are negative, the range was large. This observation led us to conclude that while in all the gastruloids profiled cell types tended to cluster together, there was variation between gastruloids in the degree of coherent clustering between types.”

      (j) The paragraph claiming organization consistency may be too strong, given the lack of statistical validation. "Overall, these data speak to the consistency of gastruloid organization...".

      We agree that quantification of variability, which we only assessed qualitatively, would help readers better evaluate claims of organizational consistency, and we appreciate the reviewer pointing this out.

      We have taken several steps to add quantification, including:

      (1) Calculating the coefficient of variation for the proportion of each cell type

      (2) Quantifying co-variation and assessing statistical significance

      (3) Quantifying the variation in AP-axis location, including comparison to a bootstrapped null, significance testing, and a new form of normalization which decreases some of the variability.

      (4) Quantifying order along the AP axis and calculating Kendall’s W to quantify how concordant this ordering is across gastruloids

      (5) Overhauling the methods used to quantify local patterning and mixing into one unified metric

      (6) Calculating significance for variation in exposure index across samples

      To address the question of whether claims of organization consistency are appropriate, we have added the following summary paragraph at the end of the results for the first two figures, and have taken the reviewer’s suggestion of only discussing statistically significant or quantified results in listing both consistent and variable features. Now rather than making an argument about whether gastruloids are consistent or not, we merely provide the readers with our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (k) Figure 3c: gene set sizes differ substantially; proportion-based normalization may be more appropriate.

      We thank the reviewer for carefully noting the gene set differences. While the NMP-only and spinal cord-only gene sets each have 10 genes, PSM-only and NMP+PSM have 6 genes, and NMP+spinal cord has 5 genes. Given the relatively small N, we feel that normalization would likely introduce a layer of abstraction that would be more confusing for the reader, especially given the qualitative nature of the claims made about the shape of the plots in question. However, we agree this point is important, so we have now indicated the size of the gene sets on the plot (revised Figure 3b, previously c).

      (l) The purpose of the first two graphs in Figure 3c is unclear.

      We thank the reviewer for pointing out that this is unclear. The purpose of this visualization is to demonstrate correlation between genes shared between cell types and genes exclusive to only one of the cell types. The text reads:

      “Figure 3b shows the total expression of each gene group versus the NMP genes for all cells in all gastruloids that we typed as NMP, presomitic mesoderm, or spinal cord. As expected, there was a clear correlation between the mixed categories and NMP genes, supporting the notion of a continuous differentiation process.”

      We have added a correlation line to revised Figure 3b (former Figure 3c) to emphasize this point.

      (m) Clustering of all genes: quantify whether clusters align with cell types.

      We appreciate this suggestion offered by the reviewer as this analysis will allow readers to quantitatively assess the overlap between cell type-based groupings of genes and our scL-score-determined clusters for this dataset. To address this, we have used a bootstrapping approach to evaluate how well our scL-score-based hierarchies capture cell type-based groupings. Specifically, following hierarchical clustering of genes using the scL-score, for each cell type, we computed the average of the minimum cophenetic distance between each pair of genes associated with that cell type. After obtaining each cell type’s average, we took the mean of these averages which we refer to as the cell type dispersion for that tree. We note that the absolute value depends on the topology of the tree. To establish a reference for this measure, we bootstrapped a null distribution by preserving the same tree topology and randomly assigning genes to leaves and calculating the resulting dispersion. We performed this 10,000 times to create a reference null distribution. We compared the true dispersion value to this null, and computed a bootstrapped p-value. In both cases (with and without cell cycle genes), the observed average cell type dispersion is less than that for all permutations (p-value=0.0001, bootstrapping), indicating that genes associated with the same cell type were significantly more likely to cluster together on the observed scL-score hierarchy than would be expected under random clustering.

      We have updated the text to reflect these quantitative comparisons:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type-specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b).

      Author response image 6.

      (n) Clarify how the distance between genes is computed for clustering; the current method is hard to interpret.

      We appreciate the reviewer’s suggestion to clarify how distances between genes were computed for hierarchical clustering. To clarify this point, we have added the following to the text:

      “Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      (o) Quantify similarity between cNMF and L-metric clusters.

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2, now in updated Figure S4.4c). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See updated Figure S4.4 (formerly Figure S4.2)).

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (p) In Figure 3e, clarify what the orange lines represent.

      We have updated the text:

      “Per-cell expression scatterplots of the two pairs of genes shown in b). The y-axis of each is the per-cell expression of Eogt. The x-axis is the per-cell expression of Pax6 (left) or Rfx4 (right). R is Pearson’s r, scL is scL-score. Count data is shown in black; smoothed 2D densities are shown in orange.”

      (q) Ensure figures and panels follow the text order. For example, Figure S1.3a is referenced earlier than 1.1 and 1.2.

      We appreciate the reviewer’s attention to detail. We have changed the order of these figures so that their reference in the text follows their numeric order.

      (r) Typo in "by covariation with any other cell type (Figure S1.1e)" did you mean S1.1.c?

      We have updated the text with this change.

      (s) For circularity (Figure 6), justify the convex-hull-based measure; thin protrusions can distort interpretation. Consider the volume difference between the convex hull and the original shape.

      We thank the reviewer for this helpful suggestion. We tested the difference method suggested by the reviewer, as well as several other methods of clustering and calculating circularity. We determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and text.

      Summary

      This work delivers a rich spatial dataset and introduces creative computational tools. The main limitations lie not in the data but in the clarity, justification, and validation of the quantitative methods. We suggest strengthening these aspects by adding formal definitions, parameter justification, benchmarking, robustness tests, and controlled interpretations. This will improve the manuscript's impact and reproducibility of the methods described, besides making it easier to understand for the readers.

      We thank the reviewer for their kind assessment and also for their many insightful comments for improvement. We feel the revised manuscript is greatly improved because of them.

      Reviewer #3 (Recommendations for the authors):

      In my view, this manuscript is well-designed and clearly written, and supports all of the claims made. I have no suggestions for major revisions for this manuscript; rather, I would suggest the following as outstanding questions for future investigation:

      We thank the reviewer for a careful reading of our manuscript and the several interesting suggestions and useful references in the literature. Including the discussion of these ideas in the text (see below) has improved the flow and scope of the manuscript.

      (1) On the NMP fate, bifurcation has been studied extensively in gastruloids and related structures (see, for example, Underhill et al 2023; Bolondi et al 2024). Can this dataset from Triandafillou and colleagues reveal new regulatory hierarchies in this process? This seems possible in principle, but I was not able to reach this interpretation (for example, it seems the authors interpret the clustering in Figure 3H as reflecting spatial patterns rather than a regulatory hierarchy).

      We agree that this is an exciting implication of the work, but we feel that with the current panel (which was originally chosen primarily to type cells and not to infer regulatory structure) we would not be able to comment on this. However we feel that with a larger set of genes such inference may be possible, so we took your suggestion in point 3) and applied the L-metric to a previously existing single-cell dataset. We looked for possible regulatory structure, and found at least one interesting case where a transcription factor showed strong co-expression with two previously unconnected genes. While more specific analyses and experimental validation would be required to establish a direct relationship, we feel that this suggestion by the reviewer represents an important potential future application of this methodology, so we’ve included a new figure and explanatory text to address this possibility:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.

      (2) On the formation of endothelial clusters ('blood islands'), this process has been observed in gastruloids (Rossi et al 2022). The observation of endoderm-associated endothelium in gastruloids is interesting, but it was also not clear how the authors interpret this finding. Are there two separate endothelial differentiation paths captured here? Or is there one path, and only some migrate towards the endoderm? The authors seem to raise possibilities, and it was slightly unclear on my reading how they interpreted their findings.

      We thank the reviewer for pointing us to this paper. We think it’s especially interesting that we also see close association between endoderm and endothelial precursors, especially given that the protocol used in the referenced paper was designed to generate blood precursors (treatment with VEGF, bFGF and ascorbic acid). In light of the comments of other we re-did this analysis — Author response image 7 shows the updated list of differentially expressed genes.

      Author response image 7.

      Some of the spatially differentially expressed genes are linked to signalling, and likely reflect overall signalling differences between the anterior (where the somite-associated endothelial cells are) and the posterior (where the endoderm-associated endothelial cells are). For example, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA is higher in the anterior). Pecam1 is involved in adhesion and cell morphology, and perhaps is higher in cells interacting with endoderm due to tighter packing/association; the same could also be true of Cdh5.

      While none of these answer the reviewer’s questions about the origin of the cells (and whether it is common), we found another instance of differential endothelial populations in embryo models: in Veenvliet 2020, they find that some endothelial cells have a mesodermal (somitic) origin. Thus we may be seeing a similar phenomenon in our samples. We have updated the text to reflect these additional lines of evidence and to clarify how we think the two populations may differ (while acknowledging that we lack the tools to confidently assign cell of origin or functional differences with this technique):

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.”

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      (3) Finally, a broader question about the analytical framework. The authors emphasize that the L-metric is parameter-free; however, much of their analysis still appears to rely on baked-in priors about known marker genes for cell type assignment. Is there a way to extend their analysis to infer cell types directly from the structure of the L-metric? The hierarchical clustering in e.g., Figure 3G suggests something like this: the hierarchy mostly (but not exactly) follows the marker gene annotation. Does this suggest that the cell type labeling should be revisited?

      Once of the initial motivations behind the creation of the L-metric (now L-score) was to have a more reliable and specific way of quantifying the interaction between genes that are known in the literature to be cell type markers — with more sophisticated (and noisy) methods of analysis like single-cell RNA sequencing, we found that there was substantial variation in the specificity and ubiquity of so-called ‘marker genes’, and that their usefulness often depended on context. We appreciate that the reviewer raises this point as well, and we think that scL-score analysis, such as that exemplified in Figure S4.4, can help identify new marker genes, or at the very least distinguish the biological context in which a marker gene is useful.

    1. eLife Assessment

      This multimodal neuroimaging study leverages fMRI, PET, and deep learning to predict memory performance. The authors introduce the brain-cognition gap to link these different imaging modalities to cognition and evaluate their results in two independent cohorts. The results are solid and provide an important contribution to the literature and will be of interest to neuroscientists working at the interface of cognition, neuroimaging.

    2. Reviewer #1 (Public review):

      [Editor's Note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have responded to the comments raised in the previous round of review.]

      Summary:

      The authors attempted to identify if a new deep learning model could be applied to both resting and task state fMRI data to predict cognition and dopaminergic signaling. They found that resting state and moving watching conditions best predict episodic memory, but only movie watching predicts both episodic and working memory. A negative 'brain gap' (where the model trained on brain connectivity predicts worse performance than what is actually observed) was associated with less physical activity, poorer cardiovascular function, and lower D1R availability.

      Strengths:

      The paper should be of broad interest to the journal's readership, with implications for cognitive neuroscience, psychiatry, and psychology fields. The paper is very well-written and clear. The authors use two independent datasets to validate their findings, including two of the largest databases of dopamine receptor availability to link brain functional connectivity/activity with neurochemical signaling.

      Comments on previous version:

      I thank the authors for their extensive efforts to revise the manuscript. I have no further concerns.

    3. Reviewer #2 (Public review):

      Summary:

      The authors developed a deep learning model based on a DenseNet CNN architecture to predict two cognitive functions: working memory and episodic memory, from functional connectivity matrices. These matrices were recorded under three conditions: during rest, a working memory task, and a movie, and were treated as images for the CNN algorithm. They tested their model's performance across different conditions and a separate dataset with a different age distribution (using the same MRI scanner, scanning configurations, and cognitive tests). They also calculated the "brain cognition gap" based on the model trained on resting functional connectivity to predict working memory. Extending from the commonly used index "brain age," the brain cognition gap was defined as the difference between the working memory score predicted by their model (predicted working memory) and the working memory score based on the working memory test itself (observed working memory). This brain cognition gap was found to be associated with physical activity, education, and cardiovascular risk. The authors also conducted additional mediation tests to examine whether regional functional variability mediated the relationship between PET-derived measures of dopamine and the brain cognition gap.

      Strengths:

      The major strength of this manuscript is the extensive effort the authors have put into creating a new 'biomarker' that links deep learning with fMRI, PET, physical activity, education, and cardiovascular risk across two studies. This effort is impressive.

      Concerns from the previous round of review:

      (1) The primary issue is still the lack of baseline models against which to benchmark the predictive performance of the proposed DenseNet model. This concern was raised independently by two reviewers. Without such benchmarks, it is difficult to interpret the reported results in the context of prior work on MRI-based cognition prediction.

      Notably, the authors state: "While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods, nor was this the aim of the present study."

      However, I could NOT find any discussion or results related to the CPM model in the manuscript. It is therefore unclear whether the DenseNet model was actually statistically compared with CPM, and, if so, how the comparison was conducted.

      Note that the statement, "While Vieira et al. show that the majority (76%) of prior studies used linear modeling approaches, including CPM and penalized regressions, these models are often vulnerable to overfitting, especially when applied to high-dimensional fMRI data," is not entirely accurate. Linear models typically have far fewer parameters than deep-learning models and are therefore often less prone to overfitting. In fact, it is well established that deep-learning models are particularly susceptible to overfitting and usually require substantially larger sample sizes to achieve stable and reliable performance. Although deep-learning models may outperform shallower models once sufficient data are available and training is well controlled, this does not justify the authors' claim as stated. I therefore disagree with the argument put forward by the authors.

      The authors further justify the absence of benchmarking by stating: "In this context, deep learning was employed as a flexible framework capable of modelling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach." However, most shallow models can likewise be applied across different brain states and cognitive targets. This rationale does not establish deep learning as a uniquely appropriate or necessary choice. If deep learning is indeed a better approach in this context, the authors should demonstrate this empirically through appropriate benchmarking against established baseline models.

      (2) Additional analysis shows that "BCG is not significantly associated with cognition itself". This is the most perplexing result. This is like saying Brain Age Gap is not related to chronological Age. It is counterintuitive since the Brain Age Gap is calculated by chronological age minus actual age, and most research has shown a strong relationship between the Brain Age Gap and age.

      If the brain cognition gap is not related to cognition, is it possible that the results found are mainly due to the predictive model not fitting well with another dataset? Regardless, the lack of association between BCG and cognition deserves a discussion.

      (3) I still do not fully understand the rationale of the mediation analysis. The analysis and findings are still not related to aims 1 and 2, since DA and entropy are not part of the prediction models. But I appreciate the explanation that this part is related to the authors' previous work, and that the authors attempted to link to them somehow.

      [Editors' note: the authors have responded to these points.]

    4. Reviewer #3 (Public review):

      Summary:

      This paper by Esmaeili and co-authors presents a connectome prediction study to predict episodic memory and relate prediction errors to other phonotypic variables.

      Strengths:

      (1) A primary and external validation dataset.

      (2) Novel use of prediction errors (i.e., brain-cognitive gap).

      (3) A wide range of data was investigated.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version:

      I thank the authors for their extensive efforts to revise the manuscript. I have no further concerns.

      Reviewer #2 (Public review):

      The authors have made several corrections to the original manuscript. For example, they revised the bootstrapping analysis to avoid arbitrarily inflating the degrees of freedom. However, most substantive concerns remain inadequately addressed.

      (1) The primary issue is still the lack of baseline models against which to benchmark the predictive performance of the proposed DenseNet model. This concern was raised independently by two reviewers. Without such benchmarks, it is difficult to interpret the reported results in the context of prior work on MRI-based cognition prediction.

      Notably, the authors state: "While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods, nor was this the aim of the present study."

      However, I could NOT find any discussion or results related to the CPM model in the manuscript. It is therefore unclear whether the DenseNet model was actually statistically compared with CPM, and, if so, how the comparison was conducted.

      Note that the statement, "While Vieira et al. show that the majority (76%) of prior studies used linear modeling approaches, including CPM and penalized regressions, these models are often vulnerable to overfitting, especially when applied to high-dimensional fMRI data," is not entirely accurate. Linear models typically have far fewer parameters than deep-learning models and are therefore often less prone to overfitting. In fact, it is well established that deep-learning models are particularly susceptible to overfitting and usually require substantially larger sample sizes to achieve stable and reliable performance. Although deep-learning models may outperform shallower models once sufficient data are available and training is well controlled, this does not justify the authors' claim as stated. I therefore disagree with the argument put forward by the authors.

      The authors further justify the absence of benchmarking by stating: "In this context, deep learning was employed as a flexible framework capable of modelling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach." However, most shallow models can likewise be applied across different brain states and cognitive targets. This rationale does not establish deep learning as a uniquely appropriate or necessary choice. If deep learning is indeed a better approach in this context, the authors should demonstrate this empirically through appropriate benchmarking against established baseline models.

      We thank the reviewer for reiterating this point. As noted in both the manuscript and our previous response, the primary goal of the present study was not to benchmark predictive algorithms, but rather to compare the predictive utility of different brain states using a consistent modeling framework.

      We did perform an exploratory CPM analysis using a conventional implementation that included correlation-based feature selection (p < 0.01), summarization of positive and negative networks, and robust regression for prediction using 3-fold validation. However, CPM performance can depend substantially on analytical choices, including feature-selection thresholds, treatment of positive and negative networks, cross-validation strategies, and model specification. Although we obtained CPM results (see below), we did not systematically evaluate how these choices influenced performance, nor did we optimize CPM to the same extent as would be required for a rigorous methodological comparison.

      For this reason, we chose not to include a CPM-versus-DenseNet comparison in the manuscript. Any direct comparison could easily be overinterpreted as evidence for the superiority of one approach over another, despite the absence of a comprehensive benchmarking framework. We therefore deliberately avoided such claims and instead focused on the scientific question motivating the study: whether predictive performance differs across brain states when the same predictive framework is applied consistently. We agree that comparisons with CPM and other predictive approaches would be valuable, but we believe such analyses are better suited to a dedicated methodological benchmarking study.

      (2) Additional analysis shows that "BCG is not significantly associated with cognition itself". This is the most perplexing result. This is like saying Brain Age Gap is not related to chronological Age. It is counterintuitive since the Brain Age Gap is calculated by chronological age minus actual age, and most research has shown a strong relationship between the Brain Age Gap and age.

      If the brain cognition gap is not related to cognition, is it possible that the results found are mainly due to the predictive model not fitting well with another dataset? Regardless, the lack of association between BCG and cognition deserves a discussion.

      We thank the reviewer for this comment. The absence of a significant BCG-cognition association might be unexpected. We agree it warrants careful interpretation. Theoretically, when a predictive model is trained and evaluated across samples with differing age distributions, the regression-to-the-mean dynamics that typically create BAG–age dependence may not transfer in the same way to BCG–cognition relationships, particularly in age-homogeneous cohorts such as COBRA. However, we acknowledge that the present findings alone are insufficient to resolve this question fully.

      We have added additional findings as supplementary material (Figure S4).

      (3) I still do not fully understand the rationale of the mediation analysis. The analysis and findings are still not related to aims 1 and 2, since DA and entropy are not part of the prediction models. But I appreciate the explanation that this part is related to the authors' previous work, and that the authors attempted to link to them somehow.

      We appreciate the reviewer's comment. The mediation analysis was not intended as a replication of our previous work, but rather as a mechanistic analysis to examine whether the association between BCG and cognition operates indirectly through brain age. Given the established links between BCG, brain age, and cognitive function, mediation analysis provides a principled framework for testing this hypothesis. Essentially, this mediation analysis supports previous theoretical model and our own empirical data where we showed lower DA contributes to more noise in FC metric, which in turn results in less accurate prediction (i.e., larger gap).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I hope the authors report CPM analysis as they claimed in the response letter, with actual statistical tests to compare the performance of DenseNet vs. CPM. It will be better to compare DenseNet with other models as well.

      We thank the reviewer for reiterating this point. As noted in both the manuscript and our previous response, the goal of the present study was not to benchmark DenseNet against alternative predictive frameworks, but rather to compare the predictive utility of different brain states using a consistent modeling approach. Although we explored CPM as a potential reference model, we found that its performance was sensitive to analytical choices, including feature-selection thresholds, summarization of positive and negative networks, and model specification. A rigorous CPM comparison would therefore require systematic optimization and validation of these choices, which would constitute a separate benchmarking analysis beyond the scope of the present study. We therefore deliberately avoided claims regarding methodological superiority and revised the manuscript accordingly. While comparisons with CPM and other machine-learning approaches would be valuable, we believe such analyses would constitute a separate benchmarking study beyond the scope of the present work.

      (2) "n-back did not significantly predict EM in DyNAMiC, and rest did not significantly predict WM. For this reason, we highlighted only the conditions that showed meaningful predictive power in the original analyses."

      These null results should also be reported; otherwise, this suggests the authors cherry-picked only the "meaningful" results, making the claims overly optimistic.

      We thank the reviewer for this comment. We agree that reporting only significant findings could create the impression of selective reporting. However, this was not our intention. All within-dataset prediction results, including both significant and non-significant findings, are reported in Tables 1 and 2. For the cross-dataset analyses, we chose to evaluate only the best-performing models identified in the DyNAMiC dataset. This decision was made a priori to test whether the most reliable predictive models generalized to an independent dataset, rather than to maximize the number of significant findings. For example, resting-state FC emerged as the strongest predictor of episodic memory in DyNAMiC and was therefore selected for external validation in COBRA. Similarly, the movie-watching model showed the strongest performance for working memory and was consequently carried forward to the cross-dataset validation. We have clarified this rationale in the manuscript.

      (3) I appreciate the correlation plots between BCG and physical activity and cardiovascular risk. The results are much weaker in COBRA (r = .17 and -.10 vs. .40 and -.27 in DyNAMIC). Perhaps this warrants discussion. Note that there are potential outliers in DyNAMIC. Perhaps the authors might like to include Spearman's rank.

      We thank the reviewer for this helpful observation. We agree that although the associations between BCG and both physical activity and cardiovascular risk were statistically significant in COBRA, the effect sizes were smaller than in DyNAMiC. To assess whether the observed association was influenced by potential outliers, we repeated the analysis using Spearman’s rank correlation. The association remained the same after controlling for age using partial Spearman rank correlation - between GAP and physical activity (DyNAMiC: r =0.40, p =0.001; COBRA: r =0.17, p=0.03) and for GAP vs. CVD risk score (DyNAMiC: r =–0.27, p =0.03; COBRA: r = –0.10, p =0.40). Nevertheless, we have now tempered the interpretation in the Discussion (P.12) to clarify that the direction and significance of the associations were consistent across datasets, but that the magnitude of the effects was weaker in COBRA. This difference may reflect cohort differences, including age range, sample composition, and differences in how physical activity was assessed.

      (4) Yes, adding figures comparing BCG and BAG in the main text would be helpful, given BAG's popularity.

      We have added additional findings as a supplementary figure.

      (5) The authors should provide this reason as a justification in the method: "We initially attempted to predict both episodic memory (EM) and working memory (WM). However, EM prediction was only reliable within and across samples for the resting state, whereas WM prediction generalized most strongly from the movie-watching condition. Because COBRA does not include a movie-watching paradigm, we could not evaluate WM prediction across datasets. For this reason, we focused on EM when examining the brain-cognition."

      We thank the reviewer’s suggestion. This clarification is included in the method section of the revised manuscript [P 21].

    1. eLife Assessment

      This study explores how complement protein C3 and its signalling may modulate immune training in alveolar macrophages. The findings are an important contribution to the field of trained immunity. The findings are convincingly supported by in vivo and ex vivo experiments, encompassing both pharmacological and genetic-based approaches.

    2. Reviewer #1 (Public review):

      [Editors' note: The revised manuscript addressed the concerns of both reviewers, who have concluded that the manuscript is convincing and important. The manuscript can move towards the Version of Record.]

      Summary:

      This study is built on the emerging knowledge of trained immunity, where innate immune cells exhibit enhanced inflammatory responses upon challenged by a prior insult. Trained immunity is now a very fast-evolving field and has been explored in diverse disease conditions and immune cell types. Earhart and the team approached the topic from a novel angle and was the first to explore a potential link to the complement system.

      The study focused on the central complement protein C3 and investigated how its signalling may modulate immune training in alveolar macrophages. The authors first performed in vivo experiments in C57BL mouse models to observe the presence of enhanced inflammation and C3a in BAL fluid following immune training. These changes were then compared with those from C3-deficient mice, which confirmed the involvement of C3a. This trained immunity was further validated in ex vivo experiments using primary alveolar macrophage, which was blunted in C3-deficiency, and, intriguingly, rescued by adding exogenous C3 protein, but not C3a. The genetic-based findings were supported by pharmacological experiments using the C3aR antagonist SB290157. Mechanistically, transcriptomic analyses suggested the involvement of metabolism-linked, particularly glycolytic, genes, which was in agreement with an upregulation of glycolytic flux in WT but not C3-deficient macrophages.

      Collectively, these data suggest that C3, possible through engaging with C3aR, contributes to trained immunity in alveolar macrophages.

      Strengths:

      The conclusions reached were well supported by in vivo and ex vivo experiments, encompassing both genetic-knockout animal models and pharmacological tools.

      The transcriptomic and cell metabolism studies provided valuable mechanistic insights.

      Weaknesses:

      For the in vivo experiments, the histopathological and other inflammatory markers (Fig 1.) were not directly linked to alveolar macrophages by experimental evidence. Other innate immune cells (e.g. dendritic cells, neutrophils) and endothelial cells could also be involved in immune training and contribute to the pathological outcomes. These cells were not examined or mentioned in the study.

      For the ex vivo experiments assessing immune training in alveolar macrophages, only the release of selected inflammatory factors were measured. Macrophage activities constitute multiple aspects (e.g. phagocytosis, ROS production, microbe killing), which should also be considered to better depict the effect of trained immunity.

      The proposed mechanism of C3 getting cleaved intracellularly then binding to lysosomal C3aR need to be further supported by experimental evidence.

      There was an absence of any validation in human-based models.

      Comments on the revised version.

      The revised manuscript now encompasses a much wider scope and stronger evidence.

      The authors have included the re-analysis of a recently published dataset of human volunteers who received aerosolized BCG exposure compared to saline. Although not proven causality, this data helped strengthen the human relevance of the findings presented in this research and directly rationalized the decision to focus on Ams. The persistence of elevated C3/C3aR1 expression to day 7 further supports the idea that complement‑associated reprogramming is not merely an acute inflammatory phenomenon. Whilst it may be outside of the scope of this current study, it would be helpful to clarify in future studies whether other complement components (C5, factor B, factor D) were also modulated in the dataset, to contextualize whether the response is uniquely centered on C3/C3aR1 or part of a broader complement activation program.

      The authors have also expanded the functional characterization of trained alveolar macrophages by including phagocytosis and ROS generation measurements. It is intriguing that HKPA training did not markedly alter the phagocytosis and ROS production by alveolar macrophages relative to the control group, however, C3 deficiency significantly dampened these responses in both trained and untrained groups. This reduction is in congruence with the cytokine release data, but there could be other factors involved.

      I appreciate the careful revision and much more expansive mechanistic interpretation regarding intracellular C3aR, and that further studies are underway to better understand the cell type-specific, subcellular localization of C3a-C3aR in alveolar macrophages.

      Overall, the revised data interpretation and discussion significantly improved in balance and contextualization of the findings.

    3. Reviewer #2 (Public review):

      Earhart et al. investigated the role of the complement system in trained innate immunity (TII) in alveolar macrophages (AM). They used a WT and C3 knockout murine model primed with locally administered heat-killed P. aeruginosa (HKPA). Additionally, they employed ex vivo AM training models using C3 knockout mice, where reconstitution of C3 and blockade of C3R were performed. The study concluded that the C3-C3R axis is essential for inducing TII in macrophages in the ex vivo model. The manuscript is well-written and easy to follow.

      Comments on revised version.

      My concerns have been addressed, and the provided data is convincing supporting the manuscript's claims.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study is built on the emerging knowledge of trained immunity, where innate immune cells exhibit enhanced inflammatory responses upon being challenged by a prior insult. Trained immunity is now a very fast-evolving field and has been explored in diverse disease conditions and immune cell types. Earhart and the team approached the topic from a novel angle and were the first to explore a potential link to the complement system.

      The study focused on the central complement protein C3 and investigated how its signalling may modulate immune training in alveolar macrophages. The authors first performed in vivo experiments in C57BL mouse models to observe the presence of enhanced inflammation and C3a in BAL fluid following immune training. These changes were then compared with those from C3-deficient mice, which confirmed the involvement of C3a. This trained immunity was further validated in ex vivo experiments using primary alveolar macrophage, which was blunted in C3-deficiency, and, intriguingly, rescued by adding exogenous C3 protein, but not C3a. The genetic-based findings were supported by pharmacological experiments using the C3aR antagonist SB290157. Mechanistically, transcriptomic analyses suggested the involvement of metabolism-linked, particularly glycolytic, genes, which was in agreement with an upregulation of glycolytic flux in WT but not C3-deficient macrophages.

      Collectively, these data suggest that C3, possibly through engaging with C3aR, contributes to trained immunity in alveolar macrophages.

      Strengths:

      The conclusions reached were well supported by in vivo and ex vivo experiments, encompassing both genetic-knockout animal models and pharmacological tools.

      The transcriptomic and cell metabolism studies provided valuable mechanistic insights.

      We thank the reviewers for acknowledging the importance of the work.

      Weaknesses:

      For the in vivo experiments, the histopathological and other inflammatory markers (Figure 1) were not directly linked to alveolar macrophages by experimental evidence. Other innate immune cells (eg. dendritic cells, neutrophils) and endothelial cells could also be involved in immune training and contribute to the pathological outcomes. These cells were not examined or mentioned in the study.

      We agree with the suggestions from Reviewer 1 that other cell types such as dendritic cells, neutrophils, and endothelial cells can also be involved in immune training. As the focus of this study was on alveolar macrophages, we specifically focused on these cell types. However,

      (1) We have re-analyzed a recently published dataset of human volunteers who received aerosolized BCG exposure compared to saline. We observe that by Day 7, aerosolized BCG exposure alters the expression of C3 and C3aR1 in alveolar macrophages in the human bronchoalveolar lavage (BAL) fluid, compared to saline. We have included this new analysis in a revised Figure 1 to clarify why our focus is on investigating the C3-C3aR1 axis in alveolar macrophages.

      (2) We have conducted new experiments where we train the mice in vivo, collect the alveolar macrophages, and then provide the second stimulus ex vivo. We observe a similar phenotype in the in vivo trained, ex vivo stimulated alveolar macrophages. We have included this new data in a new Figure S2.

      (3) We have updated our Discussion to state “However, we also acknowledge that other immune cells such as dendritic cells and neutrophils, and non-immune cells such as epithelial cells, endothelial cells and fibroblasts can also be involved in immune training (Bigot et al., 2025; Friščić et al., 2021; Moorlag et al., 2020).”

      (2) For the ex vivo experiments assessing immune training in alveolar macrophages, only the release of selected inflammatory factors were measured. Macrophage activities constitute multiple aspects (e.g. phagocytosis, ROS production, microbe killing), which should also be considered to better depict the effect of trained immunity.

      We agree with the reviewer and have conducted additional experiments to assess immune responses influenced by training in alveolar macrophages. Specifically, we show that in addition to impairing the release of proinflammatory cytokines such as TNFα and IL-6, C3-deficient alveolar macrophages exhibit significantly lower phagocytosis and ROS production compared to WT alveolar macrophages post-training with heat-killed Pseudomonas aeruginosa. Results from these additional experiments have been included in new Figure S2.

      (3) The proposed mechanism of C3 getting cleaved intracellularly and then binding to lysosomal C3aR needs to be further supported by experimental evidence. 

      The mechanism of C3 being cleaved intracellularly involves serine protease-dependent cleavage of C3 to C3a and has been experimentally demonstrated previously (Liszewski et al. Immunity 2013; Elvington et al. J Clin Invest 2017). A prior report demonstrated that intracellular C3a interacted with a lysosomal C3aR to promote CD4<sup>+</sup> T cell survival (Liszewski et al. Immunity 2013). Based on the reviewer’s suggestions, we performed confocal microscopy on alveolar macrophages. Although we clearly observed intracellular colocalization of C3a (using a monoclonal antibody to the neo-epitope) with C3aR, we observed only some colocalization with LAMP1, a lysosomal marker (see Author response image 1). Hence, we will refrain from making comments on how C3 binds to lysosomal C3aR intracellularly in alveolar macrophages, as this may be cell type-specific or stimulation-specific. We have now revised the sentence in the manuscript to remove any references to lysosomal C3aR and now state – “Upon internalization, C3 is cleaved to C3a (Elvington et al., 2017), binds to C3aR, and affects cytokine production in CD4<sup>+</sup> T cells (Liszewski et al., 2013)”. We have not incorporated the Author response image 1 in the main manuscript as we would like to explore this further to precisely define the subcellular localization of C3a-C3aR in alveolar macrophages, but have provided it for the reviewer to explain the basis of the rewording in the revision.

      Author response image 1.

      C3a-C3aR colocalization in mouse ex vivo cultured alveolar macrophages (mexAM). mexAMs were harvested and cultured as per the protocol from Gorki et al. (2022). Cells were incubated in a Millicell EZ Slide 8-well glass chamber slide overnight to allow for adherence, then fixed, permeabilized, and incubated with anti-C3a conjugated to AF555 (blue, Hycult HM1072), anti-C3aR conjugated to AF647 (red, Hycult HM1123), and anti-LAMP1 (green, Cell Signaling 99437) overnight at 4°C. Slides were washed 3X in PBS (5 min each) and mounted overnight at 4°C in ProLong Diamond Antifade Mountant with DAPI (white). Images were acquired on a Zeiss LSM 880 confocal microscope at 63X. At least 6 cells per condition imaged. Experiments were conducted in duplicate (technical replicates) and repeated (for biological replicates). Scale bar, 2 μm.

      (4) There was an absence of any validation in human-based models.

      We acknowledge that the observations need to be validated in human-based models. The focus of our manuscript is on training in alveolar macrophages. Unfortunately, we do not have access to an adequate representation of human alveolar macrophages for our ex vivo testing to account for individual-level variation in immune responses. We anticipate this work will form the basis of these future studies. In the interim, we re-analyzed a recently published publicly available dataset of human BAL specimens from human volunteers who underwent aerosolized BCG administration (Marshall et al. Nat Comm 2025). We observe an increase in C3 and C3aR1 expression at Day 2, which persists through Day 7 post-training with aerosolized BCG compared to aerosolized saline specifically in human alveolar macrophages. We have included this data in Revised Figure 1. We also validated C3 uptake in alveolar macrophages using precision-cut lung slices from human donors. We have included this additional data in new Supplementary Figure 3.

      Reviewer #2 (Public review):

      Earhart et al. investigated the role of the complement system in trained innate immunity (TII) in alveolar macrophages (AM). They used a WT and C3 knockout murine model primed with locally administered heat-killed P. aeruginosa (HKPA). Additionally, they employed ex vivo AM training models using C3 knockout mice, where reconstitution of C3 and blockade of C3R were performed. The study concluded that the C3-C3R axis is essential for inducing TII in macrophages in the ex vivo model. The manuscript is well-written and easy to follow. However, I have the following major concerns.

      (1) The secondary challenge to assess the reprogramming of innate cells in the BAL was conducted 14 days after the initial exposure to HKPA. However, no evidence is provided to confirm that homeostasis was re-established following the primary exposure. Demonstrating the resolution of acute inflammation is essential to ensure that the observed responses to the secondary challenge are not confounded by persistent inflammation from the initial exposure.

      We thank the reviewer for giving us an opportunity to clarify this point. We have now included additional data from the bronchoalveolar lavage fluid of these mice to show that the levels of protein leaked into the BAL, levels of proinflammatory cytokines (e.g., TNFα, CXCL1) and the neutrophils (all relevant to the acute phase of inflammation) were similar between the untreated and treated wildtype mice. This new data has been included in Figure S1.

      (2) In Figure 1D, cytokine production by BAL cells from WT and C3KO mice after HKPA exposure and LPS challenge is shown. However, it is unclear whether the reduced response in trained C3KO mice is due to a defect in trained immunity or an intrinsic inability of C3KO cells to respond to LPS. To clarify this, the response of trained C3KO cells should also be compared to untrained C3KO controls after the LPS challenge. This comparison is necessary to determine if the reduction is specifically related to innate immune memory or a broader impairment in LPS responsiveness. Such control should be included in all ex vivo training and LPS stimulation experiments as well.

      We thank the reviewers for their suggestions. We have conducted additional experiments and we observe no significant differences in the BAL cytokine levels between the wildtype and C3-deficient mice post-training in the absence of infection. This new data has been included in Supplementary Figure S1.

      Additionally, we came across several manuscripts, including a recent one in eLife as a part of this Series (Gu et al. Elife 2021; Zahalka et al. Mucosal Immunol 2022; Prevel et al Elife 2025) that have done in vivo training followed by an ex vivo challenge. Hence, we have conducted new experiments to compare the response of in vivo HKPA-trained wildtype (WT) and C3-deficient (C3KO) alveolar macrophages compared to untrained AMs after an ex vivo LPS challenge. This new data has been included in Figure S2.

      (3) The data presented provide evidence of alterations in the functional and metabolic activities of innate cells in the lung, indicating the induction of innate immune memory in a C3-C3R axis-dependent pathway. However, it remains to be established whether such changes can lead to altered disease outcomes. Therefore, the impact of these changes should be demonstrated, for instance, through an infection model to support the claim made in the study that C3 modulates trained immunity in AMs through C3aR signalling.

      We acknowledge this is a Limitation of our manuscript. As this is a Short Report, we focused on how C3, via the C3aR, affects the reprogramming of alveolar macrophages. Recent work demonstrated that systemically administered β-glucan induces peripheral trained immunity and aggravates lung injury (Prével et al., 2025), similar to disease in models of periodontitis and arthritis (Haacke et al., 2025). However, training with β-glucan also reduces bleomycin-induced lung fibrosis (Kang et al., 2024). Hence, our ongoing work involves optimizing relevant intrapulmonary exposures to assess how trained immune responses are modulated by the C3a-C3aR axis. We have included the Reviewer’s critique in our revised Discussion as a limitation, while referencing the abovementioned manuscripts.

      (4) Figure 3, panels B and C - stats should be shown for comparing WT-HKPA-trained and C3KO HKPA-trained.

      These suggestions have been incorporated into Revised Fig 3B and 3C (now Figure 4).

      (5) In Figure 4, where the proper untrained C3KO is included, the data presented in Figure 4C show an increase in basal and maximum glycolysis in trained C3KO compared to their untrained control counterparts. Statistical analysis should be provided for this comparison. Based on these data, it appears that metabolic reprogramming occurs even in the absence of C3. Furthermore, C3KO cells intrinsically exhibit reduced glycolytic capacity compared to WT. These observations challenge the conclusions made in the manuscript. Therefore, without the proper control (untrained C3KO) included in all experimental approaches, it is impossible to draw an evidence-based conclusion that the C3-C3R axis plays a role in the induction of innate immune memory.

      We have included the statistical comparisons for all the groups in Figure 4C (now Figure 5), as suggested by the reviewer. The data suggests that C3-deficient (C3KO) alveolar macrophages have a blunted metabolic response to training, as compared to C3-sufficient (WT) alveolar macrophages. However, the C3KO cells do not have reduced glycolytic capacity compared to WT in the absence of training. The blunted response in trained C3KO AMs is rescued by exogenous C3, but is then reversed by C3aR antagonism. We have also provided new data/analyses with proper controls (untrained C3KO) in the other Figures (for example, in Figures S1, S2 and 3B&C (now Figure 4)). Taken together, the data would suggest that the effects of C3 in AM reprogramming are C3aR-dependent.

      (6) The Results and Discussion sections should be separated, and the results should be thoroughly analyzed in the context of published literature. Separating these sections will allow for a clearer presentation of findings and ensure that the discussion provides a comprehensive interpretation of the data.

      We thank the reviewer for this suggestion. The manuscript has been submitted as a Brief Report, and hence, we adhered to the instructions to authors for this format. However, we have added an additional section towards the end of the manuscript based on the Reviewer’s suggestion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It is intriguing that whilst there was a significant elevation of C3a in the cell culture medium of alveolar macrophages, the addition of exogenous C3a failed to rescue the phenotypes of C3-deficient alveolar macrophages. Please help discuss this.

      Alveolar macrophages secrete both C3 and proteases, which can cleave C3 to C3a in the supernatant. However, we used the addition of exogenous C3a to compare it to the addition of full-length C3. C3 is internalized by multiple cell types, including alveolar macrophages (as demonstrated in new Figure 3 of our Revised Manuscript) as C3(H<sub>2</sub>O) in comparison to C3a. Hence, we propose that the internalization of C3(H<sub>2</sub>O) provides an intracellular source of C3a (previously reported in Elvington et al. J Clin Invest 2017), which engages with the C3aR to result in alveolar macrophage reprogramming. In comparison, incubating cells with C3a does not exert similar effects. We have included these comments in a separate section towards the end of the manuscript.

      Please provide details for the statement "cell-permeable C3aR antagonist (SB290157)" (Figure 3E). Could a paracrine-based mechanism also be at play?

      SB290157 does not act selectively on the cell surface, but rather, can also enter cells. Our data, along with previously published reports (Quell et al. J Immunol 2017; Zha et al. Cancer Immunol Res 2019), suggest that C3aR may be intracellular in AMs. However, SB290157 can also block any receptor that may be present on the surface. For this reason, we used exogenous C3a as a way to interrogate surface C3aR signaling, and did not observe significant changes in AM reprogramming with exogenous C3a. However, as this is an indirect approach, we cannot completely rule out a paracrine-based mechanism and have included this limitation in the Discussion section of the revised manuscript.

      For Figure 3, please also provide the statistical analysis results for WT versus C3KO HKCA-trained cells. The statistical tests described in the legend for Figure 3D seem to apply to Figure 3E. Please check the labels.

      These suggestions have been incorporated into Revised Figure 3 (now Figure 4).

    1. eLife Assessment

      This valuable study presents an analysis of the gene regulatory networks that contribute to tumour heterogeneity and tumor plasticity in Ewing sarcoma, with key implications for other fusion-driven sarcomas. The authors employed compelling orthogonal approaches, including single-cell sequencing and xenografts, to reveal the existence and plasticity of specific gene regulatory networks (e.g., TGF-beta signaling) within Ewing sarcoma, as well as significant differences that exist between cell lines and patient tumors.

    2. Reviewer #1 (Public review):

      The investigators elegantly utilized single-cell co-assay of RNA and ATAC seq to unveil the heterogeneous gene regulatory networks in Ewing sarcoma. The authors should be commended on their ability to identify multiple unique modules of gene regulation of Ewing sarcoma utilizing complex computational methods between numerous Ewing sarcoma cell lines. Additionally, they complimented their single cell findings with xenografts as well as primary Ewing sarcoma patient tumors - validating the intratumoral heterogeneous gene regulatory networks of Ewing sarcoma. More importantly, they have revealed that exogenous TGF-B may modify these distinct epigenetic and transcriptional signatures within Ewing sarcoma tumors. Overall, the manuscript highlights an important discovery of the heterogenous gene regulatory programming of Ewing sarcoma and further highlights the role that TGFB plays within the tumor microenvironment of Ewing sarcoma. There are some areas of ambiguity that require clarification to increase the impact of the manuscript.

      Comments on the latest revision:

      The responses to my review were appropriate and my comments were all addressed.

    3. Reviewer #2 (Public review):

      Summary:

      This work by Waltner, et. al. provides a comprehensive single cell multiomics analysis of plasticity in gene regulatory networks present in Ewing sarcoma using single cell RNA-sequencing (scRNA-seq) and single cell assay for transposase accessible chromatin with sequencing (scATAC-seq). They find that Ewing sarcoma cell lines models have distinct patterns of chromatin accessibility compared to non-Ewing sarcoma models, and that there is significant variability across Ewing sarcoma cell lines, and sometimes within a single cell line. These differences across models are linked to 3 distinct gene regulatory modules, 2 of which are present across the range of model systems studied here. The first modules present across models is activated when the fusion is expressed and includes genes enriched for the known EWSR1::FLI1 response element, GGAA microsatellites along with other neural crest transcription factors. The other module primarily consists of genes repressed by EWSR1::FLI1, which are activated in EWSR1::FLI1-low states. Interestingly, EWSR1::FLI1-low cells have already been tied to more migratory and metastatic phenotypes and the data here suggest these cells are more responsive to external signals from TGF-β and this may be mediated through FOSL2-mediated gene regulation. This is a technically rigorous study, with a variety of different analytical techniques used to address similar questions and this approach elevates confidence in the answers provided. This is further strengthened by the diverse set of model systems used, including patient-derived cell lines, cell line xenograft models, patient-derived xenografts, mining available single cell data from patient samples, and validation of the gene modules identified in a larger set of patient microarray samples. In whole, this study provides a valuable resource for understanding heterogeneity, plasticity, and gene expression networks in Ewing sarcoma. This may be a useful resource for future studies of metastatic disease and provide a framework for similar questions in other fusion-driven sarcomas.

      Comments on revised version.

      The authors have addressed comments from my prior review. Thank you!

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      We thank both reviewers for their thoughtful and constructive evaluations of our manuscript. We are pleased that the reviewers recognize the value of our multimodal single-cell approach and the diversity of model systems used to define gene regulatory networks underlying heterogeneity in Ewing sarcoma.

      Reviewer #1 (Public review):

      The investigators elegantly utilized a single-cell co-assay of RNA and ATAC seq to unveil the heterogeneous gene regulatory networks in Ewing sarcoma. The authors should be commended on their ability to identify multiple unique modules of gene regulation of Ewing sarcoma utilizing complex computational methods between numerous Ewing sarcoma cell lines. Additionally, they complemented their single-cell findings with xenografts as well as primary Ewing sarcoma patient tumors - validating the intratumoral heterogeneous gene regulatory networks of Ewing sarcoma. More importantly, they have revealed that exogenous TGF-β may modify these distinct epigenetic and transcriptional signatures within Ewing sarcoma tumors. Overall, the manuscript highlights an important discovery of the heterogenous gene regulatory programming of Ewing sarcoma and further highlights the role that TGFB plays within the tumor microenvironment of Ewing sarcoma. There are some areas of ambiguity that require clarification to increase the impact of the manuscript.

      We appreciate Reviewer 1's positive assessment of our work and their recognition of the importance of identifying heterogeneous gene regulatory programs in Ewing sarcoma, including the role of TGF-β in the tumour microenvironment. We have addressed the areas of ambiguity noted by the reviewer, including clarifying cluster assignments and the selection of k=3 for module identification, adding statistical comparisons to relevant figures, correcting figure cross-references, and improving figure labeling for clarity. We have also added higher-resolution images of the spatial profiling data and highlighted relevant correlations between CHLA9 and CHLA10 clusters. We believe these revisions improve the clarity and rigour of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      This work by Waltner et. al. provides a comprehensive single-cell multiomics analysis of plasticity in gene regulatory networks present in Ewing sarcoma using single-cell RNA-sequencing (scRNA-seq) and single-cell assay for transposase accessible chromatin with sequencing (scATAC-seq). They find that Ewing sarcoma cell line models have distinct patterns of chromatin accessibility compared to non-Ewing sarcoma models, and that there is significant variability across Ewing sarcoma cell lines, and sometimes within a single cell line. These differences across models are linked to 3 distinct gene regulatory modules, 2 of which are present across the range of model systems studied here. The first modules present across models are activated when the fusion is expressed and include genes enriched for the known EWSR1::FLI1 response element, GGAA microsatellites, along with other neural crest transcription factors. The other module primarily consists of genes repressed by EWSR1::FLI1, which are activated in EWSR1::FLI1-low states. Interestingly, EWSR1::FLI1-low cells have already been tied to more migratory and metastatic phenotypes, and the data here suggest these cells are more responsive to external signals from TGF-β, and this may be mediated through FOSL2-mediated gene regulation. While there are some minor additional validation studies that can be performed to strengthen a few individual analyses, this is a technically rigorous study, with a variety of different analytical techniques used to address similar questions, and this approach elevates confidence in the answers provided. This is further strengthened by the diverse set of model systems used, including patient-derived cell lines, cell line xenograft models, patient-derived xenografts, mining available single-cell data from patient samples, and validation of the gene modules identified in a larger set of patient microarray samples. In whole, this study provides a valuable resource for understanding heterogeneity, plasticity, and gene expression networks in Ewing sarcoma. This may be useful for future studies of metastatic disease and may also provide a framework for similar questions in other fusion-driven sarcomas.

      Strengths:

      There are a few core strengths in this study. First is the number and diversity of Ewing sarcoma models studied, spanning commonly used cell lines, patient-derived xenografts, and patient samples. The second is the large array of rigorous and orthogonal approaches used to uncover the identity and function of various gene modules. This includes an array of informatics techniques, as well as specific modulation of cell line models in culture. A third is confirmation that different gene expression programs are present in the same tumor using spatial transcriptomic analysis. Lastly, the authors have made all of their data and code accessible, enabling continued use of this dataset as a resource for others.

      Weaknesses:

      As highlighted by the authors, this study is somewhat limited by the small number of single-cell data from patient samples that are publicly available. Much of the analysis comes from cell lines. Additionally, they focus only on one type of signal that may modulate cell plasticity, and there are likely to be many others. Lastly, there are a few weak spots in the data. Some of this likely arises from the underlying complexity of the data, the generally sparse nature of scATAC data, and the biological heterogeneity present in the cell lines studied. The most pronounced weakness was in the analysis of transcription factors that dictate gene expression in the distinct modules, as well as the response to TGF-β. While some specific transcription factors showed module-specific expression consistent with the computational prediction in Figure 2, others did not likely due to additional factors not tested here. Likewise, the same transcription factors did not always show consistent enrichment in the gene modules that responded to TGF-β treatment when analyzed across cell lines. On the whole, these are relatively minor weaknesses and do not diminish the value of this study.

      We thank Reviewer 2 for their thorough and balanced assessment. We agree that the study's strengths lie in the breadth of model systems and orthogonal analytical approaches, and we appreciate the reviewer's acknowledgement that the identified weaknesses are relatively minor.

      In response to the reviewer's suggestions, we have made several substantive improvements. First, we have included a new Western blot panel in Figure 1 showing EWS::FLI1 protein levels across cell lines, which provides important context for interpreting the long-read fusion transcript detection data. Second, we have revised text throughout the Results section to improve precision — in particular, clarifying that our chromatin analyses assess accessibility at published EWS::FLI1 binding sites rather than binding per se, and ensuring that our stated hypotheses match the metrics presented in the corresponding figures. Third, we have improved figure color schemes and labeling to aid interpretation and corrected errors in panel labeling.

      Regarding the reviewer's observation about transcription factor enrichment patterns across cell lines (particularly RUNX3 in the TGF-β response analysis and FOSL2’s modest expression at the protein level in CHLA9), we acknowledge that not all TFs showed perfectly consistent module-specific behaviour across every cell line. As the reviewer notes, this likely reflects the underlying biological complexity and additional regulatory factors not tested here. We have made attempts to temper our language accordingly.

      We believe that the revised manuscript, with its additional experimental data, improved figures, and clarified text, addresses the concerns raised by both reviewers and strengthens the overall impact of our findings.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Specific comments:

      (1) Figure 1A: Suggest adding cell labels on the UMAP plot - difficult to tell with different colors - which cluster is which cell line.

      We have added labels to Figure 1A (UMAP plot) as suggested.

      (2) Figure 1B: Although this may have been slightly addressed later in the manuscript, since CHLA9 and 10 are from the same patient, are the authors surprised to see the differences in EWS::FLI1 motif accessibility, or are these findings further reinforcement of inherent heterogeneity? Should one expect the ChromVAR deviation z-score at least to overlap between CHLA9 and 10, since they're from the same patient, if not, can the authors explain why?

      We were indeed surprised to see differences in EWS::FLI motif accessibility inferred from our data particularly from the isogenic lines CHLA9 and 10. While it is impossible to know for sure, we suspect that some treatment related changes increased accessibility in CHLA10. However, when considering global accessibility signatures, a subset of cells from CHLA10 (C23) was most correlated with signatures from CHLA9. We have addressed this by stating “Unsurprisingly, of all cell lines, CHLA10 C23 cells exhibited the greatest correlation to profiles from CHLA9,” in the 11th paragraph of the results section.

      (3) Figure 1B: Can the authors comment on the differences in trends of fusion transcript to chromatin enrichment (e.g., A673 and TC71 inverse trends vs. other cell lines that are direct positive or negative trends).

      We are hesitant to draw too many conclusions from transcript counting vs chromatin enrichment given some outliers, but in general, we found that high-fusion transcribing cell lines A4573, SKNMC, RDES were associated with module 3, while the lower transcribing cells CHLA9/10, TC32 and PDX305 were associated with module 2. We did not choose to emphasize this correlation too strongly in the paper as there may be other factors such as additional mutations (such as BRAF<sup>V600E</sup> in A673) that may have either been present in the original tumour or after serial passaging that may play a role.

      (4) Figure S1D: Please make the pt. smaller in order to better visualize the differences and highlight the stark differences of the ChromVAR binding site score.

      We have adjusted the point size of the embeddings in Fig S1D to enhance visualization.

      (5) Figure S1D: Is it surprising to see a large proportion of PDX305 and minor proportions of CHLA10 and CHLA9 with low ChromVAR binding site score?

      Our analyses indeed show that PDX305, CHLA9, and CHLA10 utilise fusion-repressed gene programs, so lower enrichment of EWS: FLI1 microsatellites is consistent with our findings.

      (6) Figure 2A: Unclear how k = 3 was ultimately selected - was it purely a visualization of how well separated each of the cell lines is? Can the authors further clarify how k = 3 was determined to be the most optimal? Is there a UMAP that demonstrated a difference in clustering between groups 1, 2, and 3? Is there a threshold cut-off to assign groups into 1, 2, or 3?

      Thank you for pointing out this omission. In Fig 1 we show that unsupervised clustering of DA peaks across cell lines grouped EwS cell lines into 2 clusters. When using peak-2-gene linkages (Fig 2), we chose a k =3 to explore the potential genes/CREs explaining the grouping found in Fig 1 because A673 was a clear outlier. We have added the following sentence to the results section paragraph 6. “We selected k = 3 to extend the two-cluster structure observed in Fig. 1F–G, reasoning that A673 represented a clear outlier whose distinct regulatory program would be obscured at lower k.”

      (7) Figure 2G: Understanding the substantial heterogeneity of each cell and TF binding motifs - given that CHLA9 is considered group 2, is it unexpected that FOSL2 was not highly expressed compared to the other EwS cell lines within group 2: (TC32, CHAL10, PDX305).

      Thank you for astutely pointing out that CHLA9 did defy the trend for FOSL2 protein expression compared with other group 2 lines. We specifically did not comment on this finding in the manuscript as we could not account for its status as an outlier. We address this in the public response above.

      (8) Figure S3B: Did the authors perform a similar computational analysis (performed for Figure 4), looking specifically at CHLA9 (Clusters 24 and 25) to determine if there are any overlaps with CHLA10 cluster 23's pathway activity/MSigDB/GO: Biological process terms? If there are potential overlaps, can the authors potentially infer tumor clonal evolution from CHLA9 to CHLA10?

      While we didn’t do pathway analysis, Figure S3 shows strong pearson correlation of C23 from CHLA10 with both C24 and C25 from CHLA9. We have highlighted this with a red box in the figure to make this more obvious to the reader.

      (9) Figure 4: Given the heterogeneity of CHLA-10. Did the authors observe any differences in morphology within CHLA-10 between the predominant modules (modules 2 and 3), given such stark transcriptional heterogeneity?

      We did not observe any morphologic differences within CHLA10 but acknowledge this would be an interesting avenue of further investigation.

      (10) Figure 5B: Can the authors also plot out module 3 gene expression to see if the cluster enrichment is unique from module 2 gene expression with TGFB1 and vehicle?

      We thank the reviewer for this suggestion and have replaced Fig S4B with violin plots so the enrichments are clearer. Module 3 expression in cluster 3 cells from CHLA10 is among the lowest in the cell line.

      (11) Figure 5H: Can the authors provide a higher zoom/resolution of the H&E stain of the ROIs in order to see if there are indeed more stroma/fibrosis in ROI9 and ROI10, and if there are differences in tumor cell morphology within different ROIs that harbor different modules?

      We have now included higher resolution and magnification of IHC panels in Figure 5H. These images are included as Supplemental Fig 4C. It is notable that although no dramatic differences in tumour cell morphology are visualized, ROI9 and ROI10 comprise small islands of viable tumour surrounded by necrosis.

      (12) Figure 6: Within Volchenboum/Lawlor's dataset, of the 46 clinically annotated primary tumors, 10 of the COG samples contained substantial stromal elements, while all the tumors in the European cohort were >70% viable tumors. Can the authors separate out the stromal-rich (n = 10) samples and analyze the 36 tumor-enriched samples to see if the survival curve is the same as what is shown in Figure 6G/H and S4E?

      We thank the reviewer for this suggestion. We observed no differences when stratifying the patients as outlined. Indeed, in the original Volchenboum et al, manuscript it was demonstrated that no prognostic gene signature was identifiable when the stromal-rich tumours were removed from the cohort.

      (13) Figure 6: Additionally, can the authors run a similar computational analysis to determine the predominant modules within bulk sequencing of the Volchenboum/Lawlor dataset between each tumor?

      The computational approach deconvolution was developed for use with bulk RNA-seq and is based on count data. The linear mixture model at the heart of deconvolution methods requires that the measured signal is proportional to abundance across the full range, and Affymetrix microarray data violate that assumption in a gene-specific, nonlinear way that can't be fully corrected post hoc.

      (14) Page 15, line 2: Wrong GSE data cited - currently cited as GSE61357 - should be GSE63157 instead.

      This has been corrected, thank you.

      (15) Please add statistical comparisons for Figure 3E-H.

      We thank the reviewer for highlighting this omission. We used the software package ggpubr to perform Wilcoxon rank-sum tests comparing mean module scores between conditions. We have added statistical labels to the plots. All comparisons were statistically to the level indicated in the figure.

      (16) Page 11 Line 21: Figure S3 D-E is not about EMT, migration, and TGFB signaling - I believe the authors are referring to Figures 4D-E

      Thank you, we have corrected these errors

      Reviewer #2 (Recommendations for the authors):

      (1) Data

      (a) In Figure 1 and the associated text, there is an analysis of cells expressing EWSR1::FLI1 performed using a locus-specific amplification and long-range sequencing. On page 7, lines 1-5, there is some discussion about how some of these track with overall transcript levels, while others don't. Additionally, a very low fraction of cells is shown to be EWSR1::FLI1 positive. This analysis might also be strengthened by a Western blot to show protein levels and how they vary across the cell lines, which may help explain additional differences in the data. The Abcam antibody ab133485 works well for western blotting of EWSR1::FLI1. While this isn't additional single-cell data, the percentage of cells with transcript detected is not really equivalent to the total expression level. This seems particularly valuable to do, as prior publications (Pishas, et. al., Mol. Cancer. Ther., 2018) show that of the cell lines tested here, TC32 has relatively high EWSR1::FLI1 protein levels, while A673 has relatively low expression. This contrasts with the percent of positive cells here.

      We thank the reviewer for this suggestion and have included a new panel, Figure 1E containing the western results for cells harvested in log-phase growth.

      (b) Figures 5D-F are not obviously referenced in the section about Figure 5 (page 12 line 4 through pg. 13 line 20). One question to pose to the authors is how to interpret the data here for TC71 in light of the fact that they were relatively insensitive to TGF-β. This shows up obviously in the Western blot in 5G.

      Thank you, for drawing attention to this omission. We have corrected the figure reference to include Figure 5D-F. In regards to our interpretation of TC71’s lack of response to TGF-B, we direct the authors attention to our statement: “The muted response of the TC71 cell line (module-3 dominant) to the influence of TGF-b (Fig. 5A-C, Fig. S4B) suggests that pre-existing transcriptional states may condition the sensitivity to TGF-β signalling.”

      (c) For the discussion of Figure 5, the authors say on page 12, line 21, that "RUNX3 was enriched in non-responsive clusters." I'm not entirely convinced that the data support this statement as written. This is true in A673s, but there appears to be no difference between the maximally responsive and maximally non-responsive clusters in CHLA10. TC71 was not particularly responsive, but showed the opposite effect.

      We agree this was not clearly written. We have de-emphasized RUNX3 findings here. Our new conclusion in the results section paragraph 15 is: “Correlation analyses revealed a reciprocal enrichment pattern of many key TFs from Figure 2, where correlation of FOSL2 (but not RUNX3) gene expression and accessibility was generally highest in TGF-β responsive clusters.”

      Can data for A673 cells be included for Figure 5G? Like CHLA10, this cell line had the pattern of FOSL2 enrichment that is concordant with that described in the text.

      We thank the reviewer for this suggestion and have included A673 in the western blot.

      (2) Figures

      (a) In Figure 2A, it is very difficult to distinguish the colors for CHLA9 and TC32 in the left panel. Can the color scheme here be changed to make these easier to distinguish?

      We thank the reviewer for pointing this out and we have added labels to the UMAP. We hope this makes the visualization more obvious.

      (b) Similarly, the KLF4 and SP1 lines in Figure 2D are a little close and might benefit from having more distinct colors.

      We thank the reviewer for pointing this out. We have changed the KLF4 to a green hue.

      (c) The lower panel of Figure 5I has 2 samples labeled "11" and no sample labeled "12".

      Thank you for catching this error. It has been corrected.

      (3) Text

      (a) One section early in the response was a little bit confusing and could benefit from some revision to improve clarity. On page 6, lines 9-11, this reads a little bit like they looked at cell-line specific EWSR1::FLI1 binding, but that wasn't the assay that was performed. Perhaps there is a better way to describe this than simply "enrichment of EWS::FLI1 sites."

      We thank the reviewer for their efforts to improve clarity of our work.

      We changed the following sentence in the 2nd paragraph of the result section:

      "We visualized the enrichment of EWS::FLI1 sites across all cell lines and discovered distinct EwS cell-line specific usage (Fig. 1C & Fig. S1D)."

      And revised to:

      "We then assessed chromatin accessibility at these published EWS::FLI1 binding sites across all cell lines and discovered that accessibility at these loci varied in a cell-line-specific manner (Fig. 1C & Fig. S1D)."

      (b) Then at the start of the next paragraph (page 6 line 12), the authors talk about "differences in EWS::FLI1 motif and binding site enrichment" and my first thoughts were whether this was differences in which sites were bound or differences in the strength of enrichment. More precise language would be helpful.

      Here in the 3rd paragraph of the result section we changed:

      "Given the differences in EWS::FLI1 motif and binding site enrichment"

      To:

      "Given the heterogeneity in the magnitude of EWS::FLI1 motif enrichment"

      (c) Related to the comment about EWSR1::FLI1 positive cells vs. protein levels above, the hypothesis on page 6, line 13 says that there was a hypothesis that different Ewing sarcoma lines have different levels of EWSR1::FLI1 transcript. But the metric shown in Figure 2D is the percentage of cells with detectable transcript, not transcript levels. Be specific about what the data are showing here.

      We agree with the need to improve clarity. In the 3rd paragraph, we changed: "we hypothesized that EwS cell lines have different levels of EWS::FLI1 transcript."

      To:

      "We hypothesized that EwS cell lines differ in the proportion of cells expressing high levels of the EWS::FLI1 fusion transcript."

    1. eLife Assessment

      This manuscript reports an important new statistical method for calculating the significance of correlations between two time-series, which provides more accuracy than other methods when the data has few replicates. The proposed method solves a real-life problem that is frequently encountered and is broadly applicable to many realistic datasets in many experimental contexts. The technique is supported with compelling mathematical derivations as well as analysis of both computer-generated and previously published experimental data.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review: Definitions and terminology have been made more precise. Additional analysis confirms the conclusions previously stated and clarifies concerns about the computational tractability of the method.]

      Summary:

      The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore, a useful and broadly applicable statistical tool is well placed in this journal.

      Strengths:

      The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

      The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

      The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

      The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

    3. Reviewer #2 (Public review):

      Summary:

      This paper presented a hypothesis testing procedure for the independence of two time-series that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

      Strengths:

      The test offers good flexibility for different kinds of dependence (through adjusting \rho) and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used, and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore a useful and broadly applicable statistical tool is well placed in this journal.

      Strengths:

      The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

      The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

      The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

      The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

      We appreciate the reviewer’s summary of our study and its strengths.

      Weaknesses:

      A few changes could be made if the manuscript is revised. I would consider all of these points minor, but the paper could be improved if these points were addressed.

      (1) The caption of Figure 2 doesn't seem to mention panel D. Figure A-2 also does not mention C in the caption.

      We apologize for this error, and thank you for catching it! The figure legends had missing or incorrect panel labels. This error has been corrected.

      (2) Figure 2D is a little hard to follow. First, the definition of "Power" is not clear, and I couldn't find the precise definition in the text. Second, the legend for the different lines in 2D is only given in Figure A-2. Perhaps a portion of the caption for Figure 2 is missing?

      We have added a definition of power in the main text:

      “Although the permutation test, simultaneous permute-match test, and sequential permute-match test are all valid, they vary in power – the probability of detecting true dependence.”

      We have clarified the use of “power” in legend of Fig 2 and clarified that the color key for Fig 2D is in Fig 2A. The relevant excerpt of the Fig 2 legend is copied here:

      “(D) Statistical power for the permutation test and various permute-match tests as a function of the replicate number n, significance level α, and strength of dependence r<sub>X, Y</sub>. Power was estimated as the proportion of simulations in which dependence was detected, calculated from 5000 simulations at each value of r<sub>X, Y</sub> between r<sub>X, Y</sub> = 0 and 0.54 in steps of size 0.01. At r<sub>X, Y</sub> = 0, there is no dependence, so the curve at that point indicates the false positive rate rather than power. We chose the Pearson correlation coefficient as our correlation function ρ. See (A) for the color legend.”

      We have also added dotted lines connecting the legend in panel A to the curves in panel D.

      (3) The concept of circular variance for the fish data was heard to understand/visualize. The equation on line 326 did not help much. If there is a very simple picture that could be added near line 326 that helps to explain Ct and theta, that could be a big help for some readers who do not work on related systems. The analysis performed is understandable, the reader just has to accept that circular variance captions the degree of alignment of the fish.

      We have replaced references to circular concentration with “mean resultant length”, which is the standard jargon for this term in circular statistics, and we have added an illustration.

      (4) For the data discussed in Figure 3, I wasn’t 100% sure how the time windows were selected. In the caption, it says “time series to different lengths starting from the first frame”. So the 20 s time window was from t=0 to t= 20 s. Would a different result be obtained if a different 20 s window was chosen (from t = 4 min to t = 4 min 20 s just to give a specific example). I suppose by chance one of the time windows would give a pvalue less than the target 0.05, that wouldn’t be surprising. Maybe a random time window should be selected (although I am not indicating what was reported was incorrect)? A little more discussion on this aspect of the study may be helpful.

      As suggested by the reviewer, we have redone the analysis of Figure 3D with random segments. This provides a more complete picture of how the chance of detecting a significant correlation varies with segment length. The main conclusion is unchanged: Perfect match tests reliably detect dependence across a wider range of segment lengths than the naive parametric alternative.

      The relevant panel and an excerpt from the legend text are copied below.

      “(D) Permute-match tests detected a significant correlation between speed and alignment more consistently than the parametric test. For a grid of lengths between 20 and 600 seconds we sampled 500 random segments of each length, each drawn from the first 600 seconds, and determined for each segment whether the parametric test and/or the two possible permute-match tests detected a significant (p ≤ 0.05) correlation. In the edge case of the maximum 600-second length, all 500 “random” segments were identical.”

      Reviewer #2 (Public review):

      Summary:

      This paper presented a hypothesis testing procedure for the independence of two timeseries that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

      Strengths:

      The test offers good flexibility for different kinds of dependence (through adjusting \rho), and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

      We appreciate the reviewer’s summary of key aspects of our manuscript.

      Weaknesses:

      (1) The size of the test is not guaranteed to (asymptotically) equal \alpha, which may damage the power.

      We thank the reviewer for raising the issue of test size and power. We agree that a conservative test (one whose size can fall below alpha) may sacrifice power.

      Our objective is distribution-free false-positive rate (FPR) control. That is, we wish to keep the FPR at or below alpha for every distribution of X and Y, because in our regime (nonstationary time series with few independent replicates) the scientist often cannot verify distributional assumptions. Inspired by the reviewer’s comment, we now show (new Proposition 14) that the perfect match probability can be made arbitrarily close to 1/n<sup>!</sup>. As a consequence, any reported perfect match p-value below 1/n<sup>!</sup> would break the distribution-free validity of the test.

      A test that exploits distributional structure could likely access lower p-values; we have now explored how the empirical FPR of the permute-match test varies with the data-generating process (see our response to reviewer 2's recommendation 1 below).

      (2) The computational time can be an issue for a moderately large sample size when calculating the X / Y-perfect match. It will be beneficial to include discussions on the implementations of the test.

      We agree this is an important consideration. We have added the following text to the Discussion:

      “The test appears computationally tractable for relevant sample sizes: Our implementation of the permute match procedure completed a single test of dependence in the setting of Fig 2 with an average runtime of 3 seconds when n = 10 on a 2023 14-inch MacBook Pro with an M2 Pro processor and 16 GB RAM (see Source data 1). For n > 10, a standard permutation test already can report a p-value below 3 × 10<sup>−8</sup> so the perfect match test is likely unnecessary for typical applications.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A few more minor notes/comments:

      As a personal preference, I like it when figures printed in grayscale retain their meaning (when possible). Just FYI, Figure 3D in grayscale is uninterpretable. Not saying a change is needed, just pointing it out.

      We changed Fig 3D to address a comment above, and think the new version better distinguishes between the permute-match and permutation test results in greyscale.

      On line 304, is that a lower bound or an upper bound? Maybe the issue is the probability mentioned on line 305 is not clear.

      As this point is not the main focus of the investigation, we have rephrased it to make it less technical and eliminate the issue of which bound is in question.

      “It seems likely that the tests could be further modified to report an even lower p-value when an X- and Y -perfect match occur simultaneously, as in Fig 3B. However, we have not investigated further and this problem is left for future efforts.”

      I am not sure if this is a weakness, but the p-value changing depending on the choice of whether to apply the X-perfect match or Y-perfect match test first is fascinating. The authors did discuss this very issue at several points in the manuscript. It is slightly unsettling to me that there isn't an exact p-value for a given set of data. This one point gives me a new perspective on statistics.

      In full transparency, I don't believe I have to background to thoroughly review the appendix. I did read through it and did not notice any errors, but I couldn't confidently say there are not any small mathematical errors or any logical flaws in the proofs. Some sections were not easy to follow (my own shortcomings, the writing appeared sufficient for more of an expert to understand).

      We greatly appreciate the reviewer’s time and effort tackling an appendix outside their comfort zone.

      Reviewer #2 (Recommendations for the authors):

      (1) In the numerical experiment session, the authors should include the null situation, i.e., the performance of the test when X and Y are independent. This helps assess the size of the test.

      We have added a section on size to our results section, copied below:

      “The permute-match test’s false positive rate depends on the process tested. The permute-match test is conservative – meaning that its false positive rate can fall below the significance level – because both the permutation test and perfect match test are conservative. As discussed elsewhere [29], the permutation test is conservative when α is not one of its possible p-values and when ties may occur between the original correlation and shuffled correlations. Checking for a perfect match is similarly conservative. The actual probability of a false-alarm perfect match event can vary depending on the process being tested. To see this consider the permute-match test in the setting where n = 3, where α = 0.05, and where r<sub>X,Y</sub>= 0 (independent X and Y). Note that in this case, obtaining a Y -perfect match (and thus p = 1/n<sup>n</sup>) is necessary and sufficient to detect dependence since α is too low for detection by either the permutation test or the p = 2/n<sup>n</sup> leg of the permute-match test. In the linear system of Fig 2, we observed among 5000 simulations a detection rate of 0.0148, significantly below the upper bound of 1/3<sup>3</sup> (Figure 2 - Source data 1; one-tailed exact binomial test, p < 10<sup>−20</sup>). Conversely, in the nonlinear system of Fig S2, this same event (Y -perfect match under n = 3 and r<sub>X,Y</sub> = 0) occurs with a detection rate of 0.0328, not significantly 9 below the upper bound of 1/3<sup>3</sup> (Figure S2 - Source data 1; one-tailed exact binomial test, p = 0.059). Thus, depending on the underlying process studied, the actual chance of a perfect match happening under independence may be near or significantly below the theoretical upper bound.”

      (2) Some insights regarding the choice of rho should be provided. Especially, are there any examples that the classical test, such as the Pearson correlation or Granger causality test does not work?

      We have redone the example of Appendix 4 with Pearson correlation, showing that Pearson correlation has substantially lower power than cross-map skill in this case (compare figures S2 and S3).

      (3) Line 34 - 35, page 2: Correlations and causality should be separately considered. This sentence talks more about causality rather than correlation.

      We appreciate the reviewer’s perspective and agree that correlation and causality are distinct.

      We feel that pointing out the issue of spurious correlations is helpful to orient our readers, especially those from a broad scientific audience. In the text, we define “correlation” as a descriptive statistic (rather than normalized covariance), and later distinguish it from “dependence”, which has causal implications due to Reichenbach’s common cause principle. We believe this distinction provides a useful backdrop for practitioners who use statistical methods but are perhaps new to thinking deeply about statistical dependence.

      (4) Please add some discussions on the situation that X_i depends on Y_{i - j} for some j > 0, which is associated with the setting of Granger causality test.

      We have added the following to the discussion:

      “No distributional assumptions are required, and the correlation function ρ can be completely arbitrary. For instance, ρ could include a lag to detect delayed dependence, or even evaluate the correlation strength at several lags and report the strongest among them [39].”

    1. eLife Assessment

      This important study by Otgonbaatar and colleagues employs advanced live microscopy, optogenetics, and an endogenous fluorescent timer system to investigate short- and long-term stabilization dynamics of β-catenin/Armadillo (Arm) during Drosophila development. The authors identify an unexpected and functionally relevant enrichment of stabilized junctional Arm in leading-edge cells during dorsal closure, providing evidence for a stabilization mechanism that appears independent of canonical Wingless signaling. These findings are significant because they expand current understanding of β-catenin/Arm beyond its canonical signaling functions and suggest a role in tissue mechanics and force transmission during dorsal closure. The proposed model represents a key advance in the field, but the strength of evidence is currently incomplete: the main conclusions regarding Wingless independence, JNK-mediated regulation, and the mechanical role of stabilized Arm are only partially supported by the available data and would benefit from further experimental testing and corroboration.

    2. Reviewer #1 (Public review):

      In this study, Otgonbaatar and colleagues investigate the stability of Armadillo (Arm) during Drosophila development using a creative tandem fluorescent protein timer approach via endogenous tagging of Arm. The tagging strategy allows for newly synthesised and longer-term stabilised Arm pools to be distinguished from one another. Specifically, the authors address the functional relevance of and mechanism behind the stabilisation of junctional Arm during dorsal closure.

      The authors show that Arm is stabilised at the leading edge during dorsal closure. Using a sophisticated optogenetics approach, which allows for acute perturbations, they show that stabilised Arm is functionally required for dorsal closure. Increasing Wg (by overexpression) did not affect dorsal closure or Arm stability, in contrast to Axin overexpression, which reduces Wg/Arm signalling. In line with canonical signalling control of Arm levels being critical, stabilisation of Arm by N-terminal mutations disrupted dorsal closure. However, the same deletion is also expected to affect interaction with alpha-catenin. Co-localisation with E-cadherin and actin suggests a junctional role of leading-edge localised Arm. Optogenetic targeting of alpha-catenin points towards a key role of adherence junctions in dorsal closure. Allele replacement with mutant variants of Arm to affect adherence junction complex assembly further indicates an important contribution of coupling between Arm and alpha-catenin. Using overexpression approaches, the authors suggest that Dsh and Jnk contribute to dorsal closure.

      This microscopy- and optogenetics-based study is generally well-conducted and provides strong evidence for stabilised Arm during dorsal closure, as well as its functional importance. This is an important discovery relevant to morphogenesis and potentially mechanotransduction. From a technical perspective, the validated beta-catenin timer provides a valuable tool for the field. The timer has revealed that Arm stabilisation does not coincide with Wg stripes, suggesting a Wg-independent stabilisation mechanism that may instead depend on adherence junction assembly, especially the interaction of Arm with alpha-catenin. However, as N-terminal deletion within Arm and Axin overexpression also disrupted dorsal closure, substantial ambiguity remains. Can suppression of the beta-catenin degradation machinery be ruled out as a regulatory mechanism? An expansion of ArmTimer mutant variants could contribute to testing the authors' conclusion further. Structural insights into junctional interactions involving Arm (e.g., 10.1074/jbc.M114.554709) could, for example, be used for further functional exploration by mutagenesis. The direct mechanistic impact of JNK and its potential link to Dsh in dorsal closure remains less compelling.

      In summary, this is a highly relevant and important study, potentially pointing to a novel stabilisation mechanism of beta-catenin in development. Further corroboration of the mechanism, to test whether it is indeed distinct from canonical signalling, would be needed to support the conclusions.

    3. Reviewer #2 (Public review):

      Summary:

      Otgonbaatar et al. sought to investigate β-catenin/Arm protein lifetime and stabilization dynamics in vivo during embryonic development. To address this question, the authors developed an endogenous tandem fluorescent protein timer (tFP) system that enables the visualization of newly synthesized versus long-lived Arm protein in vivo. Using this approach, the authors sought to determine where stabilized Arm accumulates during development and how it contributes to dorsal closure.

      Strengths:

      A major strength of the study is the development and application of the endogenous Arm timer system, which provides a powerful approach for monitoring protein stabilization dynamics in living tissues. Using this system, the authors unexpectedly found that the strongest Arm stabilization occurs not in Wnt signaling regions, but at the leading edge cells during dorsal closure. The study combines quantitative live imaging, optogenetic perturbation, genetic analysis, and structure-function approaches to demonstrate that stabilized junctional Arm interacts with α-catenin and contributes to tissue mechanics required at the leading edge for dorsal closure. Particularly compelling is the combination of multiple perturbations, including optogenetic disruption of Arm or α-catenin, Axin overexpression, and Arm mutants, which produce consistent dorsal closure defects.

      Some conclusions are generally supported by the presented data. The work provides strong evidence that Arm plays an important role in dorsal closure. The identification of a requirement for the Dishevelled DEP domain and JNK signaling supports a non-canonical regulatory mechanism controlling dorsal closure.

      Weaknesses:

      (1) Conclusions are made regarding force transmission;(however, no experimental evidence is provided to support these conclusions.

      (2) The conclusion was made that Wingless does not affect dorsal closure. However, this was based solely on Wingless overexpression in the amnioserosa, and the level of Wingless expression was not quantified. One possibility is that this level was not sufficient to see an effect. Alternatively, Wingless may have a role in migrating epithelium rather than the amnioserosa. Indeed, it is known that wingless mutants display a defect in dorsal closure.

      (3) The effect of JNK knockdown on Arm localization maybe is indirect, and due to a secondary consequence on disruption of epithelial morphology rather than a direct effect of JNK on Arm.

      (4) Some conclusions rely on overexpression-based perturbations (e.g., Axin or Arm mutants), which may not fully recapitulate endogenous physiological regulation.

      (5) The Arm timer was not able to detect Wingless-dependent Arm stabilization in stripes. This finding demonstrates that the timer is not sensitive enough to thoroughly analyze Arm dynamics.

      Overall, this work provides important conceptual advances in understanding junctional β-catenin/Arm function during dorsal closure. The endogenous fluorescent timer approach will likely be broadly useful to the community for studying protein stability dynamics in vivo, and the findings expand current views of β-catenin by highlighting its mechanical and junctional functions during tissue morphogenesis.

    4. Author response:

      We are glad the reviewers found the tandem fluorescent timer approach valuable and the leading-edge Arm stabilization finding significant.

      We agree with the Assessment that our evidence for three specific claims Wingless-independence, JNK-mediated regulation of Arm stability, and a direct mechanical/force-transmission role for stabilized Arm is currently incomplete, and we will revise the text throughout to reflect this more precisely rather than overstating the current data. In addition, we commit to two new experiments, both using existing reagents and fly stocks, that speak directly to the two most experimentally tractable points raised by the reviewers:

      (1) Re-staining our existing JNK-RNAi and JNK-overexpression embryos for E-cadherin (reagent already validated in Figure 4), to test whether JNK acts directly on junctional architecture or only indirectly, via broader epithelial disruption.

      (2) Imaging ArmTimer in a wingless loss-of-function background, to directly test Wingless-dependence of leading-edge Arm stabilization as the reciprocal of our existing overexpression data.

      We address each public review point below and outline the accompanying text revisions.

      On Wingless independence (Reviewer #1; Reviewer #2, Weaknesses #2 and #5; Recommendation #1):

      We agree that our current evidence unquantified Wg overexpression restricted to the amnioserosa (C381-Gal4) and uniform overexpression, alongside the absence of detectable Wg-stripe-associated Arm-Timer signal supports a more limited conclusion than "Wingless-independent" as currently stated. We will revise our language throughout the Abstract, Results, and Discussion to state that canonical Wg overexpression does not detectably enhance leading-edge Arm stabilization or perturb dorsal closure under our conditions, rather than asserting pathway independence. As noted above, we commit to imaging ArmTimer in a wg mutant background to test this directly, complementing our overexpression data with the reciprocal loss-of-function manipulation.

      On the related point that the Timer's failure to detect a Wg-stripe-associated stabilization signal could reflect a sensitivity limitation rather than a true absence of stabilization (Reviewer #2, Weakness #5): we agree and will state this explicitly rather than treating absence of signal as evidence of absence. This does not undermine the positive leading-edge finding, which is not defined relative to the stripe comparison: all embryos, channels, and time points were imaged and rendered using identical laser power and brightness/sensitivity settings, and the leading-edge RFP signal clearly exceeds background under those same acquisition conditions. We also note that detection limits of this kind are a recognized challenge for endogenously tagged reporters of canonical Wnt/β-catenin signaling generally, including in mammalian systems, and cite two studies already in our bibliography that report the same class of limitation: de Man et al. (2021, eLife 10:e66440) and Ambrosi et al. (2022, eLife 11:e64498). We have added this clarification, with these citations, to the Results (paragraph describing Figure 3).

      On the mechanistic link between Arm stability and destruction-complex activity (Reviewer #1):

      We agree that because both ΔArm and Axin overexpression converge on the destruction complex, our data cannot yet fully separate "escape from degradation" from "impaired α-catenin/junctional coupling" as the operative mechanism. We will revise the Discussion to state this ambiguity explicitly and will treat the ArmTimer-AA result (partial α-catenin-binding disruption via phosphosite mutation, independent of destruction-complex regulation) as the strongest current evidence isolating the junctional-coupling mechanism. We also thank Reviewer 1 for pointing us to Pokutta, Choi, Ahlsen, Hansen & Weis (2014, J Biol Chem 289:13589-13601), which structurally and thermodynamically characterized the mammalian cadherin·β-catenin·α-catenin complex and showed that α-catenin binding to β-catenin is a distinct, allosterically regulated interface cadherin binding increases β-catenin's affinity for α-catenin roughly 10-fold, and α-catenin homodimerization independently competes with β-catenin binding. We have added this citation to the Discussion as structural support for treating cadherin engagement, α-catenin coupling, and destruction-complex regulation as mechanistically separable interfaces, and note that the crystallized β-catenin·α-catenin interface provides a structural template for future experiments for example, structure-guided point mutations at the homologous interface residues in Arm, or in vitro binding assays comparing wild-type and threonine-mutant (T111A/T121A) Arm affinity for α-catenin.

      We do not, however, believe a destruction-complex-independent stabilizing allele of Arm is a tractable experiment to close this gap directly: any allele that stabilizes Arm without engaging the destruction complex is, by definition, a Wnt pathway gain-of-function allele, since destruction-complex-mediated degradation is the very regulatory step that canonical Wnt signaling controls. Nor would restricting the allele to a transcriptionally inactive form of Arm cleanly resolve the confound: Wnt/TCF target loci include dedicated repressive TCF-binding sites (Blauwkamp, Chang & Cadigan, 2008, EMBO J 27:1436-1446), so a transcriptionally "dead" stabilized Arm could still alter transcription by disrupting TCF-mediated repression. We therefore treat this as a genuine, currently unresolvable confound of the overexpression approach, and rely instead on the CRY2 optogenetic and ArmTimer-AA results as the strongest available evidence isolating a junctional-coupling contribution. We have added this reasoning, with both citations, to the Discussion.

      On JNK acting on Arm directly vs. indirectly (Reviewer #1; Reviewer #2, Weakness #3 and Recommendation #2):

      This is the most actionable point raised by both reviewers. As noted above, we commit to re-imaging and re-staining our existing JNK-RNAi and JNK-overexpression embryos for E-cadherin to determine whether junctional/polarity architecture is broadly disrupted under these conditions (indirect mechanism) or whether E-cadherin localization is comparatively preserved while Arm stabilization is specifically altered (direct mechanism). In the meantime, we note that a direct mechanism is biochemically plausible: in mammalian cells, JNK phosphorylates β-catenin directly and regulates adherens junction integrity, and JNK activity separately controls the binding of α-catenin to the junctional complex (Lee, Koria, Qu & Andreadis, 2009, FASEB J 23:3874-3883; Lee, Padmashali, Koria & Andreadis, 2011, FASEB J 25:613-623). We cite these as precedent that a direct route from JNK to junctional β-catenin/α-catenin regulation exists in another system, while being explicit that this does not establish the same mechanism in Drosophila dorsal closure that will be tested directly by the E-cadherin re-staining experiment. We have added these citations and this caveat to the Discussion.

      On the Dsh-DEP-to-JNK mechanistic link (Reviewer #1, Public Review #2 and Recommendation #5):

      We agree that our data show the Dsh-DEP requirement and the JNK requirement for dorsal closure as parallel, independent findings rather than a demonstrated linear pathway in our system. To provide context for why we consider a DEP-to-JNK connection a reasonable working hypothesis, we searched the literature in both Drosophila and vertebrates and will cite six additional studies establishing this link: Axelrod et al. (1998) and Axelrod (2001), establishing that DEP-dependent membrane recruitment and unipolar localization of Dishevelled are specifically required for planar polarity signaling, distinct from Wingless signaling; Paricio et al. (1999) and Fanto et al. (2000), showing Dishevelled acts through Misshapen and Rac1/RhoA to the same JNK module used in dorsal closure; and Moriguchi et al. (1999) and Yamanaka et al. (2002), showing biochemically in vertebrates that the DEP domain of Dvl-1 selectively activates JNK independent of β-catenin/TCF-LEF activity, and that this JNK requirement is conserved in Xenopus convergent extension, the vertebrate process most functionally analogous to dorsal closure. We will state explicitly that this precedent, while now cross-species, comes from planar-cell-polarity and convergent-extension assays rather than dorsal closure itself, so it supports the plausibility of a Dsh/Dvl-DEP-to-JNK connection without establishing that the identical pathway operates in our system.

      On the Dsh DIX/DEP domain-separability argument (Reviewer #1, Recommendation #4):

      We thank the reviewer for pointing us to Gammons, Renko, Johnson, Rutherford & Bienz (2016, Mol Cell 64:92-104), which showed that the Wnt signalosome itself is assembled by head-to-tail DEP domain swapping between Dishevelled molecules, and that this DEP-dependent oligomerization is directly required for canonical Wnt pathway activity not restricted to the non-canonical/planar-polarity branch as we had implied. We agree this evidence undercuts our previous interpretation of the DshΔDEP dorsal closure phenotype as evidence for a strong non-canonical/polarity-specific role for the DEP-dependent branch of Dsh. We have revised the Discussion accordingly: we now state that the DEP domain is required for the morphogenetic program culminating in dorsal closure, cite Gammons et al. directly, and note that this requirement does not by itself establish a non-canonical/polarity-specific role, since we cannot rule out a contribution from DEP-dependent canonical Wnt signalosome assembly.

      On force transmission (Reviewer #2, Weakness #1):

      We agree that we have not directly measured force or tension at the leading edge, and that our current data (colocalization with actin/E-cadherin, and functional requirement shown via CRY2 optogenetics and mutant analysis) are consistent with, but do not directly demonstrate, a role in force transmission. We do not have the in-house expertise to perform direct force/tension measurements (e.g., laser ablation, junctional tension assays), so we will not be adding such an experiment in this revision. Instead, we have revised the language throughout the manuscript including two Discussion section headings that previously stated a mechanical role for stabilized Arm as established fact to consistently present the mechanical/force-transmission role as a hypothesis raised by our data, not a demonstrated conclusion, and we retain a clear statement that direct force measurement (ideally in collaboration with groups with the relevant biophysical expertise) is future work rather than a claim we are making in this manuscript.

      On the phosphomimetic threonine mutant (Reviewer #1, Recommendation #2):

      ArmTimer-AA (T111A, T121A) was generated with the expectation that the tyrosine phosphosite mutants (ArmTimer-EE, ArmTimer-FF) would be the primary drivers of any dorsal closure phenotype, given their proposed role in E-cadherin binding; the pronounced zippering defect we observed in ArmTimer-AA was therefore an unanticipated finding rather than a predicted result. We have not generated the reciprocal phosphomimetic ArmTimer-EE(Thr) (T111E, T121E) allele. Generating and characterizing this allele is a substantial undertaking we estimate over a year including allele generation, validation, and phenotypic characterization and we will state this explicitly in the Discussion as planned future work rather than part of the current revision.

      On confirmation of myristoylated-Dsh membrane targeting (Reviewer #1, Recommendation #3):

      We cannot confirm that myristoylation localizes all Dsh protein to the membrane. However, this strategy has extensive prior genetic validation using the identical Src-derived myristoylation sequence: it was originally used to tether Armadillo and shown sufficient for constitutive Wnt pathway activation (Zecca, Basler & Struhl, 1996; Tolwinski & Wieschaus, 2001, 2004), and the same approach was subsequently applied to GSK3 and Dishevelled, in each case producing the expected pathway-activation phenotypes (Mannava & Tolwinski, 2015; Kaur et al., 2017). We have added these citations to the Results where the Myr-Dsh constructs are introduced.

      On overexpression-based perturbations versus endogenous regulation (Reviewer #2, Weakness #4):

      We would like to clarify that most of the Arm alleles used in this study including all of the point-mutant Timer alleles (ArmF1a, ArmTimer-FF, ArmTimer-EE, ArmTimer-AA) central to our mechanistic conclusions were generated as knock-ins at the endogenous ‘arm’ locus via MiMIC/RMCE, not overexpressed. The two exceptions are ΔArm and ArmS56A, expressed from UAS constructs because both are gain-of-function alleles anticipated to be lethal if expressed from the endogenous locus, based on prior experience with similarly stabilizing mutations. Axin overexpression was used because no Axin mutant or knock-in allele was generated for this study; we agree an endogenous Axin allele would be the ideal complement and will state this explicitly as a limitation, while noting that our CRY2 optogenetic perturbations of Arm and α-catenin which act acutely on the endogenous proteins provide an orthogonal line of evidence supporting the same conclusions. We have added this clarification to the Discussion.

    1. eLife Assessment

      This study presents a valuable RNA velocity solution which integrates cell differentiation and gene regulation, with a balance between neuralODE and raw gene space. The evidence supporting the claims of the authors is solid, although inclusion of discussion on the challenges in capturing cell cycle transitions would have strengthened the study. The work will be of interest to scientists working in the field of computational biology and gene regulation.

    2. Reviewer #1 (Public review):

      Summary:

      In the paper, the authors propose a new RNA velocity method, TSvelo, which predicts the transcription rate linearly based on the expression of RNA levels of transcription factors. This framework is an extension of its recent work TFvelo by including unspliced reads and designing a coherent neuralODE framework. Improved performance was demonstrated in six diverse datasets.

      Strengths:

      Overall, this method introduces innovative solutions to link cell differentiation and gene regulation, with a balance between model complexity (neuralODE) and interpretability (raw gene space).

      Comments on revised version:

      I thank the authors for further revision, and I do not have any other concerns. I believe it is an important contribution to this field of trajectory inference and gene regulation.

    3. Reviewer #3 (Public review):

      Despite the abundance of RNA velocity tools, there are still major limitations, and there is strong skepticism about the results these methods lead to. In this paper, the authors try to address some limitations of current RNA velocity approaches by proposing a unified framework to jointly infer transcriptional and splicing dynamics. The method is then benchmarked on 6 real datasets against the most popular RNA velocity tools.

      Comments on revised version:

      The Authors addressed my 2 follow-up comments suitably.

      Thanks for the time you took addressing them. I have no further comments.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the paper, the authors propose a new RNA velocity method, TSvelo, which predicts the transcription rate linearly based on the expression of RNA levels of transcription factors. This framework is an extension of its recent work TFvelo by including unspliced reads and designing a coherent neuralODE framework. Improved performance was demonstrated in six diverse datasets.

      Strengths:

      Overall, this method introduces innovative solutions to link cell differentiation and gene regulation, with a balance between model complexity (neuralODE) and interpretability (raw gene space).

      Comments on revised version:

      The authors have added comprehensive analyses in this revision, and all of my concerns have been very well addressed. Here, I just want to re-emphasize the original points 1 and 3.

      (1) The analysis and clarification are very helpful - thanks! I found that Fig. R1 and R2 are very insightful, as DoRothEA-only returns much worse performance. Please consider adding these two figures to the supp figure and possibly highlighting your setting for edge pruning (down-weights); therefore, the model is more likely to be affected by false negatives than false positives in the TF-target prior.

      We thank the reviewer for the positive feedback and for recognizing the value of the additional analyses. We have added the previous Fig. R1 and Fig. R2 to the Supplementary Information as Fig. S13 and Fig. S14, respectively, and have referred to them in the revised manuscript.

      We have also expanded the description of the TF–target prior used in TSvelo in the “Acquiring Prior Knowledge of Gene Regulatory Relations” subsection of the Methods. As noted by the reviewer, TSvelo is expected to be less sensitive to false-positive TF–target interactions because unsupported edges can be down-weighted during training. In contrast, missing true regulatory interactions are not represented in the prior network and therefore cannot contribute to the learned regulatory dynamics, making the model potentially more sensitive to false negatives.

      (3) Please consider adding some discussion on the challenges in capturing cell cycle transitions.

      We thank the reviewer for this suggestion. We have added a brief discussion in the Discussion section on the challenges of modeling cell-cycle transitions. In particular, cell-cycle progression is often characterized by cyclic dynamics and overlapping transcriptional programs, which can complicate the inference of directional state transitions and regulatory relationships.

      Reviewer #3 (Public review):

      Despite the abundance of RNA velocity tools, there are still major limitations, and there is strong skepticism about the results these methods lead to. In this paper, the authors try to address some limitations of current RNA velocity approaches by proposing a unified framework to jointly infer transcriptional and splicing dynamics. The method is then benchmarked on 6 real datasets against the most popular RNA velocity tools.

      Comments on revised version.

      The Authors addressed all my comments suitably. I'd like to thank them for the time they spent addressing them: the revised paper is much more convincing.

      I have 2 very minor follow-up concerns:

      (1) I appreciated the simulation study, however, no null simulation is present.

      We know RNA velocity tools are inclined to provide false positives: trajectories even when the data doesn't have any.

      I'd be helpful to add null simulations where the data has no trajectories and see if methods erroneously identify any.

      We thank the reviewer for this helpful suggestion. We have added null simulations to evaluate TSvelo and baseline approaches on data without underlying dynamic structure. Specifically, we generated a null dataset including 200 genes and 600 cells by independently sampling spliced (S) and unspliced (U) counts, thereby removing any coherent transcriptional relationship between them.

      When applying scVelo and UniTVelo to this data, no genes passed the velocity gene selection step under the default likelihood-based filtering, and no velocity field could be obtained. We further tested TSvelo, Dynamo, and cellDancer on the same null data and observed that all three methods still produce trajectory-like patterns despite the absence of true dynamics (See Supplementary Information as Fig. S17).

      Including TSvelo, many RNA velocity and trajectory inference approaches assume that they are applied to datasets reflecting underlying dynamic biological processes. We agree that incorporating additional checks during preprocessing could help prevent applying velocity analysis to non-dynamic datasets. We have added this discussion to the revised manuscript.

      (2) Several of the novel analyses are only reported in the Supplementary material and only references in the main text (e.g., "A validation of TSvelo on simulated data is provided in Fig. S1 and Fig. S2 in the Supplementary Information."). This is pity!

      If allowed, I'd add some comments about the new analyses (simulations, computational benchmarks, etc...) also in the main text.

      We thank the reviewer for this suggestion. We agree that several analyses presented in the Supplementary Information provide important support for our conclusions. To improve their visibility, we have expanded the corresponding descriptions in the main text and briefly summarized the key findings of the relevant Supplementary Figures instead of only citing them. These revisions have been made for Fig. S1, Fig. S2, Fig. S10, Fig. S12, Fig. S13, Fig. S14 and Fig. S17. In particular, we have incorporated a summary of the simulation results at the end of the subsection “Estimate RNA Velocity with TSvelo” in the Results section, and added a discussion of the computational benchmarking analyses in the Discussion section. We hope these changes improve the accessibility of these results while maintaining a concise presentation of the main findings.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      I suggest the paper to undergo (very) minor revisions as detailed in the Public Review.

      Simone Tiberi, The University of Bologna

      We sincerely thank all reviewers for their thoughtful suggestions, which have helped improve the clarity and overall presentation of the manuscript.

    1. eLife Assessment

      This study demonstrates that endothelial toll-like receptor 4 is a central regulator of leptomeningeal inflammation in neonatal E. coli meningitis. The data are derived from cell-type-specific gene knockouts in mice and cultured endothelial cells and are convincing. This work is important as it advances our understanding of host cellular processes and molecular pathways underlying meningitis pathogenesis and specifically expands the knowledge of how breakdown of the blood brain barrier contributes to the pathogenesis.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Seegren and colleagues demonstrate that in a mouse model of neonatal E. coli meningitis, loss of toll-like receptor 4 (TLR4) in VE-cadherin+ endothelial cells and a subset of meningeal fibroblasts leads to a marked decrease in transcriptional dysregulation across multiple leptomeningeal cell types, a decrease in vascular permeability, and a decrease in macrophage abundance. In contrast, loss of macrophage TLR4 had less pronounced effects. Using cultured wildtype and TLR4-knockout endothelial cells, the authors further demonstrate that TLR4 signaling leads to reversible internalization of the tight junction protein claudin-5, establishing a potential mechanism of increased vascular permeability. Authors also show that claudin-5 internalization is independent of NF-κB. Finally, the authors use RNA-sequencing of wildtype and TLR4-knockout endothelial cells to define the TLR4-dependent cell-autonomous transcriptional response to E. coli.

      Comments on revised version.

      The authors have considerably improved and strengthened the work through the addition of new experimental data, new data analyses, and modifications to their interpretation. Notably, the authors used additional Cre-reporter mice to clarify that Cdh5-CreER is active in endothelial cells and some meningeal fibroblasts, and thus revised nomenclature and interpretation to acknowledge that the Tlr4fl/-;Cdh5-CreER cKO (Tlr4-VEKO) is not exclusively endothelial. The authors also demonstrated that Tlr4-VEKO does not affect peripheral E.coli burden, but acknowledge that changes to periphery-derived signals (e.g., cytokines) may contribute to observed leptomeningeal phenotypes.

      The authors added PCA plots to show similarity in gene expression shifts across biological replicates (mice). This provides support for the claim that Tlr4-VEKO attenuates infection-associated transcriptional changes. With respect to differential expression analysis, I agree with authors that characteristics of individual cells (e.g. heterogeneity) are of interest. I remain concerned, however, that the formal differential analysis strategy appears to consider cells as independent experimental units, which they are not because a single cell cannot be randomly assigned to an experimental group (control or cKO, uninfected or infected). The mouse is the correct experimental unit for a comparison across these groups because it can be randomized. I appreciate that many of the gene expression changes appear consistent across mice (e.g. Figure 1 - Figure supplement 7) and that there are clear infection- and genotype-associated phenotypes in other assays. I would simply caution that the authors' analysis strategy likely leads to a larger number of type I errors (false positives) than is generally accepted; a mixed (hierarchical) model or pseudo-bulk approach would be more appropriate for future studies.

    3. Reviewer #2 (Public review):

      Summary:

      The authors use a postnatal mouse model of E. coli bacterial meningitis and a mouse brain endothelioma cell line combined with cell type specific gene deletion to study the function of endothelial TLR4, a cell surface receptor that recognizes gram positive bacterial wall components, in the local leptomeningeal (LPM) response with a focus on endothelial barrier breakdown mediated by TLR4. Single cell transcriptional profiling and imaging studies using wholemount preps of the LPM support that LPM endothelial, CD206+ local macrophage and LPM fibroblast and arachnoid barrier cell inflammatory response and is abrogated in endothelial specific KO of TLR4, pointing to a role for endothelial TLR4 in local LPM response. Culture studies using Bend3.1 cells (a mouse brain endothelioma cell line) support a direct role for TLR4 in the bacteria-mediated inflammatory response and in internalization of Cldn5 via the endosomal-lysosomal pathway, resulting in loss of barrier integrity

      Strengths:

      The local LPM cell response in meningitis and the role of specific LPM cells in inflammation and CNS barrier breakdown has not been extensively studied, despite ample evidence for primary immune response in the meninges in human patients and in animal models. The authors employ a robust, multi-model approach using both in vivo and in vitro models with cell-type specific knockout to study the function of TLR4 in brain endothelial cell response. The authors nicely combine functional barrier assays with IF for junctional localization in their experimental design and they delve into potential mechanisms of Cldn5 internalization using markers of endosomal-lysomal pathway localization. The authors also describe a new type of barrier assay using a streptavidin-coated plates upon which barrier forming cell cultures can be plated, this could be a very useful alternative or complement to other size-selective barrier assays and presumably could work for other barrier forming cell types, like epithelial cells.

      Comments on revised version.

      In their revision, the authors addressed prior noted weaknesses with new data and analysis. They now show that TLR4-VE-cad cKO mice have a largely similar disease progression as control mice, including increased bacterial burden in the LPM and brain. This underscores that that the reduced vascular leakage and blunted inflammatory response is due to loss of TLR4 response to bacteria on VE-cad recombined cells and not because the mice are protected from meningitis. The authors also performed additional experiments to show that Cldn5 internalization via the endosomal-lysosomal pathway is independent of NFKB signaling. The authors also added in important discussion points about how their results fit into the broader literature on TLR4 in BBB endothelial cell junctional protein localization and prior work on meningitis in global TLR4.

    4. Reviewer #3 (Public review):

      Summary:

      This study investigates the molecular underpinnings of immune responses in the leptomeninges in neonatal bacterial meningitis. Bacterial meningitis is a major disease burden, particularly for neonates, and it has previously been noted that the meningeal immune environment in infants is permissive to opportunistic infection (Kim et al., Sci Immunol, 2023). There is less known about the contribution of the stromal compartment to meningeal immune responses. Seegren et al. interrogate the role of leptomeningeal endothelium in host defense in E. coli infected neonatal mice using mouse genetic tools to delete the LPS receptor Tlr4 from either endothelial cells/stromal cells (using Cdh5-CreER) or myeloid cells (using LysM-Cre). The authors use snRNAseq, cleared cortical mounts, and in vitro work to define the impact of E. coli infection on leptomeningeal endothelial cells. This study uses a range of innovative techniques to probe the role of the stromal compartment in meningitis. With additional experiments to confirm the specificity of their Cre models, this strengthens the interpretation of the study significantly. The only major weakness is the inability to confirm TLR4 knockout in myeloid cells.

      Strengths:

      This study makes excellent use of cleared cortical mounts to examine the biology of the leptomeninges, in particular, changes to the endothelium, with unprecedented detail. In combination with high-quality sequencing data provide new insights into the impact of meningitis on the leptomeninges. The data presented by the authors is of very high quality.

      The authors have also done substantial work to address my two major comments regarding 1) the specificity of their Cre systems and 2) peripheral impacts of the interventions.

      (1) The authors identified and acknowledged some impacts in the leptomeningeal stroma (the relatively high level of recombination in ECs vs FBs presumably reflects a single low dose being given, where other groups have done more aggressive tamoxifen regimens that drive recombination in FBs as well). Given the incomplete recombination in the leptomeningeal FBs, I agree with their conclusion that it is probably endothelial driven. Acknowledging the contributions of other myeloid cells with the L. The Cre-NLS experiments with nuclear markers provided excellent data and had beautiful staining.

      (2) The authors did not observe differences in bacterial burden in peripheral organs in either CKO model, suggesting that CNS impacts are not downstream of peripheral bacterial control.

      Weaknesses:

      (1) The inducible Cre lines used by the authors target peripheral tissues as well as CNS tissues. Although this is mollified by the lack of impact on peripheral disease burden.

      (2) The authors were not able to confirm TLR4 knockout in myeloid cells, and this caveat is acknowledged. The lack of response in TLR4 VEKO mice strongly suggests successful conditional knockout.

      (3) The cell line model (bEnd.3) is a relatively low fidelity model of BBB endothelial cells. The authors acknowledge this, and it is likely that endothelial cell responses to LPS are highly conserved.

      (4) It is perhaps not surprising that Tlr4 is required for meningitis responses with E. coli. However, it is unclear if these findings can be generalised to other, more common, meningitis infections (streptococcal/pneumococcal).

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Seegren and colleagues demonstrate that in a mouse model of neonatal E. coli meningitis, loss of endothelial toll-like receptor 4 (TLR4) leads to a marked decrease in transcriptional dysregulation across multiple leptomeningeal cell types, a decrease in vascular permeability, and a decrease in macrophage abundance. In contrast, loss of macrophage TLR4 had less pronounced effects. Using cultured wild-type and TLR4knockout endothelial cells, the authors further demonstrate that TLR4-NF-κB signaling leads to reversible internalization of the tight junction protein claudin-5, establishing a potential mechanism of increased vascular permeability. Finally, the authors use RNA sequencing of wild-type and TLR4-knockout endothelial cells to define the TLR4dependent cell-autonomous transcriptional response to E. coli.

      Strengths:

      (1) The authors address an important, well-motivated hypothesis related to the cellular and molecular mechanisms of leptomeningeal inflammation.

      (2) The authors use model systems (mouse conditional knockouts and cultured endothelial cells) that are appropriate to address their hypotheses. The data are of high quality.

      Weaknesses:

      (1) The authors perform single-nucleus RNA-seq on dissected leptomeninges from control and E. coli-infected mice across three genotypes (WT, Tlr4MKO, and Tlr4ECKO). A major discovery from this experiment, as summarized by the authors, is: "Tlr4ECKO mice exhibited a global attenuation of infection-induced transcriptional responses across all major leptomeningeal cell types, as judged by the positions of cell clusters in the UMAP." This conclusion could be considerably strengthened by improving the qualitative and quantitative analysis.

      Thank you for this comment. We agree that the UMAP-based interpretation would benefit from additional qualitative and quantitative support. We have expanded the snRNA-seq analysis with additional images and supplemental figures (Figure 1 – figure supplement 3, Figure 1 – figure supplement 5, and Figure 1 – figure supplement 6). The first and third of these new supplemental figures show dot plots for each major leptomeningeal cell type, for each genotype, for the two experimental conditions (infected vs. uninfected), and for individual genes in three immune-related gene sets (NF-kB and TNF-α, JAK-STAT, and IFN-ɣ), providing a more explicit comparison of infection-induced transcriptional responses across genotypes. The second of these new supplemental figure shows principal component analysis (PCA) of the individual snRNA-seq datasets (one mouse per dataset) for each genotype and experimental condition, demonstrating that the observed transcriptional shifts are consistent across biological replicates. Finally, Figure 1 – figure supplement 7, which was included in the original submission, shows changes in the most up- and down-regulated genes (based on adjusted p-value or fold change) in endothelial and myeloid cells across individual mice and genotypes/conditions, further supporting the genotype-dependent effects at the level of individual animals.

      (2) The authors interpret E. coli infection-induced increases in leptomeningeal sulfo-NHSbiotin as evidence of compromised BBB integrity (i.e., extravasation from the vasculature) (Results, page 7), but another possible route in this context is sulfo-NHS-biotin entry from the dura across a compromised arachnoid barrier. The complete rescue in Tlr4ECKOs is strongly suggestive that the vascular route dominates, but it would strengthen the work if the authors could assess arachnoid barrier fidelity (e.g. via immunohistochemistry). At a minimum, authors should mention that the sulfo-NHS-biotin signal in this context may represent both vascular and arachnoid barrier extravasation.

      Thank you for this comment. We agree that our data cannot rule out leakage across the arachnoid barrier during infection. While the rescue observed in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice strongly supports a dominant vascular contribution, we acknowledge that the sulfo-NHS-biotin signal may reflect permeability at both the vascular and arachnoid barriers. We do not think there is a clear way to directly test this possibility functionally, since the arachnoid barrier appears intact by confocal microscopy. Subtle differences in barrier cell morphology might be detectable by electron microscopy, but this would not definitively address whether infection permits molecular passage across the arachnoid barrier. We have followed the reviewer’s suggestion and revised the Results section to reflect this interpretation. Specifically, we added the following: “The simplest interpretation of these data is that the site of sulfo-NHS biotin leakage is primarily vascular. However, we cannot exclude some contribution from increased arachnoid barrier permeability.”

      (3) The authors state that "deletion of TLR4 prevented both NF-κB nuclear translocation and Cldn5 internalization in response to E. coli (Figure 4A-D)" (Results, page 9). In Figures 4C and D, however, there is no indicator of a statistical test directly comparing the two genotypes. A comparison of within-genotype P-values should not be used to support a genotype difference (PMID: 34726155).

      Thank you for pointing out this omission. We have updated the figures so that the between-genotype p-values are shown for those panels (including this panel) that had not previously shown them.

      (4) In the first paragraph of the Results, the authors summarize the meningeal layers as (1) pia, (2) subarachnoid space, (3) arachnoid, and (4) dura, and then state "The second and third layers constitute the leptomeninges." This definition of leptomeninges seems to omit the pia, which is widely considered part of the leptomeninges (PMID: 37776854).

      Thank you for pointing out this error, which has now been corrected.

      (5) The Cdh5-CreER/+;Tlr4 fl/- mouse lacks TLR4 in all endothelial cells (i.e., in peripheral organs as well as CNS/leptomeninges), and, as the authors note, the periphery is exposed to E. coli. It would be helpful if the authors could comment in the Discussion on the possibility that peripheral effects (e.g., peripheral endothelial cytokine production, changes to blood composition as a result of changes to peripheral endothelial permeability) may contribute to the observed leptomeningeal phenotypes.

      Thank you for raising this point. We agree that peripheral responses could contribute to the observed leptomeningeal phenotypes in this model. We have added two sentences to the second paragraph of the Discussion to address this: “We note that these experiments do not distinguish between local vs. distal anatomic sources of LPS or downstream effector molecules, such as cytokines, that activate the leptomeningeal inflammatory response (Huang et al., 2021). Histologic observations of RFP-expressing E. coli in the brain, liver, and lungs, together with positive blood cultures, indicate substantial systemic dissemination in this model. Thus, the inflammatory responses of leptomeningeal cells likely reflect exposure to bacterial products and inflammatory mediators derived from both local meningeal and peripheral sources.”

      Reviewer #2 (Public review):

      Summary:

      The authors use a postnatal mouse model of E. coli bacterial meningitis and a mouse brain endothelioma cell line combined with cell-type-specific gene deletion to study the function of endothelial TLR4, a cell surface receptor that recognizes gram positive bacterial wall components, in the local leptomeningeal (LPM) response with a focus on endothelial barrier breakdown mediated by TLR4. Single-cell transcriptional profiling and imaging studies using whole-mount preps of the LPM support that LPM endothelial, CD206+ local macrophage and LPM fibroblast and arachnoid barrier cell inflammatory response and is abrogated in endothelial-specific KO of TLR4, pointing to a role for endothelial TLR4 in local LPM response. Culture studies using Bend3.1 cells (a mouse brain endothelioma cell line) support a direct role for TLR4 in the bacteria-mediated inflammatory response and in internalization of Cldn5 via the endosomal-lysosomal pathway, resulting in loss of barrier integrity

      Strengths:

      The local LPM cell response in meningitis and the role of specific LPM cells in inflammation and CNS barrier breakdown have not been extensively studied, despite ample evidence for primary immune response in the meninges in human patients and in animal models. The authors employ a robust, multi-model approach using both in vivo and in vitro models with cell-type-specific knockout to study the function of TLR4 in brain endothelial cell response. The authors nicely combine functional barrier assays with IF for junctional localization in their experimental design, and they delve into potential mechanisms of Cldn5 internalization using markers of endosomal-lysosomal pathway localization. The authors also describe a new type of barrier assay using a streptavidin-coated plate upon which barrier-forming cell cultures can be placted, this could be a very useful alternative or complement to other size-selective barrier assays and presumably could work for other barrier forming cells types, likely epithelial cells.

      Weaknesses:

      (1) There are no measures of bacterial burden in peripheral organs, blood, in the LPM or brain in the TLR4 endothelial cKO mice. Lack of TLR4 in endothelial cells could prevent bacterial 'access' into the LPM and brain, essentially preventing meningitis and leading to a lack of inflammatory responses in the LPM-located cells simply because there is no bacteria present. Bacteremia may also be reduced, as might inflammatory responses in peripheral organs with TLR4-deficient peripheral endothelium. Bacterial counts and inflammatory measures in peripheral organs and blood are important to better understand the mechanism(s) underlying the reduced inflammatory profile in LPM cells and no LPM endothelial breakdown in the Tlr4 endothelial cKO mice. In other words, does deleting TLR4 in EC protect against the development of meningitis by somehow blocking bacteria access to the LPM (this would be supported by low or no CFU counts in infected Tlr4 endothelial cKO) or is it what the authors appear to propose in Figure 1J that TLF4 in EC is the only cell responding to the bacteria to trigger the immune cascade in the LPM? More data is needed to resolve this, as this is a major claim of the paper.

      Thank you for this comment. We agree that it is important to distinguish whether the reduced inflammatory response in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice reflects altered bacterial burden versus altered host sensing. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4CKO mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs of infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice in E. coli burden; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and the time of sacrifice 24 hours later (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid leptomeningeal cells does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (2) The authors look at the underlying cortical response (cerebral vasculature for ICAM and immune cells) but do not use markers that could identify microglia (Iba1), the primary resident immune cell (CD206 is not useful, at this stage, in perivascular macrophages that are extremely sparse in the postnatal brain). This would be important to better study the impact on CNS resident immune cell morphological activation.

      Thank you for this comment. In response, we have analyzed Iba1 staining in the cortex in infected vs. uninfected mice. This is shown in Figure 2 – figure supplement 3. These data demonstrate a several-fold increase in Iba1 immunostaining in infected compared to uninfected cortex, consistent with increased microglial activation in response to infection. There is no statistically significant difference between infected WT and infected Cdh5-CreER; Tlr4CKO mice in Iba1 staining in cortex.

      (3) The authors suggest that Cldn5 junctional localization is selectively disrupted upon bacterial exposure, mediated by TLR4 - they suggest this based on studying PECAM, GLUT1, ZO-1 and B-catenin (all normally junction or cell surface located in cultured Bend3.1) in relationship to Cldn5 localization (normally high) - it is possibly these are also impact by bacteria exposure (maybe through different mechanisms?) - a better measure would be to use the similar cyto/PM measure they do for Cldn5 in Fig. 4D and to evaluate this or to use intensity measurements.

      Thank you for this comment. As the reviewer noted, the analysis of Cldn5 localization with vs. without E. coli exposure and in WT vs. Tlr4KO bEnd.3 cells (shown in Figure 4B and D) – uses Cell Trace to partition the image into cytoplasmic vs. plasma membrane territories. For the analyses in Figure 5, we wanted to compare the localization (and potentially re-localization) behaviors of a variety of subcellular markers with the localization and re-localization of Cldn5 following E. coli exposure. By directly measuring the % overlap of the two immunostains, we get that data. We note that the goal of this analysis is to assess relative co-localization with Cldn5 rather than absolute subcellular partitioning of each marker. While this analysis could have been extended to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories, it is clear by visual inspection of Figure 5A-C that beta-catenin, ZO-1, and PECAM1 remain plasma membrane-associated with E. coli exposure, and GLUT1 goes from the part of the plasma membrane not involved in cell-cell contact without E coli exposure to cytoplasmic with E. coli exposure (as judged by the appearance of a nuclear “shadow” after E. coli exposure). Thus, we do not believe that additional cytoplasmic vs. plasma membrane quantification for these markers would alter the interpretation. The main reason that we did not extend this analysis to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories is because that would introduce the Cell Trace localization as an additional variable.

      (4) The discussion could benefit from delving more into the prior literature on E coli mediated breakdown of junctions in cultured human microvascular brain endothelial cell model and critical host-pathogen interactions of the bacteria with ECs (PMID: 14593586), and how this might involve TLR4.

      Thank you for this comment. Two paragraphs addressing the prior literature have now been added to the discussion.

      (5) It would be important to discuss how their results relate to earlier studies on TLR4-/- and TLR2-/- global knockout mice and protection vs vulnerability to development of meningitis (see PMCID: PMC3524395) - this paper showed that TLR4 global KO mice have increased susceptibility to die from meningitis and have much higher CFU counts in the CNS. In this manuscript and their prior work (Wang et al., 2023), this group shown that both global TLR4-/- mutants and their EC-specific KO have reduced barrier permeability, but we don't have any information about CFU or susceptibility to death from meningitis in their models.

      Thank you for these comments. The model we use – subcutaneous injection of E. coli (a clinical isolate from an infant with meningitis) at postnatal day (P)5 – results in the death of the infected mouse within 2 days (shown in Figure 1 – figure supplement 3 in Wang et al. 2023). Our analyses of infected mice were conducted 24 hours after infection. As noted in the reply to comment #1, in the revised manuscript we present a clinical assessment of WT vs. Cdh5-CreER; Tlr4CKO mice 24 hours after infection based on (1) a quantitative microscopic analysis of E. coli burden in the brain (visualized based on RFP fluorescence in the E. coli used here), (2) quantifying CFUs in blood and (3) mouse weights at P5 and P6, a sensitive indicator of overall health since this is a time when mice are normally gaining weight rapidly (~25% weight gain per day). These data (shown in Figure 2 figure supplement 4) indicate that bacterial burden and disease severity are similar between genotypes in our model. In Wang et al., 2023, we did not conduct a quantitative clinical assessment of WT vs. Tlr4-/- mice following infection, but by visual inspection, infected WT and Tlr4-/- mice appeared to have similar downhill clinical trajectories. We have expanded the Discussion to relate these findings to prior studies of global TLR4 and TLR2 knockout mice, noting that differences in experimental models and the distinction between global versus VECadCreER-specific deletion may account for the differing outcomes reported.

      Comment on the paper listed by the reviewer (PMCID: PMC3524395).

      The cited study demonstrates that global TLR4 deficiency leads to increased bacterial burden and mortality, indicating an essential role for TLR4 in host defense and bacterial clearance. In our study of Cdh5-CreER; Tlr4CKO mice, bacterial burden and disease severity at 24 hours post-infection are similar between WT and Cdh5-CreER; Tlr4CKO mice, indicating that Cdh5-CreER; Tlr4CKO does not alter the clinical course at this time point. This difference is noted in the Discussion section.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the molecular underpinnings of immune responses in the leptomeninges in neonatal bacterial meningitis. Bacterial meningitis is a major disease burden, particularly for neonates, and it has previously been noted that the meningeal immune environment in infants is permissive to opportunistic infection (Kim et al., Sci Immunol, 2023). There is less known about the contribution of the stromal compartment to meningeal immune responses. Seegren et al. interrogate the role of leptomeningeal endothelium in host defence in E. coli infected neonatal mice using mouse genetic tools to delete the LPS receptor Tlr4 from either endothelial cells (using Cdh5-CreER) or macrophages (using LysM-Cre). The authors use snRNAseq, cleared cortical mounts, and in vitro work to define the impact of E. coli infection on leptomeningeal endothelial cells. This study uses a range of innovative techniques to probe the role of the stromal compartment in meningitis.

      Strengths:

      This study makes excellent use of cleared cortical mounts to examine the biology of the leptomeninges, in particular, changes to the endothelium, with unprecedented detail. In combination with high-quality sequencing data provide new insights into the impact of meningitis on the leptomeninges. The data presented by the authors is of very high quality.

      Weaknesses:

      The weaknesses of the study were in terms of interpretation and perhaps study design.

      (1) Most importantly, the authors need to provide additional validation of their conditional knockout models. The authors need to confirm that the Cdh5-CreER does not impact leptomeningeal fibroblasts and to confirm gene deletion in macrophages.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclearlocalized GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite correct. The Results section of the revised manuscript includes an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      (2) The authors could also strengthen the paper by providing data on the impact of these conditional knockout models on the course of meningitis and bacterial burden.

      Thank you for this comment. We agree that these additional analyses strengthen the manuscript. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences in bacterial burden between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) Finally, it is perhaps not surprising that Tlr4 is required for meningitis responses with E. coli. However, it is unclear if these findings can be generalised to other, more common, meningitis infections (streptococcal/pneumococcal).

      At present, it is an open question whether TLR4 plays as a large a role in meningitis caused by other gram-negative bacteria and whether TLR2 plays a similarly large role in meningitis caused by gram-positive bacteria. In the Discussion, the last two sentences under “Limitations of the study” summarize this point: “Finally, the present study focused on E. coli K1, the dominant Gram-negative neonatal pathogen. Future work could assess TLR signaling in response to other bacterial pathogens, such as Group B Streptococcus.”

      (4) There are additional minor issues; for instance, the arachnoid fibroblast 2 population appears to closely resemble dural border cells.

      Thank you for this comment. That is correct, and we have changed the nomenclature to “dural border cells”.

      (5) The cell line model (bEnd.3) is a relatively low-fidelity model of BBB endothelial cells, and this should be acknowledged.

      Thank you for this comment. That is correct. Despite being brain-derived, bEnd.3 cells have lost many BBB-specific attributes. Their responses might best be considered as generic endothelial responses rather than brain-specific endothelial responses. This is now stated in the Results section: “Although they are brain-derived, bEnd.3 cells lack many BBB-specific attributes and, therefore, they likely exhibit generalized endothelial responses rather than brain-specific responses to bacterial exposure.”

      With these caveats, it is difficult to be certain that the endothelium alone is the driver of meningeal immune responses in meningitis, and what the impact of these is.

      We agree with this critique. As noted above, the expression of Cdh5-CreER in essentially all endothelial cells and in a subset of dural border cells and leptomeningeal fibroblasts means that the comparison of TLR4 CKO with Cdh5-CreER vs. Lyz2-Cre is assessing phenotypes driven by TLR4 signaling in endothelial plus a subset of other non-myeloid cells vs. TLR4 signaling in myeloid cells. We have revised the text to reflect this more precise understanding of Cdh5-CreER specificity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Transcriptomic analysis: The analysis and display of the single-nucleus RNA-seq data should be improved. The authors could perform a more granular, unbiased clustering of each cell class in the combined dataset and then compare the proportion of each experimental group (genotype x control/infected) in each cluster. At present, it appears the differentially-expressed genes (DEGs) shown in Figure 1 were identified using the Seurat FindMarkers function with default parameters (Methods). This considers each cell as an independent experimental unit and is therefore not appropriate for a comparison of control versus infected groups (see e.g., PMID 34584091, 35880426. The authors should implement a statistical analysis strategy that considers true biological replicates (mice, as shown in Supplementary File 1).

      We do not fully agree with this critique. We agree that biological replication at the level of individual mice is important for interpreting these data, but within each mouse, the characteristics of individual cells is also of interest, including the degree of heterogeneity, the sample size for a given cell cluster, and the statistical significance of any observed changes in transcript abundance. As requested, we have prepared a new supplemental figure (Figure 1 – figure supplement 5) showing a principal component analysis of the scRNA-seq data for each mouse (one mouse was used for each snRNA-seq dataset) and for each of the principal leptomeningeal cell types. This analysis shows, for example, that the three infected Cdh5-Cre; Tlr4flox/- mice have transcriptomes for each of the six cell clusters that are very similar to the transcriptomes of the two uninfected WT and the two uninfected Cdh5-CreER; Tlr4flox/- mice. Thus, the genotype- and condition-dependent effects are consistent across biological replicates. At the most granular level, Figure 1 – figure supplement 7, which was part of the original submission, shows for the most up- and down-regulated genes (based on adjusted p-value or based on fold-change) in endothelial cells and in myeloid cells how individual transcript abundances change for each mouse and for each genotype/condition.

      (2) The authors use immunohistochemistry to assess claudin-5 "disorganization and redistribution" (Results, pages 7-8 and Figures 3A-B). They state that "Tlr4ECKO mice showed minimal changes in the distribution of Cldn5, implying that cell autonomous endothelial TLR4 signaling regulates tight-junction organization." It is not clear, however, that the quantified parameter (Cldn5+ area relative to total area) would be an accurate readout of claudin-5 organization/distribution (i.e., subcellular localization) as it would also be sensitive to claudin-5 expression, vascular density, and vessel diameter. The authors use a similar assessment of ZO-1 to suggest that changes to claudin-5 are not due to a "generalized disassembly of TJs" and could also use this to argue that the above potential confounds (vascular density, vessel diameter) do not change, but the data in Figure 3 - Figure Supplement 1B, lower panel, show that infection does cause an increase in ZO-1 area relative to total area (P = 0.0004). Thus, the statement in the results "Zonula Occludens-1 (ZO-1) [...] remained unchanged during infection (Figure 3 - figure supplement 1)" is not accurate. The authors should revise this section to ensure their conclusions are aligned with the presented data.

      Thank you for this comment. The reviewer is correct that the Cldn5 area measurement is unable to deconvolve the various factors that might contribute to it (vessel density and diameter, and Cldn5 distribution). This part has been rewritten. “Consistent with prior findings (Wang et al., 2023), both WT and Tlr4<sup>MKO</sup> mice showed an increase in the area occupied by Cldn5 in the leptomeninges following infection, likely referable to both increased vessel diameter and a redistribution of Cldn5 within ECs (Figure 3A-B; Figure 3 – figure supplement 1C).”

      The reviewer is also correct about our initial description of the ZO-1 data. What we meant to write and what the revised manuscript now shows is: “The area occupied by Zonula Occludens-1 (ZO-1), a tight junction scaffold protein, showed a modest but statistically significant increase in WT leptomeningeal vessels but no significant change in Tlr4<sup>VEKO</sup> leptomeningeal vessels during infection (Figure 3 – figure supplement 1A and B).”

      (3) In the Methods, under Mouse Models and E. coli Infection, the authors state, "The Cdh5-CreER line (Monvoisin et al., 2023) was the same line used in Wang et al. (2023)." However, there is no mention of Cdh5-CreER in Wang et al. (2023). Could authors please clarify? Also, because this appears to be an inducible Cre, the authors must include details on the dose and timing of tamoxifen or 4-OHT used in this study.

      Thank you for catching that error. We meant to reference Wang et al (2025), not Wang et al (2023). [Wang et al (2025) is: Wang Y, Rattner A, Li Z, Smallwood PM, Nathans J. (2025) Vascular endothelial-specific loss of TGF-beta signaling as a model for choroidal neovascularization and central nervous system vascular inflammation. Elife 14:RP107018.] This has now been corrected.

      We have now included the details related to 4HT injection in the Methods section “Mouse Models and E. coli Infection”. These are intraperitoneal injection at P2 with 40- 50 µL of 2 mg/ml 4HT.

      (4) The legend for Figure 1A is "Schematic of the leptomeninges", but the figure shows the entire brain-skull interface, including underlying cortex, leptomeninges, dura, and skull.

      Thank you. Corrected.

      (5) Page 7, typo: "In the brain, CD206+ cell were too sparse ..." Should be "cells".

      Thank you. Corrected.

      (6) Page 14, typo: "... could represents a double-edged ..." Should be "represent".

      Thank you. Corrected.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform CFU counts from LPM, dura, brain, peripheral organs (liver) in infected v mock mice from control v TLR4 EC-cKO.

      Thank you for this comment, with which we agree. We have addressed this by quantifying bacterial burden and assessing disease severity in WT and Cdh5CreER; Tlr4floxed mice. Specifically, we performed CFU measurements in blood, monitored mouse weights at P5 and P6, and histologically surveyed the E. coli-RFP signal (i.e., E. coli burden) in brain, liver, and lung. These analyses show that bacterial burden and disease progression are comparable between WT and Cdh5-CreER; Tlr4floxed mice at 24 hours post-infection. These data are presented in Figure 2 – figure supplement 4 and described in the Results.

      (2) Lyz2Cre/+ is used to delete TLR4 from macrophages, but recombination efficiency (in LPM BAMs) is described as only partial, suggesting that TLR4-response in LPM BAMs (and potentially macrophages in the dura) is at least partially intact. It undercuts conclusions that can be made using this line.

      Thank you for this comment. We have conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The Results section text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Also, as noted in the reply to comment 4 below, a direct analysis of Tlr4 recombination efficiency is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (3) Inflammatory responses [qPCR] from peripheral organs and also physiological measures in the pups [weight post-infection, time to moribund or death curves] in control v TLR4 EC-cKO and TLR4 mac-cKO.

      Thank you for this comment. We have not conducted a qPCR analysis of inflammatory gene expression in peripheral organs because (1) the dramatic upregulation of these transcripts in the leptomeninges, (2) the presence of E. coli in blood and peripheral organs, and (3) the clinical assessment (cessation of weight gain) all predict that such an analysis would reveal a large up-regulation of inflammatory gene expression throughout the body. More specifically, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (4) The conditional macrophage line is problematic due to the partial recombination. I question the utility of including this unless they can come up with a way resolve the response of recombined TLR4 macrophages vs ones that are not (could they use the single cell data to pick this a part? Are TLR4-null cells and TLR4 'wt' cells transcriptionally similar in the infected condition, suggesting TLR4 is not doing much in the macs, potentially due to alternate TLRs?). There are good BAM Cre lines that have been described [Lyve1-cre would be good for LPM BAMS, the other is Pf4-cre, see https://pmc.ncbi.nlm.nih.gov/articles/PMC7375817/ - just as an FYI for the future].

      Thank you for this comment. As noted in the reply to point 2 (above), we have conducted a more in-depth analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      We agree that, based on Figure 6 in the cited paper [McKinsey et al (2020) A new genetic strategy for targeting microglia in development and disease eLife 9:e54590], the Pf4-Cre line may be superior to the Lyz2<sup>Cre</sup> line that we used for recombination in leptomeningeal myeloid cells. Unfortunately, we missed this paper in our literature searches, probably because it focuses on a microglial CreER line, P2ry12-CreER, and the Pf4-Cre line is not mentioned in the title or abstract. Our decision to use the Lyz2<sup>Cre</sup> line was based on an extensive comparison among myeloid Cre lines showing that Lyz2<sup>Cre</sup> was the most efficient [Abram CL, Roberge GL, Hu Y, Lowell CA. 2014. Comparative analysis of the efficiency and specificity of myeloid-Cre deleting strains using ROSA-EYFP reporter mice. J Immunol Methods 408:89-100.] However, the Abram et al study did not look at the leptomeninges. Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (5) Figure 1 - Figure Supplement 2 - the authors nicely break down the pathway response [NFKB and TNF] in EC and macs, it would be great to have similar information for the fibroblasts (in the main figure or the supplement). Does their inflammatory response show a similar pattern?

      Thank you for this suggestion. We have now done that analysis and present it in Figure 1 – figure supplement 3. For completeness, we also performed the same type of analyses for JAK-STAT signaling and IFN-gamma response and these are shown in Figure 1 – figure supplement 6. The principal conclusion is that across all major leptomeningeal cell types, the Cdh5-CreER; Tlr4floxed samples (i.e., Tlr4 KO’d in non-myeloid cells) show much reduced transcriptome changes with infection.

      (6) What is ASC and Cd206 quantification measuring, and how does this relate to 'activation' - is this the intensity of signal or a morphological change? What is the precedence for using ASC (citations)? In their prior work, they showed no change in CD206 number, so a significant increase upon infection here, it's confusing exactly what is being studied. Also, loss of Lyve1 is a well-accepted measure of activation that they have previously used, adding that it could be helpful. This is not a major issue since they have robust data that the macrophages are not transcriptionally activated. Clarification of what exactly is being measured would be sufficient (in the text).

      CD206 immunostaining, which reveals myeloid cell morphology, shows that, with E. coli infection, myeloid cells convert from a more compact morphology to a more expanded morphology. This is now explained more fully in the Results section.

      Regarding ASC, changes in the state of ASC aggregation and ASC subcellular localization have been used by others to monitor immune cell responses to inflammatory signals (Sester et al., 2016; Franklin et al., 2018). While this change in subcellular localization may explain part of the increase in immunostained area in myeloid cells in the infected mice (Figure 2D), the increase in the area of ASC immunostaining largely reflects a shift of myeloid cells from a compact to a more extended morphology. This is now explained more fully in the Results section. We have also added two references (Sester et al., 2016; Franklin et al., 2018) that described how ASC distribution changes with inflammation.

      Regarding LYVE1, we observe a decrease in LYVE1 transcript abundance in myeloid cells with infection, as predicted. Given the large amount of other data that document myeloid activation with infection, we have elected not to include this.

      (7) The authors suggest the internalization of Cldn5 is not due to NFKB downstream signaling that includes transcriptional mechanisms because it happens as early as 1 hour, prior to NFKB localization to the nucleus. However, a lot of their experiments, including on endosomal-lysosomal protein co-localization are done at 4 hours, when their RNAseq data show robust NFKB-mediated gene upregulation and (though not tested) potentially protein production of factors that can act back on the cells, including to impact endo-lysosomal processing. Without studies at earlier timepoints post-bacteria exposure, separating these two mechanisms is difficult.

      Thank you for this comment. We have explored this question by looking at Cldn5 internalization in bEnd.3 cells at 1 hour after E. coli exposure, and the data clearly show that internalization occurs within 1 hour. Additionally, we have conducted this experiment in the presence of 1 uM ACHP, an IKK inhibitor that blocks NF-кB migration to the nucleus. ACHP treatment shows no effect on the rapid internalization of Cldn5, implying a mechanism independent of NF-кB control of gene expression. These data are shown in a new figure (Figure 6) in the revised manuscript.

      (8) Figure 2 - CD206 are quite sparse however, Iba1 would work well to look at microglial activation.

      Thank you for this suggestion, which we have followed. To assess microglial activation, we have immunostained for Iba1 and quantified the data. These are now included in Figure 2 – figure supplement 3. The data show that there is an increase in Iba1 immunostaining following E. coli infection in both WT and Cdh5-CreER; Tlr4floxed mice, with more in the former than the latter, but the difference is not statistically significant.

      (9) Suggest performing the LAMP+ co-localization experiment at <1hr, prior to NFKB nuclear localization and transcriptional changes. This would better support it, this is (or is not) independent of the NFKB. Could also test this with an NFKB inhibitor, do they still see the CLDN5 internalization when NFKB is blocked?

      Thank you for these suggestions. We have done both of these analyses, and the results are presented in Figure 6. The results show that (1) Cldn5 is internalized within 1 hour and (2) its internalization is independent of NF-кB signaling inhibition by 1 uM ACHP. Since ACHP treatment shows no effect on the rapid internalization of Cldn5, that implies a mechanism independent of NF-кB control for gene expression.

      Reviewer #3 (Recommendations for the authors):

      Major points

      (1) The most important caveat is that the Cdh5-CreER model is known to recombine in leptomeningeal fibroblasts (10.1038/s41586-023-06993-7, 10.1101/2025.05.13.653681), and Cdh5 expression in these populations is now well described (10.1038/s41467-02341580-4, 10.1016/j.neuron.2023.09.002). Although the authors did not observe recombination in their reporter (details of the tamoxifen injection protocol should be provided), it is imperative to validate the specificity of their model to Tlr4 in endothelial cells, leveraging their sequencing data and providing additional IHC or ISH to confirm this. Alternatively, Tlr4 could be deleted in a more specific model, e.g., the Pdgfb-iCreERT2 or Slco1c1-CreERT2. It is also important to do the same with the LysM model, to confirm that the lack of impact of macrophage Tlr4 is not due to failure to delete the gene. This is again important to the interpretation of the study, since the authors propose that the endothelium, specifically, is the driver of the meningitis response.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclear-localised GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite right. The Results section of the revised manuscript has an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      We have also conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup>-directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text in the Results section now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. The phenotype of Cdh5-CreER; Tlr4floxed mice – a dramatically reduced infection-associated transcriptional response – argues that the floxed Tlr4 target was recombined at appreciable efficiency in those mice (Figure 1D and 1E). For Lyz2<sup>Cre</sup>; Tlr4floxed mice the principal phenotype is an up-regulation of infection-associated transcripts in a subset of dural border cells in the absence of infection; the transcriptional response to infection was largely unaffected in all leptomeningeal cell types (Figure 1D and 1E). We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4VEKO and Tlr4MKO mice.”

      (2) The authors did not examine the consequences of Tlr4 cKO on the course of meningitis or bacterial burden. Knowing the impact of this would strengthen the paper and allow us to determine if the endothelial responses are helpful or harmful in meningitis progression.

      For the revised manuscript, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) TLR4 is a known receptor for LPS. It is unsurprising (especially in the in vitro experiments) that Tlr4 knockout reduces NF-kB signalling and other downstream changes to endothelial cells. Furthermore, it is uncertain if the infection was left to continue, similar changes to the endothelium would nonetheless occur through other mediators such as IL1B and TNFa.

      We agree that it makes logical sense that Tlr4 KO decreases NF-кB signaling. The interesting next question is: what are the mechanistic underpinnings of the responses that are downstream of TLR4 and NF-кB? The cell culture experiments with WT vs. Tlr4KO bEnd.3 cells identify one set of cell biological responses related to Cldn5 and junctional integrity, and the NF-кB inhibition experiment (Figure 6) implies that rapid internalization of Cldn5 occurs in the absence of NF-кB mediated transcriptional changes. Regarding the possibility that other mediators such as IL1B or TNFα might, at least partially, make up for the lack of TLR4 signaling later in the infection, that is an open question at present.

      (3) The arachnoid fibroblast 2 cluster should be renamed to dural border cells based on their high expression of Slc4a10, Adamtsl3, Tmeff2, etc which are all highly enriched in dural border cells. I suspect this cluster is also highly enriched for Slc47a1, probably the most specific marker for these cells (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002).

      Thank you for this comment. The reviewer is correct. These are dural border cells and they express Slc47a1, as seen in a new supplemental Figure 1 – figure supplement 4, which shows UMAP plots for many leptomeningeal cell type-specific genes. We have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (4) It would be helpful to provide higher resolution images of Cldn5 in the leptomeningeal mounts. At the current resolution, it is difficult to tell if there is a similar internalisation/disruption phenotype to what is observed in vitro. Notably, this finding is similar to another recent publication on Cldn5 recycling (in the context of stroke) (10.1186/s40478-025-02125-6).

      Higher resolution images of Cldn5 in leptomeningeal vessels without or with E. coli infection are now shown in Figure 3 - figure supplement 1C. There is a visual impression of greater area occupied by Cldn5, which is confirmed by quantification (Figure 3A and B). This effect appears to be due to both an average increase in vessel diameter and a redistribution of some of the Cldn5 away from plasma membrane junctions. Thank you for pointing out the interesting and relevant Cottarelli et al (2025) paper, which we had not read. This is now referenced.

      Minor points

      (1) Typo: prominant should be spelled prominent.

      Thank you for catching that one. It is now corrected.

      (2) Strictly speaking, the arachnoid layer is not epithelial (despite Cdh1 expression). They are fibroblasts that acquire barrier-forming properties.

      Thank you for that comment. That appears to be the consensus view, and we will go along with it.

      (3) Notably, LyzM Cre will also recombine in other myeloid populations, so I wouldn't describe it as a macrophage.

      Thank you for this comment. We agree, and we have therefore changed the text and figure labels from “macrophage” to “myeloid”.

      (4) It is interesting and notable that ICAM1 expression is observed in nonendothelial populations, in the IHC, too, perhaps.

      We agree. ICAM1 may be a broader marker/mediator of inflammation than is generally recognized.

      (5) In F1B, your labels on the right image to the arachnoid barrier and pial surface are presumably meant to refer to the image on the left with DPP4 and laminin labelling? The subarachnoid should be between the laminin and DPP4 layers (although it will be collapsed in your preparations).

      Thank you for catching this error. The vertical bars were sized erroneously, and the labels were also placed erroneously. These have now been corrected.

      (6) I would reference the papers that defined leptomeningeal cell type markers (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002) when you define your cell types.

      Thank you. We have done that, and we have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (7) I would change references to the subarachnoid space in your figures to the leptomeninges (which include the SAS, but extend either side of it).

      Thank you. The labels have been changed to “leptomeninges”.

      (8) In Figure 2 - Supplement 1A, it looks like the populations are mislabelled.

      Thank you. This has been corrected to be consistent with the assignments in Figure 1B

    1. eLife Assessment

      This important study addresses how the molecular identity of a single neuron specifies its hard-wired synaptic connectivity, using repeated single-cell RNA sequencing of identified Drosophila sensory neurons together with functional perturbation of candidate cell-surface molecules. It demonstrates remarkably low transcriptomic variability across animals for the same identified neuron, defines a tractable set of differentially expressed cell-surface molecules that distinguish mechanosensory from chemosensory neurons, and links several of these molecules to axonal targeting and circuit function. The evidence is solid, with the single-neuron transcriptomic datasets and Dscam isoform repertoires offering a lasting resource for the field, though clearer articulation of the experimental logic, additional controls in the RNAi screen, and a more complete characterization of the neuronal re-wiring would further strengthen the central claims.

    2. Reviewer #1 (Public review):

      Summary:

      The authors sequence the transcriptome of three sensory neurons from D. melanogaster to study the cell-cell and animal-animal variability in these cells, with a focus on cell adhesion molecules. The work reports useful cell-specific transcriptomics datasets that will be of great interest to those studying cell types, transcriptomes, neuronal development, and cell surface proteomes. The authors also report large numbers of knockdown data (gene-by-gene or in combinations) and report neuronal wiring and behavioral phenotypes. The manuscript is highly descriptive of the system studied - in a good way, but often over-speculates in rationale or conclusions.

      Strengths:

      The manuscript is data-rich. The single-cell transcriptomics datasets, not trivial to collect, are a major strength of the work and will prove useful to the field. Also, the biased expression of Dscam is interesting, even though the authors cannot pursue the mechanism or a function for this.

      Weaknesses:

      The study lacks depth (i.e., mechanism) in explaining observations.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, dos Santos et al seek to identify cell-specific programs that drive neuronal wiring patterns. They focus on two chemosensory and mechanosensory neurons in the Drosophila nervous system, as they both display stereotyped connectivity in the ventral nerve cord. Single-neuron RNA sequencing identified cell surface molecules that distinguish the sensory neurons and may instruct their respective wiring patterns. They functionally test several of these candidates and observe miswiring phenotypes upon knockdown experiments. Additionally, they attempt to miswire the chemosensory neurons. Overall, this manuscript addresses an important question about how neurons identify appropriate synaptic partners through precise cell surface molecular codes. However, there are significant deficiencies in the experimental logic and rigor, and the manuscript can be very difficult to digest.

      Strengths:

      The use of two sensory neurons with stereotyped connectivity is a significant strength, as this enables the authors to identify genes that are required for wiring. Additionally, analyzing the transcriptomes of single neurons repeatedly could potentially be a robust approach to identifying cell-specific cell-surface molecules that drive wiring.

      Weaknesses:

      (1) The authors perform RNAseq for single identifiable neurons, as opposed to neuronal subclasses, which has been reported before. It would be beneficial to elaborate on the significance of using single neurons for answering the scientific question. This is briefly mentioned toward the end of one of the results subsections: "Repeated RNA sequencing of an identifiable neuron seeks to address the fundamental nature of variability in connectomics, axonal branching, and cellular identity." But this should be in the Introduction.

      (2) The authors chose the P14 pupal stage for one of the analyses. It is not clear why this specific stage is chosen. Does pSc and aPa connectivity occur at this stage?

      (3) This reviewer is confused as to why looking at differentially expressed CSMs between pupal and adult stages of two different neurons is useful. This does not seem like an appropriate comparison. This data might be better in the supplemental material, especially given the lack of precise age synchronization across pupal samples (as reported).

      (4) It is very difficult to follow the logic because the manuscript seems to jump around between different results and lacks a compelling through line.

      (5) "Single cell sequencing of the same neuron reveals transcriptome precision": What are the controls here? An aPa neuron is shown in Figure 3 as an example of a different neuronal subtype, but were other factors (e.g., lack of Repo expression) checked to ensure that samples were not contaminated?

      (6) "However, whether any of these exon 6 or 9 splicing specificities are biologically significant can only be determined using exon 6 and 9 isoform-specific RNAi." The authors could alternatively use CRISPR techniques to target specific isoforms that they hypothesize might be important for neural wiring, enabling them to assess isoform-specific wiring defects.

      (7) In the section "The set of cell surface receptors required to wire up the pSc mechanosensory neuron": Several previous subsections of the Results use RNAseq to identify molecules expressed in pSc neurons across different stages. It's unclear why the authors did not start with the identified list of candidate cell surface receptors identified in their RNAseq experiments.

      a. Were any of the genes screened the same as those identified by the authors as differentially expressed in pSc mechanosensory neurons, either across developmental stage (pupa vs. adult) or across neuronal subtype (pSc vs. Gr59d)? If so, it would be helpful to state this here. (They do mention later on that five CSMs identified were more highly expressed in pSc than aPa. However, changes in expression across developmental stages within the pSc neuron would still be helpful to comment on, especially since the authors identified greater transcriptomic differences across developmental stages than they did between different neuronal subtypes.)

      b. The 39 genes not expressed in pSc neurons served as their negative control, but the average axonal targeting grade was 2.3 (between moderate and severe). This calls into question the use of this method as an appropriate measure of whether a gene expressed by pSc neurons is truly required for proper axon targeting; there seems to be a strong probability of significant off-target effects. Performing a global knockdown and cell-specific rescue could potentially complement these experiments and serve as a stronger indicator of candidate receptors' roles in pSc-specific axon targeting.

      (8) It seems as though the purpose of the experiments described in the last results subsection ("Re-wiring the Gr59 chemosensory neuron") is to redirect the Gr59d neuron toward the pSc neuron's axonal targeting phenotype. However, the authors do not state whether they were able to do so effectively (i.e., whether or not there were significant differences between the rewired Gr59d neuron and the pSc neuron). This leaves the story unfinished.

      (9) At the end of the discussion, the authors state that "...if a Gr59d chemosensory neuron is functionally rewired to a pSc mechanosensory circuit, activation of the Gr59d neuron using a bitter tastant molecule should elicit a grooming (mechanosensory) response...". The authors should attempt this experiment, especially given that they have developed the PXGS technique.

    4. Author response:

      We are pleased that the reviewers found the repeated single-neuron sequencing and the finding of less than 1% transcriptomic variability to be original and striking, valued the single-neuron Dscam isoform repertoires and the scale of the functional screen, and judged the evidence solid to compelling. We provide below our provisional response and an outline of the revisions we plan.

      Overall plan: We intend to submit a revised version that addresses the public reviews and the recommendations to the authors. Because our conclusions rest on data already in the manuscript, the revisions are clarifications, added analysis of existing data, tempered language, and improved figures, rather than new experiments. Given the focused nature of these revisions, we would be happy for the editors to assess the revised version without re-involving the reviewers.

      One factual note for the Assessment and public reviews: The morphological RNAi screen comprised 213 cell-surface receptor genes; the figure of “140 genes” in one public review is the number that produced strong-to-severe phenotypes (Grade 3–5 at >40% penetrance), not the number screened. We will make this unambiguous in the revised text.

      Main changes in the revision:

      (1) We will explain that the RNAi screen was performed blind and independent of the RNA sequencing experiments. This was intentional, so that functional perturbation and transcriptomic identity would serve as independent lines of evidence, but could be compared with each other.

      (2) We will revise the Methods and Results to clarify how the morphological RNAi screen and behavioral subset should be interpreted, including conservative treatment of the negative-control distribution and mild-to-moderate phenotypes.

      (3) We will soften language that overstated certainty. Differentially expressed molecules are now described as prioritized candidates and convergent evidence, not definitive determinants.

      (4) We will reframe Gr59d/PXGS experiments as morphological rewiring and ectopic branching, and no longer imply a pSc-like conversion or functional rewiring.

      (5) We will add a limitations paragraph addressing RNAi off-target/background concerns, the absence of direct aPa functional testing, and the need for future mechanistic validation.

      (6) We will disclose or remove any figure panels that overlap with the companion PXGS manuscript and revise legends/labels to make it more clear.

      We hope these revisions make the logic of the study clearer and align the strength of the claims with the evidence.

      Below is our more detailed (provisional) response (not sure if this is required at this stage):

      Response to the eLife Assessment:

      Clearer articulation of the experimental logic. The Assessment’s central request (Reviewer 2) concerns the relationship between the transcriptomic experiments and the functional screen. The two were performed independently on purpose; the RNAi screen was assembled from a comprehensive survey of the literature rather than from the results of our differentially expressed genes from single-cell RNA sequencing. Thus, the RNAi screen was performed and graded blind in parallel with the single cell sequencing, with the gene identities unmasked only after both were complete. This was intentional, so that the sequencing (i.e., which molecules differ between neurons) and the screen (which molecules are functionally required) would provide mutually unbiased corroborating evidence rather than self-referential support/circular reasoning. We will state this more explicitly in the Introduction, in the Results where the screen is introduced, and in the Methods.

      Additional controls in the RNAi screen. We will treat the 39 genes that were identified in our single-cell RNA sequencing to be not expressed in the pSc neuron as a randomized negative control set in our RNAi experiments. We will state more explicitly the empirical RNAi false positive rate for a miswiring phenotype is 6/39 = 15%, likely due to RNAi off-target effects. We will also more clearly state that our claims about cell surface receptor functions are restricted to strong-to-severe phenotypes at high penetrance reproduced by at least two independent RNAi lines and corroborated independently (differential expression and/or single neuron qPCR).

      A more complete characterization of the re-wiring. We will state more clearly that mis-expressing the pSc-enriched cell surface receptors within Gr59d neurons partially shifts the arbour toward a pSc-like pattern (e.g., increased ectopic branching), and does not reproduce the full anatomical wiring, and that functional/behavioral re-wiring was not tested.

      Response to Reviewer 1:

      Reviewer 1 found the work valuable and data-rich, and the Dscam expression bias interesting. They noted over-confident language and asked how rigorously the differentially expressed genes were identified.

      Over-confident language. We will rewrite the two flagged sentences. The claim that the ~10 differentially expressed molecules are “likely the most important” will become a correlational statement, while also noting the lack of an aPa-specific Gal4 driver for direct testing. Our sentence that, “Our RNA sequencing data is biologically inadequate without a functional characterization of each molecule within the specific neuron” will be replaced with a clearer statement that gene expression data can nominate candidates, and functional perturbation of each gene is required to demonstrate necessity and sufficiency (i.e., biological function); which is exactly why we paired the RNA sequencing with an independent RNAi screen.

      Rigor of the differential-expression calls. We will more clearly state the statistical criteria in the Results (absolute log2 fold change ≥ 2 and Benjamini–Hochberg-adjusted p < 0.05). We will also note the small replicate numbers for the pooled pSc versus aPa comparisons, and emphasize that the central gene calls are independently supported by the blind RNAi screen and, for five genes, by single neuron qPCR. The full statistical workflow is in the Methods.

      Response to Reviewer 2:

      Reviewer 2 considered the findings potentially important but raised concerns about the experimental logic, the rigour of the screen, the completeness of the re-wiring, figure quality, and overlap with our PXGS companion paper. We will address each.

      Experimental logic. Beyond the design of our independent, blinded RNAi screen described above, we will add to the Introduction the rationale for sequencing single identified neurons (rather than subclasses) along with the two-pronged strategy, and add a summary paragraph at the start of the Discussion.

      Developmental stage choices. We will clarify our justification for the P14 pupal stage (the period when the mechanosensory neuron is actively elaborating its arbour while also enabling dissection). We will also clarify the rationale and caveats for comparing the pupal pSc neuron with the adult Gr59d neuron (i.e., the wiring occurs at the pupal stage, but the pupal Gr59d neurons could not be isolated at sufficient quality; the pSc pupal samples are less age-synchronized, so we simply used the comparison to identify the genes shared with the adult comparison).

      Transcriptome precision controls. We will state that the ten single pSc neurons passed the same quality controls for neuronal markers (elav, nSyb) and glial markers (Repo, moody < 20 reads) as all single-neuron libraries, which argues against any contamination by the attendant glial cell, and the aPa transcriptome is used as a different identity comparison.

      Off-target rate. As stated above, we will add the false positive rate for RNAi and restrict our confidence claims to those genes/cell surface receptors with multiple lines of evidence (e.g., strong phenotype, multiple RNAi lines, RNA sequencing, etc).

      Rewiring completeness and the behavioral prediction. As stated above, we will clarify that true re-wiring of the Gr59d neuron requires a future experiment, where a bitter tastant stimulus would elicit a grooming response.

      Response to Reviewer 3:

      We thank Reviewer 3 for judging our work to be fundamental in significance and the evidence compelling, with no major criticisms. Our clarifications above will further reinforce our hypothesis that the differential expression of specific cell surface receptors “do, in fact, control synaptic patterns,” which the reviewer highlighted.

      We are grateful for the reviewers’ time and for eLife’s model. We believe the planned revisions substantially clarify the experimental logic and tighten the claims, and we look forward to submitting the revised version.

    1. eLife Assessment

      This study addresses an important question in liver biology: how zonal hepatocytes balance survival and proliferation following injury? The authors propose that a mid-zone Atf4-Chop axis to Btg2 program temporarily suppresses proliferation to promote survival after a variety of chemical and surgical liver injury models. The authors provide evidence that some zones mount tailored stress responses, which ultimately promote regeneration and liver healing; however, the "mid-zone" changes with different injury models, making it difficult to conclude that the ATF4-CHOP response is specific to this zone in all injury contexts. In addition, it is possible that Atf4 and Btg2 overexpression could lead to Cyp2e1 suppression, which could reduce the extent of injury after CCl4 or APAP. To some extent, these points make the strength of the evidence incomplete, but do not entirely detract from the significance of the study, which is underscored by the helpful observation that there are zone-specific stress responses that mediate liver regeneration and survival.

    2. Reviewer #1 (Public review):

      Summary:

      The authors present evidence that during acetaminophen (APAP)-induced liver injury, mid-zone hepatocytes activate an integrated stress response (ISR) program via Atf4 and Chop, leading to induction of Btg2. This program suppresses proliferation in the early phase of injury, prioritizing hepatocyte survival before regeneration begins. The study uses spatial transcriptomics, immunohistochemistry, CUT&RUN, and AAV overexpression to support this model.

      Strengths:

      (1) Innovative use of spatial transcriptomics to capture zonal differences in hepatocyte stress responses.

      (2) Identification of a mid-zone specific ISR signature and candidate downstream regulator Btg2

      (3) Functional experiments with Atf4-Chop-Btg2 modulation provide causal evidence linking ISR activation to proliferation inhibition.

      (4) Conceptually significant model that hepatocytes actively balance survival and regeneration dynamically in a zone-specific manner.

      (5) Multiple models validation of the finding

      (6) The functional link of such zone2 phenotype is added.

    3. Reviewer #2 (Public review):

      The manuscript reports protection of midlobular hepatocytes from APAP toxicity by activation of Atf4-CHOP (Ddit3)-mediated cell cycle arrest and stress response. The authors acknowledge that their finding is unexpected because CHOP typically induces cell death. Therefore, they functionally validate several aspects of the proposed Atf4-CHOP mechanism. Along these lines, the mitigation of APAP toxicity by AAV expression of Atf4 or Btg2, the latter identified as CHOP effector, is impressive. Whether Atf4 indeed acts through CHOP and whether midlobular hepatocytes are protected because of cell cycle arrest is less clear. These and other criticisms are described in the following.

      Major points:

      (1) Starting with the basics, one wonders why midlobular hepatocytes manage to mount a defensive response to APAP, but PC hepatocytes don't. Is this because midlobular hepatocytes express the relevant Cyps (2e1 but also 1a2 and 3a11) at lower levels, which mitigates toxicity and buys them time? This would be supported by F2A but not by F3B, at least not for the most important Cyp2e1. A moderate difference is shown for Cyp1a2 expression in F3D but is that enough to explain the different fates? Or are additional post-transcriptional effects on these Cyps at work? The difference in baseline Cyp2e1 expression between F2A and F3B remains unexplained after revision.

      (2) The evidence presented in support of cell cycle arrest of midlobular hepatocytes is not fully convincing: there is no overt difference in S and G2/M gene scores in F2F; the marker genes used for S phase and G1 to S progression in F2G are unusual. Along these lines, one wonders if spatial transcriptomics confirmed the Ki67 immunostaining results in F1 also for specific zones, not only overall as shown in F2E? In contrast to the revised discussion, the abstract does not reflect that limited evidence for a cell cycle arrest in pericentral hepatocytes was found.

      (3) The authors conclude in line 364 that halting of proliferation by Btg2 favors survival, which raises the question of whether Btg2 knockout causes death in midlobular hepatocytes in F6K. Data addressing this question, that is, localization and extent of tissue necrosis and ALT levels after APAP, are missing. The efficiency of knockout of Btg2 is also not given. Additional Btg2 knockout data support its proposed role in the revised manuscript.

      (4) Related to the previous question, the BTG2 immunostaining in F6F is not convincing when compared to F6D. One also wonders if it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3? The BTG2 immunostaining remains weak, not only in in F6F but now also in F6D of the revised manuscript, which together with lack of high-resolution immunostaining of AAV-Ddit3-induced BTG2 in the absence of APAP results in limited support for the conclusion that APAP promotes nuclear localization of BTG2.

      (5) Related to the previous question, the proposed Atf4-Ddit3 axis is challenged by the lack of midlobular induction of Atf4 in the APAP scRNA-seq data published by another group presented in S4F and G. Further analysis of AAV-Atf4 samples generated for F5 could address if it is really Atf4 that acts on Ddit3 in APAP toxicity. The extended list of transcription factors (from 30 to 50) includes Atf4 but direct evidence for an interaction with Ddit3 is missing from the revised manuscript.

      (6) Related to the previous question, the ATF4 immunostaining in F5A doesn't look convincing, with many brown pigments appearing to be outside of the nucleus. The ATF4 immunostaining after APAP challenge remains weak.

      (7) It is not ruled out that AAV expression of Atf4 or Btg2 reduces hepatocyte sensitivity to APAP by affecting expression of the Cyps needed for activation. In other words, does AAV-Atf4 or AAV-Btg2 change the expression of any of the Cyps relevant to APAP in the 3 weeks before APAP application (F5B)? S5A of the revised manuscript rules out loss of Cyp2e1 expression as a confounding factor.

      (8) It is laudable that the authors tried to extend their findings to human by using snRNA-seq data from a published study (line 391) but it is unclear why they didn't analyze all 10 patients in that study but instead focused on 2 and stated that this small sample number prevented drawing definitive conclusions and could therefore only be mentioned in the discussion. The revised manuscript continues to focus on rare spatial transcriptomics analyses of patients with APAP toxicity although more snRNA-seq analyses of such patients are available which should also allow for analysis of hepatocyte zonation.

      Comments on revised version.

      After revision, the proposed role of Btg2 is substantiated but it remains unclear why midlobular hepatocytes don't proliferate after APAP challenge and whether the observed protective effects are indeed mediated by Atf4 acting directly through CHOP.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Zonation definition under injury has been shown to be sustained broadly, but is not sufficiently validated and quantified, especially considering the resolution of the 10x Visium system and the potential variation of outcomes based on how to define zones.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g.,Alb, Mup20, Cyp2f2,Pck1,Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2,Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g.,Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) The model is built entirely in APAP injury, which specifically targets pericentral hepatocytes. It remains unclear whether the proposed mechanism applies to other liver injuries (e.g., partial hepatectomy, CCl4).

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Baseline proliferation appears higher than expected in homeostasis (Figure 1B), and fold change analysis (not absolute counts) may be needed to assess zonal proliferation suppression (Figure 1D).

      We thank the reviewer for this insightful comment. The baseline proliferation rates observed in our study are consistent with previously reported zonal distributions (PMID: 33632817; PMID: 33632818), with approximately 70% of proliferating hepatocytes located in zone 2, 20% in zone 3, and 10% in zone 1 under homeostatic conditions. To further address the reviewer’s concern, we performed a fold-change analysis of Ki-67<sup>+</sup>hepatocytes across different zones. This analysis revealed that only the mid (zone 2) and pericentral regions exhibited significant changes, whereas no statistically significant differences were observed in the other zones (as shown in Author response image 1). Importantly, when considered together with the absolute cell counts, these results indicate that the apparent suppression of proliferation is most pronounced in the mid zone, likely due to its relatively higher baseline proliferation under homeostatic conditions. In contrast, this effect is less evident in the fold-change analysis, as zones with low baseline proliferation show limited dynamic range for detecting relative changes.

      Author response image 1.

      Fold changes of Ki67-positive cells across liver zones (PC, Mid, PP) at 0, 3, 6, 12 and 24 h post-APAP. Fold change was the number of Ki67-positive cells in the three regions at each time point after APAP treatment divided by the number of positive cells in each region at 0 hour post-APAP. (a) denotes significance between PC and Mid regions, (b) denotes significance between PC and PP regions, and (c) denotes significance between Mid and PP regions.

      (4) AAV-based overexpression raises potential confounds (altered CYP activity before injury) and shows incomplete penetrance that is not quantified (Figure 5 - Figure 6).

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (5) The functional link between proliferation suppression and improved survival is inferred, but direct survival /injury readouts are limited.

      We thank the reviewer for this insightful comment. To more directly evaluate the functional link between proliferation control and liver injury, we manipulated Btg2, a downstream effector of the Atf4–Chop axis and a known inhibitor of cell proliferation. Knockdown of Btg2 using AAV8–CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial knockdown efficiency, we observed a clear exacerbation of liver injury, as evidenced by an approximately 2-fold increase in serum ALT levels and a ~1.5-fold expansion of necrotic areas. In parallel, hepatocyte proliferation was significantly increased (~1.8-fold increase in Ki67⁺ hepatocytes) compared to control mice (Revised Figure 6K–N). Conversely, Btg2 overexpression produced the opposite phenotype, markedly attenuating liver injury while suppressing hepatocyte proliferation (Revised Figure 6G–J). Together, these gain- and loss-of-function data provide direct evidence linking proliferation control to injury severity, thereby supporting a causal relationship between suppressed proliferation and improved liver outcomes (Revised manuscript, page 14, lines 399–406).

      Reviewer #2 (Public Review):

      (1) Starting with the basics, one wonders why midlobular hepatocytes manage to mount a defensive response to APAP but pericentral hepatocytes don't. Is this because midlobular hepatocytes express the relevant Cyps (2e1, but also 1a2 and 3a11) at lower levels, which mitigates toxicity and buys them time? This would be supported by F2A but not by F3B, at least not for the most important Cyp2e1. A moderate difference is shown for Cyp1a2 expression in F3D, but is that enough to explain the different fates? Or are additional post-transcriptional effects on these Cyps at work?

      We thank the reviewer for this important question. We fully agree that the differential susceptibility between mid‑zone and pericentral (PC) hepatocytes is likely rooted in the zonal gradient of cytochrome P450 expression. Our spatial transcriptomics data (Revised Figure 2A) show that mid‑zone hepatocytes express Cyp2e1, Cyp1a2, and Cyp3a11 at levels intermediate between PC and periportal (PP) zones. This intermediate expression may generate sufficient NAPQI to activate stress signaling but not so much as to cause immediate mitochondrial collapse, thus “buying time” for adaptive responses. We also appreciate the reviewer’s observation that Cyp2e1 mRNA levels remain highest in the PC zone even after APAP (Revised Figure 3B). However, mRNA abundance does not necessarily reflect functional protein level. In the PC zone, massive necrosis rapidly compromises cellular integrity; as shown in Revised Figure 3D, Cyp1a2 protein declines sharply around the central vein, and we observed similar degradation for Cyp2e1 (data not shown). Consequently, despite sustained Cyp2e1 transcripts, PC hepatocytes are unable to mount an effective stress response because they are already undergoing cell death. By contrast, mid‑zone hepatocytes retain sufficient metabolic capacity to activate the Atf4‑Chop axis while preserving cellular function.

      (2) The evidence presented in support of cell cycle arrest of midlobular hepatocytes is not fully convincing: there is no overt difference in S and G2/M gene scores in F2F; the marker genes used for S phase and G1 to S progression in F2G are unusual. Along these lines, one wonders if spatial transcriptomics confirmed the Ki67 immunostaining results in F1 also for specific zones, not only overall, as shown in F2E?

      We thank the reviewer for these important observations. We agree that the current spatial transcriptomics (ST) data alone do not provide sufficiently strong support for this conclusion. The limited sensitivity of ST for detecting rare proliferative events further constrains its utility in this context. At baseline, only ~1% of ST spots are Ki67-positive (Revised Figure S1I), and this fraction becomes even lower during the early phase following APAP injury. As a result, there are insufficient Ki67+ spots to robustly assess zonal distribution using ST, which precludes a reliable spatial validation of proliferation patterns at this resolution. For this reason, our primary evidence for zonal proliferation dynamics relies on Ki67 immunohistochemistry (Revised Figure 1), which provides single-cell resolution and higher sensitivity. These data show a marked reduction in Ki67+ hepatocytes specifically in the midlobular zone at 3-6 hours post-APAP, supporting a transient suppression of proliferation in this region. In addition, we agree that the transcriptional evidence for cell cycle arrest was not strong the S and G2/M scores showed no overt difference, and the gene sets used were suboptimal. We have therefore moved these analyses to the supplement and toned down the claims. We have also clarified this limitation in the manuscript (Revised manuscript, page 18, line 518-524)

      (3) The authors conclude in line 364 that halting of proliferation by Btg2 favors survival, which raises the question of whether Btg2 knockout causes death in midlobular hepatocytes in F6K. Data addressing this question, that is, the localization and extent of tissue necrosis and ALT levels after APAP, are missing. The efficiency of the knockout of Btg2 is also not given.

      We thank the reviewer for this insightful comment. We have included the missing data. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup>hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N, Revised manuscript, page 14, line 399-406).

      (4) Related to the previous question, the BTG2 immunostaining in F6F is not convincing when compared to F6D. One also wonders if it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3?

      We thank the reviewer for this insightful comment. We have included an inset of the original image to better show BTG2 staining in revised Figure 6F. During our study, we tested BTG2 expression in mice transduced with AAV‑TBG‑EGFP or AAV‑TBG‑BTG2 for three weeks without APAP challenge. We observed that BTG2 in these non‑injured livers was predominantly cytoplasmic (Author response image 2), contrasting with the nuclear localization seen after APAP treatment (Figure 6F). Regarding whether it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3, we think Ddit3 promotes BTG2 expression (as shown in revised Figure F6F), but APAP is necessary for its nuclear translocation.

      Author response image 2.

      Immunohistochemical detection of Btg2 in liver tissue from mice transduced with AAV-TBG-EGFP or AAV-TBG-Btg2 for 3 weeks without APAP treatment.

      (5) Related to the previous question, the proposed Atf4-Ddit3 axis is challenged by the lack of midlobular induction of Atf4 in the APAP scRNA-seq data published by another group, presented in S4F and G. Further analysis of AAV-Atf4 samples generated for F5 could address whether it is really Atf4 that acts on Ddit3 in APAP toxicity.

      We thank the reviewer for this insightful comment. We agree that Atf4 was not among the top 30 active transcription factors in our initial analysis; however, when we extended the list to the top 50, Atf4 was included. We have therefore updated Revised Figures S4F and G to show the top 50 transcription factors. We also appreciate the reviewer’s suggestion to further investigate whether Atf4 directly acts on Ddit3 in the context of APAP toxicity. While this still shows a less pronounced midlobular enrichment for Atf4 compared with Ddit3, we sought additional evidence for a functional Atf4-Ddit3 link. In primary hepatocytes treated with APAP, we observed nuclear co‑localization of Atf4 and Ddit3 (Author response image 3A) and increased nuclear protein levels of both factors (Author response image 3B), supporting their potential cooperative role. We agree that direct analysis of AAV‑Atf4 samples generated for Figure 5 would provide more definitive evidence; unfortunately, co‑staining for Atf4 and Ddit3 on those tissue sections didn’t work well.

      Author response image 3.

      Subcellular localization of Atf4 and Chop in primary hepatocytes following APAP treatment. (A) Immunofluorescence staining of Atf4 and Chop in primary hepatocytes treated with 10 mM APAP for 6 hours or left untreated (UT). Nuclei were counterstained with DAPI. Scale bar as indicated. (B) Primary hepatocytes were treated with 0, 5, or 10 mM APAP for 6 hours. Cytoplasmic and nuclear fractions were isolated and analyzed by western blot. Lamin B1 and α-Tubulin were used as markers for the nucleus and cytoplasm, respectively

      (6) Related to the previous question, the ATF4 immunostaining in F5A doesn't look convincing, with many brown pigments appearing to be outside of the nucleus.

      We thank the reviewer for this helpful comment. To better demonstrate ATF4 nuclear localization, we have added enlarged insets of the original representative images in revised Figure 5A. These magnified views more clearly show nuclear ATF4 staining after APAP treatment, addressing the concern about extranuclear signal.

      (7) It is not ruled out that AAV expression of Atf4 or Btg2 reduces hepatocyte sensitivity to APAP by affecting the expression of the Cyps needed for activation. In other words, does AAV-Atf4 or AAV-Btg2 change the expression of any of the Cyps relevant to APAP in the 3 weeks before APAP application (F5B)?

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      (8) It is laudable that the authors tried to extend their findings to humans by using snRNA-seq data from a published study (line 391), but it is unclear why they didn't analyze all 10 patients in that study but instead focused on 2 and stated that this small sample number prevented drawing definitive conclusions and could therefore only be mentioned in the discussion.

      We thank the reviewer for this clarification. The analysis mentioned in line 391 originally referred to spatial transcriptomics (ST) data from two ALF patients, not snRNA-seq. For the snRNA-seq dataset, we analyzed all 10 patients, but snRNA-seq lacks spatial resolution and cannot reliably assign zonal identity. We stipulate that snRNA-seq requires viable cells and thus likely excludes necrotic/peri-necrotic areas. Therefore, direct zonal comparison with our ST data was not possible. We have now clarified this in the revised manuscript (Revised manuscript, page 18, line 510-519).

      Reviewer #3 (Public Review):

      The main concern is that the overexpression of ATF4 and DDIT3 is causing reduced cell death and damage by APAP. This makes it harder to understand if these genes are truly increasing survival or if they are just reducing the injury caused by APAP. It may be better to perform overexpression immediately after, or at the same time as APAP delivery. Alternatively, loss-of-function experiments using AAV-shRNAs against these targets could be useful.

      We thank the reviewer for raising this important point. We agree that overexpression prior to APAP administration leaves open the question of whether the observed protection reflects true cytoprotection or simply reduced initiation of injury. To address this, we pursued loss‑of‑function approaches. Due to their very low basal expression, AAV‑shRNA‑mediated knockdown of endogenous Atf4 and Ddit3 proved inefficient. We therefore targeted Btg2, a downstream mediator of Ddit3 that inhibits proliferation. Knockdown of Btg2 resulted in a significant increase in APAP‑induced liver injury, as evidenced by elevated ALT levels and expanded necrotic areas (Revised Figure 6K-N). These results indicate that the ATF4‑DDIT3‑BTG2 axis limits hepatocellular damage, consistent with a protective role. We have clarified this point in the revised manuscript (page 15, line 407-414)

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Clarify how zones were defined when necrosis disrupted pericentral areas. Provide marker validation across time and whether necrotic spots are excluded or not from zonal analysis.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g., Alb, Mup20, Cyp2f2, Pck1, Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2, Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g., Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) Test whether the ISR-Btg2 program applies in other models; even targeted validation via qPCR and IF would be valuable.

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Proliferation quantification in liver sections in Figure 1: how to define the zones and why, at the basal level, there is a high proliferation rate in the mid zone? From Figure 1B-C, all three zones showed decreased hepatocyte proliferation, although the mid zone had a higher baseline. Will the mid-zone stand out by converting to the fold change of Ki-67+ hepatocytes decrease?

      We thank the reviewer for these insightful comments. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x<sup>2</sup> + z<sup>2</sup> – y<sup>2</sup>) / (2z<sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885). Consistent with previous reports (PMID: 33632817; PMID: 33632818), we observed a higher baseline proliferation rate in the mid zone, where approximately 70% of proliferating hepatocytes reside under basal conditions, compared to 10% in zone 1 and 20% in zone 3. However, when analyzing the fold change in Ki-67+ hepatocytes, only Mid and PC region showed significant difference in Ki-67+ hepatocytes, other zones showed no significant differences (as shown in the fold-change results in Author response image 1), indicating that the mid zone does not stand out in the fold change analysis. See Author response image 1.

      (4) The authors need to strengthen the causal chain with rescue experiments, e.g., Atf4/Chop overexpression and Btg2 knockdown. Link proliferation suppression to survival/ALT directly.

      We thank the reviewer for these constructive comments. Besides existing data from Figure 5 (Atf4 overexpression), we included Btg2 knockdown data in the revised Figure. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup> hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N) (Revised manuscript, page 14, lines 399–406).

      (5) Transduction efficiency, distribution, and expression levels via the AAV overexpression need to be quantified. Key CYP genes in the APAP metabolic pathway need to be assessed to exclude confounds.

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (6) The authors claim that the requirement of the Atf4/Chop at the early stage of APAP injury protects hepatocytes from proliferation for survival. What is the consequence if we remove the protective mechanism?

      We thank the reviewer for this insightful question. In our model, early induction of Atf4 and Chop functions as a cell survival checkpoint. Removal of this protective mechanism is predicted to result in two deleterious outcomes: (1) Acute exacerbation of necrosis due to the inability of hepatocytes to manage stress-induced bioenergetic demands, and (2) Impaired long-term regeneration due to depletion of the surviving cell pool. We directly tested the acute prediction (< 24 h) in Author response image 4. We deleted Ddit3 specifically in hepatocytes. Initial attempts using AAV-CasRx failed due to negligible baseline Atf4/Chop expression in healthy liver, preventing effective knockdown. We therefore generated hepatocyte-specific Ddit3 knockout mice (Alb<sup>∆Ddit3</sup>; Author response image 4B). Immunohistochemistry confirmed APAP-induced Chop induction occurs primarily in the centrilobular zone by 6 h (Author response image 4A). Following a two-dose APAP regimen (Author response image 4C), Alb<sup>∆Ddit3</sup> mice displayed significantly larger areas of centrilobular necrosis compared to Ddit3<sup>fl/fl</sup> controls (Author response image 4D; **p < 0.01). Thus, hepatocyte-intrinsic Chop limits acute APAP injury, consistent with its proposed early protective role.

      Author response image 4.

      Hepatocyte-specific deletion of Ddit3 exacerbates APAP-induced liver injury. (A) Immunohistochemical staining of Chop in liver sections at 0,3 and 6 h post-APAP. Red arrows indicate Chop-positive hepatocytes. Scale bar = 50μm. Quantification of zonal distribution of Chop-positive cells in liver sections at 6 h post-APAP is conducted . The statistic is the percentage of Chop-positive hepatocytes in each layer over the total number of Chop-positive hepatocytes. n=3 mice. (B)The construction, genotyping strategy and genotyping results of Alb<sup>∆Ddit3</sup> mice. P: positive control; WT: Wild-type; Neg: Blank control(ddH<sub>2</sub>O). (C) Schematic figure illustrating the experimental strategy for the administration of two doses of APAP to Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice. (D) H&E staining showing liver morphology from Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice at 6 h post-second dose of APAP. Injured area is outlined by black dashed lines. Scale bars = 200 μm. The percentage of injury area is quantified. n = 3- 4 mice/group. Data are represented as means ± SD; *p < 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001; ns, not significant.

      (7) Is there any human relevance to the sensitivity of APAP injury regarding the Atf4/Chop axis?

      We thank the reviewer for this insightful comment. During our study, we analyzed a spatial transcriptomics dataset from APAP patients. In one of two analyzed patients, mid-zone hepatocytes exhibited transcriptional signatures remarkably consistent with our murine findings, including: (1) upregulation of Atf4-Chop pathways, and (2) downregulation of cell proliferation genes (Author response image 5). This suggests that this axis may also be involved in the response to APAP injury in humans. However, given the limited sample size, definitive conclusions cannot be drawn at this stage. We have now included this point in the Discussion section (Revised manuscript, page 18, line 510-519).

      Author response image 5.

      Spatial transcriptomics (GSE223561) reveals zonal gene expression changes in APAP patients. Heatmap of ISR, cell death, and cell cycle gene expression across zonal regions in healthy versus APAP‑treated human livers. 

      (8) Several IHC stainings have a weak signal and need inserts to zoom in for a clear view of the positive signals. Figure 5A, E, G, and Figure 6D, F.

      We thank the reviewer for this observation. We agree that the immunostaining signals for several target genes are relatively weak, which reflects their low endogenous expression levels. To address this, we have included higher-magnification insets in the indicated panels (Revised Figure 5A, E, G and Figure 6D, F) to show the positive signals.

      Reviewer #2 (Recommendations for the authors):

      (1) What is the functional classification of DEG in F2A based on? GO terms?

      We thank the reviewer for this constructive question. The functional classification of differentially expressed genes (DEGs) in F2A is based on Gene Ontology (GO) terms. For each DEG, we retrieved its associated GO annotations across the three main categories (biological process, cellular component, molecular function). In cases where a gene was assigned multiple GO terms, we prioritized the most representative or significantly enriched term for functional interpretation. This clarification has been incorporated into the revised figure legend and the according GO number has been included in the figure.

      (3) The rationale for focusing on CHOP is not clear because Ddit3 is not shown in the spatial transcriptomics in F2A and is not significant in F2B, contradicting what is stated in line 206.

      We thank the reviewer for raising this important point. We apologize that Ddit3 was missing from the original figure. In the revised manuscript, we have included an updated version of Figure 2A, which now shows that Ddit3 is indeed one of the differentially expressed genes (DEGs) in the Mid zone at both 3 and 6 hours post-APAP. We agree with the reviewer that, as shown in Figure S1G (previous Figure 2B), Ddit3 did not reach statistical significance, due to its relatively low expression level in that analysis. Nevertheless, when we examined transcription factor (TF) activity in the Mid zone during early AILI, Ddit3 and Atf3 ranked as the top two most highly expressed TFs among the top ten with the highest activity, whereas Atf4 ranked seventh (Revised Figure 4B and Figure S3B). Given that Ddit3 frequently co-worked with Atf4 and that the Atf4–Ddit3 axis plays a well-established role in cellular stress adaptation, we considered this pathway to be biologically relevant and worthy of further investigation.

      (3) The term "redistribution" used in line 197 to describe the expression of Cyp2e1 and other Cyps in the midlobular zone seems inappropriate, considering that they just continue to be expressed there, whereas pericentral hepatocytes are dying in F3B; the same applies to "Gene Expression Shift" in F3H.

      We thank the reviewer for this important clarification. We have revised the text (Revised manuscript, page 9, line 234-236) to state that selective loss of Cyp‑expressing pericentral hepatocytes leads to the mid‑zone becoming the primary site of residual Cyp activity. The figure label has been changed from “Gene Expression Shift” to “Peri‑necrotic Cyp retention” and the legend now explicitly notes that this is an apparent zonal shift due to necrosis, not active redistribution.

      Reviewer #3 (Recommendations for the authors):

      (1) Please do not use abbreviations like AILI. This makes the paper more difficult to read.

      We thank the reviewer for pointing this out. We have replaced AILI with the full term “APAP-induced liver injury” to ensure easiness for readers.

      (2) It will be important to clarify how pericentral, mid, and periportal were defined. In Figure 1, it appears that some of the pericentral hepatocytes that are Ki67 positive are quite mid-zonal. It would be important to have rigorous definitions for the location determination.

      We thank the reviewer for this constructive comment. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x <sup>2</sup> + z <sup>2</sup> – y <sup>2</sup>) / (2z <sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885).

      We thank the reviewers for their rigorous critique again. We thank eLife for fostering an environment of fairness and transparency that enables authors to communicate openly and present their data honestly.

    1. eLife Assessment

      This study investigates the role of Interleukin-2-inducible T cell kinase (ITK) deficiency in autoimmune lung injury using a pristane-induced pulmonary hemorrhage (PH) model, suggesting that ITK-deficient regulatory T cells (Tregs) restrict severe tissue pathology. The work represents a valuable addition to the fields of autoimmunity, inflammation, and T-cell biology in the lung. However, the experimental evidence supporting the underlying cellular and molecular mechanisms and the integration of foundational background literature to provide the necessary context are incomplete.

    2. Reviewer #1 (Public review):

      In this study, Hossain et al. investigated the role of Interleukin-2-inducible T cell kinase (ITK) in autoimmune lung injury, demonstrating that ITK-deficient (Itk-/-) mice are protected against pristane-induced pulmonary hemorrhage (PH). The authors suggest that this protection correlates with a significant remodeling of the T cell compartment in Itk-/- mice, including increased frequency of memory-like CD4+ and CD8+ T cells (CD44⁺CD62L⁺) as well as higher frequency of Treg populations. Furthermore, adoptive transfer of ITK-deficient Treg isolated from injured ITK-deficient mice confers protection against pulmonary hemorrhage in WT recipients.

      Strengths:

      The adoptive transfer of wild-type and Itk-/- Treg populations demonstrates that ITK-deficient Treg can actively rescue pre-existing lung injury and reverse systemic secondary metrics like proteinuria in wild-type recipients, providing proof-of-concept validation for the therapeutic utility of the ITK-Treg axis.

      Weaknesses:

      A primary limitation of this manuscript is its omission of foundational literature from the Schwartzberg and Littman laboratories, which originally established the indispensable role of IL-2-inducible T-cell kinase (ITK) in proximal T-cell receptor (TCR) signaling dynamics and thymic lineage commitment. Because classic studies demonstrate that ITK is a critical regulator of thymic T cell development and cellular proliferation (PMID: 8777721, 10213685), the authors' claim that "these findings indicate that ITK deficiency skews the T cell compartment toward a memory-like state, establishing a distinct immune baseline that may favor protective and regulatory responses over pathogenic inflammation" is not substantiated by evidence and requires more robust validation.

      The exclusive reliance on splenic immunophenotyping is a major limitation, as it fails to capture the local cellular dynamics within the primary organs of injury (the lung and kidney). Evaluating canonical and non-canonical Treg expansion solely in the spleen overlooks the distinct functional programming of tissue-resident subsets. The authors should extend their characterization of regulatory T cell compartments directly to the lungs and draining lymphoid structures.

      More importantly, the authors overlook key historical publications that explicitly established ITK as a negative "rheostat" or gatekeeper for regulatory T cell (Treg) differentiation. Specifically, Huang et al. (PMID: 25063868) previously demonstrated that Treg abundance is inversely correlated with ITK expression, and that ITK activity serves as a vital negative tuner of IL-2-driven Foxp3⁺ Treg expansion. Since it is already well-established that suppressing or deleting ITK promotes Treg accumulation and function, and that these cells are intrinsically vital to suppressing systemic autoimmunity, it is unclear how these findings expand upon our existing mechanistic understanding of ITK regulatory biology.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Hossaim and colleagues investigate the role of the ITK kinase in modulating inflammation in a pristane-induced lung hemorrhage model. Using a germline ITK KO mouse, they report that loss of ITK skews the T cell compartment toward a memory-like state, expanding Tregs, and conferring protection against alveolar hemorrhage, inflammatory monocyte recruitment, proteinuria, and systemic cytokine elevation. They further show that transfer of ITK-deficient Tregs into wild-type hosts with established disease attenuates injury and shifts the cytokine balance toward resolution, and that ITK-deficient Tregs carry a transcriptional signature enriched for OXPHOS, mTORC1, MYC, and cell-cycle programs. While these observations are interesting for the development of potential immunotherapies, there are several issues with the methodological approach that support the authors' claims, tempering my enthusiasm for this manuscript.

      Strengths:

      (1) The clinical motivation and potential targeted therapies are relevant.

      (2) The murine phenotype seems robust.

      Weaknesses:

      (1) All loss-of-function experiments are from a global ITK knockout. This is a major limitation and weakness of this study. The protection observed in the intact knockout, therefore, cannot be attributed to Tregs specifically. The Treg-intrinsic claim rests almost entirely on a single adoptive-transfer experiment. In order to show that this effect is Treg-specific, the authors would need to generate a Treg-specific ITK-deficient mouse

      (2) In their sufficiency experiment (adoptive Treg cell transfer), donor and/or host cells are not congenically marked, so persistence, lung trafficking, and in vivo expansion of transferred Tregs are not demonstrated.

      (3) The authors claim that ITK-deficient Tregs possess enhanced metabolic fitness. This conclusion is based on transcriptional profiling of isolated splenic Tregs from unchallenged mice, yet it concerns lung protection during active disease. A disease-state and ideally lung-relevant transcriptome would more directly support the mechanistic narrative. Additional functional validation would be needed (Seahorse assay, mitochondrial mass/potential, etc). Some of these GSEA programs enriched in ITK-deficient Tregs could reflect a more general proliferative signature.

    4. Reviewer #3 (Public review):

      Summary:

      Hossain et al. investigate the role of ITK as a central regulator of autoimmune lung injury. They used ITK-deficient mice and the pristane-induced pulmonary hemorrhage (PH) model to show that ITK deficiency confers protection against PH. The adoptive cell transfer experiment suggests a possible role for altered Treg cells in ITK-deficient mice in regulating the inflammatory response in the lungs of pristane-injected mice. This study shows that targeting the ITK axis may be beneficial by reducing systemic inflammatory injury that contributes to poor outcomes in PH.

      Strengths:

      This study highlights the importance of ITK in regulating pulmonary hemorrhage. The enrichment of Treg cells is known to confer protection in autoimmunity-mediated alveolar damage. However, ITK's involvement in regulating Treg cell function is interesting and could be explored as a novel therapeutic approach for chronic inflammation.

      Weaknesses:

      The novelty of this study lies in the association between ITK-deficient Tregs and pulmonary hemorrhage in autoimmunity. The weakness of the manuscript is the lack of sufficient experiments to support the claim that ITK-deficient mice show protection specifically mediated by Treg cells, and to demonstrate that ITK-deficient Treg cells are more efficient than WT Treg cells in regulating other immune cells that drive pulmonary damage. The authors performed all the experiments in ITK global knockout mice, in which not only T cells but all other cell types are deficient in ITK. Furthermore, they have not performed any functional analysis to demonstrate the functional differences between WT Treg and ITK-deficient Treg cells, undermining the novelty of this study.

    1. eLife assessment

      The authors present a valuable open-source tool for three-dimensional analysis of dissected slices of human brains including 3D reconstruction and high-resolution 3D segmentation. Convincing evidence is provided based on experiments on both real and synthetic data. This tool would be of use to researchers in the neuropathology and neuroimaging field.

    1. eLife assessment

      The study presents a tool for searching molecular dynamics simulation data, making such data sets accessible for open science. The authors provide convincing evidence that it is possible to identify noteworthy molecular dynamics simulation data sets and their analysis can produce valuable information.

    1. eLife Assessment

      This valuable manuscript by Alonso-Caraballo et al is a novel piece of work that examines the impact of oxycodone self-administration on neural plasticity within paraventricular thalamic (PVT) to nucleus accumbens shell (Shell) pathway - two regions shown to play a key role in cue-induced drug seeking on their own - and whether this plasticity varies based on abstinence period and biological sex. Data show that a clinically relevant long-access model of self-administration promotes dependence in both male and female rats and provide compelling data that when compared to current literature indicate that craving-induced relapse for opioids may develop faster and may be more pronounced in females compared to males. In addition to these behavioral findings, the authors provide the first evidence that glutamate signaling within the PVT-to-Shell pathway is selectively strengthened at the output medium spiny neurons by opioids following protracted, but not acute abstinence. These data highlight a potential role for these adaptations in relapse behavior and identify a potential therapeutic target during abstinence to reduce relapse risk in abstaining individuals.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have made minor revisions to address the comments raised in the previous round of review.]

      Summary:

      This manuscript by Alonso-Caraballo et al, is a novel piece of work that examines the impact of oxycodone self-administration on neural plasticity within paraventricular thalamic (PVT) to nucleus accumbens shell (Shell) pathway - two regions shown to play a key role in cue-induced drug seeking on their own, and whether this plasticity varies based on abstinence period and biological sex.

      Strengths:

      The authors show using a clinically relevant long-access model of opioid self-administration promotes dependence and acute withdrawal in both male and female rats. During subsequent cue-induced relapse tests at 1 or 14-days following the conclusion of self-administration, data show that while both male and females demonstrate drug-seeking behavior at both time points, females show a further elevation in responding on day 14 versus day 1 that is not observed in the males. When accounting for past work showing elevations in drug seeking in males after 30 days, these data indicate that craving-induced relapse for opioids may develop faster and may be more pronounced in females compared to males.

      These behavioral findings were paralleled by use of ex vivo acute slice electrophysiology and circuit-specific ex vivo optogenetics to examine the impact of oxycodone self-administration on synaptic strength within the paraventricular thalamus (PVT) to nucleus accumbens shell (NAcSh) pathway(s). Data support a time-dependent but sex independent strengthening of glutamatergic signaling at PVT-to-NAcSh medium spiny neurons (MSNs) that is only present following a relapse test at 14 days post abstinence in males versus females, providing the first evidence that opioid self-administration and/or cue-induced drug-seeking augments this pathway. Using an extensive set of physiological measures, the authors show that this increased synaptic strength reflects a upregulation of presynaptic release probability. Further, this upregulation of excitatory signaling aligned temporally with an increase in MSN excitability, as assessed by increases in action potential firing frequency. Finally, the authors provide the first evidence that similar to other inputs to the NAcSh, PVT projections innervate both MSN as well as local interneurons, promoting a GABA-A specific feedforward inhibitory circuit. Interestingly, unlike direct excitatory inputs to MSNs, no changes were observed ostensibly within this feedforward circuit, highlighting a selective enhancement of excitatory drive and output of MSNs with protracted abstinence.

      Overall, these data highlight a potential role for heightened synaptic strength within the PVT-NAcSh pathway in cue-induced relapse behavior during protracted abstinence and identify a potential therapeutic target during abstinence to reduce relapse risk in abstaining individuals.

      Weaknesses:

      Overall, the experimental approach and data provided appear rigorous and support their overall conclusions and achieve their goal of understanding how opioid self-administration impacts synaptic strength within the PVT-NAcSh pathway. Although not undermining these data, there are a few potential weaknesses that reduce the impact of the work. For example, the inability to directly assess whether cue-induced drug-seeking is in fact augmented compared to daily intake during self-administration in the maintenance face only permits the authors to denote that reexposure to cues and the context is sufficient to promote active lever pressing without demonstrating whether seeking behavior is in fact elevated further during a cue test. This is notably understandable as drug available sessions were 6-hours versus a 1hour relapse test. Importantly, it is clearly demonstrated that drug seeking is higher on average in female mice after 14 days versus 1 day.

      With regard to interpretation of electrophysiology findings, the lack of inclusion of an abstinence only group does not permit interpretations to parse out whether observed increases in synaptic strength (or the lack of) reflect abstinence or an interaction between abstinence period and re-exposure to the operant chamber, as slices were taken 30-45 min post relapse test. While much literature has shown that drug induced adaptations in the NAc requires a post drug period for plasticity to measurably emerge, studies have also shown that re-exposure to heroin-associated cues following abstinence seemingly "reverses" increases in cell excitability in prelimbic-NAc pyramidal neurons (Kokane et al., 2023) and that depotentiation of morphine-induced increases in synaptic strength in the NAc shell can be depotentiated by drug re-exopsure -- an effect also observed with cocaine re-exposure (Madayag et al., 2019). Notably, the lack of effect at 14 but not 1 day supports the likelihood that the relapse test does not in fact influence the plasticity within the PVT-NAcSh circuit.

      While the lack of effect on AMPAR:NMDAR ratio and rectification indices do support the notion that enhanced EPSC amplitudes in input-output curves do not reflect a change in AMPAR subunit expression (i.e., increased GluA2-lacking receptors that exhibit inward rectification at depolarized potential) nor a change in postsynaptic sensitivity to glutamate, without direct assessment of AMPAR-specific and NMDAR-specific input-output curves, it doesn't definitively exclude the possibility that both AMPA and NMDA receptor currents are being upregulated, thus negating an observable change in postsynaptic strength.

      Overall, these findings provide novel insight into how the PVT-NAcSh pathway is altered by opioid self-administration and whether this is unique based on abstinence period and sex. Importantly, these were the primary objectives stated by the author. Data highlight a potential role for the observed adaptations in relapse behavior and identify a potential therapeutic target during abstinence to reduce relapse risk in abstaining individuals. However, it should be noted that no causal link is demonstrated without experiments to reduce/prevent relapse.

      Comments on previous revisions:

      The authors addressed previous concerns brought up, specifically by clarifying data interpretation as well as text modifications related to potential caveats of these interpretations.

    3. Reviewer #2 (Public review):

      Summary:

      This is an interesting paper from Alonso-Caraballo and colleagues that examines the influence of opioid use, acute and prolonged abstinence, and sex on cue-induced relapse and paraventricular thalamus (PVT) to nucleus accumbens shell (NAcSh) medium spiny neurons circuit physiology. The study presents a valuable finding that following prolonged, but not acute abstinence from oxycodone self-administration, female rodents exhibit higher relapse rates to drug paired cues. Additionally, the study presents the useful finding that prolonged abstinence increased PVT-NAcSh MSN synaptic strength in both sexes, an effect that is likely due to presynaptic adaptations. While the evidence to support these two findings is solid, further experiments are required to determine the functional role of the PVT-NAcSh MSN circuit in relapse following prolonged oxycodone abstinence, and the mechanism underlying the heightened relapse vulnerability in females in this model of opioid use disorder.

      Strengths:

      The paper is interesting, well written and presented, and the experiments are well designed and conducted. The revised analysis of spike count data that models the hierarchical structure of the data is appropriate to overcome low animal numbers and the potential for oversampling. The authors are transparent in reporting the results related to this analysis in figure 5 and acknowledge the study is underpowered to confirm the trend of increased intrinsic excitability in male MSNs following prolonged oxycodone analysis.

      Impact:

      The topic is of interest to the field of substance use disorders and gives solid evidence for the need to consider targeted therapeutics aimed at relapse prevention in opioid use disorder.

    4. Reviewer #3 (Public review):

      Summary:

      Alonso-Caraballo et al. use behavioral testing and ex vivo patch-clamp electrophysiology combined with circuit-specific optogenetic stimulation of PVT terminals to examine how oxycodone self-administration and abstinence duration shape cue-induced relapse and PVT-NAcSh synaptic transmission in male and female rats. In the revision, the authors reanalyzed intrinsic excitability using nested hierarchical GLMMs, acknowledged the low power in the male prolonged-abstinence group, and expanded the discussion of relevant PVT-NAc literature. These changes improve the manuscript. That said, most of the revisions are textual and the main experimental gap remains. Both sexes show increased oxycodone seeking compared to saline at 14 days, but only females show a time-dependent incubation from 1 to 14 days, and the PVT-NAcSh synaptic strengthening is the same in both sexes. Nothing in the revision brings those two observations closer together. The excitability data also come from NAcSh MSNs with no confirmation of PVT connectivity, which limits what circuit-specific conclusions can be drawn. The study is a solid characterization of abstinence-related synaptic changes in this pathway, but some of the conclusions still go further than the data allow.

      Strengths:

      The behavioral characterization is thorough and well-executed, covering self-administration, somatic withdrawal, and cue-induced relapse across two abstinence durations in both sexes. The sex-specific escalation in oxycodone seeking from 1 to 14 days in females but not males is a clear and compelling finding. The use of circuit-specific ex vivo optogenetics to isolate PVT terminal inputs onto NAcSh neurons is a genuine methodological strength, and the demonstration of feedforward inhibitory recruitment through local GABAergic interneurons adds meaningful novelty to the circuit characterization. The reanalysis of intrinsic excitability using nested hierarchical GLMMs appropriately accounts for the non-independence of cells recorded within the same animal and is a real improvement over the original approach. The expanded discussion of prior PVT-NAc work, particularly the more accurate treatment of Keyes et al. (2020) and Paniccia et al. (2024), better situates the findings within the existing literature.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1:

      I recommend that the title be changed to not focus on sex differences to avoid misunderstanding.

      We thank Reviewer #1 for this suggestion and agree that the original title could create a misleading impression. We have updated the title from "Sex-specific behavioral and thalamo-accumbal circuit adaptations after oxycodone abstinence" to "Thalamo-accumbal circuit adaptations following extended oxycodone abstinence" to more accurately reflect the scope of the findings.

      The authors should also address the lack of difference physiologically compared to the behavior as a caveat more clearly in the discussion.

      We thank the reviewer for this important suggestion. We have revised the Discussion to explicitly address this dissociation. Specifically, we added the following to the PVT-NAcSh synaptic strength section: " The absence of sex differences in PVT-NAcSh synaptic measures suggests that this pathway, as characterized here, represents a shared neuro-adaptation to prolonged oxycodone abstinence rather than a substrate for the heightened relapse vulnerability observed in females. The mechanisms driving sex-specific relapse likely involve additional circuit elements, such as sex hormone-dependent modulation, upstream inputs, or cell-type specific plasticity." This point is also summarized in the abstract.

      Reviewer #2:

      A major weakness of this study is the disconnect between the behavioral and neurophysiological data reported. While a striking sex difference in relapse-like behavior is observed, there are no statistically significant sex differences in any of the neurophysiological data reported. Moreover, without an experiment to functionally test the role of the PVT-NAc projection in relapse-like behavior following prolonged oxycodone, these two arms of the study seem divorced.

      We respectfully disagree with the characterization that the behavioral and neurophysiological data are "divorced." The two arms of the study converge on a consistent and meaningful finding: PVT-NAcSh synaptic strength increases specifically after prolonged abstinence, this is the same time point at which enhanced cue-induced relapse is observed in both sexes. The absence of sex differences in synaptic measures does not weaken this convergence; it refines it by suggesting that circuit-level potentiation is a shared neuro-adaptation, while the sex-specific behavioral phenotype likely reflects additional modulatory mechanisms acting on this shared substrate. We have revised the Discussion to explicitly address this dissociation, as noted in our response to Reviewer #1 above. We acknowledge that functional manipulation of the PVT-NAcSh circuit would be required to establish causality, and we state this clearly in the manuscript.

      In the introduction the authors state they aim to test the hypothesis that increased synaptic strength in PVTNAcSh projections are necessary for drug-seeking. This study does not include the required experiments to test this hypothesis.

      We have revised the relevant section in the Introduction to accurately reflect the scope of our study: " We aimed to determine whether synaptic strength in PVT-NAcSh projections is affected following oxycodone abstinence and whether such changes are associated with cue-induced relapse and drug-seeking. Additionally, we examined whether there are sex-specific differences in either cue-induced relapse or PVT-NAcSh synaptic transmission after either 1 (acute) or 14 (prolonged) days of forced abstinence. Our results demonstrate that sex-specific enhancement in cue-induced relapse emerges after prolonged abstinence but not during acute abstinence from oxycodone self-administration. Although both males and females show increased cue-induced relapse after prolonged abstinence, females exhibited a greater relapse rate compared to males. Both sexes showed similar increases in PVT-NAcSh synaptic strength after prolonged abstinence, while synaptic strength was not altered after acute abstinence compared to saline controls. Together, these findings reveal a time-dependent increase in PVT-NAcSh synaptic strength and a sex-specific effect of prolonged abstinence on cue-induced relapse, while synaptic enhancements after prolonged abstinence were not sex-specific." This revision avoids implying a necessary or causal role for the circuit, which we did not test.

      Reviewer #3:

      The PVT-NAcSh synaptic strengthening after prolonged abstinence is statistically indistinguishable between sexes, while females but not males show a time-dependent escalation in oxycodone seeking from 1 to 14 days of abstinence. The Discussion proposes hormonal modulation or differences in upstream inputs as possible explanations, but none of these are tested and the gap is left unresolved.

      We agree that the mechanistic basis of the behavioral sex difference remains an open question that the current study does not resolve. As noted in our response to Reviewer #1, we have revised the Discussion to explicitly acknowledge this dissociation and to clarify that PVT-NAcSh synaptic strengthening represents a shared neuro-adaptation rather than a mechanism specific to the female behavioral phenotype. We maintain that identifying this dissociation is itself a scientifically meaningful finding.

      The intrinsic excitability recordings come from NAcSh MSNs with no confirmation that those neurons receive direct PVT input, which was raised in the original review, acknowledged in the revision, and not experimentally addressed.

      We have added the following clarification to the excitability section of the Discussion: “It should be noted that the intrinsic excitability recordings were designed to characterize general properties of NAcSh MSNs following oxycodone abstinence, independent of their synaptic inputs. As such, the excitability data should be interpreted as reflecting changes in the NAcSh MSN population broadly rather than in PVT-connected neurons specifically. The standing theory suggests that MSN excitability decreases as a homeostatic response to increased glutamatergic input [23,43,58]. Our data do not support a compensatory decrease in excitability in either sex at either abstinence time point”. These recordings were never intended to be circuit-specific; the experiment was designed to characterize NAcSh MSN excitability at the population level, which is a valid and informative question in its own right.

      The male prolonged-abstinence excitability trend has approximately 20% statistical power and is non-significant, yet the Discussion interprets it as a potential neuro-adaptation that could facilitate signal flow through the PVT-NAcSh circuit and contribute to relapse, which goes well beyond what the data support.

      We have revised the relevant Discussion text to ensure the male excitability trend is interpreted appropriately. The revised text now reads: "In males, a non-significant trend toward increased excitability was observed after prolonged abstinence, with a large effect size (Cohen's d = 1.18); however, given that this group was substantially underpowered (approximately 20% power), this finding should be interpreted with caution and cannot be taken as evidence of a neuro-adaptation. Whether this trend, if confirmed in future studies with larger cohorts, reflects a broader MSN population response or is specific to PVT-connected neurons remains an open and interesting question." The speculative mechanistic interpretation previously present in this section has been removed.

      The failure to distinguish between D1 and D2 MSNs remains a significant limitation given that cell-type specific plasticity at PVT-NAc synapses has been shown to be directly relevant to opioid seeking in prior work.

      We agree that distinguishing between D1 and D2 MSNs would provide important mechanistic insight, and we acknowledge this explicitly as a limitation and a future direction in the Discussion. The use of transgenic Cre rat lines for cell-type-specific recording in a self-administration model requires significant additional infrastructure and was beyond the scope of the present study. This is precisely the direction our laboratory is currently pursuing, and the present findings provide empirical motivation for those experiments.

      The Conclusion builds a mechanistic framework around D2 MSNs, PV interneurons, and D1 MSNs that is drawn from studies using different drugs or experimental designs, and none of these cell-type-specific mechanisms are tested in the present experiments.

      We thank the reviewer for this important critique. We have revised the opening of the Conclusion to clarify that the cell-type-specific framework is grounded in prior literature and represents a hypothesis for future investigation rather than a conclusion drawn from the present data. The revised text now reads: " When considered alongside prior work, our findings highlight the need to examine the anatomical and cell-type specific organization of PVT inputs to the NAcSh in the context of opioid relapse. Based on existing literature, PVT projections onto D2 MSNs and PV interneurons may contribute to relapse vulnerability, while adaptations involving D1 MSNs may underlie incubation of craving, though these mechanisms remain to be directly tested in the oxycodone self-administration model used here". We believe this framing accurately represents the relationship between our findings and the broader literature without overstating what the present data demonstrates.

    1. eLife Assessment

      This important study addresses the role of sphingolipid metabolism in maintaining endolysosomal membrane integrity and its impact on tau pathology in Caenorhabditis elegans and human cell culture models. The findings are convincing, and the proposed mechanisms are conceivable. The experimental evidence supports the conclusions of the study. The work will be of broad interest to cell biologists and biologists working on Alzheimer's disease and related proteinopathies.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, Tittelmeier et al. explored the role of sphingolipid metabolism in maintaining endolysosomal membrane integrity and its downstream effects on tau aggregation and toxicity, using both worms and human cell models. The authors showed that knockdown of sphingolipid metabolism genes reduced endolysosomal membrane fluidity, as revealed by FRAP and C-Laurdan imaging, leading to increased vesicle rupture. Furthermore, tau aggregates accumulated in endolysosomes and exacerbated membrane rigidity and damage, promoting seeded tau aggregation, likely by enabling tau seed escape into the cytosol. Importantly, unsaturated fatty acid supplementation restored membrane fluidity, suppressed tau propagation, and alleviated neurotoxicity in C. elegans. These findings provide insight into how lipid dysregulation contributes to tau pathology and highlight membrane fluidity restoration as a potential therapeutic avenue for Alzheimer's disease.

      Strengths:

      The study addresses the connection between sphingolipid metabolism, endolysosomal membrane integrity, and tau pathology, which is a relevant topic in the context of Alzheimer's disease and related tauopathies.

      The use of both C. elegans and human cell models provides cross-species perspectives that help frame the findings in a broader biological context.

      The combination of FRAP and C-Laurdan dye imaging offers a biophysical approach to investigate changes in membrane properties, which is a technically interesting aspect of the study.

      The observation that unsaturated fatty acid supplementation can modulate membrane fluidity and influence tau-related phenotypes adds an element of potential therapeutic interest.

      The study presents multiple experimental approaches to address the proposed mechanism, and efforts were made to examine both membrane behavior and tau aggregation dynamics.

      Comments on revised version:

      I thank the authors for their thorough revisions and detailed responses. All of my previous concerns have been satisfactorily addressed, and I have no further comments.

    3. Reviewer #2 (Public review):

      Tittelmeier et al. investigated the role of sphingolipid (SL) metabolism in the maintenance of endolysosomal vesicle integrity. They find that both impaired SL biosynthesis and degradation in C. elegans decreases the fluidity of endolysosomal membranes and promotes their rupture, while it has little effect on plasma membrane fluidity. Endolysosomal membrane fluidity is also negatively affected in human cells upon knockdown (KD) of a gene (SPHK2) involved in the SL degradation pathway. Aggregated forms of tau in both models (C. elegans and human cells) can also cause rigidification of the endolysosomal membrane, with SL homeostasis disruption having an additive effect, exacerbating endolysosomal rupture. Notably, KD of SPHK2 also increased the formation of tau foci, suggesting that compromised endolysosomal integrity may promote tau aggregation. These data provide a clearer understanding of how genetic manipulation of SL metabolism affects endolysosomal membranes and their rigidification in the context of tau aggregation. Supplementation of polyunsaturated fatty acids (PUFAs), which has a beneficial effect on Alzheimer's patients, improved membrane fluidity and reduced tau propagation in human cells and tau-associated neurotoxicity in C. elegans, suggesting a possible mechanism of action.

      Comments on revised version:

      The authors have:<br /> Corrected editorial errors (Points 1 and 2).

      Clarified the experimental rationale, added new data to rule out alternative explanation, and improved the presentation of the C. elegans model (Point 3).

      Provided experimental evidence and appropriate discussion regarding the specificity and broader physiological context of SL gene knockdown effects (Point 4).

      Overall, the authors' responses are thorough, supported by new data where appropriate, and demonstrate a clear understanding of the concerns raised. All points raised have been satisfactorily resolved.

    4. Reviewer #3 (Public review):

      Summary:

      The authors set off with an analysis of the lysosomal integrity upon knockdown of genes of the sphingolipid metabolic pathway that they identified in a previous work of an RNA screen using a new C.elegans Tau model. They then used cell culture and C.elegans experiments to study the link between lysosomal rupture and Tau propagation.

      Strengths:

      The authors use two complementary model systems and used probes to assess membrane rigidity that allow a quick assessment of the membrane dynamics and offer the opportunity to treat the cells with lipids, RNAi. Tau seeds etc.

      Comments on revised version:

      The authors have addressed the majority of my critical comments and thus I support the manuscript.

      They have still not analysed the knockdown efficiencies of their RNAi experiments. But this is their choice.

      The other publication establishing their Tau model is meanwhile published and there is no disconnect anymore between the model their analysis builds on.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Tittelmeier et al. explored the role of sphingolipid metabolism in maintaining endolysosomal membrane integrity and its downstream effects on tau aggregation and toxicity, using both worms and human cell models. The authors showed that knockdown of sphingolipid metabolism genes reduced endolysosomal membrane fluidity, as revealed by FRAP and C-Laurdan imaging, leading to increased vesicle rupture. Furthermore, tau aggregates accumulated in endolysosomes and exacerbated membrane rigidity and damage, promoting seeded tau aggregation, likely by enabling tau seed escape into the cytosol. Importantly, unsaturated fatty acid supplementation restored membrane fluidity, suppressed tau propagation, and alleviated neurotoxicity in C. elegans. These findings provide insight into how lipid dysregulation contributes to tau pathology and highlight membrane fluidity restoration as a potential therapeutic avenue for Alzheimer's disease.

      Strengths:

      The study addresses the connection between sphingolipid metabolism, endolysosomal membrane integrity, and tau pathology, which is a relevant topic in the context of Alzheimer's disease and related tauopathies.

      The use of both C. elegans and human cell models provides cross-species perspectives that help frame the findings in a broader biological context.

      The combination of FRAP and C-Laurdan dye imaging offers a biophysical approach to investigate changes in membrane properties, which is a technically interesting aspect of the study.

      The observation that unsaturated fatty acid supplementation can modulate membrane fluidity and influence tau-related phenotypes adds an element of potential therapeutic interest.

      The study presents multiple experimental approaches to address the proposed mechanism, and efforts were made to examine both membrane behavior and tau aggregation dynamics.

      We thank the reviewer for this positive assessment of the study.

      Weaknesses:

      In Figure 3, the authors used C-Laurdan imaging to assess membrane fluidity and showed that knockdown of SPHK2, the human ortholog of sphk-1, led to increased membrane rigidity. However, the authors did not co-stain with a lysosomal marker, making it unclear whether the observed effect is specific to lysosomal membranes or reflects general membrane changes. Co-staining with LysoTracker or applying segmentation masks to isolate lysosomal signals would significantly improve interpretation.

      We agree with the reviewer that it is important to isolate lysosomal signals for interpreting the C-Laurdan data. We therefore repeated and extended the C-Laurdan experiments in combination with LysoTracker staining and selectively analyzed LysoTracker-positive regions. These analyses showed pronounced increases in GP values in LysoTracker-positive vesicles after SPHK2 knockdown, supporting the conclusion that SPHK2 depletion increases endolysosomal membrane rigidity. We also performed LysoTracker-based analysis in the tau-fibril and fatty-acid experiments to better assess lysosome-associated membrane properties (see new and updated Figures 3B-E, Figures 5A and B, Figures S3A, B, F, G, Figures S4A-F, and Figures S5A-D for details). The respective Results sections have been revised accordingly.

      Line 173 states that Lipofectamine 2000 increases membrane fluidity based on GP index changes, but this is incorrect. A higher GP index indicates increased membrane order (i.e., reduced fluidity), so the statement should be revised. Additionally, Lipofectamine 2000 can itself alter membrane rigidity, posing a risk of false-positive interpretations. To confirm the role of SPHK2 in this phenotype, the authors should use a CRISPR/Cas9 knockout model instead of relying solely on siRNA transfection, which may be confounded by the delivery reagent. Without lysosomal co-staining and SPHK2 KO validation, the authors cannot conclusively claim that SPHK2 loss affects endolysosomal membrane integrity.

      We thank the reviewer for pointing out the incorrect wording regarding Lipofectamine. A higher GP index indicates increased membrane order/rigidity, not increased fluidity. Since the revised main figure now includes the SH-SY5Y data (Figure 3A-D), in which Lipofectamine alone did not significantly alter GP values (see Figure S3A, B), we removed the misleading statement from the Results.

      We also agree that Lipofectamine can affect membrane properties and therefore needs to be carefully controlled. In all siRNA-mediated experiments, SPHK2 siRNA was compared to a matched control siRNA condition exposed to the same transfection reagent. We state in the Methods that cells were transfected with either SPHK2 or scrambled control siRNA and that the medium was exchanged after 6 h “to minimize lipofectamine impact on the endolysosomal system”. Thus, the effect attributed to SPHK2 KD is assessed relative to the appropriate Lipofectamine-containing control condition.

      Importantly, we have now added lysosomal co-staining to address the reviewer’s concern about compartment specificity. This analysis showed that SPHK2 KD resulted in a pronounced increase in GP values within LysoTracker-positive compartments, demonstrating increased membrane rigidity at lysosomes. Thus, the revised data support the conclusion that SPHK2 KD increases lysosome-associated membrane rigidity, rather than only causing nonspecific effects on other cellular membranes.

      We also clarified the relationship between the current siRNA-based assay and our previous CRISPR inhibition-based analysis. In the revised Results, we now write: “While SPHK2 KD alone significantly increased galectin puncta above the matched control, its effect was more modest than in our previous CRISPR inhibition-based analysis [20]. This difference likely stems from the earlier readout required for the combined siRNA/tau fibril assay, when transient Lipofectamine-associated effects still increased the control background.”

      This addresses why the SPHK2 KD effect appears smaller in the current siRNA/tau-fibril assay than in our previous CRISPR inhibition-based analysis. The previous study, which is now peer-reviewed and published in the journal Autophagy, used a CRISPR inhibition-based strategy to reduce SPHK2 levels, which resulted in a highly significant increase in sfGFP-LGALS3 foci formation compared to the control [1]. Thus, the SPHK2 phenotype is not supported solely by the current siRNA experiment.

      In addition, we sought to genetically validate the RNAi phenotypes using mutant strains. However, mutant strains were not available for all sphingolipid metabolism hits analyzed in this study. We therefore used the sphk-1 mutant strain available at CGC (CZ24969; sphk-1(ju831)) to validate one of the key SL metabolism hits independently of RNAi. The revised manuscript states: “As genetic validation independent of RNAi, we tested an available sphk-1 mutant strain, which also showed a robust increase in hypodermal sfGFP::LGALS3 foci (Figure S1A).” This result supports the conclusion that genetic perturbation of sphingosine kinase activity compromises endolysosomal integrity in vivo.

      Together, the revised manuscript addresses the reviewer’s concerns by correcting the GP interpretation, controlling the siRNA experiments against matched Lipofectamine-treated controls, adding LysoTracker-based lysosome-associated C-Laurdan analysis, relating the current siRNA data to our previous CRISPR inhibition-based analysis, and providing genetic validation for the available sphk-1 mutant.

      The section titled "Fibrillar tau increases membrane rigidity and exacerbates endolysosomal damage" (lines 177-215) requires substantial revision. The narrative jumps abruptly between worms and cell models, making it hard to follow the logic. The use of the F3ΔK281::mCherry strain is introduced without explanation or context. It is unclear whether this strain is relevant to lysosomal membrane rupture, as no reference or justification is provided. The authors should clarify whether this reporter is intended to detect lysosomal membrane permeabilization (LMP). If so, it would be more appropriate to use established LMP reporters, such as lysosome-targeted fluorescent sensors, galectin-based reporters, or dextran leakage assays. Based on the current data in Figure 3G, it is difficult to draw firm conclusions regarding membrane rupture levels.

      We agree that this section required clarification, and we have substantially revised the Results to improve the logic and separation between model systems.

      First, we now introduce the C. elegans reporter strain earlier in the manuscript, in the first Results section. In the revised text, we explain both the tau construct and the actual lysosomal damage reporter: “In this strain, endolysosomal membrane damage is monitored in the hypodermis by expression of human galectin-3 fused to superfolder-GFP (sfGFP::LGALS3). The animals also express an aggregation-prone tau fragment fused to mCherry (F3ΔK281::mCherry) in touch receptor neurons, which is transmitted to the hypodermis, as described previously [20].” We also clarify the principle of the Galectin reporter: “Under steady-state conditions, sfGFP::LGALS3 remains diffusely distributed throughout the cytosol. Upon endolysosomal damage, luminal β-galactosides become exposed and recruit sfGFP::LGALS3 into visible puncta, providing a sensitive readout of vesicle rupture.” Thus, F3ΔK281::mCherry is not the reporter for lysosomal membrane permeabilization; the membrane-damage readout is sfGFP::LGALS3 puncta formation.

      Second, we reorganized the manuscript to separate the human cell experiments from the C. elegans experiments more clearly. The revised section “Fibrillar tau and SPHK2 KD act in concert to exacerbate endolysosomal damage and seeded tau aggregation” now focuses on human cell data. The C. elegans experiments are now presented in a separate section, “Tau transmission sensitizes endolysosomal membranes to sphingolipid perturbations in vivo.” We believe that this revised structure now clearly distinguishes the role of the Galectin reporter from the tau transmission model, separates the human cell and C. elegans data, and avoids the abrupt transitions between model systems noted by the reviewer.

      To support the conclusion that sphingolipid metabolism gene knockdown alters membrane properties, the study would benefit from direct lipidomic analysis. Measuring changes in sphingolipid profiles in both C. elegans and cell models would provide biochemical evidence for the proposed disruption of lipid homeostasis. Given the availability of lipidomics platforms, this type of analysis should be feasible in both worms and human cells and would significantly strengthen the mechanistic claims regarding membrane fluidity and integrity.

      Because we did not perform lipidomics in the present study, we have revised the wording throughout the manuscript to avoid implying that we directly measured lipid composition. Instead, we now refer to “genetic perturbation/disruption of sphingolipid metabolism” or “knockdown of enzymes involved in sphingolipid metabolism” when describing our experimental interventions.

      We agree that lipidomic analyses will be important in future studies to define how perturbation of sphingolipid metabolism changes lipid composition in C. elegans and human cells. However, lipidomics itself would not directly establish which lipid changes causally drive the membrane rigidification observed in our study. Membrane fluidity is a biophysical property determined by the combined composition of the membrane, including lipid abundance, saturation, acyl-chain length, head groups, sterol content, and membrane-associated proteins. Thus, even if lipidomics identified changes in sphingolipid profiles, these changes could not be directly translated into a predictable effect on membrane fluidity without additional biophysical validation, using Laurdan dye imaging or FRAP. Moreover, whole-cell or whole-animal lipidomics would not resolve whether the relevant lipid changes occur specifically at endolysosomal membranes, which are the focus of our study.

      We have now clarified this point in the Discussion. Specifically, we state that “even detailed lipidomics would not by itself identify which lipid changes are responsible for the observed membrane rigidification” and that future lysosome-enriched or organelle-specific lipidomic approaches should be combined with direct manipulation of candidate lipid species, followed by measurements of membrane fluidity and rupture, to determine which lipid changes causally contribute to endolysosomal membrane rigidification. In the present study, we therefore focused on direct quantitative biophysical readouts of membrane properties in C. elegans. We used FRAP of the lysosomal membrane protein LAAT-1::mCherry to assess lateral mobility within lysosomal membranes and showed that knockdown of sphingolipid-metabolism genes increased the time to half-maximal recovery, indicating reduced lysosomal membrane fluidity. Notably, knockdown of genes involved in both sphingolipid biosynthesis and sphingolipid degradation increased membrane rigidity. This makes it unlikely that the observed rigidification is caused by accumulation or depletion of a single shared lipid species. Rather, perturbations at different steps of sphingolipid metabolism may lead to distinct lipidomic changes that nevertheless converge on a common biophysical outcome: reduced endolysosomal membrane fluidity. In parallel, we used C-Laurdan imaging to quantify membrane order and found that SPHK2 knockdown in SH-SY5Y human neuroblastoma cells increased GP values, consistent with increased membrane rigidity. Two-channel thresholding of LysoTracker-positive compartments further showed that SPHK2 knockdown increased GP values in lysosome-associated regions.

      Thus, although lipidomics will be valuable to define the underlying lipid changes in future work, the current data already provide convergent quantitative evidence from independent membrane-fluidity readouts across C. elegans and human cell models. This cross-model consistency strengthens the robustness and reproducibility of the central conclusion that perturbation of sphingolipid metabolism alters endolysosomal membrane properties and promotes membrane rupture.

      The conclusions of the study rely heavily on imaging-based assays, including FRAP, C-Laurdan, and fluorescence microscopy. While these approaches provide valuable spatial and qualitative insights, they are inherently indirect and subject to interpretive limitations. To strengthen the mechanistic claims, the authors should incorporate additional biochemical or quantitative approaches. For example, lipidomics would allow direct measurement of membrane lipid composition changes, and western blotting or quantitative proteomics could assess levels of membrane-associated proteins involved in endolysosomal function or stress responses. Including such data would significantly improve the robustness and reproducibility of the study's conclusions.

      We agree that lipidomic and proteomic analyses will be important in future studies to define which sphingolipid species and/or membrane-associated proteins contribute to the observed rigidification of endolysosomal membranes. In response to this point, we have revised the wording throughout the manuscript to more precisely distinguish our experimental interventions from inferred changes in lipid composition. Because we did not directly measure lipid composition in the present study, we now refer more specifically to “genetic perturbation/disruption of sphingolipid metabolism” or “knockdown of enzymes involved in sphingolipid metabolism” when describing our data, rather than implying that global sphingolipid homeostasis was directly quantified. We retain “sphingolipid imbalance” only in interpretive or model-based statements where appropriate.

      However, we respectfully disagree that the current data are only qualitative. FRAP and C-Laurdan GP imaging are established quantitative biophysical approaches: FRAP provides quantitative parameters such as the time to half-maximal recovery and the mobile fraction, whereas C-Laurdan GP provides a ratiometric measurement of membrane lipid order and packing. Similarly, the Galectin puncta assay is an established quantitative readout of lysosomal membrane permeabilization. Thus, while these approaches are imaging-based, they provide quantitative readouts of membrane mobility, membrane order, and membrane rupture, respectively.

      We also note that lipidomic and proteomic profiling, although valuable, would not by itself establish which lipid or protein changes causally drive the membrane rigidification observed in our study. Membrane fluidity is an emergent biophysical property determined by the combined composition of the membrane, including lipid abundance, saturation, acyl-chain length, head groups, sterol content, and membrane-associated proteins. Therefore, an increase or decrease in a given lipid or protein species cannot be directly translated into a predictable change in membrane fluidity without additional biophysical validation. This point is further supported by our observation that knockdown of genes involved in both sphingolipid biosynthesis and sphingolipid degradation increased endolysosomal membrane rigidity. These perturbations would be expected to affect lipid composition in different, possibly even opposing, ways, making it unlikely that the shared rigidification phenotype is caused by accumulation or depletion of one single lipid species. Rather, distinct lipidomic changes may converge on a common biophysical outcome: reduced endolysosomal membrane fluidity.

      We have clarified this point in the Discussion and now state that future lysosome-enriched or organelle-specific lipidomic/proteomic approaches should be combined with direct manipulation of candidate lipid or protein species, followed by measurements of membrane fluidity and rupture, to determine which changes causally contribute to endolysosomal membrane rigidification. Such experiments would address the distinct question of which molecular components mediate the effect. By contrast, the central aim of the present study was to test whether genetic perturbation of enzymes involved in sphingolipid metabolism alters membrane fluidity and thereby promotes endolysosomal rupture and tau seeding.

      For this question, direct biophysical measurements of membrane fluidity and quantitative readouts of membrane rupture are the most relevant assays. We therefore used complementary quantitative approaches in two distinct model systems: FRAP of the lysosomal membrane protein LAAT-1::mCherry in C. elegans and C-Laurdan GP imaging in human cells. The fact that perturbing sphingolipid metabolism reduced endolysosomal membrane fluidity in C. elegans and increased lysosome-associated membrane rigidity in human cells supports the robustness and reproducibility of the central conclusion across independent model systems. In the revised manuscript, we further strengthened the human-cell data by adding SH-SY5Y neuroblastoma cells as a neuronal-like model and by combining C-Laurdan imaging with LysoTracker-based analysis to assess lysosome-associated membrane properties.

      To further address causality, we manipulated membrane fluidity independently of sphingolipid metabolism enzymes using fatty acid supplementation. Increasing membrane rigidity with PA exacerbated tau-induced endolysosomal rupture and seeded aggregation, whereas increasing membrane fluidity with ALA reduced tau-induced membrane rigidification, endolysosomal rupture, and seeded aggregation. Thus, the revised manuscript combines genetic perturbation of sphingolipid metabolism, quantitative membrane-fluidity measurements, whole-cell and lysosome-associated C-Laurdan analysis, and Galectin-based rupture assays across complementary models.

      Regarding lysosomal function, we agree that functional readouts are informative, but lysosomal membrane rupture and global lysosomal degradative capacity are related but not identical readouts. This distinction is supported by Yong et al., who reported that lipid dysregulation can induce lysosomal membrane permeabilization and lysosomal accumulation of endogenous protein aggregates without broadly impairing core lysosomal or proteasomal functions [2]. Accordingly, the absence of overt defects in general lysosomal activity would not necessarily exclude membrane damage.

      The human cell experiments were performed exclusively in HEK293T cells, which are not physiologically relevant for modeling Alzheimer's disease or lysosomal function in neurons. Given that the study aims to draw conclusions related to tau aggregation and lysosomal membrane integrity, the use of a more disease relevant cellular model is essential. There are several established AD-relevant cell models, including iPSCderived neurons, neuroblastoma lines expressing tau, or microglial models, which would better reflect the cellular context of tauopathies. Validation of key findings in at least one of these systems would substantially enhance the biological relevance and translational impact of the study.

      We have expanded and clarified the human cell data in the revised manuscript. Specifically, we now include SH-SY5Y human neuroblastoma cells for key C-Laurdan experiments assessing membrane rigidity after SPHK2 knockdown. We also show that recombinant tau fibrils increased membrane rigidity in SHSY5Y and HEK293T cells, including in LysoTracker-positive compartments.

      Importantly, the HEK293T cells are used for specific, established quantitative assays rather than as a model of neuronal toxicity. In particular, HEK293T sfGFP-LGALS3 cells are used to quantify galectin puncta formation as a readout of endolysosomal rupture, and HEK tau-Venus biosensor cells are used to quantify seeded tau aggregation. Thus, SH-SY5Y cells and HEK293T cells are used for complementary purposes: SHSY5Y cells provide a more neuronal-like human cell context for membrane-rigidity measurements, whereas HEK293T reporter/biosensor cells provide robust quantitative assays for galectin puncta formation and tau seeding.

      In addition, tau-associated neuronal dysfunction and toxicity were assessed in vivo, in functional C. elegans touch receptor neurons. In the revised manuscript, we show that ALA supplementation mitigated the age-dependent touch-response deficit and reduced neurotoxicity in animals expressing F3ΔK281::mCherry in touch receptor neurons. We have also revised the wording throughout the manuscript to avoid implying that HEK293T cells are used to model neuronal toxicity.

      Finally, the relevance of these hits to human neuronal tau seeding is also supported by our previous study, in which conserved hits from the C. elegans screen, including sphingosine kinase perturbation, were validated in human iPSC-derived neurons for their effect on seeded tau aggregation [1]. We now cite this published study where appropriate. Together, the revised manuscript combines neuronal-like human SH-SY5Y cells, established HEK293T tau-seeding and galectin reporter assays, in vivo neuronal readouts in C. elegans, and prior validation in human iPSC-derived neurons, thereby strengthening the biological relevance of the conclusions while using each model for the assay in which it is most informative.

      The authors reported that PUFA supplementation rescues neurotoxic phenotypes by increasing membrane fluidity. However, the data supporting this claim rely entirely on confocal imaging, shown in both the main and supplemental figures. To substantiate the mechanistic link between PUFA treatment and improved lysosomal membrane properties, the authors should include functional assays demonstrating that PUFAs are indeed incorporated into lysosomal membranes. Additionally, lipidomics analysis would be valuable to identify which lipid species are altered upon supplementation and correlate these changes with the observed phenotypic rescue. Furthermore, the conclusion that PUFAs rescue "neurotoxic phenotypes" is not appropriate based on data derived solely from HEK293T cells, which are not neuronal. To make claims about tau-related neurotoxicity, the authors should validate their findings in a more relevant neuronal model, such as SH-SY5Y neuroblastoma cells expressing tau or iPSC-derived neurons. This would better reflect the cellular environment of Alzheimer's disease and provide stronger support for the proposed therapeutic potential of PUFA supplementation.

      We agree that PUFA supplementation can have effects beyond membrane fluidity and that our data do not directly demonstrate incorporation of ALA into lysosomal membranes. We have therefore revised the Discussion to explicitly acknowledge this limitation. In the revised text, we state that “PUFAs can also influence lipid signaling, oxidative stress responses, and broader membrane remodeling” and that we “cannot exclude additional direct or indirect effects of ALA.” At the same time, we note that the opposing effects of PA and ALA, together with the sphingolipid-metabolism knockdown data, support membrane fluidity as a major determinant of endolysosomal membrane integrity and rupture in our models. To strengthen the link between ALA and lysosome-associated membrane properties, we combined CLaurdan imaging with LysoTracker-based analysis. In the revised Results, we show that ALA prevented tau-induced membrane rigidification and that LysoTracker-based analysis indicated effects on lysosome-associated membrane properties. ALA also reduced tau-induced endolysosomal rupture and seeded aggregation in human cell models.

      Regarding lipidomics, we refer to our response above and to the revised Discussion. We agree that lipidomics would be valuable to identify ALA-induced lipid changes, but such data would not by itself establish how these changes affect membrane fluidity without additional biophysical validation.

      Finally, we clarify that our conclusion regarding tau-associated neuronal dysfunction and toxicity is not based on HEK293T cells. HEK293T cells were used for established quantitative assays of Galectin puncta formation and seeded tau aggregation. The neurotoxicity experiments were performed in vivo in C. elegans touch receptor neurons, where ALA supplementation reduced galectin foci formation, mitigated age-dependent touch-response deficit and reduced neuronal toxicity.

      While the authors demonstrate that ALA supplementation mitigates neurotoxicity in C. elegans expressing aggregated tau (F3ΔK281::mCherry), the current data are not sufficient to conclude that ALA directly rescues tau aggregation toxicity via a lysosome-specific mechanism. It remains unclear how lipid composition is altered upon ALA treatment and whether these changes correlate with functional improvement of lysosomal pathways. The manuscript does not provide mechanistic insight into how ALA enhances lysosomal health or attenuates endolysosomal damage. Moreover, supplementation with PUFAs like ALA can activate a wide range of cellular processes beyond lysosomal function, including alterations in membrane fluidity, signaling cascades, and oxidative stress responses. The authors should clarify how they distinguish the lysosome-related effects from these alternative pathways. For example, did they observe specific lysosomal markers or structural improvements in lysosomes upon ALA treatment?

      Additional data or controls would be necessary to support a lysosome-specific protective mechanism and to exclude the involvement of other PUFA-responsive pathways in the observed phenotypes.

      We agree that our data do not prove that ALA acts exclusively through a lysosome-specific mechanism or that ALA is directly incorporated into lysosomal membranes. We have therefore revised the manuscript to avoid this interpretation and explicitly acknowledge alternative PUFA-responsive pathways. In the revised Discussion, we state that “PUFAs can also influence lipid signaling, oxidative stress responses, and broader membrane remodeling” and that we “cannot exclude additional direct or indirect effects of ALA.” We further conclude more cautiously that the opposing effects of PA and ALA, together with the sphingolipid metabolism perturbation data, support membrane fluidity as a major determinant of endolysosomal membrane integrity and rupture in our models.

      To strengthen the lysosome-related aspect of the mechanism, we added LysoTracker-based analysis to the C-Laurdan experiments. In the revised Results, ALA prevented tau-induced membrane rigidification, and LysoTracker-based analysis indicated that ALA also affected lysosome-associated membrane properties. ALA further reduced tau-induced Galectin puncta formation and seeded tau aggregation in human cell models. These data support an effect of ALA on lysosome-associated membrane order and rupture, while not excluding additional effects through lipid signaling, oxidative stress responses, or other PUFA-responsive pathways.

      Regarding lipid composition, we refer to the revised Discussion and our response above. We agree that lipidomics would be valuable to identify ALA-induced lipid changes, but such analyses would need to be organelle-specific and combined with biophysical validation to determine how candidate lipid changes affect membrane fluidity and rupture.

      Finally, we clarify that our conclusion regarding tau-associated neuronal dysfunction and toxicity is based on the C. elegans experiments, not on HEK293T cells. HEK293T cells were used for quantitative Galectin puncta and tau-seeding assays, whereas neuronal dysfunction and toxicity were assessed in vivo in touch receptor neurons in C. elegans. In the revised Results, we state that ALA supplementation mitigated galectin foci formation, age-dependent touch-response deficit and reduced neuronal toxicity in animals expressing F3ΔK281::mCherry in these neurons.

      Reviewer #2 (Public review):

      Tittelmeier et al. investigated the role of sphingolipid (SL) metabolism in the maintenance of endolysosomal vesicle integrity. They find that both impaired SL biosynthesis and degradation in C. elegans, decrease the fluidity of endolysosomal membranes and promote their rupture, while it has little effect on plasma membrane fluidity. Endolysosomal membrane fluidity is also negatively affected in human cells upon knockdown (KD) of a gene (SPHK2) involved in the SL degradation pathway. Aggregated forms of tau in both models (C. elegans and human cells) can also cause rigidification of the endolysosomal membrane, with SL homeostasis disruption having an additive effect, exacerbating endolysosomal rupture. Notably, KD of SPHK2 also increased the formation of tau foci, suggesting that compromised endolysosomal integrity may promote tau aggregation. These data provide a clearer understanding of how genetic manipulation of SL metabolism affects endolysosomal membranes and their rigidification in the context of tau aggregation. Supplementation of polyunsaturated fatty acids (PUFAs), which has a beneficial effect on Alzheimer's patients, improved membrane fluidity and reduced tau propagation in human cells and tau-associated neurotoxicity in C. elegans, suggesting a possible mechanism of action.

      Overall, the conclusions of this paper are supported by the data, with a few aspects requiring further clarification and elaboration.

      (1) A reference to Figure S2E-G, which shows that KD of SL biosynthesis genes do not affect the plasma membrane, is missing from the main text.

      We thank the reviewer for pointing this out. We have added the reference to the respective figures in the main text when discussing the plasma membrane FRAP experiments.

      (2) In Figure 3C, lipofectamine alone shows that it increases membrane rigidity (increased GP values), not membrane fluidity.

      We thank the reviewer for pointing out this incorrect wording. A higher GP index indicates increased membrane order/rigidity, not increased membrane fluidity. Since the revised main figure now includes the SH-SY5Y data, in which Lipofectamine alone did not significantly alter GP values, we removed the misleading statement from the Results. Importantly, all siRNA-mediated knockdown experiments were compared to matched control siRNA conditions exposed to the same transfection reagent. Thus, the effect attributed to SPHK2 KD is assessed relative to the appropriate Lipofectamine-containing control condition.

      (3) In Figure 3F, the EV cntl condition expressing F3:mCh tau should have increased LGALS3 foci compared to the mCh EV cntl according to Ref (20) and its Figure 2G (at least for Day 5 animals), which would be indicative of the tau spreading in hypodermal tissue. What C. elegans age was examined in Figure 3F? Can the authors provide evidence of the transmission of the F3:mCh tau from the touch receptor neurons to the hypodermis in the EV [similar to Figure 2C & D from Ref (20)] and compare it to the KDs? Otherwise, it seems that KD of SL genes impacts not only endolysosomal rupture but significantly affects tau accumulation/spreading as well (e.g., shown later in HEK cells, where SPHK2 KD increases the formation of tau-Venus foci).

      We thank the reviewer for raising this important point. The analysis referred to by the reviewer has now been moved to the revised C. elegans section and is presented as Figure 4A and B. The experiments were performed in the reporter strain used in our genome-wide screen In Ref (20), now published in Autophagy [1]. We clarified the purpose of the reporter strain and the relationship between tau transmission and the galectin puncta readout. In the revised manuscript, we now state: “In this strain, endolysosomal membrane damage is monitored in the hypodermis by expression of human galectin-3 fused to superfolder-GFP (sfGFP::LGALS3). The animals also express an aggregation-prone tau fragment fused to mCherry (F3ΔK281::mCherry) in touch receptor neurons, which is transmitted to the hypodermis, as described previously [20].” We further clarify that sfGFP::LGALS3 puncta formation, not F3ΔK281::mCherry, is the readout of endolysosomal rupture: “Under steady-state conditions, sfGFP::LGALS3 remains diffusely distributed throughout the cytosol. Upon endolysosomal damage, luminal β-galactosides become exposed and recruit sfGFP::LGALS3 into visible puncta, providing a sensitive readout of vesicle rupture”.

      Furthermore, we now better explain that transmitted tau sensitizes endolysosomal membranes to additional perturbations rather than necessarily inducing a strong lysosomal rupture phenotype on its own. In the experiments shown in Figure 4A and B, we compare F3ΔK281::mCherry animals with matched mCherry-only control animals that also express sfGFP::LGALS3 in the hypodermis. We now state: “We compared animals expressing F3ΔK281::mCherry in touch receptor neurons, from where it is transmitted to the hypodermis, with matched controls expressing mCherry alone in the same neurons. In both strains, sfGFP::LGALS3 is expressed in the hypodermis to monitor endolysosomal membrane damage.” We have also clarified the age of the animals in the revised figure legends.

      To experimentally address whether the enhanced rupture phenotype could be explained by altered tau transmission, we quantified hypodermal F3ΔK281::mCherry levels after sphk-1 RNAi (new Figure 4C, D). Importantly, sphk-1 RNAi did not increase hypodermal F3ΔK281::mCherry levels, arguing that the enhanced rupture phenotype is not due to increased tau transmission. Moreover, C. elegans neurons are largely refractory to systemic RNAi under the conditions used here [3]. We have added this important information to the Discussion. Specifically, the revised manuscript states that “the enhanced rupture phenotype is unlikely to result from a direct effect of RNAi on neuronal F3ΔK281::mCherry expression, as C. elegans neurons are largely refractory to systemic RNAi under the conditions used here,” supporting the interpretation that the RNAi treatments primarily affect endolysosomal integrity in the recipient tissue rather than neuronal tau expression itself.

      Finally, we would like to clarify that the increased tau-Venus foci in HEK cells should not be interpreted as a direct induction of tau aggregation by SPHK2 KD. Only upon addition of recombinant tau fibrils did SPHK2 KD significantly increase tau-Venus foci formation (Figure 3 H, I). This is consistent with the control experiments performed in human iPSCs in our previous study and supports our interpretation that perturbation of sphingolipid metabolism increases susceptibility to seeded tau aggregation by promoting endolysosomal rupture and tau seed escape, rather than by directly increasing tau aggregation or tau transmission.

      (4) Sphingolipids are essential membrane components and signaling molecules. Does KD of SL genes in C. elegans and the subsequent endolysosomal rupture cause any major, intermediate, or minor defects/phenotypes (in non-aggregation prone models, w/t.)?

      We agree that sphingolipids are essential membrane components and signaling molecules and that perturbing sphingolipid metabolism can have broader physiological consequences. In the revised manuscript, we address this point in two ways.

      First, we directly tested whether SL gene knockdown can induce endolysosomal rupture independently of aggregation-prone tau by using matched control animals expressing mCherry alone in touch receptor neurons while also expressing sfGFP::LGALS3 in the hypodermis. In these animals, knockdown of most SL-related hits resulted in nearly all animals displaying hypodermal sfGFP::LGALS3 foci, indicating that perturbation of SL metabolism can compromise endolysosomal integrity in the absence of transmitted F3ΔK281::mCherry (Figure 4A, B).

      Second, we have added a Discussion paragraph to place these findings into a broader physiological context. We now clarify that endolysosomal membrane rupture and global lysosomal function are related but not identical readouts. In support of this distinction, we discuss work showing that lipid dysregulation can induce lysosomal membrane permeabilization and lysosomal accumulation of endogenous protein aggregates without broadly impairing core lysosomal or proteasomal function [2]. Thus, membrane damage can occur even when general lysosomal activity is not overtly disrupted.

      We also discuss a recent study published during the revision of this manuscript that independently identified SPHK-1 as an important regulator of lysosomal integrity in C. elegans, showing that strong sphk1 loss-of-function causes lysosomal sphingosine accumulation, membrane rupture, impaired degradative function, cargo accumulation, developmental defects, and reduced lifespan [4].

      Importantly, while that study focused on a strong loss-of-function mutation in a single SL-metabolism gene, our data show that knockdown of multiple SL-metabolism genes, including genes involved in both SL biosynthesis and degradation, converges on reduced endolysosomal membrane fluidity and increased rupture. This suggests that the observed membrane rigidification and rupture are not specific to one mutant background but represent a broader consequence of perturbing SL metabolism at multiple points. A systematic characterization of all organismal phenotypes caused by each SL gene knockdown was beyond the scope of the present study. Therefore, the revised manuscript now makes clear that the study focuses on endolysosomal membrane fluidity and rupture because these membrane-level changes are directly linked to tau seed escape and seeded tau aggregation, while broader physiological consequences may vary depending on the strength and context of the perturbation.

      Reviewer #3 (Public review):

      Summary:

      The authors set off with an analysis of the lysosomal integrity upon knockdown of genes of the sphingolipid metabolic pathway that they identified in a previous (yet unpublished) work of an RNA screen using a new C. elegans Tau model. They then used cell culture and C. elegans experiments to study the link between lysosomal rupture and Tau propagation.

      Strengths:

      The authors use two complementary model systems and use probes to assess membrane rigidity that allow a quick assessment of the membrane dynamics and offer the opportunity to treat the cells with lipids, RNAi. Tau seeds, etc.

      Weaknesses:

      The main weakness is that this work builds on not-yet-peer-reviewed manuscript that established a new C. elegans Tau model and RNAi screen that aimed to identify genes involved in the propagation of Tau.

      This reviewer misses essential information of the C. elegans Tau strain (not included in the method section): e.g., promoter used for the expression, information on the used Tau variant, expression pattern, and aggregation, etc.

      We thank the reviewer for raising this point. The related study establishing the C. elegans tau transmission model and RNAi screen has now been peer-reviewed and published in Autophagy [1]. We now cite the published article throughout the revised manuscript instead of the previous preprint.

      We also agree that the current manuscript should be understandable without requiring the reader to consult the previous paper for the basic logic of the model. We therefore added a clearer introduction of the reporter strain in the Results. Specifically, we now explain that the strain expresses the aggregation-prone tau fragment F3ΔK281::mCherry in touch receptor neurons, that this tau fragment is transmitted to the hypodermis, and that endolysosomal membrane damage is monitored in the hypodermis using sfGFP::LGALS3. We further clarify that sfGFP::LGALS3 remains diffuse under steady-state conditions and forms puncta upon endolysosomal membrane damage, when luminal β-galactosides become exposed. Thus, the revised manuscript now provides the key information needed to understand the experimental system used here, while the published Autophagy paper is cited for the full characterization of the tau transmission model, expression pattern, aggregation properties, and original genome-wide RNAi screen.

      Throughout the study, I missed data on:

      (1) Effect of the knockdown on Tau expression, localisation (with lysosomal membrane?), aggregation, and proteotoxicity. The effect of the RNAi-mediated knockdown could also simply lead to a reduced expression of Tau that, in turn, leads to suppressed propagation.

      We agree that it is important to distinguish effects on tau expression/transmission from effects on endolysosomal membrane integrity. In the C. elegans experiments, F3ΔK281::mCherry is expressed in touch receptor neurons and transmitted to the hypodermis, where sfGFP::LGALS3 reports endolysosomal membrane damage. We now describe this more clearly in the revised Results.

      The RNAi treatments target genes involved in sphingolipid metabolism under systemic RNAi conditions. Because C. elegans neurons are largely refractory to systemic RNAi in the absence of sensitizing backgrounds [3], which we did not use, a direct RNAi-mediated reduction of neuronal F3ΔK281::mCherry expression is very unlikely. We have added this point to the Discussion, stating that “the enhanced rupture phenotype is unlikely to result from a direct effect of RNAi on neuronal F3ΔK281::mCherry expression, as C. elegans neurons are largely refractory to systemic RNAi under the conditions used here.”

      Experimentally, we also tested whether sphingolipid perturbation alters transmitted tau levels (new Figure 4C, D). Specifically, we quantified hypodermal F3ΔK281::mCherry after sphk-1 RNAi and found no increase, arguing that the enhanced rupture phenotype is not due to increased tau transmission.

      Moreover, in the cell-based tau-Venus assay, SPHK2 knockdown alone did not induce detectable tau aggregation in the absence of exogenously added tau fibrils (Figure 3H, I). Only upon addition of recombinant tau fibrils did SPHK2 knockdown significantly increase tau-Venus foci formation. We now state this explicitly in the revised Results and conclude that disruption of sphingolipid metabolism is not sufficient on its own to initiate detectable tau aggregation under the conditions tested here but rather increases cellular susceptibility to seeded tau aggregation when tau fibrils are present.

      Together, these data argue against a direct effect of sphingolipid gene knockdown on tau expression or spontaneous tau aggregation. Instead, they support our interpretation that perturbation of sphingolipid metabolism compromises endolysosomal membrane integrity, thereby facilitating tau seed escape and seeded aggregation when tau seeds are present.

      (2) A quantification of RNAi knockdown is needed to judge the efficiency of the RNAi, in particular for the combinatorial RNAi experiments involving 2 and even 4 genes in parallel. Ideally, these analyses should be validated with mutants for these genes.

      We agree that RNAi efficiency can vary between clones and that this is particularly relevant for combinatorial RNAi experiments targeting two or more potentially redundant genes. We have added this limitation to the Results section and now state: “Because KD efficiency was not assessed for the individual RNAi clones or co-RNAi combinations, these experiments do not allow comparison of relative RNAi strength or inference of the relative importance of individual genes. Thus, the conclusions drawn from these RNAi experiments are qualitative: specific single or combined KDs can promote endolysosomal rupture, whereas the absence of a detectable phenotype after RNAi cannot exclude gene involvement, as KD may have been insufficient.”

      Where mutant strains were available, we performed genetic validation. Specifically, a sphk-1 mutant available at CGC (CZ24969; sphk-1(ju831)) also showed increased hypodermal sfGFP::LGALS3 puncta (new Figure S1A), supporting the RNAi-based conclusion that genetic perturbation of sphingolipid metabolism compromises endolysosomal integrity. Corresponding mutant strains were not available for the other selected hits. Importantly, most hits also induced sfGFP::LGALS3 foci in human HEK293T cells as assessed in our previous study [1], providing additional support that the observed effects are not random RNAi artifacts.

      Further:

      (3) Figure 4 H, I: Would Tau also aggregate in the absence of externally added Tau?

      No. In the tau-Venus biosensor cell line, SPHK2 knockdown alone did not increase tau-Venus foci formation (now Figure 3H, I). Tau-Venus foci increased only after addition of recombinant tau fibrils and were further enhanced by SPHK2 knockdown. We now state this explicitly in the Results.

      (4) How specific is the effect for Tau? It would help if the authors could assess other amyloid proteins.

      We agree that similar membrane-level mechanisms may apply to other amyloid assemblies. We have therefore added recent literature to the Discussion supporting the broader concept that intralysosomal amyloid assemblies can physically deform and rupture lysosomal membranes. The revised manuscript states: “This interpretation is consistent with recent ultrastructural studies showing that intralysosomal amyloid assemblies can physically deform and rupture lysosomal membranes.” We further clarify that “whether this mechanism is specific to tau or also applies to other amyloid assemblies remains to be determined.”

      Whether perturbation of SL metabolism similarly affects endolysosomal escape and seeded aggregation of other disease-associated amyloid proteins is an important question that we plan to address in future work. However, these experiments require additional disease-specific models, aggregation assays, and validation, and are therefore beyond the scope of the present revision.

      (5) The connection between sphingolipids and AD is not new. See He et al, 2010, Neurobiol. Aging + numerous publications and also not between Tau seeding and lysosomal rupture: Rose et al., PNAS 2024 (that has been cited by the authors).

      We agree and our manuscript does not aim to establish these associations as new. We state explicitly that alterations in sphingolipid metabolism have been reported in aging and AD, and that endolysosomal rupture is increasingly recognized as a critical step in tau seed escape and propagation.

      The novelty of our study lies in mechanistically connecting these two previously established areas. Specifically, we show that genetic perturbation of enzymes involved in sphingolipid metabolism reduces endolysosomal membrane fluidity, promotes membrane rupture, and thereby increases susceptibility to tau seed escape and seeded aggregation. We have revised the Introduction and Discussion to better emphasize this mechanistic contribution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Figure formatting and annotation need improvement. Panel letters throughout the figures should be in uppercase, and gene names in pathway diagrams should be italicized for consistency. Several scale bars are missing, including in Figures 1C, 2A, and 2H, and should be clearly indicated in the figures and legends. In Figure 1C, the age of the worms used in the assay is not specified. While the Methods section mentions "age-synchronized animals," the precise age at the time of imaging or experimentation is not stated. It would strengthen the study to explore whether membrane integrity phenotypes vary between young adults (day 1) and older adults (day 7 or 10) across the different conditions. Figure 1B lacks sufficient detail describing the galectin puncta assay used. A brief explanation of the assay rationale and readout would help contextualize the findings.

      We thank the reviewer for pointing this out. We have revised the figures and figure legends accordingly by standardizing panel labels, adding scale bars where missing, and providing the age of animals used in the assays. We also expanded the description of the galectin puncta assay in the Results to explain the rationale and readout of sfGFP::LGALS3 puncta formation.

      Regarding the reviewer’s suggestion to compare young and aged animals, we agree that age-dependent changes in endolysosomal membrane integrity are an interesting question. However, the purpose of the present study was to investigate how perturbation of sphingolipid metabolism affects endolysosomal membrane fluidity and rupture under the assay conditions used in our original screen. A systematic comparison across aging is beyond the scope of the current revision. We have therefore clarified the animal ages used in the relevant figure legends and Methods.

      In Figure S1A, the authors show co-knockdown of multiple genes, including one condition with simultaneous RNAi against four targets. Because different RNAi clones can vary in knockdown efficiency, it is important to provide validation of gene knockdown levels (e.g., by qRT-PCR) shown in both panels a and b.

      We agree that RNAi efficiency can vary between clones and that this is particularly relevant for combinatorial RNAi experiments targeting two or more potentially redundant genes. We have added this limitation to the Results section and now state: “Because KD efficiency was not assessed for the individual RNAi clones or co-RNAi combinations, these experiments do not allow comparison of relative RNAi strength or inference of the relative importance of individual genes. Thus, the conclusions drawn from these RNAi experiments are qualitative: specific single or combined KDs can promote endolysosomal rupture, whereas the absence of a detectable phenotype after RNAi cannot exclude gene involvement, as KD may have been insufficient.”

      Where mutant strains were available, we performed genetic validation. Specifically, a sphk-1 mutant available at CGC (CZ24969; sphk-1(ju831)) also showed increased hypodermal sfGFP::LGALS3 puncta (new Figure S1A), supporting the RNAi-based conclusion that genetic perturbation of sphingolipid metabolism compromises endolysosomal integrity. Corresponding mutant strains were not available for the other selected hits. Importantly, most hits also induced sfGFP::LGALS3 foci in human HEK293T cells as assessed in our previous study [1], providing additional support that the observed effects are not random RNAi artifacts.

      In Figure 2E, the FRAP recovery curves show only ~60% recovery in controls after 25 seconds, and an even lower recovery (~40%) in hpo-8 and spp-10 RNAi conditions. The authors should discuss why the recovery is incomplete and what it implies about the mobile fraction of the protein or membrane components in these conditions.

      We agree that incomplete FRAP recovery is informative. For this reason, we report both the time to half-maximal recovery (thalf) and the maximal recoverable fluorescence signal. Increased thalf indicates reduced lateral mobility of LAAT-1::mCherry within the lysosomal membrane, consistent with reduced membrane fluidity. In addition, a reduced maximal recovery suggests that a larger fraction of the reporter is immobile or only slowly mobile during the time window analyzed. This may reflect stronger confinement of LAAT-1::mCherry within even more rigid membrane domains. However, because RNAi efficiency may differ between clones and we have not assessed their individual KD efficiency, we avoid overinterpreting differences in the absolute strength of recovery defects between individual KDs. Instead, we conclude that KD of sphingolipid-metabolism genes identified in our screen consistently reduces lysosomal membrane fluidity, as reflected by increased thalf and, in some cases, reduced maximal recovery.

      In Figure S3A, the Western blot for SPHK2 shows unequal loading between the control and siSPHK2 lanes. The blot should be normalized to a loading control and quantified to demonstrate knockdown efficiency.

      We have quantified SPHK2 levels relative to GAPDH across independent experiments and present the normalized quantification (Figure S3C-E).

      Key experimental details are missing from the manuscript. The strains of C. elegans and RNAi bacteria used were not described, and there is no information on biological replicates. The authors should clarify how many times each experiment was performed and provide more transparency on experimental reproducibility.

      We thank the reviewer for pointing this out. The C. elegans strains and RNAi bacterial clones used in this study were established and fully described in our previous study, which has now been published in Autophagy [1]. We now cite the published article throughout the revised manuscript and have added additional information in the Results section to explain the key features of the strains used here.

      We have also revised the Methods and figure legends to improve transparency regarding experimental details. The figure legends include the number of biological replicates, the number of animals or cells analyzed, and the statistical tests used for each experiment. In addition, the Statistical Analysis section in the Methods now summarizes how replicate numbers and sample sizes are reported across the study. Finally, the source details for the strains and RNAi clones used in this study are now provided in Tables S1 and S2, respectively. These revisions should improve the experimental clarity and reproducibility of the data shown.

      References:

      (1) Sandhof CA, Martin N, Tittelmeier J, Schlueter A, Pezzali M, Schoendorf DC, et al. A novel C. elegans model for MAPT/Tau spreading reveals genes critical for endolysosomal integrity and seeded MAPT/Tau aggregation. Autophagy. 2025;21(12):2963-81. Epub 20250904. doi: 10.1080/15548627.2025.2551676. PubMed PMID: 40851193; PubMed Central PMCID: PMCPMC12758218.

      (2) Yong J, Villalta JE, Vu N, Kukurugya MA, Olsson N, Lopez MP, et al. Impairment of lipid homeostasis causes lysosomal accumulation of endogenous protein aggregates through ESCRT disruption. eLife. 2024;12. Epub 20241223. doi: 10.7554/eLife.86194. PubMed PMID: 39713930; PubMed Central PMCID: PMCPMC11666243.

      (3) Calixto A, Chelur D, Topalidou I, Chen X, Chalfie M. Enhanced neuronal RNAi in C. elegans using SID-1. Nat Methods. 2010;7(7):554-9. doi: 10.1038/nmeth.1463. PubMed PMID: 20512143; PubMed Central PMCID: PMC2894993.

      (4) Li Y, Zhang J, Li M, Yang L, Wang X. Sphingosine kinase SPHK-1 maintains sphingolipid metabolism to protect lysosome membrane integrity in C. elegans. Mol Biol Cell. 2026;37(1):ar1. Epub 20251105. doi: 10.1091/mbc.E25-04-0182. PubMed PMID: 41191545; PubMed Central PMCID: PMCPMC12696880.

    1. eLife Assessment

      This study presents a fundamental finding that the JAK-STAT pathway (JSP) exerts context-dependent roles across distinct cellular compartments within the breast cancer microenvironment. The conclusions are supported by convincing evidence from multi-omics analyses. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

    2. Reviewer #2 (Public review):

      Summary:

      The JAK-STAT pathway (JSP) exhibits cell-type-specific functional heterogeneity in breast cancer. This study investigates the JSP in breast cancer and its response to anti-PD‑1 immunotherapy. JSP displays distinct cell‑type heterogeneity: it promotes malignant phenotypes and immunosuppression in tumor cells, while enhancing cytotoxicity and reducing exhaustion in T cells. Elevated JSP expression correlates with improved immunotherapy responses, especially in triple‑negative breast cancer. These findings highlight the paradoxical roles of JSP, indicating that broad inhibition may compromise anti‑tumor immunity.

      Strengths:

      The major strengths of this study include the comprehensive characterization JSP heterogeneity across epithelial, tumor, and T cells in breast cancer. The identification of JSP and STAT4 as predictive biomarkers for immunotherapy response, particularly in triple‑negative breast cancer, provides clinically relevant insights for patient stratification.

      Comments on revised version.

      The corresponding content has been revised.

    3. Reviewer #3 (Public review):

      Summary:

      This multi-omics study by Zhou et al elucidates the context-dependent roles of the Janus kinase-signal transducer and activator of transcription (JAK-STAT) pathway (JSP) across different cellular compartments in the breast cancer tumor microenvironment. While bulk JSP activity is associated with a favorable prognosis, single-cell analysis reveals a paradoxical landscape: high JSP in T cells drives anti-tumor cytotoxicity and reduces exhaustion, whereas high activity in tumor epithelial cells promotes malignancy and immunosuppression via the MIF-CD74 signaling axis. The JSP score (immune-related) serves as a robust predictive biomarker for response to anti-PD-1 immunotherapy, particularly in triple-negative breast cancer (TNBC). Furthermore, the study identifies the STAT4/SLC47A1 axis as a critical mechanism through which tumor cells resist ferroptosis, facilitating disease progression. These findings suggest that broad JAK-STAT inhibition may be counterproductive in cancer therapeutics; instead, therapeutic success depends on precise modulation and carefully timed interventions to preserve its T-cell-associated functions. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

      Strengths:

      Significant therapeutics implications

      Weaknesses:

      Limited molecular mechanisms

      Comments on revised version:

      The authors have addressed my comments

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In their manuscript, Zhou and colleagues present a detailed look at how the JSP functions differently in the various cells of a breast tumor. The authors have effectively shown that the JSP acts as a double-edged sword, as it helps T cells fight cancer but also allows tumor cells to grow and avoid ferroptosis. These findings are important because they identify a useful biomarker to predict how TNBC patients might respond to PD-1 inhibitors.

      Strengths:

      This work is important because it provides a clear explanation for the conflicting roles of the JSP in the tumor environment. The evidence is solid, as it combines data from thousands of patients with single-cell analysis and lab experiments to confirm the role of STAT4 in cancer progression and immunity.

      Comments on revised version:

      The authors made a significant effort to improve the manuscript. My comments were sufficiently addressed.

      We sincerely appreciate your careful review and positive feedback. We are glad to hear that you are satisfied with the revised manuscript and acknowledge the scientific value and solid evidence of our work. Thank you again for all your efforts and valuable suggestions.

      Reviewer #2 (Public review):

      Summary:

      The JAK-STAT pathway (JSP) exhibits cell-type-specific functional heterogeneity in breast cancer. This study investigates the JSP in breast cancer and its response to anti-PD‑1 immunotherapy. JSP displays distinct cell‑type heterogeneity: it promotes malignant phenotypes and immunosuppression in tumor cells, while enhancing cytotoxicity and reducing exhaustion in T cells. Elevated JSP expression correlates with improved immunotherapy responses, especially in triple‑negative breast cancer. These findings highlight the paradoxical roles of JSP, indicating that broad inhibition may compromise anti‑tumor immunity.

      Strengths:

      The major strengths of this study include the comprehensive characterization JSP heterogeneity across epithelial, tumor, and T cells in breast cancer. The identification of JSP and STAT4 as predictive biomarkers for immunotherapy response, particularly in triple‑negative breast cancer, provides clinically relevant insights for patient stratification.

      Weaknesses:

      The corresponding content has been revised.

      We sincerely thank you for your detailed review and valuable comments. We greatly appreciate your recognition of the cell-type-specific heterogeneity of the JAK-STAT pathway and the clinical value of JSP and STAT4 as predictive biomarkers for immunotherapy in triple-negative breast cancer. We have thoroughly revised the manuscript according to your previous suggestions, and all raised concerns have been fully addressed.

      Reviewer #3 (Public review):

      Summary:

      This multi-omics study by Zhou et al elucidates the context-dependent roles of the Janus kinase-signal transducer and activator of transcription (JAK-STAT) pathway (JSP) across different cellular compartments in the breast cancer tumor microenvironment. While bulk JSP activity is associated with a favorable prognosis, single-cell analysis reveals a paradoxical landscape: high JSP in T cells drives anti-tumor cytotoxicity and reduces exhaustion, whereas high activity in tumor epithelial cells promotes malignancy and immunosuppression via the MIF-CD74 signaling axis. The JSP score (immune-related) serves as a robust predictive biomarker for response to anti-PD-1 immunotherapy, particularly in triple-negative breast cancer (TNBC). Furthermore, the study identifies the STAT4/SLC47A1 axis as a critical mechanism through which tumor cells resist ferroptosis, facilitating disease progression. These findings suggest that broad JAK-STAT inhibition may be counterproductive in cancer therapeutics; instead, therapeutic success depends on precise modulation and carefully timed interventions to preserve its T-cell-associated functions. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

      Strengths:

      Significant therapeutics implications

      Weaknesses:

      Limited molecular mechanisms

      Comments on revised version:

      The authors have addressed my comments

      Many thanks for your careful evaluation and valuable suggestions. We highly appreciate your affirmation of the therapeutic significance of this study. We have fully revised the manuscript to enrich the molecular mechanisms, and all your comments have been properly resolved.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The most content has been revised.

      Minor corrections:

      (1) The icon about "prognosis" in graphic abstract is overly childish.

      The graphical abstract has been redrawn. The inappropriate prognosis icon is deleted accordingly.

      (2) Please double check the whole content to avoid typos. For instance, "2.2" and "2.3" have been repeated twice.

      We appreciate your reminder. The duplicate numbering of 2.2 and 2.3 resulted from Word’s automatic heading feature. We have disabled this function and fixed all repeated section numbers. In addition, we have carefully checked the full text and corrected all typos.

      (3) It will be more interesting if the oncogenic role of STAT4 could be verified via cell cloning assay.

      We appreciate your thoughtful comment. Considering the limited revision time, we cannot add the cell cloning assay in the current version. Our present data sufficiently validate the oncogenic function of STAT4, and the main conclusions remain reliable.

      Reviewer #3 (Recommendations for the authors):

      I recommend publishing the revised manuscript in eLife.

      We sincerely thank you for your positive evaluation and endorsement for the publication of our revised manuscript. We greatly appreciate your rigorous review and insightful comments that have substantially improved the quality and readability of this work.

      We sincerely appreciate all reviewers and editors for your thorough reviewing work and thoughtful feedback. Your suggestions have helped us greatly improve this manuscript. Thank you very much.

    1. eLife Assessment

      In this valuable study, the authors performed cell-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior (NT1) vs middle & posterior (NT2-9) cells of the C. elegans intestine, under fed, starved, or refeeding conditions. The data generated will be very helpful to the C. elegans community, and the evidence supporting the conclusions of the study is assessed to be solid. Some methodological caveats remain and are discussed.

    2. Reviewer #1 (Public review):

      Summary

      In this study, the authors have performed tissue-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior vs posterior cells of the C. elegans intestine. They have performed this analysis in fed and fasted states of the animal. The data generated will be very useful to the C. elegans community, and the role of pyruvate shown in this study will result in interesting follow-up investigations.

      However, several strong claims made in the study are solely based on in silico predictions and are not supported by experimental evidence.

      Comments on revised version.

      The authors have been responsive to the comments, but have not added new experiments in this manuscript that would have clarified and improved some of the mentioned shortcomings of the study.

      There are 3 comments that the authors should address:

      (1) In their response to reviewers, the authors agree that "the Pges-1deltaB promoter is not absolutely restricted to INT1 and that weak GFP expression can also be detected in INT2." They also mention that "because Pges-1deltaB is an engineered promoter derived from the intestine-specific Pges-1 promoter, this low-level INT2 expression is not unexpected." However, in line 93 of the revised manuscript, the authors claim that "Pges-1deltaB is strictly expressed in INT1 cells". This discrepancy should be fixed. They should instead describe this in line 93 as "Pges-1deltaB expression is very strongly enriched in INT1 cells, but low-level expression in INT2 was also detected".

      (2) In response to reviewers, the authors explained that "Our model is that fasting induces INS-7 secretion by lowering intracellular pyruvate in INT1 cells. Under this framework, blocking mitochondrial pyruvate breakdown would be expected to reduce pyruvate utilization and thus maintain intracellular pyruvate, preventing the drop in pyruvate that normally occurs during fasting. This would explain why these manipulations suppress fasting-induced INS-7 secretion." However, the effect of blocking import of pyruvate from cytosol into mitochondria (via knockdown of mitochondrial pyruvate carrier genes mpc-1 and mpc-2) does not agree with their proposed model. Blocking mitochondrial import of pyruvate should maintain cytosolic pyruvate levels and thus prevent the drop in pyruvate that normally occurs during fasting. In such a scenario, we would expect to see no increase in INS-7 secretion during fasting, which is opposite to the result in Fig.7D. If the pyruvate sensor is in the cytosol, we would expect that the mpc-1/2 RNAi treated animals would be unable to increase INS-7 secretion upon starvation. If the pyruvate sensor is in the mitochondrial matrix, we would expect that the mpc-1/2 RNAi treated animals would have higher INS-7 secretion than vector RNAi control animals in fed conditions. How do the authors explain this discrepancy between their observed results and their proposed model? Why does blocking mitochondrial import of pyruvate affect only refeeding-induced reduction in INS-7 secretion but not fasting-induced increase in INS-7 secretion? Is it possible that instead of responding to absolute intracellular concentrations of pyruvate, the pyruvate sensor increases INS-7 secretion upon detecting a relative drop in the mitochondrial levels of pyruvate (or its downstream metabolite)? This should be described in the text to better interpret the mpc-1/2 RNAi results.

      (3) Line 493: The authors refer to 'Table S4', which is not included in the manuscript.

    3. Reviewer #3 (Public review):

      In this study, Liu and colleagues utilize TRAP-seq to profile the repertoire of actively translated mRNAs in different intestinal cell types (anterior INT1 vs. posterior INT2-9 cells) in C. elegans. A key goal of this study was to identify transcripts differentially expressed/translated between these intestinal cell subtypes in the context of animals being well fed or subjected to acute (30 minutes) or chronic (3 hours) starvation, followed by refeeding.

      The authors identify a number of differentially expressed genes across all of the conditions tested. They then provide an initial survey of the landscape of translatome changes through Weighted Gene Network Correlation Analysis (WGNA), and some high-level functional surveys via Gene Ontology (GO) term analysis and protein domain analysis. The authors validate the enriched expression patterns of some of their identified candidate genes using fluorescent promoter fusion reporters, confirming INT1-specific expression. The authors further implicate the role of several other candidate genes in pathogen avoidance and in response to nutritional cues by knocking them down specifically in INT1 cells by RNAi. Finally, the authors identify pyruvate as a major nutrient signal coming from the bacterial diet that suppresses the release of a key insulin peptide (INS-7) and identify some of the genes expressed in INT1 that are required for this response.

      Strengths:

      (1) Good use of and justification for TRAP-seq, because scRNA-seq would be difficult under the varied conditions used (starvation, refeeding)

      (2) The manuscript is generally clear to read, and the data are generally well-presented with good supporting data that includes replicates, sample sizes, error measurements, and associated statistics.

      (3) The dataset will be an interesting resource to mine for future studies focusing on mechanisms of how particular intestinal cell types respond to different environmental signals.

      Weaknesses:

      (1) A limitation of TRAP-seq, although powerful, is that only relative comparisons can be made between genotypes/conditions to identify differentially-expressed genes, rather than assessing whether a given gene is expressed at a certain level in a cell type under a certain condition. This limitation is due to the non-specific association of sticky RNA species to the beads during the immunoprecipitation step. This is a minor point however, and the authors do a nice job of focusing their analysis on differentially expressed transcripts in the current study.

      (2) Another limitation of the current study is that the experiments testing the role of candidate genes identified by their profiling experiments do not dive a bit deeper into providing a mechanistic understanding of the phenotypes being studied. At present, the results are thus viewed more as a genomics-based screen with some limited follow-up on interesting hits. However, this reviewer appreciates that when placed in context of the work presented, a presentation of the profiling data along with some validation is an excellent starting point for future mechanistic studies elaborating on these interesting candidates.

      Appraisal of whether the authors achieved their aims, and whether the results support their conclusions.

      The main goal of the study was to survey the dynamic responses at the level of actively translated mRNAs of the INT1 vs INT2-9 cells in response to metabolic challenge.

      Overall, the authors use established methods to perform their genome-wide analysis, and the set of differentially regulated genes are enriched for expected molecular functions and form coherent networks in anticipated pathways.

      The validation experiments (promoter::GFP fusion reporters, INT1-specific knockdowns of highly regulated genes) further corroborate the quality of the TRAP-seq datasets generated.

      I have a few points for the authors that would further strengthen this work:

      (1) The authors rightfully focus on the top differentially-regulated candidates, but it's unclear at present how far down their fold change list would lead to expression pattern validations. It would be useful to test a few more promoter::GFP fusion reporters at different enrichment/fold-change/statistical cutoffs.

      (2) Although the INT1-specific RNAi provides a convenient strategy for rapidly perturbing and testing genes of interest for phenotypes, independently validating the knockdowns with genetic mutants, or alternatively (if genes are essential), degron alleles.

      Likely impact of the work on the field, and the utility of the methods and data to the community.

      The TRAP-seq data and list of differentially-expressed candidate genes will form an interesting set of high-priority candidates to study for their role in the reception and transduction of nutritional cues in response to food status and pathogens. This data will thus benefit the C. elegans community of researchers studying the mechanisms governing these phenomena.

      Comments on revised version:

      I think the authors have done a good job of addressing the suggestions from the previous round of review in this new version.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      In this study, the authors have performed tissue-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior vs posterior cells of the C. elegans intestine. They have performed this analysis in fed and fasted states of the animal. The data generated will be very useful to the C. elegans community, and the role of pyruvate shown in this study will result in interesting follow-up investigations.

      However, several strong claims made in the study are solely based on in silico predictions and are not supported by experimental evidence.

      Strengths:

      Several studies in the past have predicted different functions of the anterior (INT1) vs posterior (INT2-9) epithelial cells of the C. elegans intestine based on their anatomy and ultrastructure, but detailed characterization of differences in gene expression between these cell types (and whether indeed these are different 'cell types') was lacking prior to this study. The genes and drivers identified to be exclusively expressed in the anterior vs posterior segments of the intestine will be very helpful to selectively modulate different parts of the C. elegans intestine in future studies.

      Another strength of this study is the careful experimental design to test how the anterior vs posterior cell types of the intestine respond differently to food deprivation and recovery after return to food. These comparisons between 'states' of a cell in different physiological conditions are difficult to pick up in single-cell analyses due to low sequencing depth, which can fail to identify subtle modulation of gene expression.

      The TRAP-associated bulk RNA-seq approach used in this study is more suitable for such comparisons and provides additional information on post-transcriptional regulation during metabolic stress.

      A key finding of this study is that pyruvate levels modulate the translation state of anterior intestinal cells during fasting. Characterization of pyruvate metabolism genes, especially of the enzymes involved in its mitochondrial breakdown, provides novel insights into how gut epithelial cells respond to the acute absence of food.

      Weaknesses:

      Unlike previous TRAP-seq studies (PMID: 30580965, 36044259, 36977417) that reported sequencing data for both input and IP samples, this study only reports the sequencing data for IP samples. Since biochemical pulldowns are variable across replicates, it is difficult to know if the observed differences between different conditions are due to biological factors or differences in IP efficiency. More importantly, since two different TRAP lines were utilized in this study and a large proportion of the results focus on the differences between the translational profiles of INT1 vs INT2-9 cells, it is essential to know if the IP worked with similar efficiency for both TRAP strains that likely have different expression levels of the HA-tagged ribosomal protein. One way to estimate this would be to perform qRT-PCR of genes that are known to be enriched in all intestinal cells and determine whether their fold-enrichment over housekeeping genes (normalized to input) is similar in INT1 vs INT2-9 TRAP strains and across the fed vs fasted conditions. The authors, in fact, mention variability across biological replicates, due to which certain replicates were excluded from their WGCNA analysis.

      We appreciate the reviewer's comments. We agree that the lack of matched input sequencing libraries limits our ability to directly assess IP efficiency across replicates, conditions, and TRAP strains. However, several features of the dataset support the conclusion that the major differences reported here reflect biological rather than purely technical variation. First, the RPL-22-3xHA construct was integrated into each line to improve consistency across experiments. Second, although the INT2-9 TRAP strain yielded more RNA than the INT1 strain, as expected given the larger number of labeled cells, downstream analyses were performed on normalized count data rather than raw counts. Third, principal component analysis showed robust separation by promoter identity across all conditions, and expected INT1-enriched genes such as ins-7 were recovered in the INT1 dataset. Together, these observations support the interpretation that the TRAP datasets capture reproducible, cell-type-specific differences in ribosome-associated transcripts. Nonetheless, we agree that direct input-normalized measurements would further strengthen the study, and we will explicitly note this as an important limitation.

      It appears that GFP expression is also detectable in INT2 (in addition to strong expression in INT1 in Fig.1A). Compared to INT3-9, which looks red, INT2 cells appear yellow, suggesting that the expression patterns of the two TRAP drivers are not mutually exclusive, which changes the interpretation of many of the results described in the study.

      We agree that the Pges-1ΔB promoter is not absolutely restricted to INT1 and that weak GFP expression can also be detected in INT2. Because Pges-1ΔB is an engineered promoter derived from the intestine-specific Pges-11 promoter, this low-level INT2 expression is not unexpected. However, we note that the expression level in INT1 is substantially higher than in INT2. Thus, although the expression patterns of the two TRAP drivers are not completely mutually exclusive, Pges-1ΔB still provides the most selective available tool for enriching the INT1 translatome in the context of the current study.

      Some parts of the study overemphasize the differences between the INT1 vs INT2-9 cell types, which is a biased representation of the results. For example, the authors specifically point out that 270 genes are differentially expressed in opposite directions in INT1 vs INT2-9 cell types during acute (30 min) fasting without mentioning the 1,268 genes that are differentially expressed in the same direction. They also do not mention here that 96% of the genes are differentially expressed in the same direction in INT1 and INT2-9 cell types after prolonged (180 min) fasting, suggesting that the divergent translational responses of these cell types are only observed in the first 30 minutes of food deprivation. Similar results have also been reported for the effect of fasting on locomotory and feeding behaviors, where 30 min of fasting produces more variable effects, which become more consistent after longer periods of fasting (PMID: 36083280). Hence, the effects of brief food deprivation should be interpreted with caution.

      The intestine functions as a discrete and cohesive organ, so the expected result is that there would be no differences across the different cell types. For us, the surprise was that, in fact, there are differences between these cells at all. However, the point is well taken, and we have added a statement in the text to reflect that many genes change similarly in INT1 and INT2-9, while the differences reflect important functional divergence between these cell types.

      Many of the interpretations of this study primarily rely on pathway enrichment analyses, which are based on the known function of genes. The function of uncharacterized genes that were found to be differentially expressed in INT1 vs INT2-9 cell types, e.g., the ShKT proteins, was not explored in this study. In addition, overreliance on pathway enrichment tools (instead of functional validation) has resulted in several conflicting findings. For example, one of the main messages of this study is that INT1 cells specialize in immune and stress response in response to fasting, which relies on pathway analysis in Figs 5E and 5F. However, pathway analysis at a different time point (shown in Figure S5A) indicates that INT2-9 cells show a much stronger increase in translation of stress and pathogen-responsive genes compared to INT1 cells. Hence, some of the results should be interpreted as different translational effects in INT1 vs INT2-9 cells after different lengths of food deprivation, without making broad claims about selective pathways being affected only in specific cell types.

      We agree that some interpretations in the manuscript relied heavily on pathway enrichment analyses and should be stated more cautiously. In particular, we agree that the current data are most consistent with state-dependent differences in translational responses between INT1 and INT2-9 cells across different durations of food deprivation, rather than with the strongest version of a claim that specific pathways are selectively engaged only in one intestinal subset. We also agree that uncharacterized genes, including the ShKT family, were not mechanistically explored in the present study and should be presented as important candidates for future investigation.

      The authors have compared their TRAP-seq results with genes enriched in the anterior and posterior intestine clusters from a previously published whole-animal adult scRNA dataset (PMID: 37352352). They claim that their TRAP-seq results are in agreement with the findings of the scRNA study. However, among the 10 genes from the 'posterior intestine' scRNA cluster in Fig.S1E, six are downregulated in the INT1 vs INT2-9 comparison, while four are upregulated. Hence, there is no clear agreement between the two studies in terms of the top enriched genes in the anterior vs posterior intestine, which should be considered for cross-study comparisons in the future.

      We have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. We instead assessed the expression levels of our INT1 up-regulated genes in their intestinal cluster and found that they have higher expression in the anterior intestine cluster (new Figure S1C). These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9.

      The authors describe in the manuscript that they have performed INT1-specific RNAi for two C-type lectin genes that are upregulated during fasting. Due to a recent expansion of C-type lectin genes in C. elegans, there is a high chance of off-target effects of RNAi that is designed for members of this gene family. More trustworthy results could have been obtained using CRISPR-based loss-of-function alleles for these genes, one of which is publicly available. Also, the authors do not provide any explanation for why knockdown of these stress-response genes, which are activated in INT1 cells in response to food deprivation, results in improved resistance to pathogens. This, in fact, suggests a role of INT1 cells in increasing pathogen susceptibility, and not pathogen resistance, during food deprivation.

      We agree that RNAi targeting C-type lectin family members may be susceptible to off-target effects, and that validation with CRISPR null alleles, where available, would strengthen these findings. In the current study, we used INT1-specific RNAi as a cell-specific first-pass approach to test candidate gene function. We also agree that the pathogen phenotype requires cautious interpretation. Specifically, the finding that knockdown of fasting-induced INT1 lectin genes improves pathogen resistance does not support a simple protective model for these genes. Instead, it suggests that INT1-expressed stress-response genes modulate host susceptibility or host-pathogen interactions.

      Many of the studies in this field (e.g., references 2-4 in this article) have investigated the effects of food deprivation ranging from 4 hr to 24 hr, which results in activation of starvation responses in C. elegans. In contrast, the authors have used shorter time periods of fasting (30 min and 180 min), and most of their follow-up experiments have used 30 min of food deprivation. Previous work has shown that the effects of food deprivation can either accumulate over time (i.e., the effect gets stronger with longer food deprivation) or can be transient (i.e., only observed briefly after removal of food and not observed during long-term food deprivation). Starvation-induced transcription factors such as DAF-16/FoxO and HLH-30 show strong translocation to the nucleus only after 30 min of fasting. Though gene expression changes in all stages of food deprivation are of biological relevance, the authors have missed the opportunity to explore whether increased INS-7 secretion from the anterior intestine is dependent on these starvation-induced transcription factors (which can be easily tested using loss-of-function alleles) or is due to other fast-acting regulatory mechanisms induced due to the absence of food contents in the gut lumen. A previous study (PMID: 40991693) has shown that DAF-16 activation during prolonged starvation shuts down insulin peptide secretion from the intestinal epithelial cells. Hence, it is not clear if increased INS-7 secretion is only a feature of short-term food deprivation or is also a signature of long-term starvation (e.g., at 8 hr or 16 hr timepoints). Since most of the INS-7 secretion data in this study are for 30 min of fasting, it remains unknown whether the discovered regulators of INS-7 secretion can be generalized for extended food deprivation that triggers major metabolic changes, such as fat loss (e.g., conditions shown in Figure 1D).

      We agree that short-term food deprivation and prolonged starvation likely engage distinct regulatory mechanisms, and that our study primarily addresses an early phase of food deprivation rather than the full spectrum of starvation responses described in prior work. We selected the 30 min fasting condition because our previous study showed that INS-7 secretion is induced within this interval and returns to baseline upon refeeding, even before detectable intestinal fat loss. We also included a 180 min fasting condition to capture a later state associated with metabolic changes. However, we agree that the present study does not determine whether the regulators of INS-7 secretion identified here also govern secretion during more prolonged starvation (for example, 8 hr or 16 hr), nor does it test whether starvation-responsive transcription factors such as DAF-16 or HLH-30 contribute to this regulation. We appreciate that determining how this response transitions during prolonged starvation will be an important direction for future work.

      Two previous studies (PMID: 18025456, 40991693) have shown a strong reduction in the expression of ins-7 in the anterior intestine using GFP-based reporters (both promoter fusions and endogenous CRISPR-generated) and in whole-animal RNA-seq data from starved animals. These results are in contrast to the increased INS-7 secretion from INT1 cells during fasting that is reported in this study. The authors here have reported that INS-7 translation is higher in INT1 compared to INT2-9 during fed, acute fasted, and chronic fasted conditions, but they have not shown whether INS-7 translation is upregulated during acute and chronic fasting in INT1 cells in their TRAP-seq analysis. Knowing whether increased INS-7 secretion during acute fasting is due to increased transcription, translation, or secretion of INS-7 is crucial to resolve the discrepancy between these studies.

      In our dataset, INS-7 translation in INT1 tended to increase during acute fasting relative to the fed state (log<sub>2</sub>FC = 0.69), although this effect did not reach statistical significance after adjustment for multiple comparisons. Consistent with this trend, our secretion assay showed that INS-7 release from INT1 increases during fasting. However, we agree that the current data do not distinguish whether this increase in secretion is driven by enhanced synthesis, regulated release of pre-existing peptide stores, or a combination of both.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to understand whether the discrete segments of the C.elegans intestine were specialized to carry out distinct functions during an animal's exposure and adaptation to a fast-changing nutrient environment. To achieve this, the authors used a method called Translating ribosome affinity purification (TRAP), which provides a snapshot of what genes are being translated into proteins (and therefore functionally prioritized by the animal) under different fasting and re-feeding conditions. By expressing the TRAP constructs in two distinct segments of the intestine (INT1) and (INT2-9), the authors were able to identify how these segments responded to changing nutrient availability.

      Already under steady state nutrient conditions, the authors found that INT1 and INT2-9 appeared to have different 'tasks', with INT1 expressing more immune- and stress-response related genes. Exposing animals to different regimens of starvation and refeeding also showed marked differences between the intestinal segments, and the gene expression patterns in INT1 were consistent with INT1 cells playing an integrative role in linking nutrient cues to the secretion of insulin molecules that regulate fat metabolism with food intake. In summary, the data presented catalogue, for the first time, gene expression differences between two areas of the intestine, suspected to play different roles, and through clever experiments, links these gene expression changes to responses to nutrient availability.

      Strengths:

      The data presented catalogue - for the first time and in a careful manner - gene expression differences between two areas of the intestine. They strongly support the presence of intriguing differences between two areas of the intestine in immune, metabolic, and stress-response regulation, and link these gene expression changes to the responses of these regions to nutrient availability.

      Weaknesses:

      The conclusions of this paper are mostly well-supported by data, but the relevance of the changing gene expression patterns could be better clarified and extended in the discussion.

      We thank the reviewer for this constructive comment. In the revised manuscript, we have now expanded the discussion to more clearly interpret these dynamic translatomic changes in the context of intestinal subset specialization. The most pronounced difference between INT1 and INT2-9 cells is the enrichment of stress-response genes in INT1. Based on the present findings, together with our previous work identifying INS-7 as an INT1-secreted signal (PMID: 39127676), we propose that INT1 cells are sentinel enteroendocrine cells that integrate information from the luminal environment and the metabolic state of intestinal cells.

      Reviewer #3 (Public review):

      Summary:

      In this study, Liu and colleagues utilize TRAP-seq to profile the repertoire of actively translated mRNAs in different intestinal cell types (anterior INT1 vs. posterior INT2-9 cells) in C. elegans. A key goal of this study was to identify transcripts differentially expressed/translated between these intestinal cell subtypes in the context of animals being well fed or subjected to acute (30 minutes) or chronic (3 hours) starvation, followed by refeeding.

      The authors identify a number of differentially expressed genes across all of the conditions tested. They then provide an initial survey of the landscape of translatome changes through Weighted Gene Network Correlation Analysis (WGNA), and some high-level functional surveys via Gene Ontology (GO) term analysis and protein domain analysis. The authors validate the enriched expression patterns of some of their identified candidate genes using fluorescent promoter fusion reporters, confirming INT1-specific expression. The authors further implicate the role of several other candidate genes in pathogen avoidance and in response to nutritional cues by knocking them down specifically in INT1 cells by RNAi. Finally, the authors identify pyruvate as a major nutrient signal coming from the bacterial diet that suppresses the release of a key insulin peptide (INS-7), and identify some of the genes expressed in INT1 that are required for this response.

      Strengths:

      (1) Good use of and justification for TRAP-seq, because scRNA-seq would be difficult under the varied conditions used (starvation, refeeding).

      (2) The manuscript is generally clear to read, and the data are generally well-presented with good supporting data that includes replicates, sample sizes, error measurements, and associated statistics.

      (3) The dataset will be an interesting resource to mine for future studies focusing on mechanisms of how particular intestinal cell types respond to different environmental signals.

      Weaknesses:

      (1) A limitation of TRAP-seq, although powerful, is that only relative comparisons can be made between genotypes/conditions to identify differentially-expressed genes, rather than assessing whether a given gene is expressed at a certain level in a cell type under a certain condition. This limitation is due to the non-specific association of sticky RNA species with the beads during the immunoprecipitation step. This is a minor point, however, and the authors do a nice job of focusing their analysis on differentially expressed transcripts in the current study.

      We agree that a limitation of TRAP-seq is that it is best suited for relative comparisons across cell types or conditions, rather than for determining the absolute expression level of a given transcript in a specific cell type. As the reviewer notes, this limitation arises in part from nonspecific recovery of background or sticky RNAs during the immunoprecipitation step, complicating the interpretation of absolute expression levels. For this reason, our analysis was designed to focus primarily on differentially enriched transcripts between INT1 and INT2-9 cells and across feeding states, rather than on assigning absolute expression levels to individual genes. We appreciate the reviewer’s recognition of this point. Our study uses TRAP-seq specifically to define relative translatomic differences between intestinal subsets and physiological states, which is well aligned with the strengths of this approach.

      (2) Another limitation of the current study is that the experiments testing the role of candidate genes identified by their profiling experiments do not delve a bit deeper into providing a mechanistic understanding of the phenotypes being studied. At present, the results are thus viewed more as a genomics-based screen with some limited follow-up on interesting hits. However, this reviewer appreciates that when placed in the context of the work presented, a presentation of the profiling data along with some validation is an excellent starting point for future mechanistic studies elaborating on these interesting candidates.

      We agree that the current study does not fully resolve the molecular mechanisms by which the candidate genes identified by TRAP-seq regulate the phenotypes examined here. Our primary goal was to generate a spatially resolved translatomic framework for INT1 and INT2-9 cells across feeding states, and to perform focused validation of selected candidates to establish the physiological relevance of the profiling results. We therefore view the current functional analyses as an initial validation and proof of principle, rather than a comprehensive mechanistic dissection of the molecular pathways for each candidate. We appreciate the reviewer’s recognition that these findings provide an excellent starting point for future studies.

      Appraisal of whether the authors achieved their aims, and whether the results support their conclusions:

      The main goal of the study was to survey the dynamic responses at the level of actively translated mRNAs of the INT1 vs INT2-9 cells in response to metabolic challenge.

      Overall, the authors use established methods to perform their genome-wide analysis, and the set of differentially regulated genes is enriched for expected molecular functions and forms coherent networks in anticipated pathways.

      The validation experiments (promoter::GFP fusion reporters, INT1-specific knockdowns of highly regulated genes) further corroborate the quality of the TRAP-seq datasets generated.

      I have a few points for the authors that would further strengthen this work:

      (1) The authors rightfully focus on the top differentially-regulated candidates, but it's unclear at present how far down their fold change list would lead to expression pattern validations. It would be useful to test a few more promoter::GFP fusion reporters at different enrichment/fold-change/statistical cutoffs.

      Testing additional promoter::mNeonGreen reporters across a wider range of fold-change and statistical thresholds could be somewhat useful for calibrating ranked TRAP-seq candidate genes. However, given the variation in strains bearing extrachromosomal arrays, we did not consider this a stringent enough test, given that the sensitivity and dynamic range of RNA-seq far outpaces genetic fluorescence-based reporters. For these reasons, we focused on the top differentially enriched candidates to provide not only an initial validation of the dataset, but also to determine whether these candidates regulate biological functions in INT1 cells and thus serve as potentially useful biological readouts in future efforts.

      (2) Although the INT1-specific RNAi provides a convenient strategy for rapidly perturbing and testing genes of interest for phenotypes, independently validating the knockdowns with genetic mutants, or alternatively (if genes are essential), degron alleles.

      We agree that validating the INT1-specific RNAi phenotypes with independent genetic approaches, including null-allele or degron-based alleles for essential genes, would further strengthen the conclusions. In the current study, we used INT1-specific RNAi as a rapid and spatially restricted strategy to functionally test candidates identified by TRAP-seq and to determine whether these genes contribute to the specialized physiological functions of INT1 cells. We consider these experiments an initial validation of candidate function rather than a complete genetic dissection, which could be conducted in future efforts to study other aspects of INT1 function.

      Impact:

      The TRAP-seq data and list of differentially-expressed candidate genes will form an interesting set of high-priority candidates to study for their role in the reception and transduction of nutritional cues in response to food status and pathogens. This data will thus benefit the C. elegans community of researchers studying the mechanisms governing these phenomena.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments:

      (1) The authors need to describe the fasting method used in detail. Was fasting performed on unseeded NGM plates or in liquid (M9 buffer)? Were the animals washed with buffer prior to starvation? If yes, how many times? These details are critical for any researcher to follow up on their results.

      We have clarified the fasting/refeeding procedure in the revised Methods section. Briefly, worms were washed off OP50-seeded NGM plates with M9 buffer, washed three times in M9 buffer, and then transferred to unseeded NGM plates for fasting. For refeeding, worms were collected from the unseeded NGM plates with M9 buffer and transferred back to OP50-seeded NGM plates.

      (2) The authors claim that "INT1 and INT2-9 cells maintain fundamentally different molecular identities independent of any and all acute or chronic conditions", which they primarily based on Principal Component Analysis (PCA). The circles shown in Figure 2A are arbitrary, and many such circles can be drawn in the 2D space to separate the samples in different ways. The authors should show this comparison in a translatome-wide similarity heatmap with hierarchical clustering (similar to Fig.2C, but with all the experimental conditions and their replicates on both x- and y-axes).

      The ellipses shown in Figure 2A were generated using the stat_ellipse() function in ggplot2, which calculates the mean and covariance of the PC1 and PC2 for each line and draws ellipses corresponding to the 95% confidence level. The separation between lines is primarily driven by PC2, which accounts for 14% of the variance in the translatomic dataset. Although this difference is less pronounced when considering the full translatome, samples from the same line nevertheless cluster together, supporting line-specific differences in translatomic profile.

      (3) Figure 4 of the study shows 18 Venn diagrams for genes that are differentially expressed between INT1 and INT2-9 cell types in fed, fasted, and refed conditions. In the absence of any statistical comparisons, it is difficult to interpret whether the extents of overlap (higher or lower than expected) are significant. Ideally, P values for hypergeometric tests should be provided for the overlap regions.

      We appreciate the reviewer’s suggestion. We explored the use of hypergeometric testing, implemented through the SuperExactTest package in R, to assess the statistical significance of the overlaps shown in the Venn diagrams. However, this analysis yielded significant P values for essentially all overlap regions, including cases in which the degree of overlap was not especially informative and did not align with the interpretation presented in the text. This outcome likely reflects the dependence of the test on the size of the input gene sets and background universe, which can make statistical significance difficult to interpret meaningfully in this context.

      (4) The claims made in lines 242-244 (Figure 5D) need to be supported by P values from hypergeometric tests.

      Similar to the previous point.

      (5) The interpretation of Figures 6E and 6F described in lines 296-297 needs to be supported by statistical analyses. The authors claim that the undulating pattern of expression of the turquoise module genes is stronger in INT1 compared to INT2-9. However, based on Figures 2C and 6F, it appears that the expression change is not necessarily weaker in INT2-9, but instead is different, i.e., the expression of turquoise module genes goes up during fasting in INT1 and goes down after refeeding, while their expression goes up during fasting and stays up after refeeding in INT2-9 cells.

      We appreciate the reviewer’s point and agree that the turquoise module shows dynamic regulation in both cell populations. The key difference is not the presence versus absence of an undulating pattern, but rather the magnitude of that change, which is greater in INT1. Because of the limited number of biological replicates in some conditions, particularly the fasting group, we interpreted these results cautiously and used a nonparametric approach to assess differences in average module expression between states. This analysis indicated that the turquoise module changes significantly in both lines, but with a larger effect size in INT1. We have included the corresponding statistical analysis and effect size in the revised manuscript.

      (6) Since the INS-7 coelomocyte uptake assay was used extensively in this study, some representative microscopy images should be included to complement the quantification.

      We have added a new Figure 6G showing representative images corresponding to the quantification presented in Figure 6H.

      (7) The authors claim that INT1-specific fmo-2 RNAi results in reduced basal INS-7 secretion, but they do not have the direct statistical comparison for this. Were experiments shown in Figures 6G and 6K done on the same day?

      In the original Figure 6K (now Figure 6L), the data are presented as the percentage of normalized INS-7::mCherry fluorescence intensity relative to fed animals treated with vector RNAi. A statistical comparison between fed animals treated with INT1-specific fmo-2 RNAi and fed vector RNAi controls was performed and was significant. We have also clarified that the experiments shown in the original Figures 6G and 6I (now Figures 6H and 6J) were performed on the same day.

      (8) It is not clear why blocking the mitochondrial breakdown of pyruvate (Figures 7E and 7F) does not mimic the fasted state in terms of increased INS-7 secretion from INT1 cells. Doesn't this contradict the proposed model in this study? Can the authors speculate why this is the case?

      We do not interpret inhibition of pyruvate dehydrogenase or pyruvate carboxylase as equivalent to the fasted state. Rather, our model is that fasting induces INS-7 secretion by lowering intracellular pyruvate in INT1 cells. Under this framework, blocking mitochondrial pyruvate breakdown would be expected to reduce pyruvate utilization and thus maintain intracellular pyruvate, preventing the drop in pyruvate that normally occurs during fasting. This would explain why these manipulations suppress fasting-induced INS-7 secretion. To directly examine this possibility, we performed the experiment in Figure 7G, which tests whether maintaining pyruvate levels in INT1 cells during fasting is sufficient to suppress INS-7 secretion. The results are consistent with this interpretation and further support a model in which decreased intracellular pyruvate is a key determinant of fasting-induced INS-7 secretion.

      Minor comments:

      (1) Figures 2E and 2G are very similar and represent the same result in two different ways (unbiased vs guided comparison). One of these should be moved to the supplementary figures.

      Although these figures show similar patterns, they were derived from two independent analytical approaches, WGCNA and differential expression analysis. We therefore interpret the concordance between these independent methods as strengthening the robustness of the association and increasing confidence in the biological relevance of the observed pattern.

      (2) In Figures S1C, S1D, and S1E, a more significant P-value is shown with a smaller circle, and a less significant P-value is shown with a larger circle. This is confusing to the reader and should be inverted.

      We appreciate the reviewer’s comment and have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes used in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. Our further examination of marker expression across the intestinal clusters in the Ghaddar et al. (PMID: 37352352). dataset confirmed this limitation. These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9. We also note that the spatial identities of the intestinal clusters in Ghaddar et al. (PMID: 37352352) were not clearly established in the text or by spatial transcriptomic evidence, making it difficult to assign the annotated anterior, middle, and posterior clusters to specific intestinal cells. We have revised the manuscript accordingly and replaced the original figure panels.

      (3) Line 131 mentions the comprehensive characterization of the translatomic differences between INT1 and INT2-9 cells under each acute and chronic condition. However, the paragraph only discusses the differences in the fed condition. This is confusing, and the authors should mention the comparison between these cell types under acute and chronic conditions in subsequent sections where it is described.

      In this paragraph, we indeed discuss the ‘fed’ condition, but in subsequent sections we follow with details analyses of acute versus chronic, as well as regional differences across the intestine. We have clarified this in the opening sentence of the referenced paragraph.

      (4) The Venn diagrams in Figure 4 look very similar, and it is hard to differentiate between how 4A is different from 4I, how 4B is different from 4J, etc. The authors should include the labels for 'acute' or 'chronic' above each Venn diagram to guide the reader through these panels.

      We have added labels indicating the acute and chronic conditions to the left side of each Venn diagram in the revised Figure 4.

      (5) It is not clear in the figure legends how Figure 5E is different from Figure S6A, and how Figure 5F is different from Figure S7A. This should be better described in the figure legends.

      We have added a sentence to better describe this in the figure legend.

      (6) Figure 6C: Survival parameters such as median lifespan, number of animals for each condition, etc., should be reported for the different conditions.

      We have revised Figure 6C to indicate the number of animals analyzed in each condition, and the median survival for each group is now reported in the corresponding figure legend.

      (7) The colors used for control RNAi and clec-160 RNAi are very similar in Fig.6C. Easily distinguishable colors should be used.

      We have changed the colors as suggested.

      (8) The INT1-specific RNAi strain should be first described in line 285.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (9) Line 304: 'REF' should be replaced with the reference.

      We have replaced the “REF” with the reference (PMID: 39127676)

      (10) The P value for statistical comparison between the fed and 30 min refed states should be shown in Figures 6G, 6I, and 6K.

      We have now included the p value for the comparison as suggested. Figures 6G, 6I, and 6K are now labeled as 6H, 6J, and 6L, respectively.

      (11) In Figure 7, the authors should consider replacing the 'refed' label with 'recovery' because the pyruvate treatment was done in the absence of 'feeding' (= bacteria consumption).

      We appreciate the reviewer’s point. However, we chose to retain the label “refed” in Figure 7 to maintain consistency across the set of conditions examined, including 2% glucose and OP50 supernatant, which likewise do not involve bacterial consumption despite not showing effect on the refeeding response of INS-7 secretion.

      (12) The full form of DISN should be mentioned in the figure legend of Figure 7.

      We have included the full form of D1SN in the figure legend of Figure 7A.

      (13) Line 367: 'normalization' should be replaced with 'return to basal levels'. 'Normalization of INS-7 secretion' might also mean normalization of INS-7::mCherry signal to CLM::GFP signal.

      We have revised the wording per the reviewer's suggestion.

      (14) The methods section has a quantitative RT-PCR section, but it is not clear if RT-PCR data are reported in any of the figures. Also, no qPCR primers are listed in Table S3.

      We have removed the quantitative RT-PCR part from the methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors describe that the RPL-22-3xHA constructs are not integrated, at the very end, in the section "Limitations of the data". An earlier mention of this caveat would have been useful. In addition, it would help if the authors could provide their defense (which I think is very valid) of using non-integrated strains in the results section, as they describe the experimental setup. Also, some details were missing, which left me wanting to know: Were there expression differences? How were they accounted for? Was expression normalized between these two constructs, and if so, how?

      The RPL-22-3xHA construct was integrated into each line to ensure more consistent transgene expression across experiments. Because the INT2-9 construct is expressed in a larger number of cells than the INT1 construct, the INT2–9 samples yielded greater amounts of pulled-down nascent RNA, as reflected in the supplemental table and in the higher aligned RNA counts observed for the INT2-9 samples. To account for these differences, differential expression analysis was performed using DESeq2, which corrects for library size by estimating sample-specific size factors with the median-of-ratios method. Raw counts are then normalized using these size factors, thereby accounting for differences in sequencing depth and minimizing confounding effects due to variation in library size. Such differences are common in RNA-seq experiments, particularly when comparing samples derived from distinct input populations.

      (2) The 'acute' and 'chronic' exposures are thought through and carefully defined. The question I do have is whether the 3-hour fasting can be considered chronic fasting, given how surprisingly fast the animals lose their fat content. Could these kinetics indicate that the 30-minute fasting is reflective of mechanisms during which senses change in food availability, whereas the 30 minutes represents acute fasting (with chronic fasting - meaning fasting, during which the animal activated alternative pathways - occurring later)? While this may appear to be pure semantics, it could influence how the authors interpret their results. One method to more objectively separate an 'acute' from a 'chronic' stage may be to conduct a time course of fat loss-does fat loss plateau after 3 hours? The timing when the rate of decrease levels off could be more indicative of the beginning of a chronic phase.

      We appreciate this important point and agree that it should be more clearly discussed. We interpret acute fasting as a pre-fat-loss state, since it is 30 minutes off food and no difference in fat levels are detectable at this stage (Fig 1B). The translatomic changes observed under acute fasting therefore likely reflect food-sensing mechanisms and early preparatory responses that promote subsequent fat mobilization. In contrast, chronic fasting (180 minutes off food – see Fig 1D) appears to represent a post-fat-loss state, in which fat stores have already been depleted, and the corresponding translatomic changes likely reflect the effects of sustained metabolic stress.

      (3) The age of the animals used has to be more explicitly stated. Were these animals egg-laying? Or L4/young adults? This is likely to impact the changes that the animals undergo.

      Day 1 young adults were subjected to the fasting. Great care was taken to ensure consistency across biological replicates.

      (4) What is the rationale, in the authors' view, that stress response genes are apparently more enriched than metabolic or mitochondrial enzymes, and membrane receptor changes? Are the latter mostly regulated by PTMs/localization changes, etc?

      Based on our current data, we cannot exclude the possibility that metabolic or mitochondrial enzymes, as well as membrane receptors, are regulated in INT1 and INT2–9 cells through mechanisms not captured at the translatome level, including post-translational modification or changes in subcellular localization under different fasting and refeeding conditions.

      (5) The refeeding experiment with latex beads and killed OP50 is very clever. Details on when INS-7 was evaluated in the caoelomocytes would help the reader better understand and interpret these results.

      INS-7mCherry signal was evaluated in the coelomocytes immediately after refeeding; we included this information in the methods section and referenced our previous paper.

      (6) In the Discussion, I was looking for a more detailed context for how to think about the differences and similarities in the RNA-seq data between the two segments, and perhaps a discussion of whether there were any indications that the two segments communicated with each other.

      The data show that the most pronounced difference between INT1 and the rest of the intestine at the RNAseq level, is the expression of stress response genes in INT1. Although there are some nuanced differences, the prevalence of stress response genes persists across feeding and fasting conditions. This difference, combined with the evidence that INT1 cells secrete the enteroendocrine peptide INS-7 (Fig 6 and PMID: 39127676) is strongly reminiscent of the mammalian enteroendocrine cells, which also secrete peptides and show strong expression of stress response genes (PMID: 37626258 and 27148273). We suggest that this category term reflects not only a canonical stress response, but also a broader response to shifts in the luminal environment, which INT1 cells are anatomically poised to detect well before the absorption of nutrients has begun further down the intestine (INT2-9). Thus, we believe INT1 cells are a newly defined enteroendocrine cell type within the C. elegans intestine.

      Regarding communication between INT1 and INT2-9 – this is an intriguing possibility that we have considered, given that peptide genes and receptors are found in the RNAseq datasets. The extent to which the expression of these genes leads to functional effects is the subject of future investigation.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1A - It would be better to also show single fluorescent protein channels to assess the specificity of the expression patterns. A schematic or labels of where the INT1 vs. INT2-9 boundaries are located would be helpful to non-experts.

      (2) Figure 4 - At present, the Venn Diagrams are a very complicated way to visualize all of the comparisons/conditions. I would recommend that the authors consider using UpSet plots to better summarize the relevant comparisons they would like to make. The same consideration applies to Figure 5D.

      (3) Line 301 - Description of the INT1-specific RNAi strategy. I think it would be better to bring this information earlier, close to line 285, where the authors first mention performing INT1-specific RNAi experiments.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (4) Line 304 - I think the authors meant to cite a reference where the REF placeholder text is found

      We have replaced the “REF” with the reference.

    1. eLife Assessment

      This study investigated mitochondrial dysfunction and the impairment of the ciliary Sonic Hedgehog signaling in Lowe syndrome (LS), a timely topic given the limited research in this area. The data obtained from patient-derived iPSC neurons and a mouse model are solid. Although the main claims of the study are only partially supported by the current evidence, it provides a useful starting point for future functional studies investigating the link between mitochondrial defects and primary cilia in neural development.

    2. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how neural cell development is affected in Lowe syndrome. Using neural cultures differentiated from human iPSCs carrying either a LS mutation or a genetically engineered mutation in OCRL, the authors show a depletion of mitochondrial DNA and decrease in mitochondrial activities that correlate with an increased formation of astrocytes at the expense of neurons. Similar effects on mitochondria and on astrocyte development were observed in a LS mouse model. Moreover, these mutant brain cells are less likely to be ciliated and show a reduction in Sonic hedgehog signalling.

      Strengths/Weaknesses:

      The study derives strength from the analyses of two different models of Lowe syndrome, both reaching similar conclusions. However, the observed changes in mitochondrial defects, neuronal/astrocytic development and primary cilia are only correlated, with no attempt to investigate a causal relationship. Moreover, the mouse model is only analysed at the adult stage providing no insights into the development of the defects. Different brain regions are analysed with immunostainings and qRT-PCR making it challenging to draw clear correlations between these findings. The quality of the corresponding figures is often poor and the selection of markers is frequently inappropriate. Taken together, these limitations complicate the interpretations of the data and significantly limit the conclusions that can be drawn from the study.

      Although the study remains incomplete as main claims are only partially supported it can be used as a starting point for future functional studies into the link between mitochondrial defects and primary cilia in neural development.

      Comments on revised version:

      I am afraid the revised manuscript does little to address the concerns I raised in my initial review. The authors have primarily revised the text, removed over-interpretations and discussed critical points as limitations of the study. This gives the impression that key concerns have merely been rationalised, particularly as only a few new experiments are presented. My main concerns therefore remain:

      (1) The authors present three different phenotypes (altered neural differentiation, mitochondria dysfunction, alterations in primary cilia and ciliary Shh signalling) but a link between these phenotypes is not investigated. No mechanistic experiments are presented. Instead, the authors try to address the lack of a mechanism through refined wording but still use formulations that imply a direct link between these phenotypes. For example, their rebuttal letter finishes with the statement that the manuscript "provides a multi-model, cross-species framework linking mitochondrial dysfunction, ciliary signaling, and altered neural differentiation in Lowe syndrome". Similar formulations are used in the text.

      (2) The authors still claim that ciliary Shh signalling is reduced but ignore the fact that Shh mRNA in iN cells and Shh protein in the IOB mouse are significantly decreased. This reduction represents the most likely explanation for the reduced levels of Gli1 and Ptc1 mRNAs (Shh target genes), rather than dysfunction of cilia. In order to test for cilia dysfunction, the authors need to use experiments in which they quantify the response of control and OCRL mutant cells to exogenously added Shh protein or Shh agonists. Moreover, the increased Gli1 protein expression in the IOB mouse contradicts the reduced levels of Gli1 mRNA.

      (3) The analyses of the IOB mice are only done in 2 months old adult animals, nevertheless claims are made that changes in cell proportions are consequences of altered cell fate decisions. Alterations in proliferation and cell death are not addressed by experiments.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study investigated mitochondrial dysfunction and the impairment of the ciliary Sonic Hedgehog signaling in Lowe syndrome (LS), a timely topic given the limited research in this area. The data from patient iPSC-derived neurons and a mouse model were collected using solid methods, but the evidence supporting key claims is incomplete, and some technical aspects fall short of expectations. Despite these limitations, the study provides a useful foundation for exploring the relationship between mitochondrial defects and primary cilia in neural development.We appreciate the editorial assessment highlighting the importance of studying mitochondrial dysfunction and ciliary signaling in Lowe syndrome. We acknowledge that our study is largely associative, and we have revised the manuscript to clearly state this limitation, toned down causal claims, and emphasized that our work provides a foundation for future mechanistic studies.

      We appreciate the editorial assessment highlighting the importance of studying mitochondrial dysfunction and ciliary signaling in Lowe syndrome. We acknowledge that our study is largely associative, and we have revised the manuscript to clearly state this limitation, toned down causal claims, and emphasized that our work provides a foundation for future mechanistic studies.

      We have also:

      - Improved figure clarity and consistency

      - Corrected errors in gene annotations and normalization

      - Refined the mechanistic framework linking OCRL, mitochondria, and cilia

      New Experimental Data:

      Figure 5, Supplementary Figure 4. We confirmed mitochondrial defects by generating ocrl-KO zebrafish (Supplementary Figure 4). We first assessed mitochondrial reactive oxygen species (mitoROS) using MitoSOX staining. Next, we evaluated mitochondrial membrane potential (ΔΨm) using MitoTracker CMXRos. Finally, we assessed mitochondrial content via TOM20 staining. For all analyses, we focused on the ocular and cranial regions of the zebrafish to maintain consistency (see Author response image 1). Notably, previous studies have reported that ocrl-KO zebrafish exhibit seizures and brain developmental abnormalities, supporting their relevance as a model for Lowe syndrome-like phenotypes [1].

      Author response image 1.

      Figure 6 c. In addition to quantifying the proportion of ciliated cells in the IOB mouse brain, we measured cilia length and compared it between IOB and WT brain sections. Our results show that IOB mice exhibit elongated cilia compared to WT controls, suggesting that OCRL deficiency is associated with stress-related alterations in ciliary structure. These findings are consistent with previous studies reporting that cilia elongation can be associated with increased ROS levels and mitochondrial dysfunction [2,3].

      Public Reviews:

      Reviewer #1 (Public review):

      The preparation of the manuscript requires improvement. There are many errors in the presentation of data.

      We thank the reviewer for this important comment. We have carefully revised the manuscript to improve the clarity, accuracy, and consistency of data presentation.

      Specifically, we have corrected inconsistencies in gene nomenclature (e.g., CO2 vs COX2, DLOOP) across the text, figures, and legends. We standardized normalization methods and ensured consistency between figures and descriptions. We revised figure labels, legends, and annotations for clarity and accuracy. We corrected referencing errors and ensured appropriate citation of prior work. We improved overall figure quality and readability. In addition, we performed a thorough review of the entire manuscript to eliminate typographical errors and ensure consistency in terminology and data interpretation.

      The use of references needs to be re-considered. Sometimes a reference is used when in fact the results included in that paper are the opposite of what the authors intend.

      We thank the reviewer for this important comment. We have carefully re-evaluated all references throughout the manuscript to ensure that they accurately reflect the findings they are cited to support. In cases where the cited studies did not fully align with our interpretation or could be misleading, we have either revised the text to more accurately represent the original findings or replaced the references with more appropriate sources. We have also clarified instances where prior studies report differing or context-dependent results to avoid overinterpretation.

      The authors conclude the paper by claiming that mitochondrial dysfunction and impairments of the ciliary SHH contribute to abnormal neuronal differentiation in LS, but the mechanism by which this sequence of events might happen hasn't been shown.

      We thank the reviewer for this important comment. We agree that the current study does not establish a direct causal mechanism linking mitochondrial dysfunction, ciliary SHH signaling, and altered neuronal differentiation in Lowe syndrome. Our data demonstrate that these processes co-occur consistently across multiple model systems, supporting a potential functional relationship. However, we acknowledge that the precise sequence of events and mechanistic connections remains to be defined. To address this, we have revised the manuscript to clarify that our conclusions are based on associative findings rather than direct mechanistic evidence. We have also updated the Discussion to explicitly acknowledge this limitation and to frame our model (Figure 7) as a proposed working hypothesis. Future studies will be required to determine whether mitochondrial dysfunction directly impacts ciliary SHH signaling and how these pathways influence neuronal differentiation.

      Phenotype of increased astrocytes in both the IOB mouse brain or iPSC-derived cultures iN cells requires clarification as one of the markers used as an astrocyte marker, BRN2, is commonly used as a neuronal marker. As LS is a neurodevelopmental disorder, and the phenotype in question is related to differentiation, it is crucial to shed light on the developmental timeline in which this phenotype is seen in the mouse brain.

      We thank the reviewer for this important comment. We agree that the use of BRN2 as an astrocytic marker was inappropriate, as it is primarily recognized as a neuronal marker. Accordingly, we have revised the manuscript to remove BRN2 from the interpretation of astrocytic identity and now rely on GFAP expression as the primary astrocytic marker. We have also clarified this point in both the Results and figure legends to avoid misinterpretation. In addition, we have revised the text to more accurately describe our findings as an altered balance in neuronal versus astrocytic marker expression, rather than a definitive increase in astrocyte numbers.

      Regarding the developmental context, we acknowledge that Lowe syndrome is a neurodevelopmental disorder and that temporal aspects are highly relevant. In our study, the in vivo analyses were performed on adult 2-month-old IOB mouse brains, which we have now explicitly stated in the manuscript. We recognize that this limits our ability to directly assess developmental dynamics of lineage specification. We have therefore added this as a limitation in the Discussion and clarified that future studies examining earlier developmental stages will be necessary to determine when these alterations arise.

      Mitochondrial dysfunction in astrocytes has been shown to induce a ciliogenic program. However, almost the opposite is shown in this paper, with regards to ciliation. Morphology of the cilia was not assessed either, which is an important feature of ciliary homeostasis. The improper ciliary homeostasis here appears to be the improper Shh signalling, which has not been shown to be related to mitochondrial dysfunction. This leaves one wondering how exactly the different phenotypes shown in this paper are connected.

      We thank the reviewer for this important comment. We agree that the relationship between mitochondrial dysfunction, ciliogenesis, and Shh signaling is complex and not fully resolved in the current study.

      As noted by the reviewer, prior studies have reported that mitochondrial dysfunction can promote a ciliogenic program [4]. In contrast, our data show a reduced proportion of ciliated cells together with increased cilia length, indicating altered ciliary homeostasis rather than a straightforward increase in ciliogenesis. To address this point, we have revised the manuscript to describe our findings as context-dependent alterations in ciliary parameters more clearly, and we now explicitly discuss this apparent discrepancy with the literature in the Discussion. We also acknowledge the reviewer’s point regarding ciliary morphology. In the revised manuscript, we have included quantification of cilia length in addition to the proportion of ciliated cells, and we have expanded the Methods section to detail how these measurements were performed. We agree that additional ultrastructural and functional analyses would further strengthen the characterization of ciliary homeostasis, and we now include this as a limitation and future direction.

      Regarding the link between mitochondrial dysfunction, ciliary alterations, and Shh signaling, we agree that our study does not establish a direct mechanistic connection. Our data demonstrate that these phenotypes co-occur consistently across multiple models, but do not define causality. To address this concern, we have revised the manuscript to clarify that our conclusions are associative, and we now present our integrated model (Figure 7) as a working hypothesis rather than a demonstrated mechanism. We also explicitly state in the Discussion that future studies will be required to determine whether mitochondrial dysfunction directly impacts ciliary signaling and Shh pathway activity.

      This paper lacks a clear mechanistic approach. While the data validates the 3 broad phenotypes mentioned, there is a lack of connection between these phenotypes or an answer to why these phenotypes appear. While the discussion attempts to shed light on this by referencing previous studies, some of the referenced studies show contradicting results. Hence, it would be beneficial to clarify these gaps with further experiments and address the larger question of the connection between the mitochondria, Shh signalling, and astrocyte formation.

      We thank the reviewer for this important and insightful comment. We agree that the current study does not establish a direct mechanistic link connecting mitochondrial dysfunction, altered Shh signaling, and changes in neuronal versus astrocytic differentiation.

      Our primary goal in this work was to identify and validate phenotypes associated with OCRL deficiency across multiple independent model systems. We demonstrate that mitochondrial dysfunction, oxidative stress, altered ciliary/Shh signaling, and changes in neural lineage-associated markers co-occur consistently in these models. However, we acknowledge that the causal relationships between these processes remain to be defined.

      To address this concern, we have revised the manuscript to more clearly state that our conclusions are associative rather than mechanistic, and we now present our integrated model (Figure 7) as a working hypothesis that links these phenotypes through a potential mitochondria-ROS-signaling axis. We have also expanded the Discussion to explicitly acknowledge this limitation and to avoid overinterpretation of causality.

      In addition, we have carefully re-evaluated and revised the cited literature to ensure accuracy, particularly in cases where prior studies report context-dependent or seemingly contradictory effects of mitochondrial dysfunction on ciliogenesis and signaling pathways. These points are now discussed more explicitly to better position our findings within the existing literature.

      We agree that further experiments, such as targeted rescue of mitochondrial function or modulation of Shh signaling, will be necessary to establish causal relationships between these pathways. These directions are now clearly outlined in the revised Discussion as important next steps.

      Most importantly, there is no mention of how the loss of OCRL, a 5-phosphatase enzyme, results in the appearance of the mentioned phenotypes. Since there are multiple studies in the field of Lowe Syndrome that shed light on the various functions of OCRL, both catalytic and non-catalytic, it is important to address the role of OCRL in resulting in these phenotypes.

      We thank the reviewer for this important comment. We agree that the link between OCRL function and the observed phenotypes was not sufficiently developed in the original version of the manuscript.

      In the revised manuscript, we have expanded the Discussion to more clearly outline how loss of OCRL could contribute to the observed mitochondrial, ciliary, and differentiation phenotypes. OCRL encodes a PI(4,5)P₂ 5-phosphatase that regulates phosphoinositide homeostasis and membrane dynamics. Disruption of this activity is known to affect endolysosomal trafficking, actin organization, and membrane remodeling-processes that are critical for organelle maintenance and ciliary function. We now discuss how these alterations could impact mitochondrial homeostasis, for example, through defects in membrane contact sites, vesicular trafficking, or organelle quality control pathways.

      In addition, we have incorporated discussion of potential non-catalytic roles of OCRL, including protein–protein interactions and scaffolding functions, which may contribute to the coordination of intracellular trafficking and cytoskeletal organization. These aspects may provide an additional layer of regulation linking OCRL loss to both mitochondrial dysfunction and ciliary alterations.

      We emphasize that, while these mechanisms are supported by prior studies, our data do not directly test them. Therefore, we have carefully framed this section as a plausible mechanistic framework rather than a demonstrated pathway and have explicitly stated this limitation. We also outline future experiments aimed at dissecting catalytic versus non-catalytic contributions of OCRL to these phenotypes.

      There are numerous errors in the qPCR experiments performed concerning the genes that were assayed. The genes mentioned in the text section do not match those indicated in the graphs or legends. This takes away the confidence of the reader in this data.

      We thank the reviewer for this important observation. We agree that the inconsistencies between the genes described in the text and those shown in the figures and legends could reduce confidence in the data. In the revised manuscript, we have carefully rechecked all qPCR experiments and corrected the gene names across the Results, figures, and figure legends to ensure full consistency. We have also standardized the nomenclature throughout the manuscript (including consistent use of gene symbols and formatting) and verified that all plotted data correspond to the correct targets.

      In addition, we have clarified the qPCR methodology, including normalization (all data are normalized to GAPDH) and primer information, to improve transparency and reproducibility.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how neural cell development is affected in Lowe syndrome. Using neural cultures differentiated from human iPSCs carrying either an LS mutation or a genetically engineered mutation in OCRL, the authors show a depletion of mitochondrial DNA and a decrease in mitochondrial activities that correlate with an increased formation of astrocytes at the expense of neurons. Similar effects on mitochondria and on astrocyte development were observed in an LS mouse model. Moreover, these mutant brain cells are less likely to be ciliated and show a reduction in Sonic Hedgehog signalling.

      Strengths/Weaknesses:

      The study derives strength from the analyses of two different models of Lowe syndrome, both reaching similar conclusions. However, the observed changes in mitochondrial defects, neuronal/astrocytic development, and primary cilia are only correlated, with no attempt to investigate a causal relationship. Moreover, the mouse model is only analysed at the adult stage providing no insights into the development of the defects. Different brain regions are analysed with immunostainings and qRT-PCR making it challenging to draw clear correlations between these findings. The quality of the corresponding figures is often poor and the selection of markers is frequently inappropriate. Taken together, these limitations complicate the interpretations of the data and significantly limit the conclusions that can be drawn from the study.

      We have carefully revised the manuscript to address the concerns raised, and we have revised the manuscript with additional supporting data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors have checked the expression of neuronal markers NeuN and FoxG1, and apart from GFAP, they categorise Brn2 as one of the astrocytes markers that they have also checked. But Brn2 is not an astrocyte marker. It is a neuronal marker that is expressed in layer 2/3 of the cortex. In fact, Brn2 is reported to be a key driver of neurogenesis in primate telencephalon development1 and for reprogramming of astrocytes to neurons2. Hence, the only glial marker they have used here is GFAP. BRN2 is a neuronal marker. It has been used as a neuronal marker even in the reference (Zhang et al, 2013), from which the protocol for inducing iPSCs to induced neurons (iNs) was taken. Hence, the qPCR results in 1g of overexpression of BRN2 indicate an increase in expression of a neuronal marker, not an astrocyte marker.

      We thank the reviewer for this important and well-founded comment. We fully agree that BRN2 is a neuronal marker and not an astrocytic marker, and that its inclusion as an astrocyte marker in our original interpretation was incorrect.

      In the revised manuscript, we have removed BRN2 from the analysis and interpretation of astrocytic identity. We now treat BRN2 exclusively as a neuronal marker and have updated the Results, figure legends, and text accordingly. Specifically, the qPCR data previously presented in Figure 1g are now interpreted as reflecting neuronal marker expression, not astrocytic differentiation. We have also revised our conclusions to avoid overinterpretation of astrocyte abundance. Our findings are now described more accurately as an altered balance in neuronal versus astrocytic marker expression, rather than a definitive increase in astrocyte numbers. In this context, GFAP remains the primary astrocytic marker used in this study.

      We acknowledge the reviewer’s point that reliance on a single astrocytic marker is a limitation. This has now been explicitly stated in the Discussion, and we note that additional astrocyte markers will be required in future studies to more comprehensively define lineage-specific changes.

      Incorrect marker usage (BRN2 as astrocyte marker)

      We thank the reviewer for identifying this critical issue. We corrected the classification of BRN2 as a neuronal marker. Also, we re-analyzed the interpretation accordingly, revised all relevant text and figures. Importantly, Astrocyte conclusions are now based primarily on GFAP expression, and we explicitly acknowledge this limitation in the Discussion

      The graphs for the RT-PCR results indicate that gene expression values are normalized to actin whereas the legend mentions that they are normalized to GAPDH. This needs clarification.

      We thank the reviewer for pointing out this inconsistency. We confirm that all qPCR data were normalized to GAPDH, and the reference to actin was an error. This has now been corrected throughout the figures, legends, and text to ensure consistency.

      The use of wording to refer to the generation of induced neurons (iNs) should ideally be changed from "we developed"; as the protocol from Zhang et al, 2013 seems to have been directly adapted in this paper.

      We thank the reviewer for this helpful suggestion. We agree that the wording was inappropriate. In the revised manuscript, we have replaced “we developed” with language indicating that iNs were generated using an established protocol, and we now explicitly state that the method was adapted from Zhang et al., 2013 [5].

      OCRL KO iPSCs were obtained from Herbert Lachman's lab and not generated in this study. Hence, the Ran et al, 2013 reference is not necessary.

      We thank the reviewer for this clarification. We agree that the OCRL knockout iPSCs were obtained from Herbert Lachman’s laboratory and were not generated in this study. Accordingly, we have removed the Ran et al., 2013 reference and revised the manuscript to clearly state the origin of the OCRL KO iPSC line.

      The reference for Figure 1c is given as Ran et al, 2013 which is wrong. It should be Zhang et al, 2013.

      We thank the reviewer for noting this error. We have corrected the reference for Figure 1c from Ran et al., 2013 to Zhang et al., 2013 in the revised manuscript.Limited in vivo mitochondrial characterization

      Figure 1D, E: GFAP is a cytoskeletal marker but its expression here is very grainy and looks like an artifact. Is it possible to show the astrocyte phenotype using other astrocytes nuclei and cytosolic markers such as NFIA and S100B, respectively?

      We thank the reviewer for this important suggestion. We acknowledge that GFAP is a cytoskeletal marker and that the signal in the current images may appear granular. We have carefully re-evaluated the staining and image processing to ensure that the signal represents true GFAP expression and have improved the image quality and presentation in the revised figures. We agree that inclusion of additional astrocytic markers such as NFIA and S100B would further strengthen the characterization. While we were not able to include these additional markers in the current revision, we now explicitly acknowledge this as a limitation in the Discussion and note that future studies will incorporate a broader panel of astrocyte markers to more comprehensively define astrocytic identity.

      Since the authors have not used enough markers to understand the cell-type composition in WT and OCRL KO/mutant lines, it's not sufficient to conclude that the NPCs preferentially favour astrocytes over neuronal lineage. Any conclusive comments regarding the cell-state/cell-type specification necessitate evidence such as genetic lineage tracing using reporters for neuronal and astrocyte markers, and/or RNA/ATAC/scRNA sequencing.

      We thank the reviewer for this important point. We agree that the current marker panel is not sufficient to definitively determine cell-type composition or to conclude preferential lineage specification.

      In the revised manuscript, we have tempered our conclusions and now describe our findings as changes in neuronal versus astrocytic marker expression, rather than evidence of a shift in lineage fate. We also explicitly acknowledge this limitation in the Discussion. We agree that approaches such as genetic lineage tracing, reporter-based assays, and single-cell transcriptomic or epigenomic analyses (e.g., scRNA-seq or scATAC-seq) would be required to rigorously define cell-state transitions and lineage outcomes. These are important directions for future studies and are now highlighted in the revised manuscript.

      Figure 2:

      The mt-DNA gene CO2 was checked, not COX2. A typographical error in the written section, which does not match with the qPCR graph of the same.

      We thank the reviewer for noting this inconsistency. We confirm that the gene analyzed was CO2, and the reference to COX2 in the text was incorrect. This has now been corrected throughout the manuscript to ensure consistency between the text, figures, and qPCR data.

      The word "neurogenesis" is used very loosely throughout the paper. In the opinion of the reviewer, there is no evidence presented that there is a defect in neurogenesis in either of the models used in this paper.

      We thank the reviewer for this important comment. We agree that the term “neurogenesis” was used too broadly and is not directly supported by our data. In the revised manuscript, we have removed or replaced this term where appropriate and now refer more precisely to changes in neuronal versus astrocytic marker expression.

      Line 150: They say that they have examined the functional properties of mitochondria during neurogenesis but it would have been better to understand OXPHOS at various time points of neurogenesis to actually conclude reduced OXPHOS 'during neurogenesis'. Moreover, genes related to other pathways such as glycolysis could have been checked to understand the bioenergetics of LS patients. Also, oxidative stress could have been checked using more than one marker. Since they are trying to understand the functional role of mitochondrial defects during neurogenesis, they could have performed live imaging of mitochondrial potential during various stages of neurogenesis. Isolation of mitochondria from LS patients and transcriptomics/proteomics might provide further clues about mitochondrial defects.

      We thank the reviewer for these thoughtful suggestions. We agree that our data do not capture mitochondrial function across multiple stages of neurogenesis. In the revised manuscript, we have modified the wording to avoid implying temporal analysis “during neurogenesis” and instead describe mitochondrial parameters in differentiated cells. To strengthen the study, we have included additional in vivo validation in the zebrafish model, where we assessed multiple mitochondrial readouts, including mitochondrial membrane potential (ΔΨm) using MitoTracker CMXRos, oxidative mitochondrial stress ROS (mitoROS) using MitoSOX staining, and mitochondrial content (TOM20), supporting mitochondrial dysfunction across systems.

      Elevated astrocytic reaction during the differentiation of NSPCs in the Lowe syndrome (IOB) mouse model.

      Title: What does astrocyte reaction mean? This term should not be used without clear evidence of reactive astrocytes being present in the model.

      We thank the reviewer for this important comment. We agree that the term “astrocytic reaction” is not appropriate without specific evidence of reactive astrocytes. In the revised manuscript, we have removed this terminology and replaced it with more accurate wording, describing our findings as altered astrocytic marker expression. This change better reflects the data and avoids overinterpretation.

      In 1a, no quantification of the mouse brain size is given. From the given images alone, there appears to be no obvious decrease in brain size between the WT and IOB mice. This contradicts the text which indicates that the IOB mouse brain is smaller.

      We appreciate this important point. We have:

      Removed claims regarding reduced brain size

      Clarified that our analysis was limited to available sections and no definitive conclusion about global brain morphology can be made

      Figure 3E: Why is the astrocyte to neuron ratio measured using a cytoskeletal marker for astrocytes, GFAP but a nuclear marker for neurons, NeuN? Ratios to measure the percentage or proportion of astrocytes to neurons can only be checked by markers of the same nature such as GFAP to MAP2 (neuronal cytoskeletal marker) or NFIA (astrocyte nuclear marker) to NeuN.

      We thank the reviewer for this important point. We agree that comparing a cytoskeletal marker (GFAP) with a nuclear marker (NeuN) is not ideal for deriving cell-type ratios. In the revised manuscript, we have removed the astrocyte-to-neuron ratio analysis and now present these data as relative marker expression/signals rather than proportions. We have also clarified this limitation in the text and Discussion.

      (3) Lack of clarity in the experiments performed on the mice brains. PAX6 is used here as a neuronal marker along with NeuN, a mature neuronal marker. This is misleading as PAX6 is rather a marker for neural stem/progenitor cells and not neurons. The age of the mice has also not been mentioned, which is crucial considering the different markers used to characterize the mouse brain as well as since the authors are indicating that there is an abnormal neurodevelopment in the IOB mouse during development. Again, BRN2 is used here as an astrocyte marker. However, it is a neuronal marker. Hence, the phenotype of increased astrocytes currently is held by GFAP expression alone. Another astrocyte marker should be used.

      We thank the reviewer for these important points. We have revised the manuscript to correct marker interpretation, now describing PAX6 as a progenitor marker rather than neuronal, and BRN2 as a neuronal marker, removing it from astrocyte-related analysis. We have also explicitly stated the age of the mice (2 months) in the Methods and Results. In addition, we have tempered our conclusions, describing the data as changes in marker expression rather than definitive cell-type shifts, and we now acknowledge that reliance on GFAP as a single astrocytic marker is a limitation, which is discussed in the revised manuscript.

      Figure 6: Increase in astrocytes, mitochondrial dysfunction, and ciliary Shh signalling are 3 phenotypes discussed in this study. However, no experiments were done to shed light on the mechanistic connection between these phenotypes. This is reflected in the abstract shown in Figure 6. There is no comment on the mechanism behind these phenotypes.

      We thank the reviewer for this important comment. We agree that the current study does not establish a direct mechanistic link between mitochondrial dysfunction, altered ciliary Shh signaling, and changes in astrocytic markers. Our aim was to identify and validate these phenotypes across multiple models. In the revised manuscript, we have clarified that Figure 7 represents a proposed working model based on associative findings rather than a defined mechanism. We have also revised the Discussion to explicitly acknowledge this limitation and to outline future experiments required to establish causal relationships between these

      Reviewer #2 (Recommendations for the authors):

      The authors report interesting findings in two different experimental models but the manuscript would benefit significantly from an analysis of a potential causal relationship between different findings. They often mention neural stem cells or the neuron/glia switch but their analysis of the mouse mutant is restricted to the adult stage. A more consistent analysis of specific brain regions would also be beneficial.

      We thank the reviewer for this constructive comment. We agree that establishing causal relationships between the observed phenotypes is an important next step. In the revised manuscript, we have clarified that our conclusions are based on associative findings and have expanded the Discussion to outline experimental strategies that could address causality in future studies. We also acknowledge that our in vivo analysis is restricted to adult (2-month-old) IOB mouse brains, which limits our ability to assess developmental dynamics such as neural stem cell behavior or neuron-glial transitions. This limitation is now explicitly stated in the Discussion. Finally, we agree that region-specific analysis would strengthen the study. Due to the availability of samples, our analysis was not systematically performed across defined brain regions. We now acknowledge this limitation and note that future studies focusing on specific regions (e.g., cortex, hippocampus) will be important to better understand the spatial aspects of the phenotype.

      Figure 1: The authors only measured the expression of marker genes by qRT-PCR. This could reflect higher expression levels in individual cells rather than a change in the proportion of neurons and astrocytes. They need to determine the cell proportions of astrocytes and neurons in addition. Moreover, there is a poor marker choice. Loss of FOXG1 expression could indicate a loss of telencephalic identity. BRN2 is expressed by cortical neurons.

      We thank the reviewer for this important comment. We agree that qPCR-based marker analysis does not directly reflect cell-type proportions and may instead represent changes in gene expression per cell. Accordingly, we have revised the manuscript to avoid conclusions about cell proportions and now describe the data as changes in marker expression. We have also corrected marker interpretation, removing BRN2 from astrocyte analysis and clarifying that FOXG1 reflects telencephalic identity. These limitations and the need for more comprehensive cell-type characterization are now acknowledged in the Discussion.

      In Figure 2, the authors determine the properties of mitochondria and claim that functional mitochondrial activities are decreased during neurogenesis in mutant iN cells. They need to take into account that according to Figure 1 the proportion of neurons and astrocytes may be changed. Hence, the decreased mitochondrial activity may reflect a fundamental difference between neurons and astrocytes. The authors need to clearly distinguish between neurons and astrocytes in their analysis. Moreover, the use of the term neurogenesis is confusing. They are analysing the neuron-to-glial switch, not the formation of neurons.

      We thank the reviewer for this important comment. We agree that differences in cell-type composition may influence mitochondrial measurements. In the revised manuscript, we have tempered our interpretation, describing these data as changes in mitochondrial parameters at the population level rather than neuron-specific effects. We also acknowledge this limitation in the Discussion and note that cell-type-specific analyses will be required in future studies. In addition, we have revised the terminology throughout the manuscript, removing the term “neurogenesis” and instead referring to changes in neuronal versus glial marker expression to more accurately reflect the scope of our analysis.

      Figure 3: The authors claim that astrocyte numbers are elevated in the IOB mouse model, however, it seems as if the authors analysed late postnatal, potentially adult brains but no age of the brains is provided. Given the large time lag between the formation of astrocytes and their analysis, the increased number of astrocytes could be due to a number of processes including altered proliferation and cell death. The authors need to investigate the proportion of astrocytes and neurons closer to the neuron-to-glia switch. Cell fate experiments like the long-term application of BrdU would be much better suited and would provide mechanistic insights. Again, markers are not adequate to reach their conclusion. Pax6 is only expressed in a tiny subset of neurons, Brn2 on the other hand is not astrocyte-specific as it is expressed in cortical neurons as well. Moreover, qRT-PCR analyses were done in the cortex and hippocampus whereas the boxes in Figure 3D are located in the basal ganglia. It would be much more informative and provide better comparisons to perform gene expression analysis and cell counts in the same brain regions.

      We thank the reviewer for these important and constructive comments. We agree that our analysis is limited by the use of adult (2-month-old) IOB mouse brains, which do not allow direct assessment of developmental processes such as the neuron-to-glia transition. We have now explicitly stated the age of the animals and clarified this limitation in the Discussion, including the possibility that changes in astrocytic markers may reflect processes such as proliferation or survival rather than lineage specification.

      We also agree that our marker panel was insufficient for definitive conclusions. Accordingly, we have revised the manuscript to remove overinterpretation, corrected marker usage, and now describe the data as changes in marker expression rather than cell-type proportions. The need for more rigorous approaches, such as lineage tracing (e.g., BrdU) and expanded marker panels, is now acknowledged as a future direction. Finally, we thank the reviewer for pointing out the inconsistency in the brain regions analyzed. We have clarified the regions used for qPCR, and we now explicitly acknowledge this limitation, noting that future studies will aim to perform region-matched molecular and histological analyses for more accurate comparisons.

      Experiments in Figure 4 assess "whether changes in mitochondrial activity are involved in the altered differentiation of stem cells and progenitor cells in the LS mouse model" in 3-month-old brain sections. The murine adult brain only contains a few neural stem cells in the SVZ and in the dentate gyrus. Instead, this analysis needed to be done at late embryonic/early postnatal stages to capture the neuronal/glial switch. In addition, RT-PCR and immunostainings should be performed in the same brain region as stated above.

      We thank the reviewer for this important point. We agree that analysis in adult (2-month-old) brains does not capture developmental stages such as the neuron-glia transition. We have revised the manuscript to remove implications of developmental analysis and now describe these data as mitochondrial parameters in adult tissue, explicitly acknowledging this limitation in the Discussion. We also clarify the brain regions used for qPCR and immunostaining and note as a limitation that these were not fully matched; future studies will perform region-specific, developmentally timed analyses.

      Figure 5: The authors examine a potential link between mitochondrial defects and primary cilia. Mutant iN cell cultures contain lower levels of SHH mRNA and show concomitantly lower expression of the SHH target genes GLI1 and PTCH1. The authors link this finding with a reduced proportion of ciliated cells but the reduced SHH signalling is most likely explained by the decreased SHH expression. The authors also limit their analysis of primary cilia to one brain region, but they should also include the cortex and hippocampus as these regions were used for their qRT-PCR analysis. SHH signalling acts as a switch to stop the proteolytic processing of GLI3 and to promote the formation of the GLI3 activator form. It is therefore important to determine the ratio of GLI3 repressor and GLI3 activator using western blots. The authors claim that they found defective cilia formation, but cilia are poorly characterised. Are there differences in intraflagellar transport, the formation of the transition zone, etc? Is ciliary length altered? The authors only make a correlative link between mitochondrial defects and cilia but present no experiments to investigate causation. They should at least discuss potential mechanisms which could explain defects in cilia.

      We thank the reviewer for these insightful comments. We agree that reduced SHH pathway activity may be influenced by decreased SHH expression, and we have revised the text to avoid overattributing this effect to ciliary changes. Our conclusions are now framed as associative, not causal.

      We have expanded our cilia analysis to include quantification of both the proportion of ciliated cells and cilia length, and clarified these methods in the manuscript. We also acknowledge that additional characterization (e.g., intraflagellar transport, transition zone structure, GLI3 activator/repressor ratios) would further strengthen the analysis, and we now include this as a limitation and future direction. Regarding regional analysis, we agree that broader brain region coverage would be valuable. Due to sample availability, our analysis was limited, and this is now explicitly acknowledged as a limitation, with future studies aimed at region-matched analyses (e.g., cortex and hippocampus). Finally, we have expanded the Discussion to outline potential mechanisms linking mitochondrial dysfunction and ciliary alterations, while clearly stating that causal relationships remain to be established.

      (1) The methods section does not contain any information on how immunostainings on brain sections were performed.

      We agree that the description of immunostaining on brain sections was missing. We have now added a detailed protocol for brain section immunostaining in the Methods section to improve clarity and reproducibility.

      (2) The abbreviation "RT-PCR" is used for both, real-time PCR and reverse transcription PCR

      We also acknowledge the inconsistent use of the term “RT-PCR.” In the revised manuscript, we have standardized the terminology, using “qPCR” (quantitative real-time PCR) throughout to avoid confusion

      Conclusion

      We believe that these revisions significantly strengthen the manuscript. While the study remains primarily associative, it provides a multi-model, cross-species framework linking mitochondrial dysfunction, ciliary signaling, and altered neural differentiation in Lowe syndrome.

      References:

      (1) Ramirez IB-R, Pietka G, Jones DR, Divecha N, Alia A, Baraban SC, et al. Impaired neural development in a zebrafish model for Lowe syndrome. Hum Mol Genet. 2012;21:1744–59. https://doi.org/10.1093/hmg/ddr608

      (2) Kim JI, Kim J, Jang H-S, Noh MR, Lipschutz JH, Park KM. Reduction of oxidative stress during recovery accelerates normalization of primary cilia length that is altered after ischemic injury in murine kidneys. Am J Physiol Renal Physiol. 2013;304:F1283-1294. https://doi.org/10.1152/ajprenal.00427.2012

      (3) Moruzzi N, Valladolid-Acebes I, Kannabiran SA, Bulgaro S, Burtscher I, Leibiger B, et al. Mitochondrial impairment and intracellular reactive oxygen species alter primary cilia morphology. Life Sci Alliance. 2022;5:e202201505. https://doi.org/10.26508/lsa.202201505

      (4) Ignatenko O, Malinen S, Rybas S, Vihinen H, Nikkanen J, Kononov A, et al. Mitochondrial dysfunction compromises ciliary homeostasis in astrocytes. J Cell Biol. 2022;222:e202203019. https://doi.org/10.1083/jcb.202203019

      (5) Zhang Y, Pak C, Han Y, Ahlenius H, Zhang Z, Chanda S, et al. Rapid Single-Step Induction of Functional Neurons from Human Pluripotent Stem Cells. Neuron. 2013;78:785–98. https://doi.org/10.1016/j.neuron.2013.05.029

    1. eLife Assessment

      This study introduces an artificial-intelligence tool that estimates fat in skull bone marrow from routine brain scans, enabling large studies that were previously impractical. The evidence for the method's repeatability and for identifying genetic links is convincing overall. The authors identify genes, diseases, and other biological characteristics that are linked to skull marrow fat, which represents an important advance. The work will be of most interest to researchers using large imaging biobanks and those studying ageing-related changes across bone, blood, and brain.

    2. Reviewer #1 (Public review):

      The authors of this study developed a method to quantify calvarial bone marrow from MRI head scans, enabling study of its composition in large datasets of adults, usually collected to study the brain. Bone marrow intensity can be semi-quantitatively measured in T1-weighted MRI scans due to the greater signal intensity of fat than watery red marrow. This is an ingenious use of the MRI-produced information for other important phenotypes, such as bone structure and marrow content. Different head types were tested for complying to the model, which is notable.

      The model was also successfully validated using several publicly available MRI resources - real data - in (1) dataset consisting of 30 individuals that were scanned 10 times each at 3-day intervals, and (2) the monozygotic (MZ) twin data from the Human Connectome Project cohort. Then the authors applied this validated method to head-MRI scans from the UK Biobank (n=33,042) to extract information on spatial distribution of bone marrow adiposity (BMA) in the calvaria, allowing a GWAS to identify associated genes.

      The authors revealed high heritability and identified 41 genetic loci significantly associated with the BMA trait, including six sex-specific loci. Of note, statistics estimate that 99% of BMA trait-influencing variants are shared with BMD (497 of 500 variants), which may mean these results demonstrate the biological relevance to bone health. Some of the BMA genes were found related to the Wnt pathway, including WNT16, WNT4, NXN; this is a "positive control", since the Wnt/β-catenin signaling pathway was suggested as an important determinant of BMA. Also, associations in genes (BMP4, DLX5, LGR4, LRP4, SFRP4) that are known to specifically influence adiposity, are encouraging. Integrating mapped genes with bone marrow single-cell RNA-seq data revealed patterns of adipogenic lineage differentiation and lipid loading.

      The study also investigated genetic overlap between BMA and twelve (or 13) "brain and body" traits, and identified significant genetic correlations with BMI, cognitive ability and Parkinson's disease.

      In sum, since MRI head scans present a hitherto unexplored opportunity to address unresolved aspects of bone marrow biology, this study is both timely and innovative.

      Comments on revised version:

      The authors responded most of this reviewer's comments. Their explanations are convincing. Yet, upon re-reading the revised version of this paper, I still have concerns about the clarity of mostly analysis presentation, e.g.:

      Line 130-133: the sentence is still unclear: "To obtain the BM signal intensity for an individual datapoint of the calvarium, we ... averaged these BM intensities to get the (average?) BM intensity for that datapoint. Then, we averaged (again?) these datapoint intensities across the calvarium to produce the global BMA measure for the scan."

      Also, I still cannot understand whether the "overlap between the true and predicted bone marrow ...below 0.7" is concerning or not, - whether this threshold of 0.7 is arbitrary.

      Genetic correlation: pls. make sure it's clear that the Rg was calculated using SNP "effect sizes".

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to enable large-scale measurement of fat in skull bone marrow using routine structural brain MRI scans. They present a neural-network pipeline trained largely on realistic simulated examples and show that the resulting skull marrow measure is highly repeatable in test-retest data and consistent in monozygotic twins. Applying it to ~33,000 UK Biobank participants, they report expected population patterns (including sex- and menopause-related differences) and identify genetic and health-related associations, creating an important resource that can be built upon by researchers interested in BMA, imaging, bone, metabolism, neuroscience, ageing, haematology, and other fields.

      Major strengths:

      A notable methodological strength is the training strategy: by using a large, simulated dataset that captures plausible variation in skull-layer thickness and MRI intensity, the authors reduce reliance on scarce expert-labelled images. The modelling choice (using 1D intensity profiles through the skull rather than analysing the full 3D volume) appears well matched to the anatomy and offers an efficient approach for thin, layered structures. Multiple validation steps (including test-retest reliability and twin concordance) support the robustness of the measurement pipeline.

      On the biological and genetic side, the study demonstrates that the skull BMA estimate relates to known correlates of marrow fat (e.g., age/sex/menopause patterns/bone density) and integrates population imaging with large-scale genetic analysis to highlight loci and candidate genes with plausible relevance to skeletal and marrow biology. The inclusion of cross-ancestry analyses and integration with cell-type-resolved gene-expression resources further improves interpretability and usability for the community.

      Major limitations:

      The main limitation is conceptual rather than technical: the phenotype is derived from T1-weighted MRI intensity, which does not directly separate fat and water signals and can vary with scanner and sequence settings. The manuscript provides convincing evidence that the measure is reproducible and biologically meaningful, but it should still be interpreted as a semi-quantitative proxy for marrow fat rather than a direct fat-fraction measurement. Accordingly, the genetic and phenotypic associations are likely informative, but the most direct claims about "adiposity" would be stronger if anchored to established quantitative fat-measurement imaging or spectroscopy in the skull.

      The genetic "replication" analysis in a smaller, ancestrally heterogeneous non-European-ancestry sample is useful as a test of transferability, but it is not equivalent to replication in an independent cohort of similar ancestry and is expected to show reduced SNP-level reproducibility because of differences in sample size and genetic background. This should be clearly framed so readers understand what level of generalisation is supported by the current evidence.

      Likely impact and utility:

      Overall, the work provides a practical method for extracting new biological information from widely available brain MRI scans and should be particularly useful to researchers working with large imaging biobanks and those studying connections between bone, blood, metabolism, and brain ageing. The combination of a scalable measurement approach and openly reported genetic results is likely to accelerate follow-up studies, including cross-cohort comparisons and mechanistic work on candidate pathways.

    4. Reviewer #3 (Public review):

      Summary:

      This paper addresses a fundamental gap in bone biology: our near-complete ignorance of the in vivo dynamics of calvarial bone marrow adiposity (BMA) at population scale. The authors developed an elegant artificial neural network trained on simulated data to automatically localize and quantify the bone marrow layer within standard T1-weighted MRI head scans; scans originally acquired to study the brain but harboring rich, unexploited information about adjacent bone. Applying this method to over 33,000 individuals from the UK Biobank, they accomplished three things that had never been done before: (1) they precisely quantified the sex-dimorphic age trajectory of calvarial BMA, including the dramatic post-menopausal rise and the protective role of hormone replacement therapy; (2) they performed the first well-powered GWAS of this trait, identifying 41 genome-wide significant loci including six sex-specific ones, with SNP heritability of 31.5%; and (3) they revealed significant genetic correlations and overlap between BMA and traits including bone mineral density, Parkinson's disease, and general cognitive ability, a finding made all the more intriguing by the recently described direct vascular channels connecting calvarial bone marrow to the meninges. Integration of GWAS genes with single-cell RNA-sequencing data from mesenchymal lineage cells further illuminated which genes govern lineage commitment to the adipogenic pathway versus lipid loading in mature adipocytes.

      Comments on revised version.

      The reviews raised substantive points across three domains, and the authors engaged with every one of them seriously and thoroughly.

      On the validation of T1-weighted MRI as a measure of BMA: Reviewer 2 raised the strongest concern, arguing that T1-weighted signal intensity had never been formally validated as a quantitative fat-fraction measure in the calvarium. The authors responded with both a principled scientific argument and new data. They assembled existing literature demonstrating that T1-weighted signal is an established semi-quantitative proxy for marrow fat in multiple skeletal sites (Loevner et al. 2002, Shen et al. 2013, Zhang et al. 2020), and provided additional comparative analyses against quantitative T1 relaxation maps, multiple intensity normalization strategies (KDE, WhiteStripe, GMM, FCM, Z-score), DEXA-derived bone mineral density, and osteoporosis status. The biological coherence of their findings, recapitulating known sex and age profiles, identifying genes already established in cell and animal models of BMA biology, and estimating heritabilities consistent with twin data constitutes powerful, convergent evidence for construct validity. Their point that a semi-quantitative measure of a highly variable, well-demarcated biological signal can outperform a perfectly precise measure of a poorly defined entity is methodologically sound and well-argued.

      On sex differences and the role of Hyperostosis frontalis interna: Reviewer 1 raised the clinically astute concern that Hyperostosis frontalis interna (HFI), a condition of inner table thickening prevalent in up to 49% of postmenopausal women, could confound calvarial BMA measurements and drive apparent sex differences. The authors performed a dedicated new analysis, stratifying BMA-BMD associations by sex and age group. They demonstrated that (1) the BMA-BMD association is robust in both males and females, (2) it remains stable across age groups, and (3) the neural network trained on simulations incorporating wide anatomical variation including inner table thickness is inherently resistant to moderate inner table thickening. Given that HFI is restricted to the frontal bone, which represents only a fraction of the calvarial surface, and that severe cases are rare (ICD-10 prevalence ~0.02% in the UK Biobank), the authors make a convincing case that this does not materially bias their results. Their suggestion that the method could itself be used in future work to study the genetic architecture of HFI is a nice forward-looking addition.

      On genetic correlation interpretation and cross-trait pleiotropy: Reviewer 1 asked for clarification of the vertical versus horizontal pleiotropy distinction and for formal Mendelian randomization to support the possible causal effect of BMA on cognition. The authors appropriately clarified the conceptual framework in the revised text and, rather than overstating a causal claim without the supporting analysis, responsibly softened the language to "may be consistent with the hypothesis that BMA could have a causal effect on cognition." This is scientifically honest and appropriate.

      On mouse scRNAseq and its relevance to humans: The authors acknowledged that the results section had not explicitly stated the mouse origin of the scRNAseq data, corrected this, and provided a well-justified rationale for the relevance of mouse mesenchymal lineage data to human BMA biology, which is a well-established and widely accepted model system in this field.

      On GWAS replication: The claim that the study lacked replication was addressed by clarifying the a priori separation of discovery (white British, n=33,042) and replication (non-white British, n=4,958) samples, with 62% of significant discovery SNPs and 95% of lead SNPs replicating in the correct direction.

      Overall Assessment:

      This is a technically innovative, scientifically rigorous, and biologically meaningful paper. The method is genuinely novel, the sample size is among the largest ever applied to this phenotype, the genetic findings are well-powered and well-replicated, and the integration across imaging, genetics, and single-cell transcriptomics is exemplary. The authors have engaged with every substantive reviewer criticism in good faith, producing new analyses where appropriate and defending, and convincingly, with findings that were challenged without adequate basis. The revised manuscript is strengthened throughout.

      This paper opens a new window quite literally, through the skull - into bone marrow biology at a scale and resolution that has never been achieved before.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors of this study developed a method to quantify calvarial bone marrow from MRI head scans, enabling the study of its composition in large datasets of adults, usually collected to study the brain. Bone marrow intensity can be semi-quantitatively measured in T1-weighted MRI scans due to the greater signal intensity of fat than watery red marrow. This is an ingenious use of the MRI-produced information for other important phenotypes, such as bone structure and marrow content. Different head types were tested for complying with the model, which is notable.

      The model was also successfully validated using several publicly available MRI resources - real data - in (1) a dataset consisting of 30 individuals that were scanned 10 times each at 3-day intervals, and (2) the monozygotic (MZ) twin data from the Human Connectome Project cohort. Then the authors applied this validated method to head-MRI scans from the UK Biobank (n=33,042) to extract information on the spatial distribution of bone marrow adiposity (BMA) in the calvaria, allowing a GWAS to identify associated genes.

      The authors revealed high heritability and identified 41 genetic loci significantly associated with the BMA trait, including six sex-specific loci. Of note, statistics estimate that 99% of BMA trait-influencing variants are shared with BMD (497 of 500 variants), which may mean these results demonstrate the biological relevance to bone health. Some of the BMA genes were found related to the Wnt pathway, including WNT16, WNT4, NXN; this is a "positive control", since the Wnt/β-catenin signaling pathway was suggested as an important determinant of BMA. Also, associations in genes (BMP4, DLX5, LGR4, LRP4, SFRP4) that are known to specifically influence adiposity, are encouraging. Integrating mapped genes with bone marrow single-cell RNA-seq data revealed patterns of adipogenic lineage differentiation and lipid loading.

      With regards to the reviewer’s comment on the overlap between BMA and BMD trait-influencing variants, we would like to add that the correlation of effect sizes within the overlap is -0.95 as we would expect from bifurcating differentiation of mesenchymal stem cells into osteoblasts or adipocytes: the underlying common biology driving these traits results in shared traits with a negative correlation in the effects.

      The study also investigated the genetic overlap between BMA and twelve (or 13) "brain and body" traits and identified significant genetic correlations with BMI, cognitive ability, and Parkinson's disease.

      In sum, since MRI head scans present a hitherto unexplored opportunity to address unresolved aspects of bone marrow biology, this study is both timely and innovative.

      There are, however, some assumptions, findings, and their interpretation, which require more critical focus.

      Sex-specificity is well described and studied here. Men have higher BMA than women, but post-menopausal women catch up in the BMA values. The authors believe that calvarial marrow has a number of features that make it particularly well-suited to the study of BMA process - which is clinically important in other bone sites. It has a simple "sandwiched" structure that they are able to model. This is true only to some extent: a condition called "Hyperostosis frontalis interna", of unknown etiology (described by Smith & Hemphill in 1956) - is characterized by irregular overgrowth of the inner table of the frontal bone (symmetric/bilateral). Although not of clinical significance, typically benign, studies report a prevalence of 12%; However, it's most common in postmenopausal women - where prevalences up to 49% in women over the age of 65 - have been reported. Thus, sexual dimorphism is obvious and the effect of estrogen is likely shared with whichever bone - and marrow - age-related pathology. So, for women not using HRT, this new layer of the bone might interfere with the calvarial BMA readings and in turn, affect the BMA-related analyses.

      Thank you for bringing the "Hyperostosis frontalis interna" condition to our attention. It is particularly interesting to hear that the etiology is unknown and one may suspect that some kind of calvarial bone marrow dysregulation may be part of the cause. Our model for bone marrow location was trained on simulated data which included variation in the thickness of all anatomical layers (including the inner table), so it will be robust to some thickening of the inner table. It might not be robust to the most extreme cases of inner table thickening (as described in some case reports), but these are rare. Further, it should also be noted that the other calvarial bones, representing a much greater fraction of the calvarial surface, remain largely unaffected by the thickening and would therefore yield correct localisation of the bone marrow layers. In summary, although the severe cases of hyperostosis frontalis interna have the potential to affect our identification of the bone marrow layer, the low frequency of such cases and the restriction of the phenotype to the frontal bone means that the potential for bias is very limited.

      It would be interesting to develop a method for detection of thickened inner bone so that the condition’s prevalence can be quantified in a large sample like the UK Biobank and its genetic architecture be determined. This could help elucidate the etiology.

      The authors suspect that the effect of BMA on BMD may be biased in women; they should comment on those "with low BMD and high BMA" given that hyperostosis frontalis might be an issue. A strong effect of SNPs in the ESR1 chromosomal region might be akin to the above concern.

      Thank you for raising this point, which we have followed up with a new analysis.

      According to ICD-10 data in UK Biobank there are only N=105 individuals with an M85.2-diagnosed disorder. Given the total sample size of N=446,814 individuals with ICD-10 data, this would translate to a prevalence of 0.02%, which speaks for an underdiagnosis in this sample such that we cannot simply remove diagnosed individuals to control for a potential diagnostic confound.

      We have therefore taken a different approach to investigate this potential issue: As you elaborated in your previous comment, the prevalence of hyperostosis frontalis increases with age in females. The literature also suggests that prevalence rates do not differ between males and females in young age / prior to menopause. Therefore, we have studied the association between BMD and BMA for males and females separately, and in two age groups based on a median split of our sample (left plot: younger than 65, right plot: subjects older than 65). In these plots, the relatively large shift in female BMA and BMD is visible with the large yellow cloud at low BMA and high BMD in the left plot disappearing in the right plot. Despite this, we observe:

      (1) Associations in both males (blue) and females (yellow), suggesting that the associations were not driven only by females.

      (2) BMA-BMD association is largely similar across the two age groups.

      If we consider that the old age group is likely to contain more cases of hyperostosis frontalis than the young group, and if we consider that old-aged females are more likely to be in this condition than men of any age, then we would expect an impact of hyperostosis frontalis on our measures to result in observable differences in Author response image 1. This is not the case. We see global age-related shifts in BMA in women, yet the association with BMD remains similar across age groups.

      The technical properties of our neural network (trained on simulated data) makes it unlikely that frontal bone will contaminate the bone marrow detection globally (description above) and these results show that hyperostosis frontalis is not a considerable issue in our analysis.

      Author response image 1.

      Then, there is a perfect overlap of the BMA SNPs that are shared with BMD (497 of 500 variants), which may prove a "face validity" of the MRI-derived BMA. However, the BMD in the study was heel-derived eBMD - which is a good proxy for osteoporosis and is mostly driven by trabecular bone. Thus, there might be a concern that the BMA metrics capture some trabecular BMD.

      The reviewer is correct in pointing out that the BMA causal variants are a near-perfect subset of the BMD causal variants. The reviewer raises the concern that the BMA measurements may capture some trabecular BMD, however it should be noted that the correlation of effect sizes for the BMA/BMD overlapping causal SNPs is negative (-0.95). If our measure of BMA had been erroneously capturing trabecular BMD then we would expect to see a positive correlation of effect sizes for the BMA/BMD overlapping causal SNPs, not a negative one.

      Next, integrating mapped genes with existing bone marrow single-cell RNA-sequencing data revealed patterns of adipogenic lineage differentiation and lipid loading. The problem here is that the scRNAseq studies of the Bone Marrow niche are overwhelmingly mouse. The authors might wish to justify why they are relevant to humans (in the absence of the human-specific scRNAseq).

      We thank the reviewer for pointing this out. We noticed that, although Figure 4 and the Methods do explicitly state that the scRNAseq data is from mouse, it is not stated in the text of the Results. This is now corrected.

      The mouse is commonly used as the model organism for in vivo investigation of human phenotypes and bone marrow adiposity is no exception because, although mice have lower bone marrow adiposity than humans, the timing and sequence in bone marrow adiposity development are similar. BMA research makes extensive use of mouse models literature as exemplified by this review of research within the field (https://www.frontiersin.org/journals/endocrinology/articles/10.3389/fendo.2016.00127/full) and this article recent article (Koh et al. 2024. “Adult skull bone marrow is an expanding and resilient haematopoietic reservoir”. https://www.nature.com/articles/s41586-024-08163-9)

      We updated the results section (line 279):

      “Mesenchymal stem cells of the BM niche commit to either the adipogenic or the osteogenic lineage (Figure 4A) and both the number committing to the adipogenic lineage and their level of lipid-loading influences the total level of BMA. This aspect of BM biology is shared between humans and mice (29), so we made use of an existing mouse scRNAseq dataset of BM mesenchymal lineage cells (30) to study variation in the expression of BMA-associated genes as cells differentiate (Figure 4B).”

      For genetic correlation analysis, the authors selected 7 body and 6 brain traits. The latter traits reflect cognition (general cognitive ability and educational attainment) and brain-related disorders. This selection might seem arbitrary. The interpretation of genetic correlation with cognitive ability, education, and Parkinson's disease was attributed to the recently discovered vascular channels that link calvarial bone marrow to the meninges. This is a fascinating hypothesis, which requires functional proof. However, there might be simpler explanations. Thus, the diploe and the inner table of the calvarium are drained by the same veins as the dura. From the anatomy textbook, we know that diploic veins connect the pericranial and endocranial venous system through the skull.

      Whilst it is true that we did not systematically compare the results of the BMA GWAS to all potentially relevant brain and body phenotypes, we did use criteria to select the phenotypes we compared to. As stated in the manuscript (line 304):

      “We selected body traits (BMD, BMI, waist-to-hip ratio, systolic and diastolic blood pressure, type-2 diabetes, coronary artery disease) that have a logical connection to BMA given the mesenchymal stem cells origin of BM adipocytes and their role in bone, fat, and vasculature (29). For the brain, we selected traits reflecting cognition (general cognitive ability and educational attainment) and disorders that are prevalent in adulthood (insomnia, multiple sclerosis, Parkinson's disease, Alzheimer's disease) since it is primarily in adulthood that the adiposity of calvarial BM experiences a substantial change”

      We entirely agree that the suggestion that the genetic correlation between BMA and cerebral traits may be mediated by the vascular channels linking calvarial bone marrow to the meninges is merely a hypothesis. We have therefore updated the text of the Discussion (line 470):

      “We tentatively speculate that calvarial MALPs may be involved in sensing perivascular flows of CSF from the meninges to the BM and in influencing the BM’s hematopoietic response, and that this might be the basis of the observed genetic overlap between BMA and some cerebral traits. However, more conventional anatomical pathways may also be relevant, as the diploë and inner table communicate with meningeal and dural venous systems through diploic veins.”

      Reviewer #2 (Public review):

      Summary:

      This study develops a new artificial intelligence method for high-throughput analysis of skull bone marrow from MRI data, which may be useful for large-scale biological analyses. Using this method, the authors then attempt to estimate skull bone marrow adiposity (BMA) using T1-weighted signal intensity from MRI scans of ~33,000 people, followed by genome-wide association analysis; however, the approach is inadequate because T1-weighted signal intensity is not validated for measurement of bone marrow adiposity. If it could be validated, the study would be an important advance in understanding of bone marrow adiposity and skeletal biology.

      Strengths:

      This paper is well-written, and the figures are nicely presented. The neural network method used for analysing skull bone marrow is innovative, and the authors validate this through several approaches. Therefore, the authors have achieved the aim of developing a method for large-scale analysis of skull bone marrow from MRI data.

      The GWAS is reasonably well-powered and addresses potential ethnicity differences, with one GWAS done across white males and females, and a separate GWAS in non-white participants. The methodology also conforms to common GWAS standards, including for mapping genetic variants to candidate genes. Moreover, the study further investigates the biological roles of these genes by analysing their expression in single-cell RNA sequencing data.

      Weaknesses:

      The fundamental weakness is that T1-weighted MRI signal intensity (T1W) is used as an estimate of BMA, but it has never been validated for this. The authors show that this T1W parameter measures something that is heritable and can be compared between subjects, but they don't show that it actually measures (or even estimates) calvarial BMA. There is an attempt to do so by comparing the T1W parameter with data from quantitative T1 images: the authors show a reasonable correlation with some of the quantitative T1 image data. However, this still does not show that the parameter is measuring BMA; it could be measuring some other biological characteristic, but this remains unclear. So, there is a need to validate the T1W parameter against an established measure of BMA, such as the bone marrow fat-fraction or proton density fat fraction measured from multi-echo MRI analysis.

      Without validating this BMA measurement method, it is not possible to interpret the GWAS or other findings reported in the study.

      We reject this criticism.

      Although T1-weighted has not been validated as a quantitative measure of fat-fraction, there are several studies showing that it is a semi-quantitative measure of fat content (e.g. Loevner et al 2002, Shen et al 2013, Zhang et al 2020) and we also provide data that support this (figures S9-11).

      Semi-quantitative measures are used in many biomedical GWASes for instance even highly heritable neuropsychiatric disorders (such as schizophrenia and bipolar disorder) involve assessment by clinicians where the test-retest kappas are in the range 0.4-0.6.

      Further, we would suggest that the shortcoming of the imperfect correlation of T1w signal intensity with fat content is more than outweighed by our precision in identifying the calvarial BM cavity and the fact that the flat calvarial bone marrow has a wide range of adiposity in middle-aged and elderly individuals (compared to other bones). This lies at the root of why:

      We clearly recapitulate the known sex and age profiles, as well as the effect of HRT.

      We estimate high BMA heritabilities (43% in males and 23% in females)

      We find clear sex differences (which is a known feature of BMA biology)

      We identify a large number of the genes already known to affect BMA from earlier animal and cell work

      A noisy measurement of an entity with strong biological signal (a well-defined bone marrow cavity with variation in BMA across subjects) will often be more informative than a highly precise measurement of a poorly defined entity with little signal.

      A less critical weakness is that the GWAS has been done only on a single cohort, without replicating the findings in a follow-up cohort. For example, the authors could repeat their analysis on the remaining ~50,000 UK Biobank imaging participants for whom MRI data is now available. However, this would be pointless without knowing what biological characteristic(s) the T1W parameter is actually reflecting.

      We disagree with this comment. We separated the UKB data into discovery and replication sets prior to running the GWAS, so these datasets are independent:

      (1) Further, we ran the discovery (white british individuals) GWAS separately for males and females (prior to combining) and reported in the results section: “We found them to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      (2) We performed our replication GWAS in non-white British males and females. As noted in the results section: “Out of the 168 significant discovery SNPs, 62% replicated at P<.05, and 39% of the 41 lead SNPs replicated at P<.05 (Table S6). One locus replicated at genome-wide significance (P<5e-8). Furthermore, 92.7% of the lead SNPs of the discovery sample showed same effect direction in the replication sample (Table S5).”

      Reviewer #3 (Public review):

      Summary:

      This manuscript, "Estimating bone marrow adiposity from head MRI and identifying its genetic 2 architecture", brings together the groups of Drs. Kaufmann and Hughes in a tour de force work to develop an artificial neural network that localizes calvaria bone marrow in T1-weighted MRI head scans, with the goal of studying its composition in several large MRI datasets, and to model sex-dimorphic age trajectories, including the effect of menopause.

      Strengths:

      Bone marrow adiposity is a very active tissue with far-reaching implications for tissue crosstalk and human health than we had initially recognized. Although MRI has been used to measure BM, studies such as the one by these two groups are still lacking whereas very large datasets are analyzed using advanced AI machine learning tools coupled with genetic studies and a specific pathology. The groups had to develop new methods and new AI machine-learning tools for the imaging analyses.

      Weaknesses:

      Some aspects of the work that authors could add additional clarification.

      (1) Imaging Limitations: The authors provide an excellent overview and references supporting the use of MRI as a method for assessing marrow fat, particularly with some specific modifications. However, MRI images can be affected by various factors, including the presence of other tissues as well as specific MRI settings, which are much harder to precisely control when using different datasets.

      We thank the reviewer for his positive assessment of our review of methods.

      Regarding MRI settings: We agree with the reviewer that differences in scan protocols can create substantial differences in the resulting images between samples. Different tools exist for harmonization of imaging data across sites, but they usually operate on tabulated data and there is no one-size-fits-all approach yet [1]. Here, we took a different approach to prevent confounding bias: We generated a large set of simulated data for training of the neural network. The simulations circumvented potential issues emerging from confound biases in training sets that we might have seen had we had combined multiple samples with different scan protocols. Nevertheless, applied to real data the models may still face confound issues, such as better BMA estimates for some scan protocols over others. We have addressed these issues as follows: (1) Validation analyses (10 repeat scans of 30 individuals and twin pairs, figure 1d and 1e) are based fully on data that was acquired on the same scanner with the same protocol. (2) Analysis in UK Biobank included data from different scan sites albeit harmonized protocols. Here we accounted for scan site in all statistical models (including GWAS).

      (1) Dominik Kraft, Gloria Matte Bon, Édith Breton, Philipp Seidel, Tobias Kaufmann; Removing scanner effects with a multivariate latent approach: A RELIEF for the ABCD imaging data?. Imaging Neuroscience 2024; 2 1–7. doi: https://doi.org/10.1162/imag_a_00157

      Regarding the presence of other tissues: We recognise in the existing text of the results section that sometimes inner or outer table voxels are wrongly identified as bone marrow, but we show that this does not have a major impact on the correct identification of the bone marrow cavity. The existing text reads:

      “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap. However, this did not result in a corresponding fall in the ratio of the predicted intensity of BM to its true intensity, because the typical BM intensity was only marginally higher than the neighbouring bone intensity. This property of the typical relative intensities of these anatomic structures also explains why the intensity ratio at high overlap is not centred on 1: any misidentification of cortical bone as BM, will typically result in an underestimate of true BM intensity (Figure S2). The neural network performed well and intensity ratios were in the range 0.9-1.1 for the vast majority of head types (Figure S3).”

      Also note that we implement a number of QC measures to exclude scans where there is evidence that we may have failed to correctly identify the bone marrow cavity. The existing text reads:

      “We used two additional QC metrics to filter out calvaria where BM location was likely to have failed. First, we set an upper limit of 30 on the standard deviation of the intensity of the outer table as scans with higher values were clear outliers and were probably cases where the location of both outer table and BM has failed (Figure S6). Second, for each calvarium, we computed the Mahalonobis distance for all vertices in the two dimensions “first layer of the BM” and “BM intensity” (Figure S7). By manual inspection we found that data points with MD > 25 often had errors in BM layer identification, typically where the network had erroneously predicted a higher and more intense layer to be the BM. We considered a calvarium as failing this QC criterium if more than 0.5% of vertices have MD > 25. This criterium is very strict as errors on only 0.5% of data points in a calvarium would not significantly affect the average BM intensity for a calvarium.”

      (2) The specific density of cranial bones as it relates to the types of bone marrow: Cranial bones are extremely dense structures, which naturally interfere with MRI imaging. While it is thought that cranial bones have mostly "red bone marrow", this is only true for a short time in humans. How sensitive is their system in differentiating between red and yellow BM?

      We implemented several measures to ensure that our method would be robust to anatomical variation between individuals. As noted in the current version of the Methods section: “In order to train the neural network model, we generated a large synthetic dataset of intensity arrays, with known boundaries between anatomical structures, by simulating the thickness and intensity of the different structures located between the outer skin and the subarachnoid space. The simulation incorporated the following real-world complexities:

      Different anatomical architectures (skin, subcutaneous fat, aponeurosis, outer table, BM, inner table, dura mater, arachnoid space), including when a structure is not present throughout the calvarium

      Variation in thickness and intensity between vertices (on the same calvarium)

      A wide variety of different calvarium types with different combinations of levels of BM adiposity, bone thickness, and subcutaneous adiposity.”

      Further, as noted in the Results section:

      “We evaluated the performance of the neural network on simulated data using two metrics (Figure 1B): 1. the overlap between the predicted and the true BM location, and 2. the ratio between the predicted intensity of the BM and the true intensity of the BM”. The accuracy in localising the bone marrow layers was good, with the only exception being: “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap”

      We also validated our procedure on real data (see Results section, subsection “Procedure validation on real data and heritability estimate”). Briefly, we checked the accuracy of our method using a dataset from the Consortium for Reliability and Reproducibility, a twin dataset from the Human Connectome Project and by manually checking many hundreds of UKBiobank scans.

      We are thus confident that we accurately identify the bone marrow cavity irrespective of whether the bone marrow is red (low adiposity) or yellow (high adiposity).

      (3) Both items above are further complicated by aging, but aging is not a linear event as we have learned. There are specific bursts of aging in humans around the age of 45 and early 60s. How do the system and model predict or incorporate these peaks of aging? It seems from the data shown that aging is reflected more as a linear phenomenon. Is this because additional aging datasets are needed?

      We agree with the reviewer that ageing probably occurs in bursts rather than being a linear process. We do see a non-linear relationship between age and BMA in our data (see figure 2B), with a more rapid rise in BMA between the ages of 45 and 65, than later in life (in women). As a result of this, when we model BMA using regression, we use orthogonal polynomials of degree 2 which allows for a non-linear relationship. However, we cannot observe bursts of BMA increase in our data because it is cross-sectional. Longitudinal data would be required to obtain information on the nature and timing of any bursts in bone marrow adiposity.

      (4) The authors describe in richness of detail their AI learning programming and how it extracted the data from datasets. The authors also show some important correlations with specific genes, SNPs. What is not clear is how conditions such as anemia for example. An expected finding would be that patients with chronic anemia have lower bone marrow (BM) signal intensity on MRI scans than healthy people. This is because the signal intensity of BM depends on the fat-to-cell ratio in the tissue.

      We agree with the reviewer that conditions affecting the bone marrow niche have a potential to affect and be affected by bone marrow adiposity, with leukemia being a known example. This is why we believe that a method, such as the one we present here, has the potential to be useful in several biomedical fields (hematology and osteology).

      Furthermore, patients with a host of musculoskeletal disorders ranging from osteopenia to osteoporosis, sarcopenia, and osteosarcopenia will also have altered MRI scans. When using such large datasets how did the authors control or exclude these pathological conditions, or were all these conditions likely present?

      We did not exclude specific pathologies. We were careful to train our NN model on a wide variety of skull thicknesses, bone marrow adiposity, and subcutaneous adiposity and to evaluate the performance of the model on simulated and real datasets (see answer to your point 2).

      Reviewer 1 raised the issue of individuals displaying Hyperostosis frontalis interna (thickening of the inner table) and we recognize that in extreme case of this condition, where there is a major change in the anatomy of the calvarial bone, our method would probably not correctly localise the bone marrow. However, such extreme cases are rare and thus would not have a major impact on our results derived from over thirty thousand individuals. We demonstrate this with an extra analysis performed in response to the point about hyperostosis frontalis interna made by reviewer 1.

      (5) Some of the genes and SNPs although significant showed very small correlations. What is their likely physiological significance?

      Bone marrow adiposity is a polygenic trait and we have identified 41 statistically significant loci. We had a discovery sample of approximately 30k individuals which is modest for a GWAS study, so these 41 loci are a lower bound on the number of genes influencing the BMA trait. When a large number of genes influence a trait, the effect size of an individual gene is typically relatively small. However, the SNP heritability estimates of 31.5% indicates that we are able to explain approximately one third of the phenotypic variation with the effect sizes estimated by our GWAS: this is quite a high fraction relative to many other GWASs of biomedical traits.

      (6) The authors could use this excellent manuscript to expand their discussion to include the need for studies like theirs to be also complemented by multi-OMICS studies that will include proteomics and lipidomics of BM, bones, and muscles.

      We agree with the reviewer and hope that such studies will be undertaken in the future. We attempted to point in this direction in the last sentence of the Discussion (line 517): “Future studies can build on our developments to further validate the proposed measure of bone marrow composition and to study its effect on bone, blood, and brain”. Word count limits prevented us from further expanding on the specific kinds of studies that should be performed.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      More moderate concerns include:

      (1) In the "Genetic correlation and overlap" part, it is unclear why a high effect correlation with high SNP overlap is suggestive of vertical pleiotropy, while "moderate overlap and a low correlation of effect sizes ... are more indicative of horizontal pleiotropy". This is not intuitive.

      In vertical pleiotropy (genetics > phenotype A > phenotype B): genetics drive phenotype A and phenotype A drives phenotype B (B has few direct genetic drivers of its own). If we perform a GWAS of phenotype A and a GWAS of phenotype B, one would expect to see a high overlap in the causal variants (because phenotype B is largely indirectly determined by the genetics of phenotype A) and the correlation should be high (either negative or positive) because there is a cause-effect relationship between A and B (whether phenotype A has a positive or negative effect on phenotype B).

      In horizontal pleiotropy: the A and B phenotypes share the same genetic loci (but have no phenotypic influence on each other). In this case, we would observe high overlap in the associated loci, but one would not expect to see a high correlation in effect sizes (across loci) because there is no a priori reason to expect that genes associated with both phenotype A and phenotype B would have a consistent (negative or positive) effect across loci. For example, gene X may increase A and B, whereas gene Y may increase A but decrease B.

      We have made a small update to the relevant part of the Results section and otherwise rely on the explanation above (which will be publicly available along with the manuscript):

      Line 325: “These patterns of high overlap and high effect correlation within this overlap are suggestive of vertical pleiotropy i.e. a molecular mechanism influencing one trait, that in turn influences a second trait, such that most of the variants driving the first trait either have the same or the opposite direction of effect on the second trait.”

      (2) "possible causal effect of BMA on cognition" asks for a formal analysis, like Mendelian randomization.

      Given the current wording, the reviewer is justified in asking for a formal analysis. Since we did not perform this analysis, we have changed the wording:

      Line 434: “Since these are two highly correlated traits [36], a high overlap and correlation of genetic effects for both traits with BMA may be consistent with the hypothesis that BMA could have a causal effect on cognition (Table 1).”

      (3) ll. 132-134: please reword this sentence for clarity: "Using the network-estimated location of the BM within ... averaged these across all datapoints...". Please define threshold of desirable overlap between the predicted BM and the true BM (=0.7?).

      Background: The model predicts the BM localisation for a datapoint (which interval of layers of the 50 layers is bone marrow). We tested the model on a wide variety of simulated data and aim for the overlap to be as close to 1 as possible, but some error is inevitable. We found that the average overlap between the true and predicted bone marrow was only below 0.7 when the bone layer is only 4 mm thick (meaning that the bone marrow is only 1-2 mm thick). This demonstrates the high accuracy of our method in identifying a very small anatomical feature.

      When applying our method to real data, we do not know the truth and therefore cannot compute the overlap between the predicted and true value. It is therefore not possible to identify datapoints where the overlap is poor (e.g. lower than 0.7) and filter them out.

      Given the above, we struggle to understand in what way an overlap threshold is relevant to how we compute the signal intensity for a datapoint. Nevertheless, we recognize that the sentence pointed to by the reviewer is poorly formulated and have tried to make it clearer:

      Line 130-133: “To obtain the BM signal intensity for an individual datapoint of the calvarium, we used the network model to estimate the location of the BM within the datapoint’s intensity array and averaged these BM intensities to get the BM intensity for that datapoint. Then, we averaged these datapoint intensities across the calvarium to produce the global BMA measure for the scan.”

      In the GWAS Results, please clarify the phrases - what was "significantly genetically correlated (Rg=.94)" (also, l. 402, "genetic correlation between the sexes" - in what?).

      Genetic correlation is a statistical measure that quantifies the extent to which two traits (or the same trait in two different cohorts) are influenced by the same genetic factors. Simply put, it is the effect sizes of the SNPs in the two GWASs of interest that are correlated (after correcting for confounding effects, such as linkage desequilibrium). When comparing two GWASs, the standard formulation is to refer to their “genetic correlation”. We made a modification to the text to clarify this:

      Line 239: “We found the male and female GWASs to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      "a more than two-fold difference between the sexes" - in which metric?

      We feel that what is being compared is stated clearly in the original sentence:

      Line 262: “A comparison of the male and female effect sizes of the top lead SNPs of each locus revealed 6 loci in which there is a more than two-fold difference between the sexes (loci 10, 18, 26, 30, 32, 37 in Table S5)”.

      (4) Also In GWAS Results, a locus Dlx5 is called "SHFM" in the Supplementary Table.

      Background:

      We identified 41 genome-wide significant loci and named the locus after the gene closest to the top lead SNP (bold in Figure 3C). Other genes in each locus for which genome-wide significant SNPs were eQTLs, are listed below the closest gene in normal font (Figure 3C).

      In table S5, we report details of the top lead SNP for all 41 loci. We report only the nearest gene to the top lead SNP.

      For locus 14, SHFM1 is the closest gene to the top lead SNP whereas DLX5 and DLX6 are genes in the locus for which genome-wide significant SNPs were eQTLs. This explains why DLX5 appears under SHFM1 in Figure 3C, but does not appear in Table S5.

      (5) Please reword MRI jargon - "Dixon method", vertix - should be introduced, as well as abbreviation "KDE".

      We had recognised that the word “vertex” would be confusing and had replaced it by datapoint, but had unfortunately missed one occurrence in the text. This is now corrected.

      Thank you for pointing out the lack of introduction of the term “KDE”. This was only explained in the supplementary materials, but has now been added to the main text:

      Line 500-508: “To ensure between-subject comparability, we used the intensity normalised nu.mgz volume output by FreeSurfer. We validated this approach through comparison with well-established intensity normalization methods; Kernel Density Estimation (KDE), WhiteStripe (WS), Gaussian Mixture Model (GMM), Fuzzy C-Means (FCM), and Z-score normalization (ZS). We found the highest test-retest reliability with our approach (Figure S9), and, together with KDE (based on reference signal intensity in WM), the highest correlation with quantitative T1 relaxation maps (Figure S10).”

      Reviewer #2 (Recommendations for the authors):

      (1) This would be an extremely useful advance for the bone and BMA fields if only it could be confirmed that the T1W signal intensity is actually measuring BMA in some meaningful way. Or, even if not BMA, to confirm what other biological characteristic(s) it is in fact capturing. This is essential for interpreting the findings.

      (2) I note that you have compared the normalized T1W parameter with quantitative T1 data (e.g. Figure S10). However, this doesn't address the fundamental issue, because even these quantitative T1 data (e.g. from MP2RAGE) may not be measuring calvarial BMA. T1W sequences have been used to estimate BM cellularity (if not BMA directly) but are not nearly as precise as water-fat imaging. For example, one study found a reasonable correlation (0.71) between T1 relaxation times and BM fat (https://www.nature.com/articles/s41598-019-57030-5). So, if your normalized T1W parameter shows a correlation of -0.44 with the T1 MP2RAGE MRI signal (Figure S10), what does this mean in terms of how well your parameter reflects the actual BMA adiposity? We can't know this, because we also don't know if the T1 MP2RAGE signal reflects calvarial BMA.

      (3) I think my recommendations are clear from the public review. Ideally, you would be able to compare the skull BM normalized T1W parameter with PDFF data that have T2* correction (since the skull BM cavity is quite small and so may suffer from T2* effects relating to tissue inhomogeneity). But even if you had only dual-echo BMFF data, this would still be much more informative than relying only on T1 data. I hope this can be done so that the findings of the study can be properly interpreted.

      As explained above, we reject this reviewer’s claim that T1-weighted signal intensity cannot be used to perform a GWAS of BMA: other studies have shown that T1-weighted signal intensity is a semi-quantitative measure of fat fraction, we have performed extra analyses that confirm this, and our results further demonstrate this. For further detail on why we reject this criticism, see our response to this reviewer’s comments.

    1. eLife Assessment

      This work presents valuable new data on the role of D-Serine and how it competes with its stereoisomer L-Serine to influence metabolism. The work presents a variety of convincing experimental data combined with simulated results to investigate the mechanisms focused on one-carbon metabolism, which is relevant for several research fields. However, some claims are only partially supported by data, and critical areas comparing L- vs D-Serine and further mechanistic studies are required. Furthermore, while the work has potential for various fields, the work has only been studied in a limited cell type and context.

    2. Reviewer #1 (Public review):

      Summary:

      The authors demonstrate the stereoselective role of D-serine in 1C metabolism showing that D-serine competes with L-serine and inhibits mitochondrial L-serine transport. They observe expression of 1C metabolites in their metabolomics approach in primary cortical neurons treated with L-serine, D-serine and mixture of both. Their conclusions are based on the reduction in levels of glycine, polyamines and their intermediates and formate. Single cell RNA sequencing of N2a cells showed that cells treated with D-serine enhanced expression of genes associated with mitochondrial functions such as respiratory chain complex assembly and mitochondrial functions with downregulation of genes related to amino acid transport, cellular growth and neuron projection extension. Their work demonstrates that D-serine inhibits tumor cell proliferation and induces apoptosis in neural progenitor cells highlighting the importance of D-serine in neurodevelopment.

      Strengths:

      D-amino acids do not merely function as ligands at receptors but have underlying roles in signaling and metabolism. These roles are just beginning to be uncovered. The authors elucidate the metabolic role of D-serine in the context of neuronal maturation by its suppression of mitochondrial L-serine availability for SHMT2 and 1C flux. This is the strength of the manuscript. The implications for the metabolic role of D-serine in neurons is a highlight and underlines its roles in neuronal metabolism.

      Weaknesses:

      These are some minor issues that come up on critical assessment of the manuscript and is only intended to strengthen the manuscript. The comments below are based on the revisions made by the authors including the justification of their approach and rebuttal.

      (1) Kinetic assessment of D-serine versus L-serine: The authors have made reference to prior work by Miyamoto et al. and justify their rationale. This is acceptable.

      (2) Molecular Dynamics simulations while a good first step in modeling interactions at the active site, relies on force fields. The authors state that any elaborate study into longer simulations is beyond the scope and their simulations data are supported by other experimental work. This is justified.

      (3) The use of N2a cell line is also justified to reflect the proliferative nature of immature neurons.

      (4) With regards to caspase 3 comment, the whole blot is convincing and shows cleaved caspase-3 band at approx. 15 kDa.

      (5) Scale Bars are clearly visible and Fig S6 which was earlier S5 is legible. If possible, the authors can include an magnified inset in the merged image to show the clear activation of caspase-3.

      (6) Issue of phosphatidyl serine standard in LC-MS is justified by the use of L-serine standard due to lack of availability.

      (7) The authors mention about enantiomeric shift of serine metabolism during neural development which appears to be a discussion of prior published data from Hubbard et al 2013, Burk et al 2020, and Bella et al 2021 in Supplementary Figure panels 8 A-E.<br /> The authors justify by citing references to the work which may be acceptable and also the current norms of publication. This reviewer felt contrary to the fact, however it is left to the editors to make a decision on this.

      (8) The discussion section has been substantially revised and now reads well.

      (9) The relevant references have been cited. In doing so, the work integrates and elucidates a mechanistic and functional role of D-serine in neurons.

      (10) Figure S7A in the revised manuscript shows the specificity of D-serine in the cleaved caspase-3 assay which is informative.

      Comments on revised version.

      This reviewer is satisfied by the effort made by the authors based on the prior comments raised.

    3. Reviewer #2 (Public review):

      Summary:

      This study by Suzuki et al. reports an interesting stereo-selective role of D-serine in regulating one-carbon metabolism during neurodevelopment to adapt the functional transition, probably through the competition with mitochondrial transport of L-serine. The authors provide a multi-layered set of evidence, including metabolomics, enzyme assays, mitochondrial transport competition and functional assays in immature/neural progenitor cells, to build up a conceptual integration of D-serine as both a neurotransmitter and a metabolic regulator in central neural system, which raises a broad potential interest to the neuroscience and metabolism communities.

      Strengths:

      This work provides a conceptual advance that D-serine is not only serves as a traditional neurotransmitter in central neural system but also critically contributes to metabolic regulation of neural cells. The authors performed solid metabolomic assays to validate the suppressive effect of D-serine on one-carbon metabolic pathway, providing some evidence that D-serine competitively inhibits mitochondrial serine transport, but not directly impairs SHMT2 enzymatic activity. All these data indicate a critical role of D-serine synthesis during neural maturation and suggest a potential translational strategy for targeting serine metabolism in neural tumors.

      Comments on revised version.

      My previous concerns have been appropriately addressed or discussed in this revised version of manuscript. I have to say that, at this stage, I have no further questions.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript presents a comprehensive and well-executed investigation into the metabolic role of D-serine in the central nervous system. The authors provide solid evidence that D-serine competitively inhibits mitochondrial L-serine transport, thereby impairing one-carbon metabolism. This stereoselective mechanism reduces glycine and formate production, suppresses cellular proliferation, and induces apoptosis in immature neural cells and glioblastoma stem cells. Developmental analyses further reveal a physiological enantiomeric shift in serine metabolism during neurogenesis, aligning with the transition from proliferation to maturation. Overall, the study bridges developmental neurobiology, cancer metabolism, and amino acid transport, uncovering a previously unrecognized metabolic function of D-serine beyond its role in neurotransmission.

      Strengths:

      (1) The discovery that D-serine inhibits one-carbon metabolism by competing for mitochondrial L-serine transport-rather than through enzymatic inhibition or receptor-mediated signaling-represents a significant and previously underappreciated mechanism. This finding has broad implications for understanding metabolic regulation during neurodevelopment and offers potential relevance for targeting metabolic vulnerabilities in cancer.

      (2) The authors integrate metabolomics, mitochondrial transport assays, molecular dynamics simulations, genetic and pharmacologic perturbations, transcriptomics, and both in vitro and ex vivo models. The breadth of experimental approaches, combined with the coherence of the findings across systems, provides strong support for the central conclusions and enhances the overall impact of the study.

      (3) The temporal shift in D-/L-serine levels during neurodevelopment is elegantly linked to the transition from proliferative to mature neuronal states. The selective vulnerability of neural progenitors and tumor cells-contrasted with the resistance of mature neurons-highlights a biologically meaningful and potentially targetable metabolic distinction.

      Weaknesses:

      (1) While the authors attribute D-serine's metabolic effects to competition with mitochondrial L-serine transport, the specific identity of the transporter(s) mediating this process remains undefined. This represents a meaningful mechanistic gap, as the central conclusion depends on D-serine limiting mitochondrial L-serine availability to inhibit one-carbon metabolism.

      (2) The effective concentrations of D-serine used in vitro (IC₅₀ ≈ 1-2 mM) exceed typical brain levels (~0.3 mM). While the authors acknowledge this, a more focused discussion on whether higher local D-serine concentrations could arise in specific microenvironments-such as synaptic compartments, tumor niches, or pathological states-would help contextualize the in vitro findings and strengthen their physiological relevance. For example, disruptions in D-serine clearance or altered expression of serine racemase and transporters in disease contexts could lead to localized accumulation. Moreover, differences between extracellular and intracellular D-serine pools-and the mechanisms governing their regulation-may further influence its metabolic impact in vivo.

      (3) While the manuscript focuses on neural stem/progenitor cells and neural tumors, it remains unclear whether the anti-proliferative effects of D-serine are specific to neural lineages or extend to other highly proliferative non-neural cell types. A brief discussion addressing this point would help clarify the scope of D-serine's metabolic impact and whether its mechanism of action reflects a unique vulnerability in neural cells or a more general feature of proliferative metabolism. This distinction is particularly relevant for assessing the broader therapeutic potential of targeting mitochondrial L-serine transport.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors demonstrate the stereoselective role of D-serine in 1C metabolism, showing that D-serine competes with L-serine and inhibits mitochondrial L-serine transport. They observe expression of 1C metabolites in their metabolomics approach in primary cortical neurons treated with L-serine, D-serine, and a mixture of both. Their conclusions are based on the reduction in levels of glycine, polyamines, and their intermediates and formate. Single-cell RNA sequencing of N2a cells showed that cells treated with D-serine enhanced expression of genes associated with mitochondrial functions, such as respiratory chain complex assembly, and mitochondrial functions, with downregulation of genes related to amino acid transport, cellular growth, and neuron projection extension. Their work demonstrates that D-serine inhibits tumor cell proliferation and induces apoptosis in neural progenitor cells, highlighting the importance of D-serine in neurodevelopment.

      Strengths:

      D-amino acids are a marvel of nature. It is fascinating that nature decided to make two versions of the same molecule, in this case, an amino acid. While the L-stereoisomer plays well-known roles in biology, the D-stereoisomer seems to function in obscurity. Research into these novel signaling molecules is gathering momentum, with newer stereoisomers being discovered. D-serine has been the most well-studied among the different stereoisomers, and we still continue to learn about this novel neurotransmitter. The roles of these molecules in the context of metabolism is not well studied. The authors aim to elucidate the metabolic role of D-serine in the context of neuronal maturation with implications for 1C metabolism and in cell proliferation. The metabolic role of these molecules is just beginning to be uncovered, especially in the context of mammalian biology. This is the strength of the manuscript. The authors have done important work in prior publications elucidating the role of D-amino acids. The advancement of the field of D-amino acids in mammalian biology is significant, as not much is known. The presentation of RNA seq data is a valuable resource to the community, however, with caveats as mentioned below.

      Weaknesses:

      The following are some of the issues that come out in a critical reading of the manuscript. Addressing these would only strengthen and clarify the work.

      (1) Kinetic assessment of D-serine versus L-serine: While the authors mention that D-serine is not a good substrate for SHMT2 compared to L-serine, the kinetic data are presented for only D-serine. In a substrate comparison with an enzyme, data must be presented for L-serine as well to make the conclusion about substrate specificity and affinity. Since the authors talk about one versus another substrate, there needs to be a kinetic comparison of both with Km (affinity). (Ref Figure 2 panel).

      We agree with the reviewer that kinetic parameters for l-serine are important for evaluating the substrate specificity of SHMT2. The hydroxymethyl-transferase activity of SHMT2 toward l-serine has been previously characterized by our co-author Tetsuya Miyamoto (Miyamoto et al., FEBS Journal, 2024; PMID: 37700610), which is appropriately cited in the manuscript (line 143). In that study, the kinetic parameters for l-serine were determined, with a Km of 0.07 ± 0.009 mM and a Kcat of 33.5 ± 0.9 min<sup>-1</sup>. The strong chiral selectivity of SHMT2 for the l-enantiomer in the hydroxymethyl-transferase reaction was also demonstrated in that work. On the other hand, the primary aim of the present study is different from characterizing d-serine as a catalytic substrate for SHMT2. Rather, our goal was to determine whether d-serine interferes with the hydroxymethyl-transferase reaction of SHMT2 by interacting with the l-serine binding site. Accordingly, the analyses shown in Fig. 2B-D were designed to evaluate whether d-serine could structurally occupy or interfere with the l-serine binding pocket of SHMT2. Therefore, our experiments focused on assessing the potential inhibitory effect of d-serine rather than performing a full kinetic comparison of d-serine and l-serine as substrates.

      (2) Molecular Dynamics simulations, while a good first step in modeling interactions at the active site, rely on force fields. These force fields are approximations and do not represent all interactions occurring in the natural world. Setting up the initial conditions in the simulations can impact the final results in non-equilibrium scenarios. The basic question here is this: Is the simulated trajectory long enough so that the system reaches thermodynamic equilibrium and the measured properties converge? Prior studies have shown mixed results with the conclusion that properties of biological systems tend to converge in multi-second trajectories (not nanosecond scales as reported by the authors) and transition rates to low probability conformations require more time. (Ref Figure 2C).

      We thank the reviewer for raising the important point regarding the limitations of molecular dynamics (MD) simulations, including the dependence on force fields and the potential effects of simulation length and initial conditions. We agree that MD simulations represent approximations of molecular behavior and that longer trajectories may be required to fully explore rare conformational states in biological systems.

      In the present study, however, the MD simulations were not intended to provide a comprehensive thermodynamic description of SHMT2 conformational dynamics. Rather, they were used as a structural assessment to evaluate whether d-serine could plausibly occupy the canonical l-serine binding site of SHMT2. As shown in Fig. 2BC and supplementary movie 1, the simulations did not support stable occupation of the l-serine binding pocket by d-serine. Importantly, this structural observation is consistent with our biochemical data showing that d-serine does not inhibit the hydroxymethyl-transferase activity of SHMT2 when l-serine is used as the substrate (Fig. 2D). Together, these results indicate that the inhibitory effect of d-serine on one-carbon metabolism is unlikely to be mediated through direct inhibition of SHMT2.

      As the reviewer correctly notes, it remains possible that d-serine interacts with SHMT2 at sites distinct from the canonical l-serine binding pocket and could exert potential allosteric effects. Indeed, previous work by Miyamoto et al. (FEBS Journal, 2024) demonstrated that SHMT2 exhibits dehydratase activity toward d-serine. However, this reaction is not directly linked to mitochondrial one-carbon metabolism. Therefore, further extensive simulations exploring alternative conformational states or potential allosteric interactions would extend beyond the scope of the present study.

      Importantly, the key conclusions of this study do not rely solely on MD simulations but are supported by multiple independent experimental approaches, including metabolomics, enzymatic assays, and mitochondrial transport analyses.

      (3) The authors use N2a cell line to demonstrate D-serine burden on primary cortical neurons. N2a is an immortalized cell line, and its properties are very different from primary neurons. The authors need to mention a rationale for the use of an immortalized cell line versus primary neurons. The transcriptomic profile of an immortalized cell line is different compared to a primary cell. Hence, the response to D-serine may vary between the two different cell types.

      We thank the reviewer for raising this important point regarding the differences between immortalized cell lines and primary neurons. As the reviewer notes, N2a cells and primary cortical neurons (PCNs) differ in several aspects, including their degree of differentiation and proliferative capacity. We appreciate the opportunity to clarify the rationale for using both systems in this study.

      One-carbon metabolism is known to be particularly active in highly proliferative or relatively undifferentiated cells. Primary cortical neurons are initially obtained as immature neuronal populations and gradually undergo maturation during culture (Fig. S6). In our experiments, we observed that sensitivity to d-serine and dependence on one-carbon metabolism were primarily evident in immature neuronal states rather than in fully mature neurons (Fig. 4EF).

      In this context, immature PCNs share certain metabolic characteristics with proliferative neural cell lines such as N2a cells. Consistent with this idea, the inhibitory effects of d-serine on one-carbon metabolism and cell proliferation were observed in both immature PCNs and N2a cells (Fig. 1E–G, Fig. 2E, Fig. 3AB, and Fig. 4F). Thus, the use of N2a cells provides a complementary experimental model for studying the metabolic vulnerability of immature neural cells that depend on one-carbon metabolism. Importantly, the key findings were consistently reproduced in primary cortical neurons, supporting the physiological relevance of the observations made in N2a cells. For clarity, we added descriptions in lines 107-108 and 167-168 in our revised manuscript.

      (4) In Figure 4D, the authors mention that D-serine activates the cleavage of caspase 3. Figure 4D shows only cleaved caspase 3 as a single band. They need to show the full blot that contains the cleaved fragments along with the major caspase 3 band.

      In our experiments, we used an antibody that specifically recognizes cleaved caspase-3 and does not recognize full-length caspase-3 (Cell Signaling Technology, anti-cleaved caspase-3 antibody, clone 5A1E). Therefore, the Western blot detects only the cleaved caspase-3 fragment at approximately 17 kDa, which appears as a single band in the blot (please see Author response image 1). For clarity, we have also included the antibody information in the Western blot section in the revised Materials and Methods (lines 501-502).

      Author response image 1.

      An original image of western blot for Fig. 4D. An arrow indicates the bands of cleaved caspase-3 (17 kDa).

      (5) In Figure panel 4, the authors use neural progenitor cells (NPCs). They need to demonstrate that the population they are working with is NPCs and not primary neurons. There must be a figure panel staining for NPC markers like SOX2 and PAX6. Also, Figure S5 needs to be properly labeled. It is confusing from the legend what panels B-E refer to? Also, scale bars are not indicated.

      We thank the reviewer for pointing out these issues. First, we have relabeled and rearranged the figure panels in revised Figure S6B-E, and revised the figure legend to provide a clearer description of each panel. We have also added scale bars to the revised images.

      To confirm the identity of neural progenitor cells (NPCs), we performed immunostaining for Nestin, a well-established marker for NPCs, instead of Sox2 and Pax6 (Fig. S6B). Nestin is widely used as a marker for NPCs(Bernal and Arranz, 2018; Bott et al., 2019; Lendahl et al., 1990), and is known to be co-expressed with Sox2 in the mouse embryonic brain(Graham et al., 2003). In addition, Nestin has been identified as a Pax6-bound gene associated with the transcriptional program regulating neural progenitor identity (Thakurela et al., 2016). Consistent with these observations, analysis of published scRNA-seq data from the developing mouse brain (Bella et al., 2021) shows that the RNA expression profile of Nestin (Nes) closely parallels those of Sox2 and Pax6 (new Fig. S9C). Based on these lines of evidence, we consider Nestin-positive cells in our cultures to represent NPCs. Importantly, Nestin-positive cells were co-stained with cleaved caspase-3, suggesting that the apoptotic population corresponds to NPCs (Fig. S6B). Together, these observations support that the apoptotic cells observe in our culture correspond to NPCs rather than differentiated neurons.

      (6) In Supplementary Figure panel 7F, the authors mention phosphatidyl L-serine and phosphatidyl D-serine. A chromatogram of the two species would clarify their presence as they used 2D-HPLC. On an MS platform, these 2 species are not distinguishable. Including a chromatogram of the 2 species would be helpful to the readers.

      We thank the reviewer for this helpful suggestion. As the reviewer correctly noted, mass spectrometry alone cannot distinguish phosphatidyl-d-serine and phosphatidyl-l-serine. To quantify these species separately, lipids were first extracted from cells using the Bligh and Dyer method, followed by phospholipase D treatment to cleave the serine moiety from phosphatidylserine (a schematic of the procedure is shown in Fig. S8D). The released serine was then derivatized with NBD-F and analyzed by 2D-HPLC for enantioselective separation and quantification of d- and l-serine, as described in the Materials and Methods section (“Quantification of glycine and serine enantiomers”). Phosphatidyl-l-serine (Sigma-Aldrich: P0474) was used to generate a standard curve for quantification (Fig. S8E). Because phosphatidyl-d-serine is not commercially available, and because the peak height of free d-serine is equivalent to that of free l-serine in the chromatograms of our 2D-HPLC system, the same standard curve was used to estimate phosphatidyl-d-serine levels.

      As requested by the reviewer, we have now added representative chromatograms of d- and l-serine derived from phosphatidyl-serine in NPC samples (new Fig. S8F).

      (7) The authors mention about enantiomeric shift of serine metabolism during neural development, which appears to be a discussion of prior published data from Hubbard et al, 2013, Burk et al, 2020, and Bella et al, 2021, in Supplementary Figure panels 8 A-E. This should not be presented as a figure panel, as it gives the false impression that the authors have performed the experiment, which is clearly not the case. However, its discussion can well serve as part of the manuscript in the discussion section.

      We appreciate this helpful suggestion. The datasets used in Fig. 4K and Fig. S9 were derived from previously published transcriptomic studies (Hubbard et al, 2013; Burk et al, 2020; Bella et al, 2021). While these studies reported transcriptomic profiles during neuronal development in vitro or in vivo, they did not specifically analyze serine metabolism, one-carbon metabolism, or d-serine biosynthesis, which are the focus of the present study. Therefore, we reanalyzed these publicly available datasets from a metabolic perspective, focusing on genes involved in serine metabolism and one-carbon metabolism. This re-analysis allowed us to examine a developmental shift in serine enantiomer metabolism, we presented the results as figure panels rather than simply citing the datasets in the Discussion. Re-analysis of publicly available transcriptomic datasets to address new biological questions has become a common approach in genomics and transcriptomics studies.

      To avoid the impression that these experiments were performed in this study, we have clearly indicated the original references and clarified this point in the figure legends (Fig. 4K and Fig. S9). These panels are intended to provide supportive evidence for the developmental shift in serine enantiomer metabolism discussed in this study.

      (8) The entire presentation of the section on enantiomeric shift of serine metabolism during neural development (lines 274-312) is a discussion and should be part of the discussion section and not in the results section. This is misleading.

      We thank the reviewer for this comment. This point overlaps with the concern raised in comment 7. Please see our response to comment 7 for a detailed explanation and the revisions made in the manuscript.

      (9) The discussion section is not well written. There is no mention of recent work related to D-serine that has a direct bearing on its metabolic properties. In the discussion section, paragraph 1, the authors mention that their work demonstrates the selective synthesis of D-serine in mature neurons as opposed to neural progenitor cells. This concept has been referred to in prior publications:

      (a) Spatiotemporal relationships among D-serine, serine racemase, and D-amino acid oxidase during mouse postnatal development. PMID:14531937.

      (b) D-cysteine is an endogenous regulator of neural progenitor cell dynamics in the mammalian brain. PMID:34556581.

      We thank the reviewer for this helpful suggestion and for drawing our attention to these studies. We have revised the Discussion to better place our findings in the context of previous work related to d-serine.

      Specifically, we have added references describing the spatiotemporal relationship between d-serine and serine racemase (Srr) during brain development (PMIDs 14531937 and 33592203) (line 333). These studies highlighted the tissue-level (PMID 14531937) and cellular-level (PMID: 33592203) relationship between Srr expression and development. These studies highlighted the role of d-serine in supporting the functional maturation of neurons during postnatal development. In contrast, our study addresses a complementary question of why d-serine is NOT present during embryonic and early postnatal stages of brain development, when proliferative metabolic activity is high. This question is fundamentally different from the previous reports. Our point is that d-serine is not favorable because it interferes with one-carbon metabolism, which is essential for cell proliferation. Therefore, this concept has not been referred to in prior publications. To clarify this point, we added the following sentence to the Discussion (line 325-328): “In addition to the known role of d-serine in the functional maturation of differentiated neurons, our findings highlight a previously unrecognized, stereoselective regulation of cellular metabolism by d-serine, and provide a rationale for its selective synthesis in mature neurons where proliferative metabolic activity is no longer required’.

      We also appreciate the reviewer bringing our attention to the study describing d-cysteine as a regulator of neural progenitor cell dynamics (PMID: 34556581). We have added the description regarding the overlapping and distinct functions of d-cysteine and d-serine to the Discussion (lines 418-436).

      (10) In the abstract, in lines 101 and 102, the authors mention "how d-serine contributes to cellular metabolism beyond neurotransmission remains largely unknown". In 2023, a paper in Stem Cell Reports by Roychaudhuri et al (PMID:37352848) showed that d and l-serine availability impacts lipid metabolism in the subventricular zone in mice, affecting proliferative properties of stem-cell derived neurons using a comprehensive lipidomics approach. There is no mention of this work even in the discussion section, as it bears directly on l and d-serine availability in neurons, which the authors are investigating. In the discussion section in lines 410-411, the authors mention the role of d-serine in neurogenesis, but surprisingly don't refer to the above reference. The role of d-serine in neurogenesis has been demonstrated in the Sultan et al (lines 855-857) and Roychaudhuri et al references.

      We thank the reviewer for highlighting these relevant studies. We have revised the statement in line 99 to avoid overgeneralization and to reflect that the metabolic roles of d-serine are incompletely understood rather than largely unknown. In addition, we have incorporated discussion of previous works, including the work by Roychaudhuri et al. (2023), into the Discussion section (lines 418-436) of our revised manuscript.

      (11) Both D-serine and the structurally similar stereoisomer D-cysteine (sulfur versus oxygen atom) have a bearing on 1C metabolism and the folate cycle. With reference to the folate cycle, Roychaudhuri et al in 2024 (PMID:39368613) have shown in rescue experiments in mice that supplementing a higher methionine diet provides folate cycle precursors to rescue the high insulin phenotype in SR-deficient mice. Since 1C metabolism is being discussed in this manuscript, the authors seem to overlook prior work in the field and not include it in their discussion, even when it is the same enzyme (SR) that synthesizes both serine and cysteine. Since the field of D-amino acid research is in its infancy, the authors must make it a point to include prior work related to D-serine at least, and not claim that it is not known. The known D-stereoisomers are not many, hence any progress in the area must include at least a discussion of the other structurally related stereoisomers.

      We are grateful to the reviewer for drawing our attention to this relevant study.

      We have now incorporated the findings from Roychaudhuri et al. 2024 into the Discussion. In that study, Srr-/- mice exhibited reduced levels of DNMT1 and DNMT3A, resulting in reduced DNA methylation activity. Notably, supplementation with a methyl-donor diet (containing choline, betaine, and methionine) restored the aberrant insulin phenotype in Srr-/- mice. These findings are relevant to our study, as DNA methylation depends on S-adenosyl-methionine (SAM), which is generated through one-carbon metabolism.

      We have expanded the Discussion to include the relationship between d-serine and d-cysteine, both of which are synthesized by Srr in the revised manuscript (lines 418-436).

      (12) Racemases (serine and aspartate) in general are promiscuous enzymes and known to synthesize other stereoisomers in addition to D-serine, D-cysteine, and D-aspartate. A few controls, like D-aspartate, D-cysteine, or even D-alanine must be included in their study to demonstrate the specific actions of D-serine, especially in the N2a cell treatment experiments. Cysteine and Serine are almost identical in structure (sulfur versus oxygen atom), and both are synthesized by serine racemase (published). Cysteine has also been very recently shown to inhibit tumor growth and neural progenitor cell proliferation. (PMIDs: 40797101 and 34556581). How the authors' work relates to the existing findings must be discussed, and this would put things in perspective for the reader.

      We thank the reviewer for this comment regarding the need to demonstrate the specificity of d-serine. As shown in Fig. S7A, d-serine, but not other d-amino acids commonly detected in mammals (Gonda et al., 2023), including d-aspartate, d-alanine, and d-proline, induced the cleavage of caspase-3 under the same experimental conditions, supporting the specific effect of d-serine in our system. We agree that d-cysteine shares structural similarity with d-serine and is also synthesized by Srr, suggesting potential functional overlap with d-serine. To place our findings in this context, we have added a paragraph in the Discussion (lines 418-436) describing the similarities of d-serine and d-cysteine. While both molecules may exert anti-proliferative effects, their underlying mechanisms appear to differ. Notably, supplementation with SAM, methionine, or glutathione did not rescue d-serine-induced growth inhibition in our system (Fig. S5J), suggesting that its effects are not primarily mediated through methylation or sulfur metabolic pathways. Instead, d-serine suppresses cellular proliferation by limiting mitochondrial l-serine availability and one-carbon metabolism. These observations highlight mechanistic divergence between d-serine and d-cysteine.

      Reviewer #2 (Public review):

      Summary:

      This study by Suzuki et al. reports an interesting stereo-selective role of D-serine in regulating one-carbon metabolism during neurodevelopment to adapt the functional transition, probably through the competition with mitochondrial transport of L-serine. The authors provide a multi-layered set of evidence, including metabolomics, enzyme assays, mitochondrial transport competition, and functional assays in immature/neural progenitor cells, to build up a conceptual integration of D-serine as both a neurotransmitter and a metabolic regulator in the central neural system, which raises a broad potential interest to the neuroscience and metabolism communities.

      Strengths:

      This work provides a conceptual advance that D-serine not only serves as a traditional neurotransmitter in the central neural system but also critically contributes to metabolic regulation of neural cells. The authors performed solid metabolomic assays to validate the suppressive effect of D-serine on the one-carbon metabolic pathway, providing some evidence that D-serine competitively inhibits mitochondrial serine transport, but not directly impairs SHMT2 enzymatic activity. All these data indicate a critical role of D-serine synthesis during neural maturation and suggest a potential translational strategy for targeting serine metabolism in neural tumors.

      Weaknesses:

      (1) The detailed mechanism by which D-serine competes with L-serine for its mitochondrial transport is not investigated. For example, although the authors made some discussion, they did not provide direct genetic or biochemical evidence linking these effects to the specific transporters, such as SFXN1.

      We thank the reviewer for this important comment regarding the mitochondrial l-serine transport mechanism. To address this point, we performed additional experiments using N2a cells in which Sfxn1 was knocked down by siRNA. Under semi-permeabilized cell conditions, we newly examined the effect of d-serine on mitochondrial L-serine transport.

      Interestingly, even under conditions where Sfxn1 expression was markedly suppressed, d-serine still inhibited mitochondrial d-serine transport. Given that SFXN1 is known to function redundantly with its paralogs (SFXN2–SFXN5) in mitochondrial serine transport (Kory et al., 2018), these findings suggest that d-serine may interfere with l-serine transport not only through SFXN1 but potentially through multiple members of the SFXN transporter family. These new data have been added as Fig. S3, and the corresponding results and discussion have been incorporated into the revised manuscript (lines 158–165 and 379-385).

      (2) Unlike tumor cells, where SHMT2 usually plays a predominant role in catalyzing serine/THF-derived one-carbon metabolism, normal cells may employ both SHMT1 and SHMT2 to do the work. Even under certain conditions that SHMT2-mediated one-carbon metabolism is suppressed, the activity of SHMT1 could be elevated for compensation. Thus, it is important to investigate whether D-serine affects SHMT1 activity or changes the balance between SHMT1- and SHMT2-mediated one-carbon metabolism. To this aim, the authors are strongly encouraged to perform a metabolic flux assay (MFA) by using 13C-labeled L-serine in the model cells in the presence and absence of D-serine.

      We thank the reviewer for this thoughtful comment regarding the potential contribution of cytosolic SHMT1 to one-carbon metabolism. As the reviewer notes, while mitochondrial SHMT2 is generally considered the predominant enzyme supporting one-carbon metabolism in proliferating or tumor cells, SHMT1 in the cytosol may function in a complementary manner in normal cells. In our experiments using primary cortical neurons, we indeed observed that the sensitivity to d-serine differed between immature and mature neuronal states (Fig. 4E). This observation suggests that the relative contribution of SHMT1- and SHMT2-mediated one-carbon metabolism may vary depending on the differentiation status of the cells.

      However, the primary focus of the present study was to investigate the mechanism underlying the anti-proliferative effect of d-serine in proliferative or undifferentiated neural cells. Our data demonstrate that d-serine inhibits one-carbon metabolism primarily by limiting mitochondrial l-serine availability through inhibition of mitochondrial l-serine transport. Therefore, a detailed analysis of the compensatory balance between SHMT1 and SHMT2 after disruption of mitochondrial serine transport falls beyond the central scope of the present study. Importantly, previous work by Miyamoto et al. (FEBS Journal, 2024) demonstrated that both SHMT1 and SHMT2 exhibit strong stereoselectivity for l-serine in their hydroxymethyl-transferase activity and do not utilize d-serine as a substrate. These findings make a direct effect of d-serine on hydroxymethyl-transferase activity of SHMT1 unlikely.

      Nevertheless, we agree that differences in the anti-proliferative effects of d-serine across cell types or differentiation states could reflect variations in cellular dependence on one-carbon metabolism or potential compensation by SHMT1. To address this point, we have expanded the Discussion section to clarify the possible contribution of SHMT1 and the limitations of the present study (lines 367–376). While isotope tracing analysis would be valuable for further dissecting compartmentalized one-carbon fluxes, such analyses would primarily address the relative contributions of SHMT1 and SHMT2 rather than the mitochondrial serine transport step that constitutes the central mechanism identified in this study.

      (3) A defect in serine-derived one-carbon metabolism may cause multiple cellular stress responses. It is valuable to detect whether cellular NADPH/NADH, GSH, or ROS is altered before and after D-serine treatment.

      We appreciate this insightful comment regarding potential cellular stress responses associated with impaired one-carbon metabolism. Consistent with the reviewer’s suggestion, we examined markers related to redox status. Transcriptomic analysis revealed compensatory changes in genes involved in mitochondrial metabolic function, including components of the NADH dehydrogenase complex (Fig. 2IJ and Fig. S4A), suggesting metabolic adaptation to d-serine treatment. In response to the reviewer’s comment, we measured GSH levels in NPCs and found that d-serine treatment led to a reduction of GSH (new Fig. S7F), indicating altered redox balance. However, supplementation with exogenous GSH did not rescue d-serine-induced cell death (Fig. S7E). These results suggest that while d-serine induces changes in cellular redox status, including GSH depletion, redox imbalance alone is unlikely to be the primary driver of cell death in this context.

      (4) The physiological relevance between D-serine and neural cell maturation/death should be further tested and discussed, since the dosage of D-serine used in the in vitro assay is much higher than that in physiological conditions.

      We thank the reviewer for this comment regarding the physiological relevance of the d-serine concentrations used in our study. We agree that the concentrations of d-serine required to compete with l-serine for mitochondrial transport are higher than those typically observed under physiological conditions, which we acknowledged in the Discussion of our manuscript (lines 386-388). Importantly, however, in vivo, d-serine levels are tightly regulated in a spatiotemporal manner during brain development, with low levels during embryonic and early postnatal stages and increased levels upon neural maturation. This temporal regulation coincides with the transition from proliferative neural progenitor states to differentiated neurons, thereby limiting the potential for d-serine to interfere with one-carbon metabolism during periods of active cell proliferation. Thus, while the concentrations used in vitro may exceed physiological levels, they allow us to uncover a latent metabolic effect of d-serine that may become relevant under specific cellular or developmental contexts. To clarify these points, we have revised the Discussion (lines 388-395) in the revised manuscript to more explicitly address the relationship between d-serine dosage and physiological relevance.

      Reviewer #3 (Public review):

      Summary:

      This manuscript presents a comprehensive and well-executed investigation into the metabolic role of D-serine in the central nervous system. The authors provide solid evidence that D-serine competitively inhibits mitochondrial L-serine transport, thereby impairing one-carbon metabolism. This stereoselective mechanism reduces glycine and formate production, suppresses cellular proliferation, and induces apoptosis in immature neural cells and glioblastoma stem cells. Developmental analyses further reveal a physiological enantiomeric shift in serine metabolism during neurogenesis, aligning with the transition from proliferation to maturation. Overall, the study bridges developmental neurobiology, cancer metabolism, and amino acid transport, uncovering a previously unrecognized metabolic function of D-serine beyond its role in neurotransmission.

      Strengths:

      (1) The discovery that D-serine inhibits one-carbon metabolism by competing for mitochondrial L-serine transport-rather than through enzymatic inhibition or receptor-mediated signaling-represents a significant and previously underappreciated mechanism. This finding has broad implications for understanding metabolic regulation during neurodevelopment and offers potential relevance for targeting metabolic vulnerabilities in cancer.

      (2) The authors integrate metabolomics, mitochondrial transport assays, molecular dynamics simulations, genetic and pharmacologic perturbations, transcriptomics, and both in vitro and ex vivo models. The breadth of experimental approaches, combined with the coherence of the findings across systems, provides strong support for the central conclusions and enhances the overall impact of the study.

      (3) The temporal shift in D-/L-serine levels during neurodevelopment is elegantly linked to the transition from proliferative to mature neuronal states. The selective vulnerability of neural progenitors and tumor cells-contrasted with the resistance of mature neurons-highlights a biologically meaningful and potentially targetable metabolic distinction.

      Weaknesses:

      (1) While the authors attribute D-serine's metabolic effects to competition with mitochondrial L-serine transport, the specific identity of the transporter(s) mediating this process remains undefined. This represents a meaningful mechanistic gap, as the central conclusion depends on D-serine limiting mitochondrial L-serine availability to inhibit one-carbon metabolism.

      We thank the reviewer for this insightful comment regarding the identity of the mitochondrial l-serine transporter. As this concern overlaps with the point raised by Reviewer 2 (Weakness 1), we refer the reviewer to our response there for a detailed description of the additional experiments performed. Briefly, we conducted new experiments using siRNA-mediated knockdown of Sfxn1 in N2a cells and examined mitochondrial l-serine transport under semi-permeabilized conditions. Notably, even with marked suppression of Sfxn1 expression, d-serine continued to inhibit mitochondrial l-serine transport. Given that SFXN family members (SFXN1–SFXN5) are reported to function redundantly in mitochondrial serine transport (Kory et al., 2018), these findings suggest that d-serine may interfere with l-serine transport not only via SFXN1 but potentially across multiple SFXN paralogs. These results have been incorporated into the revised manuscript and are presented in Fig. S3, with the corresponding discussion added to the Results and Discussion sections (lines 158–164 and 379-385).

      (2) The effective concentrations of D-serine used in vitro (IC<sub>50</sub> ≈ 1-2 mM) exceed typical brain levels (~0.3 mM). While the authors acknowledge this, a more focused discussion on whether higher local D-serine concentrations could arise in specific microenvironments - such as synaptic compartments, tumor niches, or pathological states-would help contextualize the in vitro findings and strengthen their physiological relevance. For example, disruptions in D-serine clearance or altered expression of serine racemase and transporters in disease contexts could lead to localized accumulation. Moreover, differences between extracellular and intracellular D-serine pools - and the mechanisms governing their regulation - may further influence its metabolic impact in vivo.

      We appreciate this insightful comment regarding the physiological relevance of the d-serine concentrations used in vitro. We agree that the effective concentrations observed in our assays exceed typical bulk brain levels. To address this point, we have expanded the Discussion (lines 386-399) to consider conditions under which locally elevated d-serine concentrations may arise in vivo. These additions provide a more nuanced interpretation of the relationship between the concentrations used in vitro and the potential physiological contexts in which d-serine may exert metabolic effects.

      (3) While the manuscript focuses on neural stem/progenitor cells and neural tumors, it remains unclear whether the anti-proliferative effects of D-serine are specific to neural lineages or extend to other highly proliferative non-neural cell types. A brief discussion addressing this point would help clarify the scope of D-serine's metabolic impact and whether its mechanism of action reflects a unique vulnerability in neural cells or a more general feature of proliferative metabolism. This distinction is particularly relevant for assessing the broader therapeutic potential of targeting mitochondrial L-serine transport.

      We thank the reviewer for this comment regarding the potential generality of the anti-proliferative effects of d-serine. We agree that it is important to clarify whether the observed effects are specific to neural lineages or reflect a broader vulnerability of proliferative cells. In our study, we focused on neural progenitor cells and neural tumour models. However, the mechanism identified here, namely limitation of mitochondrial L-serine availability, targets a fundamental metabolic pathway that supports cell proliferation. Therefore, it is possible that similar effects may extend to other highly proliferative cell types beyond the neural lineage. To address this point, we have expanded the Discussion (lines 408-410) to clarify that the observed effects of d-serine may reflect a general metabolic vulnerability associated with proliferative states, while also noting that the degree of sensitivity is likely to depend on cell-type-specific reliance on mitochondrial one-carbon metabolism.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor issues:

      The authors mention terms like neural tissue (line 109) and neural tumor cells (line 182). Neural tissue can mean anything under the sun. They need to mention the specific tissue being studied or investigated.

      We have changed the terms.

      Reviewer #2 (Recommendations for the authors):

      (1) The figure items were not well organized, and the legends were difficult to read since they apparently lacked key information. For example, in Figure 2D, did the author intend to show SHMT2 activity? It is not clear how this assay was performed (in vitro or in vivo experiment?). Additionally, more detailed information should be provided in the Methods section.

      We have improved the organization of figures and revised figure legends for better readability.

      (2) The writing of the manuscript should be significantly improved, using a professional editing service, in order to increase the readability.

      We appreciate the reviewer’s comment. The manuscript was professionally edited prior to submission. Nevertheless, we have made targeted revisions to the Introduction and Discussion sections to improve overall readability. We hope that these revisions have improved the clarity of the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) The core mechanism centers on competition for mitochondrial L-serine transport, yet the identity of the transporter(s) involved remains speculative. While Kory et al. (2018) identified SFXN1 as a mitochondrial L-serine transporter, this connection is not directly addressed in the current study. It would strengthen the manuscript to clarify whether SFXN1 or related isoforms are expressed in the neural cell models used and whether their expression patterns correspond with the observed D-serine sensitivity. Even if functional validation is beyond the current scope, a more detailed discussion of potential transporter candidates and the limitations of existing data would provide important mechanistic context and help frame future directions.

      We thank the reviewer for this insightful comment regarding the mitochondrial l-serine transporter and the potential involvement of SFXN family members. As this concern overlaps with the point raised by Reviewer 2 (Weakness 1), we refer the reviewer to our response there for a detailed description of the additional experiments performed.

      Briefly, we performed siRNA-mediated knockdown of Sfxn1 in N2a cells and examined mitochondrial l-serine transport under semi-permeabilized conditions. Even under conditions of marked suppression of Sfxn1 expression, d-serine continued to inhibit mitochondrial l-serine transport. Given that members of the SFXN family have been reported to function redundantly in mitochondrial serine transport (e.g., Kory et al., 2018), these findings suggest that d-serine may affect l-serine transport not only through SFXN1 but potentially across multiple SFXN paralogs.

      These new results have been incorporated into the revised manuscript (Fig. S3), and the relevant discussion has been expanded to clarify the potential roles of SFXN family transporters and the current limitations in defining the exact transporter responsible (lines 158–165 and 379-385).

      (2) The authors use ex vivo brain slice cultures with tumor xenografts to demonstrate tissue-level relevance, which is a valuable strength of the study. However, additional context would enhance its translational significance. It would be helpful to discuss whether in vivo D-serine administration (e.g., ICV or systemic) is feasible and safe, especially given the high concentrations required in vitro. Briefly addressing whether genetic models, such as Srr knockout mice, support a role for D-serine in tumor progression or neurodevelopment would also strengthen the interpretation.

      We thank the reviewer for this important comment regarding the translational relevance of d-serine administration in vivo. To address this point, we performed additional exploratory experiments using a subcutaneous tumour xenograft model in nude mice. Because the inhibitory effect of d-serine on one-carbon metabolism becomes evident under l-serine–limited conditions, tumor-bearing mice were fed an l-serine/glycine–deficient diet and administered d-serine in drinking water.

      However, when d-serine was provided at high concentrations (≥ 500 mM) in drinking water, the mice exhibited marked behavioral abnormalities, including increased aggression and other abnormal behaviors. Due to these adverse effects, the experiment was ethically terminated. While our study focuses on the inhibitory effect of d-serine on one-carbon metabolism, d-serine is also a physiological co-agonist of the NMDA receptors, and therefore high systemic concentrations may influence neuronal excitability in vivo. These observations suggest that the concentrations required to achieve antitumor effects may be associated with significant neurological side effects.

      Based on these findings, we consider that direct administration of d-serine itself may have limited therapeutic applicability as an antitumor reagent. Instead, future development of derivatives or strategies that retain the metabolic inhibitory effect while minimizing NMDA receptor–mediated effects may be required.

      Regarding serine racemase (Srr) knockout models, xenograft tumour experiments would require an immunodeficient background, which makes the generation and use of such compound models technically challenging. Therefore, we did not pursue this approach in the present study. We have incorporated these considerations into the revised Discussion (lines 413-417) to clarify the translational implications and current limitations of our findings.

      Bella DJD, Habibi E, Stickels RR, Scalia G, Brown J, Yadollahpour P, Yang SM, Abbate C, Biancalani T, Macosko EZ, Chen F, Regev A, Arlotta P. 2021. Molecular logic of cellular diversification in the mouse cerebral cortex. Nature 595:554–559. DOI: https://doi.org/10.1038/s41586-021-03670-5, PMID: 34163074

      Bernal A, Arranz L. 2018. Nestin-expressing progenitor cells: function, identity and therapeutic implications. Cellular and Molecular Life Sciences 75:2177–2195. DOI: https://doi.org/10.1007/s00018-018-2794-z, PMID: 29541793

      Bott CJ, Johnson CG, Yap CC, Dwyer ND, Litwa KA, Winckler B. 2019. Nestin in immature embryonic neurons affects axon growth cone morphology and Semaphorin3a sensitivity. Molecular Biology of the Cell 30:1214–1229. DOI: https://doi.org/10.1091/mbc.e18-06-0361, PMID: 30840538

      Graham V, Khudyakov J, Ellis P, Pevny L. 2003. SOX2 Functions to Maintain Neural Progenitor Identity. Neuron 39:749–765. DOI: https://doi.org/10.1016/s0896-6273(03)00497-5, PMID: 12948443

      Lendahl U, Zimmerman LB, McKay RDG. 1990. CNS stem cells express a new class of intermediate filament protein. Cell 60:585–595. DOI: https://doi.org/10.1016/0092-8674(90)90662-x, PMID: 1689217

      Thakurela S, Tiwari N, Schick S, Garding A, Ivanek R, Berninger B, Tiwari VK. 2016. Mapping gene regulatory circuitry of Pax6 during neurogenesis. Cell Discovery 2:15045. DOI: https://doi.org/10.1038/celldisc.2015.45, PMID: 27462442

    1. eLife Assessment

      This useful study investigates how plasticity and homeostatic adaptation can lead to the emergence of synchronization patterns associated with different sleep phases. The ideas are novel and interesting, but the evidence at present remains incomplete. Further work is needed to determine whether the reported states persist for larger system sizes, longer integration times, and how robust the results are to different initial conditions.

    2. Reviewer #1 (Public review):

      Summary:

      The authors aim to understand how changes in the balance between excitatory and inhibitory interactions influence the stability and reorganization of network connections. To address this question, they extend a coupled-phase-oscillator model by adding plasticity rules. The central finding is that stronger inhibitory interactions lead to relatively stable and desynchronized network dynamics, whereas weaker inhibitory interactions produce a bistable regime in which intermediate-strength connections fluctuate while stronger connections are preserved.

      Strengths:

      This study offers a simple theoretical framework for linking network state, coupling stability, and reorganization. The model produces clear qualitative results, showing that different dynamical regimes are associated with different balances of excitatory and inhibitory interactions. This could be useful as a conceptual starting point for considering how network states may regulate the stability and flexibility of connections. The manuscript also explores several model parameters.

      Weaknesses:

      The evidence is incomplete in supporting the biological interpretations. The model is a highly simplified coupled-phase-oscillator system and does not directly represent spiking activity, membrane potentials, synaptic currents, conduction delays, cellular excitability, or detailed biological plasticity mechanisms. Although the authors clarify that the model units are not actual neurons or synapses, the discussion often interprets the results in terms of neuronal inhibition, synaptic stability, sleep-related reorganization, and preservation of strong biological connections. This creates a gap between the abstract model and the biological conclusions. In particular, the manuscript does not sufficiently discuss what biological oscillatory activity the modeled phases are intended to represent, such as population-level activity reflected in electroencephalography or local field potentials. In several places, the manuscript appears to assume that neurons can generally be treated as oscillators, but this is not always a valid assumption. The authors should more clearly distinguish between rhythmic or phase-like activity at the population level and the dynamics of individual neurons, and should frame the model more cautiously as a phenomenological description of collective synchronization rather than a mechanistic model of spiking neuronal circuits.

      There are also important methodological limitations. Although the manuscript presents the model equations, parameter values, time step, simulation duration, and coupling update rules, several other essential details are not clearly specified, including the number of simulation runs, the procedure for setting initial conditions, and the numerical method used to solve the ordinary differential equations. Critically, technical details such as the integration scheme, solver settings, initialization procedure, and random seed handling are essential for reproducibility. Because the main findings depend on the interaction between phase dynamics and adaptive coupling, even small implementation differences could affect the reported dynamical regimes and coupling fluctuations.

      A further concern is the presentation of the mathematical formulation. Several equations appear to contain notation errors and inconsistencies, making it difficult to follow the exact model definition. The authors should carefully revise the mathematical notation throughout the manuscript to ensure that the model can be understood and reproduced unambiguously.

      Overall, the study provides a useful but limited theoretical account of how network dynamics may regulate coupling stability and reorganization. The results support the internal behavior of the proposed model, but the broader biological claims are not yet fully convincing. The likely impact of the work is therefore mainly conceptual: it may stimulate further modeling studies, but additional methodological detail, stronger justification of the modeling assumptions, and comparison with more biologically grounded models would be needed before the conclusions can be applied confidently to neuronal circuit dynamics or sleep-related synaptic reorganization.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the impact of plasticity mechanisms in an excitation-inhibition (EI) network model on the emergence of synchronization patterns that the authors associate with different sleep phases. The model consists of an EI Kuramoto network in which recurrent excitatory couplings evolve according to Hebbian and homeostatic adaptation rules. Through numerical simulations, the authors analyze how these plasticity mechanisms modify both the collective dynamics and the structure of the coupling matrix.

      Strengths:

      The topic addressed in the manuscript is timely and potentially relevant, as understanding the interplay between synaptic adaptation and collective neural dynamics remains an important challenge in theoretical neuroscience.

      Weaknesses:

      In its current form, the work suffers from substantial conceptual, methodological, and technical limitations that significantly weaken the conclusions.

      From a biological perspective, the model is highly abstract and qualitative. The connection between the model variables and the physiological processes that the authors aim to describe remains unclear. Consequently, the manuscript does not provide sufficient evidence to support biologically meaningful conclusions regarding sleep dynamics. In my opinion, the work is more naturally positioned within the framework of theoretical or computational dynamical systems than within the scope of a biology-oriented journal.

      From a mathematical and dynamical-systems perspective, the analysis is incomplete, and several important technical aspects are either missing or inadequately addressed. In particular, the characterization of the dynamical regimes is often imprecise, the numerical evidence is not sufficiently robust, and little effort is made to interpret the results within the broader context of synchronization theory, adaptive networks, or collective dynamics.

      More specifically:

      (1) The biological interpretation of the model variables is ambiguous throughout the manuscript. At several points, the authors suggest that individual oscillators should not be interpreted as neurons but rather as abstract biological units (lines 78-82, 86-89, 418-420). However, other parts of the manuscript refer to the coupling matrix entries, particularly $J_{ee}$, as synaptic weights (e.g., line 118). These two interpretations are not obviously compatible. If the oscillators represent coarse-grained or abstract units, the biological meaning of the adaptive couplings should be carefully justified. More generally, the manuscript lacks a clear discussion of what aspects of neural circuits are captured by the model and which aspects are intentionally neglected.

      (2) More fundamentally, the manuscript inherits the well-known limitations associated with interpreting Kuramoto oscillators as neural elements. Kuramoto phase oscillators provide a minimal description of synchronization phenomena, but they do not explicitly represent membrane dynamics, firing rates, spiking activity, synaptic currents, or realistic neuronal timescales.

      Under certain assumptions, Kuramoto-like models can be rigorously derived from more detailed neuronal models through phase-reduction techniques (see, for instance, Chapter 10 of Izhikevich's \textit{Dynamical Systems in Neuroscience}). However, the authors do not employ such a reduction procedure, nor do they establish a formal connection between their model variables and the dynamics of neuronal populations. As a consequence, the biological interpretation of the model remains unclear.

      The authors should therefore explicitly discuss these limitations and carefully justify why the synchronization patterns observed in such a highly reduced model can be related to neural sleep states. At present, the biological interpretation appears considerably stronger than what the model itself can support, for instance, the claims in lines 321-322 or 330-337. In particular, it remains unclear whether the reported dynamical regimes should be interpreted as genuine mechanisms underlying sleep rhythms or merely as generic synchronization phenomena arising in adaptive oscillator networks.

      (3) The numerical methodology raises serious concerns regarding the robustness of the reported results. According to the Methods section, simulations are performed using only $N=100$ oscillators and integration times of approximately 50 time units. Such choices may be sufficient for illustrative purposes but are generally inadequate for drawing conclusions about asymptotic collective behavior in adaptive dynamical systems. Finite-size fluctuations can strongly affect synchronization measures, and adaptive networks are well known to exhibit extremely long transients, metastability, and slow convergence processes. No systematic finite-size analysis is provided, nor is there any demonstration that the reported states persist for larger system sizes or longer integration times.

      (4) The use of the term "bistable regime" to describe the dynamics shown in Figure 1C is incorrect. A bistable regime refers to a parameter region in which multiple attractors coexist and the asymptotic state depends on the initial condition. The figure instead presents a single trajectory displaying oscillatory dynamics. No evidence is provided for the coexistence of attractors, nor are multiple initial conditions explored. Furthermore, the displayed time series are too short to determine whether the observed dynamics correspond to a stable limit cycle, quasiperiodic motion, intermittent behavior, or a long transient approaching another attractor. The authors should perform a proper dynamical characterization of this regime. Similar collective oscillatory states have been extensively studied in synchronization and adaptive-network models and should be discussed in relation to the existing literature.

      In particular, several time series shown in the manuscript (e.g., Figures 1C and 2F) exhibit trends that suggest the possibility of unresolved transient dynamics. The authors should demonstrate convergence of the reported regimes by substantially extending simulation times and by performing finite-size analyses. Simulations with at least one order of magnitude more units ($N\gtrsim 1000$) and integration times sufficient to establish asymptotic behavior would be expected in a study whose main claims rely on collective dynamical phenomena.

      (5) The manuscript lacks several standard tools routinely employed in the analysis of nonlinear dynamical systems. The conclusions are largely based on visual inspection of time series and order parameters. However, no bifurcation analysis, stability analysis, phase-space characterization, attractor reconstruction, or systematic exploration of parameter dependence is provided. As a consequence, many of the identified "phases" or "regimes" remain only qualitatively described. A more rigorous dynamical-systems treatment would substantially strengthen the work and would help distinguish genuine asymptotic states from finite-size or transient phenomena.

      (6) A substantial fraction of the results appears to extend the authors' previous work by incorporating plastic adaptation mechanisms. While incremental advances are acceptable, the manuscript would benefit from a broader theoretical context. The discussion is heavily centered on previous studies by the same authors, whereas there exists an extensive literature on synchronization, adaptive networks, neural mass models, balanced EI systems, and sleep-related oscillations that is largely absent from the discussion. The novelty and significance of the present contribution would be easier to assess if the results were more carefully compared with alternative theoretical approaches.

    1. eLife Assessment

      This important study introduces a non-perturbative pulse-labeling strategy for yeast nuclear pore complexes (NPCs), employing a nanobody-based approach in order to selectively capture Nup84-containing complexes for imaging and biochemical analysis. The data convincingly demonstrate that a short induction period (20 minutes to 1 hour) yields a strong and sustained signal, enabling affinity purification that faithfully recapitulates the endogenous Nup84 interactome. This tool offers a powerful framework for investigating NPC dynamics and associated interactomes through both imaging and biochemical assays.

    2. Reviewer #1 (Public review):

      Summary:

      The authors present a nanobody-based pulse-labeling system to track yeast NPCs. Transient expression of a nanobody targeting Nup84 (fused to NeonGreen or an affinity tag) permits selective visualization and biochemical capture of NPCs. Short induction effectively labels NPCs, and the resulting purifications match those from conventional Nup84 tagging. Crucially, when induction is repressed, dilution of the labeled pool through successive cell cycles allows the visualization of "old" NPCs (and potentially individual NPCs) providing a powerful view of NPC lifespan and turnover without permanently modifying a core scaffold protein.

      Strengths:

      (1) A brief expression pulse labels NPCs, and subsequent repression allows dilution-based tracking of older (and possibly single) NPCs over multiple cell cycles.

      (2) The affinity-purified complexes closely match known Nup84-associated proteins, indicating specificity and supporting utility for proteomics.

      Weakness:

      Reliance on GAL induction introduces metabolic shifts (raffinose → galactose → glucose) that could subtly alter cell physiology or the kinetics of NPC assembly. As acknowledged by the authors, alternative induction systems (e.g., β-estradiol-responsive GAL4-ER-VP16) could be implemented as a way to avoid carbon-source changes.

      Comments on revised version.

      The authors have thoughtfully addressed all of my concerns. In particular, they have updated the proteomic analysis in Figure 1I, showing that they recover most NPC components (including basket Nups), including non-NPC proteins as controls, and providing all data as a supplementary table. These changes strengthen the authors conclusion and improve transparency. I have no further recommendations and congratulate the authors for their exciting work.

    3. Reviewer #2 (Public review):

      Summary:

      This preprint describes a practical and useful approach for labeling and tracking NPCs in situ, using a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity. Useful applications including timelapse imaging, affinity purification, and proximity labeling are envisioned.

      Strengths:

      Clever use of a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity.

    4. Reviewer #3 (Public review):

      Summary:

      Submitted to the Tools and Resources series, this study reports on the use of a single-domain antibody targeting the nucleoporin Nup84 to probe and track NPCs in budding yeast. The authors demonstrate their ability to rapidly label or pull down NPCs by inducing the expression of a tagged version of the nanobody (Fig. 1).

      Strengths:

      This tool's main strength is its versatility as an inexpensive, easy-to-set-up alternative to metabolic labelling or optical switching. This same rationale could, in principle, be applied to the study of other multiprotein complexes using similar strategies, provided that single-chain antibodies are available.

      Weaknesses:

      This approach has no inherent weaknesses, but it would be useful to verify in the future that this pulse labelling strategy can also be used to detect assembly intermediates, structural variants, or damaged NPCs, e.g. NPC clusters formed in some nucleoporin mutants.

      Overall, the data clearly shows that Nup84 nanobodies are a valuable tool for imaging NPC dynamics and investigating their interactomes through affinity purification.

      Comments on revised version.

      None at this stage.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors present a nanobody-based pulse-labeling system to track yeast NPCs. Transient expression of a nanobody targeting Nup84 (fused to NeonGreen or an affinity tag) permits selective visualization and biochemical capture of NPCs. Short induction effectively labels NPCs, and the resulting purifications match those from conventional Nup84 tagging. Crucially, when induction is repressed, dilution of the labeled pool through successive cell cycles allows the visualization of "old" NPCs (and potentially individual NPCs), providing a powerful view of NPC lifespan and turnover without permanently modifying a core scaffold protein.

      Strengths:

      (1) A brief expression pulse labels NPCs, and subsequent repression allows dilution-based tracking of older (and possibly single) NPCs over multiple cell cycles.

      (2) The affinity-purified complexes closely match known Nup84-associated proteins, indicating specificity and supporting utility for proteomics.

      We thank the reviewer for this evaluation

      Weaknesses:

      (1) Reliance on GAL induction introduces metabolic shifts (raffinose -> galactose -> glucose) that could subtly alter cell physiology or the kinetics of NPC assembly. Alternative induction systems (e.g., β-estradiol-responsive GAL4-ER-VP16) could be discussed as a way to avoid carbon-source changes.

      Indeed, this could be an improvement, and we mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3.

      (2) While proteomics is solid, a comprehensive supplementary table listing all identified proteins (with enrichment and statistics) would enhance transparency.

      Indeed, we now provide source data showing LFQ intensities, fold-enrichment and statistics for all detected proteins.

      (3) Importantly, the authors note that the method is particularly useful "in conditions where direct tagging of Nup84 interferes with its function, while sub-stoichiometric nanobody binding does not." After this sentence, it would be valuable to add concrete examples, such as experiments examining NPC integrity in aging or stress conditions where epitope tags can exacerbate phenotypes. These examples will help readers identify situations in which this approach offers clear advantages.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants, and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks, and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #2 (Public review):

      Summary:

      This preprint describes a practical and useful approach for labeling and tracking NPCs in situ. While useful applications including timelapse imaging, affinity purification, or proximity labeling are envisioned, addressing some outstanding technical questions would give a clearer picture of the sensitivity and temporal resolution of this approach.

      Strengths:

      Clever use of a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity.

      We thank the reviewer for this evaluation

      Weaknesses:

      The decrease in nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over during this time. However, it is also possible that the nanobody: Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #3 (Public review):

      Summary:

      Submitted to the Tools and Resources series, this study reports on the use of a single-domain antibody targeting the nucleoporin Nup84 to probe and track NPCs in budding yeast. The authors demonstrate their ability to rapidly label or pull down NPCs by inducing the expression of a tagged version of the nanobody (Figure 1).

      Strengths:

      This tool's main strength is its versatility as an inexpensive, easy-to-set-up alternative to metabolic labelling or optical switching. This same rationale could, in principle, be applied to the study of other multiprotein complexes using similar strategies, provided that single-chain antibodies are available.

      We thank the reviewer for this evaluation

      Weaknesses:

      This approach has no inherent weaknesses, but it would be useful for the authors to verify that their pulse labelling strategy can also be used to detect assembly intermediates, structural variants, or damaged NPCs.

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure but consider such studies to be beyond the scope of this study.

      Overall, the data clearly show that Nup84 nanobodies are a valuable tool for imaging NPC dynamics and investigating their interactomes through affinity purification.

      We thank the reviewer again for the constructive feedback and thoughts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 1A, and although it is partially mentioned in the legend, it would be helpful to indicate precisely when cells are grown in raffinose, when galactose is added for induction, and when glucose is used to terminate expression.

      We included “galactose” and “glucose” to Panel A to indicate induction and termination of expression, respectively.

      (2) Related to the previous point, consider mentioning the GAL4-ER-VP16 (ADGEV) estradiol-inducible system as an optional strategy to avoid carbon shifts and potentially reduce cell-to-cell variability.

      We mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3

      (3) Add a brief sentence explaining that the ZZ tag is derived from Protein A and binds IgG Fc.

      This information is now added on p.2

      (4) The statement "all Nups significantly coenriched with VHH[Nup84]-ZZ..." is likely inaccurate, since not all Nups are labeled in panels F-H, and some basket components are missing in panel I (particularly basket components such as Nup60, Nup1). Consider revising to "most Nups significantly coenriched...". In panel I, please include a clearly non-enriched protein as a visual reference for the color scale.

      We are very grateful to the reviewer for pointing this out. We accidentally used a faulty filtering on the dataset to generate figure panel I, omitting several Nups that were reproducibly found in all replicas. All Nups, except for Gle1 and Pom33, were detected reproducibly.

      We have made the following adjustments to the figure panel and accompanying text:

      In Fig. 1I, we included the missing Nups and 5 proteins that co-purified with VHH[Nup84] but not specifically enriched, as the reviewer suggested. They cluster in a separate group and their abundance is not going up in time. We randomly selected these 5 proteins from the list of genes that were reproducibly found in all four timepoints.

      For clarity, we removed the NTRs

      We changed the text to “we found that all Nups, except Gle1 and Pom33, significantly coenriched with VHH[Nup84]-ZZ” on p.2.

      We updated the methods section, describing the clustering method and how we selected the 5 random proteins

      (5) Provide a supplementary spreadsheet with LFQ intensities, fold-enrichment, and statistics for all detected proteins. This will address questions about missing Nups and support transparency.

      This information is now added as Source data Figure 1.

      (6) Directly after the statement "Amongst others this is useful in conditions where direct tagging of Nup84 interferes with its function, while sub stoichiometric nanobody binding does not," it would be useful to include concrete instances, such as stress or aging conditions, where Nup84 tagging may sensitize NPC integrity.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      (7) In panels K and L, since individual points correspond to biological replicates, overlaying a box plot obscures much of the data. Consider overlaying the means per replicate instead of box plots: see the "SuperPlots" approach for a clear explanation of how to present this (PMID: 32346721).

      We thank the reviewer for the “SuperPlots” suggestion, and we agree that representing the data in this way improves the visualization of individual biological replicates. We have updated the summarizing overlay in figures in panel K and L to represent the means per replicate instead of boxplots.

      (8) I spotted a few typos ("Lasty" ? "Lastly"; "in maintained" vs. "is maintained").

      Thank you, these are corrected

      Overall, this is a neat, well-executed methodological advance with clear value to the NPC field and potentially other complex assemblies. I look forward to seeing a revised version.

      Thank you!

      Reviewer #2 (Recommendations for the authors):

      Based on the recent structural analyses and NPC modeling using this nanobody, how accessible is the Nup84 epitope expected to be within the fully assembled NPC? While the data shown indicate that nanobody labeling of NPCs is readily detectable, stating this clearly would help motivate the approach and interpret the resulting data.

      We now included such a statement in the introduction on p.1.

      The decrease of nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over due to cell division during this time window. However, it is also possible that nanobody:Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      Reviewer #3 (Recommendations for the authors):

      (1) As mentioned above, to assess the general relevance of this tool, it would be informative to verify whether the VHH[Nup84] nanobody can access and detect NPC species under conditions that challenge their structural organization or biogenesis, for example, in nucleoporin mutants or under stress. The authors could, for instance, analyze the localization of VHH[Nup84] in yeast strains harboring clustered NPCs (nup133Δ), or following stresses known to impact NPC organization (e.g., osmotic stress or energy depletion; PMID: 34762489).

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure and tried to include such data. Unfortunately, this was not successful, and further efforts are beyond the scope of his study. Following the reviewer’s suggestion, we expressed VHH[Nup84] in nup133∆N (nup133∆2-300) (Doye, Wepf, and Hurt 1994) following the experimental set-up in panel A and examined its localization. However, at t=2hrs hardly any nanobody signal was detectable in nup133∆N (see Author response image 1, upper panel A) and only after overnight expression nanobody-labelled NPC clusters are detectable (bottom panel A). Considering that expression levels of free mNG are also lower at t=2hrs in nup133∆N cells compared to WT cells (Author response image 1, panel B), it appears that protein expression under the Gal system is generally reduced in a nup133∆N background. These expression level differences between nup133∆N and WT preclude statements about the accessibility of the Nup84 epitope in nup133∆N. We note that nup133∆N cells do not have general mRNA export defects (Doye, Wepf, and Hurt 1994), so other inducible systems may be better suited for such analysis.

      Author response image 1.

      Expression level differences in WT and Nup133∆N cells. Left: localization of VHH[Nup84]-mNG in Nup133∆N cells at t=2hr following a 20-minute induction pulse and after overnight 0.5% galactose (ON) induction. Right: mNG levels in WT and Nup133∆N cells at t=2hr following a 20-minute induction pulse. Brightness/contrast settings are identical between the two panels. All panels are sum slices projections from 30 z-slices of 0.1µm. Scale bar = 5 µm.

      (2) Since outer rings are found on both sides of NPCs (i.e., the cytoplasmic and nuclear faces), could the authors indicate whether the VHH[Nup84] nanobody can enter the nucleus and probe the nuclear outer rings? Along these lines, it would be useful to provide a summary of the structural organization of NPCs in the introduction.

      Thank you, we have added a sentence on the localization of Nup84 in NPCs in the introduction. Based on what is known about influx (nuclear transport receptor-independent nuclear entry) of proteins with similar size and surface properties (Popken et al. 2015; Timney et al. 2016), the nanobody can rapidly enter the nucleus and hence bind Nup84 on both the nuclear and cytoplasmic side. We have no data to answer if binding might initially be biased towards cytosolic VHH[Nup84] binding the cytoplasmic outer rings.

      (3) The authors state that VHH[Nup84] and direct Nup84 detection are indistinguishable (p. 2). Could they provide images of the endogenously tagged Nup84-GFP strain for comparison?

      We have now included a pairwise comparison in a Figure 1 – supplement 1.

      Minor corrections:

      (1) There are a few typos that need correcting: 'Nup84Δ' (p. 1; should read 'nup84Δ') and 'promotor' (p. 2; should read 'promoter').

      Thank you, these are corrected

      (2) The reference 'Veldsink et al. 2025' (quoted in the PunctaFinder analysis description on page 8) does not appear in the References section.

      Thank you, these are corrected.

      We thank the reviewer again for the constructive feedback and thoughts.

      References

      Doye, V., R. Wepf, and E. C. Hurt. 1994. 'A novel nuclear pore protein Nup133p with distinct roles in poly(A)+ RNA transport and nuclear pore distribution', EMBO J, 13: 6062-75.

      Khmelinskii, Anton, Philipp J. Keller, Holger Lorenz, Elmar Schiebel, and Michael Knop. 2010. 'Segregation of yeast nuclear pores', Nature, 466: E1-E1.

      Popken, Petra, Ali Ghavami, Patrick R. Onck, Bert Poolman, and Liesbeth M. Veenhoff. 2015. 'Size-dependent leak of soluble and membrane proteins through the yeast nuclear pore complex', Molecular Biology of the Cell, 26: 1386-94.

      Timney, Benjamin L., Barak Raveh, Roxana Mironska, Jill M. Trivedi, Seung Joong Kim, Daniel Russel, Susan R. Wente, Andrej Sali, and Michael P. Rout. 2016. 'Simple rules for passive diffusion through the nuclear pore complex', Journal of Cell Biology, 215: 57-76.

      Zsok, J., F. Simon, G. Bayrak, L. Isaki, N. Kerff, Y. Kicheva, A. Wolstenholme, L. E. Weiss, and E. Dultz. 2024. 'Nuclear basket proteins regulate the distribution and mobility of nuclear pore complexes in budding yeast', Mol Biol Cell, 35: ar143.

    1. eLife Assessment

      This valuable study presents a comparative analysis of the transcriptomic features underlying C. elegans longevity, providing insights into how different changes in gene expression can promote longevity. The authors present solid evidence with analysis and selected functional validation showing that some long-lived animals share common changes while others appear to use opposing strategies. The datasets and analyses contained within and the user-friendly website developed will be of interest to researchers interested in complicated transcriptomic analyses and/or the biology of aging.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      This manuscript by Rudich ZD et al. systematically profiled the transcriptomic changes in nine long-lived C. elegans mutants and presented a careful and informative comparative analysis of these aging-related changes. In addition to these valuable datasets and bioinformatics analyses, the authors performed a large-scale RNAi screen to assess the role of the differentially expressed genes (DEGs) in these mutants and identify several potential targets to promote healthy aging. Moreover, the authors have provided a user-friendly website to examine genes of interest in those longevity mutants from their datasets.

      Strengths:

      Compared to previous transcriptomic analyses of these mutants in different reports, this study minimized the technical variations and benefitted from the advances in RNA-Seq technology and bioinformatics tools. Therefore, it should provide a more consistent and comprehensive view of the molecular mechanisms underlying the longevity of these mutants. The datasets in this manuscript are valuable to other researchers in the biology of aging.

      Weaknesses:

      Meanwhile, since these mutants have been extensively studied, the advance of this study in unknown ageing mechanisms remains limited.

      Comments on revised version.

      In the revised manuscript, the authors have addressed most of my concerns. In the text of this manuscript, the authors should still include more discussion on why osm-5 and daf-2 are categorized into two different groups.

    3. Reviewer #2 (Public review):

      Summary:

      In the manuscript titled "Multiple Molecular Pathways to Longevity: Opposing Gene Expression Programs Define Distinct Aging Strategies", the authors investigated diverse genetic pathways that contribute to lifespan extension in Caenorhabditis elegans and aimed to identify shared and distinct molecular mechanisms among various longevity mutants. Through comprehensive RNA sequencing of different longevity mutants representing seven distinct pathways, the authors showed that these mutants cluster into three primary groups based on their gene expression profiles. This transcriptomic analysis revealed that while some longevity genes are commonly regulated across multiple pathways, others exhibit opposing expression patterns, suggesting that distinct molecular strategies can lead to increased lifespan. Specifically, they identified a set of 196 genes that are consistently upregulated in most longevity mutants, many of which are involved in innate immunity and stress defense. By performing RNAi-based screening, the authors further validated the functional roles of several candidates, including C08F11.7, ugt-62, and K05C4.9, supporting their contributions to longevity and stress resistance. The authors conclude that longevity is mediated through multiple molecular pathways and provide a public online tool to study these complex transcriptomic landscapes.

      Significance:

      This study provides a systematic, side-by-side transcriptomic comparison of nine genetically distinct long-lived C. elegans mutants, revealing that lifespan extension arises from both shared and opposing gene expression programs. By identifying three distinct longevity groups and demonstrating that key pathways can be modulated in opposite directions to achieve long life, the work challenges the notion of a single universal transcriptional signature of aging. Importantly, functional validation shows that select commonly regulated genes can directly modulate lifespan and stress resistance, highlighting actionable molecular targets for promoting healthy aging.

      Comments on revised version:

      The authors addressed my concerns successfully.

    4. Author response:

      Reviewer #1:

      Major comments

      (1) Although I myself believe that the datasets in this study should be more consistent and comprehensive, the authors should perform a data mining analysis of previously reported transcriptomic changes of these mutants or similar mutants in the same longevity pathway and compare the reported changes with their findings to highlight the necessity and advances of this study.

      According to this suggestion, we have compared the differentially expressed genes identified in this study to previous gene expression studies involving these long-lived mutant strains. To our knowledge no previous studies have examined gene expression in sod-2 or ife-2 mutants, and at the time that we performed the RNA sequencing gene expression in osm-5 worms had not been examined (it took us a long time to complete this paper). We have included weighted Venn diagrams to illustrate the overlap and supplemental tables to list the overlapping di erentially expressed genes. For our current study, we felt it was important to compare RNA-seq data generated under exactly the same experimental and analysis paradigms in order to best compare across the nine long-lived mutants. These new analyses are included in Figures S19 – S25 and Table S2 . Please see lines 111-114, Figure S19-25, and Table S2.  

      (2) This manuscript does not perform any regulon or transcription factor (TF) analyses. TFs are the drivers of the transcriptomic changes and multiple conserved TFs (e.g., daf-16) have already been identified in these pathways. Therefore, it is necessary to examine and compare the regulons/TFs in these new datasets by bioinformatics. Such analyses can: a) provide more information of the driving force of these transcriptomic changes; b) show the role of these known longevity TFs; c) propose new TFs driving longevity; d) support the findings of 'longevity strategies' and 'longevity groups' from the perspective of TFs.

      According to this suggestion, we have now performed transcription factor analysis on the RNA-seq data to determine which transcription factors might be driving the longevity-associated transcriptional changes. To do this we used two complementary approaches: (1) transcription factor inference, which is based on the coordinated expression changes of known transcription factors; and (2) motif enrichment analysis, which is based on identifying transcription factor binding motifs in the promoters of di erentially expressed genes. After identifying which transcription factors were identified for each individual mutant, we then compared the identified transcription factors across all nine mutants. Interestingly, while 33 of the same transcription factors were implicated in group 1 and group 2 longevity mutants, 25 are modulated in different directions (activated in group 1, repressed in group 2 or vice versa) while only 5 are modulated in the same direction. This indicates that although group 1 and group 2 longevity mutants may modulate overlapping pathways to achieve long lifespan, in most cases these pathways are modulated in opposite directions. These new analyses are included in Figure S31 and Table S5. Please see lines 194-208, Figure S31, and Table S5.  

      (3) osm-5 and daf-2 are categorized into two different groups in this study. Since the longevity of cilia (-) mutants is through daf-16, the same master TF driving daf-2 longevity, please perform further analyses or discussion to clarify this issue.

      Loss of daf-16 is generally detrimental to lifespan. Disruption of daf-16 decreases the lifespan of all nine long-lived mutants that we examined (see supplemental table in our review paper PMID:37127095). However, loss of daf-16 also decreases wild-type lifespan. Thus, without further evidence it is hard to distinguish between the loss of daf-16 non-specifically decreasing lifespan verse activation of DAF-16 actually contributing to lifespan extension. In daf-2 mutants and the long-lived mitochondrial mutants there is increased nuclear localization of DAF-16 and upregulation of DAF-16 target genes. The differentially expressed genes in the long-lived mitochondrial mutants exhibit about a 50% overlap with the differentially expressed genes in daf-2 mutants (see Author response image 1). In contrast, osm-5 mutants show upregulation of some DAF-16 upregulated genes, no change in some DAF-16 upregulated genes and downregulation of other DAF-16 upregulated genes (see Author response image 1). Only about 10% of the differentially expressed genes in osm-5 mutants overlap with differentially expressed genes in daf-2 mutants. We believe that these results are consistent with loss of DAF-16 causing a general decrease in lifespan and not specifically contributing to osm-5 longevity. These comparisons will be included in a manuscript that we are currently preparing on osm-5 mutant longevity.

      Author response image 1.

      (4) This manuscript focused on genes whose RNAi suppressed the mutants longevity. Please also use bioinformatics to analyze the functions of those whose RNAi extends the mutants longevity, because these genes could tell the health price these mutants pay and help improve ageing interventions by reducing side effects.

      We perform enrichment analysis for both genes upregulated and downregulated in the long-lived mutant strains. The downregulated genes are involved in translation, ribosome biogenesis and gene expression. For the RNAi screen, we aimed to identify genes that are contributing to longevity and so we looked for a decrease in the lifespan of long-lived mutants when treated with RNAi. We did not screen for genes that extend the long-lived mutants longevity. While we did, nonetheless, identify multiple RNAi clones that increased either daf-2 or nuo-6 lifespan, there were not enough genes to identify any patterns of enrichment.

      (5) (OPTIONAL) I strongly suggest a comprehensive comparison of these transcriptomic changes in long-lived mutants with published age-related transcriptomic changes in wild type worms.

      According to this suggestion, we have now compared the differentially expressed genes that we identified in the nine long-lived mutants with genes that were found to be differentially expressed with aging. Interestingly, the group 2 long-lived mutants show a larger overlap for genes modulated in the opposite direction as aging (genes downregulated during aging are upregulated in eat-2 and osm-5 mutants). We have added this new analysis to our manuscript. Please see lines 210-223, Figure S32 and Table S6.

      Minor comments

      (1) Please further clarify the analysis of DEGs correlated with lifespan extension in Fig. 2 by a depiction. In Fig. 2C and D, please label data dots from different strains with different colors.

      According to this suggestion, each strain has been labelled a different colour.

      (2) In Fig. 3 and S20, please label the percentage of overlapping genes on top of each bars.

      We have now labelled the percentage of overlapping genes in Figure 3 and S20 (now S27).

      Reviewer #2:

      Major comments

      (1) While the authors identified a set of 196 upregulated genes, the rationale for narrowing these down to the three final candidates (C08F11.7, ugt-62, and K05C4.9) is not clearly described. The authors show that genetic inhibition of several genes, including DC2.5, C05B5.5, T07C4.5, and W03B1.7, decreases lifespan in both nuo-6 mutants and wild-type animals. However, the authors did not describe why these additional validated candidates, which also showed significant effects on longevity, were not pursued for further

      characterization. The authors should explicitly state the criteria used to prioritize these three genes over the other validated genes.

      Due to the costs and time involved in generating and characterizing new strains, we decided that we would select three strains to study further as a proof-of-principle. When deciding which genes to study further, we considered several approaches. In the end, we chose to use the strength/reproducibility of the increase in weighted mortality to identify genes with a clear, consistent impact. C05B5.5 and T07C4.5 were ruled out because they had an inconsistent impact on weighted mortality (Figure S28). W03B1.7 was ruled out because it did not have a strong enough e ect on weighted mortality (Figure S28). That narrowed it down to C08F11.7, ugt-62, DC2.5, and K05C4.9. Of those 4, C08F11.7, ugt-62, and K05C4.9 have the greatest consistent impact on weighted mortality (Figure S28) and so these genes were chosen. We have updated the manuscript to include this justification for focussing on C08F11.7, ugt-62, and K05C4.9. Please see lines 273-278.

      (2) The authors conclude that longevity can be mediated by multiple molecular pathways. However, it remains unclear whether these distinct strategies can operate simultaneously or are mutually exclusive. The authors need to test whether lifespan extension in a Group 1 mutant is further enhanced or suppressed by the knockdown of a key Group 2-specific genes. These experiments would help determine these pathways act additively, antagonistically, or as partially redundant survival programs.

      This is an excellent suggestion. While our data identify several genes that are regulated in opposite directions in group 1 and group 2 longevity mutants, we do not yet know the extent to which each of these genes contribute to the longevity of group 1 and group 2 mutants. The three genes that we focused on for further characterization (C08F11.7, ugt-62 and K05C4.9) are upregulated in group 1 longevity mutants but not group 2 mutants. Contrary to what might be expected, RNAi knockdown of these genes does not decrease the lifespan of the group 1 longevity mutant daf-2 but does decrease the lifespan of the group 2 longevity mutant eat-2. We recently reviewed the e ect of di erent resilience pathways on the lifespan of long-lived genetic mutants. Disruption of daf-16, sek-1, skn-1, hsf-1, ire-1 and trx-1 can decrease lifespan in both group 1 and group 2 longevity mutants, but also decreases lifespan in wild-type worms suggesting that at least in some mutants the e ect on longevity may be non-specific. Disruption of hif-1 does not a ect the longevity of group 2 mutants, but does a ect the lifespan of some group 1 mutants (clk-1, isp-1, nuo-6) but not others (daf-2, glp-1). To more definitively answer the question, it would be interesting to cross different combinations of group 1 and group 2 longevity mutants to see the extent to which different longevity groups synergize. This is something we are currently working on for a separate manuscript. We have added these points to the revised manuscript. Please see lines 363-381.

      (3) The authors provide interesting data on overexpression of the three candidate genes. However, whereas C08F11.7 clearly demonstrates both necessity and sufficiency for lifespan extension, overexpression of ugt-62 and K05C4.9 does not independently extend lifespan. To strengthen the manuscript, the authors should expand the discussion of these divergent results and clarify possible explanations.

      According to this suggestion, we have expanded our discussion to discuss possibilities of why these genes might be having different effects on lifespan. Please see lines 411-423.

      (4) Key citations are missing and the authors should add multiple citations including the following ones. Please cite the following paper and discuss the authors' finding with respect to the related work (Lee et al PMID: 40814218). Add citations in the sentence describing changes in the transcriptome of C. elegans associated with age (Lee et al., PMID: 38508494). Furthermore, please cite papers describing the overviews of survival assay using C. elegans (Kwon et al., PMID: 40436148, Hwang et al., PMID: 40436147).

      We have added the suggested citations to the revised manuscript. Please see lines 211 (Ref #39), 307 (Ref #41), 423 (Ref #54) and 436 (Ref #55).

      Minor comments

      (1) To improve readability, please provide the full names for all abbreviations at their first appearance in the manuscript.

      We have added the full names for each abbreviation on first appearance.

      (2) Please ensure that the labels in the figures match the text exactly. For instance, if different promoters are used for generating overexpression animals, it may be helpful to indicate the specific promoter in the figure panel or legend for clarity.

      We have ensured that the nomenclature in the text and the figures is the same. We have noted the promoter used for the overexpression strains in the figure legend.

      (3) For all lifespan and stress resistance assays, please include the total number of animals (n) and the number of independent biological replicates (N) in the figure legends to confirm statistical reliability.

      We have added the number of animals and independent biological replicates to the figures and figure legends.

      (4) Please clearly specify the exact developmental stage of the animals used for the survival assays in the Materials and Methods section.

      We have updated the methods to describe the developmental stages used for the survival assays.

    1. eLife Assessment

      This valuable study provides insights into the role of MATR3 in oocyte maturation and folliculogenesis, using conditional knockout mice and in vitro follicle culture systems to show that MATR3 is required for oocyte growth and gene transcription, with downstream effects on follicle development. The evidence is solid, but some minor inadequacies in replication of key methods and independent validation reduce confidence in the conclusions. The work will be of interest to researchers in reproductive biology and fertility.

    2. Reviewer #1 (Public review):

      Summary:

      This study aims to clarify MATR3's function and molecular mechanism in oocyte growth and maturation, explore its association with OMA and its potential as a diagnostic and therapeutic target using specific knockout mouse models, human OMA samples and multi-omics technologies. And it has fully achieved preset objectives with results strongly supporting conclusions. Specifically, it addresses the gap in the synergistic mechanism of epigenetic and secretory signals regulated by RNA-binding proteins (RBPs) in oocyte growth and enriches the molecular etiological spectrum of oocyte maturation disorders. It is the first time to reveal the conservative function of MATR3 in multiple species, providing a paradigm for cross-species research on RBPs in the field of reproductive biology. And it provides a new candidate target for OMA, a clinically refractory infertility disease, and is expected to promote the optimization of assisted reproductive technology and the development of precision medicine.

      Strengths:

      The strengths of this study are significant and prominent. First, the research system is comprehensive, integrating knockout mouse models, in vitro knockdown models, multi-species (mouse, porcine and human) verification, combined with scRNA-seq, LACE-seq, CO-IP and other multi-omics and molecular biology technologies, forming a complete and progressive evidence chain. Second, the mechanism analysis is in-depth, clarifying the dual molecular mechanisms of MATR3 regulating the transcriptional synthesis and secretion of GDF9 through "recruiting KDM3B to regulate H3K9me2 demethylation" and "directly binding to Rdx mRNA", with a clear logical closed loop. Third, the clinical correlation is close. It is the first time to find abnormal nuclear localization of MATR3 in oocytes of OMA patients, providing new clues for clinical disease mechanism research, and verifying the downstream function of GDF9 through rescue experiments, effectively enhancing the translational value of the results.

      Weaknesses:

      This study included only one OMA patient's oocyte sample. Without clinical screening for MATR3 mutations or abnormal expression, establishing a causal relationship between MATR3 and OMA remains difficult.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout mouse models together with in vitro follicle culture and molecular analyses. The authors aim to determine whether MATR3 regulates oocyte maturation and follicle development and to explore potential mechanisms linking MATR3 function to transcriptional and epigenetic regulation in growing oocytes.

      Strengths:

      A major strength of the work is the use of a conditional knockout mouse model combined with complementary in vitro follicle culture approaches, which together provide a useful framework for examining gene function during oocyte development. The study also attempts to integrate cellular phenotypes with molecular analyses of transcriptional activity and epigenetic markers.

      Weaknesses:

      Several weaknesses limit the strength of the conclusions. These include insufficient validation of key experimental manipulations (such as the efficiency of MATR3 knockdown in siRNA experiments), limited quantification or statistical analysis for some datasets, inconsistencies between the text and presented data in certain figures, and incomplete methodological descriptions that make it difficult to fully evaluate reproducibility.

      Comments on revised version.

      Thank you for submitting the revised manuscript. I believe the revisions have substantially improved the quality and clarity of the study, and the authors have addressed the major concerns raised during the initial review.

    4. Reviewer #3 (Public review):

      Summary:

      The study aims to elucidate the dual molecular mechanisms of the RNA-binding protein MATR3 in oocyte growth and maturation. The authors propose that MATR3, highly expressed in growing oocytes (GOs), regulates oocyte quality through two pathways: epigenetically, by recruiting KDM3B to remove the repressive H3K9me2 mark at the Gdf9 locus to activate transcription; and post-transcriptionally, by binding Rdx mRNA to maintain microvillus structure for GDF9 secretion. This mechanism ensures oocyte-granulosa cell communication and female fertility. The study also explores the link between MATR3 and human oocyte maturation arrest (OMA).

      Strengths:

      The study proposes an innovative dual-mechanism model encompassing "epigenetic transcriptional activation and cytoskeletal regulation," which not only expands the functional understanding of RNA-binding proteins in chromatin regulation but also reveals the coordination between nuclear transcription and organelle structure. By integrating scRNA-seq and LACE-seq, the authors constructed a comprehensive regulatory network for MATR3, identifying both key targets and numerous potential molecules, thereby providing rich resources for future mechanistic studies. Furthermore, the inclusion of oocyte samples from human OMA patients directly links the basic findings to clinical reproductive disorders. Despite the limited sample size, this approach demonstrates strong translational potential.

      Weaknesses:

      The partial phenotypic improvement achieved by exogenous GDF9 supplementation suggests that the downstream effector pathways may involve a more complex network regulation, implying that the current interpretation of GDF9 central role could be further explored. Regarding the developmental abnormalities of granulosa cells in the conditional knockout model, their pathological origins require in-depth analysis to determine whether they represent primary alterations or secondary adaptive responses resulting from the loss of oocyte signaling.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study aims to clarify MATR3's function and molecular mechanism in oocyte growth and maturation, explore its association with OMA, and its potential as a diagnostic and therapeutic target using specific knockout mouse models, human OMA samples, and multi-omics technologies. And it has fully achieved preset objectives with results strongly supporting conclusions. Specifically, it addresses the gap in the synergistic mechanism of epigenetic and secretory signals regulated by RNA-binding proteins (RBPs) in oocyte growth and enriches the molecular etiological spectrum of oocyte maturation disorders. It is the first time the conservative function of MATR3 has been revealed in multiple species, providing a paradigm for cross-species research on RBPs in the field of reproductive biology. It also provides a new candidate target for OMA, a clinically refractory infertility disease, and is expected to promote the optimization of assisted reproductive technology and the development of precision medicine.

      Strengths:

      The strengths of this study are significant and prominent. First, the research system is comprehensive, integrating knockout mouse models, in vitro knockdown models, multi-species (mouse, porcine, and human) verification, combined with scRNA-seq, LACE-seq, CO-IP, and other multi-omics and molecular biology technologies, forming a complete and progressive evidence chain. Second, the mechanism analysis is in-depth, clarifying the dual molecular mechanisms of MATR3 regulating the transcriptional synthesis and secretion of GDF9 through "recruiting KDM3B to regulate H3K9me2 demethylation" and "directly binding to Rdx mRNA", with a clear logical closed loop. Third, the clinical correlation is close. It is the first time to find abnormal nuclear localization of MATR3 in oocytes of OMA patients, providing new clues for clinical disease mechanism research, and verifying the downstream function of GDF9 through rescue experiments, effectively enhancing the translational value of the results.

      Weaknesses:

      This study included only one OMA patient's oocyte sample. Without clinical screening for MATR3 mutations or abnormal expression, establishing a causal relationship between MATR3 and OMA remains difficult.

      We greatly appreciate positive comments and constructive feedback on our manuscript.

      We are encouraged that you recognize the novelty, rigour, and clinical relevance of our study on MATR3 in oocyte development and OMA. We have carefully considered your comments and revised the manuscript accordingly. We will further expand the OMA patient cohort in future studies to verify the causal relationship between MATR3 and OMA.

      Reviewer #2 (Public review):

      Summary:

      This study investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout mouse models together with in vitro follicle culture and molecular analyses. The authors aim to determine whether MATR3 regulates oocyte maturation and follicle development and to explore potential mechanisms linking MATR3 function to transcriptional and epigenetic regulation in growing oocytes.

      Strengths:

      A major strength of the work is the use of a conditional knockout mouse model combined with complementary in vitro follicle culture approaches, which together provide a useful framework for examining gene function during oocyte development. The study also attempts to integrate cellular phenotypes with molecular analyses of transcriptional activity and epigenetic markers.

      Weaknesses:

      Several weaknesses limit the strength of the conclusions. These include insufficient validation of key experimental manipulations (such as the efficiency of MATR3 knockdown in siRNA experiments), limited quantification or statistical analysis for some datasets, inconsistencies between the text and presented data in certain figures, and incomplete methodological descriptions that make it difficult to fully evaluate reproducibility.

      We greatly appreciate your constructive comments and suggestions. We are grateful for the recognition of our conditional knockout mouse model and experimental design. We have carefully addressed all the weaknesses mentioned by the reviewer, including the validation of key experiments, quantitative and statistical analysis, consistency between text and figures, and detailed methodological descriptions. Details are described point-by-point below.

      Reviewer #3 (Public review):

      Summary:

      The study aims to elucidate the dual molecular mechanisms of the RNA-binding protein MATR3 in oocyte growth and maturation. The authors propose that MATR3, highly expressed in growing oocytes (GOs), regulates oocyte quality through two pathways: epigenetically, by recruiting KDM3B to remove the repressive H3K9me2 mark at the Gdf9 locus to activate transcription; and post-transcriptionally, by binding Rdx mRNA to maintain microvillus structure for GDF9 secretion. This mechanism ensures oocyte-granulosa cell communication and female fertility. The study also explores the link between MATR3 and human oocyte maturation arrest (OMA).

      Strengths:

      The study proposes an innovative dual-mechanism model encompassing "epigenetic transcriptional activation and cytoskeletal regulation," which not only expands the functional understanding of RNA-binding proteins in chromatin regulation but also reveals the coordination between nuclear transcription and organelle structure. By integrating scRNA-seq and LACE-seq, the authors constructed a comprehensive regulatory network for MATR3, identifying both key targets and numerous potential molecules, thereby providing rich resources for future mechanistic studies. Furthermore, the inclusion of oocyte samples from human OMA patients directly links the basic findings to clinical reproductive disorders. Despite the limited sample size, this approach demonstrates strong translational potential.

      Weaknesses:

      The partial phenotypic improvement achieved by exogenous GDF9 supplementation suggests that the downstream effector pathways may involve a more complex network regulation, implying that the current interpretation of GDF9's central role could be further explored. Regarding the developmental abnormalities of granulosa cells in the conditional knockout model, their pathological origins require in-depth analysis to determine whether they represent primary alterations or secondary adaptive responses resulting from the loss of oocyte signaling.

      We greatly appreciate your positive and insightful comments on our study. We are grateful for the recognition of our novel dual-mechanism model, comprehensive multi-omics analysis, and translational potential from basic research to clinical OMA. We have carefully addressed the weaknesses raised by the reviewer, including in-depth discussion of the GDF9-centered regulatory network and clarification of the origin of granulosa cell abnormalities. More details point-by-point responses are provided below.

      Recommendations for the authors:

      Point-by-point responses to reviewers’ comments

      We thank the reviewer very much for his/her reviewing of our work, and we appreciate the constructive comments and suggestions that have helped us to prepare an improved revision. Based on the comments of the reviewer, we have carefully revised the manuscript by performing some new experiments.

      Reviewer #1 (Recommendations for the authors):

      (1) Did most of the follicles cultured in vitro reach the antral follicle stage after 6 days?

      We greatly appreciate your insightful question. We statistically analyzed the survival rate and antral follicle ratio of in vitro-cultured follicles after 6 days of culture. Due to differences in culture systems and protocols, the follicle survival rate in our study (57.43 ± 3.11%) was different from that reported in previous literature (92 ± 10%). However, the proportion of antral follicles among surviving follicles was highly consistent between our results and published data (83 ± 13% vs 80.87 ± 3.27%) (Cortvrindt and Smitz 2002).

      Author response image 1.

      Ratio and survival rate of antral follicles after 6 days of culture. n = 3. Data are represented as mean ± SD.

      (2) In Figure 2F, at which stage did MATR3 begin to affect oocyte diameter?

      Thank you for your careful observation. Our morphological analysis of oocytes collected from PD14 and PD23 mice showed no significant difference in oocyte diameter between the cKO and Ctrl groups at the GO stage (Fig. S3D, E). However, oocytes in the cKO group became significantly smaller than those in the Ctrl group once they reached the FGO stage (Fig. 2E, F). Taken together, these results indicate that the growth defect caused by MATR3 deletion begins to manifest during the transition from the GO to FGO stage, with significant reduction in oocyte diameter clearly observed at the FGO stage as shown in Figure 2F.

      (3) What was the developmental potential of oocytes in Matr3-knockout mice?

      Thank you for this important question. Compared with the Ctrl group, oocytes derived from cKO mice showed a drastically reduced fertilization rate (91.55 ± 1.96% vs 10.55 ± 4.78%) and almost completely failed to develop to the blastocyst stage (76.62 ± 7.56% vs 3.67 ± 3.38%). These results clearly demonstrate that maternal deletion of Matr3 severely compromises the developmental potential of mouse oocytes, including fertilization capacity and subsequent early embryonic development.

      Author response image 2.

      Results of in vitro fertilization of oocytes. 2-cell: 2 days after fertilization; blastocyst: 4 days after fertilization. Data are represented as mean ± SD. ***P < 0.001.

      (4) The legend labels in the figures should not be bold.

      Thank you for your valuable suggestion. We have revised all the figures accordingly.

      Reviewer #2 (Recommendations for the authors):

      This manuscript investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout (cKO) mouse models combined with in vitro follicle culture approaches. The topic is relevant to the field of reproductive biology and provides potentially important insights into the molecular mechanisms regulating oocyte maturation and follicle development.

      While the study presents interesting observations and utilizes both in vivo and in vitro experimental systems, several issues need to be addressed before the manuscript can meet the expected standards. These include concerns related to data interpretation, validation of experimental approaches, completeness of methodological descriptions, and clarity in data presentation. In addition, the manuscript requires substantial language editing to improve clarity and readability.

      The comments below outline major issues that should be addressed to strengthen the manuscript, as well as specific minor points regarding presentation and clarity.

      (1) The manuscript requires substantial revision to improve the written language and grammar. Numerous sentences are unclear or awkwardly phrased, which makes interpretation of the results difficult in several sections. The authors are strongly encouraged to have the manuscript professionally edited or thoroughly revised for language and clarity before making a resubmission.

      We sincerely appreciate the careful and constructive comments on the language quality and clarity of the manuscript. We fully agree that the written language, grammar, and sentence structure need substantial improvement to ensure the results are presented clearly and accurately.

      To address these concerns thoroughly, we have carefully revised the entire manuscript, including correcting grammatical errors, refining awkward phrasing, and restructuring unclear sentences to enhance readability and logical flow. In addition, we have sought professional language editing support to further polish the English expression and ensure the manuscript meets the linguistic standards of the journal.

      All revisions related to language and clarity have been completed, and we believe the revised version is significantly improved in terms of readability and precision.

      (2) Interpretation of oocyte maturation results (Line 140; Figure 2E, H). The manuscript states: "During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig. 2E, H)." However, Figure 2H appears to show that a small proportion of knockout oocytes do reach the MII stage. Therefore, the description in the text seems inconsistent with the data presented. The authors should clarify the exact maturation rates in both groups, revise the text to accurately reflect the data, and provide statistical analysis to support the stated conclusions.

      Thank you for your valuable comment. A small proportion of knockout oocytes from PD23 cKO mice can indeed develop to the MII stage. We have revised the corresponding description and supplemented the statistical analysis of maturation rates to support our conclusion.

      Line 140: “During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig.2E, H).” have been replaced by “During in vitro maturation, the proportion of oocytes from PD23 cKO mice developing to metaphase II stage was significantly reduced (Fig.2E, H, 54.9±2.08% vs 9.57±1.11%).”

      (3) Human oocyte sample size: In Figure 1D, it is unclear how many human oocytes were analyzed. It is important to specify the sample size (n) for all experiments. The authors should clearly indicate the number of oocytes analyzed in this experiment. Provide this information either in the figure legend or in the main text.

      Thank you for this important comment. We agree that the altered subcellular localization of MATR3 in human OMA oocytes is of great physiological significance for understanding the functional role of MATR3 during oocyte development.

      Unfortunately, during a 3‑month period of sample collection, we examined MATR3 localization in immature oocytes that failed to reach the MII stage, obtained from 11 women undergoing IVF treatment. Among these samples, only one donor’s oocytes exhibited the NSN chromatin configuration. Excitingly, these NSN‑stage oocytes from this donor clearly showed the loss of MATR3 nuclear localization, which strongly supports the critical role of MATR3 during oocyte growth and maturation. We have now clearly stated the sample size (n = 11) in the figure legend and main text as suggested. In future studies, we will continue to collect more human oocyte samples to further validate these observations with an expanded sample size.

      (4) Figure annotation issue: The figure legend for Figure 1F refers to an arrow, but no arrow is visible in the figure panel. Please correct this inconsistency by either adding the appropriate arrow to the figure or revising the legend accordingly.

      Thank you for pointing out this error. We have revised the figure legend for Figure 1F accordingly to correct this inconsistency.

      (5) Description of follicle analysis (Line 146): The sentence: "This was reinforced by the data of available follicles within the follicles of mice on PD35 (Fig. 2I, J)." is incorrect or poorly phrased. It should likely read: "...available follicles within the ovaries of mice at PD35...".

      We really appreciate your constructive suggestion on the phrasing. We have revised this sentence in the revised manuscript accordingly.

      (6) Quantification of proliferating cells: Figure 2K shows Ki-positive cells, but quantitative analysis is not provided. The authors should quantify the number or proportion of Ki-positive cells in both control and cKO groups and include statistical analysis to support any claims regarding differences in proliferation.

      Thank you for your valuable suggestion. We have quantified the number of Ki‑67‑positive granulosa cells in both control and cKO groups and performed the corresponding statistical analysis.

      The quantitative results have been added to Fig. S3G, and the relevant description has been supplemented in the main text at Line 150 to support our conclusion regarding cell proliferation differences.

      Line 150: “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K) in cKO mice were lower than those found in the Ctrl.” have been replaced by “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K, Fig. S3G) in cKO mice were lower than those found in the Ctrl (73.55±13.29% vs 24.65±7.80%).”

      (7) Validation of findings in the in vivo cKO model (Figure 3): The development of an in vitro follicle culture system is an interesting and valuable component of the study. However, several key analyses performed in vitro (e.g., transcription assays and analysis of epigenetic markers) should ideally also be validated in oocytes derived from the in vivo cKO model.

      Thanks for the valuable concern. We fully agree with you that the in vivo cKO model should be used to validate several key analyses performed in vitro. We collected growing oocytes from Ctrl and cKO mice and conducted transcription assays as well as analysis of epigenetic markers. The results showed that Matr3 knockout significantly downregulated transcriptional activity in GO and increased H3K9me2 levels (Author response image 3), which is consistent with our in vitro findings (Fig 3D E I J). These in vivo results confirm that MATR3 plays a critical role in regulating GO transcriptional activity and H3K9me2 levels.

      Author response image 3.

      Matr3 knockout results in the reduction of transcriptional activity. A EU staining (green) in GO collected from Ctrl and cKO. n = 15. B Quantification of the mean fluorescence intensity of EU in oocytes. C H3K9me2 staining (red) in GO collected from Ctrl and cKO. n = 15. D Quantification of the mean fluorescence intensity of H3K9me2 in oocytes. Scale bar: 20 μm. Data are represented as mean ± S.D. ***P < 0.001.

      (8) To strengthen the conclusions, the authors should consider repeating key experiments using oocytes directly isolated from the cKO mice. This would help confirm that the observed effects are not artifacts of the in vitro culture system.

      Thank you for this valuable and constructive suggestion. We fully agree that the conditional knockout mouse model is essential for verifying the physiological significance of MATR3 in vivo.

      To address this point, we have validated multiple key in vitro findings using oocytes directly isolated from cKO mice. For instances, the changes in oocyte transcriptional activity (EU staining) (in Comments 7), H3K9me2 levels (in Comments 7), GDF9 levels (Fig 4.B C E F), and OO-Mvi (Fig 6.A B, Author response image 4) all showed consistent trends with our in vitro knockdown results. In addition, the complete infertility phenotype of cKO female mice further demonstrates that MATR3 is indispensable for oocyte growth and meiotic maturation. We have also provided supplemental data from GDF9 rescue experiments and sequencing analysis performed in the mouse model.

      Author response image 4.

      Matr3 knockdown impairs the structural integrity of oocyte OO-MVi. A p-ERM staining (green) showing the OO-Mvi in oocyte from NC and si-Matr3. B Quantification of the number of Oo-Mvi vesicles (n = 6). Scale bar: 20 μm. Data are represented as mean ± SD. ***P < 0.001.

      In conclusion, the core conclusions of this study are supported by the mutual validation of key experimental results from MATR3-specific knockdown in vitro and Matr3 conditional knockout mouse models in vivo.

      (9) (1) Validation of MATR3 knockdown: The in vitro MATR3 knockdown experiment presented in Figure 4G raises an important concern: it is unclear whether Matr3 knockdown was effectively achieved in the oocytes analyzed. The authors should provide direct evidence of knockdown efficiency, for example, immunostaining for MATR3 protein on the oocytes. Without such validation, it is difficult to interpret the functional outcomes observed.

      As requested, we have provided direct evidences of the knockdown efficiency via immunostaining, which is now presented in Fig.S2D.

      In this experiment, oocytes from early growing follicles (approximately 150 μm in diameter) were microinjected with Matr3 siRNA. Following 5 days of continuous in vitro culture, oocytes from both the NC and si-Matr3 groups were isolated and subjected to immunofluorescence staining to assess protein levels. As shown in the figure, the oocytes at this stage exhibited the characteristic non-surrounded nucleolus (NSN) chromatin configuration. We observed robust MATR3 protein expression within the nucleus of NC oocytes, whereas the MATR3 protein levels were markedly reduced in the si-Matr3 group. These results confirm the successful construction of the Matr3 knockdown model in early growing follicle oocytes.

      (9) (2) Furthermore, it would be more convincing if the authors could perform the Gdf9 supplementation experiments using follicles isolated from the cKO mice, rather than relying solely on siRNA knockdown in vitro. Such experiments would provide clearer and more physiologically relevant evidence. If these experiments were attempted but did not produce similar results, this should be discussed.

      Thank you for your valuable and insightful suggestion. We fully agree that performing GDF9 supplementation experiments using follicles isolated from cKO mice would provide more direct and physiologically relevant evidence to strengthen our conclusions.

      Unfortunately, when we attempted to conduct GDF9 rescue experiments on follicles from cKO mice, neither the control nor cKO follicles were able to develop to the antral follicle stage (n=3). We speculate that this was caused by insufficient bioactivity of the veterinary-grade FSH been used, as compared to the imported FSH been provided by NHPP. Unfortunately, this particular FSH product has been discontinued. We are currently actively seeking and attempting to purchase new, qualified FSH reagents to repeat these experiments and further validate our findings in future work.

      (10) Figure citation order: Figures are not cited sequentially in the text. For example, Figure 6J is described first (line 250), followed by Figure 6A. Figures should be discussed in logical order, typically starting from panel A. Please revise the text to ensure that figure panels are introduced sequentially.

      Thank you for this careful and important comment.

      We have carefully revised the citation order of all figure panels in the main text, especially for Figure 6, to ensure they are introduced sequentially from panel A to the last panel in logical and numerical order, rather than being cited out of sequence.

      The corresponding adjustments have been made in the revised version of the manuscript.

      (11) Figure 6J interpretation: The purpose of the images shown in this figure is unclear. The authors should provide higher magnification images to clearly visualize the Oo-Mvi structures and include quantification of the observed phenotype to support the interpretation. It is important because the main findings of the paper heavily rely on these results.

      Thank you for this valuable and constructive suggestion. We fully agree that higher‑magnification images and quantitative analysis are essential to clearly demonstrate the Oo‑Mvi structures and reliably support our conclusions, especially given the importance of these results to the main findings of this study.

      Accordingly, we have replaced the original panels in Figure 6A with higher‑magnification images to better visualize Oo‑Mvi structures. In addition, we have supplemented the corresponding quantitative analysis of the observed phenotype to strengthen the interpretation of this figure (Fig 6B). All revisions have been incorporated into the revised manuscript.

      (12) Incomplete Materials and Methods section: The Materials and Methods section lacks important experimental details required for reproducibility. Specifically, the Matr3 flox mouse model. Either provide the appropriate reference describing the Matr3 floxed mice or include details on how the floxed allele was generated.

      Thank you for your valuable and careful comment. We fully agree that detailed experimental information in the Materials and Methods section is crucial for ensuring the reproducibility of the study. And we apologize for the omission of key details regarding the Matr3 flox mouse model.

      In response to your suggestion, we have thoroughly supplemented the relevant experimental details in the Materials and Methods section of the revised manuscript, including the specific construction strategy of the Matr3 floxed allele. These detailed descriptions will enable other researchers to reproduce our mouse model and verify the experimental results.

      All supplementary information has been integrated into the revised manuscript to meet the requirements of experimental reproducibility. We greatly appreciate your guidance in helping us improve the completeness and rigour of our study.

      (13) Follicle isolation: The manuscript does not describe how growing follicles were isolated. Please specify whether follicles were isolated using enzymatic digestion or mechanical dissection and provide sufficient methodological detail so that other researchers can reproduce the experiments.

      Thank you for this valuable comment. We agree that adding this information is essential for ensuring the reproducibility of our experiments. We have supplemented the corresponding description in the Materials and Methods section. Briefly, growing follicles were isolated by mechanical dissection using insulin syringes under a stereomicroscope, without any enzymatic digestion.

      Reviewer #3 (Recommendations for the authors):

      (1) Since KDM3B and MATR3 interact in cell lines, does this relationship affect the functional localization of KDM3B within oocytes? Specifically, does the localization of KDM3B change in cKO mice (e.g., nuclear export or aggregation)?

      Thank you for this insightful and constructive question. To address whether the interaction between KDM3B and MATR3 influences the functional localization of KDM3B in oocytes, we performed immunofluorescence staining to examine the subcellular distribution of KDM3B in cKO oocytes. Our results demonstrated that the nuclear localization of KDM3B remained unaltered; no obvious nuclear export or abnormal aggregation was observed in MATR3-deficient oocytes (Author response image 5).

      Based on these observations combined with our other experimental data, we propose that MATR3 regulates oocyte transcriptional activity through its physical interaction with KDM3B, rather than by controlling the nuclear targeting of KDM3B. Notably, despite unchanged nuclear localization of KDM3B in MATR3 cKO oocytes, we detected significantly elevated global levels of H3K9me2 and markedly reduced transcriptional activity (in Comments 7). These findings indicate that KDM3B loses its physiological function of demethylating H3K9me2 and promoting transcription in the absence of MATR3.

      In line with this mechanism, previous results showed that KDM3B knockout in female mice leads to follicle arrest at the secondary follicle stage and consequent infertility (Liu et al. 2015). Collectively, we conclude that in MATR3 cKO oocytes, although KDM3B is properly localized in the nucleus, it fails to execute its H3K9me2 demethylase activity, thereby impairing normal transcriptional regulation during oocyte development.

      Author response image 5.

      Matr3 knockout has no effect on KDM3B localization. KDM3B staining (red) in GO collected from Ctrl and cKO. n   = 50. Scale bar: 40 μm.

      (2) As the GDF9 rescue experiment only partially restores the phenotype, it is suggested to select 2-3 novel targets with high binding intensity and significant expression changes from the LACE-seq data (e.g., Igf2bp2 or Ccnb1 mentioned in the text) to further illustrate that MATR3 regulates a network.

      Thanks for this meaningful and constructive suggestion.

      We have supplemented the relevant data in Fig. S9 and further elaborated on these findings in the Discussion section in our revised manuscript, following your advice. Briefly, we collected growing oocytes from Ctrl and cKO mice and performed RT‑qPCR analysis to verify the expression of Igf2bp2 and Ccnb1 - two representative novel targets with strong binding intensity and significant expression changes identified from our LACE‑seq bioinformatics analysis. The results showed that both genes were significantly downregulated in cKO oocytes compared with controls, supporting the notion that MATR3 regulates a functional RNA network during oocyte development.

      (3) Are the granulosa cell defects primary or secondary? It is recommended to collect ovaries from earlier-stage cKO mice (e.g., PD7 or PD10) to examine the levels of FOXL2 and PCNA in granulosa cells.

      Thank you for your valuable comment. To clarify whether the granulosa cell defects are primary or secondary, we further investigated the temporal effect of MATR3 deficiency on granulosa cells by collecting ovaries from PD7, which is a critical period for primordial follicle activation and early follicular development.

      To evaluate the status of granulosa cells, we performed immunofluorescence staining on PD7 ovarian sections using FOXL2 (a specific marker for granulosa cells) to quantify the number of granulosa cells, and Ki67 (a proliferation-related marker) to assess the proliferative capacity of granulosa cells. The results, as shown in Author response image 6, demonstrated that there were no significant differences in either the number of granulosa cells or their proliferation levels in primary follicles between cKO and Ctrl.

      These findings are consistent with the data in our supplementary Fig S3F, where we observed no significant differences in the number of primordial follicles and growing follicles at PD7 between the two groups. Collectively, these results indicate that the activation of primordial follicles is not affected by MATR3 deficiency in oocytes, and the impairment of granulosa cells caused by oocyte-specific Matr3 knockout occurs at the secondary follicle stage rather than the early stages of follicular development. Therefore, we conclude that the granulosa cell defects in cKO mice are secondary to the oocyte dysfunction induced by MATR3 deficiency.

      Author response image 6.

      Loss of MATR3 in oocytes does not affect the number and proliferation of granulosa cells in primordial follicles. A Immunohistochemistry results showing granulosa cells in PD7 ovaries from Ctrl and cKO. B Quantification of granulosa cell number in the largest cross-section of primary follicles. C Quantification of the proliferation rate of granulosa cells in the largest cross-section of primary follicles. n = 15. Data are represented as mean ± SD. n.s., not significant.

      References:

      (1) Cortvrindt RG, Smitz JE. 2002. Follicle culture in reproductive toxicology: a tool for in-vitro testing of ovarian function? Human reproduction update 8: 243-254.

      (2) Liu Z, Chen X, Zhou S, Liao L, Jiang R, Xu J. 2015. The histone H3K9 demethylase Kdm3b is required for somatic growth and female reproductive function. International journal of biological sciences 11: 494-507.

    1. eLife Assessment

      One-carbon tetrahydrofolate metabolism plays a crucial role in producing essential metabolic intermediates. In this valuable study, the authors employ a solid genetics-based approach to demonstrate that three distinct metabolic pathways are essential for synthesising 1C-tetrahydrofolates (1C-THF). Disrupting any of these pathways impairs both growth and virulence.

    2. Reviewer #1 (Public review):

      Summary:

      This study identifies three redundant pathways-glycine cleavage system (GCS), serine hydroxymethyltransferase (GlyA), and formate-tetrahydrofolate ligase/FolD-that feed the one-carbon tetrahydrofolate (1C-THF) pool essential for Listeria monocytogenes growth and virulence. Reactivation of the normally inactive fhs gene rescues 1C-THF deficiency, revealing metabolic plasticity and vulnerability for potential antimicrobial targeting.

      Strengths:

      (1) Novel evolutionary insight-Reversible reactivation of a pseudogene (fhs) shows adaptive metabolic plasticity, relevant for pathogen evolution.

      (2) They systematically combine targeted gene deletions with suppressor screening to dissect the folate/one-carbon network (GCS, GlyA, Fhs/FolD).

    3. Reviewer #3 (Public review):

      Summary:

      In this study, Freier et al., demonstrate that 3 distinct metabolic pathways are critical for the synthesis of 1C-THF, a metabolite that is crucial for the growth and virulence of Listeria monocytogenes. Using an elegant suppressor screen, they also demonstrate the hierarchical importance of these metabolic pathways with respect to the biosynthesis of 1C-THF.

      Strengths:

      This study uses elegant bacterial genetics to confirm that 3 distinct metabolic pathways are critical for 1C-THF synthesis in L. monocytogenes and lack of either one of these pathways compromises bacterial growth and virulence. The study uses a combination of in vitro growth assays, macrophage-CFU assays and murine infection models to demonstrate this.

      Comments on revisions:

      The revised manuscript is improved, and the additional genetic experiments provide further support for the proposed metabolic model. However, the central conclusion is not fully established without direct measurement of 1C-THF levels. While I appreciate the authors' explanation regarding the technical limitations, quantitative metabolite measurements (e.g., by mass spectrometry) would have provided much stronger evidence linking the genetic perturbations to altered 1C-THF pools.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study identifies three redundant pathways-glycine cleavage system (GCS), serine hydroxymethyltransferase (GlyA), and formate-tetrahydrofolate ligase/FolD-that feed the one-carbon tetrahydrofolate (1C-THF) pool essential for Listeria monocytogenes growth and virulence. Reactivation of the normally inactive fhs gene rescues 1C-THF deficiency, revealing metabolic plasticity and vulnerability for potential antimicrobial targeting

      Strengths:

      (1) Novel evolutionary insight - reversible reactivation of a pseudogene (fhs) shows adaptive metabolic plasticity, relevant for pathogen evolution.

      (2) They systematically combine targeted gene deletions with suppressor screening to dissect the folate/one-carbon network (GCS, GlyA, Fhs/FolD).

      Weaknesses:

      (1) The study infers 1C-THF depletion mostly genetically and indirectly (growth rescue with adenine) without direct quantification of folate intermediates or fluxes. Biochemical confirmation, LC-MS-based metabolomics of folates/1C donors, or isotopic tracing would strengthen mechanistic claims.

      We agree with the reviewer that quantification of 1C-THF intermediates would strengthen our conclusions. However, the chemical methodologies to extract folates from L. monocytogenes are not established in our lab. Moreover, quantification of C1-substituted folates requires comprehensive biochemical and analytical expertise that we also do not have and which we cannot cover though co-operations. However, to further strengthen our arguments, we have introduced an experiment in the updated manuscript that demonstrates synthetic lethality of a ΔgcvPAB ΔglyA mutant with a deletion of fold (Fig. 7A). This gene encodes 5,10-methylene-tetrahydrofolate dehydrogenase/ 5,10-methylene-tetrahydrofolate cyclohydrolase, which is the third enzyme involved in N5, N10-methylene-THF generation next to GlyA and GcvPBA. Synthetic lethality of a ΔgcvPAB ΔglyA double mutant with a fold deletion is best explained by GcvPAB and GlyA also feeding the N5,N10-methylene-THF pool.

      (2) In multiple result sections, the authors report data from technical triplicates but do not mention independent biological replicates (e.g., Figure 2C, Figure 4A-B, Figure 6D). In addition, some results mention statistical significance but without a detailed description of the specific statistical tests used or replicates, such as Figure 2A-C, Figure 2E, and Figure 2G-I.

      We thank the reviewer for this helpful comment. Experiments were usually repeated three independent times, with each repetition including three technical replicates. Mean values and standard deviations were usually calculated from the technical replicates of a representative run. Statistical significance was calculated using t-tests for pairwise comparisons or t-tests using Bonferroni-Holm correction for multiple comparisons. We made sure that this is explicitly explained for each experiment in the figure legends. Wherever other calculations were used, we also clarified this in the legends.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Freier et al examines the impact of deletion of the glycine cleavage system (GCS) GcvPAB enzyme complex in the facultative intracellular bacterial pathogen Listeria monocytogenes. GcvPAB mediates the oxidative decarboxylation of glycine as a first step in a pathway that leads to the generation of N5, N10-methylene-Tetrahydrofolate (THF) to replenish the 1-carbon THF (1C-THF) pool. 1C-THF species are important for the biosynthesis of purines and pyrimidines as well as for the formation of serine, methionine, and N-formylmethionine, and the authors have previously demonstrated that gcvPAB is important for bacterial replication within macrophages. A significant defect for growth is observed for the gcvPAB deletion mutant in defined media, and this growth defect appears to stem from the sensitivity of the mutant strain to excess glycine, which is hypothesized to further deplete the 1C-THF pool. Selection of suppressor mutations that restored growth of gcvPAB deletion mutants in synthetic media with high glycine yielded mutants that reversed stop codon inactivation of the formatetetrahydrofolate ligase (fhs) gene, supporting the premise that generation of N10-formyl-THF can restore growth. Mutations within the folk, codY, and glyA genes, encoding serine hydroxymethyltransferase, were also identified, although the functional impact of these mutations is somewhat less clear. Overall, the authors report that their work identifies three pathways that feed the 1C-THF pool to support the growth and virulence of L. monocytogenes and that this work represents the first example of the spontaneous reactivation of a L. monocytogenes gene that is inactivated by a premature stop codon.

      Strengths:

      This is an interesting study that takes advantage of a naturally existing fhs mutant Listeria strain to reveal the contributions of different pathways leading to 1C-THF synthesis. The defects observed for the gcvPAB mutant in terms of intracellular growth and virulence are somewhat subtle, indicating that bacteria must be able to access host sources (such as adenine?) to compensate for the loss of purine and fMet synthesis. Overall, the authors do a nice job of assessing the importance of the pathways identified for 1C-THF synthesis.

      Weaknesses:

      (1) Line 114 and Figure 1: The authors indicate that the gcvPAB deletion forms significantly fewer plaques in addition to forming smaller plaques (although this is a bit hard to see in the plaque images). A reduction in the overall number of plaques sounds like a bacterial invasion defect - has this been carefully assessed? The smaller plaque size makes sense with reduced bacterial replication, but I'm not sure I understand the reduction in plaque number.

      The observation that the ΔgcvPAB mutant forms fewer plaques was not our claim, and we have already addressed the possibility of an invasion defect by quantifying bacterial numbers during infection of 3T3 cells. As shown in Fig. 2A, the ΔgcvPAB mutant invades 3T3 cells (the same cells used in the plaque formation assays) as efficiently as the wild type but exhibits reduced intracellular growth. Therefore, the plaquing defect is not due to impaired invasion. Furthermore, we also have analyzed the intracellular dissemination of the ΔgcvPAB mutant in 3T3 fibroblasts compared to a ΔactA mutant by microscopy. This shows that the ΔgcvPAB is evenly distributed throughout the infected host cells as the wild type and unlike the ΔactA mutant, which forms clusters (Fig. S1). Both experiments indicate that the reduced plaque area results from impaired intracellular growth rather than defects in invasion or cell-to-cell spread. The apparent reduction in plaque numbers in ΔgcvPAB-infected 3T3 cells is likely due to a strong reduction in plaque size, with only the largest plaques remaining visible. We have rephrased the relevant section to clarify this point and avoid any potential confusion:

      “In agreement with our previous results, only small plaques were formed in 3T3 cells upon infection with the ΔgcvPAB mutant (plaque area: 15±19% of wild-type level) and small plaques were also formed by the complemented strain in the absence of IPTG (50±19%).”

      (2) Do other Listeria strains contain the stop codon in fhs? How common is this mutation? That would be interesting to know.

      We determined the frequency of inactivated fhs genes among 30,000 publicly available L. monocytogenes genomes. The analysis identified only 10 isolates carrying truncated fhs alleles. These isolates fell into two groups: (i) EGD-e and its descendants, and (ii) a cluster of five ST2 food isolates. These findings have been added as a separate results section.

      (3) Based on the observation that fhs+ ΔgcvPAB ΔglyA mutant is only possible to isolate in complex media, and fhs is responsible for converting formate to 1C-THF with the addition of FolD, have the authors thought of supplementing synthetic media with formate and assessing mutant growth?

      No, we did not test formate supplementation. However, we included additional experiments testing adenine and thymine supplementation (Fig. 6E and 7B). These results show that purine and thymine become limiting in mutants lacking 1C-THF synthesizing pathways.

      Reviewer #3 (Public review):

      Summary:

      In this study, Freier et al. demonstrate that 3 distinct metabolic pathways are critical for the synthesis of 1C-THF, a metabolite that is crucial for the growth and virulence of Listeria monocytogenes. Using an elegant suppressor screen, they also demonstrate the hierarchical importance of these metabolic pathways with respect to the biosynthesis of 1C-THF.

      Strengths:

      This study uses elegant bacterial genetics to confirm that 3 distinct metabolic pathways are critical for 1CTHF synthesis in L. monocytogenes, and the lack of either one of these pathways compromises bacterial growth and virulence. The study uses a combination of in vitro growth assays, macrophage-CFU assays, and murine infection models to demonstrate this.

      Weaknesses:

      (1) The primary finding of the study is that the perturbation of any of the 3 metabolic pathways important for the synthesis of 1C-THF results in reduced growth and virulence of L. monocytogenes. However, there is no evidence demonstrating the levels of 1C-THF in the various knockouts and suppressor mutants used in this study. It is important to measure the levels of this metabolite (ideally using mass spectrometry) in the various knockouts and suppressor mutants, to provide strong causality.

      As already outlined above, we do not have any experimental possibilities to measure 1C-substituted THF in L. monocytogenes extracts directly. However, to provide additional evidence for our interpretation that “Three pathways feed the 1C-THF pool…”, we included additional genetic experiments.

      The first experiment demonstrates that the growth defect of the fhs- ΔglyA igcvPAB strain in synthetic medium lacking IPTG—which reflects the synthetic lethality of fhs with glyA and gcvPAB—can be rescued by the addition of adenine (Fig. 6E). This indicates that the fhs/fold pathway, GlyA, and the glycine cleavage system are essential due to their combined contribution to purine biosynthesis. This result confirms the metabolic model presented in Fig. 1A and thus supports the hypothesis that all three pathways contribute to 1C-THF biosynthesis.

      The second experiment additionally demonstrates synthetic lethality of gcvPAB and glyA with the fold gene. FolD acts downstream of Fhs and is one of the three enzymes synthesizing N5, N10methylene-THF shown in Fig. 1A. The fold gene is essential in EGD-e (PMID: 36114002), most likely explained by fhs inactivation. However, we were able to delete fold in an EGD-e background carrying a reconstituted fhs gene and the resulting fhs<sup>+</sup> Δfold strain was as viable as a fhs<sup>+</sup> ΔglyA ΔgcvPAB strain (Fig. 7A). However, a fhs<sup>+</sup> Δfold ΔglyA igcvPAB strain required IPTG for growth in BHI medium (Fig. 7A), indicating that the simultaneous deletion of glyA and gcvPAB is not tolerated in the absence of fold, similar to what is observed in the absence of fhs. Notably, this growth defect was not rescued by adenine supplementation (Fig. 7B), but was alleviated by thymine addition, which is also consistent with the metabolic model shown in Fig. 1A.

      Even though we are unable to directly demonstrate reduced 1C-THF levels, we hope that these two genetic approaches together with the revised title and heading of the relevant paragraph will, in the reviewers’ eyes, support our hypotheses.

      (2) The story becomes a little hard to follow since macrophage-CFU assays and murine infection model data precede the in vitro growth assays. The manuscript would benefit from a reorganization of Figures 2,3, and 4 for better readability and flow of data.

      We respectfully disagree with the reviewer. The attenuation of the ΔgcvPAB mutant in macrophages and fibroblasts was the primary motivation for further investigating its phenotype. Therefore, we chose to begin the manuscript with results from various virulence studies before presenting the sections that provide mechanistic explanations. In our view, this sequence represents a more logical and coherent way to present the data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Synthetic medium assumptions: LSM "mimics" intracellular limitation but isn't chemically validated against host cytosolic composition. Nutrient availability conclusions could be biased.

      This is correct. We added this information:

      “LSM broth is a chemically defined medium that contains all components required for growth at defined concentrations, but it has not been chemically validated to reflect host cytosolic conditions…”

      (2) Glycine toxicity: The paper introduces the concept of glycine toxicity in ΔgcvPAB mutants. However, the conditions under which glycine becomes toxic could be further elucidated. Why glycine causes toxicity despite being an essential metabolite in other contexts requires a more in-depth mechanistic explanation.

      The concept of glycine toxicity in GCS mutants has been described previously by other researchers. We have added further details to better explain glycine toxicity and how it accounts for the growth phenotype of the ΔgcvPAB mutant:

      “If glycine cannot be catabolized (and 1C-THF cannot be generated) by the GCS due to deletion of gcvPAB, glycine might be re-routed to the serine hydroxymethyl transferase GlyA for serine formation, even though this would consume 1C-THF and therefore even further deplete the cell for 1C-THF.” and

      “In the complete absence of glycine, growth of the ΔgcvPAB mutant was largely unaffected, presumably because glycine cannot be converted to serine by GlyA anymore, thereby conserving the 1C-THF pool.”

      (3) Figures:

      (a) Figure 1B, scale bar?

      A scale bar was added.

      (b) Figure 1C, the standard error bars for igcvPAB (both with and without IPTG) are relatively wide, indicating high variability in the data. This suggests that the results in the igcvPAB group are not as consistent as the wild-type (wt) or ΔgcvPAB groups. Please show the original data and perform a statistical test (e.g., t-test or ANOVA).

      The original data have been added to Fig. 1B. t-test results (with Bonferroni-Holm correction) are now included for all samples.

      (c) Figure 2D, scale bar?

      These are sections of agar plates. From our point of view, a scale bar does not add relevant information.

      (d) Figure 3, please check the labels of the Figure 3 legend. (D) and (E) or A-B?

      Thanks, corrected.

      (e) Figure 5E, quantification of plaque areas?

      The plaque areas were quantified. A blot showing these quantitative data is now presented in Fig. 5F.

      Reviewer #2 (Recommendations for the authors):

      Line 58: There are published studies that indicate that syncytiotrophoblasts are actually resistant to Listeria infection and that it is extravillous trophoblasts that are likely to serve as entry points for Listeria into the placenta (see, for example, Lowe et al, Infect Immun. 2018 Volume 86 Issue 6 e00801-17).

      We thank the reviewer for this comment and have removed our statement claiming that syncytiothrophoblasts are the entry point as this is not relevant to the understanding of the work presented here.

      Reviewer #3 (Recommendations for the authors):

      (1) Please mention the number of times experiments were performed as independent biological replicates, wherever applicable.

      We added this information to the figure legends wherever this was necessary.

      (2) Please provide the details of the type of statistical analysis used for the various graphs, either in the figure legends or in the materials & methods section.

      This information was also added to the figure legends wherever it still was missing.

      (3) Can the authors comment on how the weight-loss phenotype of animals and the variation in the size of the spleen between animals infected with wild-type and mutants in Figure 2 can be explained without any significant changes in the CFU? Additionally, I did not see details regarding the number of animals used in the murine infection model experiments. Please mention this along with the type of statistical analyses used.

      We do not see a contradiction here, as the apparent differences are explained by the distinct time points at which CFU numbers (day 3 post-infection) and spleen size (day 9 postinfection) were measured. Starting from day 8, the difference in weight between animals infected with the wild-type strain and those infected with the ΔgcvPAB mutant becomes clear for the first time. At day 3, no significant differences are detected in either CFU numbers or weight. By day 9, when the weight difference has become apparent, differences in spleen size are also observed. To improve clarity, the time points at which each analysis was performed have been added to Fig. 2G and Fig. 2I. The number of infected animals and the type of statistical analysis used are now specified in the figure legend.

    1. eLife Assessment

      This is a detailed and well-designed simulation study of the utility of replication metrics in animal-to-human study translations in bridging the gap between laboratory discoveries and health practice, a critical consideration in turning laboratory scientific research findings into tangible, real-world applications, to directly help human health. The study approaches are convincing, and the findings are important, as they offer insights into clinical research translations to advance health decision-making.

    2. Reviewer #1 (Public review):

      [Editors' note: This revised version of your article has been assessed by the Reviewing Editor without further input from the original reviewers. The comments raised by the original reviewers in the earlier round of review have been addressed. The study findings are quite insightful and important, and the evidence is strong, convincing, and a substantial addition to the evidence base.]

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

    3. Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodated for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

    4. Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

      Weaknesses:

      (1) The conceptual distinction between replication and translation could be more clearly emphasized.

      (2) Interpretation of results is dense and can be challenging to follow without a clear and summarized.

      (3) Some simulation parameters (effect sizes, heterogeneity, and number of animal studies) require more substantial justification.

      (4) Practical recommendations could be more explicit to guide applied researchers.

      We thank Reviewer 1 for the general positive assessment of our study and for the constructive feedback. We have addressed all of the four identified weaknesses in the revised manuscript. Specifically,

      (1) Conceptual distinction between replication and translation. We have reinforced this distinction at multiple points in the manuscript: in Section 2.7 (just before introducing the translation success metrics), in the Discussion, and in a new working definition of translation success added to the Introduction. We further explicitly acknowledge that statistical translation success, as defined here, is narrower than biological translation.

      (2) The dense result section. We have added a summary Table (Table 3) at the end of the Results section that compares all metrics on key properties (strengths and weaknesses, overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch). We also direct readers to this table early in Section 3.2, so that readers less interested in the technical details can obtain the key take-home messages without reading the full section.

      (3) Further justification of simulation parameters. We have substantially extended the rationale for our parameter choices in Section 2.4 and the Limitations section. We explain that our parameters are grounded in an empirical meta-analytic dataset, contextualise the large effect size and heterogeneity value against published benchmarks from preclinical research, and clarify that our main goal was to explore directional trends rather than absolute performance under specific values. We have also added an invitation for others to explore alternative parameter spaces using our openly available code.

      (4) Practical recommendations. We have extended the Recommendations section (pages 21–22) with more explicit scenario-specific guidance, supported by the new summary table.

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodate for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

      We thank Reviewer 2 for the thoughtful summary and for recognising the strengths of our study. We believe that both weaknesses were addressed in the revised version of our manuscript. Specifically,

      (1) Heterogeneity in animal studies. We agree that high heterogeneity in animal studies is an important limitation, and we address it directly in our simulation design by including a wide range of heterogeneity values (including very high levels, as observed in animal studies). Our results show clearly how heterogeneity affects the performance of each metric, and we highlight this in both the new summary Table (Table 3) and the Recommendations section which was extended. We also caution applied researchers to pay special attention to heterogeneity when interpreting translation results.

      (2) Focus on the statistical component of translation. We fully agree that statistical translation success is only one aspect of a broader decision-making process. We have elaborated on this in the revised manuscript, both in a new working definition of translation success in the Introduction (which explicitly distinguishes statistical from biological translation) and in a new paragraph in the Discussion section where we situate our metrics within translational decision-making frameworks such as PATH. They make it clear that progression to human trials depends on a suite of evidence of which statistical translation is only one part.

      Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Weaknesses:

      The weaknesses of the study are the limited focus on specific metrics, the assumptions, particularly in the limited number of human study variables, and the less-than-ideal approachable summary of findings for a non-technical audience.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

      We thank Reviewer 3 for the positive assessment and for the constructive suggestions.

      We have addressed the identified weaknesses as follows:

      (1) The assumptions around human study variables. We acknowledge these as inherent constraints of the simulation design. We have added a note in the Limitations section about the fixed human sample size (N = 107 per group), clarifying that while this value is grounded in a power analysis as per regulatory standards, it represents one particular scenario and may not generalise to all contexts. Further, we have contextualised and motivated the other simulation parameters better. We also invite readers to explore alternative conditions using our openly available code.

      (2) Approachability of the summary of findings for a non-technical audience. We have added a summary Table (Table 3) at the end of the Results section, comparing the metrics on key properties including overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch. We direct readers to this table early in Section 3.2 so that those less interested in the technical details can obtain the main take-home messages without reading the full section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) Conceptual framing: clearer distinction between replication vs translation

      The Introduction correctly points out the conceptual difference between replication and translation (animal to human), but this distinction needs to be reinforced repeatedly, especially when interpreting metric performance. For instance, several metrics (e.g., meta-analysis, replication BF) inherently assume exchangeability of findings, which is rarely justified in translation because species differ biologically.

      The manuscript should explicitly state why treating animal findings as "original studies" and human findings as "replications" can be misleading. Add a subsection in the Discussion: Why replication metrics behave differently in translation settings. This will help provide a more straightforward interpretation of the results beyond the numerical findings.

      Thank you for your feedback. While the purpose of our study is to assess the applicability of the replication success metrics in the translation context, we agree that the reader should be reminded that these two concepts differ and metrics’ assumptions might not always hold. We have reiterated the difference between replication and translation in Section 2.7, just before we introduce the translation success metrics (see end of page 7). We also reiterate it in the Discussion (see end of page 20). Here, we emphasise that while some of the metrics assume that both studies investigate the same effect, this is unlikely to be the case in translation, leading to some of the metric’s assumptions being violated which impacts the performance of the metrics.

      (2) Stronger justification of the simulation parameters is needed

      The simulation factors are comprehensively presented (Table 1), but certain choices appear arbitrary or oversimplified.

      - Effect sizes: The three levels (0, −4.44, −24.37) are derived from the motivating dataset, but the paper should explain that these represent extremely large effects in many biomedical contexts.

      - Heterogeneity values: τ<sup>2</sup> = 291.1 is enormous; adding context about real-world heterogeneity distributions would help.

      - Number of animal studies (k): Using only 2-5 studies may not reflect reality; many preclinical fields have >30 studies before clinical translation.

      These choices should be more explicitly defended in Section 4 (Limitations), beyond the brief mention already there. Provide a sensitivity analysis, or explain why extrapolation beyond this parameter space is reasonable.

      We agree that the choice of the parameter values might sometimes appear arbitrary. However, instead of arbitrarily choosing parameter values, we base our choice on data from a meta-analysis. This particular meta-analysis might not be representative of all of pre-clinical and clinical research, but because Terstappen included both animal and human studies investigating the same research question it was particularly well suited. They further used an outcome (maternal blood pressure) that is comparable between rats and humans, which is quite rare. We have specified this further in Section 2.4 (Motivating dataset, page 5). In the Limitations section, we acknowledge any possibly unrealistic simulation conditions again, and emphasize that our main goal was to explore trends in the metrics’ behavior as the conditions changed rather than their absolute performance under specific values. Further, the effect sizes (0, −4.44, −24.37 mmHg) span a meaningful range on the unstandardized mean difference scale for blood pressure measurements: from no effect to a modest but clinically relevant reduction to a large effect typical of animal studies. The large heterogeneity value corresponds to a relative heterogeneity of I^2 of 95.33% in the animal meta-analysis. While this appears high, it is frequently observed in preclinical research: Hooijmans et al (2022) showed that 55% of animal study meta-analyses using mean differences as effect size measure have I^2>75%. We also added a footnote reiterating the fact that such high effect sizes (on the raw mean difference scale) are indeed common in animal studies (on page 6). Regarding k, we acknowledge that pooling only 2 to 5 animal studies may not reflect common practice. However, the directional trends in type 1 error and power are clearly visible in our Figures. Larger k decreases the type 1 error of the animal studies, while the power is increased unless there is high heterogeneity between animal studies and there is only a small effect. Extending the range further is unlikely to change the conclusions. Moreover, in practice, the decision to advance to human trials considers evidence well beyond the statistical considerations we simulate. All of the above is now emphasized more explicitly in both the methods, where we have substantially extended the reasoning for choosing the simulation conditions, and the limitations section. Finally, we added an invitation to others to use our open material (i.e., code) and explore the behaviour of the metrics under other conditions (see top of page 21).

      (3) Decision criteria (strict/lenient/no criterion) need a clearer rationale

      The three continuation rules are a strength of the study, but:

      - The lenient criterion (any negative estimate is considered "beneficial") is unrealistic and should be reframed.

      - The strict criterion (p < 0.025) heavily inflates effect sizes (in Figure 1b) and may distort interpretation.

      It would be helpful to provide a table showing, for each criterion, its real-world analogue (e.g., regulatory requirement, exploratory progression, mechanistic plausibility).

      We have followed your suggestion and added a Table (Table 2) with the description of the criterion and a description of its real-world analogue. No criterion represents an important reference scenario used to evaluate metric behaviour independent of progression decisions. The strict criterion is the closest to regulatory-style evidence. It is also highly selective and therefore might induce biases (e.g., inflated effect sizes). We link lenient to an exploratory decision-making where efficacy evidence is considered in addition to other factors (e.g., safety), but not intended to represent a certain regulatory standard.

      (4) Interpretation of simulation results needs more focus

      The Results section is extremely detailed, making it challenging to identify the central take-home messages. The authors should consider adding a concise summary table comparing metrics on key properties:

      - T1E control robustness.

      - Sensitivity to heterogeneity.

      - Dependence on animal sample size.

      - Dependence on k.

      - Bias under asymmetric effects.

      Moving some nested-loop plot descriptions to the Supplement. Right now, descriptions are technically correct but cognitively heavy.

      We agree with your comment and have attempted to implement it in our summary Table 3, at the end of the results section. We also point readers early on to the Table, so that they can skip the more technical and detailed description if they want (see first paragraph section 3.2, page 12). After some trial and error, we agreed that the chosen columns are the most useful for an applied researcher to get a quick overview. Our table now summarises for each metric its main strengths and weaknesses, its behaviour with increasing heterogeneity, its sensitivity to more animal data (i.e., larger k and larger animal sample size), and its behaviour under effect mismatch (i.e., when the true effect in the animal and human study are dissimilar).

      (5) The discussion should provide explicit recommendations.

      The authors provide high-level recommendations, but the recommendations lack specific guidance. When heterogeneity is low, controlled sceptical p-value works well. When effect sizes differ: weighted Edgington is stable. The authors should avoid using replication BF when the animal effect ≠ human effect. Meta-analysis should not be used when human heterogeneity is high, because of inflated T1E.

      We agree that explicit recommendations would be helpful to the applied researcher. As mentioned in the reply to the previous comment, we have added a summary table which lists the strengths and weaknesses of each metric. We also extended the paragraph in the Recommendations section (on page 21 and 22) to give some examples of scenarios in which certain metrics would be recommended over others.

      (6) Recommendations for applied researchers

      The study is missing an explicit definition of "translation success". The manuscript implicitly defines translation success as: "Both animal and human results show a beneficial treatment effect according to metric X". But this is different from biological translation, which concerns underlying mechanisms. The authors briefly mention this conceptual challenge, but this should be elaborated, as it is central to interpretation.

      Thank you for this comment. We agree that “translation success” was not explicitly defined. We have now added a working definition in the Introduction, clarifying that, in this paper, translation success is defined statistically, and depends on the metric. We now explicitly acknowledge that this is a narrower definition than biological translation. We also elaborate on this distinction in the Discussion where we note that the appropriate metric and interpretation of translation success depends on the translation goal and that statistical translation is distinct from biological translation.

      Minor points:

      (1) The abstract could include a direct sentence on the main conclusion. For example, no metric was uniformly optimal; controlled sceptical p-value and weighted Edgington performed most consistently.

      Our abstract already included main conclusions. We added the word “However” to emphasize the sentence “no metric was uniformly optimal” a bit more.

      (2) The figures are informative, but nested loop plots are very dense. Consider providing a guided example in the figure caption explaining how to read them (as partially done in Figure 1a, but repeat for all).

      We agree that the Figures can be very overwhelming at first. We did not want to add specific helping elements as we did in Figure 1 to not make the figures even busier. The goal was to introduce the reader gently to the nested loop plots via Figure 1 before having them look at the remaining figures. We hope that with the added summary Table and the more detailed recommendations, applied researchers less interested in the statistical details will still find the information most relevant for them easily.

      (3) Methods: Section 2.4 could clearly state that effect sizes are in units of mmHg (blood pressure) from the dataset.

      Thank you for pointing this out. This has been added.

      (4) Results: This section is long; consider adding a brief summary paragraph at the end of 3.2.

      We added a summary table, allowing interested readers to skip the long section entirely.

      (5) Limitations: Add a note about publication bias in animal studies (you mention it in the Introduction, but not in Limitations). Add a statement about effect direction consistency (i.e., animal effect negative but human positive), which is not explored in the simulation grid.

      Thank you for pointing out this inconsistency. A note about publication bias in animal studies was added to the Limitations section (that this was not investigated). A note about opposite animal and human effects was added to Section 2.5 (Simulation conditions) under “Animal and human effect sizes”.

      Reviewer #2 (Recommendations for the authors):

      Animal studies are typically highly controlled, using animal models that are either outbred to provide higher genetic variability or inbred with very little genetic variability and with a specific phenotype. Additionally, many rodent models are incomplete models of the overall human phenotype and are typically used to investigate only one aspect of the condition/disease. Some of the rat animal models that the Terstappen et al. (2020) systematic review used as the simulation parameters for the study included outbred (Sprague-Dawley, Wistar) and inbred Spontaneous Hypertensive Rats (SHR), which have different mechanisms in which hypertensive onset can occur, especially if inducing preeclampsia in outbred animals. Is it feasible to reduce heterogeneity in the animal results if only outbred or only SHR are considered instead? I realize this may reduce the sample size even further.

      You raise an important point differentiating biological (rather than statistical) translation. We have added a sentence about differences between rat models and humans to the new paragraph in the Limitations section (bottom page 20 and top page 21) on the distinction between biological and statistical translation. As for reducing heterogeneity in the animal results by focusing on one type of rats, we agree focusing on one type of rats might reduce heterogeneity. We however consider this reduction to be very small (because the results of the study with SHR are actually comparable to the results with Wistar and SD rats). Therefore, rerunning the simulation would not yield results that differ in any meaningful way from those already reported and the substantial computational effort required to do so is not warranted.

      Reviewer #3 (Recommendations for the authors):

      Overall, I found this a very detailed study. However, my recommendation is to provide a more approachable overview of the results to reach a wider audience. Currently, the article is much more technical and statistically focused. I think two additions could help.

      (1) A summary table of each of the metrics and their strengths and weaknesses under the various conditions (e.g., animal and human study characteristics). Currently, this is done via text, but I think a high-level summary via a table could be a compelling way to make the simulations more approachable for a non-technical audience.

      As requested also by reviewer 1, we have added a summary table.

      (2) Contextualize the findings within the decision-making process a little more. The authors have a well-written limitations section that acknowledges this; however, I think the discussion (and maybe the introduction) could be enriched by putting the simulation findings into context. For example, this paper suggests a framework that includes replication as part of the decision-making process for human trials (https://www.cell.com/med/fulltext/S2666-6340(24)00296-4).

      We agree that situating our metrics within existing translational decision-making frameworks adds important context. We have added a paragraph in the Discussion (before the Limitations section on page 21) clarifying that the metrics evaluated here should not be viewed as standalone decision rules for progression from animal studies to human trials. Several frameworks have recently emerged precisely to guide such decisions in a more structured, multidimensional way. We refer to PATH and also to the GALENOS approach [DOI: 10.1186/s12874-026-02891-4]. Within such frameworks, translation success metrics of the kind evaluated here may provide a quantitative assessment of the consistency between animal and human efficacy findings, thereby informing one component of a broader translational evidence assessment. We have also briefly mentioned at the end of the Introduction (page 4) that frameworks for structuring the use of preclinical evidence in translational decisions are being developed, further motivating the need for quantitative tools such as those evaluated here.

      Below are some additional minor comments for the authors to consider:

      (1) In the abstract (4th line), there is an extra 'l' in failure.

      Thank you for the detailed review. We have fixed this.

      (2) I think since the study is completed, the objectives in the introduction should be past tense, not future.

      We have fixed this.

      (3) The limitations section should include the fixed human sample size. N=107 per group is grounded in the literature, but this varies widely based on the effect size of interest. Again, not material to the point of translation under simulated conditions (of which this would have increased the simulations well above the 648 already included), but given the impact this has on insights, this limits this investigation to a degree and should be acknowledged.

      Thank you for your comment. We have added a note about the human sample size to the paragraph about the simulation conditions in the Limitations section. The human sample size was computed via power analysis as per regulations, but we realize this could change depending on the effect size.

      (4) I appreciate how shrinkage was calculated. Though it is worth noting that the Reproducibility Project: Cancer Biology found much higher rates, which are similar to reports from biotech and pharma (e.g., 11% and 20-25% for Amgen and Bayer).

      We already mentioned the high rates of shrinkage in the Replication Project Cancer Biology (see page 10). We have now also emphasised that one could adapt these levels further depending on the situation.

    1. eLife Assessment

      This fundamental study uses simultaneous EEG and fMRI recordings to shed light on the relationship between alpha and gamma oscillations and specific cortical layers. The sophisticated methodology provides compelling evidence for correlations between oscillatory power and the strength and contents of fMRI signals in different cortical layers. This paper will be of interest to neuroscientists studying the role and mechanisms of alpha and gamma oscillations.

    2. Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two a higher frequency within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state of the art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts is in line with recent findings and theories. In sum, this paper makes a very nice contribution to the literature. In particular, it provides a novel characterization of how oscillatory signals orchestrate the coding of visual contents in the visual cortex and provides a starting point for further research examining how this oscillatory coding changes across visual contents and tasks.

    3. Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. The authors defined both feature-specific and feature-unspecific contrasts, further subdivided by EEG power regressors, to examine how orientation information is signalled across cortical layers. Feature specific contrast was established via comparing trials where stimulus orientation (respectively) was left with those where the stimulus orientation was right. Further specifying it depending on EEG power regressors as congruent where stimulus orientation of EEG regressor matches voxel preference or incongruent (stimulus orientation of EEG regressor does not match voxel preference).

      Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence. On the other hand, the lower-frequency findings are compelling, the functional specificity observed within the alpha frequency band is a particularly interesting result and further highlights the importance of distinguishing feature specificity in order to reveal more nuanced characteristics of neuroimaging data.

      Overall, main findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. The paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings.

      The evidence for gamma involvement in the observed effects is selective: no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 5C,F), with significant effects emerging only in positively responding voxels and only for the contrast between congruent and incongruent conditions in the feature-specific BOLD response. The authors address this in the Discussion, noting that the stimulus may have elicited a weaker gamma response overall, and the contrast of EEG congruent vs. incongruent is necessary in order to achieve the largest contrast-to-noise ratio.

      Authors reported negative relationship between the alpha frequency band and the feature specific BOLD signal increases (for congruent condition, Figure 5A,D). Furthermore, testing for the functional specificity between lower vs. upper alpha (Figure 5B,E), the authors included statistical test on the mixed effects model coefficients, and found significant interaction between alpha frequencies in the congruent condition. This interaction was mainly driven by the upper alpha band (which was later confirmed with simple effects analysis) and revealed stronger negative relationship of upper alpha and the BOLD signal for the subtraction of congruent over incongruent conditions. These are exciting results, which further advocate for a more active role of upper alpha band involvement (relative to lower alpha band) in processing visual features.

      Expanding on this, the authors have also conducted an exploratory analysis of the relationship between the behavioural findings and underlying neural activity for non-oddball trials (Figure S12 in Supplementary Figures). This confirmed a positive relationship between task performance and alpha frequency, suggesting that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power.

      This study provides a valuable and exciting contribution to the literature on oscillatory dynamics and laminar fMRI.

      Comments on revised version.

      Thank you for the thorough revision and for addressing the comments so carefully. The new figures are super beautiful and make the results considerably easier to interpret, they are a real improvement to the paper.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two higher frequencies within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state-of-the-art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts are in line with recent findings and theories. In sum, this paper makes a very nice contribution. I still have a few comments that I outline below, regarding the data visualization, some methodological aspects, and a couple of theoretical points.

      The authors put a lot of effort into the figure design. For instance, I really like Figure 1, which conveys a lot of information in a nice way. Figures 3 and 4, however, seem over engineered, and it takes a lot of time to distill the contents from them. The fact that they have a supplementary figure explaining the composition of these figures already indicates that the authors realized this is not particularly intuitive. First of all, the ordering of the conditions is not really intuitive. Second, the indication of significance through saturation does not really work; I have a hard time discerning the more and less saturated colors. And finally, the white dots do not really help either. I don't fully understand why they are placed where they are placed (e.g., in Figure 3). My suggestion would be to get rid of one of the factors (I think the voxel selection threshold could go: the authors could run with one of the stricter ones, and the rest could go into the supplement?) and then turn this into a few line plots. That would be so much easier to digest.

      We thank the reviewer for their insightful comments. Below we will address each point separately and highlight the changes made to the manuscript. In agreement with the reviewer we have recompiled Figures 4 and 5 (previously Figures 3 and 4). The new figures only present results for the 10% voxel selection threshold (with 5% and 25% moved to Supplementary Figures, see Figures S1-S9). Instead of the radially arranged layout, we opted for a more traditional figure layout, which significantly improved readability.

      (2) The division between high- and low-frequency alpha in the feature-specific signal correspondence is very interesting. I am wondering whether there is an opposite effect in the feature-unspecific signal correspondence. Would the high-frequency alpha show less of a feature-unspecific correlation with the BOLD?

      Following the reviewer’s interesting suggestion, we added the low/high frequency alpha analysis to the feature-unspecific analysis. Indeed, we have found a significant interaction between the sign of the signal change for selected voxel (positive vs. negative BOLD) and alpha sub-band (low vs high frequency alpha). An analysis of simple effects did not reveal any significant effects, however we found a trend level difference (p=0.097) between low and high-frequency alpha for the positive voxel sub-selection. This indicates a stronger negative relationship between upper alpha and the positive BOLD signal as compared to lower alpha. We interpret this result as partial evidence for a feature-related contribution of the upper alpha band. “Active” cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. The negative relationship between alpha and negative BOLD could thus be interpreted as an indirect effect, resulting from a reduced alpha decrease in cortical patches responding to non-attended receptive field locations. However, the involvement of attention-related processes remains speculative, since attention was not explicitly manipulated as part of the experiment.

      We have added Figure 4 B.

      We have also added this section to the Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub><0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      And the following section of the Discussion was extended:

      “The significant interaction between the sign of the BOLD signal deflection and upper or lower α bands (see Figure 4 B) further indicates that multiple α-related processes contribute differentially to positive or negative BOLD. "Active" cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. This hypothesis receives additional support from the trend-level difference in α sub-bands for positive BOLD, indicating that lower α is less related to the active, possibly feature-related processes. The absence of this difference for negative BOLD again indicates a broader, more general process. Future experiments manipulating visual features and attention might reveal a differential upper and lower α response to attended visual features and a more general relationship between α (and possibly superficial layer cortical activity) for suppressed (unattended) receptive fields.”

      (3) In the discussion (line 330 onwards), the authors mention that low-frequency alpha is predominantly related to superficial layers, referencing Figure 4A. I have a hard time appreciating this pattern there. Can the authors provide some more information on where to look?

      We thank the reviewer for pointing out the lack of clarity of this section in the Discussion. We have now rephrased the Discussion, focusing more on the laminar difference and keeping the frequency difference to a separate paragraph. Our main argument for possibly multiple alpha-related processes are twofold: a difference in alpha frequency depending on the underlying analysis (low vs high frequency alpha) and a different layer distribution (superficial layers vs. superficial and deep layers, depending on the analysis). The respective section in the Discussion now focuses on the laminar difference only. We find a negative relationship between alpha and the BOLD signal most prominently in superficial layers (feature-unspecific contrast for the BOLD signal with negative t-values; Figure 4). In addition, we find a superficial and deep layer contribution for the feature-specific contrast (congruent - incongruent; Figure 5A). While the here presented experiment was set out to investigate feature-specific processes, the meaning of the feature-unspecific results are of speculative nature. Future experiments should target the laminar difference between feature-specific and unspecific processes with respect to alpha frequency and layer distribution directly. 

      We have modified the respective sections in the Discussion:

      “Furthermore, we observed that the relationship between the feature-specific BOLD signal and α is predominantly linked to frequencies above 11 Hz (see Figure 5A). An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies. Since individual frequency variations have been included as a random slope in the linear mixed-effects model, these effects cannot be explained by a subset of participants driving lower or upper α separately. Specifically our findings on upper α indicate that α is not exclusively linked to global signal modulations, which has been the traditional perspective [...]”

      “Not only did we find a dissociation in the frequency domain between the relationship of α and the BOLD signal, but furthermore found that the laminar activation patterns provide further evidence for potentially multiple α-related processes. The association between α and the BOLD signal was strongest in superficial layers for negative BOLD activity and feature-specific activity (see Figure 4A and 5A). However, deep layer-related α effects were limited to feature-specific processes only (see Figure 5 A Co-Inco). These findings suggest that superficial layer α reflects are broader, more general process, while deep layer α operates more narrowly, linked to the processing of the visual features themselves. Previous findings using laminar fMRI (which did not include the investigation of oscillatory activity), indicate that superficial layer activity might be more related to the modulation of attention [...]”

      (4) How did the authors deal with the signal-to-noise ratio (SNR) across layers, where the presence of larger drain veins typically increases BOLD (and thereby SNR) in superficial layers? This may explain the pattern of feature-unspecific effects in the alpha (Figure 3). Can the authors perform some type of SNR estimate (e.g., split-half reliability of voxel activations or similar) across layers to check whether SNR plays a role in this general pattern?

      We agree with the reviewer that the vascular draining effect typically leads to increased signal change in superficial layers, the effect on (t)SNR however might be less straightforward. We did not include any counteracting measures, because we were not interested in the amplitude of the signal change, but now include an estimate of tSNR (See Figure S10 in Supplementary Figures). We found that in fact the signal-to-noise ratio is higher in deep layers. Most importantly however, the tSNR layer profiles we identified do not reflect the correlation layer result patterns of the combined EEG-fMRI analysis. This indicates that our results are most likely not the result of tSNR differences. In order to confirm our tSNR pattern we have also conducted a second layer analysis based on the LAYNII toolbox, which assigns voxels between pial and white matter to distinct layers (as compared to our fraction-based approach) and found a similar profile as with our initial analysis. However, absolute tSNR values were found to be higher for our weighted layer analysis. We speculate that while functionally relevant components of the BOLD signal drain towards superficial layers, physiological noise components will drain towards superficial layers as well.

      It is furthermore worth pointing out that for the contrast (congruent - incongruent), the vascular draining effect would cancel out between the conditions. Our findings on superficial and deep layers for those contrasts can hence not be explained by vascular draining at all.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      (5) The GLM used for modelling the fMRI data included lots of regressors, and the scanning was intermittent. How much data was available in the end for sensibly estimating the baseline? This was not really clear to me from the methods (or I might have missed it). This seems relevant here, as the sign of the beta estimates plays a major role in interpreting the results here.

      This is a very important remark and we would like to apologise for the confusion. It was not clear in the manuscript that the GLM was computed on z-transformed fMRI data. We have not specifically collected any “baseline volumes”. A positive beta value would indicate that the sign of the predictor matches the sign of the BOLD signal deflection (and vice versa).

      We have added or modified the following sections in Results and Methods respectively:

      “Before the GLM was computed, the fMRI data was z-transformed across time, separately for each block and voxel.”

      “A general linear model (GLM) has been computed with predictors for each TF bin separately for all voxels in V1 that later have been sub-selected according to the respective condition. Time courses for each voxel have been z-transformed before the GLM was computed for each voxel and experimental block separately. Afterwards, each of the resulting regression coefficients (β values) were multiplied with the voxel-specific layer weights that have been obtained as described above.”

      (6) Some recent research suggests that gamma activity, much in contrast to the prevailing view of the mechanism for feedforward information propagation, relates to the feedback process (e.g., Vinck et al., 2025, TiCS). This view kind of fits with the localization of gamma to the deep layer here?

      (7) Another recent review (Stecher et al., 2025, TiNS) discusses feature-specific codes in visual alpha rhythms quite a bit, and it might be worth discussing how your results align with the results reported there.

      We would like to thank the reviewer for pointing out these papers. Yes, we believe that those could be very related to the effects reported here. At the time of writing the initial manuscript we were not aware of the mentioned publications. 

      We have now included these papers in the Discussion:

      “Recent publications on the information exchange within and between primary visual cortex areas of macaques also reported deep layer γ band activity depending on the stimulus material (Gieselmann et al., 2022; Ferro et al., 2021). Those publications challenge the feed-forward exclusivity of γ altogether by revealing intra-area feedback communication in V1 from layer 5 to layer 6 and layer 6 to supra-granular layers. Possibly, the relationship between γ and deep layer BOLD we observed is also related to similar processes (Vinck et al., 2025).”

      “Similarly, in a recent opinion article, Stecher et al. (2025) promote the idea of "content-aware" α-oscillations. In agreement with our results, the authors argue that α-oscillations are related to content-specific feedback signals, reflected in increased decoding performance based on α power of top-down related processes, even prior to the onset of the stimulus (Hetenyi et al., 2025).. Accordingly, we interpret the lower α effect [...]”

      Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence, but most seem more like a suggestive trend. Nonetheless, the paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings. Further investigation will be important to strengthen and expand upon these insights.

      One of the main strengths of the paper lies in the use of a well-established and straightforward experimental paradigm (the visual oddball task). As a result, the behavioural effects reported were largely expected and reassuring to see replicated. The acquisition technique used is very novel, and while this may introduce challenges for data analysis, the authors appear to have addressed these appropriately.

      Later findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. However, the layer weight calculation is lacking in the manuscript. While it is discussed in the methods, it would help to briefly explain in the results how these weights are calculated, so that the reader can better follow what is being interpreted.

      Line 104 states there is one virtual channel per hemisphere for low and high frequencies. It may be helpful to include the number of channels (n=4) in the results section, as specified in the methods. Also, this raises the question of whether a single virtual channel (i.e., voxel) provides sufficient information for reproducibility.

      We thank the reviewer for encouraging us to clarify the virtual channel selection and we agree that the current description could be misleading. Indeed, we selected 4 virtual channels in total: 1 for each frequency band (alpha/gamma), for each hemisphere separately. The main goal of this selection was to find the clearest response of that frequency band to the task. Previous publications used a supervised (ICA-based) approach to extract those responses. To increase reproducibility, we have chosen an unsupervised beamformer-based approach. The reconstruction of time or frequency-resolved sources in the brain typically yields spatially highly correlated results. Publications focusing on this type of analyses report a spatial extent of typically multiple centimetres, which here is the case as well (see Figure 3A of the updated manuscript). As such, the single voxel selection boils down to selecting the peak response within a large patch of very similarly responding voxels. Using this approach we were able to select the frequency response with the highest possible SNR. We do not however claim that the respective single voxel is exclusively carrying this information. In addition we have added a short explanation to the Discussion, since we believe that virtual channel selection with a different objective (e.g. maximising the difference between conditions or maximising cross-frequency coupling, etc.) could indeed profoundly impact the EEG-fMRI correlation, which would open up opportunities for interesting analyses that are however beyond the scope of this project.

      We have added the following section to the Discussion:

      “Future work might also vary the exact virtual channel selection for obtaining EEG-based regressors. Here, we focused on the grid points (voxel locations) with the strongest α or γ response for each frequency band in each hemisphere, derived from the average frequency response to maximise SNR. However, selecting the respective virtual channels based on the response to specific stimulus features or the interaction between high and low frequency bands are possibilities worth exploring in future work.”

      One area that would benefit from further clarification is the interpretation of gamma oscillations. The evidence for gamma involvement in the observed effects appears somewhat limited. For example, no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 2). Significant effects emerged only when the analysis was restricted to positively responding voxels, and even then, only for the contrast between EEG-coherent and EEG-incoherent conditions in the feature-specific BOLD response. It remains unclear how to interpret this selective emergence of gamma-related effects. Given previous literature linking gamma to feedforward processing, one might expect more robust involvement in broader, feature-unspecific contrasts. The current discussion presents the gamma-related findings with some confidence, and the manuscript would benefit from a more nuanced reflection on why these effects may not have appeared more broadly. The explanation provided in line 230, that restricting the analysis to positively responding voxels may have increased the SNR, is reasonable, but it may not fully account for the absence of gamma effects in V1's feature-unspecific response. Including the actual beta values from Figure 4 in the legend or main text would also help readers better assess the strength and specificity of the reported effects.

      We agree with the reviewer that the missing gamma-band response for the feature-unspecific signal, as well as the limitation of the effect solely to the feature-specific contrast for positive voxel selections only was unexpected. In fact, based on previous literature, we were expecting a feature-unspecific effect in the gamma band as well. However, the literature on laminar level EEG-fMRI is sparse and previous experiments used tasks that did not allow for the separation into distinct features (here left or right-oriented gratings). While we cannot fully explain the absence of the gamma effect for the feature-unspecific condition, we reasoned that our stimuli evoked weaker gamma band responses compared to previous literature. 

      The fact that we only see a significant gamma band response for the contrast for positive voxel selections can be interpreted twofold: First, previous experiments limit their analyses to positive BOLD responses only, for which we find an effect as well. Second, the fact that a significant effect could only be obtained for the contrast, might indicate that gamma band activity is related to the actual features themselves. A cortical column responding to left-oriented gratings would then be related to a gamma band response linked to that orientation. If this response to a single orientation could not be fully captured due to SNR-related issues, we would not see this effect in the congruent-only condition and also not in the feature-unspecific condition (because this boils down to both congruent conditions combined). If gamma-band oscillations are actually reflecting the response of a column to a certain orientation, then the lowest possible response would be found for the exact orthogonal orientation (here the incongruent condition). The contrast between most preferred and most not-preferred orientation might have helped to overcome the inherently low SNR, explaining the results for the contrast.

      Lastly, we did not include actual beta values in the main text, because those might be misleading. We compute the relationship between EEG power and the BOLD signal for every voxel separately, then weighted the result with the respective layer weight and lastly aggregated across voxels.This means that the beta values express the strength of the association between EEG and fMRI for an average voxel. For this reason the values are tiny and the values themselves are less meaningful than “typical” beta values.

      We have added or modified the following sections in the Discussion or Methods respectively:

      “Based on previous literature, we expected a γ band effect for the congruent condition of the feature-specific analysis (Scheeringa et al., 2016), which we did not observe. A possible explanation could be the used stimulus material in our experiment as compared to Scheeringa et al., (2016). Muthukumaraswamy et al., (2013) found that stationary gratings evoke a weaker γ band response as compared to moving annular stimuli that have been used by Scheeringa and colleagues. If γ is related to the processing of the actual features themselves (e.g. to a column preferably responding to left-oriented gratings), then contrasting congruent and incongruent voxel selections provides the largest possible contrast-to-noise ratio (CNR). In turn annular stimuli as previously used might have activated all possible orientations and thus might have greatly boosted γ SNR.”

      “The described procedure of computing a GLM based on z-transformed data using z-transformed predictors yields β-coefficients that reflect the average relationship of a single voxel's BOLD response for a given layer (fraction of the single voxel's β) with EEG power changes of a specified frequency.”

      Relating to behavioural findings for underlying neural activity, could the authors test on a trial-by-trial basis how behavioural performance relates to the BOLD signal / oscillatory activity change? Line 305 states that "Since behavioural performance in the present study was consistently high at 94% on average and participants were instructed to respond quickly to potential oddball stimuli, a higher alpha frequency might reflect a more successful stimulus encoding and hence faster and more accurate behavioural performance." Also, this might help to relate the findings to the lower vs upper alpha functionality difference.

      This is a very interesting suggestion. We now include an exploratory analysis of the relationship between frequency and behavioural performance in the Supplementary Figures (see Figure S12). We did not perform a correlation between behavioural performance and alpha over trials because of the low numbers of oddball trials (N=40) and very limited number of false responses (94% response accuracy on average). However, we computed a correlation across participants. After averaging the alpha time-frequency spectrum across non-oddball trials, the individual alpha frequency was determined by the frequency where the alpha decrease (between 0.1 and 0.8 s post-stimulus) was largest. The correlation between alpha frequency and either reaction times and d’ (as a measure for accuracy), yields a significantly positive relationship between d’ and alpha frequency. This indicates that alpha frequency is related to task performance. We interpret those exploratory findings such that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power. 

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      In Figure 4, the EEG alpha specificity plot shows relatively large error bars, and there is visible overlap between the lower and upper alpha in both congruent and incongruent conditions. While upper alpha shows a positive slope across conditions and lower alpha remains flat, the interaction appears to be driven by the change from congruent to incongruent in upper alpha. It is worth clarifying whether the simple effects (e.g., lower vs upper within each condition) were tested, given the visual similarity at the incongruent condition. Overall, the significant interaction (p < 0.001, FDR-corrected) is consistent with diverging trends, but a breakdown of simple effects would help interpret the result more clearly. Was there a significant difference between lower and upper alpha in congruent or incongruent conditions?

      We thank the reviewer for this important remark and have added a simple effects analysis (see Figures 4 b and 5 b, e). We found that the main driver for the interaction between congruence condition and alpha frequency is upper alpha. Specifically the negative relationship between upper alpha and the BOLD signal is significantly stronger for the congruent condition and weaker for the incongruent condition. This indicates the upper alpha indeed is related to the processing of visual features.

      We have added a simple effects analysis (See Figures 4 and 5).

      We have added or modified the following in Results, Discussion and Methods respectively:

      In Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub> < 0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      “After correcting for multiple comparisons, we found a significant interaction (p<sub>FDR</sub> < 0.001). This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis. We found a significantly stronger negative relationship of upper α and the BOLD signal for congruent selections (p<sub>FDR</sub> < 0.01) and the reverse for the incongruent condition (p<sub>FDR</sub> < 0.01), as well as a significantly stronger negative relationship within the upper α sub-band for congruent over incongruent voxel selections (p<sub>FDR</sub> < 0.01).”

      “This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis, which revealed a significantly stronger negative relationship of upper α and the BOLD signal for congruent over incongruent selections (p<sub>FDR</sub> < 0.001).”

      In Discussion:

      “An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies.”

      In Methods:

      “Significant interactions were decomposed into simple effects using Wald tests on the model coefficients, ensuring that post-hoc comparisons were derived from the same statistical global variance as the primary interaction.”

      Overall, this study provides a valuable contribution to the literature on oscillatory dynamics and laminar fMRI, though some interpretations would benefit from further clarification or qualification.

      Reviewer #3 (Public review):

      Summary:

      Clausner et al. investigate the relationship between cortical oscillations in the alpha and gamma bands and the feature-specific and feature-unspecific BOLD signals across cortical layers. Using a well-designed stimulus and GLM, they show a method by which different BOLD signals can be differentiated and investigated alongside multiple cortical oscillatory frequencies. In addition to the previously reported positive relationship between gamma and BOLD signals in superficial layers, they show a relationship between gamma and feature-specific BOLD in the deeper layers. Alpha-band power is shown to have a negative relationship with the negative BOLD response for both feature-specific and feature-unspecific contrasts. When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency, and can therefore be interpreted as more feature-specific than lower frequency alpha.

      Strengths:

      The use of interleaved EEG-fMRI has provided a rich dataset that can be used to evaluate the relationship of cortical layer BOLD signals with multiple EEG frequencies. The EEG data were of sufficient quality to see the modulation of both alpha-band and gamma-band oscillations in the group mean VE-channel TFS. The good EEG data quality is backed up with a highly technical analysis pipeline that ultimately enables the interpretation of the cortical layer relationship of the BOLD signal with a range of frequencies in the alpha and gamma bands. The stimulus design allowed for the generation of multiple contrasts for the BOLD signal and the alpha/gamma oscillations in the GLM analysis. Feature-specific and unspecific BOLD contrasts are used with congruently or incongruently selected EEG power regressors to delineate between local and global alpha modulations. A transparent approach is used for the selection of voxels contributing to the final layer profiles, for which statistical analysis is comprehensive but uses an alternative statistical test, which I have not seen in previous layer-fMRI literature.

      A significant negative relationship between alpha-band power and the BOLD signal was seen in congruently (EEGco) selected voxels (predominantly in superficial layers) and in feature-contrast (EEGco-inco) selected (superficial and deep layers). When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency than lower frequency alpha. This is interpreted as a frequency dissociation in the alpha-BOLD relationship, with upper frequency alpha being feature-specific and lower frequency alpha corresponding to general modulation. These results are a valuable addition to the current literature and improve our current understanding of the role of cortical alpha oscillations.

      There is not much work in the literature on the relationship between alpha power and the negative BOLD response (NBR), so the data provided here are particularly valuable. The negative relationship between the NBR and alpha power shown here suggests that there is a reduction in alpha power, linked to locally reduced BOLD activity, which is in line with the previously hypothesized inhibitory nature of alpha.

      Weaknesses:

      It is not entirely clear how the draining vein effect seen in GE-BOLD layer-fMRI data has been accounted for in the analysis. For the contrast of congruent-incongruent, it is assumed that the underlying draining effect will be the same for both conditions, and so should be cancelled out. However, for the other contrasts, it is unclear how the final layer profiles aren't confounded by the bias in BOLD signal towards the superficial layers. Many of the profiles in Figure 3 and Figure 4A show an increased negative correlation between alpha power and the BOLD signal towards the superficial layers.

      We thank the reviewer for this important remark. Reviewer 1 raised a similar concern and I would like to refer you to our response to Reviewer 1, point 4. The veinal draining typically results in a higher signal change closer to the cortical surface. We did not take any measures to counteract this effect, but provide an analysis of tSNR in Supplementary Figures (see Figure S10). Possibly due to the drainage of physiological noise towards the surface, we found the highest tSNR in deep, followed by middle and superficial layers. To verify those results we computed the same analysis using a second layering algorithm, which resulted in the same profile, but overall less tSNR. Crucially the tSNR profile is not reflected in our EEG-fMRI results.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      When investigating if high alpha (8-10 Hz) and low alpha (11-13 Hz) are two different sources of alpha, it would be beneficial to show if this effect is only seen at the group level or can be seen in any single subjects. Inter-subject variability in peak alpha power could result in some subjects having a single low alpha peak and some a single high alpha peak rather than two peaks from different sources.

      We agree with the reviewer that a bias in a subset of participants to generally higher or lower alpha frequencies could potentially skew the presented results. While the initially computed model included a random intercept for the frequencies, we have now added the random slope as well. This ensures that the difference between low and high frequency alpha is indeed only driven by the difference in condition and not the result of individual differences across conditions themselves.

      In order to verify that not a small subset of participants is driving the result pattern, we also computed the fraction of participants that either show the dual alpha pattern (i.e. follow the exact pattern of the group average), contribute to the group average with a single peak or contradict the pattern entirely. Thereby, 40.4% of all participants show a dual alpha pattern, 38.4% a single alpha pattern in the direction of the group average and 21.2% contradict the group average. See Author response image 1:

      Author response image 1.

      Alpha Response Patterns with Example Subjects: V1 Feature Specific Contrast

      We would also like to highlight our added exploratory analysis of the relationship between alpha frequency and behavioural performance, which was requested by Reviewer 2, point 3. We find a significant positive correlation between alpha frequency and task performance on a group level. This indicates that higher alpha frequencies might be related to better discrimination of visual features. We speculate that participants with better task performance are capable of modulating their upper alpha more than participants with worse performance.

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      The figure layout used to present the main findings throughout is an innovative way to present so much information, but it is difficult to decipher the main findings described in the text. The readability would be improved if the example (Appendix 0 - Figure 1) in the supplementary material is included as a second panel inside Figure 3, or, if this is not possible, the example (Appendix 0 - Figure 1) should be clearly referred to in the figure caption. 

      Since Reviewer 1 suggested using an entirely different figure layout, we now opted to remove some information from the main text figures (we only show the 10% threshold, but 5% and 25% is in Supplementary Figures) and chose a more common figure layout. See Figures 4 and 5.

      Recommendations for authors:

      Reviewer #2 (Recommendations for the authors):

      The contrasts used in the analysis are not clearly introduced in the main text. While the methods section explains them more thoroughly, some of this explanation would be better placed in the results section, where the contrasts are first used. Specifically, the concepts of "feature-specific" vs. "feature-unspecific" BOLD signals are introduced with a very brief definition, which could be confusing for readers. The same applies to the terms EEG co and EEG inco; it would help to briefly explain these when they are first mentioned in the results. The supplementary figures and legends are helpful, so it is clear that the authors were prioritising clarity overall.

      The respective analyses are now also explained in the Results section:

      “During each trial either a left or a right-oriented grating was presented, from which two types of analyses have been derived: feature-unspecific BOLD activation (i.e. the response to any stimulus orientation), and feature-specific BOLD activation (i.e. the response to a specific stimulus orientation or the contrast between them). Thereby, fMRI data and EEG-based regressors could either be combined congruently (Co) by combining the BOLD signal of orientation-selective voxels with EEG-based regressors built from the same orientation trials, or incongruently (Inco), by combining the orientation-specific BOLD signal with EEG-based regressors built from the other orientation trials. Finally, those two congruency conditions have been contrasted (Co-Inco).”

      Figures are overall clear and illustrative of the results. For Figure 4, however, the use of dotted elements makes it somewhat harder to interpret what's being shown. While the supplementary figure clarifies the findings, rephrasing the figure legend to explain what the dotted lines represent would be helpful.

      Figures 4 and 5 have been replaced with a new layout and legends have been improved.

      The reported ranges overlap (e.g., alpha: 2-32 Hz; gamma: 20-120 Hz). It would be helpful to explain why such overlapping bands were chosen.

      Both frequency bands of interest differ slightly in their later time-frequency analysis (i.e. number of tapers and filter type). The overlap itself is not meaningful per se and results from the selection of a wide band for each respective sub-band. This wide selection was chosen to avoid filter artefacts. For the alpha sub-band, we also wanted to ensure that the beta spectrum is covered which also includes the alpha harmonic and for the gamma band that the full range of high-frequency activity is captured (e.g. EMG activity).

      Only a single time point was used for baseline correction of the low alpha band. Is this typical? The authors note that due to the gradient artefact arising in the pre-stimulus period, the baseline correction is somewhat difficult, although further clarification would be useful here.

      Relatedly, was pilot scanning conducted? If so, was the presence of strong gradient artefacts unexpected? More details about this would strengthen the methodological transparency.

      Indeed only a single time bin was used as the baseline for the alpha sub-band. After the piloting phase a slight adjustment to the final fMRI sequence has been made which was not expected to introduce gradient artefacts so close to the onset of the stimulus. Unexpectedly, those artefacts were visible until 300 ms before the onset of the stimulus. Similarly, a pre-stimulus alpha was observed (starting 250 ms before the onset of the stimulus), which we also aimed to exclude from the baseline period. In the end only the time bin centered at 300 ms prior to stimulus onset was chosen. However, this time bin contains 400 ms of data (the width of the window for the time frequency analysis). Thus, the term time point was misleading, because the actual time window that made up the baseline is 500 ms to 100 ms prior to the onset of the stimulus. 

      We have adjusted our wording in Methods to make this more clear:

      “For this reason, the low frequency baseline period comprised only a single 400 ms time bin centred around -0.3 s, because a pre-stimulus α decrease was expected starting around 0.25 s prior to stimulus onset.”

      Including a one-sentence explanation of the AROS test in the main text for clarity. As line 796 in the methods: "Each significant cluster has been further processed by means of an auto-regressive rank order similarity (aros) test (Clausner and Gentili, 2022). The fundamental idea behind the AROS test is whether group averages (i.e. averages of the signal of cortical layer in the present case), can be ranked such that the rank order is explained significantly better by the data than it would if the average data could not be meaningfully sorted (i.e. is shuffled)."

      An explanation has been added to the Results section:

      “Each significant cluster was then averaged along the frequency dimension at the widest point to enable an auto-regressive rank order similarity (aros) test Clausner & Gentili (2022), testing the laminar activation profile. The aros test transforms the layer averages into a rank order and tests - using a permutation procedure - if the rank order of the layer averages explains the data better than a random rank order (shuffled layer labels) would.”

      Line 223: "In fact, an analysis of the relationship between the EEG signal and the BOLD signal that focused on the feature contrast only (L - R; independent of the comparison to baseline) revealed a trend-level result with an even stronger deep layer contribution as compared to superficial layers." Could you point to which figure represents this finding - Figure 4B?

      This refers to Figure 4A in the old manuscript, for the 25% threshold for the gamma band. Since now the new figures do not include the 25% threshold anymore, it refers to Figure S4i.

      The number of participants is missing from the main text. Including this in the results section would improve clarity.

      The description of our sample has been moved from Methods to Results.

      Given the complexity of the data acquisition and analysis, the well-designed and easy-to-follow analysis pipeline figure (currently in the supplement) would be better placed in the main text.

      The mentioned Figure has been moved to the main text (now Figure 2).

      Also, simply out of curiosity, what do the authors think about the theta blob around 200ms post-stimulus?

      The theta blob most likely reflects the post-stimulus ERP as often observed in response to visual stimuli. We hypothesise that it is stronger in the middle and superficial layers, but we did not want to extend too much the scope of this paper. Additional analyses could be performed in the future on this evoked activity.

      Reviewer #3 (Recommendations for the authors):

      (1) Minor Corrections to the text and figures:

      We would like to thank the reviewer for the very valuable recommendations. Below we shortly describe how each suggestion has been implemented.

      We have made the white box more clear (see Figure 3 B).

      (b) Page 10: Top of 2nd paragraph - 'The full experimental protocol comprised a high resolution anatomical T1 scan lasting for 8 min'. The methods state this scan is 6 min 31 sec.

      The confusion results from the fact that the T1 scan was recorded during a short practice block that the participants performed inside the scanner. This block lasted 8min during which the 6 min 31 sec T1 scan was recorded. We have made this more clear:

      “Once prepared, the participant was placed inside the scanner and performed an 8 min practice block. A T1-weighted scan was acquired during this time in the sagittal orientation using a 3D MPRAGE sequence Brant-Zawadzki et al., (1992) with the following parameters: TR/TI = 2.2/1.1 s, 11° flip angle, FOV 256 x 256 x 180 mm and an 0.8 mm isotropic resolution. Parallel imaging (iPAT = 2) was used to accelerate the acquisition, resulting in an acquisition time of 6 min and 31s.”

      (c) Page 10: 'Stimulus presentation' paragraph - 'Stimuli were projected onto a screen behind the subject's head using'. The use of 'subject' should be replaced with 'participant' throughout.

      We have corrected the phrasing.

      (d) Page 14: Figures 2A and 2B are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (e) Figure 5 caption: 'Regressors are build for each time-frequency bin separately.' should be 'built'

      We have corrected the mistake.

      (f) Page 16, final paragraph: 'Afterwards, each of the resulting regression coefficients (B coefficients) was multiplied with the voxel specific layer weights that have been obtained as described above.' Should be 'were multiplied'

      We have corrected the mistake.

      (g) Page 17: 'Subsequently, separate analyses were done for two frequency of interest (FOI) ranges centerd around' - typo

      We have corrected the mistake.

      (h) Page 17 - 'Within these frequency ranges inferential statistics based a cluster level' - missing word. Should be 'based on a cluster level'

      We have corrected the mistake.

      (i) Page 14 Figure 2B and 5D are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (2) fMRI data pre-processing:

      Please provide a comment on the EEG-fMRI data quality - e.g. tSNR of EPI data. Perhaps example EPI data could be shown in the supplementary information.

      We included the below Figure S11 in Supplementary Figures showing an example EPI. We have also included an illustration of the result of our layering approach. Furthermore, we included a tSNR analysis (see Response to Reviewer 1, point 4).

      On a practical note - with 14-minute long runs whilst wearing an EEG cap, I would expect participant motion to be a concern. Could you provide some metrics on perhaps the average of the mean and maximum per subject displacement/rotation?

      We ensured that participants receive tactile feedback for their respective head motion from a strip of tape span across their foreheads. This resulted in overall manageable motion during each experimental block. During the main experiment, the average framewise displacement was 0.3 mm, with an average total translation of 1.6 mm and an average total rotation of 1.6 deg within each block. 

      We have added Figure S13 to Supplementary Figures.

      We have added a section to Methods:

      “Subject motion per block was low, with a mean (SD) frame-wise displacement Power et al. (2012) of 0.34 mm (0.24 mm) for the main experiment and 0.23 mm (0.22 mm) for the retinotopy (see also Figure S13 in Supplementary Figures).”

    1. eLife Assessment

      The report by Liu and colleagues provides a valuable analysis of environmental adaptation across diverse lineages of the grass Phragmites australis differing by their level of ploidy. The analysis reports solid evidence that lineages with distinct levels of ploidy occupy different climate niches. The use in tandem of regional survey and common garden experiment represents a convincing approach to suggest a correlation between ploidy and climate adaptation. This manuscript will be of interest to a broad community of ecological genomicists interested in how structural variation in gene dosage potentially affects the pattern of adaptation.

    2. Reviewer #1 (Public review):

      Summary:

      The article is testing the relative advantages of plant lineages with differing ploidy and admixture across environmental gradients. The results show that intraspecific variation in ploidy and admixture between lineages impacts plant traits that may enable persistence and range expansion.

      Strengths:

      Suitable marker panel size and convincing results that include attempts to analyse mixed ploidy level data, which is a challenge.

      Weaknesses:

      (1) Inadequate explanation of allele dosage for ploidy levels, some of which do not match the allele counts expected for genome copy number.

      (2) The setup and sample sizes of the common garden experiments are very unclear. The numbers implied are extremely low to draw robust conclusions.

      (3) Unclear how allele dosage is determined. Given it's so central to many analyses, it would be useful to see how this is done rather than use a citation.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript describes a combination of species distribution mapping experimental data from common garden and physiological experiments to project the future distribution of genetic subgroups with the widespread grass Phragmites australis. Overall, the sample sizes seem appropriate for the questions being asked, and the key results regarding projected change in distribution of the focal lineages are well supported. However, at this point, it is difficult to evaluate the broader impact of the work on the field or the utility of the data for the broader community outside of those studying the focal species, P. austrina.

      Strengths:

      A key strength of the paper is the use of common garden and physiological experiments in conjunction with species distribution modeling. The experiments provide a mechanistic basis for the correlations between interspecific lineage and climatic data, suggesting that the distributional patterns are more likely to result from genetic differences rather than limited dispersal among regions. I would, in fact, emphasize the experimental validation of modeling efforts even more in the introduction.

      Weaknesses:

      I see two weaknesses with the framing of the ms and the presentation of the results. First, no data support the claims that polyploidy has any causal effect. The ploidy levels are, in fact, completely confounded with other genetic differences, so it is not possible to eliminate genetic variation, independent of ploidy, as the causative factor. As the authors note, ploidy was not manipulated in the reported experiments. Thus, the focus on polyploidy in the introduction and elsewhere distracts from the novel and informative experiments that were conducted. Second, the manuscript indicates that intraspecific variation is critical for the evolutionary potential of a species to respond to environmental change, but intraspecific variation is seldom considered in species distribution models. To me, an assessment of evolutionary potential requires estimates of heritable genetic variation and responses to selection. The sample sizes presented here are modest to estimate heritabilities, but the manuscript could be framed with this perspective in mind. However, instead, the manuscript performs species distribution modeling on a small number of sub-specific lineages, essentially treating them as homogeneous "species" - thus the analysis commits the same oversimplification that the manuscript highlights, but does so at a finer evolutionary scale than species. Not acknowledging this simplification (or better, examining phenotypic variation within the genetically defined lineages) hinders what would otherwise be a strength of the manuscript.

      The title suggests that asymmetric introgression and thermal tolerance are the most important findings of the work. However, the introduction contains no explanation of the potential importance of gene flow (other than to say that asymmetric gene flow was suggested by some preliminary analyses), and the discussion offers only a limited explanation of either the potential mechanisms underlying the asymmetric gene flow or its importance for the long-term evolution of the species. Similarly, the novelty of combining experiments and species distribution modeling is scarcely mentioned, and there is no exploration of the connection between tolerance alleles and gene flow. Could introgression of heat tolerance alleles alter the spread of the hybridizing lineages, for example? A greater emphasis on these general population genetic parameters could potentially highlight the broader impact of this work.

    4. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback.

      We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them.

      Reviewer #1 raised two important concerns regarding our methodology. First, the determination of allele dosage is insufficiently explained, which is central to our ploidy assignment and downstream analyses. Second, the setup and sample sizes of the common garden experiments are unclear, raising questions about the robustness of our conclusions. We accept these criticisms and will address them as follows.

      Regarding allele dosage, we will add a detailed step-by-step description of our calling pipeline in the Methods section, including the criteria for peak height ratios and thresholds used to assign copy numbers. We will also clarify a crucial biological detail: the common reed (Phragmites australis) is an allotetraploid in its origin. As a consequence, many molecular markers, including the widely used SSR markers in previous studies, behave as disomic markers (i.e., two homeologous copies inherited in a Mendelian manner). Therefore, observing more than two alleles at a locus is indeed indicative of higher-level ploidy (hexaploidy or octoploidy) in this system. We will explicitly state this to resolve any confusion about why tetraploids in our dataset are treated as having a maximum of two alleles, while hexaploids and octoploids can carry more.

      Regarding the common garden experiment, we will explicitly report the replication number for each lineage-by-treatment combination and clarify the experimental design. We will also discuss the statistical approaches used given the sample sizes, while acknowledging that the consistency between experimental results and distributional patterns lends additional support to our conclusions.

      Reviewer #2 raised three substantive framing issues. First, ploidy is completely confounded with genetic background, yet our manuscript places undue emphasis on polyploidy as a causal factor. Second, our species distribution models treat each lineage as a homogeneous entity, failing to capture within-lineage variation and thus repeating the oversimplification we criticize. Third, we insufficiently explore the evolutionary significance of asymmetric introgression, gene flow, and the novelty of combining SDM with experiments. We fully agree with these points and will revise accordingly.

      To address the confounding issue, we will substantially reframe the manuscript to de-emphasize claims about polyploidy as a causal driver, and instead focus on the adaptive differentiation among distinct genetic lineages that happen to differ in ploidy. The Discussion will explicitly state that dissecting ploidy effects from background genetic effects will require future experimental approaches.

      To address the simplification in SDMs, we will add a clear acknowledgment of this limitation, discussing how it may affect predictive accuracy and suggesting that future studies incorporating population-level genomic data could more directly assess evolutionary potential.

      To address the insufficient exploration of introgression and the novelty of our approach, we will expand the Introduction to better highlight the value of coupling controlled experiments with SDMs at the intraspecific level. In the Discussion, we will elaborate on the evolutionary significance of asymmetric introgression, including testable hypotheses about how gene flow might mediate the spread of heat-tolerance alleles and influence lineage geographical limits under climate change.<br /> We also thank the reviewer for the suggestion to emphasize the experimental validation of SDM efforts, which we will incorporate into a revised Introduction.

      Looking beyond the present study, we envision three complementary directions that build upon our current findings. Expanding common garden experiments to include admixed individuals would test whether introgressed genomic blocks confer fitness advantages under thermal stress. Leveraging the population genomic framework established here, we will transition to whole-genome resequencing for selection scans and genotype-environment association analyses to pinpoint adaptive loci and reveal whether heat-tolerance alleles are preferentially transferred via asymmetric introgression. We will also integrate transcriptomic profiling with phenotypic measurements to identify candidate genes whose expression correlates with thermal performance and introgressed ancestry, helping to disentangle ploidy effects from genetic background. Together, these directions span expanded phenotyping, whole-genome resequencing, and transcriptome-guided discovery, forming an integrated framework that moves from the correlative patterns reported here toward mechanistic understanding. These perspectives are briefly outlined in our Discussion, and we hope the present study will serve as a foundation for these future investigations, which we plan to pursue in subsequent work.

      We believe these revisions will substantially strengthen the manuscript.

    1. eLife Assessment

      This important study combines peptide engineering, molecular docking, and functional assays to define the molecular basis of ligand recognition and activation of the human Y4 receptor and to identify three novel small-molecule agonists. The evidence supporting the conclusions is convincing, with complementary experimental and computational approaches providing strong support for the proposed receptor-ligand interactions. While concentration-response analyses of the small-molecule agonists and additional structural or mutagenesis studies would further strengthen the work, these are not essential to support the main conclusions. The work will be of interest to researchers studying GPCR pharmacology, structural biology, and ligand discovery.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes an investigation of peptide analogue agonists selective for the human Y4 receptor for pancreatic polypeptide over Y1, Y2 and Y5 receptors. After studies of mutated Y4R in transiently transfected COS-7 cells, binding models were calculated. Then, screening of a virtual library identified three non-peptidergic (albeit somewhat peptide-like) compounds with potential agonist activity that were subsequently confirmed and furthermore were found to have receptor interactions similar to the peptide analogues. This study provides fundamental new information that improves understanding of the Y4R structure and mechanism of activation by the native agonist and the selective peptide analogues. The non-peptide agonists have potential for future pharmacotherapy.

      Strengths:

      All of the experiments seem to be well performed, using state-of-the-art methods. The manuscript is quite comprehensive and has used a broad range of methods. The conclusions are convincingly supported by the experimental results.

      Weaknesses:

      The mutagenesis was almost exclusively based on the replacement of potentially interesting amino acid residues with alanine. Replacement with other residues, based on modelling and docking, could have refined the model further. Neither molecular dynamics nor cryo-EM was used to study the agonists' interactions with the Y4 receptor and these are therefore likely next steps in the characterization of the Y4R mechanism of activation.

    3. Reviewer #2 (Public review):

      Summary:

      Pelczyk et al. investigated the binding site of the neuropeptide Y Y4 receptor with the aim of identifying novel small-molecule agonists. The authors first assessed small cyclic peptides as tool compounds and then identified interactions between peptides and receptor residues, which were confirmed by single-point mutagenesis combined with functional assays for intracellular signalling. It is interesting that a peptide receptor can be activated by the relatively small cyclic peptides used in the study. The authors identified both common and peptide-specific interactions. The identified interactions guided ultra-large library screening, which yielded 53 compounds, 3 of which were confirmed as Y4R-specific agonists in an IP-one accumulation assay.

      Strengths:

      The combination of techniques (docking, mutagenesis and functional assays) strongly supports the identification and evaluation of small molecules as agonists at the neuropeptide Y Y4 receptor. Functional assays highlight residues that are important for the binding of all tested peptides, as well as residues with peptide-specific importance.

      The structure-activity relationship component of the study nicely highlights which components of the peptide are important for binding to the different members of the neuropeptide Y receptor family.

      Weaknesses:

      It would have been great to see concentration-response curves for the three identified small-molecule agonists, as this would have stengthened the case for these agonists.

    1. eLife Assessment

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

    3. Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The main strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses:

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

    5. Author response:

      eLife Assessment:

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We will make major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We will change the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions will greatly improve our work and presentation and we are grateful to the reviewers and our editors.

      We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will include an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS will strengthen our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results will be reported in a new Supplemental Table.

      We will include a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results will be reported in the revised manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will temper our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We will strengthen our control of clinical confounders, and add clearer statistical correction in the revision. We look forward to conducting future studies to test causality and understand mechanism.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We will note in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single center, cross-sectional. We will make further modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. eLife Assessment

      This is an important study that applies a new chromatin profiling technique to the study of cellular responses to low oxygen. The authors provide convincing evidence for distinct kinetic phases of the response and identify many new putative regulators of the response. This work will be of broad interest to those studying low oxygen responses and transcriptional regulation.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have satisfactorily addressed the comments raised in the previous round of review with textual revisions.]

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

    3. Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major comments from the first round of review:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

      Weaknesses:

      There are several areas where the clarity of presentation could be improved:

      (1) Given the importance of the methodology, the methods section needs more detail on how the extent of MNase digestion is chosen to achieve optimal results with MOA-seq. This is described to some extent in the description of control library preparation, but not for the experimental samples.

      We thank the reviewer for noting this unintended omission. We have not updated the Methods section to specify as follows:

      "Digestion patterns were assessed via gel electrophoresis, and the light digest levels ideal for MOA-seq (as per Savadel et al., 2021) were selected as the lightest digest levels that give a pattern of a nucleosomal ladder spanning the entire DNA fragment size range from undigested to mononucleosome bands, as indicated in Figure 1 with the asterisk-marked gel lanes."

      (2) The abstract describes this approach as "native cistrome profiling" but this is misleading since formaldehyde fixation is used.

      We believe the formaldehyde fixation captures native chromatin structure, but indeed we are digesting fixed chromatin and have updated the wording to read as “in situ cistrome profiling.”

      (3) Species- and field-specific jargon and abbreviations need to be clarified on first usage. For example, on page 9: "Downsampling analysis was carried out for two sets of published reference peaks; the CTCF cCRE peak midpoints and for the ERG motif under the ERG ReMap ChIP-seq peaks." The different categories of cCREs were not clearly defined, nor will it be clear what the term ReMap refers to for those outside the field. The sentence after this refers to IDR, which also should be defined.

      We thank the reviewer for highlighting the need for clearer definitions of field-specific terminology and abbreviations. In response, we have revised the manuscript to explicitly define all relevant terms at first mention. Specifically, we now describe the ENCODE candidate cis-regulatory element (cCRE) catalogue and define the individual cCRE categories, including promoter-like (PLS), proximal enhancer-like (pELS), distal enhancer-like (dELS), DNase I–H3K4me3 (K4m3), and CTCF-only regions. We also clarify that ReMap is a curated database of human transcriptional regulator binding peaks derived from ChIP-seq, ChIP-exo, and DAP-seq experiments. Additionally, we now define IDR as the Irreproducible Discovery Rate framework upon first use.

      (4) Figure 4C: Are these motifs examined under MOA sites specifically or anywhere in the genes in question?

      Leading up to and including Figure 4C, we have not yet examined any motifs. Instead, Figure 4C compares gene sets, one defined by our diff-MOA, and those from GO libraries, in this case the "target genes" which are defined by TF-specific studies, primarily ChIP-seq but also related immuno-based mapping techniques. Consequently, the analysis shown in Fig. 4C is not a motif enrichment analysis. Instead, we used the ENRICHR gene set enrichment analysis tool with ENCODE and ChEA consensus transcription factor target gene sets. Thus, the analysis was performed at the gene-set level, and transcription factor motifs were not examined within diff-MOA peaks or elsewhere in the associated genes for Fig. 4C. We note that motif enrichment within diff-MOA peaks was subsequently examined separately in Fig. 6. In Fig. 7, we further examined differentially expressed genes associated with diff-MOA peaks containing enriched transcription factor motifs and used clustering analyses to investigate their regulatory relationships. We have clarified these distinctions in the revised manuscript.

      If the question is about the location of MOA footprints relative to gene structure, we did not examine any MOA sites at any specific location, just overlapping the gene +/- 200 bp, as indicated in Fig. 4B.

      (5) Figure 5B shows that up-DEGs with diff-MOA footprints tend to show more losses of footprints. Do the authors interpret this as a loss of repressor binding?

      Not exclusively, but yes, that is one plausible explanation. That is, the activation (defined by increased RNA levels) via de-repression could be happening. But we also expect these dynamic footprints to be but one component. In other words, we interpret the relationship as consistent with that possibility, but not only that possibility. A logical explanation is that loss of footprint occupancy associated with upregulated genes could be based on displacement of repressive DNA-binding factors, thereby contributing to transcriptional activation. Thus, while loss of repressor binding is a plausible explanation for a subset of these events, additional factor-specific experiments would be required to know for sure in each case. We have added text to the Discussion acknowledging this possibility.

      Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major Comments:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Regarding additional biological replicates, we have acknowledged this point in the discussion. Importantly, we did subject the replicates to IDR analysis, which we explain in the methods as "In accordance with ENCODE ChIP-seq guidelines (Landt et al., 2012), we further evaluated data quality by assessing pooled pseudo-replicate consistency and self-consistency for each individual replicate (Supplementary Table S2)." This IDR analysis demonstrated consistent peaks between our bioreplicates, meeting ENCODE guidelines. In addition, downsampling analysis demonstrated that our sequencing depth of coverage (Supp Fig 1) was over 10-fold greater than required. We do appreciate that it will be useful to have more biological replicates from other cell types, tissues, or organisms, and hope this study prompts just such future research.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

      Live-cell imaging could indeed be interesting, but it is beyond our current capacity to add to this study and consider this an exciting future direction, but presence in the nucleus could include both bound and unbound HIF1, so the results may not easily track the DNA-bound HIF1 only.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In Figure 3B, the x-axis is not labeled.

      Thank you for pointing this out. We have revised Figure 3B by adding the previously missing x-axis label.

      Reviewer #2 (Recommendations for the authors):

      In the abstract, it would be good to define what MOA-seq is and what the cistrome is.

      Thank you for this suggestion. We have revised the abstract to define both MOA-seq (MNase-defined cistrome-Occupancy Analysis sequencing) and the cistrome upon first mention to improve accessibility for readers who may be unfamiliar with these terms.

    1. eLife Assessment

      The manuscript concerns a fundamental and controversial question in Trypanosoma brucei biology and the parasite life cycle, whether or not dividing slender bloodstream forms must transition to growth-arrested stumpy forms before differentiating to the procyclic form in the Tsetse midgut. The authors provide further evidence that slender bloodstream forms can infect Tsetse flies, and that although their differentiation is considerably delayed, they do not become classical stumpy forms during the process. The study is solid in design and execution, and addresses several criticisms made of the authors' earlier work, although discrepancies with results from other laboratories remain.

    2. Reviewer #2 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

      We outline below the 3 points that, in our opinion, remain to be clarified.

      (1) Effect of NAG on slender-form infections in tsetse flies

      The conclusion that "NAG has a negligible effect on slender infections in tsetse flies" based on Figure 1, cannot be fully supported in the absence of a positive control. A relevant positive control is well established in the literature, namely that NAG promotes Tsetse infection by stumpy forms. Without such a control, it is not possible to exclude technical issues (for example, an ineffective NAG treatment), which would yield results similar to those presented in Figure 1.

      We agree that an internal stumpy-form positive control would provide an additional technical reference. However, the enhancing effect of NAG on stumpy-form midgut infections is well established and was also demonstrated under the experimental framework of our original study (Schuster et al. 2021, Figure 2A).

      The purpose of the present Research Advance was therefore not to re-establish the known effect of NAG on stumpy infections, but to test whether slender-form infections require NAG supplementation. Under the conditions tested here, slender bloodstream forms established midgut, proventriculus and salivary-gland infections also in the absence of NAG. We have revised the text accordingly to avoid implying a general absence of NAG effects and to make clear that our conclusion is restricted to slender-form infections under the conditions tested (line 128).

      (2) Infection of non-teneral flies

      Because the experiments shown in Figure 1 (teneral flies) and Figure 2 (non-teneral flies) were not conducted in parallel or under identical conditions, it is important that the figure legends clearly state the parasite numbers used in each case. Specifically, infections of teneral flies were performed with 200 parasites/mL (approximately 4 parasites per bloodmeal), whereas non-teneral infections used 1 × 10<sup>6</sup> parasites/mL (approximately 20,000 parasites per bloodmeal?). At present, this information is scattered across the Methods and Supplementary Tables 1 and 2, making it difficult for readers to immediately appreciate that the parasite load differs by roughly 5,000-fold between these conditions.

      As previously shown by the authors (Schuster et al., 2021) and in the Rotureau laboratory (Tsagmo Ngoune et al.), and as generally expected, the initial parasite dose strongly influences infection outcomes in teneral flies. In this context, it would be informative to know whether the authors have attempted infections of non-teneral flies using lower parasite numbers (noting that Tsagmo Ngoune et al. used a maximum of 10,000 parasites) and what the infection rate was.

      Relatedly, the statement in line 370 appears to be an overgeneralization, as fly age was not directly tested under matched experimental conditions:

      Line 370 - "Here, we unambiguously show that, in the absence of immunosuppressive treatment, slender forms can establish infections in tsetse flies, irrespective of the fly's age or sex."

      We thank the reviewer for highlighting the inconsistent presentation of parasite doses between Figure 1 and 2. We agree this is confusing and have revised the figure legends to clearly state both the parasite concentration (cells/mL) and estimated fly uptake per bloodmeal for each experiment (Lines 143 and 206).

      Regarding experiments with non-teneral flies using lower parasite numbers: We have not tested intermediate doses (e.g., 10,000 parasites/bloodmeal as used by Ngoune et al.) in non-teneral flies. Given that teneral flies already show relatively low infection rates even under optimal conditions, we chose the higher parasite dose (20,000 parasites/bloodmeal) for non-teneral flies to ensure sufficient statistical power for meaningful analysis of infection outcomes across different fly compartments.

      We acknowledge the reviewer's concern regarding the statement in line 370 and have revised this sentence (line 375) to more accurately reflect our experimental conditions, avoiding overgeneralization beyond the specific parameters tested.

      This reads now: “Here, we demonstrate that slender forms can establish infections without immunosuppressive treatment under the conditions tested. This infectivity was observed in both teneral and non-teneral, as well as in both male and female flies, indicating that slender forms retain transmission potential across different fly demographics. However, direct age comparisons under identical parasite doses remain to be tested.”

      (3) Transcriptomic analysis

      Supplementary Figure 8 lacks statistical analysis, which limits its interpretability. Two types of comparisons would be particularly helpful:

      (i) a comparison of PAD1/2 expression levels between slender and stumpy forms at 0 h; and

      (ii) for each gene, a comparison of the overall change in expression (from 0 to 72 h) between infections initiated with slender versus stumpy forms.

      In addition, the figure legend should clarify what "expression levels" refer to. TPM? Normalized counts?

      We appreciate this helpful comment and included statistical analysis for the expression of PAD1 and PAD2 (Supplementary Figure 8) between the two forms for the baseline (0 h) as well as during the differentiation to procyclic forms (0 h to 72 h) by using Welch´s t-test.

      While PAD1 did not show a statistically significant difference in this analysis, PAD2 displayed significant differences in expression dynamics over time. This supports the broader transcriptomic observation that slender- and stumpy-initiated differentiation follow distinct transcriptional trajectories before converging at the procyclic stage.

      We also clarified the figure legends showing the mean log2 counts per million (CPM) values.

      Finally, for the benefit of the field, eLife could encourage publishing a collaborative study in which the Engstler and Rotureau laboratories exchange parasite lines and culture protocols (including media with and without methylcellulose) and perform tsetse fly infections in parallel in their respective laboratories. Such an approach could help resolve the remaining discrepancies and provide a valuable reference for the community.

      We appreciate this constructive suggestion. A collaborative inter-laboratory study in which parasite lines, culture conditions and infection protocols are exchanged between the Engstler and Rotureau laboratories would be a valuable way to address the remaining discrepancies in the field. In particular, parallel infections using matched parasite lines and culture conditions, including media with and without methylcellulose, could provide a useful reference dataset for the community.

      At the same time, such a study would require substantial coordination, reciprocal strain exchange, protocol harmonization and new infection series in two laboratories. It therefore goes beyond the scope of the present Research Advance, which was designed to address the specific methodological concerns raised in response to our original publication. We have restricted our conclusions accordingly and view the proposed collaborative benchmark study as an important direction for future work.

    1. eLife Assessment

      This important study reports the development of the first tankyrase degrader and demonstrates its enhanced ability to inhibit β-catenin signaling compared to conventional tankyrase inhibitors. The evidence supporting the conclusions is comprehensive and convincing, based on rigorous biochemical and cellular analyses. The findings will be of broad interest to researchers studying Wnt signaling, protein degradation, and cancer biology.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation-thereby impairing β-catenin degradation-the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Comments on revised version:

      I had a favorable opinion of the manuscript in the first round of review. I don't have additional comments on the revised manuscript. The manuscript looks fine to me.

    3. Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Comments on revised version:

      I thank the authors for responding to the queries raised in the original review, most of which have now been addressed. This further strengthens this well-conducted study and well-presented manuscript. I congratulate the authors for this interesting and insightful work.

      A few minor points remain:

      I appreciate the authors acknowledge that testing the physical properties of the degradasome puncta is necessary to explore whether they indeed represent condensates. The term "condensates" implies liquid-liquid phase separation (rightly or wrongly). However, this question has not yet been resolved in the case of degradasomes. I therefore suggest the term "condensates" to be avoided. A simple morphological description as "puncta" may suffice.

      I thank the authors for including the additional data comparing tankyrase binding by IWR and IWR-POMA. I agree that using the BRET signal of IWR-POMA is informative. Adding the IC50 values directly to the figure panels (S3E, S3G) would help the reader to quickly assess binding. The comparison between these two panels is insightful.

      Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue?

    4. Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

    5. Reviewer #4 (Public review):

      From the Reviewing Editor:

      This important study reports the development of the first PROTACs targeting the ADP-ribosyltransferases tankyrase 1 and 2, with the goal of inhibiting Wnt/β-catenin signaling more completely than is possible with catalytic tankyrase inhibitors. The work addresses a significant limitation of existing tankyrase inhibitors: although catalytic inhibition stabilizes AXIN1/2 and suppresses Wnt signaling, it also stabilizes tankyrase itself, potentially enhancing non-catalytic scaffolding functions and promoting accumulation of degradasome-like puncta.

      The evidence is convincing. The authors use appropriate and well-validated approaches, including chemical biology, cellular assays, and proteomic profiling, to show that PROTAC-mediated degradation of tankyrase avoids tankyrase accumulation while still stabilizing AXIN and inhibiting Wnt/β-catenin signaling. The data support the conclusion that degradation of tankyrase can separate pathway inhibition from the confounding effects of stabilized tankyrase protein and may therefore offer advantages over conventional catalytic inhibitors.

      A strength of the study is the clear mechanistic comparison between tankyrase degradation and catalytic inhibition. The manuscript provides convincing evidence that the PROTAC and catalytic inhibitors act through distinct mechanisms, with the PROTAC targeting both catalytic and scaffolding roles of tankyrase. The study is well conducted and clearly presented, and the authors have addressed most concerns raised during review.

      A remaining limitation is that the therapeutic potential of the compound is not tested in vivo, for example in APC-mutant colorectal cancer models, APCmin mice, or patient-derived xenografts. Such experiments would strengthen claims about practical efficacy, although they are not essential for the main mechanistic conclusions of the manuscript.

      Overall, this is an important and insightful contribution. It advances the tankyrase and Wnt signaling fields by providing a new chemical strategy to suppress tankyrase function more completely than catalytic inhibition alone, and it offers a useful framework for future therapeutic exploration of tankyrase degradation.

    6. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation - thereby impairing β-catenin degradation - the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Weaknesses:

      (1) A key caveat is that the identified tankyrase degrader also targets GSPT1 for degradation. This raises the possibility that GSPT1 degradation may contribute to the observed β-catenin and tumor growth inhibition.

      (2) The authors address this concern reasonably by showing that DLD1 cells resistant to GSPT1 degradation remain sensitive to the tankyrase degraded.

      (3) To further strengthen this point, the authors might consider generating TNKS1/2 double knockout cells (e.g., in DLD1 or SW480 backgrounds) and demonstrating that the degrader loses its growth-inhibitory effect in these models. However, given the technical challenges of creating double knockouts in cancer cell lines, such experiments could be considered optional.

      We thank the Reviewer for the favorable feedback. The major concern is the collateral degradation of GSPT1. As the Reviewer noted, IWR1-POMA was able to suppress colony formation in DLD-1 cells resistant to a GSPT1/2 degrader (DLD-1R, Figure 6B and S9F), suggesting that TNKS but not GSPT degradation is responsible for growth inhibition.

      We also appreciate that the Reviewer brought it to our attention an important early observation of the TNKS scaffolding effects. Cong reported in 2009 that overexpression of TNKS induced AXIN puncta formation in a SAM but not PARP domain-dependent manner (PMID: 19759537, Ref. 12). We have added this reference to the introduction of TNKS scaffolding in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Weaknesses:

      Although the results are promising and mostly compelling, the claim that the PROTACs provide "a deeper suppression of the WNT/β-catenin pathway activity" requires further corroboration, particularly at endogenous tankyrase levels.

      We thank the Reviewer for the encouraging and insightful comments. The major critique concerns whether TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels. IWR1-POMA reduced the level of cytosolic β-catenin more effectively than IWR1 in Wnt3A-stimulated HEK293 cells without protein overexpression (Figure 1D). IWR1POMA also suppressed STF activity more effectively than IWR1 in DLD-1 cells (Figure S8C) and reduced the expression levels of several WNT/β-catenin targets more effectively than IWR1 (Figure 1G and S8D). These results support that TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels.

      There are also some other points that, if considered, would further improve the manuscript, as detailed below.

      (1) Abstract and line 62: Many catalytic tankyrase inhibitors tend to display toxicity, which is likely on-target (e.g., 10.1177/0192623315621192; 10.1158/0008-5472). This constitutes the main limiting factor for these compounds. An incomplete inhibition of Wnt/beta-catenin signalling may contribute to the challenges, but this does not appear to be the dominant problem. A more prominent introduction to this important challenge is probably expected by the field.

      A previous study showed that G007-LK, a selective TNKS inhibitor, exhibited weak efficacy and dose-limiting toxicity at 5‒30 mg/kg BID or 10‒60 mg/kg QD in various mouse xenograft models (PMID: 23539443, Ref. 28). Similarly, G-631, another TNKS inhibitor, also showed dose-limiting toxicity without significant efficacy at 25‒100 mg/kg QD in mice (PMID: 26692561, Ref. 60). However, other studies showed that G007-LK was well-tolerated at 200 mg/kg QD over 3 weeks in mice (PMID: 29316982, Ref. 61), and treating mice with G007-LK at 10 mg/kg QD over 6 months also improved glucose tolerance without notable toxicity (PMID: 26631215, Ref. 62). Importantly, basroparib, a selective TNKS inhibitor, was well tolerated in a recent clinical trial (PMID: 40964966, Ref. 64), and constitutive silencing of both TNKS1 and TNKS2 for 150 days in APC-null mice prevented tumorigenesis without damaging the intestines (PMID: 31337618, Ref. 8). We have included some discussion of the toxicity issue associated with TNKS targeting at the end of the Discussion section.

      (2) The authors do a good job in setting the scene for the need for tankyrase degraders. Their observations relating to the formation of puncta (degradasomes) being tankyrase-dependent are compatible with a previous study by Martino-Echarri et al. 2016 (10.1371/journal.pone.0150484): simultaneous silencing of TNKS and TNKS2 by RNAi abolishes degradasome formation. The paper is cited as reference 17, but only in passing, and deserves more prominence. (It includes an entire paragraph titled "Expression of tankyrases 1 and 2 is required for TNKSi-induced formation of axin puncta").

      Indeed, Henderson’s 2016 paper (PMID: 26930278, previously Ref. 17, now Ref. 18) shed important light on the role of TNKS scaffolding in the DC. However, whereas this study demonstrated that knocking down both TNKS1 and TNKS2 by siRNA prevented G007-LK to induce AXIN puncta, it concluded that “puncta formation requires both the expression and the inactivation of TNKS,” which is inconsistent with our observations that accumulation of either catalytically active or inactive TNKS can promote AXIN puncta formation. The function roles of TNKS scaffolding in the DC also remained unaddressed. We have included additional discussion of Henderson’s findings in the first paragraph the Discussion section.

      (3) Moreover, the scaffolding concept has been discussed comprehensively in other studies: 10.1111/bph.14038 and more recently 10.1042/BCJ20230230. There are also a few studies that focus on targeting the ankyrin repeat clusters of tankyrase to disengage substrates (10.1038/s41598-020-69229-y; 10.1038/s41598-019-55240-5) that illustrate the concept of blocking the scaffolding function. In that sense, the hypotheses are mature, and it is interesting to see some of them supported in this study. The authors could improve how they set their work into the context of these other efforts and proposals.

      Indeed, Guettler demonstrated in 2016 that TNKS scaffolding could promote WNT/β-catenin signaling, which forms the basis of the current work. Meanwhile, whereas there have been efforts to target the SAM or ARC domain to address TNKS scaffolding by Guettler and Lehtiö, our approach of targeting TNKS for degradation is complementary. We have included in the last paragraph of the Discussion section information on efforts to target the ARC or SAM domains as an alternative approach to suppress WNT/β-catenin signaling without promoting TNKS oligomerization (PMID: 31836723 and 32704068, Ref. 66 and 67).

      (4) In several places in the manuscript, the DC is referred to as "biomolecular condensate", at times even as a "classic example", implying that it operates through phase separation. This has not been demonstrated. In fact, super-resolution microscopy indicates that the puncta are not droplet-like (10.7554/eLife.08022), which would argue against the condensate hypothesis.

      Biomolecular condensates are membraneless cellular compartments formed by phase separation of biomolecules, regardless of their physical/material properties (PMID: 28935776 and 28225081, Ref. 22 and 23). Super-resolution microscopy studies by Stenmark (PMID: 26124443, Ref. 17) showed that AXIN, APC, TNKS, and β-catenin interacted with each other to assemble into membraneless complexes, wherein AXIN and APC formed filaments throughout the DC. Peifer has also summarized evidence that supports the condensate nature of the DC (PMID: 30782412, Ref. 9; see also PMID: 26393419). However, we acknowledge that testing the physical properties of reconstituted DC (for example, PMID: 34352208) with TNKS will provide a better understanding of the nature, for example liquid vs. gel, of these condensates.

      (5) It is beautiful to be able to use IWR1 and IWR1-POMA at identical concentrations for direct comparisons. However, this requires the two compounds to bind to tankyrase similarly well and reach the target to a comparable extent. How sure are authors that target engagement is comparable? Has this been evaluated?

      Using a BRET assay, we have confirmed that IWR1-POMA binds to TNKS1 with affinity comparable to that of IWR1. Details of this study is now included in the Results sections, and the data are presented in the Supplementary Information (Fig. S3E–G).

      (6) Figure 1F: It is not immediately apparent how IWR1-POMA shows more complete containment of Wnt/beta-catenin signalling. Most Wnt/beta-catenin targets lie close to the perfect diagonal, so I do not see how the statement "that IWR1-POMA controlled WNT/β-catenin signaling more effectively than IWR1" (in the legend of Figure 1F) is supported. Minimally, an expanded explanation would benefit the reader. Providing the colour-coding legend directly in the figure would help improve clarity. Also, the panel is very small and may benefit from a different presentation in the figure.

      We have updated Fig. 1F to include an inset of Quadrant III for improved clarity and readability. We have also moved Fig. S7C to the main text as Fig. 1G and added an expanded explanation for these figures.

      (7) Figure 2: The conclusion of a "deeper suppression" of signalling relies on overexpression of tankyrase in an otherwise tankyrase-null background. Have the authors attempted to measure reporter activity or endogenous gene expression without tankyrase overexpression, in Wnt3a-stimulated cells (in the context of a normal Wnt/beta-catenin pathway) or CRC cells at the basal level? Non-catalytic activity in a similar assay has previously been observed upon tankyrase overexpression (10.1016/j.molcel.2016.06.019). Whether or not there is a substantial scaffolding effect at endogenous tankyrase levels after tankyrase inhibition remains unconfirmed, and the PROTAC is a valuable tool to address this important question. The findings presented in Figure S7C and D go some way towards answering this question - these data could be presented more prominently, and similar assays could be performed in other cell systems.

      IWR1-POMA suppressed STF activity more effectively than IWR1 in APC-mut DLD-1 and SW480 CRC cells without TNKS overexpression (Fig. S8C). Similarly, IWR1-POMA provided a deeper suppression of STF signals in HeLa cells transfected with AXIN1 and β-catenin while expressing endogenous TNKS (Fig. 4G). These results suggest that inhibitor-induced TNKS scaffolding plays a significant role at endogenous TNKS expression levels. Following the reviewer’s suggestion, Fig. S7C is now Fig. 1G.

      (8) Line 237/238: "TNKS accumulation negatively impacts the catalytic activity of the DC (Figure 5D)" - the data do not show this. Beta-catenin levels are a surrogate readout for DC function (phosphorylation and ubiquitylation). Minimally, this requires rewording, with reference to beta-catenin levels.

      We have rephrased "TNKS accumulation negatively impacts the catalytic activity of the DC" as "TNKS accumulation negatively impacts the exchange of β-catenin in the DC."

      (9) Line 303-304: Beta-catenin is thought to exchange at beta-catenin degradasomes; this is clear from previous FRAP assays and the observation that phospho-beta-catenin accumulates in degradasomes upon proteasome inhibition (10.1158/1541-7786.MCR-15-0125). However, degradasome size hasn't, to my knowledge, been related to activity. Can this be clarified, please?

      We apologize for confusing β-catenin phosphorylation with β-catenin abundance. Here, we refer the catalytic activity of the DC to as the ability of the DC to promote β-catenin degradation rather than the kinetics of β-catenin phosphorylation. It is commonly observed that AXIN stabilization by TNKS inhibitors increases the DC size and reduces the β-catenin levels. As such, the induction of AXIN puncta by TNKS inhibitors is frequently used as an indicator of WNT/β-catenin pathway inhibition. However, we have found that, TNKS inhibition drives TNKS accumulation, which reduces the ability of the DC to promote β-catenin degradation. We agree that the DC only primes β-catenin but does not catalyze its degradation. We have revised our manuscript as follows: "increasing the local concentration of the DC components improves its 'effective activity'[50,51]."

      (10) There are previous hypotheses/proposals that the sensitivity of CRC cells to tankyrase inhibition correlates with APC truncation or PIK3CA status (10.1158/1535-7163.MCT-16-0578; 10.1038/s41416-023-02484-8). Have the authors considered expanding their cell line panel (Figure S7) to sample a wider range of cell lines, including some that are wild-type with regard to APC or Wnt/beta-catenin signalling in general? This would be a valuable addition to the work. Quantitated colony formation data could be moved to the main body of the manuscript.

      We have so far tested the effects of IWR1-POMA on the proliferation of DLD-1, SW480, HT-29, HCT116, and RKO cells (Fig. 6A and 6B). While a heterozygous Ser45 deletion in CTNNB1 confers resistance to IWR1-POMA, we did not observe sensitivity associated with APC or PIK3CA status. The ability of IWR1-POMA to suppress the growth of RKO cells expressing wild-type APC is consistent with a previous report that knockdown of both TNKS1 and TNKS2 stabilized PTEN to suppress cell proliferation and glycolysis in vitro and tumor growth in vivo (PMID: 25547115, Ref. 48) independently of the β-catenin pathway. We have added this new information as well as quantification of the colony growth results (Fig. S8A, S8B, S9A, S9F, and S9G) to the revised manuscript.

      (11) The manuscript only mentions toxicity (i.e., therapeutic window) in the last sentence of the Discussion section. As this is THE main challenge with tankyrase inhibitors (as mentioned above), can the authors expand their discussion of this aspect? Is there an expectation that PROTACs may be less toxic?

      As discussed above, evidence for on-target toxicity of WNT/β-catenin inhibition is mixed. Yet, the absence of dose-limiting toxicity for basroparib at doses up to 360 mg QD in human (PMID: 40964966, Ref. 64) is encouraging. PROTAC works by catalyzing target degradation, which is different from traditional catalytic inhibitors that require continuous target occupancy at a high level. It remains unclear whether the observed on-target toxicity of TNKSi is associated with TNKS accumulation at high doses, akin to the cytotoxicity induced by PARP1-trapping upon catalytic inhibition. We have included a brief discussion of the toxicity issue in the final paragraph of the Discussion section.

      (12) Figures 3, 4, 5A: For fluorescence microscopy experiments, can these be quantified, and can repeat data be included?

      We have included quantification data and replicate information for Fig. 3–5.

      (13) Figure 4, S6: An additional channel illustrating the distribution of cells (e.g., nuclei, cytoskeleton, or membrane) would be helpful for orientation and context for the AXIN1 signal.

      We have included cell outlines or nuclear staining for Fig. 3, 4, S6, and S7.

      (14) How were cytosolic fractions of cells prepared to assess cytosolic beta-catenin levels? This detail is missing from the methods.

      We have updated the Methods section to include additional details on the preparation of the cytosolic fractions of cells.

      Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

      We thank the Reviewer for the helpful suggestions.

      I have some recommendations listed below for consideration by the authors to strengthen their study:

      (1) The title is slightly misleading, as it is already known that the scaffolding function of TNKS is important within the DC. The authors should consider incorporating the PROTAC targeting aspect in the title (e.g., PROTAC-mediated targeting of tankyrase leads to increased inhibition of betacat signaling and CRC growth inhibition).

      We have modified the title accordingly to "Targeting tankyrase scaffolding in the β-catenin destruction complex by PROTAC overcomes the limitation of catalytic inhibitors in cancer."

      (2) The authors should comment in the manuscript on the bell-shaped curve obtained with treatment of cells with the PROTACs (Figure S2C). This likely indicates tittering of the targets within a bifunctional molecule with increasing concentration (and likely reveals the auto-inhibition conferred by the catalytic inhibition alone).

      As suggested by the Reviewer, the bell-shaped dose-response likely originated from the formation of non-productive binary protein-ligand complexes at high PROTAC concentrations. We have added a sentence to clarify this unique behavior of PROTAC molecules.

      (3) The authors comment that using G007-LK as warehead was unsuccessful, but they do not show data. Do the authors know why this was the case?

      The structure-activity relationship of PROTACs is often unpredictable, as both the kinetics and thermodynamics of target and E3 ligase binding play important roles in promoting efficient target degradation. We have include data on G007-LK based PROTACs (Fig. S2D) in the revised manuscript.

      (4) Throughout the manuscript, the authors need to do a better job at quantifying their results (i.e., the western blots and the IF). For example, the degradation of TNKS1/2 in Figure 1D is not overly convincing. Similarly, the IF data in Figure 3 needs to be quantified in some ways. Along the same lines, the effect of IWR1-POMA treatments on the proliferation of cells and organoids should be quantified using viability assays... There is also no indication of how many times these experiments were performed and whether the blots shown are representative experiments. The quantification should include all experiments.

      We have included quantification of the immunofluorescence images, colony formation data, and Western blots in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) For clarity, can the authors use the official gene names, TNKS and TNKS2?

      We favor using TNKS1 and TNKS2 when referring to the protein for clarity and use TNKS for simplicity when referring to both proteins.

      (2) Line 92: The authors refer to TNKS2 "induction" - it remains unclear what is meant by "induction".

      We have changed "without induction" to "under basal conditions".

      (3) Can the authors please display molecular weight markers for Western blots throughout?

      (4) Line 144: The description "significantly more effectively" refers to Figure S5A, which shows a single, non-quantified Western blot. I don't think significance has been tested, and this statement should be reworded, or quantified aggregate data provided.

      We have added a Supplementary Information file showing molecular weight markers and quantification of Western blots.

      (5) Line 226: "plateaued at a much lower level" - can this be expressed more quantitatively in the text?

      We have included more quantitative information on the FRAP results.

      (6) Line 249: Can the authors repeat the cross-reference to Figure S7A here?

      We have repeated the cross-reference to the figures.

      (7) Line 266: The description of the experiment using the GSPT1/2 degrader CC-90009 would benefit from a brief recap of the purpose as not every reader will be familiar with this common PROTAC off-target. This is a very thorough analysis, though, and commendable.

      We have added background information on GSPT1 degradation to the revised manuscript.

      (8) Figure 1A: Can the number of repeats and the type of repeats be indicated, please?

      (9) Figure 2: Does n refer to biological or technical repeats?

      (10) Figure 5B, D: How many separate experiments are the data based on?

      (12) Figure S3D, S9A, D: number and types of repeats and the nature of the displayed data and error bars need to be included, please.

      (13) Figure S6B, S7B: I can see three data points, but it would still be helpful to state the number and type of repeats in the legend.

      (14) Figures S9A, S9D: There is value in showing the cumulative data from several repeats in the main figure (Figure 6, which currently is only qualitative) rather than the supplementary material.

      (15) Where single Western blots are shown, can the authors indicate how many experiments they are representative of?

      We have included the number of biological repeats for all data.

      (11) Figure S2C: For most graphs, the main response of interest occurs at low compound concentrations. The y-axis scale does not always help the reader to appreciate the effects, as the response seems small against the magnitude of the hook effect. Interrupting the y-axis as in the final panel may help, with y-axis scales consistent over all panels in the figure.

      We have updated Fig. S2C to emphasize on the degradation efficacy.

      (16) The authors may want to give further method details for some of their assays to facilitate replication of their experiments in the future. For example, the STF assay description is currently quite minimalistic. I assume the assay is fairly robust, though. Other details include cell media (general media details and specific additives and their concentrations in the 3D spheroid formation assay), etc. A general look at the methods section will likely be beneficial.

      We have updated the Methods section to provide more detailed experimental information.

      Reviewer #3 (Recommendations for the authors):

      (1) In Figure 2A, one of the most important findings of the manuscript is that IWR1-POMA induced promoted deeper suppression of beta-catenin-mediated transcription. This seems to be the case only at 3.2uM. Is it statistically significant? What are the data points on this graph? What are the error bars?

      We have included statistical analysis as Fig. S5G.

      (2) On Figure 2C and 2D, do the authors know why the TNKS20M1054V mutant is much better at promoting signaling than the TNKS1-PD ? Is it expression levels?

      It is indeed interesting that TNKS2-M1054V promoted significantly stronger WNT signaling than TNKS1-PD. The basis for its strong scaffolding effect is unclear.

      (3) In Figure 4C, the authors claim that when cells are treated with IWR1-POMA, AXIN1 is distributed diffusely throughout the cytoplasm. It appears that small punctae are visible.

      Quantitative analysis (Fig. 4F) suggest that the size of AXIN1 puncta upon IWR1-POMA is rather insignificant.

      (4) Label on Figure 1D has a spelling error TNKS1/2.

      Corrected.

    1. eLife Assessment

      This manuscript presents openretina, an open-source platform that integrates retinal datasets, model training, benchmarking, and in silico analysis tools within a unified framework. The resource is valuable because it addresses long-standing challenges in reproducibility, accessibility, and cross-study comparison in computational retina research, while providing a foundation for community-driven model development and evaluation. The supporting evidence is solid, with the authors demonstrating a functional and well-documented platform across multiple datasets and species, although a clearer discussion of model interpretability, current performance limitations, and data quality standards would strengthen the resource.

    2. Reviewer #1 (Public review):

      Summary:

      This "Tools and Resources" submission describes a platform for the modeling of stimulus-response relationships in the retina. It includes a repository for experimental data sets with standardized programmatic access, and a suite of software for constructing stimulus-response models and evaluating them.

      Strengths:

      (1) The paper is well written.

      (2) The platform could serve an integrative function by connecting different research programs and offering a common baseline for evaluating stimulus-response models.

      (3) The finding that there is "substantial explainable variance remains uncaptured by current models" is a useful insight to motivate further work and measure progress.

      Weaknesses:

      (1) The modeling supported by the package focuses on predictive accuracy at the cost of less interpretability.

      (2) The article needs to make a stronger argument that this style of modeling is fruitful, especially when applied to the retina.

      Main comments:

      (1) Abstract machine learning vs mechanistic models. The "Core + Readout" architecture advocated here seems to be divorced from all the neurobiological detail that is already known in the retina. It mostly aims at prediction, not interpretation. Such a black-box modeling framework is useful in brain regions where we know very little about connectivity, or mechanisms, or even about the primary function being performed there, like in the mammalian cortex. In those cases, any model that can deliver a prediction is a step forward, even if it does not connect to biological mechanisms. But that's decidedly not the situation in the retina, where so many mechanistic details are known: from consensus cell types, to synaptic detail, to single-neuron biophysics, to circuit motifs. How can one connect this ML modeling approach with the extensive mechanistic knowledge available in retinal neuroscience? And can the combination somehow lead to a better understanding? The authors seem to recognize this tension (e.g. line 215ff and 370ff) but don't give it much weight. A stronger case needs to be made here for how this kind of modeling will advance the field.

      (2) The "gradient field" approach. Figure 4c illustrates a case of this dissonance. The gradient field of the response increases with contrast in multiple directions. This is obvious a priori (see line 274) from the more mechanistic model we already have of this On-Off cell. These are the W3 cells described in www.pnas.org/cgi/doi/10.1073/pnas.1211547109. The circuit-based model from that paper, with rectifying on and off subunits from bipolar cells, gives a much more compact explanation for what the neuron does. Because each of the subunits has a spatio-temporal receptive field, this model can predict the entire dynamics to arbitrary stimuli, rather than just 2 dimensions of static stimuli as in the present analysis. So what is the value added here? Again, a stronger case needs to be made that these "Core + Readout" modeling activities enhance understanding.

      (3) The "most exciting input" approach (Line 193ff):

      - Presumably, some power constraint must be put on the stimulus? Otherwise, increasing the contrast will make it more exciting. What are these constraints?

      - Presumably, this optimal stimulus is computed from the model based on non-optimal stimuli? What are the assumptions going into that?

      - The most exciting stimulus is not necessarily the most useful characterization. Near its maximal firing rate, the neuron doesn't discriminate stimuli much, because the slope there is zero (line 237). Instead (or in addition), one would like to know along which stimulus axis the neuron is most sensitive. See e.g. discussion in Dayan & Abbott 2000, Figure 3.11.

    3. Reviewer #2 (Public review):

      Summary

      openretina is a Python package for training and applying convolutional neural network-based models of retinal ganglion cell responses. The package integrates dataloading, model training, and evaluation in a unified framework built on PyTorch Lightning and Hydra, and ships with pre-trained model checkpoints and publicly available datasets (whitenoise, natural scenes) spanning multiple species (marmoset, mouse, axolotl, salamander) and recording modalities (multielectrode array recordings or 2-p calcium imaging). Beyond predictive modelling, openretina includes a suite of in silico analysis tools for probing learned representations, including maximally exciting input synthesis, discriminatory stimulus optimisation, and model weight visualisation. The broader openretina initiative aims to establish a community-driven platform for computational retina research, lowering barriers to entry and facilitating cross-dataset model benchmarking. This is a valuable contribution given the longstanding fragmentation of datasets, codebases, and analysis practices across retina laboratories.

      Strengths:

      The tool has several strengths. By providing a framework built on deep learning infrastructure, the package substantially lowers the barrier to entry for researchers without extensive machine learning backgrounds. The inclusion of pre-trained model checkpoints across multiple species and recording modalities will allow users to apply state-of-the-art models. The in silico toolkit - and in particular the MEI synthesis pipeline - has already demonstrated its scientific potential, with prior work using optimised stimuli to discover a previously uncharacterised RGC type confirmed experimentally, illustrating what becomes possible when these tools are made broadly accessible. The current modular Core + Readout architecture is a well-suited architecture for modeling retina responses. The HDF5-based data standard provides a sensible common format for contributing new datasets. Overall, the initiative is well-motivated, the engineering is competent, and the vision of a collaborative, community-driven platform for retina modelling is one that the retina community would benefit from.

      Weaknesses:

      (1) The in silico tools provided are valuable, but users should interpret their outputs in light of the performance of the underlying models. The predictive performances of current models and datasets in the package are far from performance ceilings.

      (2) The authors appropriately note that optimised stimuli reveal what a neuron responds to but not how the computation is implemented. I would encourage readers to keep this distinction in mind when using the weight visualization tools as well - convolutional filters in a shared, unconstrained core do not map onto retinal circuit elements, and should be treated as model descriptors rather than circuit proxies. For example, RGCs of the same type may appear to sample inputs from two different filters, which should have been a single filter. Or a single RGC may be sampling from two filters, which under more constrained conditions could be approximated with a single filter. These are degeneracies in the CNN modeling framework that should be kept in mind when drawing circuit-level interpretations.

      (3) The datasets currently distributed with the package vary in recording quality, and users should be aware that model performance may not only reflect architectural limitations but may also be limited by noise and data artifacts, including spike sorting errors.

      (4) As the platform grows and community-contributed datasets are added, explicit data quality standards will be essential. I encourage the authors to develop dataset standards to ensure that their resource provides access to highly curated datasets, which I believe is an important step in having high-fidelity models whose functional interpretations can be trusted.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript presents openretina, a Python-based platform designed to facilitate collaborative retinal modeling across datasets, laboratories, species, and recording modalities. The package provides standardized model architectures, evaluation metrics, and analysis tools, while also integrating several publicly available retinal datasets. The authors further demonstrate the platform through examples of in silico analyses and model benchmarking.

      Strengths:

      (1) Emphasis on standardization and reproducibility. Retinal modeling has become increasingly dependent on deep learning approaches, yet datasets and evaluation procedures remain fragmented across laboratories. By providing a unified framework, the authors lower barriers to entry and create opportunities for more systematic comparisons of models and datasets.

      (2) The manuscript is clearly written, and the examples effectively illustrate the range of analyses supported by the platform.

      (3) The benchmarking results are useful, particularly because they reveal substantial remaining gaps between current model performance and explainable variance ceilings.

      Weaknesses:

      Not a weakness per se, but rather a limitation, is that the manuscript focuses on software infrastructure rather than new biological or computational insights. While this is appropriate for a resource paper, some of the scientific examples, such as the gradient-field analysis of ON-OFF cells, function more as demonstrations than as rigorous validations of novel hypotheses. It might be useful to add a few sentences discussing potential scientific projects that can be immediately facilitated by the openretina (the current text in the Discussion focuses more on advancements in the technical/social aspects of science that will be supported by openretina).

      Overall, this is a valuable and timely resource that is likely to benefit the retinal and computational neuroscience communities.

    5. Author response:

      We thank the editors and reviewers for their thoughtful assessment of our manuscript, and for recognizing openretina as a valuable and timely resource for the retinal modelling community.

      We are especially glad that the reviewers appreciated the motivation of the project, the focus on standardization and reproducibility, and the potential of the platform to support systematic benchmarking and community-driven model development.

      We also understand the concerns raised. In the revision of the manuscript, we will strengthen the conceptual discussion of how predictive models, including the current “Core + Readout” models, can contribute to retinal neuroscience alongside more mechanistic and circuit-based approaches. This is a central matter for us, and one that some of us have recently addressed in a broader review on current trends in retina modelling (see https://doi.org/10.1016/j.visres.2026.108854). We will draw on this perspective to better articulate when predictive models are useful, where their limitations lie, and how openretina can provide infrastructure for comparing functional, normative and mechanistic models within a shared framework.

      We will also clarify the scope and limitations of the in-silico analysis methods provided within openretina. This will include a more explicit discussion of how MEIs, gradient-field analyses, and model-weight visualisations should be interpreted.

      Furthermore, we will add more information that will help the reader better judge different aspects of dataset quality, including, for example, spike-sorting or calcium-processing information and explainable-variance distributions. We note, however, that there are many subtle details about experimental workflows that are difficult to capture in compact indicators. In addition, we will make it clearer that the manuscript represents a snapshot of a living resource: The website, dataset cards, documentation, and repository will be the primary source of this information, especially as new datasets are contributed.

      Finally, we will of course address the technical clarifications raised by the reviewers, with the aim of making the manuscript more accessible overall.

      We are grateful for the reviewers’ constructive comments and believe that addressing these points will make our presentation of openretina clearer and more useful to the community.

    1. eLife Assessment

      This manuscript describes an important development of several variants of optogenetic tools to control endogenous p53 activity. They are based on peptides competing with Mdm2/MdmX for binding to p53, thus releasing p53 from its negative regulators and stabilizing its cellular levels. In principle, the data are convincing but should be complemented by investigations of p53 target genes at endogenous levels (instead of only reporter constructs). The study therefore remains incomplete but will be of interest to scientists working on optogenetics as well as the p53 field.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors apply the AsLOV2 domain to control the localisation and the exposure of two peptides (PMI and PMI-M3) that compete with Mdm2/MdmX for binding to p53, thus freeing p53 from these negative regulators and allowing its levels to rise. The authors follow an established strategy in optogenetics, which is to combine two layers of regulation for tighter control: (1) caging the peptide into the Ja helix of AsLOV2; 2) sequestration of the peptide away from its site of action using the LOVTRAP system.

      Strengths:

      The authors show that a reporter is activated when cells are exposed to light. A strength is in the lower background that was achieved after adding the second layer of regulation.

      Weaknesses:

      This study claims to be focused on the control of endogenous p53; however, endogenous p53 levels are not quantified. Moreover, endogenous p53 target genes are also not analysed. Only a synthetic reporter is quantified, which has been placed in the genome of HCT116 cells after the creation of a stable cell line. Microscopy images show only one or a maximum of two cells. Finally, the authors claim their strategy is a general one that can be applied to control other peptides, but they do not show this generality in this paper.

    3. Reviewer #2 (Public review):

      The authors developed Opto-MDMi, an optogenetic system for light-controlled activation of endogenous p53. The main idea is to target the p53-MDM2/MDMX regulatory interaction using PMI inhibitory peptides. This is a nice strategy because it avoids overexpression of p53, which can have adverse effects that might confound the study of p53 activity. The authors first tested a LOVTRAP-based localization strategy, which showed some efficacy but also showed basal activation. They then developed a LOV2-PMI peptide-caging module to control the activity of the PMI peptide itself, testing for interactions first in vitro and then in vivo. Finally, they combined the two systems into a dual-lock design, where LOVTRAP controls localization and LOV2-PMI controls peptide activity. This combination led to somewhat more potent stimulation of p53 activity.

      Another useful aspect of the paper is the detailed description of the development and testing of the LOV2-PMI peptide-caging module, which may aid in the design of other LOV2-based peptide-caging designs.

      Strengths

      Overall, the paper is novel and rigorous, and the claims are supported by the data. The optoMDMi tool seems ready for implementation, for example, to manipulate and study the role of p53 signaling dynamics. A few points of clarification would strengthen the work.

      Weaknesses

      The authors develop many tool variants, but there is some lack of clarity over how all of these tools compare to each other, and which ones interested users should use. The work would also be strengthened by showing modulation of endogenous p53 in more than one cell line.

    1. eLife Assessment

      Verma and colleagues interrogate the mechanisms of phagosome maturation arrest during Mycobacterium tuberculosis infection. While cellular events that culminate in this arrest have been largely elucidated, involvement of other organelles, such as mitochondria, has not been highlighted mechanistically. In this valuable study, elements of mitochondrial quality control, such as mitophagy and mitochondrial-derived vesicles involvement, are shown to be paramount in the host-pathogen tussle. The evidence supporting the main conclusions is solid, based on multiple complementary approaches and appropriate controls, although some central mechanistic aspects of the proposed pathway remain only partially resolved.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important and interesting manuscript that uncovers the cross-talk between mitochondrial quality control and phagosome maturation arrest imposed by Mtb.

      A broader host pathogen (intracellular) question pertains to evading phagosomal maturation/arrest. While cellular events that culminate in this arrest have been largely elucidated, involvement of other organelles, such as mitochondria, has not been highlighted mechanistically. This manuscript paints a larger picture than the well-known conventional endolysosomal pathway and portrays a larger landscape involving elements of the mitochondrial quality control, such as mitophagy and mitochondrial-derived vesicles' involvement in the host-pathogen tussle.

      Strengths:

      The systematic characterisation to unravel the interplay between mitochondrial-related pathways and the endolysosomal system allows the authors to unearth some important findings.

      Weaknesses:

      The conclusions drawn require more robust experimentation and analysis.

    3. Reviewer #2 (Public review):

      This manuscript examines the role of autophagy receptor proteins, particularly p62/SQSTM1, in regulating intracellular Mtb survival in human macrophages. Counterintuitively, depleting p62 reduces bacterial survival rather than enhancing it, pointing to a previously unrecognised mechanism. The authors demonstrate that in the absence of p62, mitochondrial quality is maintained through enhanced TOM20⁺ mitochondria-derived vesicle (MDV) biogenesis, dependent on MIRO1/MIRO2. During Mtb infection, these MDVs are redirected to bacterial phagosomes, promoting RAB7 recruitment, overcoming phagosome maturation arrest and facilitating lysosomal targeting of Mtb. In parallel, bacteria experience increased oxidative stress, further contributing to bacterial killing.

      Strengths:

      The mechanistic chain is built using multiple complementary approaches, including genetic perturbation, redox biosensors, metabolic assays and microscopy. The use of primary human macrophages from multiple donors alongside established cell lines increases confidence that the phenotype is not cell-line specific. The replication clock experiment is particularly elegant and clearly demonstrates that the reduction in bacterial burden reflects enhanced killing rather than impaired bacterial replication. Overall, the study identifies an unexpected connection between mitochondrial quality control and phagosome maturation and provides a potentially important advance in our understanding of host-pathogen interactions.

      Weaknesses:

      The study remains entirely in vitro, and the phenotype is absent in mouse macrophages, limiting the immediate physiological and translational relevance of the findings. In addition, many of the central mechanistic conclusions rely heavily on colocalisation analyses, making it difficult to distinguish direct mechanistic relationships from associated trafficking events.

      Overall, this is an interesting and technically strong study that uncovers a novel link between mitochondrial quality control and anti-mycobacterial defence. The mechanistic model is plausible and supported by substantial experimental work. However, several aspects of the proposed pathway require stronger experimental support before some of the broader conclusions can be fully justified.

      Major points

      (1) The central conclusion that TOM20⁺ MDVs are recruited to Mtb-containing phagosomes is based largely on microscopy and colocalisation analyses. Additional orthogonal approaches would strengthen this key aspect of the study and help establish the nature of the vesicles recruited to bacterial phagosomes.

      (2) The proposed mechanism whereby TOM20⁺ MDVs facilitate RAB7 recruitment and reverse phagosome maturation arrest remains incompletely demonstrated. While the MIRO1/2 and RAB7 knockdown experiments support the model, they do not directly establish a causal link between MDV recruitment and phagosomal RAB7 acquisition. Additional experiments addressing this step would considerably strengthen the manuscript.

      (3) The absence of a phenotype in mouse macrophages raises important questions regarding the conservation and physiological relevance of the proposed mechanism. The authors should discuss possible explanations for this species-specific effect and, if feasible, provide additional experimental insight into the basis of this difference.

      (4) The conclusion that mitochondrial quality is maintained despite impaired p62-dependent mitochondrial turnover is based primarily on mitochondrial content, membrane potential, ROS measurements and Seahorse analysis. These are informative but relatively indirect measurements. Additional assessment of mitochondrial turnover by mitophagy would strengthen this aspect of the study.

      (5) The proteins studied throughout the manuscript (p62/SQSTM1, NDP52, OPTN, TAX1BP1 and NBR1) are generally classified as selective autophagy receptors rather than adaptors. The terminology should be corrected throughout the manuscript.

      Minor points:

      (1) Several conclusions throughout the manuscript are based primarily on colocalisation analyses. The limitations of these approaches should be acknowledged explicitly.

      (2) The discussion would benefit from a clearer consideration of how the proposed mechanism relates to established pathways regulating phagosome maturation arrest during Mtb infection.

      (3) The authors may wish to comment on whether enhanced MDV biogenesis could represent a broader host defence mechanism against intracellular pathogens beyond Mtb.