10,000 Matching Annotations
  1. Last 7 days
    1. eLife Assessment

      In this manuscript, the authors describe a cell-specific mechanism by which glutamate transporters regulate the fidelity with which T-stellate cells in the mouse ventral cochlear nucleus relay information from auditory nerve inputs. The study is supported by solid electrophysiological data. It provides valuable insights into how the rapid binding of glutamate to transporters shapes auditory information processing at specific synapses.

    2. Reviewer #1 (Public review):

      In this article, the authors investigate how glutamate transporter function regulates excitability and synaptic coding in T-stellate cells in the mouse ventral cochlear nucleus. They test this in acute brain slices using whole-cell electrophysiology and artificially raise the relative local concentration of glutamate via pharmacological inhibition of transporter proteins. The main finding is that when sub-saturating doses of DL-TBOA are applied, cells become much more sensitive to synaptic input, diminishing the normally high fidelity of EPSP-spike coupling in these neurons. Notably, high-frequency stimulation in the presence of DL-TBOA reveals a large and slowly decaying AMPA receptor component that underlies persistent/rebound firing in earlier recordings. These effects are not seen in other ventral cochlear neurons, suggesting that rapid glutamate clearance in T-stellate cells, particularly, is important for auditory intensity coding. Overall, these experiments are well-performed, and the findings are robust, though there are some aspects that could be expanded to make the work more impactful. These include a better understanding of the relative contribution of neuronal vs glial transporters and an ability to separate the relative contributions of tonic glutamate concentrations in the cleft vs changes in membrane potential in action potential output. Additionally, there were some minor issues of clarity in both the figure presentation and the main text language that should be addressed.

      Major Points:

      (1) Given the dramatic effect of saturating DL-TBOA on tonic leak/RMP and that the sub-maximal concentration used in most of the experiments still varied between 25-50 uM, Figure 1 would be strengthened substantially by a dose-response curve. Ideally, 5 or 6 concentrations, plotting the effect on tonic current or RMP increase.

      (2) Examining the contribution of glial (EAAT1/2) vs. neuronal (EAAT3) transporters (Fig 8) is intriguing but comes across as incomplete here, especially given the small number of recordings. Using a different non-selective EAAT inhibitor (TFB-TBOA) to chase the EAAT1/2 blocker combo seems like an odd choice, given that you have already characterized the effects of DL-TBOA well. One could also try a lower concentration (~50-100 nM) of TFB-TBOA since it is somewhat selective itself for glial EAAT1/2. Given the data presented, neuronal transporters (presumably EAAT3) appear to dominate the rapid clearance of glutamate at this synapse, but this point isn't emphasized or explored sufficiently.

      (3) Separating the effects of depolarization vs. glutamate clearance was never explored. What effect does depolarizing the cell ~10 mV in control conditions (i.e., without TBOA) have on AP number/fidelity during synaptic stimulation experiments? The authors state that submaximal DL-TBOA generally causes no more than a 5 mV change in RMP, but tonic depolarization could also influence spike fidelity. This experiment could demonstrate that the increase in excitability during/after stimulation is not due to increased engagement of voltage-gated channels.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and mechanistically interesting question: whether plasma membrane glutamate transporters contribute only to slow clearance of ambient glutamate or whether they can rapidly shape synaptic signaling during high-frequency auditory activity. This manuscript provides important evidence that EAAT-mediated glutamate uptake is not merely a slow background clearance mechanism but is essential for maintaining reliable synaptic transmission and linear stimulus-intensity coding in ventral cochlear nucleus T-stellate cells during sustained auditory nerve activity.

      Strengths:

      The finding that EAATs may be required for rapid, local control of glutamate during high-frequency auditory nerve activity is interesting and could have broad relevance to auditory processing. The electrophysiological evidence is generally strong, particularly the use of patch-clamp recordings, stimulus trains, partial versus complete EAAT blockade, and comparison with bushy cell/endbulb synapses. The comparison between T-stellate cells and bushy cells/endbulb synapses strengthens the manuscript. The authors demonstrate that EAAT blockade disrupts coding in T-stellate cells but has little effect on bushy cell spike transmission, supporting a cell-type- and synapse-specific role of glutamate uptake.

      Weaknesses:

      However, some mechanistic conclusions, especially the specific contribution of neuronal versus glial EAATs and the absence of glutamate crosstalk between auditory nerve inputs, rely mainly on pharmacological and indirect electrophysiological inference and would be strengthened by additional anatomical, genetic, or direct glutamate-sensing evidence.

      (1) Clarification of DL-TBOA concentration.

      The authors used bath application of 200 µM TBOA and 25-50 µM in the other experiments, stating that "sub-maximal concentrations (25-50 µM)". The authors should provide a clearer rationale for why different concentrations were used across experiments rather than a fixed concentration.

      The reversibility of DL-TBOA effects should be demonstrated by washout experiments. In addition, potential off-target effects of DL-TBOA on postsynaptic receptors, intrinsic membrane excitability, or presynaptic release (e.g., PPR measurement) should be carefully considered. It would also be useful to test the effects of the submaximal DL-TBOA concentrations (25-50 µM) on membrane potential and inward currents, shown in Figure 1, to determine whether these concentrations depolarize the membrane potential in current-clamp mode or induce inward currents under voltage-clamp conditions.

      (2) Potential contribution of altered intrinsic excitability.

      In Figures 3B and 3C, DL-TBOA appears to induce additional action potentials even immediately after the first stimulation, whereas Figures 6 and 7 suggest that the first EPSC is not substantially altered. This raises the possibility that the enhanced firing may partly result from a modest depolarization caused by background glutamate accumulation or from other changes in intrinsic membrane properties after drug treatment. To address this, the authors should provide a quantitative analysis of physiological parameters under submaximal DL-TBOA conditions, including spontaneous action potential frequency, resting membrane potential, input resistance, and spike threshold.

      (3) Spillover/ crosstalk between AN-fiber-synpases.

      The authors should provide more explanation of how altering the number of active auditory nerve fibers demonstrates the absence of glutamate spillover/crosstalk between bouton synapses. Strong stimulation likely recruits more AN fibers, but it may also change release probability, axonal synchrony, or stimulation spread. The authors should more clearly justify the interpretation that strong stimulation recruits additional independent AN fibers rather than altering release probability or activating fibers with different intrinsic properties.

      (4) Interpretation of glial versus neuronal EAAT contributions.

      The authors claim that both neuronal and glial transporters contribute to rapid uptake using pharmacological approaches. The pharmacological data demonstrate that glial EAATs play a major role in glutamate clearance at T-stellate cell synapses. The strong increase in EPSC decay time and synaptic charge after UCPH-101/DHK application supports the conclusion that glial transporters contribute substantially to limiting glutamate accumulation during sustained auditory nerve activity. However, the conclusion that neuronal EAATs contribute directly should be stated with some caution. The evidence for neuronal EAAT involvement is indirect and depends on the pharmacological specificity and completeness of glial EAAT blockade. The conclusion would be strengthened by additional evidence, such as EAAT subtype expression/localization in T-stellate cells or auditory nerve terminals, transporter current recordings, immunohistochemistry, or genetic manipulation of neuronal EAATs. In addition, fitting the decay phase with a double-exponential model may help determine whether glial and neuronal EAATs contribute over distinct temporal windows.

    1. eLife Assessment

      This manuscript describes an important development of several variants of optogenetic tools to control endogenous p53 activity. They are based on peptides competing with Mdm2/MdmX for binding to p53, thus releasing p53 from its negative regulators and stabilizing its cellular levels. In principle, the data are convincing but should be complemented by investigations of p53 target genes at endogenous levels (instead of only reporter constructs). The study therefore remains incomplete but will be of interest to scientists working on optogenetics as well as the p53 field.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors apply the AsLOV2 domain to control the localisation and the exposure of two peptides (PMI and PMI-M3) that compete with Mdm2/MdmX for binding to p53, thus freeing p53 from these negative regulators and allowing its levels to rise. The authors follow an established strategy in optogenetics, which is to combine two layers of regulation for tighter control: (1) caging the peptide into the Ja helix of AsLOV2; 2) sequestration of the peptide away from its site of action using the LOVTRAP system.

      Strengths:

      The authors show that a reporter is activated when cells are exposed to light. A strength is in the lower background that was achieved after adding the second layer of regulation.

      Weaknesses:

      This study claims to be focused on the control of endogenous p53; however, endogenous p53 levels are not quantified. Moreover, endogenous p53 target genes are also not analysed. Only a synthetic reporter is quantified, which has been placed in the genome of HCT116 cells after the creation of a stable cell line. Microscopy images show only one or a maximum of two cells. Finally, the authors claim their strategy is a general one that can be applied to control other peptides, but they do not show this generality in this paper.

    3. Reviewer #2 (Public review):

      The authors developed Opto-MDMi, an optogenetic system for light-controlled activation of endogenous p53. The main idea is to target the p53-MDM2/MDMX regulatory interaction using PMI inhibitory peptides. This is a nice strategy because it avoids overexpression of p53, which can have adverse effects that might confound the study of p53 activity. The authors first tested a LOVTRAP-based localization strategy, which showed some efficacy but also showed basal activation. They then developed a LOV2-PMI peptide-caging module to control the activity of the PMI peptide itself, testing for interactions first in vitro and then in vivo. Finally, they combined the two systems into a dual-lock design, where LOVTRAP controls localization and LOV2-PMI controls peptide activity. This combination led to somewhat more potent stimulation of p53 activity.

      Another useful aspect of the paper is the detailed description of the development and testing of the LOV2-PMI peptide-caging module, which may aid in the design of other LOV2-based peptide-caging designs.

      Strengths

      Overall, the paper is novel and rigorous, and the claims are supported by the data. The optoMDMi tool seems ready for implementation, for example, to manipulate and study the role of p53 signaling dynamics. A few points of clarification would strengthen the work.

      Weaknesses

      The authors develop many tool variants, but there is some lack of clarity over how all of these tools compare to each other, and which ones interested users should use. The work would also be strengthened by showing modulation of endogenous p53 in more than one cell line.

    1. eLife Assessment

      Verma and colleagues interrogate the mechanisms of phagosome maturation arrest during Mycobacterium tuberculosis infection. While cellular events that culminate in this arrest have been largely elucidated, involvement of other organelles, such as mitochondria, has not been highlighted mechanistically. In this valuable study, elements of mitochondrial quality control, such as mitophagy and mitochondrial-derived vesicles involvement, are shown to be paramount in the host-pathogen tussle. The evidence supporting the main conclusions is solid, based on multiple complementary approaches and appropriate controls, although some central mechanistic aspects of the proposed pathway remain only partially resolved.

    2. Reviewer #1 (Public review):

      Summary:

      This is an important and interesting manuscript that uncovers the cross-talk between mitochondrial quality control and phagosome maturation arrest imposed by Mtb.

      A broader host pathogen (intracellular) question pertains to evading phagosomal maturation/arrest. While cellular events that culminate in this arrest have been largely elucidated, involvement of other organelles, such as mitochondria, has not been highlighted mechanistically. This manuscript paints a larger picture than the well-known conventional endolysosomal pathway and portrays a larger landscape involving elements of the mitochondrial quality control, such as mitophagy and mitochondrial-derived vesicles' involvement in the host-pathogen tussle.

      Strengths:

      The systematic characterisation to unravel the interplay between mitochondrial-related pathways and the endolysosomal system allows the authors to unearth some important findings.

      Weaknesses:

      The conclusions drawn require more robust experimentation and analysis.

    3. Reviewer #2 (Public review):

      This manuscript examines the role of autophagy receptor proteins, particularly p62/SQSTM1, in regulating intracellular Mtb survival in human macrophages. Counterintuitively, depleting p62 reduces bacterial survival rather than enhancing it, pointing to a previously unrecognised mechanism. The authors demonstrate that in the absence of p62, mitochondrial quality is maintained through enhanced TOM20⁺ mitochondria-derived vesicle (MDV) biogenesis, dependent on MIRO1/MIRO2. During Mtb infection, these MDVs are redirected to bacterial phagosomes, promoting RAB7 recruitment, overcoming phagosome maturation arrest and facilitating lysosomal targeting of Mtb. In parallel, bacteria experience increased oxidative stress, further contributing to bacterial killing.

      Strengths:

      The mechanistic chain is built using multiple complementary approaches, including genetic perturbation, redox biosensors, metabolic assays and microscopy. The use of primary human macrophages from multiple donors alongside established cell lines increases confidence that the phenotype is not cell-line specific. The replication clock experiment is particularly elegant and clearly demonstrates that the reduction in bacterial burden reflects enhanced killing rather than impaired bacterial replication. Overall, the study identifies an unexpected connection between mitochondrial quality control and phagosome maturation and provides a potentially important advance in our understanding of host-pathogen interactions.

      Weaknesses:

      The study remains entirely in vitro, and the phenotype is absent in mouse macrophages, limiting the immediate physiological and translational relevance of the findings. In addition, many of the central mechanistic conclusions rely heavily on colocalisation analyses, making it difficult to distinguish direct mechanistic relationships from associated trafficking events.

      Overall, this is an interesting and technically strong study that uncovers a novel link between mitochondrial quality control and anti-mycobacterial defence. The mechanistic model is plausible and supported by substantial experimental work. However, several aspects of the proposed pathway require stronger experimental support before some of the broader conclusions can be fully justified.

      Major points

      (1) The central conclusion that TOM20⁺ MDVs are recruited to Mtb-containing phagosomes is based largely on microscopy and colocalisation analyses. Additional orthogonal approaches would strengthen this key aspect of the study and help establish the nature of the vesicles recruited to bacterial phagosomes.

      (2) The proposed mechanism whereby TOM20⁺ MDVs facilitate RAB7 recruitment and reverse phagosome maturation arrest remains incompletely demonstrated. While the MIRO1/2 and RAB7 knockdown experiments support the model, they do not directly establish a causal link between MDV recruitment and phagosomal RAB7 acquisition. Additional experiments addressing this step would considerably strengthen the manuscript.

      (3) The absence of a phenotype in mouse macrophages raises important questions regarding the conservation and physiological relevance of the proposed mechanism. The authors should discuss possible explanations for this species-specific effect and, if feasible, provide additional experimental insight into the basis of this difference.

      (4) The conclusion that mitochondrial quality is maintained despite impaired p62-dependent mitochondrial turnover is based primarily on mitochondrial content, membrane potential, ROS measurements and Seahorse analysis. These are informative but relatively indirect measurements. Additional assessment of mitochondrial turnover by mitophagy would strengthen this aspect of the study.

      (5) The proteins studied throughout the manuscript (p62/SQSTM1, NDP52, OPTN, TAX1BP1 and NBR1) are generally classified as selective autophagy receptors rather than adaptors. The terminology should be corrected throughout the manuscript.

      Minor points:

      (1) Several conclusions throughout the manuscript are based primarily on colocalisation analyses. The limitations of these approaches should be acknowledged explicitly.

      (2) The discussion would benefit from a clearer consideration of how the proposed mechanism relates to established pathways regulating phagosome maturation arrest during Mtb infection.

      (3) The authors may wish to comment on whether enhanced MDV biogenesis could represent a broader host defence mechanism against intracellular pathogens beyond Mtb.

    1. eLife Assessment

      This valuable study provides a cross-species single-cell transcriptomic resource for early female gonadal development in mammals. The data supporting the main conclusion remain incomplete, and experimental validation is needed to strengthen the conclusions. The work will be of interest to reproductive biologists and developmental biologists.

    2. Reviewer #1 (Public review):

      Summary:

      Fang et al. characterize the cellular basis of early ovarian development through a comparative analysis of single-cell transcriptomic data. The authors integrate a novel bovine scRNA-seq dataset, spanning six gestational stages (E38-E112), with stage-matched human (PCW6-16) and mouse (E11.5-E18.5) counterparts. Beyond identifying shared gonadal cell types across these three species, the study uncovers a previously uncharacterized bovine-specific cell population with steroidogenic features. Their analysis highlights conserved, dynamically expressed regulators, including TFAP2C and ZCWPW1 in germ cells and FOS and JUNB in granulosa cells. Furthermore, by employing a machine learning Support Vector Machine (SVM) model, the authors quantify cell-type conservation, demonstrating that while immune and germ cells are highly conserved across species, granulosa cells exhibit substantial evolutionary divergence. This study makes a significant contribution to developmental biology by establishing a comprehensive, cross-species single-cell roadmap of fetal ovarian development. By integrating livestock data with human and rodent models, the authors identify novel cellular states and provide a framework for assessing transcriptional conservation across species.

      Strengths:

      (1) While human and mouse fetal ovaries have been mapped, the inclusion of a high-resolution bovine dataset (107,930 cells total across the study) provides a critical "large mammal" perspective that is often missing from comparative studies.

      (2) The identification of a bovine-specific cell population is an important finding. It suggests that ruminants may have a different developmental timeline for steroidogenic precursors (potentially theca cell ancestors) compared to rodents or humans.

      (3) Training a Support Vector Machine (SVM) to quantitatively assess cell-type similarity is a major strength. It moves beyond qualitative UMAP "eye-balling" to provide a statistical probability of conservation.

      (4) The study links gene expression to higher-order biological processes like epigenetic reprogramming and cell-cell communication (CellChat), providing a holistic view of the gonadal niche.

      Weaknesses:

      (1) The authors integrated publicly available scRNA-seq datasets generated across different laboratories and technical platforms. However, the specific methods used to control for and evaluate batch effects are not clearly described. It is critical to clarify whether the observed species-specific differences are purely biological or partly influenced by technical variation between datasets.

      (2) A challenge inherent to all single-cell studies is the reliance on manual marker-gene-based annotation. While this is standard practice, it remains unclear how robust these assignments are, particularly for the novel "bovine-specific" population. Further evidence or cross-validation (e.g., through varied clustering resolutions or automated annotation tools) is required to ensure these clusters represent true biological states rather than computational artifacts.

      (3) The authors utilized a linear SVM to assess cross-species similarity. However, it is not clear how this model performs compared to established single-cell mapping and comparative tools (e.g., MetaNeighbor or Seurat v5). Providing a justification for this specific SVM-based approach, or a brief comparison with existing benchmarks, would strengthen the methodological rigor of the study.

      (4) While the computational evidence is compelling, the study would be significantly enhanced by independent validation of the "unclassified bovine-specific" cell population. To confirm the biological reality and reproducibility of this novel cell state, the authors should provide additional evidence. This could include in situ validation (e.g., immunofluorescence or in situ hybridization) to determine its physical location and morphology within the gonad, or demonstrating the presence of this specific cell population within an independent, non-overlapping bovine dataset.

    3. Reviewer #2 (Public review):

      Summary:

      The authors generate a comparative single-cell transcriptomic atlas of fetal ovarian development in cattle, human, and mouse, with the goal of identifying conserved and species-specific cellular and molecular features of early ovarian differentiation. The study provides a valuable resource for the field and reveals potentially interesting species-specific characteristics, including a putative bovine steroidogenic cell population. While the dataset is substantial and the computational analyses are generally appropriate, several major conclusions rely primarily on computational inference without independent experimental validation, limiting the strength of evidence supporting some of the central claims.

      Strengths:

      This study provides a valuable cross-species single-cell atlas of fetal ovarian development by integrating newly generated bovine data with human and mouse datasets. The work fills an important gap in reproductive biology and offers a useful resource for investigating conserved and species-specific features of ovarian development.

      The analyses are comprehensive and combine developmental trajectory reconstruction, regulatory network inference, cell-cell communication analysis, and cross-species classification. The identification of a putative bovine-specific steroidogenic cell population is particularly intriguing and may provide a basis for future studies of species-specific ovarian development.

      Weaknesses:

      The main limitation is that several key conclusions rely primarily on computational analyses without independent experimental validation. In particular, the proposed bovine-specific steroidogenic cell population, which represents the major novel finding of the study, is supported only by transcriptomic evidence.

      In addition, many mechanistic interpretations derived from trajectory, regulatory network, and cell-cell communication analyses remain speculative. While the study succeeds as a comparative resource, the evidence supporting several of the central biological claims remains incomplete, and the biological significance of some cross-species differences is not fully explored.

    1. eLife Assessment

      The Review Article by Bal and co-workers presents an overview of skeletal muscle physiology, focusing on Sarcoplasmic Reticulum and Mitochondria-Associated Membranes (MAMs) in relation to calcium handling, ROS, and signaling. It provides a foundation based on the current literature but could have gone further by identifying future research directions and potential avenues for therapeutic intervention.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript is a narrative review addressing age-related alterations in SR-mitochondria interactions in skeletal muscle and their contribution to sarcopenia. It synthesizes existing literature on calcium signaling, mitochondrial dynamics, redox balance, and structural remodeling, and discusses potential interventions including exercise and pharmacological strategies. While the topic is timely and relevant, the manuscript largely reiterates established concepts without providing sufficient conceptual novelty, critical synthesis, or mechanistic insight beyond the current literature.

      Strengths:

      (1) Timely topic: The focus on SR-mitochondria communication in aging muscle is relevant and of growing interest.

      (2) Broad coverage: The review compiles a wide range of literature spanning calcium handling, mitochondrial biology, ROS signaling, and exercise physiology.

      (3) Clear organization: The manuscript is structured logically with thematic sections (SR, mitochondria, MAMs, aging, interventions).

      (4) Didactic value: Could serve as a general overview for non-specialists entering the field.

      Weaknesses:

      (1) Lack of novelty and conceptual advance: The manuscript does not offer new hypotheses, frameworks, or critical reinterpretation of the field. Most statements summarize already well-established knowledge, and no unifying model or novel perspective is developed to justify publication in a high-impact journal like eLife.

      (2) Limited critical analysis: The review is predominantly descriptive rather than analytical. Conflicting findings (e.g., MFN2 roles, MAM density changes, Ca²⁺ overload vs deficiency) are mentioned but not critically evaluated or reconciled. There is little discussion of limitations in the cited studies or gaps in the field.

      (3) Overgeneralization and speculative claims: Several assertions are presented with insufficient nuance (e.g., causal links between MAM disruption and sarcopenia, or therapeutic efficacy of interventions). The distinction between correlation and causation is often unclear, reducing scientific rigor.

      (4) Insufficient depth for a specialist audience: Despite its length, the manuscript lacks mechanistic depth in key areas (e.g., precise molecular regulation of MAMs in vivo, tissue-specific differences, quantitative aspects of Ca²⁺ flux). It reads more like a textbook summary than a high-level scholarly review.

      (5) Redundancy and verbosity: Many sections repeat similar concepts (Ca²⁺ dysregulation, ROS effects, mitochondrial dysfunction) without adding new insight, leading to an unnecessarily long manuscript with limited added value.

      (6) Weak integration of recent literature into a coherent narrative: Although many references are cited, they are not effectively synthesized into a cohesive argument. The manuscript lacks a strong central thesis or clearly defined take-home messages.

      (7) Limited translational or experimental perspective: The section on therapeutic targeting is largely speculative and does not critically assess feasibility, limitations, or current clinical evidence.

    3. Reviewer #2 (Public review):

      This review addresses a highly relevant and timely topic, namely the role of sarcoplasmic reticulum-mitochondria communication and mitochondria-associated membranes (MAMs) in skeletal muscle aging. The manuscript successfully brings together literature from several interconnected fields, including calcium signaling, mitochondrial biology, excitation-contraction coupling, muscle metabolism, and sarcopenia. Given the growing interest in organelle crosstalk as a determinant of muscle health and disease, the topic is undoubtedly of considerable interest to the readership and has the potential to make a valuable contribution to the field.

      However, in its current form, the manuscript devotes a substantial proportion of its content to the description of well-established concepts that are already extensively covered in the literature. Large sections are dedicated to general skeletal muscle physiology, excitation-contraction coupling, calcium handling, mitochondrial biology, and MAM structure and composition. While this background information is useful, the level of detail is often excessive for a review that aims to focus on aging-induced alterations in SR-mitochondria interactions. As a consequence, the central theme of the manuscript becomes diluted, and the review reads more like a broad overview of skeletal muscle physiology than a focused analysis of aging-related MAM remodeling.

      In contrast, the sections specifically dedicated to aging and MAM dysfunction, which represent the most novel and potentially impactful aspects of the review, are comparatively brief and largely descriptive. The discussion of how aging alters MAM architecture, calcium microdomains, mitochondrial calcium signaling, and organelle communication would benefit from substantially greater depth. For example, although the manuscript highlights alterations in proteins such as MFN2, IP3R, VDAC, and MCU, the mechanistic implications of these changes for sarcopenia and age-associated muscle dysfunction are not critically developed. Similarly, the review would be strengthened by a more comprehensive discussion of the evidence linking MAM disruption to impaired muscle performance, metabolic inflexibility, denervation, and mitochondrial dysfunction during aging.

      Another limitation is that much of the manuscript summarizes published findings without sufficiently evaluating the strength of the evidence or discussing existing controversies. Several statements imply causal relationships between MAM disruption and sarcopenia, whereas in many cases, the available data remain largely correlative. The authors should more clearly distinguish between established mechanisms, experimental observations, and emerging hypotheses. A more critical assessment of conflicting findings, particularly regarding the role of MFN2 and the dual consequences of altered mitochondrial calcium uptake, would considerably improve the scientific rigor of the review.

      A major omission concerns the role of mitochondrial Ca²⁺ uptake in skeletal muscle physiology and aging. Throughout the manuscript, mitochondrial Ca²⁺ uptake is presented as a central determinant of muscle function and as a key mechanism linking MAM disruption to sarcopenia. However, the authors do not adequately discuss evidence that challenges this view. In particular, genetic mouse models lacking MCU exhibit surprisingly mild skeletal muscle phenotypes under basal conditions despite a near-complete abolition of rapid mitochondrial Ca²⁺ uptake. These findings have generated considerable debate regarding the physiological importance of mitochondrial Ca²⁺ uptake for muscle function and metabolic regulation. While MCU deletion clearly affects exercise adaptation and certain stress responses, the relatively modest baseline phenotype suggests the existence of compensatory pathways and raises important questions regarding the extent to which impaired mitochondrial Ca²⁺ uptake alone can explain age-associated muscle dysfunction. A balanced review should acknowledge these observations and discuss the ongoing debate regarding the relative contributions of mitochondrial Ca²⁺ deficiency versus mitochondrial Ca²⁺ overload in aging skeletal muscle.

      Similarly, the discussion of MFN2 would benefit from greater nuance. The manuscript largely presents MFN2 as a structural tether linking the sarcoplasmic reticulum and mitochondria. However, MFN2 is a multifunctional protein with well-established roles in mitochondrial fusion, mitochondrial network organization, mitophagy regulation, and metabolic signaling. Consequently, many of the phenotypes associated with altered MFN2 expression cannot be unequivocally attributed to changes in MAM formation. The review does not sufficiently distinguish between the effects of MFN2 on organelle tethering and its effects on mitochondrial dynamics. This distinction is particularly important because several studies have questioned whether MFN2 acts primarily as a positive tether, a negative regulator of contacts, or whether its influence on organelle communication is secondary to its role in controlling mitochondrial morphology. As a result, attributing age-related alterations in SR-mitochondria communication solely to changes in MFN2-mediated tethering may oversimplify a considerably more complex biological scenario.

      The manuscript's organization could also be improved. The sections discussing aging-related alterations, mitochondrial dysfunction, calcium dysregulation, oxidative stress, and therapeutic interventions contain significant overlap and repetition. Streamlining some background sections and reallocating space to a more detailed discussion of aging-specific mechanisms would help maintain focus and improve readability. In particular, the manuscript would benefit from expanding the sections on aging-induced MAM remodeling, age-dependent changes in MAM composition and ultrastructure, and the potential of MAM-targeted interventions as therapeutic strategies for sarcopenia.

      Finally, the review would gain from a stronger future perspectives section. Several important questions remain unresolved, including whether MAM disruption is a primary driver of muscle aging or a secondary consequence of mitochondrial dysfunction, how MAM architecture differs among muscle fiber types during aging, and whether MAM-associated proteins could serve as reliable biomarkers or therapeutic targets in human sarcopenia. Highlighting these knowledge gaps would further enhance the review's impact.

      Overall, the manuscript covers an important and emerging area of research and contains a valuable compilation of the relevant literature. Nevertheless, substantial revision is required to reduce the emphasis on well-established background information, deepen and critically analyze the aging-specific sections, and provide a more focused discussion of the role of MAMs in skeletal muscle aging and sarcopenia.

    1. eLife Assessment

      This study investigates the cellular mechanisms underlying theta-nested gamma oscillations in the medial entorhinal cortex; the experiments are rigorous, and the analyses and modeling provide potentially useful insights into cell-type-dependent circuit dynamics. However, the evidence supporting several key conclusions remains incomplete. The study is limited by conceptual constraints in the experimental design and a modeling approach that does not fully address underlying physiological mechanisms. Overall, this is a careful study that addresses how distinct neuronal populations in superficial MEC participate in theta-gamma coordination and provides new data linking cell-type-specific activity patterns to oscillatory network structure.

    2. Reviewer #1 (Public review):

      Summary:

      The question posed on cell-type-dependent relationships to theta-nested gamma rhythms is an important one. The authors use a variety of ontogenetic, imaging, electrophysiology, and computational techniques to show that reciprocal interactions between excitatory neurons and interneurons in the medial entorhinal cortex generate gamma oscillations. They measure LFP gamma, gamma power of postsynaptic currents in different neurons, spike phases with reference to LFP gamma, and spatial correlations of membrane potentials across a large population of neurons. Arguing (correctly) that gamma rhythm in this setting is generated through a pyramidal-interneuron network gamma (PING) mechanism, they demonstrate cell-type-specific differences in gamma phase-locking. While they show spatial dependencies of sub-threshold voltages and even argue for topographic clustering, these could simply be reflections of the synchronous stimulation paradigm that they use.

      Overall, I appreciate the methodology and rigor, but would have expected more from the study in terms of relevance to physiological stimulation conditions as well as in terms of mechanisms underlying the differences that they report here..

      Strengths:

      The authors are rigorous in how they conduct the experiments, report the data, and perform the analyses. The modeling respects the heterogeneities and is truthful to the experimental design. The conclusions on PING mechanisms are fine, but are not unexpected given the circuitry of the mEC.

      Weaknesses:

      The interpretation of the conclusions, while for the most part is fine, could have been better, especially given the conceptual limitations of the experimental design. The modeling part could have gone beyond simple descriptive matching and addressed mechanistic questions.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors studied the cellular mechanism of theta-nested gamma oscillations in the medial entorhinal cortex (MEC) in vitro. The theta-nested gamma activity was induced by theta-modulated optogenetic stimulation of CaMKII+ neurons. In Figures 1 through 4, they describe the firing phase, synaptic input, and LFP-IPSC coupling of stellate cells, pyramidal cells, and interneurons. They then conducted voltage imaging, capturing the simultaneous activity of 41 cells, and found that subthreshold membrane potentials cluster in a weakly distance-dependent manner (Figure 5). The experiments and analysis are done rigorously for the most part.

      However, the results described in Figures 1 to 4 are largely descriptive and highly similar to those in their recent publication, which utilized almost identical experiments. While the voltage imaging data during theta-nested gamma oscillations are novel, the authors report data from only a single experiment, leaving it unclear whether the results are reproducible. Furthermore, without a comparison to in vivo data, it remains unclear what novel insights this manuscript provides to advance our understanding of the cellular mechanisms underlying theta-nested gamma oscillations.

      (1) The authors recently published another paper on the topic of theta-nested gamma oscillations in the MEC (Williams et al., eNeuro, 2026). In that study, they utilized a Thy1 promoter instead of the CaMKII promoter used here. The motivation for testing the CaMKII promoter in the current manuscript, as well as the novel insights expected from this experimental setup, remains unclear. Given that existing literature suggests inhibitory MEC cells play a critical role in theta activity (e.g., Gonzalez-Sulser et al., 2014)-implying that theta modulation should drive inhibitory rather than excitatory cells-the previous use of the Thy1 promoter appears closer to in vivo conditions than the CaMKII promoter used here.

      The overall conclusion of the current manuscript is that excitatory-inhibitory (E-I) interactions dominate the generation of theta-nested gamma oscillations. However, in their previous eNeuro paper, the authors demonstrated that the interneuron network gamma (ING) mechanism can sustain gamma oscillations without excitatory synaptic transmission. It seems expected that excitatory cells would be involved when the optogenetic stimulation selectively drives excitatory cells. If CaMKII stimulation is less physiological and artificially forces the theta-nested gamma activity to rely on excitatory connections, this conclusion could be misleading. It may potentially describe a mechanism that is irrelevant to physiological processes in vivo. Please see my comment 3, which is related to this point.

      In addition, Figures 1 and 2 heavily overlap with the authors' previous eNeuro publication. The differences in experimental settings and the motivation for performing almost identical experiments must be clearly articulated prior to these figures to avoid confusion. The authors must also justify why it is necessary to present such similar data, and explicitly point out the novel findings in the current paper compared to their previous work.

      (2) Using voltage imaging to investigate theta-nested gamma oscillations is novel. However, the impact of the findings from this experiment appears minimal in the manuscript's current state. The most novel and interesting observation is likely presented in Figure 6, where the authors identified clustered voltage correlations. However, this appears to be an n=1 experiment, and these findings should be replicated at least in a few experiments. Furthermore, the manuscript lacks a discussion or interpretation of this observation, making it unclear whether the result is biologically meaningful. Please find specific suggestions regarding this point below.

      (3) The authors' primary motivation for investigating the mechanisms underlying theta-modulated gamma oscillations is their potential role in grid cell firing. Therefore, it is critical that the mechanisms studied here in vitro accurately reflect in vivo processes. For this reason, greater effort should be made to better link this in vitro study with existing in vivo data. Numerous public in vivo datasets are available that detail the firing activity of putative principal cells and interneurons during exploratory behavior in mice. Intracellular recordings in awake animals have also been published, some of which the authors already cite. The data presented in Figures 1 and 4, for example, could be straightforwardly compared with those existing in vivo metrics. Furthermore, available in vivo silicon probe recordings could provide a reliable estimate of the spatial distribution of gamma-related spike activity. Such data should be compared with the voltage imaging results presented in this study.

      This limitation connects back to the first point. In this manuscript, the authors tested a different method for inducing theta-nested gamma oscillations (via the CaMKII promoter) than in their recent eNeuro paper (via the Thy1 promoter). The outcomes of these two induction methods must be systematically compared against in vivo data to determine which approach aligns more closely with physiological conditions. Without such a comparison, the scientific justification for testing a different promoter in this study remains unclear.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Williams et al. combine optogenetics, whole-cell electrophysiology, local field potential recordings, large-scale voltage imaging, and computational modeling to investigate the cellular and circuit mechanisms underlying theta-nested gamma oscillations in superficial medial entorhinal cortex (mEC). The authors propose that fast-spiking interneurons receive strong gamma-frequency excitatory drive and provide rhythmic inhibition onto principal neurons, supporting a pyramidal-interneuron network gamma (PING) mechanism. They further report cell-type-specific differences in gamma phase locking, spatial clustering of subthreshold voltage signals, and a network model reproducing several observed features, including interneuron bursting and gamma-cycle skipping in excitatory neurons.

      Strengths:

      The study is technically sophisticated and addresses an important question in entorhinal circuit function. The combination of intracellular recordings, voltage imaging, and computational modeling is a clear strength.

      Weaknesses:

      Several key conclusions developed from experimental results require additional raw data, statistical support, clearer methodological description, and more cautious interpretation. The computational modeling focuses primarily on stellate cells, whereas the experimental results suggest an important role for pyramidal neurons in PING dynamics. This creates inconsistency between theory and experiments.

    5. Author response:

      We thank the editors and reviewers for their thoughtful comments. Below, we list our provisional responses to the reviewers’ major points:

      On the rationale for CaMKIIα versus Thy1-driven stimulation and physiological relevance: We agree that we did not make clear the motivation for using CaMKIIα-driven stimulation, distinct from the Thy1-driven paradigm in our previous work (Williams et al., 2026). Using the Thy1 driver, both excitatory and inhibitory cells received direct theta drive. In contrast, CaMKIIα expression is largely restricted to principal neurons. Comparing these models lets us isolate a "driven I-cell" PING mechanism from the "E cell recovers first" mechanism relevant when interneurons are also directly driven.

      Regarding physiological relevance, Gonzalez-Sulser et al. (2014) found that septal GABAergic projections selectively and directly inhibit mEC interneurons, rather than exciting either principal cells or interneurons, implying that theta drive in vivo likely acts through rhythmic disinhibition of interneurons rather than direct excitation of any cell type. Neither the Thy1 nor the CaMKIIα paradigm reproduces this disinhibitory mechanism: both rely on excitatory optogenetic drive rather than rhythmic inhibition of interneurons, and replicating the natural drive (tonic excitatory tone plus rhythmic, interneuron-selective inhibition) is technically difficult in acute slices, which are largely quiescent without exogenous stimulation. We therefore view CaMKIIα and Thy1 as complementary approximations, each isolating a different circuit interaction. If forced to choose, we’d argue that the CaMKIIα is a better model of disinhibition of excitatory neurons. We will revise the Discussion regarding this point.

      On reproducibility of the voltage imaging findings: We thank the reviewer for this comment and agree that clarification is warranted.

      The voltage imaging dataset combines two levels of analysis with different sample sizes. The population-level firing and spike-correlation analyses (Fig. 5F–H) are pooled across multiple imaging sessions (n = 240 neurons). The spatial clustering analysis of subthreshold voltage correlations (Fig. 6, and the corresponding example traces in Fig. 5A–E) are drawn from a single representative recording session, as the reviewer correctly notes. We have voltage imaging data from 14 fields of view (1 FOV per slice) across 6 mice (240 neurons total; 3–41 neurons per FOV). In revision, we will extend the clustering and spatial-correlation analysis from Fig. 6 across sessions to assess whether the reported organization is reproducible, rather than relying on a single example. We will also revise the text to distinguish clearly which analyses are single-session versus pooled.

      On restricting the computational model of excitatory neurons to stellate cells: We modeled stellate cells as the excitatory population because they are the principal cells reciprocally connected to fast-spiking PV+ interneurons (Fuchs et al., 2016), the interneuron class most directly implicated in theta-nested gamma. Pyramidal cells, by contrast, are primarily connected via 5-HT3a-positive interneurons (Fuchs et al., 2016), with the exception of a subset of "intermediate" pyramidal cells that do show reciprocal PV+ connectivity. Our model, which captures the full measured heterogeneity of stellate cell and PV+ interneuron intrinsic properties and their reciprocal connectivity, is, to our knowledge, the most biophysically constrained implementation of this specific microcircuit to date. Incorporating the PV+-connected intermediate pyramidal population is a natural next step. Because this refinement, which requires more experimental data, is nontrivial and beyond the scope of this study, we will note this explicitly as a limitation of the current model in the revised Discussion.

      In vivo comparison (temporal/phase-locking): We agree that grounding our findings in existing in vivo data strengthens the study and will add these comparisons to the revision.

      Our whole-cell recordings reproduce the temporal organization in vivo and provide further insights into cell-type differences between the principal cells. All cell types were strongly phase-locked to theta, while gamma phase-locking declined across successive spikes, with stellate cells decoupling after the first spike and pyramidal cells after the second. This earlier decoupling in stellate cells may contribute to their weaker theta rhythmicity reported in freely moving rats (Ray et al., 2014; Tang et al., 2014). In extracellular recordings from behaving mice, spike-train cross-correlation identifies putative monosynaptic excitatory connections (1–4 ms) from principal cells onto fast-spiking interneurons (Latuske et al., 2015); the excitation-to-inhibition offset we measured is of comparable magnitude, here resolved as a synaptic-current delay in electrophysiologically classified cell types.

      We note that bursting and theta engagement have been assigned inconsistently across in vivo datasets. Bursty cells are preferentially classified as putative stellate by spikepattern classifiers (Latuske et al., 2015), while anatomically identified pyramidal cells are reported as the bursty, theta-rhythmic population in other work (Ebbesen et al., 2016). Because our cell-type assignments are based on subthreshold intrinsic properties (membrane sag, time constant) rather than spike patterning, our phase-locking results are independent of this classification ambiguity.

      In vivo comparison (spatial organization): We agree high-density silicon-probe datasets are the appropriate reference here. To our knowledge, the anatomical distribution of gamma-locked spiking in superficial mEC has not been characterized in vivo. The highest-density available recordings (Gardner et al., 2022) analyze population activity in the decoded state rather than tissue coordinates, do not examine gamma, and are restricted to grid cells. We regard the dissociation we observe between spatially clustered subthreshold input and spatially distributed spiking as a principal advance of the present study, and as a testable prediction for future high-density recordings.

      Ebbesen CL, Reifenstein ET, Tang Q, Burgalossi A, Ray S, Schreiber S, Kempter R, Brecht M. 2016. Cell Type-Specific Differences in Spike Timing and Spike Shape in the Rat Parasubiculum and Superficial Medial Entorhinal Cortex. Cell Reports 16:1005–1015. DOI: https://doi.org/10.1016/j.celrep.2016.06.057

      Fuchs EC, Neitz A, Pinna R, Melzer S, Caputi A, Monyer H. 2016. Local and Distant Input Controlling Excitation in Layer II of the Medial Entorhinal Cortex. Neuron 89:194–208. DOI: https://doi.org/10.1016/j.neuron.2015.11.029

      Gardner RJ, Hermansen E, Pachitariu M, Burak Y, Baas NA, Dunn BA, Moser M-B, Moser EI. 2022. Toroidal topology of population activity in grid cells. Nature 602:123–128. DOI: https://doi.org/10.1038/s41586-021-04268-7

      Gonzalez-Sulser A, Parthier D, Candela A, McClure C, Pastoll H, Garden D, Sürmeli G, Nolan MF. 2014. Gabaergic projections from the medial septum selectively inhibit interneurons in the medial entorhinal cortex. Journal of Neuroscience 34:16739–16743. DOI: https://doi.org/10.1523/JNEUROSCI.1612-14.2014, PMID: 25505326

      Latuske P, Toader O, Allen K. 2015. Interspike Intervals Reveal Functionally Distinct Cell Populations in the Medial Entorhinal Cortex. Journal of Neuroscience 35:10963–10976. DOI: https://doi.org/10.1523/JNEUROSCI.0276-15.2015

      Ray S, Naumann R, Burgalossi A, Tang Q, Schmidt H, Brecht M. 2014. Grid-Layout and Theta-Modulation of Layer 2 Pyramidal Neurons in Medial Entorhinal Cortex. Science 343:891–896. DOI: https://doi.org/10.1126/science.1243028

      Tang Q, Burgalossi A, Ebbesen CL, Ray S, Naumann R, Schmidt H, Spicher D, Brecht M. 2014. Pyramidal and Stellate Cell Specificity of Grid and Border Representations in Layer 2 of Medial Entorhinal Cortex. Neuron 84:1191–1197. DOI: https://doi.org/10.1016/j.neuron.2014.11.009

      Williams B, Vedururu Srinivas A, Baravalle R, Fernandez FR, Canavier CC, White JohnA. 2026. Fast spiking interneurons autonomously generate fast gamma oscillations in the medial entorhinal cortex with excitation strength tuning ING–PING transitions. eneuro ENEURO.0452-25.2026. DOI: https://doi.org/10.1523/ENEURO.0452-25.2026

    1. eLife Assessment

      This study makes a solid and valuable contribution to elucidating the intricate relationship between mitochondrial calcium and neuronal survival. Well-controlled experiments show that homeostatic mitochondrial calcium correlates with the most resilient neuronal subtypes after optic nerve injury. However, altering mitochondrial calcium levels does not affect neuronal survival as initially predicted by this correlation.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      (2) Some findings can have alternate interpretations that are not considered.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

    4. Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

      We appreciate Reviewer #1’s assessment of our manuscript. We also agree that we should have more clearly indicated that our mitochondrial Ca2+ sensor (Cox8-Twitch2b) is localized to the mitochondrial matrix. The Cox8-mitochondrial localization peptide is a well-established tool first identified in 1992 by Rizzuto and colleagues (Rizzuto, Simpson and Pozzan, 1992). We should have cited this work in our manuscript and will add it to our references. Further, as discussed in our submission, Cox8-Twitch2b has previously been validated for mitochondrial Ca2+ measurements in CNS axons (Witte et al., 2019). Thus, given the decades of use and characterization for this toolset, and the fact that we have pharmacological data supporting mitochondrial matrix localization of Cox8-Twitch2b, we do not feel it is strongly necessary to demonstrate mitochondrial matrix versus inner membrane space localization. However, we could attempt immuno-electron microscopy if this is deemed critical.

      We also agree that the mechanism by which reducing mitochondrial Ca2+ is protective would be satisfying and strengthen this study. But we feel it is beyond the scope of this project. It is likely manifold since mitochondrial Ca2+ impacts many vital cellular functions relevant to pathology including metabolism and apoptosis. We ultimately believe that an adequate investigation of these mechanisms would significantly slow down the dissemination of the core novel findings presented herein.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      We agree with Reviewer #2 that our AAV manipulations of shMCU and MCU overexpression should be analyzed to verify how they alter mitochondrial Ca2+. To do this, we will co-express gene therapy vectors to lower and raise MCU expression with mito-Twitch2b biosensor and perform direct measurements of mitochondrial Ca2+. We will then determine if there is a relationship between gene expression level (inferred by mCherry intensity) and mitochondrial Ca2+ within samples, and if mean mitochondrial Ca2+ levels in treatments are higher or lower than mCherry reporter only controls.

      (2) Some findings can have alternate interpretations that are not considered.

      We will expand our Results and Discussion sections to broaden the interpretations of our data.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

      We agree that a more fine-grained understanding of RGC mitochondrial Ca2+ diversity would make interpretations of our data stronger. In our revisions, we will thus expand the number of RGC families in which we directly measure homeostatic mitochondrial Ca2+ levels. To do this, we will perform in vivo mito-Twitch2b measurements, collect and fix retinal wholemounts and immunostain for ON-OFF-direction selective RGCs using the marker CART and F-RGCs using the marker Foxp2. This will provide a complement of well-surviving RGC types (alpha and intrinsically photosensitive RGCs already examined) and poorly-surviving types.

      Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

      We agree with the feedback from Reviewer #3, especially as it aligns with input from other reviewers. As these points agree with aspects above we will briefly reiterate our proposed revisions. We will validate the true effects on mitochondrial Ca2+ levels after gene therapy treatments by co-injecting AAV-mito-Twitch2b and AAV-shMCU or AAV-MCU. We will measure mitochondrial Ca2+ levels and correlate these levels with mCherry reporter expression intensity to determine the effect size of these treatments, and compare sample mean mitochondrial Ca2+ levels with those of mCherry control AAV.

      To further map the variance in homeostatic mitochondrial Ca2+ levels to RGC types we will perform in vivo mito-Twitch2b imaging, and then immunostain for ON-OFF-direction selective RGCs (CART) and F-RGCs (Foxp2), two poorly surviving RGC types.

      Lastly, we agree with Reviewer #3 that finer delineation between mitochondrial Ca2+ levels and their relationship to survival may be informative. We will split RGCs into smaller subgroups based on homeostatic mitochondrial Ca2+ levels and examine their survival outcome.

      Overall, we thank the Reviewers for their feedback, and believe the suggested changes will greatly strengthen our study.

      REFERENCES

      Rizzuto R., Simpson A.W. and Pozzan T. (1992). Rapid changes of mitochondrial Ca2+ revealed by specifically targeted recombinant aequorin. Nature, 358 (6384): 325-327.

      Witte M.E., Schumacher A-M., Mahler C.F., Bewersdorf J.P., Lehmitz J., Scheiter A., Sanchez P., Williams P.R., Griesbeck O., Naumann R., Misgeld T. and Kerschensteiner M. (2019). Calcium influx through plasma-membrane nanoruptures drives axon degeneration in a model of multiple sclerosis. Neuron, 101(4): 615-624.

    1. eLife Assessment

      This article describes the comprehensive metabolic phenotype of a mouse model of Down Syndrome, together with supporting transcriptomic, metabolomic, and biochemical data. The evidence presented is compelling and highlights several core phenotypes including insulin resistance, dyslipidemia, and tissue signatures indicating inflammatory and cellular stress pathways. Similarities and differences in male and female mice are highlighted. This important study provides essential groundwork for the further genetic dissection of dosage-sensitive genes causing metabolic dysregulation in Down Syndrome.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

    3. Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

    4. Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

      Weaknesses:

      Many aspects of this study have been described in other Down syndrome mouse models, though there are certainly aspects that are new. It would be useful if the authors could do a direct critique and comparison with previous publications in the area, utilizing the same Down Syndrome mouse model. There are also a few limitations in the number of animals used and the interpretation of the data that should be acknowledged.

      We have cited all relevant publications using Down syndrome mouse models. Regarding the Dp16 model, we have cited and discussed the only other study addressing metabolic aspects beyond body weight (Reference #138; PMID: 39803786). While that study reported glucose intolerance, insulin resistance, and defective insulin secretion, we did not measure pancreatic insulin content in our mice. Crucially, while the previous study found no sexual dimorphism, our study observed extensive sexual dimorphism in body weight gain, tissue-specific gene expression, and serum and liver metabolite changes.

      Regarding sample size, we used 6 mice per genotype per sex for transcriptomic and metabolomic analyses; this is constrained by the cost of performing these omics-type analyses. For mitochondrial respiration assays, we used 9–10 mice, and for most other in vivo and ex vivo assays, we utilized 12–15 mice, with some assays exceeding 20. We believe these sample sizes are robust and appropriate for this study.

      Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

      Weaknesses:

      (1) Assessing metabolism using dynamic testing is a strength. ITT, GTT and LTTs are included.

      (2) The dosing for GTTs, ITTs and LTTs was performed per body weight. But the mice under chow and HFD had different body weights. This may compromise the interpretation of the data. Further, ITTs are presented as percentage change, and this can be heavily influenced by baseline glucose measures. The changes appear quite dramatic, so can the authors plot the raw data instead?

      We have updated the ITT data plots to show raw glucose values instead of percentage change. Regarding the dosing, we believe basing it on body weight is an appropriate approach. This method is consistent with nearly all published rodent studies, as blood volume and metabolic tissues such as skeletal muscle and adipose tissue scale with body weight. Adjusting for weight prevents potentially erroneous conclusions. As for the diet groups, we compared WT and Dp16 mice only within the same diet group (Chow or HFD) rather than across different diets. We believe this ensures a valid and appropriate comparison for our study.

      (3) In addition, throughout the manuscript, it is not clear which tissues are the most dominant in disrupting metabolism. The ITT and GTT are composite measures across tissues. Tissue-specific analyses using a clamp technique or isolated tissues may provide more clarity here.

      Our data suggest a systemic metabolic deficit across multiple tissues, supported by tolerance tests, pan-tissue transcriptomic analyses, and liver and serum metabolite profiling. This is consistent with the triplication of genes in Down syndrome, several of which have known metabolic roles as highlighted in our discussion. We do not have evidence to support the role of a dominant tissue that contributes to the systemic metabolic dysfunction.

      Regarding the suggestion to use a clamp technique, we agree this would effectively determine whether insulin resistance is localized in the liver or skeletal muscle. However, we do not currently have the necessary equipment at Johns Hopkins University to perform these experiments. Conducting this work would require sending separate cohorts of WT and Dp16 male and female mice (on both chow and HFD) to an NIH-funded Mouse Metabolic Phenotyping Centre (MMPC). While we appreciate the value of this approach, we believe such labor-intensive experimentation falls beyond the scope of the present study.

      (4) One of the aims of the study was "to help inform the selection of mouse model that best reflects the metabolic profile of DS". The discussion does not contain a comparison between the previous work on different strains and relative to known human data.

      We chose not to include a comparison of different mouse models in the "Discussion" section because we previously highlighted the widely used Down syndrome models (Ts65Dn, Tc1, and TcMAC21) and their associated caveats in the "Introduction." Given the significant limitations of those models such as hypermetabolism in TcMAC21 and the presence of 41 triplicated protein-coding genes unrelated to human chromosome 21 we focused our in-depth metabolic analyses on the Dp16 model, which does not share these issues. We felt that restating this information in the "Discussion" would be unnecessarily repetitive.

      (5) Data availability. Raw metabolomic data should be made available.

      We have uploaded all metabolomics data, along with details regarding sample processing and data analysis, to the Metabolomics Workbench, an NIH-funded public repository. We have updated the "Methods" and "Data Availability" sections of the manuscript to include this information and the corresponding access link.

      Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

      Weaknesses:

      It might have been helpful to have included blood pressure measurements, given the differences in 19-Nor-deoxycorticosterone. The discussion references several articles that describe sex-dependent differences in metabolic phenotypes in humans with Down syndrome, and it might have been helpful to state more explicitly whether these differences correlate with those observed here in mice.

      We appreciate the suggestion of blood pressure measurements. While we agree this is an important metric, given the metabolic focus of the present study and the significant volume of data already presented, we feel that blood pressure analysis is beyond the current scope and better suited for a follow-up study.

      Our study highlights sex differences in metabolic phenotypes in individuals with Down syndrome. While most published human studies focus on a limited set of parameters such as body weight, adiposity, serum lipoprotein profile, and fasting lipid/glucose levels our mouse data remain generally concordant with these findings. Beyond these standard measurements, we also observed substantial sex differences in pan-tissue transcriptomes as well as serum and liver metabolites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A major question is how these findings compare to data that have previously been published. For example, Lamantia et al. Bone 2024 and Dard et al. European Journal of Pharmacology 2025 both report no changes in body weight using the same Dp(16)1Yey Down syndrome mouse model? There is also a recent publication on liver dysfunction in Down Syndrome using the same mouse model. It would be useful to understand some of the similarities and differences of what is being reported by Dunn et al. Cell Rep 2026. In this assessment, there is an in-serum alanine transaminase (ALT) level, which was not the case in Dunn et al?

      For the Lamantia et al. Bone 2024 study, the authors only measured the body weights of Dp16 mice at 6 weeks of age. Our findings at 6 weeks align with Lamantia et al., showing no weight differences between Dp16 and WT mice of either sex (Fig. 2A and C). For the Dard et al. 2025 study, the authors only measured the body weights of Dp16 mice at 12 weeks old (P90) and observed no differences in body weights between genotype of either sex. At 12 weeks of age, we also did not observe body weight differences between Dp16 male mice and WT littermates (Fig. 2A). However, at 12 weeks of age, the Dp16 female mice clearly gained more weight compared to WT littermates (Fig. 2C). Our study tracked weights weekly from 6 to 16 weeks, revealing that while Dp16 females start at weights similar to WT littermates, the groups diverge over time. The reason for the difference between our findings and the single-point measurement by Dard et al. is unclear. Notable variables include:

      Mouse Sourcing: We obtained all cohorts and littermate controls from Jackson Laboratory, while Dard et al. bred their mice in-house.

      Diet: We used Envigo standard chow (catalogue # 2018SX). Dard et al. did not specify the chow used in their study.

      It remains uncertain whether these or other environmental factors contribute to the observed weight differences in female mice.

      In the Dunn et al study (Cell Rep 2026), they also performed metabolic analyses on serum and liver tissue in Dp16 mice. Consistent with their metabolic analyses of serum and liver tissue in Dp16 mice, we also observed the upregulation of multiple bile acids, including taurochenodeoxycholic, tauromuricholic, taurolithocholic, and lithocholic acids. Furthermore, our findings align with theirs regarding the transcriptomic and biochemical signatures of hepatic inflammation and fibrosis. However, there are two notable differences between our studies:

      (1) Liver Injury Markers: We observed an elevation in serum ALT, whereas the Dunn et al. study did not.

      (2) Sex Differences: We identified significant sex differences in the Dp16 transcriptome and metabolome. In contrast, Dunn et al. reported minimal to no sex differences and consequently combined male and female data for all analyses.

      Because Dunn et al. combined male and female data, a sex-stratified comparison between our results (separated by sex) and theirs was not feasible.

      (2) It would be important to understand trends in wild-type animals compared to Dp16 mice. For example, the sex specific and non-specific features - are any of these described in obesogenic wild-type animals fed on a high-fat diet? I.e., are the same features at play and just exacerbated in Dp16, or is this a Dp16-specific feature of systemic metabolism?

      Published literature indicates that WT females typically gain significantly less weight on a high-fat diet (HFD) than WT males. However, our data suggest that the weight gain patterns observed in Figure 6A and C are specific to the Dp16 genotype. Dp16 females gained substantially more weight during the first six weeks of HFD before WT females caught up. In contrast, Dp16 males showed robust initial weight gain comparable to WT controls, but their weight plateaued after seven weeks while WT controls continued to gain, leading to a clear divergence (Fig. 6A).

      Other metabolic parameters also appear specific to the Dp16 model. On a standard chow diet, WT mice of both sexes generally do not exhibit glucose intolerance, insulin resistance, dysregulated lipoprotein profiles (VLDL-TG), or an impaired capacity to handle lipid loads. We observed all of these features in our Dp16 male and female mice (Fig. 3). Furthermore, transcriptomic analyses of Dp16 mice on standard chow revealed gene signatures of inflammation, fibrosis, and oxidative stress that are absent in WT mice.

      When challenged with HFD, while WT mice typically develop glucose intolerance and insulin resistance, the triplicated genes in Dp16 mice significantly exacerbated this metabolic deterioration. This is reflected in the worsening of glucose control and insulin sensitivity observed in our tolerance tests.

      In summary, most of these metabolic features are specific to Dp16 mice on a standard chow diet and are further exacerbated when combined with a high-fat diet.

      (3) Food intake data is difficult to interpret when weight has already diverged, as bigger animals will eat more food. Hence, the higher food may be a consequence rather than a cause of the weight gain (data in Figure 1).

      The reviewer makes a valid point. Since physical activity and energy expenditure do not differ significantly between Dp16 females and WT controls (Fig. 2F), the observed increase in food intake may indeed contribute to the higher body weights in Dp16 female mice.

      To rigorously confirm this, food intake would need to be measured between 6 and 8 weeks of age, prior to the divergence in body weight. Unfortunately, we did not measure food intake at that earlier time point.

      (4) The n numbers seem to vary significantly. For example, the use of n=6 for metabolic studies is generally rather small and underpowered. For the seahorse data, another concern is the snap freezing of samples before Seahorse assessment. For example, snap freezing of samples has been shown to increase certain metabolites. Freeze-thaw tissues often show a significant reduction in optical redox ratio.

      Regarding the transcriptomics and metabolomics studies, we utilized N=6 mice per tissue per sex. While we agree that a larger sample size is always preferable, the high cost of OMICS analyses covering 144 RNA-seq and 48 metabolomics samples limited our capacity to increase this number. However, N=6 remains a robust and standard approach for these specific assays. For the majority of our other in vivo and ex vivo data, we employed a higher sample size of 12-15 mice per genotype per sex to ensure statistical rigour. For a few assays, we have sample size of over 20.

      Regarding the respirometry analysis, we acknowledge the limitations of using frozen tissue. We chose this method because it allowed us to perform Seahorse assays on multiple tissues from 9-10 mice, which is a significant sample size for this type of analysis. The alternative isolating mitochondria from fresh tissue would have restricted our ability to process multiple tissues from a large number of animals on the same day due to the length of the protocol. We believe this trade-off was necessary to maintain a high sample size across various tissues.

      (5) For oestradiol measurements, were the samples taken at the same times within the estrous cycle? This may affect the comparability of female Dp16 and WT mice?

      Regarding our protocol, blood samples were collected between 11:00 AM and noon, with food removed two hours prior. While we did not specifically monitor the oestrous cycle of the female mice, serum samples for both the Dp16 females and WT littermates were collected on the same day and at the same time to ensure comparability across the groups.

      (6) Body weight reduction and organ size reduction on an HFD are especially interesting. Could enhanced inflammation and fibrosis be the root cause of this? Are there other mouse models where this is the reason?

      On a high-fat diet, we observed a reduction in iWAT and gWAT fat depot weights in both male and female Dp16 mice, which is consistent with their lower overall body weights (Fig. 6 - figure supplement 3). Conversely, Dp16 females fed a high-fat diet showed increased heart and kidney weights. Despite their lower adiposity, the Dp16 mice on this diet exhibited greater insulin resistance and glucose intolerance (Fig. 7). This suggests that the worsening of glucose control is independent of obesity. While we observed signatures of inflammation and fibrosis, we do not yet have direct mechanistic evidence demonstrating that these factors causally impaired glucose and lipid metabolism.

      (7) The authors are circumspect throughout to avoid over-claiming, as the majority of data is observational. One exception: "Many bile acids serve as ligands for nuclear hormone receptors (e.g., FRX and TGR5) that control various aspects of glucose and lipid metabolism (74, 75), and extensive changes in circulating bile acids are contributing, at least in part, to the systemic metabolic phenotypes in Dp16 mice." The authors have not shown a direct link between bile acids and metabolism in this model. Please edit.

      We have edited the text accordingly.

      Minor:

      (1)"Most human studies at the whole-body level are limited to assessing the impact of trisomy 21 on food intake, adiposity, physical activity level, and energy expenditure in adolescents or adults with DS"

      While we were uncertain of the reviewer's specific intent regarding the suggested changes, we have rephrased the sentence for clarity.

      (2) It is somewhat surprising that T3 is elevated, although there are reports of T3 elevation in visceral obesity in humans (e.g., Sun Nam et al., Obes Res Clin Pract, 2010).

      We observed that T3 levels did not differ by genotype in mice of either sex when fed a standard chow (Fig. 2 - figure supplement 5). However, we noted elevated T3 levels in both male and female Dp16 mice on a high-fat diet (Fig. 6 - figure supplement 2). While increased T3 levels correlated with higher physical activity and a modest increase in metabolic rate in Dp16 females, this was not observed in males (Fig. 6). We do not currently have a clear explanation for these findings. Given that individuals with Down syndrome often present with hypothyroidism and lower T3 levels, this discrepancy may reflect a species-specific difference between humans and mice.

      (3) Please can the authors clarify the percentage gene coverage, as this is quoted as ~58% of Hsa21 gene orthologs or ~65% of the Hsa21 gene orthologs, where the same reference is used.

      We apologize for the confusion. The number of triplicated genes in Dp16 mice corresponds to ~58% of Hsa21 genes (PMID: 26765563). We have corrected the typographical error in the text.

      (4) "segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs" for this given percentage majority sounds too strong, and the use of percentage is recommended.

      We have modified the text accordingly.

      (5) It is puzzling that in female gWAT with 7 triplicated Hsa21 gene orthologs (Rbm11, Chodl, Cldn8, Sh3bgr, Igsf5, Itgb2l, and Tmprss2). Could this be a technical issue? Was the reduced expression quantified by RT-Q-PCR?

      We have examined the normalized counts in the RNA-seq data for the seven genes in question, and the results do not appear to be an artifact. The sample size for this data is six mice per tissue per sex. In general, we prefer utilizing raw and normalized counts from RNA sequencing because there is a linear relationship between transcript amount and raw counts that is independent of housekeeping genes. In contrast, RT-qPCR involves mRNA amplification and requires expression to be normalized by one or more housekeeping genes (such as GAPDH, β-actin, 36B4, or ubiquitin) under the assumption that their levels remain constant.

      (6) The difference in body temperature is of interest. In male Dp16 mice, there is an increase in core temperature and a lowering of body temperature in females. In female Dp16 mice, higher estradiol levels have been stated by the authors to contribute to lower body temperature and higher physical activity (69-72). I am uncertain if the references are all relevant, as some relate to ovariectomized animals. No explanation is given for males.

      We currently do not have an explanation for why Dp16 males on a chow diet exhibit higher core body temperature, while Dp16 females show lower body temperatures. Although elevated T3 levels can increase body temperature, we have ruled this out; our data indicates there are no significant differences in T3 levels between genotypes for either sex on a chow diet.

      (7) The authors find a higher percentage heart weight in Dp16 mice on HFD and comment in the discussion that this is in keeping with "high-fat diet-induced cardiac hypertrophy". From what I can see, no histology has been performed to justify this statement. Furthermore, it would be useful to understand which animals had congenital heart disease in the first instance.

      We have modified the text accordingly. Unfortunately, we do not have histology data on the heart to inform us on whether some of our mice had congenital heart disease.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should comment on the dosing method of glucose/insulin/lipid in the tolerance tests to acknowledge that differences in body weight may affect these tests. In addition, I encourage the authors to present ITT data as raw data, and not % change.

      In response to the reviewer’s comments, we have updated the ITT data plots to show raw data rather than percentage change. Regarding the dosing methodology, we maintain that basing dosage on body weight is appropriate. This approach is consistent with the vast majority of published rodent studies, as blood volume and metabolic tissues—such as skeletal muscle and adipose tissue—scale with body weight. Standardizing dose independently of body weight could lead to erroneous conclusions.

      (2) It would be useful for the authors to include a discussion on the likely specific tissue involvement in the whole-body metabolic disturbance. From my reading of the manuscript, there seems to be data suggesting functional and transcriptional dysfunction across most tissues, but do the authors suggest there is a dominant tissue in this regard?

      Due to the triplication of large number of genes on human chromosome 21, people with Down syndrome exhibit deficits across most organ systems (PMID: 32029743). Metabolic homeostasis also involves multiple tissues and cell types (adipose tissues, liver, skeletal muscle, pancreas, gut, hypothalamus, and immune cells). Most of the triplicated genes do express across these tissues. Our data indicate metabolic dysregulation across adipose tissues (white and brown), liver, skeletal muscle, and hypothalamus. Given the complex genetic perturbations of the Down syndrome mouse model, we do not think that there is a dominant tissue that contributes disproportionately to the systemic metabolic dysfunction phenotypes we observed in the Dp16 mice. Rather, we think that the metabolic phenotype is due to the combined deficits across multiple organs and tissues. As we do not have data to support the disproportionate contribution of any one tissue, we therefore did not speculate on the dominant contribution of any single tissue in the Discussion.

      (3) Related to this, muscle lipid is thought to be a major driver of muscle insulin resistance. Do the authors have measures of muscle lipid accumulation? This might be particularly interesting in the HFD models.

      Unfortunately, we did not measure lipid content in the skeletal muscle during this study. For the chow-fed mice, the entire gastrocnemius muscle was used for RNA isolation to perform RNA sequencing, and no tissue remains for additional analysis. Regarding the HFD-fed group, skeletal muscle was not collected at the termination of the study. As a result, we are unable to provide the requested lipid analysis data.

      (4) For mitochondrial analyses - do the authors have measures of total tissue mitochondria, and might changes in mitochondria abundance be driving some of these differences?

      For all our mitochondrial respiration analyses, we normalized the data to mitochondrial content as quantified by the MTDR assay (PMID: 32432379; PMID: 39704485). These results indicate that for a given amount of mitochondrial content, respiration as measured by the Seahorse assay is reduced in Dp16 mouse tissues, specifically in the BAT and liver.

      (5) To broaden the scope and interest, can the authors compare the transcriptional or metabolomic data to what has been found in non-DS insulin resistance (humans or mice), for example? This may help to highlight the key changes in metabolism that are causal for specific phenotypes.

      Overall, this is a comprehensive assessment of metabolism in a DS model.

      We appreciate the reviewer’s suggestion. However, given the vast number of published datasets on non-DS insulin resistance in both humans and mice, comparisons would yield varying results depending on the specific datasets selected. Consequently, we feel that such an analysis is beyond the scope of this study. We would like to highlight that many of the processes dysregulated in Dp16 mice as identified through our pan-tissue transcriptomes and metabolomes align with those frequently observed in non-DS insulin resistance. These include signatures of chronic low-grade inflammation, fibrosis, ER and oxidative stress, and impaired glucose and lipid metabolism.

      Reviewer #3 (Recommendations for the authors):

      It is slightly disconcerting that Figure 5 - Figure Supplements 2-5 are referred to in the text before the data in Figure 5 are discussed. It might make sense to indicate that the data are discussed further below (assuming that the authors do not wish to renumber these figures).

      We have fixed this issue raised by the reviewer.

    1. eLife Assessment

      This study presents an important large-scale behavioral and transcriptomic analysis of Drosophila that are heterozygous for putative loss-of-function alleles of homologs of human genes that have been linked to autism spectrum disorders. The authors consider 48 genes as hits from their screen, which show significant behavioral alterations in sleep, basal activity, and/or social behavior, and significant sexual dimorphism. The authors then focus on the domino/SRCAP gene as a candidate regulator of sleep, social behavior, transcriptional programs, and RNA splicing. The work generates a solid dataset and applies quantitative analytical approaches that will be of interest to researchers in the field, yet the evidence presented remains incomplete because issues of genetic background need to be further addressed.

    2. Joint public review:

      Summary:

      In this study, Stirtz et al., performed a targeted screen of 80 Drosophila strains carrying heterozygous MiMIC insertions in genes that are homologous to human genes that have been linked to autism spectrum disorders (ASD). This is an important and timely topic, as human genetic studies have identified a large number of ASD risk genes, yet the functional characterization of many of these candidates remains limited. The authors identify 48 putative mutants with altered sleep, activity, or social behavior. They then focus on one hit, domino (the orthologue of human SRCAP), for which the heterozygous MiMIC mutants show altered behavior in males but not in females. They show that domino is a candidate regulator of sleep, activity, social behavior, transcriptional programs, and RNA splicing. The authors molecularly validate that the heterozygous MiMIC insertion in domino causes a 50% reduction in gene expression, and use RNA-seq to show that the heterozygous MiMIC males and females have altered gene expression profiles and splicing patterns. Finally, they use immunostaining against the commonly used synaptic marker, Bruchpilot, to show that both males and female heterozygous domino flies express a higher immunosignal compared to the wild-type control.

      Strengths:

      This work provides potential genetic links between human ASD genes and fly behavioral phenotypes. Overall, it represents an ambitious and technically valuable effort that generates a substantial behavioral dataset across a large number of ASD-associated orthologues and develops quantitative analytical approaches to extract information from complex phenotypes. One strength of this study is its focus on heterozygous mutants, which is more representative of human scenarios. The study also provides a potentially useful resource for the field, particularly through the identification of candidate genes and behavioral signatures that may warrant future mechanistic investigations. The screening experiments and analysis are well conceived, the manuscript is very clearly written and is easily understandable, and the concise, accurate interpretations for each result, aided by clear graphic representation of multiple dimensions in the behaviors tested, allow the reader to understand the paper with ease.

      Weaknesses:

      The work presents a few important weaknesses, especially with regard to the genetic and molecular validation of the mutants identified.

      (1) The authors validate that the MiMIC insertion affects the gene of interest only for the domino gene. The original MiMIC study (PMID: 25824290, eLife) reported that ~8% (5/63) MiMIC lines do not function as strong loss-of-function alleles. Thus, of the 48 hits identified here, one would estimate that ~4 of them may not cause the loss of function of the gene defined by the MiMIC insertion. To strengthen their claim, the authors would need to confirm that all of the MiMIC lines that they consider as hits do indeed significantly reduce the expression of the target genes.

      (2) Although the authors document that they validated the phenotype seen in the domino MiMIC line using a second mutant allele (Trojan), these two mutants share the same genetic background because the Trojan line was made from the MiMIC line via recombinase-mediated cassette exchange. Thus, the phenotype seen in the MiMIC and Trojan lines would need to be confirmed using a completely independent mutant in order to demonstrate that the reported behavioral, molecular, and synaptic defects reported can be fully attributed to the partial loss of domino function. Also, while the authors performed an RNA-seq experiment in both the MiMIC and Trojan lines, they do not show whether the Bruchpilot phenotype is also seen in the Trojan allele. Thus, this phenotype would also need to be examined in the Trojan allele or, preferably, in a mutant allele that is independent of the MiMIC line.

      (3) The RNA-seq results would benefit from a discussion of potential compensatory or secondary transcriptional effects resulting from the constitutive domino reduction, particularly since the expected global bias toward transcriptional downregulation was not observed. In addition, some neurobiological interpretations appear stronger than currently justified by the literature or the data presented, particularly regarding the Bruchpilot immunoreactivity analyses and their relationship to sleep-regulatory circuits. Additional validation using better-established sleep-related neuronal populations, together with a clearer discussion of sex-specific effects and alternative interpretations of the observed phenotypes, would substantially strengthen the manuscript.

      (4) An explanation of the extensive PCA analyses performed would help the naïve reader.

    1. eLife Assessment

      In this useful Tools & Resources article, the authors describe a new cryogenic light microscopy design and characterize its temperature and spatial stability. This compelling system avoids the challenges associated with vacuum-based designs, particularly vacuum transfer systems, which are difficult to engineer. A key advantage of the system is that it reduces ice contamination and drift, which are the primary challenges in open cryostat systems.

    2. Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      We thank the reviewer for their positive assessment and strive to address the weaknesses they have constructively raised below.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we will include a new supplementary figure (Fig. S2) providing detailed 3D views of the copper adapter, microscope objective, and plastic cap. The figure will show that the rim surrounding the front lens of the objective is covered by the plastic cap to provide thermal insulation between the objective housing and the cryogenic environment (Fig. S2b). We will also clarify that the front surface of the cap is levelled with the front objective lens to maintain the full working distance of the objective while allowing for axial movement of the z-stage. Finally, we will explicitly state that the front objective lens is exposed to the cryogenic environment (cold nitrogen gas).

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      To minimise vibrations arising from the nitrogen refilling pumps, the cryostat and liquid nitrogen tubing are mechanically decoupled from the microscope cage system, objective, translation stages, and sample. Specifically, the cryostat and nitrogen tubing are supported independently on a laboratory jack and surround the cage system without rigid mechanical contact. In the revised manuscript, we will update Fig. S1a, b to illustrate the liquid nitrogen tubing and refilling system more clearly. In addition, we will include a new supplementary figure (Fig. S3) to show the detailed cryostat design, refilling tubing, and temperature sensor position.

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      We agree that the purpose of the rectangular aperture should be made clearer. In line with the suggestion from the reviewer, in the revised manuscript, we will briefly explain that the aperture is intended for selective illumination, such as in applications to cryo-FIB lamellae, and will cite our recent preprint describing this approach.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      We thank the reviewer for this suggestion. In the revised manuscript, we will include an additional supplementary figure (Fig. S4) showing the axial drift measured without focus-lock correction to illustrate the intrinsic mechanical stability of the microscope. We will also clarify that the periodic lateral displacement observed along the x-direction (approximately 300 nm amplitude with a period of ~22 minutes) arises from slight lateral repositioning accompanying z-stage stepping during focus-lock operation, likely due to mechanical coupling between the axes of the translation stage. We will revise the lower panel of Fig. 2d by reducing the y-axis range to improve data visibility.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      We agree that this would be useful. In the revised manuscript, we will update Fig. 3 to include two additional low and intermediate-magnification cryo-EM images showing a representative grid square and a zoomed-in region of it, including a few grid holes. These images provide an overview of the ice contamination across a substantially larger field of view.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      The y-axis in Fig. 4b represents the localization rate (number of detected localizations per frame) within the selected ROI in Fig.4c and was not normalized. The values were calculated in SMAP by binning the localization frames into 100 temporal bins and dividing the number of localizations in each bin by the corresponding bin width, resulting in units of localizations per frame. Therefore, values close to 1 indicate approximately one localization detected per frame at that time point. To avoid potential confusion regarding the interpretation of this representation, we will replace this plot in the revised manuscript with a more explicit visualization showing the number of detected localizations per defined number of frames as a function of time (frame number) for the specific ROI shown in Fig. 4c.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

      In the revised manuscript, we will add fluorescence intensity traces corresponding to the single-molecule events shown in Fig. 4c and Fig. 4f to further demonstrate their single-molecule emission characteristics.

      Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      We thank the reviewer for their critical assessment, and for their suggestions below which we have used to improve the manuscript.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      We thank the reviewer for highlighting this point. We agree that focus stabilization strategies for cryo-FIB lamellae are likely to differ from those used for the individual bacterial cell samples. For lateral drift correction, the presence of a single continuously detectable bright feature within the field of view is sufficient. Importantly, this feature does not need to be a fluorescent bead; any stable signal that can be continuously detected by the camera can serve as a suitable reference for drift correction. We will expand the Discussion to describe potential strategies for stable cryo-SMLM imaging, including the use of intrinsic sample or lamella features for autofocus, minimal fiducial-based approaches, and the practical implications of dedicating the far-red channel to focus stabilization.

      Furthermore, in the revised manuscript, we will include a new supplementary figure (Fig. S4) demonstrating the intrinsic axial stability of the microscope in the absence of active focus-lock correction. These measurements show that the system remains within the objective's depth of focus for a relatively long time, providing adequate stability for experiments in which far-red fluorescent fiducial beads are unavailable, such as cryo-FIB lamella imaging.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      We thank the reviewer for this thoughtful comment. We agree that a correlative example displaying more pronounced structural features would further illustrate the capabilities of cryo-SMLM. However, the primary aim of the present work is the development and characterization of a robust cryogenic super-resolution microscope for reliable cryo-SMLM and correlative cryo-CLEM, rather than the demonstration of new biological applications. The utility of correlative cryo-SMLM/cryo-ET for resolving cellular structures has already been established in previous studies, including those employing rsEGFP2-labelled targets.

      The correlative dataset presented here is intended to demonstrate the compatibility of the microscope with cryo-CLEM workflows rather than to provide detailed biological insight. Moreover, the use of intact E. coli cells imposes inherent limitations on the ultrastructural information accessible by cryo-electron tomography; overcoming these limitations would typically require specimen thinning, for example, by cryo-focused ion beam (cryo-FIB) milling, which is beyond the scope of the present work.

      Regarding the concern about auto-fluorescence, elevated background fluorescence is not unique to bacterial samples or green fluorescent proteins but is a general consideration in cryo-SMLM that depends on the specimen and imaging conditions. While auto-fluorescence may reduce image contrast, it does not affect the conclusions of this work, which focuses on the design and performance of the microscope.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

      Thank you for this useful suggestion. In the revised manuscript, we will make available the complete SolidWorks CAD files for all custom-designed components, together with a comprehensive parts list and the full assembly corresponding to Fig. S1 as supplementary materials.

    1. eLife Assessment

      This manuscript describes a valuable study of the mechanism by which acetylation on the histone H3 core domain regulates RNA polymerase II transcription passing through nucleosomes. The authors provide convincing evidence that acetylation influences transcription in a context-specific fashion. Some questions relating to the static nucleosome structures and the polymerase passage remain, but this manuscript will be of considerable interest to researchers in the chromatin and transcription fields.

    2. Reviewer #1 (Public review):

      Summary:

      The authors investigate how site-specific acetylation within the histone H3 folded domain affects RNA polymerase II transcription through nucleosomes. They focus on H3K56ac, H3K64ac, and H3K122ac, prepare chemically defined nucleosomes carrying each modification, and compare their effects using an in vitro transcription assay, cryo-electron microscopy structures, and micrococcal nuclease sensitivity assays.

      The main finding is that H3K56ac and H3K122ac increase production of full-length run-off transcripts and reduce pausing near the nucleosomal dyad region, whereas H3K64ac has little detectable effect under the same reconstituted conditions. The structural analyses suggest that H3K56ac weakens or destabilizes DNA near the entry/exit region, while H3K122ac alters histone-DNA contacts near the dyad. These observations support a model in which different acetylation sites within the H3 folded domain influence nucleosomal transcription barriers through distinct local effects on histone-DNA interactions.

      This is a useful study because it examines histone core-domain acetylation using chemically defined nucleosomes and directly compares several modifications in the same experimental system. However, the broader cellular context of these modifications is not sufficiently developed, and some mechanistic conclusions rely on correlations between static nucleosome structures and endpoint transcription assays rather than direct observation of polymerase passage through modified nucleosomes.

      Strengths:

      (1) The study uses site-specifically acetylated H3 proteins and reconstituted nucleosomes, allowing direct comparison of H3K56ac, H3K64ac, and H3K122ac under controlled conditions.

      (2) The combination of transcription assays, cryo-electron microscopy, and nuclease sensitivity assays provides multiple lines of evidence, particularly for increased DNA end flexibility in H3K56ac nucleosomes.

      (3) The authors analyze unmodified, H3K56ac, H3K64ac, and H3K122ac nucleosomes in parallel, with reported structural resolutions of approximately 3 Angstroms and accompanying validation materials.

      (4) The negative result for H3K64ac is informative, because it distinguishes the direct effect of this modification in a minimal reconstituted system from prior cellular associations with active chromatin and histone eviction.<br /> The comparison with H3 N-terminal acetylation highlights that acetylation within the folded domain may affect transcription at different positions or by different mechanisms than tail acetylation.

      Weaknesses:

      The rationale for focusing on H3K56ac, H3K64ac, and H3K122ac has not been developed sufficiently. The manuscript would benefit from a clearer summary of what is known about the abundance of these modifications in cells, the enzymes or histone metabolic pathways that may introduce or remove them, and whether they are thought to occur before histone deposition, on assembled nucleosomes, or during nucleosome remodeling.

      The central mechanistic model is based mainly on correlations between structures of free nucleosomes and endpoint transcription assays. The study does not directly observe RNA polymerase II paused at or passing through the relevant nucleosomal positions, so the proposed link between local structural changes and reduced pausing should be stated with appropriate caution.

      The H3K56ac interpretation is supported by both structural observations and nuclease sensitivity data, but the map comparison underlying the reduced entry/exit DNA density is still mostly qualitative. The manuscript should more clearly state the map comparison conditions, such as contouring and local map quality, so that non-specialist readers can judge how robust the local density differences are.

      The H3K122ac mechanism is plausible, but the evidence for dyad destabilization is more indirect. The main support comes from the orientation of the K122 side chain and its distance from DNA, while an independent biochemical test of dyad-region destabilization is not provided.

      The transcription assay appears to include statistical testing, but the figure legend and methods should more clearly state which tests were used, what comparisons were made, how n was defined, and whether multiple-comparison correction was applied.

      The relationship between the 198 bp transcription template, the linker DNA, the 9-base mismatched region, and the DNA regions modeled in the cryo-electron microscopy structures is somewhat difficult to follow. This does not necessarily require new experiments, but a clearer explanation would help readers connect the transcription assay design with the structural models.

      The use of H3.2 C110A for chemical ligation and the use of the PL2-6 single-chain antibody fragment for cryo-electron microscopy sample stabilization are reasonable technical choices, but their purposes and possible effects on interpretation should be explained more clearly for readers outside structural biology.

      Because the work uses a minimal in vitro system with human nucleosomes and Komagataella phaffii RNA polymerase II/TFIIS, the conclusions should be limited to direct physical effects on nucleosome transcription barriers unless cellular cofactors, remodelers, histone chaperones, additional modifications, and nucleosome positioning are addressed or discussed.

    3. Reviewer #2 (Public review):

      Summary:

      Chromatin regulates a wide range of biological processes. The nucleosome, composed of 147 bp of DNA wrapped around a histone octamer containing histones H2A, H2B, H3, and H4, is the fundamental unit of chromatin. Post-translational modifications of histone proteins regulate the dynamic properties of nucleosomes and thereby influence chromatin accessibility and gene expression. Among these modifications, lysine acetylation on histone H3 is closely associated with transcriptional activation. While the epigenetic functions of acetylation on the histone H3 N-terminal tail have been extensively studied, the molecular mechanisms by which acetylation within the histone H3 core domain, particularly at Lys56, Lys64, and Lys122, modulates nucleosome architecture to facilitate RNA polymerase II (RNAPII) transcription remain unclear.

      In this study, Oishi et al. investigated the effects of histone H3 acetylation at K56, K64, and K122 on RNAPII transcription using in vitro transcription assays. Furthermore, the authors determined the three-dimensional structures of nucleosomes containing these acetylation marks by cryo-electron microscopy single-particle analysis, revealing distinct structural dynamics depending on the acetylation site. Overall, this study advances our understanding of the molecular mechanisms linking histone H3 core acetylation to transcriptional regulation.

      Strengths:

      (1) Site-specifically acetylated histone H3 proteins were chemically synthesized using a unique and rational peptide ligation strategy, representing a major technical strength of this study.

      (2) The in vitro transcription assays demonstrated that H3K56ac and H3K122ac increase the production of run-off transcripts, whereas H3K64ac has little effect on transcription efficiency. These findings highlight the distinct functional roles of individual acetylation sites within the histone H3 core domain.

      (3) The cryo-EM structures of nucleosomes containing either H3K56ac or H3K122ac revealed that H3 acetylation weakens histone-DNA interactions, providing a structural basis for the observed effects on transcription.

      Weaknesses:

      (1) Although the biochemical and structural data are convincing and sufficiently support the authors' conclusions, complementary cellular experiments would further strengthen the physiological relevance of the in vitro findings. While such experiments are not essential for supporting the main claims of the study, they would enhance the overall impact and biological significance of the work.

      (2) Although the authors demonstrate the structural consequences of individual H3 core acetylation events, the study does not investigate potential synergistic effects among multiple acetylated lysine residues within the H3 core domain. Consequently, the relationship between combinatorial acetylation patterns and their collective impact on RNA polymerase II-mediated transcription remains unclear.

    4. Reviewer #3 (Public review):

      This is a short and punchy manuscript that nicely summarises the 4 structures that are determined and provides a basis for the differences seen for acetylation sites shown for RNAPII activity.

      The authors build on previous biochemical work that determined the functional outcomes of H3 core acetylation, adapting an assay they have previously used extensively to investigate RNAPII transcription on nucleosomes and, indeed, even H3 N-terminal tail acetylation. This assay is as such well set up and has a wealth of confirmatory previous studies from this lab and the authors are careful not to overanalyse their results, leading to robust and well-considered results. The structures are determined to a high resolution, allowing the interpretation put forward about side chain orientations, with clear densities shown for the regions of interest.

      Further discussion or experiments would strengthen the conclusions further:

      (1) The conclusion on the role of H3K56Acetylation could be strengthened, especially as the results are somewhat counterintuitive. It is conceptually surprising that acetylation near the entry/exit DNA that destabilises this region also leads to a reduced stall propensity at the dyad but has a limited effect at SHL5? While it can be explained by the clash at the dyad pause being reduced, the more direct effect of DNA breathing amplification would be expected to have a larger effect at SHL 5. Indeed, the density for DNA at SHL5 appears to be weaker in Figure 2A, suggesting the entry/exit DNA flexibility is amplified past this region.

      Perhaps another assay that looks more directly at the flexibility of the entry/exit DNA would be useful, either through restriction enzyme-mediated cleavage or FRET (DNA ends and H2AK119 labels), providing stronger evidence of this effect. MNase is rather indirect and similar to the RNAPII assay itself.

      Similarly, were the authors surprised by the modest effect (less than 2-fold) in transcriptional pause at SHL 0 for the K122Ac? Presumably, based on the model in Figure 4, this would be expected to be the area with the largest effect? The results of K56Ac and K122Ac almost seem swapped to what would be expected in Figure 1H. Further discussion of this observation would be useful.

      (2) Could the local weakening of DNA, especially at the dyad, be observed in the cryo-EM structures? Perhaps comparison of local resolution estimation differences in this region compared to unmodified would be useful.

      (3) Caution should be taken, and discussion should include that the structural data presented is after extensive processing. Many nucleosome averaging classes were discarded in the 3D classification steps (nicely summarised in Table 1 as "particles for 3d classification" and "particles in final map"). Indeed, it is likely that higher DNA flexibility particles would be thrown away during this processing step. This can be observed for K56Ac DNA ordering, for example, in Supplementary Figure S4, yellow and cyan classes from the round of 3D classification look to be high resolution and have a higher order of DNA, so there has been some selection here. How was this done? While this is not fully quantifiable, it gives an idea of the extent of wrapping. We would suggest discussing the methodological limitations and showing the models after the first auto refinement to see if the features discussed on end flexibility and dan ordering are retained.

      (4) Di Cerbo et al. (reference 13) showed acetylation at K64 alters salt stability and affects transcription. Why do the authors think there is a discrepancy, albeit with different assays? Direct reference and discussion of this in the text should be included.

      (5) Why was H3.2 used, while this is relatively abundant in mouse cells, human protein was used, and this appears to be less common than H3.1 and H3.3. We are sure that the effect is not likely to be substantive on structure (as shown by the Kurumizaka lab previously), but should be addressed in the text

    1. eLife Assessment

      This important study provides a mechanistic view of how antibody affinity maturation can reshape encounter-state landscapes and association pathways, with implications for understanding HIV antibody maturation and vaccine design. The results are solid, supported by a coherent integration of adaptive molecular dynamics, Markov state modeling, SPR kinetics, mutagenesis, and double-mutant cycle analysis, although aspects of the kinetic validation, MSM-state robustness, and causal interpretation would benefit from further support. The work will be of interest to immunologists, structural biologists, and computational biophysicists.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript uses simulations and MSMs paired with experimental binding assays to examine the binding mechanisms of different antibodies to their targets. The authors argue that contacts in encounter complexes play an important role in determining the association rates and binding affinities that distinguish more mature antibodies from less efficacious antibodies from earlier in the maturation process.

      Strengths:

      The idea is interesting, and the combination of computational models and experiments is a good direction.

      Weaknesses:

      The manuscript focuses heavily on kinetics, but it is not clear whether the simulations recapitulate the relative rates of binding of the two antibodies. The relationship between the simulated binding behavior and the experimentally observed kinetic differences is therefore not fully established.

      The comparison of committor probabilities or fluxes between the two antibodies may not be appropriate. These properties are related to the barrier height the system has to cross to move forward vs back to the starting state, under the simplifying assumption that the properties of other states aren't critical. Even in this simplified case, the same flux or committor probability could occur with very different barrier heights, e.g., rates or transition probabilities.

      Some claims are presented in a very qualitative way that people who aren't experts in MSMs may have difficulty tying to the results in Figure 1.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript addresses an important and underexplored question: how affinity maturation alters antibody encounter-state landscapes rather than simply improving bound-state affinity. The authors combine adaptive MD, Markov State Models (MSMs), transition path theory, mutagenesis, SPR kinetics, and double-mutant cycle analysis into a coherent story.

      Strengths:

      This manuscript presents a compelling computational and experimental analysis of antibody affinity maturation in the HIV-1 DH270 lineage. The main finding is that somatic mutations reshape encounter-state pathways through glycan-mediated steering rather than simply stabilizing the final bound state. This is novel and potentially important for vaccine design. The combination of adaptive MD, MSMs, SPR kinetics, and double-mutant cycle analysis is a major strength.

      Weaknesses:

      The proposed sequence that somatic mutations cause glycan capture, which causes reorientation, which causes enhanced association, is based on correlation rather than direct causality.

      The four MSM states are not convincingly explained, and the robustness of these states is unclear.

      The productive collision surface area analysis needs more quantitative data.

      The coupling energy values are near the uncertainty range. Some conclusions about long-range communication networks appear stronger than the data justify. The data support coupling, but they do not necessarily support detailed mechanistic networks.

      The study investigates one lineage, one epitope class, and one viral system. Hence, the generalization is limited.

    4. Reviewer #3 (Public review):

      Summary:

      In this work, the authors set out to characterise how encounter states between antibodies and antigens evolve during affinity maturation through molecular dynamics simulations and Markov state modeling. They demonstrate how early glycan-mediated interactions increased association rates rather than modifying the final bound state.

      Strengths:

      The computational approach is backed up by experimental results and allows for visualising otherwise too short-lived association states, thus allowing to discriminate between different lineages.

      Weaknesses:

      The figures and captions are not always clear about what they are trying to show. The choice of CVs is not sufficiently discussed.

    1. eLife Assessment

      Muetter et al. provide an important argument that luminescence is a reliable, high-throughput alternative to colony-forming units (CFU) for super-MIC investigations, particularly when the quantity of interest is biomass. By examining 20 antimicrobials spanning 11 classes, the work shows that discrepancies between CFU and luminescence are often biological (filamentation, Viable But Not Culturable). The work provides a convincing view of how these three common measurements (luminescence, optical density, and CFU) relate to one another across a range of drug treatments, although testing on clinical isolates could be of further benefit.

    2. Reviewer #2 (Public review):

      Summary:

      In antibiotic research, accurately measuring decreases in bacterial populations is essential. The authors conducted a comprehensive evaluation of the luminescence assay, a commonly used but previously under-quantified method, benchmarking it against the gold-standard CFU counting approach. They found that luminescence measurements generally aligned with CFU results but sometimes reported slower decline rates for certain antimicrobials. These discrepancies were linked to differences in how the two methods capture biomass and colony formation, which vary with the antimicrobial's mechanism of action. The study demonstrates that luminescence assays can serve as a high-throughput alternative to labor-intensive CFU counting, provided their limitations are understood and corrected.

      Strengths:

      The authors developed a mathematical model to partially correct luminescence-based measurements, making the approach broadly applicable to several commonly used antibiotics. They also analyzed antibiotic-treated single-cell morphologies and linked filamentation to bulk luminescence signals. This analysis helped define the range of drug conditions under which luminescence assays provide reliable estimates of bacterial dynamics.

      They extensively evaluated the method using 20 antibiotics and one antimicrobial peptide, encompassing many of the most commonly used agents and experimental factors (e.g. treatment time) typically considered in antibiotic research.

      Comments on revised version:

      No further comments. The authors have adequately addressed my concerns.

    3. Reviewer #3 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability / carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorization of drug-specific assay behaviors.

      The study critically exposes flaws in the "gold standard" CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimize plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      In summary:<br /> Muetter et al. provide a compelling argument that luminescence is a reliable, high-throughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      Comments on revised version:

      The revised version addressed my comments well.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines how luminescence can be used to measure bacterial population dynamics during antimicrobial treatment by comparing it directly with optical density and colony counts. The authors aim to determine when luminescence reflects changes in population size and when it instead captures metabolic or physiological states induced by drug exposure. By generating parallel datasets under controlled conditions, the work provides a detailed view of how these three common measurements relate to one another across a range of drug treatments.

      Strengths

      The study is technically strong and thoughtfully designed. Measuring luminescence, optical density, and colony counts from the same cultures allows the authors to make clear and informative comparisons between methods. The data are compelling, and the analyses highlight both agreements and divergences in a way that is easy to interpret. The manuscript also succeeds in showing why these divergences arise. For example, the observation that filamentation and metabolic shifts can sustain luminescence even when colony counts drop provides valuable information on how different readouts capture distinct aspects of bacterial physiology. The writing is clear, the figures are effective, and the work will be useful for researchers who need high-throughput approaches to quantify microbial population dynamics experimentally.

      Weaknesses:

      The study also exposes some inherent limitations of luminescence-based measurements. Because luminescence depends on metabolic activity, it can remain high when cells are damaged or unable to resume growth, and it can fall quickly when drugs disrupt energy production, even if cells remain physically intact. These properties complicate interpretation in conditions that induce strong stress re-sponses or heterogeneous survival states.

      In addition, the use of drug-free plates for colony counts may overestimate survival when filamented or stressed cells recover once the antibiotic is removed, making differences between luminescence and colony counts harder to attribute to killing alone. Finally, while the authors discuss luminescence in the context of clinically relevant concentration ranges, the current implementation relies on engineered laboratory strains and does not directly demonstrate applicability to clinical isolates. These limitations do not detract from the technical value of the work but should be kept in mind by readers who wish to apply the method more broadly.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Luminescence limitations. We agree that the lack of a direct link between light intensity and a population property such as biomass or cell number is the main limitation of the luminescence method. To further emphasise this, we have expanded the Discussion in the revised manuscript.

      Drug-free plates. The use of drug-free plates is intentional. As we measure a time series, the question at each point is how many cells are alive at each time point. Cells that are stressed but viable at time t contribute correctly to the count at t. How long they survive under the respective treatment is captured by the subsequent timepoints.

      Filaments. Recovery of plated filamented cells should not inflate this estimate. A single plated filamentous cell is expected to yield either zero (death before division) or one single colony, regardless of in how many parts it separates, as all descendants are part of the same cluster. However, if the cells divide before plating, CFU can overestimate survival. Having that said, we have no indication that this occurred in our experiments, since in all observed discrepancies, CFU-based estimates were equal to or lower than those obtained from luminescence and the time cells spent in dilution was kept short.

      Clinical applicability. We agree with the reviewer that the method is not practical for ad-hoc pharmacodynamic studies of clinical isolates. What we instead provide is an E. coli-based model system to explore clinically relevant treatment conditions, which we address in the revised manuscript. We believe that constructing analogous bioluminescent model strains in other clinically relevant species would be a valuable direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability/carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorisation of drug-specific assay behaviours.

      The study critically exposes flaws in the “gold standard” CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimise plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      Weaknesses:

      The study is conducted exclusively using Escherichia coli. While E. coli is a standard model organism, the paper claims to evaluate luminescence as a generalisable high-throughput tool. Many of the discrepancies observed are driven by filamentation. However, distinct morphological responses occur in other critical pathogens (e.g., Staphylococcus aureus does not filament in the same way).

      The authors propose that luminescence data can be corrected using microscopyderived volume data to better align with CFU counts. The primary appeal of luminescence is high-throughput efficiency. If a researcher must perform timelapse microscopy to calculate cell volume changes to “correct” their luminescence data, the high-throughput advantage is lost.

      The paper argues that for ciprofloxacin, CFU underestimates viability because cells remain intact and impermeable to propidium iodide. While the cells are metabolically active and membrane-intact, if they cannot divide to form a colony (even after drug removal/dilution), their clinical relevance as “living” pathogens is debatable.

      Some other comments:

      The use of a population dynamical model to simulate filamentation effects is excellent. The finding that light intensity tracks volume ($\psi_V$) better than cell number ($\psi_B$) is a key theoretical contribution.

      The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      The use of bootstrapping to estimate rate distributions is appropriate and robust.

      Conclusion:

      Muetter et al. provide a compelling argument that luminescence is a reliable, highthroughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Generalisability. We agree that the alignments and divergences reported for specific drugs may not transfer directly to other species, which may elongate differently (e.g. cocci) or show different physiological responses to treatment. Constructing analogous model strains — for example based on S. aureus to cover a broader range of morphologies and clinically relevant species would therefore be an interesting follow-up project, and we have adjusted the Discussion to make this clearer. We nevertheless believe that the broader conclusions (larger cells emit more light) of the paper likely hold across species.

      Volume correction. We agree that requiring microscopy would undermine the high-throughput advantage of the luminescence assay. It was not our intention to propose this as a practical approach, nor to imply that the luminescence signal needs a correction. Taken on its own, the signal can be interpreted as the cumulative metabolic output of the population, which is closely linked to biomass, and that measure is valuable in itself for many applications. We used the volume correction only to demonstrate that luminescence tracks biomass more closely than cell number: by adjusting the luminescence distribution with the measured volume change, it moves towards the CFU distribution. We have revised the Discussion to prevent this from being misunderstood as a required step.

      Culturability vs. clinical relevance. We agree that the dynamics of culturable cells are highly relevant, especially in a clinical context. Our aim was to explain the observed differences between CFU and luminescence by highlighting that culturability and viability are not always identical, without implying that one measure is inherently superior to the other — we leave it to the reader to decide which metric best suits their needs.

      Linear elongation. The model assumes linear elongation for mathematical convenience, which, depending on the specific strain and drug mechanism, could be incorrect. Its purpose is to demonstrate that a shift of the mean cell volume to a new, higher equilibrium under treatment can cause an initial peak in the luminescence signal despite a declining population. This remains true for non-linear elongation models, though the shape, height and position of the peak may change. We have adjusted the Results to make this clearer.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors present luminescence as a practical measurement of population decline under antibiotic exposure. One aspect that could be clarified is how the method behaves when tolerance arises from phenotypic heterogeneity, such as the presence of small, metabolically quiet survivors. Because luminescence reflects metabolic activity and biomass, the signal will be dominated by metabolically active cells, making rare tolerant subpopulations difficult to detect. A short discussion of how luminescence performs in these heterogeneous scenarios, and whether complementary assays are needed to capture long-lived tolerant cells, would strengthen the manuscript.

      Yes, that is a valid concern and we thank the reviewer for raising this point.

      Heterogeneity in cell-specific luminosity alone does not bias population-level rate estimates. A bias can arise, however, when specific luminosity correlates with a second factor — most importantly, the decline rate under treatment.

      We agree with the reviewer’s suggestion that brighter cells plausibly die faster than tolerant, metabolically quiet ones. When one subpopulation dominates the light signal, we expect minimal bias, as the rate estimate primarily reflects that subpopulation. However, in a transition phase when both subpopulations contribute roughly equally to the light signal, luminescence likely overestimates the decline.

      We added a corresponding caveat to the Discussion (lines 581–583).

      (2) The manuscript shows that filamentation can influence ψ_I by altering biomass and metabolic activity independently of cell number. However, antibiotic exposure can also trigger other stress responses and metabolic shifts that change energy fluxes, redox balance, and biosynthetic activity. Since luminescence depends on metabolic state and substrate availability, these additional physiological transitions may also affect ψ_I in ways not directly tied to birth or death processes. It would be useful to comment on whether such responses, beyond filamentation, are likely to influence luminescence dynamics across different drug classes or treatment conditions.

      We thank the reviewer for raising this point and agree that there is no biological law strictly linking luminosity to a single population property such as biomass or cell number, and changes in the metabolism most likely affect Ψ<sub>I</sub> as well.

      Transitioning to a new metabolic steady state biases Ψ<sub>I</sub>; once the new steady state is reached, however, the rate estimate should no longer be affected.

      Looking across drug classes, drugs that primarily lyse cells (polymyxins and, to a lesser degree, beta-lactams targeting PBP1) did not show noticeable deviations between Ψ<sub>I</sub> and Ψ<sub>CFU</sub>, and — perhaps counterintuitively — neither did ribosome-inhibiting drugs.

      For the remaining cases, we were able to attribute part of the discrepancy between Ψ<sub>I</sub> and Ψ<sub>CFU</sub> to changes in biomass or loss of culturability, though drug-induced metabolic changes may also contribute to the residual differences.

      We clarify this in the Discussion (lines 569–579).

      (3) The authors quantify survival using colony counts on drug-free medium. Because filamentation can be a reversible state that persists during antibiotic exposure, plating on drug-free medium may capture recovery potential rather than in-treatment viability. Filamented or stressed cells that cannot divide in the presence of a drug may nevertheless form colonies once the drug is removed. Clarifying how this recovery step affects ψ_CFU would help readers interpret differences between luminescence-based and colony-based measurements, especially in cases where transient tolerant states are present.

      We thank the reviewer for raising this point.

      Our CFU assay estimates the number of culturable cells at each time point; the rate Ψ<sub>CFU</sub> is then inferred from how this number changes across time points. Plating on drug-free medium is intentional, as it maximises the probability that a culturable cell is detected at each snapshot. Whether those cells would have continued dividing or died under continued treatment is captured by the subsequent time points.

      Filamentation interacts with the probability of colony formation in several, partly opposing ways:

      (1) It can increase the death rate, as for ceftazidime and cefepime, which is part of the kill effect captured by Ψ<sub>CFU</sub>;

      (2) Entanglement between filaments may reduce the number of colonies per plated bacterium;

      (3) Conversely, if a filament divides upon drug removal, its fragments form a cluster that — stochastically — is very likely to produce one (but not multiple) colony.

      The only scenario in which CFU could overestimate bacterial density is if a filament separates into individual cells in the liquid phase before plating; we have no indication that this occurred in our experiments.

      We addressed this concern in our response to the public comment.

      (4) A brief discussion comparing luminescence to fluorescent reporter systems could be helpful. Fluorescent proteins typically require a chromophore maturation step before becoming detectable, which introduces a delay between the underlying cellular event and the appearance of the signal. In contrast, as far as I understand, lux reporters emit light immediately once the enzymatic components and substrates are present, without a maturation stage. Highlighting this distinction may help readers understand why luminescence is well-suited for tracking rapid changes in population physiology under antibiotic exposure. However, the manuscript also notes that luminescence can lag slightly behind very rapid killing (particularly for AMPs), but the temporal dynamics of signal shutdown are not explored in detail. Because lux reflects metabolic activity rather than viability, a short delay between irreversible damage and the loss of light is biologically expected. It may help readers if the authors could expand on the mechanism underlying this delay in order to clarify when ψ_I is likely to track true biomass decline and when residual metabolic activity might mask early killing events.

      On fluorescent reporters:

      We thank the reviewer for this suggestion.

      Under some conditions, change rates can also be measured using fluorescence, provided the number of fluorescent molecules per bacterium remains constant. This requires a balance between production, maturation, degradation and dilution, which is only established if the growth rate and conditions remain constant over a sufficiently long period (typically hours).

      For measuring population decline, however, the key issue is that cell death does not inactivate fluorescent proteins: once matured, they emit independently of the cell’s metabolic state and decay only with the protein’s half-life, which is typically slower than the kill rates of interest.

      We added a clarification to the Introduction (lines 58–60).

      On the lux signal lag:

      We thank the reviewer for raising this point. The short lag between luminescence and CFU decline could in principle arise from two mechanisms: (i) luminescence declining more slowly than the actual cell number (residual light from dead cells), or (ii) CFU declining more steeply than the actual cell number (damaged but still viable cells failing to form colonies).

      Mechanism (i) splits into two sub-cases:

      (i.a) Dead but impermeable — the lux reaction could in principle continue for a short while if enough components are retained in the cell. However, a metabolically active, impermeable cell is difficult to classify as dead in the first place, making this scenario conceptually awkward.

      (i.b) Dead and permeable (lysed) — the lux components dilute into the medium, and by mass-action the reaction rate should drop rapidly (though not instantly). Any residual signal after lysis should therefore be short-lived.

      Mechanism (ii) — damaged (e.g. permeable) cells may be particularly sensitive to plating on agar (e.g. due to oxidative stress), resulting in a declining probability of colony formation.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse, making (i.b) and/or (ii) the likely explanations. Based on our experimental data, we cannot distinguish between these possibilities and therefore limit ourselves to reporting the observed discrepancy.

      We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded the discussion there.

      (5) In lines 85–89, the authors state that “high-throughput OD and luminescence measurements at sub-MIC concentrations provide valuable insights into drug effects on growth rates, [but] the super-MIC range is clinically more relevant,” and they present luminescence as a way to investigate super-MIC population dynamics. While super-MIC behaviour is indeed important for pharmacodynamics and resistance evolution, it is not clear that the specific luminescence implementation used here has direct clinical relevance. The study relies on a chromosomally integrated reporter in a laboratory strain, and the manuscript does not demonstrate that this approach can be applied to clinical isolates or diagnostic workflows. It may be helpful to moderate the claim of “clinical relevance” and frame the method more clearly as a high-throughput experimental tool that can inform clinically relevant questions, rather than as an assay ready for clinical application.

      We agree and have moderated the framing accordingly (lines 100–104).

      Reviewer #2 (Recommendations for the authors):

      (1) The conclusions regarding “biomass vs. cell number” may not apply equally to non-rod-shaped bacteria or species with different stress responses. The authors must explicitly discuss this limitation in the Discussion.

      The broad conclusion that bigger cells emit more light likely holds across morphologies, since it rests on the principle that more cellular material means more metabolic activity and therefore more light. The quantitative relationship between cell size and luminosity, however, may differ across species, shapes and conditions, for two reasons. First, chromosome copy number: whether drug-induced morphological changes are accompanied by chromosome replication and therefore an increase in lux operon copy number — varies across species and drug mechanisms. Second, the surface-to-volume ratio likely modulates mass-specific metabolism; some morphological changes preserve it (e.g. purely lateral elongation) while others do not.

      The more specific conclusions about which drug classes produce alignment or divergence between CFU and luminescence may also not transfer directly, as drug mechanisms can act differently across species.

      We already note this limitation in the Discussion (lines 594– 597) and have expanded the wording there.

      (2) The manuscript should clarify that luminescence is a superior metric for biomass without correction, rather than framing the volume correction as a necessary step to mimic CFU. The divergence should be embraced as a feature (biomass tracking), not a bug that needs fixing via labor-intensive microscopy.

      We agree with the framing and will make it clearer; it was actually our intention to clarify which method does what, rather than judge one as better or worse.

      We removed the “correction” sentence from the Discussion to make this clearer.

      (3) The authors should refrain from definitively stating CFU “underestimates” viability and instead use more precise terminology, such as “reproductive capability” vs. “metabolic integrity.”

      We agree with the reviewer that measuring culturability is a property, not a flaw, of CFU. Our intention was to emphasise that when CFU is used as a proxy for viability (which it often is), it can yield lower values than the actual number of survivors. We tried to make that distinction explicit in the manuscript (e.g. in lines 317–322).

      We would also like to note that in the case of antimicrobial carryover, CFU can genuinely underestimate culturability itself, not only viability.

      Regarding the suggested reproductive capability vs. metabolic integrity framing: we agree that metabolism and luminescence are closely linked. What held us back from drawing that link directly is that metabolism is hard to quantify, being the cumulative output of a diverse set of processes.

      (4) The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      Linear elongation is a mathematically convenient simplification whose only purpose in the model is to allow the population to converge to a new equilibrium volume under treatment. Assuming constant volume-specific luminosity, we showed that this produces an initial peak in light intensity before the signal declines in parallel with Ψ<sub>B</sub>. The exact shape, height and position of this peak depend on the volume growth model used, but the qualitative pattern — peak followed by parallel decline — holds for other growth models as well. We now clarify this in lines 230–235.

      (5) The authors suggest the carryover effect is due to a delay between cell death and cessation of luminescence. This “lag time” is a critical physical constraint of the lux system (likely related to ATP depletion or enzyme decay) and should be quantified or discussed in more detail as a fundamental “speed limit” for the assay.

      The origin of the lag between luminescence and CFU is an interesting question, but one we cannot definitively answer. We can, however, discuss the potential mechanisms:

      A dead but impermeable cell could in principle continue to emit residual light for some time. We note, though, that calling a metabolically active, impermeable cell “dead” is a question of definition we would rather not discuss here.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse. Under lysis, the lux components dilute quickly into the medium, and by mass-action the reaction rate should drop rapidly — though not necessarily instantaneously.

      A plausible alternative to a delayed cessation of the light signal is that the probability of colony formation drops rapidly after permeabilisation, for example because permeable cells are sensitive to oxidative stress when plated on agar.

      Based on our experimental data we cannot distinguish between these mechanisms, so we limit ourselves to reporting the observed discrepancy. We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded on the candidate mechanisms there.

      Additional revisions

      Beyond the changes prompted by the reviewers’ comments, we made the following revisions to the supplementary information:

      We corrected the Λ matrix (converted row 2, col 4 from 0 → 2)

      We removed the line numbering

    1. eLife Assessment

      This valuable study introduces MULTI i<sup>2</sup>, a robust and high-throughput method to measure Plasmodium falciparum viability in the presence of drugs. This new assay offers significant time savings over the traditional Parasite Reduction Rate (PRR) assay and should enable faster screening of drug combinations, which is urgently needed in the field. The assay is well validated, with convincing data showing it can reproduce known drug interactions and identify new interaction patterns.

    2. Reviewer #1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRRv2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRRv2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRRv2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      There are a number of areas for improvement:

      (1) Many antimalarials have quite specific times of action. Are these MULTI-i2 assays, and the comparator PRRv2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      (2) The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULTI-i2 method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      (3) It would be helpful for authors to provide some indication of the cost comparison between the PPRv2 and MULTI-i2.

      (4) Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      (5) The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

    3. Reviewer #2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRRv2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i2 assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i2 assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i2 assay.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i2 assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRRv2 assay?

      The addition of an inducible element is an improvement of their earlier lacZ/β-galSENSOR (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRRv2, they fail to compare it to their own non-inducible lacZ/β-galSENSOR system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved? How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

    4. Reviewer #3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULTI-i2, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i2 assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULTI-i2 assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULTI-i2 provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULTI-i2 methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1) The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      (2) Related to that above, how would MULTI-i2 perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      (3) Given the stated cost and labor efficiency of MULTI-i2, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i2 method more attractive. In particular, it would be nice to see if one could use MULTI-i2 for studies of triple combinations as enthusiastically suggested.

      (4) Throughout the manuscript, the authors claim that MULTI-i2 is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

    5. Author response:

      Reviewer 1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we will also explicitly state this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we will mention and describe alternative approaches for assessing parasite viability. This will also include the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We will revise the text to explicitly mention the use of dual straining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al, 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer 2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We will clarify this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We will add this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We will revise the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      Reviewer 3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1): The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We will add more explanations to the Discussion including the strengths and limitations.

      (2) Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      (3) Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      (4) Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. eLife Assessment

      This important study combines anatomical tracing, tissue clearing, and functional manipulations to demonstrate lateralized brainstem control of hepatic glucose metabolism and identify a site of sympathetic nerve crossover supplying the liver. The evidence supporting the anatomical organization of hepatic sympathetic innervation is compelling, and the functional studies provide solid support for a role of asymmetric sympathetic outflow in regulating glucose homeostasis. While some uncertainty remains regarding the contribution of sensory innervation and the extent to which these findings generalize beyond mice, the work provides an invaluable advance in understanding neural regulation of liver metabolism.

    2. Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether hepatic glucose metabolism is differentially regulated by the left and right sides of the LPGi and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi, which were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, and in changes in protein expression in the liver lobes. These data suggested lobe-specific modulation of HGP. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) Image clarity was improved in some cases, but not in others. For example, Figure 3I, showing c-Fos expression, is not convincing due to the image quality and lack of orientation.

      (2) The methods section states that 8-weeks-old male mice were used in the experiments without specifying the experiments (e.g., brain injection with AAVs or PRV organ inoculation). The authors should include these details.

      (3) The authors should use the exact location of pre- and postganglionic neurons as they often refer to neurons in the sympathetic chain. Their findings should be compared with the existing literature on the location of preganglionic cells.

      (4) Figure legends should be revised and matched with the text.

    3. Reviewer #4 (Public review):

      Summary of General Strengths & Weaknesses:

      The studies here are highly informative for anatomical tracing and sympathetic nerve function in the liver in relation to glucose levels, but because they are conducted in a single species, it is challenging to translate them to humans or determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies provides mechanistically informative. Denervation studies lack proper controls, and sensory innervation in the liver is overlooked.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also labels a portion of sensory fibers that need to be ruled out in whole-mount imaging data.

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases in tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though that is clearly assumed. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      Comments on the revised version.

      Across all reviewer comments, the revised resubmission has adequately addressed all concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Wang et al. reports the potential involvement of an asymmetric neurocircuit in the sympathetic control of liver glucose metabolism.

      Strengths:

      The concept that the contralateral brain-liver neurocircuit preferentially regulates each liver lobe may be interesting.

      Weaknesses:

      However, the experimental evidence presented did not support the study's central conclusion.

      We thank the reviewer for recognizing the conceptual novelty of our work and for constructive comments aimed at enhancing its rigor and clarity. In response, we carried out targeted experiments to address the points raised, including: (i) further characterization of LPGi projections to vagal and sympathetic circuits; (ii) evaluation of potential pancreatic involvement; and (iii) validation of the specificity of chemogenetic activation within the proposed circuit. All new experiments, figures, and text have been incorporated, and corresponding revisions are highlighted for ease of review.

      (1) Pseudorabies virus (PRV) tracing experiment:

      The liver not only possesses sympathetic innervations but also vagal sensory innervations. The experimental setup failed to distinguish whether the PRV-labeling of LPGi (Lateral Paragigantocellular Nucleus) is derived from sympathetic or vagal sensory inputs to the liver.

      Thank you for raising this important point. We fully agree that the liver receives both sympathetic and vagal sensory innervation, and we acknowledge that PRV-based tracing alone does not definitively distinguish between these two pathways. This represented a limitation of the original experimental design.

      Based on established anatomical literature as well as our experimental observations, vagal sensory neuron cell bodies reside in the nodose ganglion (NG), and their central projections terminate predominantly in the nucleus of the solitary tract (NTS) (Nature. 2023;623(7986):387-396; Curr Biol. 2020;30(20):3986-3998.e5.), which is located in the dorsomedial medulla. In contrast, the LPGi, together with other sympathetic-related nuclei, is predominantly distributed in the ventral medulla (Cell Metab. 2025;37(11):2264-2279.e10; Nat Commun. 2022;13(1):5079).

      To determine whether the LPGi contains neurons that modulate the liver via vagal sensory pathways, we performed two complementary experiments.

      First, we conducted CGRP immunohistochemistry on brainstem sections, using the NTS, a well-established visceral sensory centre, as a positive control. While abundant CGRP-positive cell bodies were detected in the NTS as expected, few to no CGRP-positive cell bodies were observed in the LPGi (Figure S1G). These results strongly support that the LPGi neurons labeled in our PRV tracing predominantly belong to sympathetic efferent circuits rather than vagal sensory pathways.

      Second, to examine whether LPGi neurons send axonal projections to sensory ganglia, we injected hSyn-Cre combined with DIO-Axon-EGFP into the LPGi and examined both the dorsal root ganglia (DRG) and nodose ganglia (NG). No Axon-EGFP-positive signals were detected in either ganglion (Figures S1H-S1J), indicating that LPGi neurons do not directly innervate sensory ganglia. In other words, PRV cannot retrogradely trace to the LPGi via the NG or DRG.

      These additions have been incorporated into the revised manuscript, with the Result 1 clearly documenting that these findings confirm that the LPGi specifically regulates sympathetic, rather than vagal sensory, inputs to the liver.

      (2) Impact on pancreas:

      The celiac ganglia not only provide sympathetic innervations to the liver but also to the pancreas, the central endocrine organ for glucose metabolism. The chemogenetic manipulation of LPGi failed to consider a direct impact on the secretion of insulin and glucagon from the pancreas.

      Thank you for this important comment. We agree that the celiac ganglia (CG) provide sympathetic innervation not only to the liver but also to the pancreas, which plays a central role in glucose homeostasis through the secretion of both insulin and glucagon. Therefore, the potential pancreatic implications associated with LPGi chemogenetic manipulation are worth careful consideration.

      To address this concern, we measured circulating glucagon and insulin levels following chemogenetic manipulation of the LPGi<sup>GAD1</sup> neurons. We found that neither glucagon nor insulin levels changed significantly under our experimental conditions, which indicated that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated by changes in pancreatic hormone secretion (Figure S2G).

      These additions have been incorporated into the revised manuscript, with the Result 2 clearly documenting that these findings confirm that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated indirectly via altered pancreatic endocrine output.

      (3) Neuroanatomy of the brain-liver neurocircuit:

      The current study and its conclusion are based on a speculative brain-liver sympathetic circuit without the necessary anatomical information downstream of LPGi.

      Thank you for raising this important point. A clear anatomical definition of the downstream pathways linking the brain to the liver was essential for interpreting the proposed brain-liver sympathetic circuit.

      The present study (Figure 4A) provides direct anatomical evidence supporting the organization of the brain–liver sympathetic neurocircuit. These observations are consistent with our recent detailed characterization of the brain-liver sympathetic circuit published in Cell Metabolism (Cell Metab. 2025;37(11):2264–2279). In that study, we showed that LPGi GABAergic neurons inhibit GABAergic neurons in the caudal ventrolateral medulla (CVLM). Disinhibition of CVLM reduced GABAergic suppression of rostral ventrolateral medulla (RVLM) neurons, which are key excitatory drivers of sympathetic tone. RVLM neurons project to sympathetic preganglionic neurons in the sympathetic chain (Syc). These neurons synapse with postganglionic sympathetic neurons in ganglia such as the celiac-superior mesenteric ganglion (CG-SMG). Postganglionic sympathetic fibers then innervate the liver, releasing norepinephrine (NE) to activate hepatic β<sub>2</sub>-adrenergic receptors and stimulate hepatic glucose production (HGP).

      Together, these data establish a coherent anatomical basis for the proposed brain-liver sympathetic pathway and clarify the downstream organization relevant to the functional experiments presented in figure 4A and Author response image 1..

      Author response image 1.

      Tracing scheme (Left) and whole-mount imaging (Right) of PRV-labeled brain-liver neurocircuit. Scale bars, 3,000 (whole mount) or 1,000 (optical sections) μm.

      (4) Local manipulation of the celiac ganglia:

      The left and right ganglia of mice are not separate from each other but rather anatomically connected. The claim that the local injection of AAV in the left or right ganglion without affecting the other side is against this basic anatomical feature.

      Thank you for raising this important anatomical point. We fully acknowledge that the left and right CG in mice are interconnected, and that unilateral viral injection could theoretically affect the contralateral side. The CG-SMG complex serves as a major sympathetic hub that regulates visceral organ functions. Recent transcriptomic, anatomical, and functional studies have revealed that the CG-SMG is not a homogeneous structure but is composed of molecularly and functionally distinct neuronal populations. These populations exhibit specialized projection patterns and regulate different aspects of gastrointestinal physiology, supporting a model of modular sympathetic control. (Nature. 2025 Jan;637(8047):895-902). Therefore, we were aware of this phenomenon during the initial stages of these experiments.

      To minimize unintended spread to the contralateral CG, we took two complementary approaches.

      First, we optimized the injection strategy by using an extremely small injection volume (100 nL per site), with a very slow infusion rate (50 nL/min), and fine glass micropipettes. With these refinements, contralateral viral spread was rarely observed.

      Second, and importantly, all animals included in the final analyses were subjected to post hoc anatomical verification. After completion of the experiments, CGs were collected, sectioned, and examined for viral expression. As shown in Supplementary Figure 5F, only mice in which viral expression was strictly confined to the targeted CG, with no detectable infection in the contralateral ganglion, were included in the presented data.

      Together, these measures ensure that our local manipulation of the intended CG produced the reported effects. We have revised the Methods section to more explicitly detail these technical precautions, and the legend for Figure S5F clearly states its role in validating injection specificity.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether the left and right LPGi differentially regulate hepatic glucose metabolism and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, as well as changes in protein expression in the liver lobes. These data suggested modulation of HGP (hepatic glucose production) in a lobe-specific manner. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      We thank the reviewer for the thorough and constructive evaluation of our manuscript. In direct response, we undertook comprehensive revisions to enhance the rigor and clarity of the study, including: (i) correcting ambiguous or misleading terminology about anatomical resolution and sympathetic circuit organization; (ii) expanding the Methods section with complete experimental details, improved image presentation, and explicit justification of our viral and genetic approaches; and (iii) strengthening data interpretation by addressing issues related to sparse PRV labeling, projection heterogeneity, and the functional implications of double-labeled neurons.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) The wording/terminology used in the manuscript is misleading, and it is not used in the proper context. For instance, the goal of the study is "to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism..." (see abstract); however, the authors focus on the brainstem (a single structure without hemispheres). Similarly, symmetric is not the best word for the projections.

      We thank the reviewer for raising these critical points regarding terminology and conceptual framing. We acknowledge that certain phrases in our original manuscript may have been overly broad or ambiguous, particularly in describing the scope of sympathetic heterogeneity and the specificity of neural projections. Due to practical constraints and the scope of our study, our investigation focused on the brainstem, which represents the final common pathway for these lateralized commands. We acknowledge that terms referring to the cerebral hemispheres do not accurately describe our study. We have revised the manuscript to ensure accurate and consistent terminology.

      Below are specific examples:

      Original 1: This study aims to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism and localize the site of sympathetic crossover to the liver.

      Revised 1: “This study aimed to determine whether the central nervous system exerts lateralized control over hepatic glucose metabolism and to localize the site of peripheral sympathetic crossover to the liver.”

      Original 2: These findings demonstrate that the brain exerts lobe-specific, lateralized control of hepatic glucose metabolism via symmetric brain-liver sympathetic pathways.

      Revised 2: “These findings demonstrate that the brainstem can exert lobe-specific, lateralized control of hepatic glucose metabolism via bilaterally projecting brain–liver sympathetic pathways.”

      Original 3: The cerebral hemispheres exhibit pronounced functional asymmetry, [1,2] a phenomenon traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation.[3,4]

      Revised 3: “Pronounced functional lateralization within the central nervous system (CNS) is a well-documented phenomenon, [1,2] traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation [3,4].

      (2) Sparse labeling of liver-related neurons was shown in the LPGi (Figure 1). It would be ideal to have lower magnification images to show the area. Higher quality images would be necessary, as it is difficult to identify brainstem areas. The low number of labeled neurons in the LPGi after five days of inoculation is surprising. Previous findings showed extensive labeling in the ventral brainstem at four days post-inoculation (Desmoulins et al., 2025). Unfortunately, it is not possible to compare the injection paradigm/methods because the PRV inoculation is missing from the methods section. If the PRV is different from the previously published viral tracers, time-dependent studies to determine the order of neurons and the time course of infection would be necessary.

      We sincerely thank the reviewer for these detailed and constructive comments regarding the PRV tracing experiments. We fully agree that careful presentation and interpretation of the anatomical data are essential for ensuring rigor and transparency. We address each point in detail below.

      (1) Image magnification and anatomical context of LPGi labeling

      We agree that the original images did not sufficiently convey the broader anatomical context of the LPGi. Due to fluorescence quenching in previous sections, we repeated the PRV retrograde tracing experiment and performed statistical analysis. In the revised manuscript, we replaced the original panels in Figure 1 and Figure S1 with new images that include lower-magnification overviews of the brainstem, alongside higher-magnification views of the LPGi (Figure 1). These images clearly delineate the LPGi with respect to established anatomical landmarks and atlas boundaries. Image contrast and resolution were optimized to allow unambiguous identification of PRV-labeled neurons and surrounding structures.

      (2) Sparse LPGi labeling at 5 days post-injection and methodological details

      We apologize for the omission of the detailed PRV injection protocol in the original Methods section. We deliberately used small-volume, local injections (1 µL per liver lobe) to minimize viral spread and to restrict labeling to circuits specifically connected to the targeted hepatic region. This sparse labeling was consistent with the use of small, spatially restricted injections designed to minimize off-target spread and preferentially label higher-order upstream neurons. This information has now been added, including the PRV strain, viral titer, injection volume, precise injection coordinates, and surgical procedures. All new figures, legends, and Method details have been incorporated, with changes clearly highlighted for ease of review.

      These additions have been incorporated into the revised manuscript, in Figure 1B-D, Figure S1C-F. Furthermore, we also added details of the Methods.

      (3) Not all LPGi cells are liver-related. Was the entire LPGi population stimulated, or was it done in a cell-type-specific manner? What was the strain, sex, and age of the mice? What was the rationale for using the particular viral constructs?

      We thank the reviewer for this insightful and important question. We agree that not all neurons within the LPGi are liver-related, and we apologize that our rationale was not clearly articulated in the original manuscript.

      (1) Our decision to target GABAergic neurons in the LPGi using GAD1-Cre mice was based on prior experimental evidence rather than an assumption about the entire LPGi population. In our previous study (Cell Metab. 2025;37(11):2264-2279.e10), we performed single-cell RNA sequencing on retrogradely labeled LPGi neurons following liver tracing. These analyses revealed that the majority of liver-projecting LPGi neurons are GABAergic in nature. Based on these findings, we chose to selectively manipulate GABAergic neurons in the LPGi rather than the entire LPGi neuronal population, to achieve greater cellular specificity and to minimize potential confounding effects arising from heterogeneous neuron types within this region. We regret that this rationale was not clearly described in the original submission and have now revised the manuscript to explicitly state this reasoning (Results section 2, paragraph 2: “Prior single-nucleus RNA sequencing and immunofluorescence analyses demonstrated that liver-projecting LPGi neurons are predominantly GABAergic.”).

      (2) In addition, we apologize for the omission of mouse strain, sex, and age information in the Methods section. These details have been fully added.

      (3) We selected AAV-based viral vectors, specifically the AAV9 serotype, due to their well-established efficiency in transducing neurons in the brainstem, relatively low toxicity, and widespread use in circuit-level chemogenetic and optogenetic studies. When combined with Cre-dependent viral constructs in GAD1-Cre mice, this approach enabled selective and reliable manipulation of LPGi GABAergic neurons.

      (4) The authors should consider the effect of stimulation of double-labeled neurons (innervating more than one lobe) and potential confounding effects regarding other physiological functions.

      We thank the reviewer for raising this important point. We agree that neurons innervating more than one liver lobe could, in principle, introduce potential confounding effects and may reflect higher-order integrative autonomic neurons.

      This consideration is consistent with a key finding of the cited study: the CG-SMG contains molecularly distinct sympathetic neuron populations (e.g., RXFP1<sup>+</sup> vs. SHOX2<sup>+</sup>) that exhibit complementary organ projections and separate, non‑overlapping functions. Specifically, RXFP1<sup>+</sup> neurons innervate secretory organs (pancreas, bile duct) to regulate secretion, while SHOX2<sup>+</sup> neurons innervate the gastrointestinal tract to control motility. This functional segregation supports the concept of specialized autonomic modules rather than a uniform, “fight-or-flight” response, reinforcing the need for careful interpretation of circuit-specific manipulations. (Nature. 2025;637(8047):895-902; Neuron. 2026;114(3):463-478.e7).

      In our PRV tracing experiments, the proportion of double-labeled neurons was relatively small, suggesting that the majority of labeled LPGi neurons preferentially associate with individual hepatic lobes. Nevertheless, we recognize that activation of this minority population could contribute to broader physiological effects beyond strictly lobe-specific regulation. We have therefore added a paragraph in the second paragraph of the Discussion (Paragraph 2: “A small subset of LPGi neurons was double-labeled after bilateral PRV injections, suggesting a fraction of these neurons projects bilaterally to both sides of the liver. Such neurons may support interlobar coordination.”).

      (5) The authors state that "central projections directly descend along the sympathetic chain to the celiac-superior mesenteric ganglia". What they mean is unclear. Do the authors refer to pre-ganglionic neurons or premotor neurons? How does it fit with the previous literature?

      We thank the reviewer for pointing out this imprecise wording. We agree that the original phrasing was anatomically inaccurate and potentially confusing. The pathways we intended to describe involve brainstem premotor neurons that project to sympathetic preganglionic neurons in the spinal cord. These preganglionic neurons then innervate neurons in the CG-SMG, which in turn provide postganglionic input to the liver.

      We have revised the manuscript to clearly distinguish premotor from preganglionic neurons (Results section 4, paragraph 1: “Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the spinal cord send descending fibers through the sympathetic chain (SyC) to innervate postganglionic neurons in the CG-SMG (Figure 4A). Further whole-mount TH immunostaining showed that TH-positive sympathetic cell bodies within the CG-SMG project to the liver along the hepatic vasculature (Figure 4B and Figure S5F). These observations suggest that the nerve bundles likely decussate at the porta hepatis before entering the individual hepatic lobes.”).

      (6) How was the chemical denervation completed for the individual lobes?

      We thank the reviewer for raising this important methodological concern. We agree that potential diffusion of 6-OHDA is a critical issue when performing lobe-specific chemical denervation, and we apologize that our original description did not sufficiently clarify how this was controlled.

      In the revised Methods section, we provided a detailed description of the denervation procedure, including the injection volume and concentration of 6-OHDA, as well as the physical separation and isolation of individual hepatic lobes during application to minimize diffusion to adjacent tissue.

      To directly assess the specificity of the chemical denervation, we included immunofluorescence and Western blot analyses demonstrating a selective reduction of sympathetic markers in the targeted lobe (Figure 3C), with minimal effects on non-targeted lobes. These results support the effectiveness and relative spatial confinement of the 6-OHDA treatment under our experimental conditions.

      We thank the reviewer for highlighting this point, which has helped us improve both the clarity and rigor of the manuscript.

      (7) The Western Blot images look like they are from different blots, but there are no details provided regarding protein amount (loading) or housekeeping. What was the reason to switch beta-actin and alpha-tubulin? In Figures 3F -G, the GS expression is not a good representative image. Were chemiluminescence or fluorescence antibodies used? Were the membranes reused?

      We thank the reviewer for this careful and detailed evaluation of the Western blot data. We apologize that insufficient methodological detail was provided in the original submission.

      (1) We would like to clarify that the protein bands shown within each panel were derived from the same membrane. To improve transparency, we provided full, uncropped images of the corresponding membranes in the supplementary materials. In addition, detailed information regarding protein loading amounts, gel conditions, and housekeeping controls has also been added to the Methods section.

      (2) The use of different loading controls (β-actin or α-tubulin) reflects a technical consideration rather than an experimental inconsistency. In our experiments, the molecular weight of the TH (62kDa) was too close to that of α-tubulin (55kDa), and β-actin (42kDa) was therefore used to avoid band overlap and to ensure accurate quantification.

      (3) Regarding the GS signal shown in Figures 3F–G, we agree that the original representative image was suboptimal. This appears to be related to antibody performance rather than sample quality. To address this, we repeated the Western blot from Figures 3F–G using a newly validated antibody. The original tissue samples had been aliquoted and stored at −80 °C, allowing reliable re-analysis.

      (4) All Western blot experiments were detected using chemiluminescence, and membrane stripping and reprobing procedures are now explicitly described in the Methods section.

      We thank the reviewer for highlighting these issues, which significantly improve the rigor and clarity of our data presentation. All new figures and legends have been incorporated, with changes clearly highlighted for ease of review.

      (8) Key references using PRV for liver innervation studies are missing (Stanley et al, 2010 [PMID: 20351287]; Torres et al., 2021 [PMID: 34231420]; Desmoulins et al., 2025 [PMID: 39647176]).

      We thank the reviewer for pointing out these important and highly relevant references that were inadvertently omitted in our initial submission. The studies by Stanley et al. (Proc Natl Acad Sci U S A, 2010), Torres et al. (Am J Physiol Regul Integr Comp Physiol, 2021), and Desmoulins et al. (Auton Neurosci, 2025) represent key PRV-based retrograde tracing work that has mapped central neural circuits innervating the liver and thus provide essential context for our anatomical analyses.

      We agree that the inclusion of these studies is necessary to properly situate our findings within the existing literature. Accordingly, we incorporated citations to these references in the revised manuscript and discussed their relationship to our results.

      Reviewer #3 (Public review):

      Summary:

      This study found a lobe-specific, lateralized control of hepatic glucose metabolism by the brain and provides anatomical evidence for sympathetic crossover at the porta hepatis. The findings are particularly insightful to the researchers in the field of liver metabolism, regeneration, and tumors.

      Strengths:

      Increasing evidence suggests spatial heterogeneity of the liver across many aspects of metabolism and regenerative capacity. The current study has provided interesting findings: neuronal innervation of the liver also shows anatomical differences across lobes. The findings could be particularly useful for understanding liver pathophysiology and treatment, such as metabolic interventions or transplantation.

      Weaknesses:

      Inclusion of detailed method and Discussion:

      We sincerely thank the reviewer for the positive and constructive feedback, which significantly enhances both the methodological rigor and the broader biological interpretation of our study. In direct response, we revised the Discussion to elaborate on the potential physiological advantages of a lateralized and lobe-specific pattern of liver innervation. Furthermore, we expanded the Methods section to include a comprehensive description of the quantitative analysis applied to PRV-labeled neurons. Together, these revisions strengthened the manuscript’s clarity, depth, and relevance to researchers in hepatic metabolism, regeneration, and disease.

      (1) The quantitative results of PRV-labeled neurons are presented, and please include the specific quantitative methods.

      We thank the reviewer for this helpful suggestion. We have added a detailed description of the quantitative methods used to analyze PRV-labeled neurons in the revised Methods section. We have now provided detailed information in the Methods section, including the criteria used for cell counting, the anatomical boundaries of the brain regions analyzed, the delineation of regions of interest, and the normalization procedures applied to derive the reported neuron counts. These additions have been incorporated into the revised Methods, with all changes clearly indicated for ease of review.

      (2) The Discussion can be expanded to include potential biological advantages of this complex lateralized innervation pattern.

      We appreciated the reviewer’s suggestion. We have expanded the Discussion to include a paragraph addressing the potential biological significance of lateralized liver innervation. We highlight that this asymmetric organization could allow for more precise, lobe-specific regulation of hepatic metabolism, enable integration of distinct physiological signals, and potentially provide robustness against perturbations. The additional discussion content has been highlighted in the revised version as indicated (Discussion section, paragraph 3: “Bilateral LPGi activation produced additive effects, indicating that both sides of the brainstem can cooperatively regulate hepatic metabolism in a spatially segregated manner. This pattern suggests that hepatic glucose output can be modulated in a lobe-specific, rather than uniform whole-organ, manner.”).

      Reviewer #4 (Public review):

      Summary:

      The studies here are highly informative in terms of anatomical tracing and sympathetic nerve function in the liver related to glucose levels, but given that they are performed in a single species, it is challenging to translated them to humans, or to determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies is mechanistically informative. Denervation studies lack appropriate controls, and the role of sensory innervation in the liver is overlooked.

      We sincerely appreciate the reviewer's thoughtful evaluation and fully agree that findings derived from a single-species model must be interpreted with caution in relation to human physiology. In direct response, we revised the manuscript to explicitly clarify that all experimental data were obtained in mice and to provide a discussion of the limitations regarding direct extrapolation to humans. Concurrently, we expanded the Discussion section by integrating our findings with recent human and translational studies, including a multicenter clinical trial demonstrating that catheter-based endovascular denervation of the celiac and hepatic arteries significantly improved glycemic control in patients with poorly controlled type 2 diabetes, without major adverse events (Signal Transduct Target Ther. 2025;10(1):371). While our current work focuses on defining the anatomical organization and functional asymmetry of this circuit in mice, the clinical findings suggest that the core principles, sympathetic control of hepatic glucose metabolism via CG-liver pathways, may be conserved and of translational relevance. Additionally, we clarified the interpretation of TH labeling and expanded the discussion of hepatic sensory and parasympathetic innervation, acknowledging their important roles in liver-brain communication and identifying them as key directions for future research. Collectively, these revisions provide a more balanced, clinically informed, and rigorous framework for interpreting our findings.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      We thank the reviewer for this suggestion. We agree that the species should be clearly indicated. The findings presented in this study were obtained in mice using tissue clearing and whole-organ imaging approaches. Due to technical limitations, these observations are currently restricted to the mouse strain. We have updated the title (Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice) and clarified the species used throughout the manuscript.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also hits a portion of sensory fibers that need to be ruled out in whole-mount imaging data

      We thank the reviewer for pointing this out. We acknowledge that TH labels not only sympathetic fibers but also a subset of sensory fibers. We have added a limitation of this point in the revised manuscript. In addition, using SyGlass (2.4.0) three-dimensional reconstruction, we observed TH-positive nerve fibers originating from the CG-SMG extending along the porta hepatis and penetrating into the liver parenchyma. Given that the CG-SMG is a well-established sympathetic ganglion innervating visceral organs (Nature. 2025 Jan;637(8047):895-902.), these nerve fibers can be definitively identified as sympathetic. In parallel, we collected DRG from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While T7-12 DRG are known to contain sensory neurons innervating the liver, only a sparse number of PRV-positive neurons were detected in these segments (Anat Rec A Discov Mol Cell Evol Biol. 2004 Sep;280(1):827-35. Auton Neurosci. 2024 Jun;253:103174). The additional figure and discussion content have been highlighted in the revised version as indicated (Discussion section, paragraph 6: “Third, although whole-mount TH immunostaining with three-dimensional reconstruction revealed sympathetic nerve bundles projecting from the CG to the liver, TH is not entirely specific and can also label a subset of sensory neurons. More selective approaches, such as genetic targeting of sympathetic lineages, will be important for further validation.”).

      Author response image 2.

      Representative immunofluorescence images of PRV-labeled neurons (EGFP) in DRG from the spinal segments T1-6 (bottom) and T7-12 (top) following PRV injections into the liver lobes. Scale bars, 200μm

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      We thank the reviewer for this suggestion. Previous studies largely relied on electrical stimulation to modulate liver innervation, which provides relatively coarse control of neural activity (Eur J Biochem. 1992;207(2):399-411). By contrast, our use of chemogenetic and optogenetic approaches allows selective, cell-type-specific manipulation of LPGi neurons. We revised the Discussion to place our functional data in the context of prior work, highlighting how these more precise approaches improve understanding of the contribution of liver-innervating neurons to hyperglycemia. The newly added discussion has been clearly labeled in the response to facilitate your review (Discussion section, paragraph 3: “This spatial organization is likely obscured by conventional electrical stimulation, which indiscriminately activates heterogeneous sympathetic fibers. By contrast, chemogenetic and optogenetic approaches permit selective, cell type-specific manipulation of LPGi neurons, thereby revealing the contralateral and lobe-specific architecture of brain-liver sympathetic control”).

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases to tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though it is clearly assumed to be. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We thank the reviewer for this insightful and important comment, which highlights a potential alternative interpretation of our findings. We agree that chemical sympathetic denervation with 6-OHDA may induce compensatory changes in non-sympathetic inputs, including sensory and parasympathetic (vagal) innervation of the liver.

      Conceptually, we agree with the reviewer’s perspective that the central nervous system operates as a highly integrated homeostatic regulatory system, continuously receiving and integrating a broad range of afferent signals. These inputs include, as noted by the reviewer, hepatic sensory and vagal afferents (Science. 2024;386(6722):673-677), as well as centrally derived interoceptive signals such as brain glucose, temperature sensing, even the pulsation of cerebral vascular system (Cell Metab. 2025;37(11):2264-2279.e10.; Cell Metab. 2022;34(6):888-901.e5; Science. 2024;383(6682):eadk8511). The CNS integrates these diverse signals and generates coordinated efferent outputs to maintain systemic homeostasis.

      From this viewpoint, the changes in c-FOS activity that we observe in the LPGi likely represent only a limited snapshot of this broader integrative process, rather than evidence of a single dominant pathway. We acknowledge that compensatory sensory or parasympathetic mechanisms, in addition to altered sympathetic drive, contributed to the observed LPGi activation following hepatic sympathetic denervation.

      We further acknowledge that, due to limitations in scope and experimental focus, we did not directly assess sensory or parasympathetic innervation of the liver in the present study. As appropriately pointed out by the reviewer, a more comprehensive characterization of hepatic neural inputs would provide a more complete picture of the underlying neurocircuitry. To address this, we expanded the Discussion and explicitly noted this limitation, including a more balanced discussion of potential crosstalk among sympathetic, sensory, and parasympathetic pathways and how these may collectively influence LPGi activity. For your convenience, the newly added discussion text has been distinctly marked in the manuscript (Discussion section, paragraph 4: “Although enhanced sympathetic output appears to mediate much of this compensation, our findings suggest that the underlying regulation extends beyond a purely descending pathway. In particular, c-FOS activation in the contralateral LPGi after unilateral 6-OHDA-mediated denervation suggests that the loss of peripheral input may be sensed through an ascending neural pathway, centrally integrated, and translated into compensatory sympathetic output to the intact hepatic lobes. These results therefore support a model in which hepatic glucose production is regulated by an integrated afferent-central-efferent loop, with our current analyses primarily resolving its efferent component.”).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Although the findings are interesting, this reviewer has major concerns about the experimental design, methodology, results, and interpretation of the data. Experimental details are lacking, including basic information (age, sex, strain of mice, procedures, magnification, etc.).

      We thank the reviewer for this important recommendation. We agree that comprehensive reporting of experimental details is essential for rigor and reproducibility.

      In the revised manuscript, we added complete information regarding mouse strain, sex, age, and sample size for each experiment. In addition, detailed descriptions of surgical procedures, viral constructs, injection parameters, imaging magnification, and analysis methods have been incorporated into the Methods section.

      These revisions ensured that all experiments are described with sufficient technical detail and clarity to allow accurate interpretation and replication of our findings. Experimental details have been incorporated, and corresponding revisions are highlighted for ease of review.

      Reviewer #3 (Recommendations for the authors):

      Addressing a few questions might help:

      (1) The study found that liver-associated LPGi neurons are predominantly GABAergic. It would be informative to molecularly characterize the PRV-traced, liver-projecting LPGi neurons to determine their neurochemical phenotypes.

      We thank the reviewer for this insightful suggestion. We agree that molecular characterization of liver-projecting LPGi neurons is important for understanding their functional identity.

      This issue has been addressed in detail in our recent study (Cell Metab. 2025;37(11):2264-2279.e10), in which we performed single-cell RNA sequencing on retrogradely traced LPGi neurons connected to the liver. These analyses demonstrated that the majority of liver-projecting LPGi neurons are GABAergic, with a defined transcriptional profile distinct from neighboring non–liver-related populations.

      Based on these findings, the current study selectively targeted GABAergic LPGi neurons using GAD1-Cre mice. We have explicitly cited these molecular results in the revised manuscript to clarify the neurochemical identity of the PRV-traced LPGi neurons. New text has been incorporated, and corresponding revisions are highlighted for ease of review.

      (2) Is it possible to do a local microinjection of a sodium channel blocker (e.g., lidocaine) or an adrenergic receptor antagonist into the porta hepatis? That would potentially provide additional evidence for the porta hepatis as the functional crossover point.

      We appreciated the reviewer’s thoughtful suggestion. Although pharmacological blockade at the porta hepatis can modulate local neural activity, this approach is inherently limited in its ability to distinguish between ipsilateral and contralateral inputs. Consequently, it may not provide definitive evidence for neural crossover at this specific site.

      In our view, the anatomical evidence provided by whole-mount tissue clearing, dual-labeled tracing, and direct visualization of decussating nerve bundles at the porta hepatis offers a more definitive demonstration of sympathetic crossover. Pharmacological blockade would affect both crossed and uncrossed fibers simultaneously and therefore would not specifically resolve the anatomical organization of this decussation.

      Nevertheless, we agree that functional interrogation of the porta hepatis represents an interesting direction for future work, and we acknowledge this possibility in the Discussion (Paragraph 6: “Fourth, although our data support a peripheral decussation at the porta hepatis, direct validation of this crossover site was not feasible with local pharmacological blockade, as currently available approaches lack sufficient spatial specificity and would likely perturb multiple neural components. Future studies employing more selective inhibitory strategies will be required to directly test this possibility.”).

      (3) It is possible to investigate the effects of unilateral LPGi manipulation or ablation of one side of CG/SMG on liver metabolism, such as hyperglycemia?

      We thank the reviewer for this important suggestion. Because unilateral LPGi manipulation was already examined in our study (Figure 2D), we focused here on unilateral ablation of the CG to further assess lateralized sympathetic control of hepatic metabolism. We successfully performed unilateral CG ablation without LPGi manipulation, but observed no significant change in blood glucose compared with the sham group (Author response image 3A and 3B). To determine whether glucose homeostasis was nonetheless affected, we further performed glucose tolerance tests (GTT) and insulin tolerance tests (ITT) (Author response image 3C and 3D). Neither test showed significant impairment after unilateral ablation, suggesting that compensatory neural mechanisms and/or hormonal homeostatic regulation may be recruited to preserve systemic glucose homeostasis.

      Author response image 3.

      (A) Blood glucose levels in mice subjected to left- or right-sided CG ablation via 6-OHDA treatment (n = 6). (B) Representative images of ablation of CG. Scale bars, 100 μm. (C and D) Blood glucose levels during GTT (C, n = 6) and ITT (D, n = 6) in mice with left- or right-sided CG ablation.

      Reviewer #4 (Recommendations for the authors):

      In the abstract and elsewhere, the use of the term 'sympathetic release' is unclear - do you mean release of nerve products, such as the neurotransmitter norepinephrine? This should be more clearly defined.

      We thank the reviewer for pointing out this ambiguity. We agree that the term “sympathetic release” was imprecise. In the revised manuscript, we explicitly referred to the release of sympathetic neurotransmitters, primarily norepinephrine, from postganglionic sympathetic fibers.

      We revised the wording throughout the manuscript to ensure accurate and consistent terminology and to avoid potential confusion regarding the underlying neurobiological mechanisms.

      Original: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased sympathetic release, glucose production, and glycogen depletion.”

      Revised: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased norepinephrine release, glucose production, and glycogen depletion.”

    1. eLife Assessment

      This valuable study demonstrates that self-motion strongly affects neural responses to visual stimuli, comparing humans moving through a virtual environment to passive viewing. The evidence for visuomotor mismatch responses is solid, although the interpretation in terms of prediction remains somewhat preliminary. This study bridges human and rodent studies on the role of prediction in sensory processing, and is therefore expected to be of interest to a large community of neuroscientists.

    2. Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference cannot be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features, but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course, the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course, this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information *from perception*. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Similarly, a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      Comments on latest version.

      Nice to see the added extra analyses. Can't see any more will be achieved via further rounds and happy with the summary to stand as is.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      - Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      - The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      - Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Comments on latest version.

      The authors added a brief discussion paragraph which addresses my previous comment.

    4. Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      Comments on latest version:

      The authors added useful points to the discussion and also included time frequency analyses to the paper formally, which strengthens the translational potential, in addition to the bolstering their claims slightly.

    5. Author response:

      The following is the authors’ response to the previous reviews

      We thank you for the time you took to review our work and for your feedback! The main changes to the manuscript are:

      We added a paragraph to the Discussion addressing differences in visuomotor mismatch responses recorded over frontal and occipital electrodes, and their possible interpretation.

      We added time-frequency power and phase-locking analysis as supplementary figures to the manuscript.

      We added a statement in the Discussion emphasizing the importance of performing these experiments with denser EEG channel coverage.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference can not be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active. I have a couple of questions/comments.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      The reviewer is correct. But one should keep in mind that our result here stands in the context of a considerable amount of work on mouse cortex investigating responses to very similar visuomotor mismatches. There we can we have much additional evidence to argue that the cortical response to a visuomotor mismatch is a prediction error. We would argue, it is the best one can do in human experiments.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      We have added a paragraph to the Discussion addressing the differences between signals recorded over frontal and occipital electrodes, as suggested.

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      We are not sure what the reviewer is referring to here. We certainly do not think (or are aware of having argued) that a visuomotor prediction error is the only possibly interpretation of the response. In the last paragraph of our response to the reviewers point 3 in the last revision, we explicitly discuss that the interpretation of the responses as a prediction error is only one possible interpretation. Our argument is that it is the most likely given the evidence.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      We would be happy to add additional considerations to other processes. If the reviewer has anything specific in mind, we can add that, but it would need to be somewhat concrete with some theoretical basis. We share the reviewer’s intuition, but unless this can be formalized to the point of being experimentally testable, we do not see any value in discussing it in the manuscript.

      Regarding the reason for a difference in visual responses in walking vs sitting state is, this is not entirely clear to us. Predictive processing would provide one possible explanation. Assuming the precision weighting of predictions is higher during walking, the sudden appearance of a visual stimulus might lead to stronger stimulus history prediction errors than when just sitting. But this is rather speculative.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information ‘from perception’. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Please excuse, the reviewer is correct, the orthodoxy of an idea is not relevant. For dubious reasons, we chose to euphemize what we actually meant to say here. With regards to circuit implementations of predictive processing (we cannot and do not intend to speak to interpretations of predictive processing that relate to conscious perception much of V1 activity is likely not consciously perceived – we assume this is what the reviewer is referring to by “we have one percept”) – the reviewers interpretation was not unorthodox, but rather incorrect (which is what we should have said). The statement that “the brain predictively ‘cancels’ expected action outcomes from perception” is incorrect in the context of sensory processing – based on both theoretical models of predictive processing, and more importantly physiological evidence. If the point was only in regards to conscious perception, we also suspect the statement is wrong, but even if it were correct, don’t see how it pertains to our work.

      Similarly a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      We believe it is absolutely essential for the progress of the field that we start to emphasize the differences between something that is “predictable in principle” and “predicted by the brain”. There is likely indeed a hierarchy of predictability that looks something like this:

      (1) Sensorimotor coupling

      (2) Laws of physics

      (3) Behavior of other living things

      (4) Artificial, human-made statistical relationships

      Almost all published experiments are based on the fourth type of prediction. Indeed, why not use physics simulations instead of oddballs and MMN? We absolutely should! But the field tends to revert to artificial couplings. As a direct consequence of this, the number of papers appearing recently (from both human and mouse fields), that are built on the following premise:

      (1) Expose an animal or human to an artificial coupling between A and B (e.g. an oddball, or a global oddball, or any of a myriad other constructions).

      (2) Probe for prediction error responses to the violation of the artificial coupling.

      (3) Find no prediction error responses and conclude predictive processing is wrong.

      Is utterly baffling. The fallacy here is of course the assumption that if something is predictable in principle, the brain must predict it. Thus, we are, and will continue to be strong on this point, and we think it is essential that we – as a field – are.

      Hope these comments are helpful.

      Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      We have added a discussion of this.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice, and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      The reference electrode placed at Cz makes it is difficult to interpret relative differences between frontal and occipital electrode responses, as the occipital electrodes are placed farther away from the Cz reference than the frontal electrodes. Similarly, signal occuring cortically near the Cz reference might only appear as though it is occipitally distributed in this montage. It is common in EEG research to remontage the data to an averaged common reference in order to better interpret the scalp distributions. As the electrode coverage was sparse for some subjects, this could be challenging, and this reviewer does not feel that it is necessary to do this analysis step, or even to drastically rewrite the body of the paper. We only request that some discussion, however brief, is included in the discussion section or the methods that recommend more dense electrode coverage in the future to better interpret scalp distributions and potential meso-scale sources.

      We have added a discussion of this as suggested.

      This is just a suggestion. The authors are encouraged to analyse (and report) time-frequency power and phase locking for these mismatch responses, as is common in much of the literature (see Roach et al 2008 Schizophrenia Bulletin). This is not to say that doing so will yield insights into oscillations per se, but converting the data to the time-frequency domain provides another perspective that has some advantages. fosters translations to rodent models, as ERP peaks do not map well between species, but e.g. delta-theta power does (see Lee et al 2018 Neuropsychopharmacology; Javitt et all 2018 Schizophrenia research; Gallimore et al 2023 Cereb Ctx). Further, ERP peaks can be influenced by the actual neuroanatomy of an individual (especially for quantifying V1 responses). Time frequency analyses may aid in interpreting the "early negative deflection with a peak latency of 48 ms " finding as well. As it stands, the report is complete, and it would be acceptable if the authors chose to save this type of analysis for a future publication.

      We have added this as suggested.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed most of my concerns by providing additional analyses, partly based on new data. The volume conduction issue is partly addressed based on the result showing latency differences, however to confidently assign responses to visual regions, one would need to perform recordings with a larger number of electrodes, sufficient to perform source localization. Nevertheless, the manuscript is now more solid than the previous version.

      We have now added this point to the Discussion.

      Reviewer #3 (Recommendations for the authors):

      The reviewer appreciates that the authors have carried out time-frequency analyses, and are ok with them leaving this out of this paper.

      We have now added this to the manuscript.

      Finally, in response to the participant quote "are you printing this? hi mom!" - this reviewer concedes that it does not significantly detract from the report, and, in the interest of amusement and joy, would abide its reinstatement.

      We greatly appreciate the reviewers entertaining our attempts at humor but will leave it out as originally suggested.

    1. eLife Assessment

      This study presents a key finding: self-generated mechanical stresses enable collective protocell proliferation without dedicated division machinery, offering insight into primitive life's population growth. While quantitative imaging, membrane tension measurements, and computational modeling support the mechanism, establishing causal links between deformation and division and testing sensitivity assumptions would strengthen the work. Overall, the work reports important findings, and although the evidence in support of the conclusions is largely solid, some incomplete elements need to be addressed.

    2. Reviewer #1 (Public review):

      Li and Wu, in this article, explore the proliferation of wall-less L-forms derived from Bacillus subtilis as mimics for protocells and report an interesting new mechanism for their proliferation. The authors carry out live-cell imaging of the L-forms and find that the clusters of cells forming proto-colonies proliferate better than the isolated single cells of L-forms. They further examine the causes for this indefinite proliferation of proto-colonies of L-forms, as compared to the isolated cells, which lyse and die out sooner. The authors show that when L-forms exist as isolated single cells, the growth in volume exceeds the rates at which surface area increases, leading to lysis. The authors further quantify the circularity and effective radius in growing proto-colonies, qualitatively estimate membrane tension and suggest that the confined space allows for mechanical shear in these cells. They propose that the mechanical stress on the membranes from adjacent cells in confined spaces deforms membranes and supports cell division to keep the population growing. These findings are also supported by modelling the proto-colonies in quasi-2D planes.

      The study is quite interesting and significant as it has implications for both evolutionary aspects as well as clinical importance, given the proliferation of certain pathogens as L-forms. The aspect of carrying out long-term imaging of colonies of L-forms as spatially constrained entities and the findings are fascinating. While the conclusions presented are backed by experiments, I only have a few questions concerning the proposed mechanism of division and proliferation of these proto-colonies.

      (1) The authors propose that the growth of neighbours leads to shearing forces in membranes and show that membrane tension increases at the periphery of the proto-colonies. They suggest that the increased membrane tension leads to a greater chance of deformation, enabling cell division. However, it is not quite clear how greater membrane tension could lead to cell division. Studies have suggested that membrane fluidisation is important for the cytokinesis event, which includes FtsZ-based division (Ramirez-Diaz, 2025).

      (2) Thus, it becomes quite important to rule out any role for the cytoskeletal proteins in the observed division with an increase in membrane tension. The authors note in line 188 that the division in protocells is independent of FtsZ, but this independence is for protocells that divide by extrusions and resolution, where the membrane is highly fluidised (Mercier et al., 2012).

      (3) The authors may use the L-form derivative where the FtsZ protein can be depleted and assess the proliferation of the proto-colonies. Likewise, authors should rule out the role of MreB as well.

      (4) Although the growth rates have been shown to be similar for proto-cells and the proto-colonies, and only the membrane tension has been shown to be higher at the periphery, it is also important that the authors rule out any increased lipid synthesis in the fraction of dividing cells in these proto-colonies. Without this, one could also envisage a model where membranes are fluidised due to an increase in lipid biosynthesis in a fraction of cells in these confined spaces, leading to increased vesiculations which experience membrane shear and deform. The authors can also consider examining proto-colonies of L-forms of branched-chain fatty acid-deficient strains.

      (5) Lastly, why does CellROX stain the proto-colonies? Are these tightly packed cells experiencing higher oxidative stress, and could that also contribute to membrane tension? This should at least be discussed.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Mechanical interaction enables a collective mode of protocell proliferation" addresses an interesting and potentially high-impact question about protocell proliferation in prebiotic environments. The central observation that wall-deficient B. Subtilis proliferate in dense colonies but die by membrane rupture in isolation is striking and a fundamental contribution to the field. However, the data and the mechanistic explanation offered for this observation are incomplete. The measurement and analyses used to build the mechanistic case raise methodological questions that may be difficult to fully resolve with the existing data and approach, and the authors should therefore consider whether additional independent experiments are needed to support the mechanical shearing hypothesis.

      Strengths:

      The central observation that wall-deficient B. Subtilis proliferate in dense colonies but die by membrane rupture in isolation is convincing and a significant contribution to the field interested in the growth of protocells. This adds an important aspect of collective growth that is different from individual dynamics.

      Weaknesses:

      (1) The surface-volume balance ratio η is an elegant concept and provides an intuitively reasonable framework for understanding why isolated cells lyse. However, its application here rests on treating cells as flat discs of uniform thickness, and Figure S4 makes clear that the cells are highly irregular and lobulated in ways that make this approximation questionable. The authors should clarify whether they have validated this assumption, for instance, through direct thickness measurements or sensitivity analysis. However, even with such validation, the modest quantitative differences between aggregated and isolated η trajectories, combined with the inherent difficulty of accurate perimeter measurement in these morphologically complex cells, mean that η measurements are unlikely to provide robust quantitative support for the mechanism. The authors should therefore consider whether η is better presented as a motivating conceptual framework rather than primary quantitative evidence and seek more direct experimental support for the surface-volume balance argument through independent means. For instance, osmotic pressure manipulation to test whether reducing volume expansion pressure preferentially rescues isolated cells.

      (2) The comparison of circularity between colony and isolated cells is complicated by the fact that the segmentation approach is fundamentally different in the two conditions; isolated cell boundaries are detected against a clear background, while colony boundaries are detected from inter-cell fluorescence gradients. The authors should address whether this introduces systematic bias. However, this may be difficult to fully resolve given the inherent complexity of the system, and that the deformation-division correlation in Figure 3C, while suggestive, would be substantially strengthened by a more direct perturbative approach. Specifically, can cell deformation be mechanically induced in isolated cells, for instance, using micromanipulation, external flow, or confinement in fabricated microstructures, to test whether artificially deformed isolated cells gain the ability to divide? Such an experiment would provide direct evidence for the deformation-division link that the correlational analysis cannot.

      (3) The interpretation of FliptR lifetime as a direct membrane tension readout is complicated in this system because cell-cell interfaces contain two apposed bilayers in proximity, potentially altering FliptR photophysics through changes in local membrane density and dielectric environment independently of tension. The authors should address whether they have considered this possibility and what controls were performed. Disambiguating tension-dependent from environment-dependent lifetime changes is technically challenging and suggests that the membrane tension argument would be more convincingly supported by an independent measurement approach. For instance, tether-pulling experiments using optical tweezers on isolated versus colony cell membranes, or testing whether membrane tension-modulating interventions such as osmotic shifts produce the predicted changes in cell fate, would provide more direct evidence. The current FLIM data should be regarded as suggestive rather than conclusive.

      (4) The Cellular Potts Model reproduces the experimental observations, but since its key parameters, particularly the substrate-pinning energy, were calibrated against those same observations, this demonstrates internal consistency rather than independent validation. The η-based lysis criterion is implemented as a model input, meaning the model cannot independently confirm the η hypothesis. The authors should clarify the extent to which model parameters were fitted to data versus independently motivated and be explicit that the model is best understood as a mechanistic illustration rather than independent evidence.

    4. Reviewer #3 (Public review):

      Summary

      This manuscript reports that protocells derived from wall-deficient B. subtilis proliferate well when densely packed but fail to divide and eventually lyse when isolated. The authors attribute this density-dependent proliferation to mechanical shearing between growing neighbors, which deforms cells and increases the likelihood of membrane stalk formation and subsequent scission, enabling division without any dedicated molecular machinery. Through a combination of quantitative imaging, membrane tension measurements, and Cellular Potts Model simulations, the authors make a compelling case that self-generated mechanical stresses are critical for sustaining population growth in protocolonies. The findings have implications for understanding the lifestyles of primitive life forms, L-form bacterial pathogenesis, and the design of synthetic cells.

      Strengths

      The central finding is both surprising and counterintuitive: crowding is not just tolerated by protocells but is required for sustained population growth. The mechanism the authors propose is interesting: mechanical shearing between growing neighbors deforms cells, increasing the likelihood of membrane stalk formation and thus division, all without dedicated molecular machinery. Conceptually, this is a type of biophysical "scaffold" (Jacobeen et al. 2018, Nat. Phys.; Day et al. 2022, Biophys. Rev.) in which key elements of a Darwinian loop, namely a life cycle involving growth and reproduction, are provided "for free" by physics, enabling open-ended Darwinian evolution that can eventually bring these life cycle components under developmental control. Such scaffolds, both biophysical and ecological (Black et al. 2020, Nat. Ecol. Evol.; Libby & Rainey 2013, Phys. Biol.), are likely key mechanisms in the origin of life and in evolutionary transitions in individuality, and this paper provides a nice example of how they can work in a protocell context.

      The combination of experiments and modeling works well. The membrane tension measurements are the strongest piece of evidence for the proposed mechanism, showing directly that tension is elevated in protocolonies and concentrated at cell-cell interfaces. The Cellular Potts Model captures the key experimental features. The discussion is nicely balanced, particularly the note about Gram-negative L-forms, whose rigid outer membrane may preclude this mechanism, which is a testable prediction for future work. I would suggest the authors also discuss the connection to biophysical scaffolding, as I think this is conceptually important and would help situate their work within a broader framework for understanding how primitive life cycles can arise from physical processes (see also Zamani-Dahaj et al. 2023, Genes; Hammerschmidt et al. 2014, Nature).

      Weaknesses

      The surface-volume balance analysis is central to the argument, and it depends on the assumption that cells have a fixed thickness of 0.8 µm, taken from the width of walled cells. But these are wall-deficient cells, which are mechanically quite different, and their thickness could plausibly vary during growth or under compression. I think the paper would benefit from either a direct measurement of cell thickness or a sensitivity analysis showing how η responds to plausible variation in this parameter. If the results are robust, that would put the analysis on much firmer ground.

      The positive correlation between cell shape deformation and division rate (Figure 3C) is central to the proposed mechanism, but I think the paper needs to be more careful about the jump from correlation to causation. The authors propose that deformation increases the likelihood of membrane stalk formation, leading to scission. That is plausible, but an alternative is that cells with higher local growth rates both deform more and divide more frequently, with the two outcomes driven independently by the same underlying cause. The paper does show that average volume growth rates are indistinguishable between aggregated and isolated cells, which argues against a simple "faster growth explains everything" interpretation, but this does not rule out local variation within protocolonies driving the correlation. I think the most convincing experiment would be to apply external mechanical stress to isolated cells and see if that alone can drive division, decoupling deformation from growth. I realize that this may be technically very difficult, but at a minimum, the paper should acknowledge this as an alternative hypothesis.

      The Cellular Potts Model has quite a few free parameters (Table S1), and it is not clear how tightly these are constrained by the data. A sensitivity analysis would go a long way toward showing that the results are robust and not overly dependent on specific parameter choices.

      In any case, this is a strong paper with a cool finding and an interesting mechanistic explanation. I think it will be of broad interest, particularly to people thinking about the origins of life and synthetic cell design.

    1. eLife Assessment

      This important study combines experimental evidence with computational modeling to suggest that neurons in the frontal eye field form two overlapping topographical maps, one for visual inputs and the other for eye movement directions. The experiments leverage the smooth cortex of the marmoset to provide solid evidence that two maps exist at different scales. The modelling work provides a potential mechanistic insight into how these patterns may be used for flexible visuomotor transformations.

    2. Reviewer #1 (Public review):

      Summary:

      Flexible natural behavior requires flexible sensory-motor mapping. In the visual domain, a visual stimulus at one location can guide a saccade toward another. How the receptive field (RF) and motor field (MF) properties of oculomotor structures support this flexibility is not known. Dotson and Reynolds address this question in the marmoset, using oblique Neuropixels penetrations across horizontal segments of the frontal eye field+, supplemented by electrical microstimulation. They report that visual RF and saccade MF vector angles each change smoothly with occasional abrupt jumps, that the two maps are organized as mosaics at distinct preferred spatial scales, and that a moiré interference pattern arising from a constrained spatial-scale mismatch between partially correlated mosaics reproduces the empirical distribution of RF-MF angular differences. They conclude that visuomotor flexibility is embedded in the geometry of mismatch and matches between visual and motor maps.

      Strengths:

      (1) The question is well-motivated. Sensory-motor mapping is known to be flexible, and asking whether the topographic relationship between the two maps itself supports that dissociation is a fresh reframing of a long-standing problem in oculomotor control.

      (2) FEF+ lies on the smooth marmoset cortical surface, which permits high-density horizontal sampling that would be difficult in the macaque arcuate sulcus, and oblique penetrations are a sensible way to track tuning across the surface. The dataset is substantial by the standards of the field (39 sites of high-density recordings across two animals, several thousand isolated units).

      (3) The data are thorough, and the convergence of three independent lines of evidence is the strongest feature of the paper. Unit recordings, electrical microstimulation, and two architecturally distinct generative models point to the same organization.

      (4) The central idea is conceptually novel. The proposal that flexibility can reside in the geometry of the maps, rather than only in time-varying activity, is original, and it generates concrete, testable predictions for tasks that require flexible visuomotor routing.

      Weaknesses:

      Major concerns

      (1) The analysis collapses each oblique penetration onto a single horizontal axis and pools angles across all cortical layers, treating cortical distance as purely tangential. Because the trajectory is angled, horizontal distance and depth are confounded, so some of the apparent RF-MF drift along a penetration could reflect a laminar transition, in addition to tangential mosaic structure.

      (2) RFs and MFs are estimated from the same free-viewing sessions in temporally adjacent epochs, leaving each measurement open to contamination by the other. Activity near a saccade can reflect peri-saccadic remapping rather than the stable retinotopic RF, and saccade-aligned activity following a recent flash can carry a residual visual component, given the long-lasting visual responses in FEF+ (>500 ms). Residual cross-contamination of this kind would tend to make RF and MF angles look more similar than they are, inflating the apparent local coupling and biasing the RF-MF difference distribution that the moiré model is fit to.

      (3) The paper claims that visual and saccade mosaics occupy distinct spatial scales, but the two preferred spatial frequencies are close, and the separation is summarized by overlapping "failed-test" bands rather than by a statistical test or confidence interval on the preferred frequency itself. The reliability of this separation is not established.

      (4) It is not clear whether the moiré model is a better model than the non-mosaic alternative. The moiré models are shown to be consistent with the data through failure to reject a Kolmogorov-Smirnov null, which is a weak form of evidence, and they are not benchmarked against a non-mosaic alternative or null model. The AM/NM convergence demonstrates architecture independence, but not that a mosaic organization is required.

      Minor concerns

      (1) The link from topography to behavioral flexibility (such as anti-saccades and other context-dependent transformations) is presented as a prediction but is not tested with any task manipulation. The work establishes an organizational principle and a plausible generative mechanism; whether that organization is actually exploited during flexible behavior remains open, and the framing should make this clear so the functional claim is not over-read.

      (2) It is unclear how relative depth (depth 0) is defined and how layer boundaries were assigned. The Methods mention common-average re-referencing for CSD and local field analyses, but no CSD or power-depth profile is shown to anchor the layer IV / depth 0 reference across penetrations.

      (3) The Discussion is brief relative to the strength of the claims. It would be helpful to address the concerns and alternative explanations above, where these cannot be fully resolved by the data.

    3. Reviewer #2 (Public review):

      Summary:

      The authors asked how the visuomotor system can keep visual selection and saccade targeting related but not rigidly coupled-the flexibility required for tasks like anti-saccades. They recorded visual receptive fields (RFs) and saccade motor fields (MFs) from individual neurons in marmoset frontal eye fields and adjacent premotor eye fields (FEF+) using obliquely inserted Neuropixels probes that traverse horizontal segments of the smooth marmoset cortex. They reported that visual and motor vector angles each changed smoothly with occasional abrupt jumps (a mosaic, rather than retinotopic, organization), that the two maps drifted with respect to one another, and that they were best described by mosaic-map models tuned to different preferred spatial frequencies. They then proposed that the offset in spatial scale, combined with partial shared structure between the maps, produced a moiré interference pattern, in which the distribution of local visual-motor angle differences matched the data.

      Strengths:

      (1) Unlike in the macaque brain, where FEF is buried in the arcuate sulcus, the marmoset cerebral cortex is lissencephalic (smooth). The authors used this feature to their great advantage and sampled horizontal mesoscale structure with oblique penetrations of ultra-high-density Neuropixels probes.

      (2) The mosaic framing was grounded in previous studies on direction maps in ferret V1 and MT, and the rate-of-change analysis (Supplementary Figure 2) plausibly reproduced the fracture-line phenomenon of those maps.

      (3) The authors confirmed the robustness of the finding by reaching the same conclusions using two architecturally distinct generators: the Fourier-based annulus model and the Gaussian-noise model.

      (4) They additionally provided an independent confirmation of the mosaic saccade-vector organization using electrical microstimulation. They reconciled the lower microstimulation spatial frequency by matching the spatial-averaging footprint (Supplementary Figure 6).

      (5) The Noise Mosaic (NM) model was used thoughtfully to decouple two properties that the Annulus Mosaic (AM) model confounds - spatial-scale offset and inter-map correlation - and to show that both an intermediate correlation (ρ ≈ 0.6) and a scale offset are required. Conceptually, a structural (topographic) substrate for visuomotor flexibility is a fresh alternative to the standard account in which flexibility lives entirely in time-varying activity on a single map.

      Weaknesses:

      The two claims here are not of the same strength. The first claim that RF and MF angles are organized as mosaics at distinct spatial scales was well supported. The second claim that a moiré interference pattern is the substrate for visuomotor flexibility was an inference rather than a direct observation, and several features of the design contribute to how strongly it can be held.

      First, the recordings were one-dimensional. Each oblique penetration yielded a line through the cortex, so the two-dimensional moiré pattern (Figure 4A) existed only in simulations; it was not reconstructed from the data. What the data provided was the marginal distribution of local angular differences, and the model was accepted when its simulated distribution was statistically indistinguishable from the empirical one. Matching a low-dimensional summary statistic is necessary but not strongly sufficient - multiple underlying architectures could produce similar 1-D difference distributions - so the moiré interpretation is best read as a plausible and parsimonious interpretation of the data rather than a confirmed mechanism.

      Second, the inferential logic needs to be strengthened. The "preferred" spatial frequencies were those at which a two-sample Kolmogorov-Smirnov test fails to reject equality between model and data. Failure to reject is not confirmation, and the width of the accepted band depends on statistical power, which depends on sample size. The authors did show that most parameter combinations were rejected, so the test did discriminate. That said, a continuous goodness-of-fit landscape with confidence intervals on the preferred SF, and a direct test that the visual and saccade preferred SFs differ, would better support the "distinct spatial scales" claim than visual inspection of two overlapping troughs.

      Third, the dataset was from two male marmosets, with 18 of 39 sites contributing to the core analyses, and the angular-difference distributions were pooled across penetrations and animals. This is standard for primate electrophysiology, but it means the spatial statistics were assumed stationary across the region, and individual variability in map layout was averaged over. The oblique-penetration geometry also added some uncertainty: the spatial-frequency estimates (in cycles/mm) are only as accurate as the reconstructed penetration angles (18.2{degree sign} {plus minus} 8.3{degree sign}), and angle error would propagate directly into the inferred scales.

      Fourth, while the authors suggested the moiré interference pattern can serve flexible routing for behaviours such as anti-saccades, this was not directly tested. Instead, they used free viewing and natural saccades, so the paper demonstrated a candidate substrate without testing whether behaviour employs it. This does not undercut the main findings, but readers should treat the functional narrative in the introduction and discussion as a set of predictions rather than results.

    1. eLife Assessment

      In this valuable study, Bhojappa et al. investigate the function of four septin-associated kinases (Elm1p, Gin4p, Hsl1p and Kcc4p) in controlling septin ring stability, actomyosin ring organization and constriction, and cytokinesis in budding yeast. The evidence supporting the conclusions is convincing. The results are a meaningful addition to previously published work and provide a clearer picture of a complex mechanism involving proteins of partially redundant functions.

    2. Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weakness:

      Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

    3. Reviewer #2 (Public review):

      Summary and strengths:

      In this study, Bhojappa et al. investigate the roles of the septin-associated kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and cytokinesis in budding yeast. Through quantitative analyses of kinase localization dynamics, septin organization, actomyosin ring (AMR) constriction, and cell morphology, the authors demonstrate that Elm1 and Gin4 play particularly important roles in maintaining proper septin architecture and cytokinetic progression. The work further identifies an interaction between the Gin4 KA1 domain and the Hof1 F-BAR domain and provides evidence that several cytokinesis-related functions of Gin4 are independent of its kinase activity. Artificial tethering approaches further supports that spatial organization at the bud neck is critical for the execution of septin-dependent cytokinetic processes. The authors combine live-cell imaging, quantitative analyses, biochemical interaction assays, and genetic perturbations to build a comprehensive framework for understanding how these kinases contribute to cytokinesis.

      Comments on revised version:

      The revised manuscript has been substantially improved. The authors have carefully addressed the concerns raised during review by providing additional experiments, analyses, quantifications, clarifications, and improved presentation of the data. The conclusions are now well supported by the experimental evidence. The study advances our understanding of the mechanisms linking septin organization to cytokinesis and will be of interest to researchers studying septins, cell division, and cytoskeletal regulation.

      Overall Assessment:

      This work provides valuable mechanistic insight into the coordination of septin organization and cytokinesis by septin-associated kinases. The experiments are carefully executed, the analyses are thorough, and the conclusions are supported by the data presented. The manuscript represents a useful contribution to the fields of cytokinesis and septin biology.

    4. Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p and Kcc4p. Elm1 and Gin4 show the strong knock-out phenotypes, whereas Hsl1p and Kcc4p show the weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Comments on revised version:

      I thank the authors for their thorough and thoughtful response to my review. The revised manuscript clearly reflects their efforts to provide rigorous and high-quality science. Addressing the concerns raised in the review required significant effort, but the improvements in the manuscript make it clear that the work was well worth it.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weaknesses:

      A control for non-specific binding is missing from the in vitro binding assay. The figures suffer sometimes from the very small text in the labels, which obscures understanding. Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

      We thank the reviewer for highlighting the strengths of our imaging pipelines and protein-protein interaction data. We have now included appropriate controls for the in vitro binding assays, which demonstrate that the observed interactions are specific (Fig. 2H). We have also revised all figures to improve clarity, including increasing font sizes to enhance visibility across panels. We agree that mapping the specific phosphorylation sites regulated by these septin kinases would provide valuable mechanistic insights. However, only a few studies have addressed this direction so far (Mortensen et al., 2002; Asano et al., 2006 and Marquardt et al., 2024) [1-3]. The current study focuses on the interplay among septin-associated kinases and their role in regulation of the actomyosin machinery (AMR). In this context, we highlight several key findings:

      (i) a molecular link between the septin kinase network and AMR through physical interaction between the KA1 domain of Gin4 and F-BAR domain of Hof1 (Fig. 2F-2H and S3H);

      (ii) a kinase-independent role for Gin4 in coordinating septin organization and AMR dynamics (Fig. 3A-3E, S3F & S3G, S3I and 4F-4H);

      (iii) a novel role for Hsl1 in regulating septins and the AMR downstream of Gin4 and Elm1, potentially through plasma-membrane binding (Fig. 4I-4K, 5A-5F, 8A-8G and S9H-S9J); and

      (iv) crosstalk between Gin4 and Hsl1 that is independent of their role in the morphogenetic checkpoint (Fig. 6A-6D and S6A-S6C).

      We have clarified this scope in the Discussion section and explicitly stated that mapping these phosphorylation sites will be an important direction for future work (Lines 623-626).

      Reviewer #2 (Public review):

      Summary:

      In this paper, Bhojappa et al. provide insights into the function of septin-related kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and actomyosin ring (AMR) structure and constriction. Their findings are both corroborative of and complementary to previous related studies.

      First, the authors provide a comparative analysis of the dynamic localization of these kinases at the bud neck, as well as a comparative analysis of defects in septin localization, splitting dynamics, AMR constriction rates, and cell morphology in kinase deficient cells. They find that septin localization and splitting kinetics, as well as AMR constriction rates, are significantly perturbed in elm1∆ and gin4∆ mutants but remain largely unaffected in hsl1∆ and kcc4∆. A similar trend is observed in terms of cell morphology and viability.

      Next, the authors focus on elm1∆ and gin4∆ cells, demonstrating that the residence time of the F-BAR protein Hof1 is significantly increased and defective in these mutants. Using yeast two-hybrid (Y2H) and in vitro binding assays, they show that the KA1 domain of Gin4 interacts with the F-BAR domain of Hof1, which may explain the cytokinesis-related functions of Elm1 and Gin4. Supporting this, they find that Gin4's role in septin localization, AMR constriction kinetics, and Hof1 bud neck localization is kinase-independent.

      The authors then conduct a series of artificial tethering experiments given their bud neck localization is mostly interdependent. They first demonstrate that artificially tethering Gin4 to the bud neck rescues the morphology defects of elm1∆ cells, with the strongest rescue observed when Gin4 was forced to interact with Hsl1-an effect that was also kinase-independent. Additionally, artificial tethering of Hsl1 to the bud neck restores the morphology of elm1∆ cells in a KA1 domain-dependent manner, suggesting that Hsl1 functions downstream of Elm1 to maintain normal cell morphology. Consistently, artificial tethering of Elm1 to the bud neck in gin4∆ cells rescues morphology defects, as well as defects in Myo1 localization and AMR constriction, but only in the presence of full-length Hsl1. The rescue fails in the absence of Hsl1 or when using a version of Hsl1 lacking the KA1 domain, which supports the role of Hsl1 downstream to Elm1 in cytokinesis.

      Strengths:

      Altogether, this study offers valuable insights into the mode of cytokinesis regulation mediated by the septin-related kinases, mainly Elm1, Gin4, and Hsl1, and would be an important contribution to the field of septins and cytokinesis after addressing current weaknesses.

      We thank the reviewer for the detailed summary and for highlighting the novel findings of our study.

      Weaknesses:

      (1) When assessing rescue of the elm1∆ phenotype, it needs to become clearer whether only morphology or also cytokinesis and septin organization are rescued.

      To clarify the extent of rescue observed in elm1Δ cells, we extended our analysis beyond morphological parameters by quantifying septin organization and AMR constriction dynamics. These analyses now show that artificial tethering partially restores septin organization and AMR constriction kinetics in elm1Δ cells in addition to improving cell morphology. These results are now described in detail in Fig. 5 and lines 373-406.

      (2) The quantification of the microscopy data does not always match up with the example images, and it's not always clear how the authors quantitatively analyzed their data.

      We revised the manuscript to clearly outline the quantification methods used for microscopy data analysis and specified the statistical tests, number of cells analyzed, and number of experimental replicates in the figure legends. We also clarified the criteria used for phenotype scoring and quantification in the Materials and Methods section. In addition, We replaced representative images where necessary to accurately reflect the quantified data throughout the revised manuscript.

      (3) The forced tethering data are key to the paper, but the lack of a summarizing table makes it difficult to grasp the full picture.

      We agree with the reviewer and have now included a new summary table (Table 1) that compiles the results of all artificial tethering experiments presented in this study, including the percentage of rescue in cellular morphology observed upon forced tethering of these kinases to the bud neck, thereby providing a clearer overview of these experiments.

      (4) Novel results and those confirming earlier results could be better distinguished.

      We have improved the overall clarity of the manuscript to distinguish novel findings from the results that corroborate previous studies, and have cited the appropriate literature throughout the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p, and Kcc4p. Elm1 and Gin4 show strong knock-out phenotypes, whereas Hsl1p and Kcc4p show weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Strengths:

      The combination of genetics, cell imaging, and biochemical characterization of proteinprotein interactions is attractive.

      We thank the reviewer for recognizing the significance of our findings and for the helpful suggestions.

      Weaknesses:

      (1) Imaging and data analysis is the main weakness of this manuscript. The authors must avoid manual counting and selection when easy analysis software can be used to limit bias. Instead of presenting unclear statistics of "percentage phenotypes", they need to define clear metrics to offer meaningful phenotype analysis.

      We agree that improving the quantitative rigour of the image analysis is essential for this study. Accordingly, we implemented a semi-automated image analysis workflow in the revised manuscript that defines reproducible metrics, such as aspect ratio, and reduces reliance on subjective phenotypic scoring. The inclusion of these parametric measurements enables clearer and more objective comparison of the rescued phenotypes.

      (2) This manuscript examines a very complex mechanism with four kinases of overlapping function using new data and existing literature. A clearer picture/model at the end of the manuscript that synthesizes the current knowledge would be beneficial:

      We incorporated a new representative model (Fig. 9) that integrates current knowledge in the field with our findings. This model highlights crosstalk among Elm1, Gin4, and Hsl1 as a key mechanism coordinating septin architectural transitions with AMR constriction during cytokinesis and is discussed in lines 520-536 of the revised manuscript.

      We sincerely thank all the reviewers for their insightful comments. We incorporated new results in Fig. S3A, S3B, S4A-S4F, 5A-5F, 6A-6D, S6A-S6C, 8E, 8F, S10B and 9, along with Table 1 summarizing the artificial tethering experiments in the revised manuscript. We believe that these revisions have improved the rigor of our analyses and enhanced the overall clarity of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 70: " abnormal cytokinesis defects": this language is redundant, either "abnormal" or "defects" would suffice:

      We thank the reviewer for identifying the redundant wording. We have rephrased the final paragraph of the Introduction section.

      Please refer to line numbers 86-97.

      (2) Currently, the final paragraph of the Introduction is an extensive, detailed summary of the results. This is unnecessary, as the Abstract and the Results sections summarize the results. Better would be a short statement of the questions addressed in the manuscript:

      We thank the reviewer for this suggestion. We have revised the final paragraph of the Introduction to remove the detailed summary of results and instead outline the key questions addressed in this study. The revised paragraph now emphasizes the knowledge gap regarding how septin-associated kinases regulate septin organization and coordinate cytokinesis. Please refer to line numbers 86-97.

      (3) In the Results, the wording of the heading associated with the first section is confusing. "and their defects during the cell cycle": it is not expected that the kinases themselves will have defects; defects may be observed in cells upon mutation of the kinases, but that is not clear from this wording:

      We thank the reviewer for this comment. The heading has been revised from “and their defects during the cell cycle” to “and defects associated with their deletions” for improved clarity.

      Please refer to line numbers 99-100.

      (4) Lines 107-109 (" they may act as molecular signals for the transition and a trigger for crosstalk of cytokinesis") are redundant with earlier lines 104-105 (" suggesting them to be a possible trigger for septin remodelling"):

      We have rephrased the text to improve clarity and avoid redundancy.

      Please refer to line number 123-125.

      (5) Lines 121-122: "Septin-associated kinases are believed to play an essential regulatory role": is the role essential, or is it regulatory? Since these kinases are not individually essential for cytokinesis, "essential" doesn't seem appropriate here:

      This has been corrected in the revised manuscript.

      (6) Figure 1 panels C and F: the font size is exceedingly small for the labels and should be greatly increased. This is true for multiple panels in Figures 2-5 and the Supplemental Figures as well:

      We thank the reviewer for highlighting this issue. We have significantly increased the font sizes of the text and the x- and y-axis labels across all figures in the manuscript.

      (7) The timing of mitotic spindle breakdown was used as a timepoint for comparison but it is not made clear in the manuscript how this was determined. Presumably, the mRuby2-Tub1 marker was visualized and the timepoint when the mitotic spindle separated into two discrete entities was called the "breakpoint" timepoint, but it would be important to better describe (and ideally show an example) of how that timepoint was determined. The only images I find with labelled tubulin are Tub1-GFP and these do not show spindle breakdown:

      We thank the reviewer for raising this point. We used GFP-Tub1 (pAFS125-GFPTUB1) and mRuby2-Tub1 (pHIS3p:mRuby2-Tub1+3′UTR::URA3) plasmids to visualize spindle dynamics across the cell cycle, and in both cases defined the spindle breakpoint as time zero. As suggested, we have now included time-lapse images of Cdc3-mCherry and GFP-Tub1 in wild-type cells in Fig. S2D to illustrate the spindle breakpoint event used for temporal alignment. Corresponding changes have also been made in the Materials and Methods section to explicitly describe this analysis.

      Please refer to line numbers 760-761.

      (8) Line 165: "F-BAR protein Hof1, which senses and induces membrane curvature" and lines 170-172 "the F-BAR protein Hof1, which is known to be associated with septin hourglass and transit to AMR during split ring trigger". It is awkward to introduce the same protein twice, in two different ways, within a few lines of each:

      We thank the reviewer for noting the redundant description. We have rephrased the text to introduce the F-BAR protein Hof1 in a more concise and streamlined manner while retaining the relevant functional information.

      Please refer to line numbers 206-208.

      (9) Throughout, it would be helpful to introduce more line breaks and organize the Results into smaller paragraphs:

      We thank the reviewer for this suggestion. We have revised the layout of the Results section and introduced additional line breaks to improve readability.

      (10) This is somewhat of a personal preference, but in the interest of transparency (and with the understanding that the 0.05 value is entirely arbitrary), would authors be willing to show actual P values in the figure panels rather than "**" or "ns", for example? Some readers may wish to apply a different standard of significance than 0.05, and not showing the P values makes this impossible. Furthermore, some readers (like this one) may interpret differently a P value of 0.044 versus "*" or 0.051 vs "ns":

      We agree with the reviewer’s suggestion. To improve the transparency of the quantitative analyses, we have modified the graphs across all main and supplementary figures to display the “actual p-values” and specified the corresponding statistical tests in the figure legends, with significance indicated by asterisks. An example is provided in the attached image showing the residence time of Inn1-mNG in wild-type and gin4Δ cells complemented with kinase-active and kinase-dead Gin4 constructs (Fig. 3E in the revised manuscript). For this analysis, significance was assessed using the nonparametric Kruskal-Wallis statistical test.

      (11) It is written that the authors "performed a Yeast Two Hybrid screen to find novel interacting partners at the bud neck" but I do not find evidence anywhere of a "screen" being performed, i.e., an unbiased search of many proteins to find a few interactors. Instead, it appears that the authors performed a two-hybrid assay to visualize interactions between a small, specific set of proteins. "Screen" should be replaced with "assay", as is currently the case in the Methods section. It is probably also valuable to point out in the text that using a yeast two-hybrid assay to assess interactions between yeast proteins has the caveat that any interactions observed could be indirect, as they may be "bridged" by endogenous yeast proteins.

      We thank the reviewer for raising this point. Our Yeast Two-Hybrid experiments were performed using a small subset of septin-associated proteins rather than as an unbiased screen. Accordingly, we have replaced the term “screen” with “assay” throughout the revised manuscript and in the Materials and Methods section.

      Please refer to line 241.

      We also agree with the limitations inherent to this assay and have now explicitly stated it in the revised manuscript, please see lines 248-255.

      (12) Panel 2H: I do not understand what the middle (as opposed to the top and bottom) blot segment represents. It is labelled "anti-HIS", like the one above it, but it is not associated with any molecular weight/ladder marker and I do not know what other species in the binding reaction in that lane would be recognized by the anti-HIS antibody. Perhaps the top segment is an "input" sample, and below it (in the middle segment) is what was bound to the beads. The figure legend is uninformative in this regard. Also, panel I in this figure is unnecessary to show, assuming that the bands shown in H are what I think they are. The blot makes the point without the need for quantification.

      As suggested by the reviewer, we have added molecular weight markers for each blot panel showing the input and bead-bound fractions. The figure legend has also been updated accordingly to clearly describe the different blot segments and experimental conditions.

      Please refer to figure legend 2H.

      In addition, as suggested by the reviewer, we have removed the quantification graph corresponding to the in-vitro binding assay from the revised manuscript.

      (13) The results in Figure 2H demonstrate that the Gin4-KA1 fragment is not non-specifically "sticky", because it does not bind GST alone, but there is no demonstration that binding by the Hof1 fragment is specific because there is no equivalent negative control for binding:

      We thank the reviewer for this suggestion. We repeated the in-vitro binding assay using 6His-bdSUMO as a negative control alongside 6His-bdSUMO-Gin4<sup>KA1</sup> to demonstrate binding specificity of the Hof1 fragment. The Hof1 N-terminal F-BAR fragment did not pull down the control 6His-bdSUMO fragment but specifically pulled down 6His-bdSUMO-Gin4<sup>KA1</sup> under identical experimental conditions (Fig. 2H), confirming the specificity of the interaction between Hof1 F-BAR domain and Gin4KA1.

      The corresponding text and results have been updated in the revised manuscript.

      Please refer to Fig. 2H and lines 259-263.

      (14) Line 257-258: "Localisation via Hsl1 is necessary to rescue the morphological defects exhibited by Δelm1 cells partially": what does "partially" refer to here? To the rescue, or the defects?

      The term “partially” refers to the extent of rescue. Specifically, elongated cell morphology was rescued in 63.75% of the elm1Δ cell population, rather than in all cells, upon artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP.

      (15) Lines 292-293: "can restore the morphological defects": this wording is unclear. "Restore" means "return to a former condition", which in this case would be normal cellular morphology, not defective cellular morphology. Similarly, see line 304: "While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering": presumably the proper localization was restored, not the mislocalization:

      In Lines 292-293, by phrase “can restore the morphological defects” was intended to indicate rescue of the elongated/clumped morphology associated with gin4Δ cells upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP. We have now rephrased this sentence as: “can rescue the elongated/clumped phenotype exhibited by gin4Δ cells”.

      Please refer to line numbers 481-482.

      Similarly, in Line 304, the statement “While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering” referred to rescue of the Myo1 mislocalization phenotype observed in gin4Δ cells. We have rephrased this sentence as: “However, Myo1-3xmCherry localization was restored to the bud neck upon artificial tethering of Elm1-GFP in gin4Δ cells”.

      Please refer to lines 506-508 in the revised manuscript.

      (16) Lines 295-296: "We find that Elm1-GFP tethering via Shs1-GBP, Bud4-GBP, and Hsl1-GBP" this should be "or", not "and":

      Thank you for pointing this out. We have now corrected the text accordingly.

      Please refer to line number 485.

      (17) Lines 311 and 312 refer to "Inn1-3xmCherry lifetime" but previously Inn1 residence time was measured. Since fluorescence lifetime is a distinct kind of measurement/ assay, it seems important to clarify here what kind of experimental data are being referred to:

      The term “Inn1-3xmCherry lifetime” was intended to describe the residence time of Inn1 at the cell division site, defined as the interval between the initial appearance of the Inn1 fluorescence signal and its complete disappearance during cytokinesis. For clarity and consistency, we have replaced the term “lifetime” with “residence time” in the revised manuscript.

      Please refer to the line numbers 511 and 512.

      (18) Discussion: "Septins are considered as the fourth cytoskeletal elements due to their extensive structural and functional diversity." This sentence is confusing, as it seems to imply that what defines a protein as being "cytoskeletal" is structural and functional diversity rather than anything to do with forming filaments, etc. The rest of this first paragraph of the Discussion also sounds like a summary of the background, is quite redundant with the Introduction, and should be shortened:

      We thank the reviewer for this suggestion. We have rewritten the first paragraph of the Discussion to improve clarity, reduce redundancy with the Introduction, and better emphasize the main findings of the study.

      Please refer to the line numbers 538-552 in the Discussion section.

      (19) Lines 365-366: "Gin4 and Hof1 are synthetic lethal": this should be revised to "gin4∆ and hof1∆ are synthetic lethal". This is a good place to point out that standard yeast nomenclature inserts the ∆ symbol after the gene name, not before (as is done in E. coli genetics, for example):

      We thank the reviewer for this suggestion. We have replaced “Gin4 and Hof1 are synthetic lethal” with “gin4Δ and hof1Δ are synthetic lethal” in the revised manuscript.

      Please refer to line number 576.

      In addition, we have now consistently placed the Δ symbol after the gene throughout the manuscript in accordance with standard yeast nomenclature.

      (20) Line 399: "We also performed an extensive GFP-GBP screens": again, here "screen" implies that a large collection of genes/proteins were assayed, perhaps in an unbiased way, which does not accurately portray what was actually done, which was an extensive tethering study using GFP-GBP:

      We thank the reviewer for this suggestion. We have replaced the term “GFP-GBP tethering screen” with “GFP-GBP tethering assay” throughout the revised manuscript. In addition, we have included a brief description of the specificity and functionality of GBP nanobody and its application in the GFP-GBP tethering strategy, extensively used in this study.

      Please refer to line numbers 326-336.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Analysis of morphological defects of elm1∆ does not directly reflect the defects in cytokinesis and septin organization. For example, the deletion of SWE1 rescues the morphological defects of elm1∆ cells but not the cytokinesis or septin mislocalization (Bouquin et al., 2000). Considering this, the authors should address whether artificial tethering of Hsl1 to Gin4 in elm1∆ cells rescues the septin and cytokinesis defects or just the morphology. Is the role of Hsl1 in cytokinesis dependent on its role in the morphogenesis checkpoint? Can the authors comment on how much the defects observed by Hsl1 tethering to the bud neck may be a result of bypassing the morphogenesis checkpoint?:

      We thank the reviewer for this important point. We performed time-lapse imaging of Cdc3-mCherry in strains where Gin4-GFP partially rescued the elongated phenotype of elm1Δ cells (63.75%) when tethered to the bud neck via Hsl1-GBP. Under these conditions, 64.29% of cells showed rescue of Cdc3-mCherry mislocalization, and Gin4 localization itself was restored to the bud neck in 57.85% of cells. We also examined Myo1-ymScarletI dynamics, while 76.92% of untethered elm1Δ cells displayed Myo1 mislocalization to the bud cortex, this was reduced to 11.36% upon Gin4-GFP tethering via Hsl1-GBP. Together, these results indicate that the morphological rescue observed in elm1Δ cells is accompanied by restoration of normal septin organization and AMR dynamics.

      Previous work (Bouquin et al., 2000) [4] showed that Swe1 deletion rescues cell elongation in elm1Δ cells without restoring septin organization. Consistent with this, we found that 60.66% of elm1Δ swe1Δ cells exhibited a round morphology, but tethering of Gin4-GFP to the bud neck via Hsl1-GBP in elm1Δ swe1Δ background did not further enhance morphological rescue. These results suggest that the rescue of cellular morphology observed in our tethering experiments may, atleast in part, depend on Hsl1-mediated regulation of the morphogenesis checkpoint.

      Importantly, despite the lack of additional morphological rescue, a clear restoration of septin localization was observed when Gin4-GFP was tethered to the bud neck via Hsl1-GBP in elm1Δ swe1Δ cells. Overall, these results suggest that while Hsl1-dependent morphogenesis checkpoint regulation may contribute to cell shape rescue, the restoration of septin organization is independent of Hsl1’s function in morphogenesis checkpoint and instead reflects a direct requirement for Gin4 and Hsl1 at the bud neck.

      Please refer to Figures 5, 6, and S6 of the revised manuscript for these additional data.

      (2) As suggested by the authors, the interaction of the Gin4-KA1 domain with the FBAR domain of Hof1 may explain the cytokinesis-related functions of Gin4. As an orthogonal approach, how does KA1 domain deletion of Gin4 affect cytokinesis and Hof1 bud neck localization?

      We thank the reviewer for this suggestion. We first examined the bud neck localization of Gin4-ka1Δ-GFP in comparison with full-length Gin4-GFP. We observed that the localization kinetics of Gin4-ka1Δ-GFP were significantly altered relative to the full-length protein, with reduced recruitment and earlier removal from the bud neck. We also analyzed the localization kinetics of Hof1-mNG in both gin4-ka1Δ and gin4Δ cells. Our results show that Hof1-mNG displays increased residence time and altered accumulation kinetics during cytokinesis in both genetic backgrounds. Thus, loss of the KA1 domain phenocopies loss of the full-length Gin4 and is consistent with disruption of the physical interaction between Gin4 and Hof1.

      Please refer to Figure S4 for these results in the revised manuscript.

      (3) The authors state that "Elm1 and Kcc4 were present at lower abundance at the bud neck (Fig S1A-D) compared to the higher abundance of Gin4 and Hsl1, as observed in their fluorescence intensities (Figures S1B-C) ". This is not evident in the figures. The authors should show a quantification of how they judged abundance at the bud neck:

      We thank the reviewer for this question. Quantification of septin kinase fluorescence intensity at the bud neck was performed using the established protocol for measuring protein accumulation kinetics described by Okada et. al. 2020 [5]. Time-lapse imaging for kinetic analysis of septin-associated kinases during bud emergence shown in Fig. S1A-S1D was carried out using a point-scanning confocal microscope with a 100×oilimmersion objective. Different laser intensities were required because the fluorescence signals of Kcc4 and Elm1 were comparatively weak and not readily detectable above cellular background under the imaging conditions used for Gin4 and Hsl1. The images shown in Fig. S1A-S1D are therefore displayed using differential contrast settings to facilitate visualization. We have now explicitly clarified this in the figure legend.

      A more direct comparison of septin kinase abundance at the bud neck is now provided in Fig. S1E-F, where localization kinetics during the HDR transition/septin remodelling stage were captured using a laser-scanning spinning-disk microscope under similar imaging conditions. We have additionally included raw fluorescence intensity profiles during the HDR transition to better illustrate the relative abundance of these kinases at the bud neck during cytokinesis.

      Please refer to Figures S1E and S1F in the revised manuscript for the updated images and quantitative analyses.

      (4) How did the authors determine G1 and M-phase in the experiments shown in Figures S1A-D? Can the authors mark these phases on the timelapse images? Also, How do the authors explain the different behaviour of Cdc3 in graphs S1A-D among different strains?

      We thank the reviewer for this comment. Cell cycle stages were initially inferred based on bud size, where kinase accumulation at the bud neck correspond to bud emergence (small bud, G1), and kinase disappearance coincided with septin splitting (large bud, M phase). However, we agree that accurate assignment of cell cycle stages would require specific cell cycle markers. To avoid confusion, we have removed the cellcycle-specific stage assignments from the Results section and describe the kinetics relative to t=0 (bud emergence).

      The differential dynamics observed in the Cdc3-mCherry kinetic profiles likely reflect heterogeneity within the the cellular population. To address this, we combined the normalized fluorescence intensity profiles of Cdc3-mCherry from the strains expressing GFP-tagged septin kinases during bud emergence and have included this data as reference (Author response image 1).

      Author response image 1.

      Plot showing spatiotemporal kinetics of Cdc3-mCherry in strains expressing either Elm1-GFP, or Gin4-GFP, or Hsl1-GFP, or Kcc4-GFP.

      (5) Forced tethering based experiments are one of the key sets of experiments for this work, but it is difficult to have a comprehensive understanding of all the data considering how large the data set is and how dispersed it is in the supplemental and main figures (Figures 3-4-5 and Figures S4-S5-S6). It would be helpful to provide a table summarizing the tested forced tethering’s and the phenotypic outcome in the tested yeast strains (Wt/mutant):

      We thank the reviewer for recognising the extensive dataset generated from the artificial tethering experiments and for suggesting the inclusion of a summary table. We have now added Table 1, which summarizes the proteins used in the GFP-GBP artificial tethering experiments, their genetic backgrounds, the total number of cells quantified across three independent replicates, and the phenotypic outcomes associated with bud neck tethering under each condition.

      Please refer to Table 1 and lines 342, 361, 451, 453, 456 and 492 in the revised manuscript.

      (6) In the introduction section, it would help the reader to provide more information on already known molecular roles of septin-associated kinases in septin organization and AMR. Later in the results section (i.e. Figure S1, S2, and S6), the authors extensively explain and show data that independently corroborate some earlier findings, which makes it difficult for the reader to distinguish novel findings from the repeated findings. I suggest shortening the text for the corroborative results, which will help to put more emphasis on their novel findings:

      We thank the reviewer for this suggestion. We have now included the canonical roles of these four septin-associated kinases in the Introduction section of the revised manuscript.

      Please refer to line numbers 79-85.

      We have also revised sections describing corroborative findings and explicitly cited previous studies wherever relevant in the Results section to better distinguish previously established observations from the novel findings presented in this work.

      Other minor comments:

      (1) In Figure 1D-1F, also show the data for hls1∆ and kcc4∆ - which are shown in S3AC in the current version:

      We have now included the Inn1-mNG residence time in hsl1Δ and kcc4Δ cells, alongside elm1Δ and gin4Δ in Fig. 1E.

      Please refer to Fig. 1E in the revised manuscript.

      (2) The authors should be more careful in interpreting their negative Y2H data in Figure 2F.

      We have now explicitly discussed the caveats and inherent limitations of the Yeast Two-Hybrid assay in the manuscript.

      Please refer to line numbers 248-255.

      (3) Please provide quantification for Figure 5A, Figure S6I:

      We have added quantitative analyses showing rescue of Cdc3-mCherry and Myo13xmCherry mislocalization in gin4Δ and gin4Δ hsl1Δ strains upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP.

      Please refer to Figures 8E, 8F, and S10B in the revised manuscript.

      (4) In Figure S6: label is missing "∆" in front of hsl1:

      We thank the reviewer for pointing out this error. We have corrected the labels accordingly.

      (5) As common consensus on yeast gene nomenclature, I suggest the use of "gene∆" instead of "∆gene":

      We thank the reviewer for this suggestion. We have now consistently placed the Δ symbol after deleted gene names throughout the manuscript in accordance with standard yeast nomenclature.

      (6) Lines (535-536): min(distribution) and max(distribution) in the formula is confusing. Clarify it or if possible use "minimum value", "maximum value" instead:

      We thank the reviewer for this suggestion. We have replaced the term “distribution” with “value” in the formula for protein accumulation kinetics analysis.

      Please refer to the updated formula in the Materials and Methods section (Lines 752753).

      (7) In line 86, "Dynamics of Septin-associated kinases and their defects during the cell cycle": Change the title as it is not clear what is meant by "their defects" given the discussed results under this title:

      We thank the reviewer for this suggestion. We have revised the section heading from “and their defects during the cell cycle” to “and defects associated with their deletions”.

      Please refer to line numbers 99-100.

      (8) On the Hof1-mNG image (Fig2C), show the line used for the line scan profile. Additionally, a similar line-scan profile could be useful in Figure S3G:

      As suggested by Reviewer 3, we removed the line-scan analysis from Figure 2 in the revised manuscript because Hof1 ring organization showed substantial heterogeneity across cells, making line-scan analysis difficult to interpret reliably.

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The % phenotype units are terrible. With no explanation, we do not really know whether they represent the percentage of cells that have a particular phenotype, or whether they correspond to a metric that measures some deviation between normal and extreme phenotypes. I would strongly recommend using precise quantitative metrics systematically (i.e. intensities, aspect ratios, division times, etc.) to properly quantify phenotypes:

      We thank the reviewer for suggesting the inclusion of precise quantitative metrics to assess phenotypic differences in the GFP-GBP tethering experiments. In the revised manuscript, we adopted quantification workflows that have been extensively validated and widely used in the literature, including those reported by Marquardt et al., 2024 (Fig. 7B and 7D) [2] from the Bi Lab. In response to the reviewer’s suggestion, we have now incorporated additional quantitative measurements, including cell area and aspect ratio (defined as the ratio of the cell’s major axis to the minor axis), for the experimental datasets presented in the manuscript.

      In the main figures, we now include aspect ratio quantification, while additional parameters are provided for the reviewer’s reference. We also quantified the fluorescence intensity of tethered proteins at the large bud neck and present these data together with the aspect ratio analysis in Figure 4 for elm1Δ cells in which Gin4GFP is artificially tethered to the bud neck via Hsl1-GBP. These quantitative analysis corroborates our qualitative observations and further strengthens our conclusions. Please refer to Figures 4D and 4E as representative examples.

      We have also changed the y-axis labels throughout the revised manuscript. For example, the y-axis in Fig. 7B is now labelled as “Cells exhibiting round morphology (%)”. Please refer to Fig. 4H, 4K, 7C, 8D, S5G, S6C, S7C, S9G and S9J for inclusion of aspect ratio quantification. Statistical analyses for the represented graphs were performed using Kruskal-Wallis nonparametric test, (N=3, n>150 cells/strain) (*: p<0.05, **: p<0.01, ****: p<0.0001, ns: p>0.05).

      Author response image 2.

      (2) Some quantitative analyses were performed manually where simple automated analysis should be performed to provide unbiased, accurate quantification:

      We fully agree with the reviewer that automated image analysis approaches, such as segmentation-based methods, are generally preferred for minimizing bias in morphological quantification. However, elm1Δ and gin4Δ cells exhibit severe phenotypes, including pronounced elongation and clumping, which makes reliable automated segmentation technically challenging for accurate quantification of parameters such as aspect ratio and cell area. For this reason, we used manual annotation for these analyses, as this approach enabled accurate delineation of individual cell and reliable measurements of morphological parameters such as cell area and size across the datasets despite being more time-consuming.

      Please refer lines 769-775 in the Materials and Methods section.

      (3) The "tethering" data also lack clear quantification. The authors should properly quantify the average intensity of Hsl1-GFP at the bud neck in each condition and correlate the results with cell aspect ratios or any other relevant parameters. For example, when comparing elm1null and elm1null Bud4-GBP with elm1null Kcc4-GBP, the visual impression is that as much Hsl1-GFP protein is recruited to the bud neck, whereas the phenotypes are dramatically different:

      We thank the reviewer for pointing this out. We have revised the image representation to facilitate clearer interpretation of the tethering experiments. In addition, we performed the key GFP-GBP tethering experiments using GBP-ymScarletI constructs, allowing direct visualisation of both the GFP-tagged protein and the GBP-tagged partner at the bud neck following tethering. We have included quantitative analyses of cellular morphology, raw fluorescence intensities of GFP-tagged proteins at the large bud neck, and the corresponding aspect ratio measurements for these updated datasets (Author response images 3, 4, 5). We have also included a summary table (Author response table 1) compiling these quantitative results for easier comparision.

      (4) I am very confused by Figure 3 which shows normal localization of Gin4-GFP in elm1null cells and seems to contradict other claims in the manuscript. This is very problematic for the interpretation of most of the "tethering" data:

      We understand the reviewer’s concern and have replaced the representative images of Gin4-GFP in elm1Δ cells in Figure 4B. Although Gin4-GFP is initially recruited to the presumptive bud neck during bud emergence in elm1Δ cells, it subsequently becomes mislocalized to the bud cortex during early cell cycle stages, resulting in reduced bud neck localization. As the cell cycle progresses, the Gin4-GFP signal at the bud neck decreases substantially in elm1Δ cells while remaining stable in wild-type cells until its departure prior to septin HDR remodelling (Fig. S5A-S5D).

      (5) Figure 6 is neither explained in the text nor in its legend. Could the authors explain the model and offer a comprehensive picture of the current knowledge?

      We have simplified the representative model to more clearly distinguish previously established knowledge from the findings presented in this study. Based on our results, We propose that Hsl1 functions both downstream of and in coordination with Elm1 and Gin4 to regulate septin stability and the timely execution of cytokinesis. Deletion of Elm1 disrupts the normal localization and crosstalk between Gin4 and Hsl1 at the bud neck, leading to septin mislocalization and misregulation of AMR dynamics, thereby revealing a previously uncharacterized role for Hsl1 in cytokinesis. The Results section has also been updated to reflect the revised model.

      Please refer to Figure 9 and lines 520-536.

      (6) Could the authors provide information about the double/triple mutant kinase phenotypes to clarify the overlap of functions among them?:

      Barral et al., 1999 [6] reported that individual deletions of Hsl1 and Gin4 result in mild cytokinetic defects, whereas deletion of Kcc4 does not produce any striking phenotype compared to wild-type cells. In contrast, the hsl1Δ gin4Δ kcc4Δ triple mutant remains viable but exhibits severe morphological abnormalities, including branched chains of elongated cells with defective cell separation. These mutants also display aberrant septin organization at the bud neck, characterized by irregular patch-like structures. Analysis of double mutants (hsl1Δ gin4Δ, gin4Δ kcc4Δ, and hsl1Δ kcc4Δ) revealed intermediate phenotypes between the corresponding single and triple mutants, with the hsl1Δ gin4Δ combination showing the strongest defects. Together, these findings suggest that the Nim1-related kinases function redundantly to regulate Swe1 activity and maintain septin architecture at the bud neck.

      Further supporting this model, Bouquin et al., 2000 [4] demonstrated that Elm1 operates independently of the Nim1-related kinases in controlling septin organization. The hsl1Δ gin4Δ kcc4Δ elm1Δ quadruple mutant exhibits severe septin localization defects and strong growth defects, in contrast to the elm1Δ single mutant, which primarily displays septin mislocalization from the bud neck to the bud cortex. These findings indicate that the combined activity of these kinases is essential for proper septin anchorage at the division plane and for assembly of the septin ring.

      Importantly, the progressively stronger phenotypes observed in double, triple and quadruple mutants also suggest that these kinases retain partially specialized functions at the bud neck. Consistent with this framework, our results support a model in which Nim1-related kinases function redundantly to regulate septin architecture and cytokinesis, likely through modulation of the AMR machinery. Because the localization of these kinases appears interdependent, as reported previously (Marquardt et al., 2020; Marquardt et al., 2024) [2,7] and corroborated by our findings, interpretation of mutant phenotypes remains complex and future studies will be required to delineate their individual contributions more precisely.

      Minor points:

      (1) Abbreviations are not defined in the manuscript.

      We have now expanded and defined all abbreviations throughout the manuscript.

      (2) Some of the writing in the figures is too small. Please make sure that a minimal size of letters/numbers is respected:

      We thank the reviewer for raising this issue. We have enlarged the figure labels and axis labels throughout the revised manuscript to improve readability.

      (3) Figure 2D. I am not sure that the line scans bring any useful information as the rings in mutant cells are quite heterogenous. This panel is also not cited in the text. Please make sure that every panel is cited at least once:

      We agree with the reviewer regarding the heterogeneity observed in the Hof1 ring organization in mutant cells and have therefore removed the line-scan analysis from the revised manuscript.

      (4) Knocked-out genes are written incorrectly. Please use the usual yeast nomenclature:

      We thank the reviewer for this suggestion. We have corrected the nomenclature for all deleted genes throughout the manuscript in accordance with standard yeast nomenclature.

      Author response image 3.

      Artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Gin4-GFP via Shs1-GBP-ymScarletI, Hsl1-GBP-ymScarletI, Bud4-GBPymScarletI and Bni5-GBP-ymScarletI in elm1Δ cells. DC*=Differential contrast. Scale bar5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple-comparison test (**: p<0.01, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=397, elm1Δ: n=408, elm1Δ-Shs1-GBPymScarletI: n=467, elm1Δ-Hsl1-GBP-ymScarletI: n=462, elm1Δ-Bud4-GBP-ymScarletI: n=401 and elm1Δ-Bni5-GBP-ymScarletI: n=317 cells). (C) Quantification of aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, n>170 cells/strain). (D) Graph depicting the raw fluorescence intensity of Gin4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=166, elm1Δ: n=177, elm1Δ-Shs1-GBP-YmScarletI: n=176, elm1Δ-Hsl1-GBPymScarletI: n=188, elm1Δ-Bud4-GBP-ymScarletI: n=185 and elm1Δ-Bni5-GBP-ymScarletI: n=151 cells).

      Author response image 4.

      Artificial tethering of Hsl1-GFP to the bud neck via septins or Nim1-related kinases rescues cellular morphology in elm1Δ cells. (A) Representative images showing the relocalization of Hsl1-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI and Gin4GBP-ymScarletI. Scale bar-5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=508, elm1Δ: n=418, elm1ΔShs1-GBP-ymScarletI: n=535 and elm1Δ-Gin4-GBP-ymScarletI: n=482 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (*: p<0.05, ****: p<0.0001), (N=3, n>165 cells/strain). (D) Quantification of raw fluorescence intensity of Hsl1-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=156, elm1Δ: n=166, elm1Δ-Shs1-GBP-ymScarletI: n=169 and elm1Δ-Gin4-GBP-ymScarletI: n=165 cells.

      Author response image 5.

      Targeted localization of Kcc4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Kcc4-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI, Hsl1-GBPymScarletI and Gin4-GBP-ymScarletI. Scale bar-5µm. (B) Quantitative analysis representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), oneway ANOVA Tukey’s multiple-comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=443, elm1Δ: n=489, elm1Δ-Shs1-GBP-ymScarletI: n=312, elm1Δ-Hsl1-GBP-ymScarletI: n=563 and elm1Δ-Gin4-GBP-ymScarletI: n=337 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (****: p<0.0001, ns: p>0.05), (N=3, n>165 cells/strain). (D) Quantification for the raw fluorescence intensity of Kcc4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (**: p<0.01, ***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=163, elm1Δ: n=168 elm1Δ-Shs1-GBP-ymScarletI: n=161, elm1Δ-Hsl1-GBP-ymScarletI: n=172 and elm1Δ-Gin4-GBP-ymScarletI: n=150 cells).

      Author response table 1.

      Summary table showing rescue of elongated morphology in elm1Δ cells upon forced recruitment of Nim1-related kinases via septins or its related kinases tagged with GBPymScarletI.

      Additional changes:

      The graph in Fig. S3D (revised preprint) has been updated to reflect a slight increase in the Chs2-mNG residence time in both the elm1Δ and gin4Δ strains, whereas our previous version indicated a delay only in the gin4Δ strain. Because the elm1Δ strain exhibited a more pronounced phenotype than the gin4Δ strain, we re-examined the analysis. The revised results show that the residence time of Chs2 during cytokinesis is modestly prolonged by approximately 2 minutes in both backgrounds. Accordingly, the graph and statistical analyses have been updated.

      References:

      (1) Mortensen, E.M., McDonald, H., Yates, J., and Kellogg, D.R. (2002). Cell Cycle-dependent Assembly of a Gin4-Septin Complex. Molecular Biology of the Cell 13, 2091-2105. 10.1091/mbc.01-10-0500.

      (2) Marquardt, J., Chen, X., and Bi, E. (2024). Reciprocal regulation by Elm1 and Gin4 controls septin hourglass assembly and remodeling. J Cell Biol 223. 10.1083/jcb.202308143.

      (3) Asano, S., Park, J.E., Yu, L.R., Zhou, M., Sakchaisri, K., Park, C.J., Kang, Y.H., Thorner, J., Veenstra, T.D., and Lee, K.S. (2006). Direct phosphorylation and activation of a Nim1-related kinase Gin4 by Elm1 in budding yeast. J Biol Chem 281, 2709027098. 10.1074/jbc.M601483200.

      (4) Bouquin, N., Barral, Y., Courbeyrette, R., Blondel, M., Snyder, M., and Mann, C. (2000). Regulation of cytokinesis by the Elm1 protein kinase in Saccharomyces cerevisiae. Journal of Cell Science 113, 1435-1445. 10.1242/jcs.113.8.1435.

      (5) Okada, H., MacTaggart, B., and Bi, E. (2021). Analysis of local protein accumulation kinetics by live-cell imaging in yeast systems. STAR Protoc 2, 100733. 10.1016/j.xpro.2021.100733.

      (6) Barral, Y., Parra, M., Bidlingmaier, S., and Snyder, M. (1999 Jan 15). Nim1-related kinases coordinate cell cycle progression with the organization of the peripheral cytoskeleton in yeast. Genes & Development 13. 10.1101/gad.13.2.176.

      (7) Marquardt, J., Yao, L.L., Okada, H., Svitkina, T., and Bi, E. (2020). The LKB1-like Kinase Elm1 Controls Septin Hourglass Assembly and Stability by Regulating Filament Pairing. Curr Biol 30, 2386-2394 e2384. 10.1016/j.cub.2020.04.035.

    1. eLife Assessment

      This useful study provides new insights into the liver-stage antigen LSA3, its export to erythrocytes, and its role in liver-stage development. While the functional importance of LSA3 is well demonstrated, the data underlying the conclusions regarding antibody specificity, liver-stage localization, and phenotype remain incomplete. A key strength of the study is the use of mosquito and humanized mouse models to access life-cycle stages that are rarely studied in most laboratories.

    2. Reviewer #1 (Public review):

      Summary:

      The extent P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood stage exported proteins tested in liver stages were not exported. An exception is LISP2 that is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages but in liver stages no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studied the localization of LSA3 in blood stages and used a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also showed that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 are thorough and mostly overcame adverse issues such as a cross-reactive antibody and negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      Weakness:

      The cross-reactivity of the antibody together with the co-infection strategy prevents reliable assessment of LSA3 localization in liver stages. Despite of this it seems LSA3 is not exported in liver stages and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. However, this was now done in a separate study focussing on blood stage parasites (PMID: 41135800).

      Due to the difficulty of working with humanised mice, it was not possible to refine some of the conclusions and questions remain:<br /> The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remain largely unknown. The co-infection used (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. It was also not possible to determine if the cross-reactive protein is expressed at that stage. While this might have been evident from the mixed WT/mutant infection (if all cells are positive for LSA3 there is cross-reaction; if about half of the cells are negative, there isn't) but assessing this failed.

      Significance:

      It is important information that LSA3 contributes to efficient liver stage development. However, neither LISP2 nor LSA3 seem to be exported in P. falciparum liver stages and can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

    3. Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood-stage LSA3 undergoes PEXEL processing and a portion of the protein is exported into the erythrocyte where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      (1) The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepesin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also reacts with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      (2) This study provides the first analysis of LSA3 exoerythrocytic function, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectible export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      Weaknesses:

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3-C reactivity at this stage could not be determined and the localization of LSA3 in the liver stage remains somewhat unclear.

      Comment on revisions:

      The authors thoughtfully addressed the reviewer comments.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      Comments on revised version.

      I appreciate the authors' efforts to revise the manuscript and to clarify several aspects of the study.

      However, I remain concerned that some conclusions extend beyond the data presented. In particular, the authors acknowledge in their rebuttal letter that antibody specificity in liver stages could not be validated and that cross-reactivity cannot be excluded. Consequently, the localization data shown in Figure 5 cannot currently be considered definitive evidence for liver stage localization of LSA3 itself.

      Similarly, the revised manuscript appropriately states that LSA3 was not detected beyond the PVM in liver stages and that export into the hepatocyte remains unresolved. Nevertheless, several statements continue to imply a role for liver stage protein export. At present, the possibility that a domain of LSA3 may face the host-cell side of the PVM remains speculative and is not supported by direct experimental evidence.

      The liver stage fitness phenotype is convincing and supports the conclusion that LSA3 contributes to normal liver stage development. However, the current data do not establish the developmental process affected nor connect the phenotype to export beyond the PVM.

      I therefore recommend that the manuscript consistently distinguish between (i) demonstrated export of LSA3 during blood stage infection and (ii) the unresolved localization and trafficking of LSA3 during liver stage infection. I would also encourage the authors to consider revising the title to better reflect the findings presented, as the current title may be interpreted as demonstrating liver stage export, which has not been shown.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The extent to which P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood-stage exported proteins tested in liver stages were not exported. An exception is LISP2, which is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, an incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum, it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages, but in liver stages, no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studies the localization of LSA3 in blood stages and uses a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also shows that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 is thorough and mostly overcame adverse issues such as a cross-reactive antibody and the negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable, as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      We thank the reviewer for their perspective on the strengths of the study.

      Weaknesses:

      The cross-reactivity of the antibody, together with the co-infection strategy, prevents reliable assessment of LSA3 localization in liver stages. Despite this, it seems LSA3 is not exported in liver stages, and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. Given that a previous publication found some inhibitory effect of LSA3 antibodies on blood stage growth, a comparison of the growth of the LSA3 disruption clones with the parent would have been very welcome and easy to do. At this point, LSA3 is one more of many proteins exported in blood stages for which the function remains unclear.

      It might be possible to refine some of the conclusions. The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remains largely unknown. The co-infection (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver, but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. The authors write that they do not know if the cross-reactive protein is expressed at that stage. But this should be immediately evident from the mixed WT/mutant infection. If all cells are positive for LSA3, there is a cross-reaction. If about half of the cells are negative, there isn't. In the latter case, the localization shown in the paper is indeed LSA3, and morphological differences between WT and LSA3 disruption could be assessed without additional experiments.

      We thank the reviewer for their comments. While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Significance:

      The conclusion from the paper that "our study presents just the second PEXEL protein so far identified as important for normal P. falciparum liver-stage development and confirms the hypothesized potential of exported proteins as malaria vaccine candidates" is partially misleading. Neither LISP2 nor LSA3 seems to be exported in P. falciparum liver stages, and we can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

      We thank the reviewer for the comment. We would like to emphasize the possibility that proteins localized at the PVM may be considered exported ‘if’ part or all of the protein (eg, a domain) faces the host cell lumen from the hepatocyte. We have not shown this to be the case for LSA3 or LISP2 but that possibility remains open. Nonetheless, LISP2 is exported (by P. berghei liver stages) and LSA3 is exported (by P. falciparum blood stages); both are exported proteins.

      Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif, but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al, PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood stage LSA3 undergoes PEXEL processing, and a portion of the protein is exported into the erythrocyte, where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepsin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also react with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage, where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      This study provides the first direct analysis of LSA3 function by reverse genetics, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage, but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectable export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      We thank the reviewer for their feedback regarding the strengths of the paper. 

      Weaknesses:

      A previous study reported that anti-LSA3 antibodies inhibit blood-stage growth, suggesting a role for LSA3 during erythrocyte infection. While the authors carefully evaluate the LSA3 mutant in mosquito and liver stages, the impact on blood stage fitness is not tested. While the knockout shows LSA3 is not essential in the blood stage, its importance during erythrocyte infection remains unclear.

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage, and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3C reactivity at this stage could not be determined, and the localization of LSA3 in the liver stage remains somewhat unclear.

      We thank the reviewer for their comments. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      We thank the reviewer for their comments.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      We understand the reviewer’s points. We agree that there is no evidence provided that LSA3 is targeted beyond the PVM; whether any of the protein faces the hepatocyte cytosol is unknown (and challenging to conduct) but this possibility remains plausible. LISP2- and LSA3deficient liver stages are less fit than parental controls and thus we stand by the conclusion that they are the two so far identified P. falciparum PEXEL proteins that are important for liver-stage development.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 163 says: "Altogether, this demonstrates that LSA3 is important but not critical for blood stage growth of P. falciparum": this is based on the cited Morita et al., 2017. However, previously LSA3 was considered dispensable based on a knock out in 3D7 (Maier et al., 2008; PMID: 18614010). Given that the authors generated a mutant for this work, it would be straightforward to test growth and clarify the importance of LSA3 in blood stages. If important, the analysis of the location and transport of LSA3 in blood stages would immediately become more relevant.  Maybe the data for this is already in the paper: the number of stage V gams was similar between mutant and control (Figure 4A). If this was calculated from the total number of asexual starting parasitemia, it includes blood stage growth, and it can be assumed that there is no growth defect in the mutant in the blood stages. If the number of stage 5 gams was calculated from the number of committed schizonts/rings, nothing can be said about blood stage growth, and asexual blood stage growth should be tested in specific experiments.

      We thank the reviewer for raising the function of LSA3 in blood stages and agree it was an obvious omission, though for good reason - a separate, collaborative study was underway. While this eLife preprint was in revision, our accompanying manuscript on the blood stage was published, showing the characterization of our NF54 DLSA3 mutant during blood-stage growth (PMID:41135800). The findings are now summarized and the citation included in the revised version of this preprint.

      Manuscript line 105: "although, notably, functional characterization of lsa3 deletion mutants has not yet been reported to confirm an important function": at least in blood stages, it was reported to be dispensable, see above. The corresponding study (Maier et al., 2008, PMID: 18614010) could be cited in that context. 

      The citation of PMID18614010 and 39913589 have now been added and we thank the reviewer.

      (2) Some questions central to the conclusions of this paper remain because it was unclear whether the serum did indeed detect LSA3 in the liver or not. It would be easy to check if all cells from the WT/Mutant mix experiment show LSA3 signal (this would mean it cross-reacts) or if only about half are positive (the mutants would be negative if there is no cross-reaction). This would be important to mention for Figure 5 because, at present, it is not known that what is labeled by the LSA3-C antibody in these images is (only) LSA3. 

      We thank the reviewer for this point and completely understand. We did check this via microscopy of liver sections co-infected with LSA3 mutant and control liver-stage parasites as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5-day post-infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (3) It is also unclear which parasites were imaged in Figure 5. The text of the results states that NF54 liver stages were used, but later: "As we employed a co-infection strategy to assess the essentiality of LSA3 versus NF54 in mice, we could not perform IFAs on individually infected mice in this study to validate the specificity of LSA3-C at the liver-stage". The legend says NF54 sporozoites on day 5 post-infection were used. I suspect it was a WT/mutant mix, in which case the above applies, and in the absence of cross-reactivity, half of the cells should be LSA3-C negative. If this is not the case, the localization in the liver becomes dubious.

      We apologize for the confusion and have corrected this. In Figure 5, we utilized liver sections from NF54-infected humanized mice that were stored at -80 C from a previously published study (McConville et al, PNAS 2024). Ideally, we would validate the specificity of LSA3 antibodies at the liver-stage using liver sections containing only DLSA3 parasites however the number of mice available was limited and the samples available to us also contained the Control line for qRT-PCR analyses (the co-infection strategy). As mentioned above, we couldn’t distinguish between these two strains by IFA at the time point analysed and this precluded us unequivocally validating the LSA3-C specificity in the liver-stage; however it cannot be excluded that the signal observed at the PVM is indeed LSA3. We are currently focusing research efforts on obtaining more humanised mice to answer this.

      Minor:

      (1) Introduction: Before the part on the PEXEL motifs, there are almost no references; please add references for all statements.

      We have added references.

      (2) Figure 1B is unclear regarding which part of the gene was deleted. The system used would permit a complete gene deletion, but the homology flanks seem to be within LSA3. If parts of the gene are left, the 75 kDa on the western blots might be a degradation product arising from both the truncated and the full-length protein. Please clarify in the sketch exactly where the homology flanks are, with respect to the start and stop of the gene. 

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (3) Line 161: Replace was with were.

      Corrected.

      (4) Figure 2, 224: Why do the authors think LSA3 must be in the luminal leaflet of the PVM as opposed to the outer leaflet of the plasma membrane?

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte.

      (5) Line 245: GFP core "derived from digestion of the reporter in the food vacuole, which confirmed it was secreted from the parasite". I wonder if the amount of GFP "core" really can be used as evidence for secretion, and its amount can be compared between experiments. Did the author quantify this for the full-length protein to get a proportion per sample?

      Use of GFP core to measure defects in P. falciparum GFP reporter secretion has been described previously (for example PMID:23387285 and 35906227). The comparison the reviewer asked for is an interesting and important question: however the control would be to compare the ratio of GFP core to uncleaved in the control lanes as well, which is not possible to do since the full-length protein is digested by plasmepsin V in the native PEXEL versions of the experiments (mLSA3-GFP in the first blot, Vehicle in the second blot) leaving no full-length protein to compare to. It stands to reason that inhibition of N-terminal processing results in less protein removal from the membrane (ER or COPII vesicle or PM) resulting in less secretion out of the parasite for retrograde transport to the food vacuole with cytostomal vacuoles (analogous to plasmepsin II). In the food vacuole, the chimeras are in normal cases digested by proteases back to the GFP core that is resistant to cleavage and evident as GFP core on the immunoblots (PMID:10775264 and 14709539 and 19055692 and 20130643). 

      (6) Figure 3 has the word plasmid in two lanes. In Figure 3E, amend the labelling of the blots.

      We apologize for the formatting error in converting the figures to PDF during the original submission and thank the reviewer for the suggestion. This has now been corrected.

      (7) Lines 266/271/284: "live IFAs", live immunofluorescence assay. Does this mean an antibody was given to living   parasites?

      The correct term is live microscopy and this has been corrected.

      (8) Does Figure 6A fit with the data in Figure 6B? It seems 6B has a milder phenotype than 6A.

      We thank the reviewer for the question. Yes the data directly correspond to each other and are represented in two ways: Panel A shows the qRT-PCR raw data for liver load of each parasite strain per humanized mouse using a scientific scale on the y-axis. Panel B shows that magnitude of the DLSA3 defect as a percentage of the total liver load per mouse:

      % total parasite liver load  = ( strain 1 or strain 2 liver load ) x100

      sum of strain 1 + strain 2 liver loads

      The intent of showing both data is to convey the correct magnitude of the difference in two ways to assist the reader in understanding the true defect, both are accurate and both are statistically significant. In revision we detected mislabelling of humanized mouse 2 and 3 in the original graphs that has now been corrected and we sincerely thank the reviewer for helping us identify this error.

      (9) Line 482: Please add references for this debate. 

      These have been added.

      Reviewer #2 (Recommendations for the authors):

      Major Comments: 

      (1) In general, the authors have taken care not to overstate conclusions from their study. Nonetheless, while not technically inaccurate, the title might misleadingly suggest LSA3 is exported in the liver stage (this was my initial impression on reading it until I looked at the data). I suggest the authors revise the title to avoid confusion by clarifying that export was only observed in the blood stage.

      We sincerely appreciate the reviewer’s point. As this article was posted as a preprint that has now been cited several times, we have carefully weighed the comment and in the end decided to retain the current title for the above reason.

      (2) While the ability to generate the ∆LSA3 parasites clearly shows that the protein is not essential in the blood stage, the impact on parasite fitness is never tested but simply assumed (for instance, in lines 163-164: "...this demonstrates that LSA3 is important...for blood-stage growth..."). Do the ∆LSA3 parasites have a fitness defect in the blood stage consistent with the previous GIA data that would support this claim? Since the rabbit anti-LSA3-C antibodies produced by Morita et al did not have GIA activity against the blood stage, it is possible that the GIA observed with the human and mouse antibodies might have been due to reactivity with a different protein. If ∆LSA3 does cause a fitness defect, it would be interesting to know if the endogenous GFP-tagged line, which alters protein trafficking/membrane association, also produces this effect.

      We agree with the reviewer and would like to clarify that this omission was not intended to create confusion but was by design, due to a separate collaborative study that was underway to address such questions. While this eLife preprint was in revision, our accompanying manuscript on characterising NF54 DLSA3 at the blood stage was published (PMID:41135800). The findings are now summarized and the citation included in the revised version of this eLife preprint. In sum, LSA3 is not critical for erythrocyte invasion but its deletion perturbs the rate and efficiency of merozoite invasion, at the step(s) of resealing of the PVM/host cell, resulting in aberrant accole forms that protrude from the infected erythrocyte.

      (2) Figure 1D: While the images are compelling and I don't doubt the claim that LSA3 is exported in the blood stage (also supported by the fractionation/Pk experiments), the authors should provide quantification of the difference in exported signal between the WT and ∆LSA3 parasites in these IFAs to rigorously support this conclusion. Also, please include details about how many independent experiments are represented by the microscopy data throughout the manuscript (Figures 1, 2, 3, and 5).

      We understand the reviewer’s request and wish to indicate that the export signal was absent in all cells infected with DLSA3 that was imaged. The microscopy performed was from n=2-3 experiments except for Figure 5 which was from n=1 humanized mouse per time point in which multiple EEFs from the liver were imaged. This has been indicated in the figure legends. 

      (3) Careful inspection of the z-series images in Figure 5A shows that most of the LSA3-C signal seen outside the PVM (beyond the boundary delineated by EXP1) is closely associated with DAPI puncta, suggesting these are merozoites. Together with the prominent gap in the EXP1 signal, this suggests the schizont has already ruptured. Thus, anti-LSA3-C signal beyond the PV seems best explained as coming from merozoites or other material released by PV rupture, not from export across the PVM, and this should be added to the text in place of comments about localization to PV extensions or potential export (lines 358-359, 422-423).

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have added the comment as requested.

      Minor Comments:

      (1) The authors may want to denote the disordered repeat region in the LSA3 schematic in Figure 1A that is mentioned in the text.

      We have added the residue boundaries of the predicted domain from AlphaFold into both the schematic and the text and included a link to the LSA3 pages in PlasmoDB and

      AlphaFold in the Methods section.

      (2) The authors use rabbit anti-LSA3-C antibodies previously generated by Morita et al. These polyclonal antibodies were raised against a recombinant C-terminal region of LSA3 (residues 750-1433), but the schematic in Figure 1A indicates the antibodies recognize a smaller region between residues 1154-1433. Please adjust the figure accordingly, or if this is not the same antiLSA3-C antibody reported by Morita, please provide details about its production.

      The figure is corrected.

      (3) The authors use Alphafold to identify a region of LSA3 with similarity to the substrate binding domain of DnaK, but the data is not shown. Please include the Alphafold prediction in supplementary figures and provide information about how the predicted structural homology was determined.

      We have added a link to the AlphaFold page for PF3D7_0220000 in the methods.

      (4) The schematic in Figure 1B indicates that the DHFR cassette was inserted at an internal site within the lsa3 gene. If this is the case, it seems possible that an N-terminal portion of the protein is still expressed, but I was unable to find details about the boundaries of the homology flanks to determine the precise insertion site. Please clarify the knockout strategy and indicate the specific insertion site.

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (5) Line 162: I think this should read "antibodies that react with LSA3 were...".

      Corrected.

      (6) Figure 1D: The merge with the transmitted light channel is missing for the third panel in the ∆LSA3 IFAs. Also, please define the scale bar length in the legend.

      Corrected.

      (7) Lines 744-746: The IFA fixation panel order description (top, bottom) in the Figure 2A legend is reversed from what is shown in the actual figure. Also, please define the scale bar length. 

      Corrected.

      (8) Lines 184-186: Since the fractionation/PK protection assays suggest most of LSA3 is in the PV, it would be interesting to know if the strong peripheral/PV signal observed in the PFA-fixed IFAs in Figure 2A is also present in the ∆LSA3 parasites, or is this non-specific? 

      Thank you for the suggestion. We agree this would be an interesting result to know but do not have the capacity at the present time.

      (9) Lines 219-225: It is unclear to me why these results are interpreted to suggest that the majority of LSA3 is peripherally associated with the luminal leaflet of the PVM. Wouldn't an integral membrane configuration in the PVM (with the C-terminus facing the host cytosol) or PPM (with the C-terminus facing the parasite cytosol) also account for the data? Adding a carbonate extraction would help clarify this point.

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM-associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte. If the question is whether LSA3 is an integral PVM protein, we agree that use of carbonate in the future would answer that question.

      (10) Figures 3D and E: There are some problems with some of the text wrapping in these panels.

      We apologise, this was a formatting issue as the manuscript was converted to PDF.

      We have corrected this error.

      (11) Line 422-423: In fact, the Z-sections shown in Figure 5 appear to indicate that the LSA3-C signal is predominantly located within the parasite, not at the PVM.

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have corrected the final conclusion to be more accommodating of this.

      (12) Lines 468-470: Since cross reactivity of anti-LSA3-C is substantial in the blood stage but was not defined in the liver stage by analysis of unmixed infections, how do the authors know that they were not observing ∆LSA3 parasites in their IFAs? I think what they mean here is that parasites lacking anti-LSA3-C reactivity were not observed, which is an important distinction.

      The reviewer is correct and this has been corrected.

      (13) Lines 478-479: The authors should also mention that the P. berghei PEXEL proteins evaluated in Fougere et al are exported in the blood stage, similar to LSA3. Moreover, other studies have shown something similar for additional endogenous PEXEL proteins or reporters in P. berghei (PMIDs 22329949, 26347246, 34956312).

      We have added the additional text regarding export into the infected erythrocyte and the reference to IBIS1.

      (14) Line 491: The data here don't support that LSA3 is "required" for liver stage development, only that it is important to it. Since the authors have not defined the cross-reactivity of anti-LSA3C in unmixed infections, it is not clear that ∆LSA3 parasites are arrested early in the liver stage, only that they show a reduced number of genome copies relative to the parental control. 

      We have amended the sentence to “required for normal liver stage development”.

      (15) Line 530: I think NGF54 should be NF54.

      Corrected.

      Reviewer #3 (Recommendations for the authors):

      (1) Antibody specificity in liver stage IFA experiments:

      The specificity of the anti-LSA3 antiserum (LSA3-C) used in liver stage IFA is not fully convincing. While the KO parasites were used effectively to validate specificity in blood stages, the same is not true for liver stages. 

      (a) It is essential to repeat IFA with ΔLSA3 parasites in liver stage infections to determine whether the observed PVM staining is truly specific.

      We appreciate the reviewer’s point, however at a cost of over $5000 per humanized mouse, we do not have the capacity to conduct this experiment at the present time. We highlight that, as the blood stage IFAs confirmed the specificity of LSA3-C for LSA3, the possibility remains open that LSA3 is specifically recognized at the PVM.

      (b) If the antibody is the same polyclonal serum used in Morita et al. (2017), why did the authors not employ a monoclonal antibody, which they presumably have access to and which would provide greater specificity? 

      We have included new data confirming that LSA3 is exported using LSA3-T, in addition to LSA3-C.

      (c) Given that rabbit antisera often show non-specific staining at the PVM in liver stage parasites, co-localization with PVM markers is not sufficient. Inclusion of the ΔLSA3 parasites in liver stage IFA is critical. It will also show whether there is any cross-reaction of the antiserum in liver stage parasites, as seen by IFA for blood stage parasites. 

      We thank the reviewer for their feedback.

      (d) To validate the serum further, the authors should infect HC-04 cells in vitro with GFP-LSA3 parasites and stain with LSA3-C to confirm overlap between the tagged protein and the antibody signal.

      We thank the reviewer for their feedback.

      (e) For higher-resolution co-localization, expansion microscopy - now commonly used even in malaria research - would substantially improve the analysis. 

      We thank the reviewer for their feedback.

      (2) The localization of LSA3 in this study differs notably from Morita et al. 2017, who reported localization to dense granules in merozoites and staining in ring-stage parasites at the PVM. 

      (a) The authors confirm DG localization, but they do not examine ring-stage parasites. They should include the IFA of ring stages to clarify whether they can replicate the previous findings.

      We thank the reviewer for their feedback.

      (b) Additionally, the differences in Western blot banding patterns between the two studies should be addressed. Do the authors have an explanation for these discrepancies? 

      We thank the reviewer for their feedback.

      (3) The authors report a ~40% reduction in liver parasite load using qPCR, which is statistically significant. However, this phenotype is modest and should not be interpreted as showing that LSA3 is essential.

      (a) Please avoid terms like "required" or "essential" and instead describe the protein as "contributing to normal development" or "influencing fitness."

      We have used the term “required for normal liver stage development”.

      (b) Since the authors generated liver sections, they should take advantage of these to quantify the number and size of liver stage parasites, which would help determine whether the phenotype reflects fewer infected cells or reduced parasite growth.

      We did check this via microscopy of liver sections, but all mice were co-infected with LSA3 mutant and control liver-stage parasites, as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5 day post infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (c) It would also be valuable to include IFA from singly infected ΔLSA3 livers (rather than co-infected), and possibly at earlier timepoints, to identify the developmental window affected.

      We agree it would be valuable.

      (4) The manuscript suggests that LSA3 may be exported beyond the PVM into the hepatocyte, based on a small number of peripheral puncta.

      (a) This claim is not convincingly supported by the data. The punctate signals shown in Figure 5 are weak and may rather reflect PVM extensions or TVN. In fact, one punctum even overlaps with the DAPI signal (figure 5, middle panel), which raises further doubt about the localization.

      We appreciate the reviewer’s careful eye and caution and have added the comment regarding DAPI.

      (b) Given the lack of KO controls in these liver stage IFAs, the authors should not describe LSA3 as "exported beyond the PVM". The language should be revised to reflect that the protein localizes predominantly to the PVM, and any extra-PVM signal remains unconfirmed and could be non-specific. 

      (c) This is especially important given the well-known tendency of rabbit antisera to produce background PVM staining in liver stage parasites. 

      Corrected.

      (e) In an earlier report (McConville et al, 2024, PNAS), they clearly state that LSA3 is NOT exported beyond the PVM. Actually, the staining in the previous report looks quite different from the images provided for Figure 5. The authors might wish to comment on this. 

      We thank the reviewer for their feedback.

      Minor comments:

      In some sections, the manuscript uses "exported" to refer to trafficking to the PVM. This terminology should be used more carefully and consistently, since "export" often implies translocation into the host cytosol

      We understand that export involves a protein localizing within the host cell and so protrusion through the PVM may also be considered exported, however, we have not confirmed this for LSA3 in liver stages.

    1. eLife Assessment

      This study offers an important set of literature-based parameter ranges for calibrating cardiac electromechanical models. The key scientific claims are largely supported by convincing evidence and carefully-documented methodologies, but a few specific issues remain only partially supported, in particular relating to the locking-free formulation, reporting of adequate details regarding the electrophysiological model (Eikonal approach), and the fact that physiological accuracy is predominantly achieved in a single chamber (the LV).

    2. Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains and other biomarkers. The overarching aim of the study is to give credibility to the underlying computational electromechanics framework as a step towards the cardiac Digital Twin vision.

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study relies on several simplified modeling assumptions that do not reflect the current state-of-the-art:

      a) Isotropic scaling of the ventricular mesh to generate an unloaded reference geometry.

      b) Simplified afterload and preload models that do not consistently capture the full range of physiological responses.

      c) Simplified epicardial boundary conditions.

      These limitations are appropriately acknowledged and discussed by the authors in a dedicated Limitations section.

      (2) Numerical Framework:

      a) The numerical framework used for the mechanical part of the model may be susceptible to locking effects that could contribute to artificially stiff and less contractile behavior. This is indicated by a ten-fold scaling of the peak active contractile force parameter relative to literature values and notable sensitivity of the model to the tissue compressibility parameter. While - as acknowledged by the authors - part of this can be attributed to simplified modeling choices, other comparable studies have not reported similar issues.

      b) The human electrophysiology model is not described in enough detail. Currently, it is not mentioned in the manuscript that an Eikonal model was used to compute activation times on the endocardial surface to be robust against coarse mesh resolutions.

      (3) Geometrical model and digital twin: The model presented combines anatomical data, electrical measurements, and physiological reference values from different individuals or population averages, rather than being derived from a single patient. The authors have appropriately moderated their claims in the revision, framing the work as a step towards the cardiac Digital Twin vision rather than asserting that the model itself constitutes a digital twin.

      (4) Calibration procedure: The description of the calibration procedure has been substantially improved in the revision. The authors now provide explicit rationale for each calibration step and clarify that the procedure targets multiple physiological biomarkers in sequence. Verification that the calibrated model produces physiological cellular dynamics, including intracellular calcium transients, is now provided. The revised manuscript also shows the simulated electrocardiogram alongside population reference ranges, which partially addresses the question of calibration quality. However, a direct comparison of the simulated electrocardiogram with the individual measured signal used for calibration is not provided, which would give a more stringent and direct assessment of how well that specific calibration target was achieved.

      Comments on revised version.

      The revision represents a genuine improvement. The calibration procedure is now more transparently described, physiological cellular dynamics are verified, the digital twin framing has been appropriately moderated to reflect a step towards that vision rather than a claim of having achieved it, and an expanded limitations section identifies where the framework falls short.

      Several of the concerns raised in the first round have been addressed, but some issues remain:<br /> The model still requires a ten-fold scaling of the peak active contractile force relative to literature values, and the imbalance between left and right ventricular output persists, along with non-physiological right ventricular pressures and ejection fraction. These are partly fundamental limitations of the current modelling approach that may not be fully resolvable within the scope of this paper, and the authors are to be credited for acknowledging them. However, they do constrain the conclusions that can be drawn about the credibility of the framework for reproducing healthy cardiac physiology.

      The population-averaged reference dataset and the calibration and validation framework remain contributions of value to the community, and the revised limitations section adds useful transparency about the current state of the art.

    3. Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Comments on revised version.

      The authors have addressed my prior comments

    4. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the reviewers for their efforts in reviewing our manuscript. We highlight that the key contributions of our paper are to provide a framework for calibration and validation of high-fidelity cardiac electromechanical models based on a diverse compilation of clinical datasets, and that we provide one example of such an evaluation of our own baseline electromechanical model. The comments raised by the reviewers were chiefly focused on the second goal, which is our specific model and the outcomes of the evaluation process, rather than on the evaluation framework itself. As such, we have made improvements to our model implementation and to provide additional confidence in our specific modelling framework through this review process. Specifically, we have strengthened the verification component of this evaluation, provided additional quantitative measures, and included a more in-depth discussion of the remaining limitations in our modelling framework. We hope that the updated version of the manuscript and our efforts to improve it are well-received by our reviewers and editors, as well as by the modelling and simulation community at large.

      eLife Assessment

      This is a potentially important study that explores the relevant range of parameter values for calibration and validation of cardiac electromechanics in ventricular models. Although much of the work presented is solid, the evidence provided to support the authors' key scientific claims is incomplete, especially as it relates to the emphasis on standardized validation and verification approaches. Notably, the level of model personalization presented in this work falls short of the threshold for what could reasonably be called a "digital twin", even by the relatively relaxed standards that have emerged in computational physiology and related fields in recent years.

      We appreciate the eLife assessment for identifying the potential importance of our study. Regarding the threshold for 'digital twin', we note that a cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning, which is a goal that to our knowledge no published electromechanical study has simultaneously fulfilled. It is for this reason that the community refers to the 'digital twin vision' rather than its realisation, and we adopt this framing consistently throughout the manuscript.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers within a single study. In our revision, we have clarified that the model evaluation presented here is an example application of that framework, which was designed not to certify a model as complete, but to provide a transparent audit of current capability that identifies where confidence is established and where further development is needed. We have updated the title and language throughout the manuscript to reflect this framing consistently.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction, and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains, and other biomarkers. The overarching aim of the study is to give "credibility to the underlying computational electromechanics framework" and to "pave the way towards credible cardiac electromechanical Digital Twins."

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study employs simplified modeling assumptions that are not state-of-the-art, e.g.,

      (a) Isotropic scaling of the mesh to generate an unloaded reference geometry.

      (b) Simple afterload and preload models that fail to produce physiological results.

      (c) Simplified epicardial boundary conditions.

      While our model was able to broadly achieve physiological behaviour based on the calibration and validation datasets, it also contains several simplifications that can be expanded with more sophisticated techniques to allow explorations in specific areas. We have added a dedicated Limitations subsection to the Discussion section of the manuscript to address these and to provide references to relevant studies.

      (2) Numerical Framework:

      (a) The mesh resolution and/or the numerical framework used for the mechanical part appears to suffer from known numerical artifacts (locking effects), leading to overly stiff or inaccurate behavior in finite element analysis. This results in an artificially stiff response to deformation, which is compensated by setting active contraction to ten times the value reported in the literature. The authors attribute this to limitations in using ex vivo tissue measurements to represent in vivo function, although similar issues were not observed in previous works.

      We thank the reviewer for raising this point and have investigated it carefully. We have added a verification section as well as discussions to the manuscript to more comprehensively discuss this point. In short, through various tests against benchmark (Land) and comparing stress-strain curves in cube simulations with the same mesh resolution, we could not identify evidence of volumetric locking effects. The elevation in contractile force was also necessary in a simplified ellipsoid version of the model in a previous publication [ref 7, Levrero-Florencio, et al. 2020]. We note that these benchmarks were performed in the incompressible transversely isotropic regime; whether analogous locking effects exist in the dynamic orthotropic active contraction framework used in the full biventricular simulations remains an open question, which we have identified as a priority for future benchmarking, for example against the Arostica et al. 2025 benchmark.

      We agree that the explanation of this as ex vivo vs in vivo difference in contractile force is too simple, and other contributing factors are better understood through comparison with similar studies in the field. Strocchi et al. (2023) used a four-chamber model with explicit atrial mechanics, and in her history matching varied Tref within +- 33-55% of a reference value of 120-150 kPa, targeting a peak active tension of 160 +- 15 kPa, which was a considerably more modest adjustment than applied here, likely reflecting differences in model geometry, pericardial constraint, and circulatory model between the two studies. Gerach et al. (2021, Mathematics) applied manual parameter adjustments informed by in vivo active tension measurements of 120 – 150 kPa and achieved ejection fractions of approximately 63%; however, they reported that systolic pressures in both ventricles were too high for a healthy heart, and similarly reported elevated peak ejection rates compared to MRI measurements, a difficulty we also encountered. Notably, Gerach et al., report that atrial contraction contributes approximately 11-13% of end-diastolic volume, which in a biventricular-only model would directly reduce the achievable LVEF and necessitate compensating adjustments to active tension. Zingaro et al. (2024, Journal of Computational Physics), using an alternative active tension model (RDQ20), similarly found it necessary to increase contractility parameter (a_XB) to achieve sufficient ejection, and explicitly report that no single parameter configuration simultaneously achieved physiological peak ejection rate and LVEF, a fundamental tension we also encountered. Together, these comparisons suggest that the elevated Tref in our model most likely reflects a combination of the absence of atrial filling, simplifications in pericardial constraint, and the lack of poroelastic behaviour, rather than volumetric locking alone. Additional investigations are needed as explained in the manuscript, and these factors are identified as open priorities for future development within our framework.

      We added a comparison with these three studies to the discussion section of the manuscript and have updated the limitations section to reflect this more nuanced account of the factors contributing to the elevated active tension scaling. Further work will be required to address these points.

      (b) Further, the authors employ the monodomain model for the simulation of the electrical excitation and relaxation on a relatively coarse grid with an approximate edge length of 1mm. This resolution is known to be insufficient for reliable results in organ-scale electrophysiology modeling.

      Our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached, as performed in Camps et al. (2024) (ref 33) using the tuneCV tool in monoAlg3D (https://github.com/rsachetto/MonoAlg3D_C/tree/master/scripts/tuneCV), which is similar to the tool in openCARP (described here: https://opencarp.org/documentation/examples/02_ep_tissue/03a_study_prep_tunecv).

      While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for simulations of ECGs in this study. We have noted this in the methods section.

      (3) Geometrical model and digital twin: The geometrical model, taken from a public cohort and calibrated to an ECG of another individual along with population-averaged values from a databank (UK Biobank), and unrelated measurements from surgical procedures, can hardly be considered a digital twin. Further, validation of the model was then performed against data from yet another cohort.

      We thank the reviewer for this point and welcome the opportunity to clarify our dataset choices. The use of multiple data sources was a deliberate methodological decision. While an ideal dataset for electromechanical model evaluation would combine full biventricular geometry, 12-lead ECG, invasive pressure measurements, and myocardial strain data from a single individual, no such dataset currently exists in the public domain, and acquiring it routinely would be impractical in clinical settings. Multi-source integration therefore reflects the realistic deployment scenario for future clinical translation of these tools.

      The specific choice of geometry was principled: the mesh associated with the ECG dataset that was available to us was truncated at the base due to the clinical acquisition protocol, which would have prevented physiologically realistic basal boundary conditions. The female Rodero geometry we chose provides full ventricular coverage and was selected on that basis.

      Demonstrating that a coherent, systematically evaluated framework can be constructed from compiled multi-modal data is itself a contribution because it makes the tools accessible to the wider community without requiring a single ideally acquired dataset.

      (4) Calibration procedure: There are apparent flaws in the calibration procedure, or it is not described in sufficient detail. The authors dedicate significant effort to motivating parameter ranges, but in the end they use mostly other parameters for the calibration process, aiming to maximize left ventricular ejection fraction. It is not clear whether the chosen parameters result in, e.g., physiological calcium traces or calibrated parameters that are within physiological ranges.

      Thank you for raising this point, which we have now clarified in the manuscript. The parameters that were chosen for the calibration process were based on the results of the sensitivity analyses.

      In addition, we have supplemented results Figure 1 with a subfigure F showing that the calcium transient and action potential durations fall within physiological ranges after calibration.

      (5) Goodness of fits, e.g., a direct comparison of the measured and the simulated ECG, are not provided to assess calibration quality.

      The calibrated model achieves QRS duration of 89 ms and QT interval of 360 ms, both of which fall within the healthy reference ranges compiled in Table 2, providing a biomarker-level assessment of calibration quality (Figure 1A). A full quantitative goodness of fit analysis of the simulated ECG morphology was performed following the methodology of Camps et al. [52], in which the same beat-averaged ECG was processed; we direct the reader to that work for full details rather than reproducing the analysis here.

      (6) Due to these limitations and weaknesses, the authors fall short of achieving some of their goals, particularly establishing credibility for the underlying computational framework and in reproducing healthy pressure-volume loops, and in achieving physiological simulations while using physiological or reported ranges for the calibrated parameters.

      For example, a key physiological requirement is that the right and left ventricular stroke volumes are approximately equal in a heart beating at a limit cycle, as the blood pumped by the right ventricle into the pulmonary circulation must match the amount pumped by the left ventricle into the systemic circulation. This balance is not achieved in this study.

      We thank the reviewer for identifying the stroke volume imbalance. We acknowledge the physiological requirement that, in a steady-state limit cycle, the right ventricular stroke volume must approximately equal the left ventricular stroke volume. However, since our model does not explicitly prescribe volumes, to achieve this, we would need to either explicitly tune active tension for the left and right ventricles separately, such as done in https://www.frontiersin.org/journals/physiology/articles/10.3389/fphys.2021.716597/full or develop a more sophisticated circulatory model and employ a multistep procedure that sequentially tunes circulatory dynamics, passive mechanics, and active contraction, such as done in https://www.biorxiv.org/content/10.64898/2025.12.11.693778v1.full. Both of which are beyond the scope of this paper.

      We note that, despite the absence of explicit RV calibration, the RV volumetric measures and pressures remain within physiological ranges, suggesting that the coupled biventricular mechanics are broadly plausible. As such, we have noted this limitation in our discussion section, and sign-posted to other studies where the stroke volume match is achieved.

      (7) The conclusive claim that "the study paves the way towards credible electromechanical cardiac Digital Twins" is not supported. The model exhibits non-physiological behavior, requires unsupported parameter alterations (such as a 10-fold active stress scaling), and does not represent a digital twin, as model data are drawn from various unrelated, non-patient-specific sources.

      We thank the reviewer for this comment, which gives us the opportunity to clarify our use of the term 'digital twin'. A cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning. This is a transformative goal for precision cardiology that the field is actively working towards, with credible, systematically validated electromechanical models as its essential foundation. To our knowledge, no published study in cardiac electromechanical modelling has simultaneously fulfilled all three requirements, and it is for this reason that the community often refers to the 'digital twin vision' rather than its realisation.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers in a single study. The model evaluation presented here is an example application of that framework. Importantly, the framework is not designed to certify a model as complete, but to provide a transparent audit of current capability by identifying where confidence is established and where further development is needed. In this sense, the limitations surfaced through this evaluation are themselves a contribution: they define open problems and priorities for the field.

      We have added a definition of the digital twin concept and the roadmap towards its realisation to the introduction and have updated the language throughout the manuscript to consistently reflect the distinction between the framework contribution and the model evaluation. We maintain that this transparent approach represents a meaningful step towards the digital twin vision.

      The specific limitations of the current model implementation are addressed in detail in the relevant sections of this response and in the updated manuscript, where we have substantially strengthened the verification and discussion components.

      Conclusion:

      Overall, this reviewer considers that the study requires a major revision, including improvements in numerical methods, modeling choices, and checks for physiological behavior. Nevertheless, the provided tables with averaged values from the UK Biobank and the presented validation strategy could be valuable to the research community.

      Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      We thank the reviewer for this comment.

      We have strengthened the quantitative aspect of the validation by reporting the peak simulated strain values for each component and comparing them against the physiological ranges compiled in Table 2. Specifically, the simulated peak strains were: E_ff ≈ -0.20, E_cc ≈ -0.15, E_rr ≈ +0.15, and E_ll ≈ -0.23. These show broad agreement with the in vivo reference ranges from Moulin et al. (2021), noting that the reference ranges are derived from a cohort of 30 subjects and therefore represent a relatively narrow population sample. Shortening strains (fibre and circumferential) are in good agreement, while radial strain is underestimated. We have noted this as a limitation. We have also indicated that a further validation would include a fully quantitative spatial comparison, such as a displacement error norm or voxel-wise strain difference map. This would require access to the raw image data and patient-specific geometry registration, which is beyond the scope of the current study.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      Thank you for raising this point. We have now included a section on model verification to the manuscript at where we perform benchmarking simulation using the Land (2015) passive inflation benchmark. We also provide a mesh subdivision analysis of the final calibrated model. Additional verifications and previous sensitivity analyses using the same numerical scheme with idealised ellipsoid geometries are also referenced in the verification section, to provide additionally confidence.

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Numerical analyses for the Alya solver used in this study has previously been published in works including Levrero et al (2021), which performed sensitivity analyses in a truncated ellipsoid geometry, and Santiago et al (2018), which demonstrated mesh convergence in a cantilever. We have updated the manuscript to point the reader to these studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major concerns:

      (1) Active stress scaling:

      The initial value for T_ref appears to be 120kPa * 10, which would be ten times the literature value fitted to human contraction data. Additionally, Table 2 lists a range of [1200-2400], which is 10 to 20 times the literature value.

      This discrepancy suggests that other model parameters, model assumptions, or the numerical scheme may be inadequate. In contrast, similar calibrations using comparable models (ToRORd-Land) in other works, such as Strocchi et al. [29], yielded T_ref values close to the literature value.

      We thank the reviewer for this comment. As discussed in our response to the public review comment 3a, the elevated T_ref scaling warrants explanation.

      We note that Strocchi et al. use a four-chamber geometry include atrial mechanics and a different pericardial constraint, any of which could contribute to differences in the required T_ref scaling. The elevated scaling in our model likely reflects a combination of factors including the absence of poro-elastic behaviour, simplifications in pericardial constraint, and the lack of atrial mechanics, rather than volumetric locking alone. We have added text to the discussion acknowledging this more explicitly and have flagged planned additional benchmarking of the dynamic orthotropic scheme as future work.

      (2) Non-physiological results, see Figure 1:

      In a healthy heart, RV stroke volume should approx. match LV stroke volume. This is clearly not the case in Figure 1B, where the RV EF is also notably low at 35%.

      Consequently, the study fails to reproduce healthy pressure-volume loops, undermining its claim to create a credible cardiac electromechanical digital twin. Hence, also the "Question of interest" posed in line 206 must be answered with a clear "No".

      Matching stroke volumes should be a primary calibration goal.

      We thank the reviewer for this comment. We agree that stroke volume balance is an important physiological criterion, and we have added it explicitly to the framework criteria in the updated manuscript, noting that our current model evaluation does not satisfy it. This is precisely the kind of transparent appraisal the V&V40 framework is designed to produce: a systematic accounting of which criteria are met and which require further development. A framework that only gets applied to models that pass all criteria would be selection-biased and less informative to the community.

      However, we respectfully disagree that the question of interest must be answered with a clear 'No'. We draw the reviewer's attention to the quantities of interest defined in the paper, which are predominantly left ventricular biomarkers, reflecting the intended scope of the calibration framework. The framework successfully reproduces these defined quantities of interest, and the LV pressure-volume loops, strain, volumes and ejection fraction are all within physiological ranges and well-matched to reference data. These quantities of interest were selected based on their clinical implications in cardiac diseases, as detailed in Table 3.

      We agree that stroke volume balance is an important physiological requirement for a fully calibrated biventricular model, and we have strengthened the future work and limitations section accordingly.

      (3) Inadequate numerical framework:

      (a) Monodomain model: The geometries from Rodero et al. [25] have an average edge length of 1mm. It is known that such a coarse resolution leads to inaccurate EP results. It is not mentioned if the authors refined that geometry to an appropriate resolution or used an Eikonal model to mitigate this issue.

      As explained earlier, our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface, as performed in Camps et al. (2024), and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached. We have added a figure in the appendix of this manuscript to show that by increasing the mesh resolution by one subdivision, we get virtually identical ECG simulations. While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for this study. We have noted this in the methods section.

      (b) Material law:

      - recent publications show that an unsplit deformation gradient for the anisotropic contribution is beneficial to reduce locking effects, see, e.g., Gueltekin et al. Computational Mechanics 63, no. 3 (2019): 443-53. https://doi.org/10.1007/s00466-018-1602-9.

      - K_ct is a penalty parameter to enforce some degree of incompressibility. Results are highly dependent on the grid size and the finite element formulation due to locking effects.

      As the authors write: "In our simulations, we saw that the LVEF was strongly sensitive to changes in the incompressibility of the tissue (Kct), such that an increase in compressibility of the myocardial tissue helped to increase LVEF." Which exactly points to the issue of locking effects.

      So an option would be to use a finer grid or a more adequate numerical scheme with quadratic finite elements, as eg. in [5] Fedele et al., or [6] Gerach et al,. or stabilized elements as in Karabelas et al. CMAME 394 (2022) https://doi.org/10.1016/j.cma.2022.114887.

      Overall, this does not point to "limitations in using ex vivo tissue measurements to represent in vivo function" but to limitations in the numerical setup. In fact, with an adequate numerical scheme, the simulations should be largely insensitive to the choice of this penalty parameter K_ct. See, e.g., Karabelas et al. above, where the authors varied K_ct from 650kPa to infinity (representing an incompressible material), and there is no visible influence on the PV loops.

      We investigated this point using the Alya solver, and we found that the mesh resolution did not alter the LVEF, and our benchmark simulations against Land (2015) did not show the existence of the volumetric locking issue that the reviewer refers to. It is possible, however, that such an effect exists in the elastodynamic orthotropic framework but not in the incompressible and transversely isotropic framework that the Land (2015) benchmarks were set up in. Future analyses could focus on performing additional benchmarking against more recent elastodynamic benchmarks, such as presented in Arostica (2025). We have updated the limitations text in our manuscript to reflect this and to cite relevant literature on this issue.

      (4) Boundary conditions:

      "This was a simplified version of the method [28], which uses an exponential decay formulation at the 'edge' of the pericardial constraint rather than a step function": I don't really see this in the cited work [28] which gives a spatially varying Robin-type boundary condition at the whole epicardium (i.e. regional scaling of normal springs stiffness based on image-derived motion from CT images) and not only at the edge.

      This is motivated by the fact that the pericardial tissue is in contact with various organs of different material properties. Not using spatially varying pericardial parameters is a limitation that might lead to non-physiological deformations, see also Pfaller et al. Biomechanics and Modeling in Mechanobiology 18 (2019): 503-29. https://doi.org/10.1007/s10237-018-1098-4.

      We thank the reviewer for this point and we have corrected the manuscript accordingly. To clarify: our implementation applies a uniform Robin spring constraint along the majority of the epicardial surface with zero constraint at the base, which is conceptually similar to Strocchi et al. [28]. The key difference is that Strocchi et al. use a smooth gradient transition from uniform constraint to zero constraint near the base, whereas our implementation uses an abrupt step transition. We acknowledge that a smooth spatially varying transition would more accurately represent the frictionless pericardial contact and have noted this as a limitation in the manuscript with reference to Pfaller et al. [41].

      Also check:

      - line 117: Gamma_valve_epi is introduced but not used. Was there any boundary condition defined on this valve plane?

      - the third equation, maybe (0,T] missing.

      - line 120: epicardium instead of endocardium.

      These errors have been corrected in the updated manuscript. No boundary conditions were applied on the epicardial surface of the valve plugs, the reference to gamma_valve_epi has been removed.

      (5) Reference geometry:

      The choice to scale the mesh to a lower volume for the unloading procedure seems questionable. This approach does not ensure that the reloaded mesh aligns with the mesh derived from image data. As a result, the geometry used for the simulations is no longer truly patient-specific.

      This mismatch is a significant limitation, as there are established methods available to achieve a proper unloaded configuration, as, e.g., in

      Marx et al. Journal of Computational Physics 463 (2022): 111266. https://doi.org/10.1016/j.jcp.2022.111266, and

      Regazzoni et al. Journal of Computational Physics 457 (2022): 111083. https://doi.org/10.1016/j.jcp.2022.111083.

      As our study aimed at creating a framework for calibration and validation in data-scarce scenarios such as it is often the case in the clinical context, using a compilation of multi-modal data from difference sources, rather than a specific method of personalisation, we did not feel it appropriate to invest significant energy to identify a patient-specific resting geometry, but rather felt that it was important for the resting geometry to fall within population values in terms of diastasis volume. We have clarified this issue in the manuscript and softened claims to Digital Twins in this study. The limitation has been addressed in the updated manuscript, and future work could further address this point.

      (6) Calibration procedure:

      There are apparent flaws in the calibration procedure, or it is not described in sufficient detail.

      We thank the reviewer for raising this point and we have substantially revised the calibration description in the manuscript to clarify the rationale behind each step.

      (a) Step 1: "Sample..." Why? kws and Cal50 are not the most significant parameters in the sensitivity analysis. Kct is a penalty parameter dependent on the numerical framework as described above; "ejection pressure threshold" was never mentioned, is it "P ejection LV" in Table 2? Aiming just for the highest LVEF might neglect non-physiological responses to parameter changes.

      While kws and Cal50 are not the single most significant parameters for LVEF in isolation, they were grouped in Step 1 because they affect both LVEF and peak systolic pressure simultaneously through cross-bridge cycling rate and residual active tension, making it necessary to sample them jointly rather than sequentially. Kct was included because myocardial compressibility affects wall thickening and therefore stroke volume. The ejection pressure threshold is P_ejection_LV in Table 1 and has now been described explicitly in the methods section. Regarding the concern about non-physiological responses: the action potential duration and active tension were monitored throughout calibration and verified to remain within physiological ranges, as now noted in the manuscript.

      (b) Step 2: As systolic pressure is directly dependent on arterial resistance for a 2-element Windkessel model, a uniform sampling approach might not be the best choice here.

      We acknowledge that uniform sampling may not be the most efficient approach for Step 2. However, since arterial resistance influences not only peak systolic pressure but also stroke volume and therefore LVEF, a more targeted approach focusing solely on pressure matching could compromise the LVEF achieved in previous steps. Uniform sampling allowed us to select the value that best balanced both quantities simultaneously.

      (c) Step 3: The authors mention in line 527: "A four-fold increase in GCaL caused an eight-fold increase in cellular active tension peak". An increase in active tension peak results in higher LVEF. So this step is likely to yield the upper boundary of the GCaL interval.

      The reviewer is correct that Step 3 tends to yield a high GCaL value. This was intentional — GCaL was used as a last resort to achieve physiological LVEF after Steps 1 and 2, since the model consistently undershot the target. The upper boundary of the sampled GCaL interval corresponds to a two-fold increase, which remains within the physiological variability bounds applied in previous studies. The resulting action potential duration was verified to remain within physiological ranges.

      (d) Step 4: Why again k_ws? It is not the most significant parameter in the SA.

      kws was resampled in Step 4 not to increase LVEF further, but to specifically target peak ejection rate and dP/dtmax, which were not adequately matched after Step 3. kws is the dominant parameter affecting these ejection dynamics biomarkers in the sensitivity analysis. Resampling at this stage allowed fine-tuning of ejection dynamics while maintaining the LVEF achieved in previous steps.

      (e) Step 5: As far as I can tell, the "diastolic volume change parameter" was mentioned the first time here.

      The diastolic volume change parameter C_pLAV has now been described in the methods section in the Phase 5 passive filling description, where it appears as the inverse of the penalty term controlling the rate of return to diastasis volume in the left ventricle.

      The whole calibration procedure seems to aim for the highest LVEF, and final values of the calibration parameters are not given.

      We note that the calibration procedure does not aim solely for the highest LVEF. As described above, the sequential strategy targets multiple quantities of interest in order of clinical importance: LVEF, peak systolic pressure, peak ejection rate, and peak filling rate, with each step designed to improve a specific subset of biomarkers without compromising those already matched. The final calibrated parameter values are reported in Figure 1F of the revised manuscript.

      (7) Novel features in this paper are actually scarce. A way more advanced calibration strategy with a whole heart model, emulators, and also the ToRORd-Land model was already presented in the study by Strocchi et al. [29]. The calibration to ECGs was presented by some of the same authors in Camps et al. [15], and the analysis of cellular effects was already published in several studies by the same group and in other publications, e.g., by the groups of Severi et al.

      The systematic compilation of credibility criteria spanning ECG morphology, pressure-volume characteristics, strain and displacement represents a novel contribution in itself, providing the field with a reusable evaluation framework. Furthermore, the present study is designed to yield mechanistic insight into how parameters at different scales influence both electrical and mechanical outputs simultaneously. This goal was not tackled in previous publications, which covered individual components, including ECG calibration in Camps et al. [15] and global sensitivity analysis with whole-heart models in Strocchi et al. [29] with no ECG consideration.

      Thus, the work by Camps et al. on ECG calibration was purely electrophysiological and did not investigate the influence of mechanical or haemodynamic parameters on ECG morphology in a fully coupled electromechanical framework. While the effect of mechanical parameters on ECG has been explored by others (e.g. Favino, 2016), this has not previously been examined alongside the relative importance of cellular, mechanical and haemodynamic parameters on pressure-volume characteristics within a single coupled framework. While Strocchi et al. present an emulation strategy, they did not address ECG biomarkers. This distinction is now stated explicitly in the introduction, where we position the present study relative to Camps et al. and Strocchi et al.

      (8) How could the calcium sensitivity Cal50 have such a drastic effect on diastolic function, i.e., filling and end-diastolic volume? As far as I understand from the description, the simulation starts with Phase 0 (loading), Phase 1 (atrial filling), and then in Phase 2, electrical activation ensues and active contraction develops, see also the section starting in line 165. Based on this description, I would expect the end-diastolic volumes to be identical across all Cal50 values. Or are the PV loops shown actually limit cycles established over simulations with multiple beats? This point wasn't explicitly clarified in the manuscript.

      Calcium sensitivity (Cal50) affects not only systolic active tension development but also diastolic residual active tension, i.e. the degree to which the muscle remains partially activated at end diastole. Higher Cal50 values increase this residual tone, effectively stiffening the myocardium during diastolic filling and reducing end-diastolic volume. This mechanism is well established as a contributor to diastolic dysfunction in heart failure [88]. We have clarified this in the manuscript and also clarified that the PV loops shown are single-beat simulations, not limit cycles, with the end-diastolic volume determined by the prescribed filling pressure alongside the passive and residual active stiffness of the myocardium.

      Minor concerns:

      (9) Line 29: The values provided: LVEF of 51%, EDV of 110 mL, and ESV of 50 mL are inconsistent. If these values are all related to the LV, the calculated LVEF should be approximately 54.55%, not 51%.

      The values quoted in the original abstract were rounded approximation, this has been corrected to report EDV=105 mL and ESV=51 mL, which are consistent with the simulated LVEF of 51%.

      (10) "Electromechanical cardiac Digital Twins have had broad applicability ..."

      Many of the cited works here are not true "Digital Twins" but rather static, non-patient-specific models of cardiac electromechanics. In some cases, the geometry may be derived from patient data, but this alone does not qualify the model as a digital twin.

      This sentence in the introduction has been rephrased as ‘Electromechanical cardiac models have had broad applicability...’. Furthermore, as stated earlier, we have removed explicit claims of Digital Twin from the paper while retaining the fact that this study provides a significant step towards rigorous credibility assessment of the high-fidelity electromechanical models that make Digital Twin construction possible.

      (11) While in the abstract and the conclusion, the authors mention "uncertainty quantification", it is mostly a sensitivity analysis that was performed in the paper.

      We have updated the text to say ‘sensitivity analysis’ where appropriate in the abstract, results, and conclusion, and replaced ‘uncertainty ranges’ with ‘variability ranges’ throughout. However, since the sensitivity analyses were performed over biologically informed ranges derived from population variability in the literature, the results are informative about how uncertainty in model inputs propagates to uncertainty in simulated biomarkers. We have therefore retained the framing of sensitivity analysis as a first step towards uncertainty quantification in the abstract and conclusion, and have added a clarifying sentence to the methods to this effect.

      (12) Line 98: As far as I can tell, the conduction velocity assigned to the endocardial surface - intended to mimic the Purkinje fiber network - is never specified. In the section beginning at line 294, only the transmural conduction velocities are reported.

      The endocardial conduction velocity has been specified in the methods section: Purkinje-myocardial junctions were modelled using a fast endocardial activation layer with isotropic conduction velocity of 300 cm/s.

      (13) Line 198, Table1:

      (a) "21/02/2025 11:09:00 AM" on two occasions is maybe not intended

      This has been removed.

      (b) For easing up comparisons, units should be consistent between the initial value and the literature ranges, e.g., PV control parameters, heart rate.

      Units have been made consistent between the initial values and literature ranges throughout Table 1.

      (14) Line 232: "... have already been used to calibrate and validation ...".

      This has been corrected.

      (15) Line 279, Table 2: This table of variability ranges is not entirely clear and could be improved:

      Table 2 has been combined with Table 1 such that the variability ranges sit next to the literature values, for ease of comparison.

      (a) "21/02/2025 11:09:00 AM" is maybe not intended.

      This has been removed.

      (b) use of units should be improved; sometimes it's given in the first column, sometimes in the second column (arterial resistance, compliance), then for k_epi it should be either kPa or kPa/cm.

      Units have been made consistent and the units for k_epi has been added in Table 1.

      (c) units should also be consistent throughout the paper, e.g. in Figure 1 E arterial resistance is Barye.ms/mL while in Table 2 it is mmHg.ms/mL.

      Barye has been removed and replaced by corresponding kPa values throughout the manuscript. This was in the original manuscript since the Alya simulation software were in units of cm, s, g, Barye.

      (d) it is also not clear how variability ranges were chosen; e.g., for arterial compliance,e literature ranges are 0.2-2.73 while the chosen range is [0.1,0.2].

      The previous ranges were chosen to achieve better LVEF. We have now updated the variability ranges to be purely based on literature values and updated the sensitivity analysis results. The ranges are now presented in Table 1 alongside the literature values for ease of comparison.

      (e) For Kct, the initial value in Table 1 is 5000kPa, the literature values are between 10 and 3333, and then the variability range is [10,500]? I guess there is a typo in one of these values.

      This has been corrected in the new Table 1.

      (f) Table 1 and 2 are in parts redundant.

      Table 1 and 2 have been combined into a single new Table 1.

      (15) Line 290: It should be uvc_l for the longitudinal coordinate.

      This has been corrected.

      (16) Line 388, Table 3, regarding values for pressure volume from reference [49]:

      (a) the number of participants is 800, including males and females; not only females, see also Table 12 https://jcmr-online.biomedcentral.com/articles/10.1186/s12968-017-0327-9/tables/12

      (b) why using female values here while having mixed sex for most of the others? Because the model is female?

      The reviewer is correct that reference [49] reports values from a mixed-sex cohort of approximately 800 participants. We used the female-specific values from Table 12 of that reference because the biventricular mesh used in this study was derived from a female subject, making sex-matched reference values the most appropriate comparison. This has been clarified in the manuscript.

      (17) Figure 5: What is Jup; why did you choose 0.93 x Jup as reference? Also in Figure 4, why did you use 0.93 x GCal as a reference?

      J_up refers to the SERCA<sup2+</sup> reuptake current, which has been relabelled as SERCA throughout the manuscript for consistency. The reference value of 0.93× was used because the sensitivity analysis sampled parameters uniformly between 50% and 200% of baseline using a fixed number of samples, and no sample fell exactly at 1.0×. The closest sampled value was 0.93×, which was therefore used as the reference. This has been clarified in the figure caption.

      (18) Line 377: The link to the GitHub repository does not work.

      This link has now been made publicly available.

      (19) Line 397, Table 3: for the sake of completeness, all abbreviations should be included: e.g., SVL, ESP, EDV, ESV are not included.

      This has been written out in full in the new Table 2.

      (20) Tick marks in many figures are not readable, e.g., Figure 5 and all the Figures in the appendix.

      Tick mark sizes and line widths have been increased across Figures 4, 5, and all appendix figures. The figures have been replotted and updated in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor Comments:

      (1) The provided GitHub link https://github.com/jennyhelyanwe/Alya_input_setup/ does not work, potentially because the repository is private. It would be nice to see the repository during the review.

      This link has now been made publicly available.

      (2) Table 3: Can you include the simulation outputs obtained for validation (with an error indication)? This would summarize the validation that's currently spread out over the results section.

      A new Table 3 has been added to the manuscript under the validation section, summarising the simulated values for all deformation and strain biomarkers alongside their reference ranges. The calibration and validate datasets are now reported separately in Tables 2 and 3, respectively.

      (3) Figure 1: Add axis labels to all plots.

      Axis labels have been added to all subplots in Figure 1 in the revised manuscript. Simulated pseudo-ECG amplitudes are normalised and therefore dimensionless.

      (4) Figure 2: Simulated and in vivo strains with exactly the same axes (size, range, ticks) and add grid lines to enable a comparison. Add the mean values of each in the other plot.

      The revised Figure 2 now includes the median in vivo strain values from Moulin et al. overlaid as a red dashed reference line on the simulation panels, enabling direct visual comparison. The simulated mean could not be overlaid on the in vivo panels as the original Moulin et al. figure data are not publicly available for replotting. Exact axis matching was not applied as this would cause some simulated curves to fall outside the visible range, obscuring the model behaviour.

      (5) Figure 3: The thickness (relative importance of the connections) is impossible to see in this plot. Instead of having gray background connections, remove them entirely below a certain threshold. Make the differences in thickness more pronounced or introduce a continuous color scale for the magnitude of the positive or negative correlation. Alternatively, you could rank the parameters from least to most important in each subfigure A-D and/or provide some numeric values.

      Figure 3 has been updated. All non-significant connections (|r| < 0.6 or p > 0.05) have been removed entirely, and gray lines have been removed in each subfigure, making the significant relationships clearer. A continuous blue-to-red colour scale has been applied to indicate the direction of correlation (blue: negative, red: positive), with line thickness proportional to the magnitude of the r-value.

      (6) Figure 3 and Table 2: Why were material parameters b, bf, bs, and bfs omitted from this study (but included a, af, as, and afs)?

      The b parameters (b, bf, bs, bfs) appear in the exponent of the Holzapfel-Ogden constitutive law and are strongly coupled to the a parameters (a, af, as, afs), which carry units of kPa. In practice, the b parameters can only be reliably identified from ex vivo multiaxial stretch experiments, whereas the a parameters can be estimated from clinical imaging data. Since our study focuses on calibration and validation in a clinical data setting, we included only the a parameters in the sensitivity analysis, consistent with previous personalisation studies.

      (7) Figure 4: What do the dotted lines represent?

      The dotted lines in Figure 4E highlight the increased longitudinal shortening with increasing GCaL, showing the basal plane moving towards the apex while the apical position remains unchanged due to the pericardial constraint. This has been clarified in the figure caption.

      (8) Figures 4, 5, A2-45: Can you use a continuous color scale (e.g., from blue to red) for low to high parameter uncertainty?

      A continuous blue-to-red colour scale has been applied to Figures 4, 5, and all appendix figures A2–A6, where blue indicates the lowest parameter value and red indicates the highest. A colour bar has been added to each figure for reference.

    1. eLife Assessment

      The study presents important findings from a very rich EEG-fMRI dataset including 107 participants, which was collected during nocturnal naps. Using overall solid methods, the authors link activity in memory related brain regions (e.g., hippocampus, thalamus, and medial prefrontal cortex), as well as their functional connectivity, to the occurrence of canonical sleep rhythms (spindles and slow oscillations) during non-rapid eye movement (NREM) sleep. This work will be of broad interest to researchers studying sleep, memory, and related domains.

    2. Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Comments on latest version:

      The authors have substantially revised their manuscript and sufficiently addressed all of my previous concerns. I have no further comments.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Thank you for this encouraging assessment. We appreciate your recognition of the value of the dataset and of the questions it allows us to address. Below, we respond to each of your points directly and revise the manuscript accordingly.

      Comments on revisions:

      Re 1: The revised introduction now cites a couple of papers but discusses them only very superficially, lumping together several studies with very different key results. This is still not very informative for the reader and does not sufficiently acknowledge previously published work. Here are two examples to illustrate this:

      (a) "These studies have generally reported that slow oscillations are associated with widespread cortical and subcortical BOLD changes, whereas spindles elicit activation in the thalamus, as well as in several cortical and paralimbic regions." Several studies even showed e.g., a clear activation of the hippocampus and parahippocampal gyrus associated with spindles, not just the thalamus

      Thank you for this comment. We agree that our previous sentence was too broad and did not sufficiently reflect the range of findings in the sleep literature. We have therefore rewritten the Introduction to state explicitly that spindle-related BOLD changes have been reported not only in the thalamus, but also in cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).

      Introduction, Page 3-4, Lines 58-62

      “Consistent with this view, prior human EEG-fMRI studies have reported spindle-related activation not only in the thalamus, but also in the hippocampus and adjacent parahippocampal gyrus (Bergmann et al., 2012; Schabus et al., 2007). Spindle-related activity has also been linked to striatal engagement, suggesting a broader network that may support memory-related processing during sleep (Fogel et al., 2017).”

      Introduction, Page 4, Lines 71-78

      “Previous EEG-fMRI studies on sleep have examined both global sleep characteristics (Hale et al., 2016; Moehlman et al., 2019) and the neural correlates of specific waves, including slow oscillations and spindles. These studies have generally shown that slow oscillations are associated with widespread cortical and subcortical BOLD changes (Czisch et al., 2009; Ilhan-Bayrakcı et al., 2022; Picchioni et al., 2011), whereas spindles have been linked not only to thalamic activation but also to cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      Introduction, Page 5, Lines 103-106

      “This coupling was associated with increased activation in both the thalamus and hippocampus, with functional connectivity patterns suggesting thalamic coordination of hippocampal-cortical communication, in line with prior EEG-fMRI studies of spindle-related activity (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      (b) "Although these findings provide valuable insights into the BOLD correlates of sleep rhythms, they often do not employ sophisticated temporal modeling (Huang et al., 2024) [, ...]." - previous studies have used e.g., spindle event-related regressors with individual spindle amplitudes as parametric modulators, first and second order derivatives of the HRF function, as well as PPI connectivity analyses, which I would consider rather sophisticated temporal modelling.

      We agree that several previous studies have already employed sophisticated modelling approaches, including parametric modulation, HRF derivatives, and PPI analyses (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Our intention was not to suggest that such methods are absent from the literature.

      Rather, we aimed to highlight that most prior work has focused on modelling individual SO or spindle events, whereas explicit modelling of their temporal interaction (e.g., SO-spindle coupling) has been less commonly addressed. We have revised the sentence to clarify this point more precisely.

      Introduction, Page 4 Lines 78-82

      “Although these findings provide important insight into the BOLD correlates of sleep rhythms, most previous studies have focused on individual oscillatory events rather than explicitly modelling their temporal interaction (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Only a few recent studies have begun to examine coupling between rhythms directly, for example Huang et al. (2024).”

      Re 4+9: The short overall recordings in some subjects on the one hand and the large number of spindles and SOs detected in N1 sleep stages are still highly concerning, in fact even more so, now that the actual numbers have been provided in the Supplementary Tables. Either the sleep staging or the detection of SO and spindle events must be incorrect. I understand that for specific EEG analysis and fMRI modelling purposes sometimes slightly different thresholds are used as compared to clinical sleep staging, but several parameters here are alarmingly off.

      (a) Given that proper NREM sleep (N2+N3) is the relevant stage for the analyses conducted in this paper, some of the N2+N3 durations are very short (eg 7-8 min) while those subjects' results have the same impact on the group level analyses as those with >100 min of N2+N3. Either subjects with very little relevant data (not overall recording time but N2+N3 time) should be excluded or weighting subject data for the group analyses according to the amount od contributed data should be done.

      Thank you for the suggestion. It is true that participants with very little N2/3 sleep could contribute noisier subject-level estimates to the group analysis. We therefore checked this directly. Only three participants contributed less than 10 min of N2/3 sleep, and excluding them did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant after exclusion, t<sub>(103)</sub> =2.50, p = 0.0071, compared with t<sub>(106)</sub> = 2.50, p = 0.0070 in the full sample. We have added this control analysis to the Results so that the robustness of the group findings is explicit in the manuscript.

      Results, Page 11-12, Lines 238-250

      To ensure the results were not driven by individual differences or parameter selection, we conducted a series of control analyses. First, we excluded participants with less than 10 minutes of N2/3 sleep. Only three participants met this criterion, and their exclusion did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant (t<sub>(103)</sub> = 2.50, p = 0.0071), comparable to the full sample (t<sub>(106)</sub> = 2.50, p = 0.0070). Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st-80th percentile; Fig. S6). Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).

      (b) The authors argue that the SO and spindle detection algorithms are valid since widely used and that they were developed for N2+N3 stages, which is why they will also detect events in other stages: "While, because the detection methods for SO and spindle are based on percentiles, this method will always detect a certain number of events when used for other stages (N1 and REM) sleep data, but the differences between these events and those detected in stage N23 remain unclear." I do agree that with very liberal thresholds, also SO and spindle vents may be detected in other stages, but it shouldn't be that many. If the percentiles of amplitude thresholds were defined based on properly scored N2+N3 stages only, very few events should be detected (erroneously!) in N1, as the occurrence of K-complexes (isolated SOs) and spindles per definition makes it N2, and during REM sleep only very few spindles and SOs are allowed to occur, without scoring it NREM instead. For the first subject (just as example, but with similar numbers for the rest of the sample), reveals as many as 60 SOs and 31 spindles within 8 min of N1 sleep (Table S2) as well as 13 SOs and 7 spindles within 2 min of REM sleep (Table S4). These numbers are completely unrealistic and question the correctness of the sleep staging as well as the physiological relevance of the EEG graphoelements identified as SO and spindles. It also completely undermines the interpretability of the respective event regressors for the fMRI analyses.

      (c) Likely, given the large numbers of coupled SO-spindle events and the apparently very low amplitude criteria for event identification, also the number of SO-spindle couplings is likely severely overestimated.

      We thank the reviewer for raising this important point. We agree with you that the original stage-wise percentile thresholding could inflate the apparent number of SOs and spindles outside N2/3 sleep. In the original analysis, the thresholds were estimated separately within each sleep stage. As you point out, this procedure can force the detector to label a relatively large number of events in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 SOs or spindles. We have therefore revised the detection procedure. Following your concern and Reviewer 2’s suggestion, the SO and spindle thresholds are now defined only from N2/3 sleep within each participant, where SOs and spindles are most abundant and physiologically expected to occur. These fixed N2/3-derived thresholds were then applied unchanged to N1 and REM for descriptive reporting. This avoids the artificial normalisation of event detection across sleep stages that can arise when each stage has its own percentile threshold. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We agree with you that the remaining detections in N1 and REM should not be treated as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We therefore report them only as descriptive detector outputs obtained under a fixed N2/3-derived threshold (see Table S2, S4 in the revised manuscript). We do not use them to support any physiological claim about SO-spindle coupling in N1 or REM.

      This point is also important for the fMRI analyses. You are right that inflated N1 or REM detections would undermine the interpretability of event regressors if those detections entered the EEG-informed fMRI models. They did not. All EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep, where SOs, spindles, and their coupling are physiologically expected and where the detection thresholds were defined. Thus, the central fMRI event regressors were based only on N2/3 events, not on detections from N1 or REM.

      We also agree with you that the absolute number of detected SO-spindle couplings depends on the chosen detection threshold. For this reason, we tested whether the main EEG-fMRI result depended on the specific detector setting. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We therefore do not argue that the detector provides a uniquely correct absolute count of SOs, spindles, or coupling events in every sleep stage. Our conclusion is more specific. The main N2/3 EEG-fMRI finding is robust across a reasonable range of SO detection thresholds, detections in N1 and REM are reported only descriptively, and the physiological interpretation of SO-spindle coupling is restricted to N2/3 sleep.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Re 10: The rationale for using a lateralized frontal electrode (F3) for both SO (should have been at least bilateral or central) and spindle detection (should have been a centro-parietal electrode) is not convincing. Other EEG-fMRI spindle or SO papers have used a number of frontal (SO) or centro-parietal (spindles) electrodes averaged or even approaches including all EEG electrodes. Searching events with low thresholds at suboptimal recording sites does not dot this highly valuable dataset justice.

      We thank the reviewer for this important comment. We agree that this choice is more sensitive to frontal SOs than to the centro-parietal fast spindle component. Our choice of F3 was driven by the practical constraints of prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to FCz can have reduced signal contrast relative to the reference, and electrodes near the vertex are also more vulnerable to prolonged pressure against the MRI head coil when participants sleep supine for several hours. In this setting, frontal electrodes provided more stable signal quality across the recording. Because the EEG events were used primarily as temporal markers for fMRI modelling, our priority was to obtain reliable event timing during N2/3 sleep rather than to estimate the full scalp topography of SOs and spindles.

      We also agree with your concern that this valuable dataset would ideally be analysed with multichannel detection strategies. To test whether the main result depended on the single lateralised F3 site, we repeated the main EEG-informed fMRI analysis using Fz, a midline frontal electrode. The result was unchanged. Hippocampal activation during SO-spindle coupling remained significant when events were detected from Fz, t<sub>(106)</sub> = 2.47, p = 0.0076, closely matching the original F3-based result, t<sub>(106)</sub> = 2.50, p = 0.0070. This control analysis does not remove the limitation that centro-parietal fast spindles may be underrepresented, and we do not claim that it does. It does show, however, that the main hippocampal fMRI finding is not driven by idiosyncratic detections from one lateralised frontal electrode. We have made this clearer in the revised manuscript.

      Finally, your concern about low thresholds is also important. As described in our response above, the revised analysis now defines SO and spindle thresholds from N2/3 sleep and applies these thresholds uniformly for descriptive comparisons across stages. We also tested the robustness of the hippocampal fMRI result across SO detection thresholds, and the effect remained significant across the 71st to 80th percentile range. We have therefore narrowed the interpretation in the revised manuscript. The main EEG-fMRI result reflects BOLD activity associated with frontal-channel-detected SO-spindle coupling during N2/3 sleep, rather than a full multichannel characterisation of all SO and spindle topographies.

      Results, Page 12, Lines 247-250

      “Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).”

      Discussion, Page 18, Lines 380-394

      “Third, sleep oscillation detection was based on a single frontal electrode. This choice improved signal stability and event timing in the prolonged simultaneous EEG-fMRI setting, but it did not exploit the full multichannel EEG information and cannot characterise the full spatial distribution of SOs and spindles. In particular, F3-based detection may be more sensitive to frontal SOs and frontal sigma activity than to the centro-parietal fast spindle component. We therefore interpret the EEG-informed fMRI results as reflecting BOLD activity associated with frontal-channel-detected SOs, spindles, and their coupling during N2/3 sleep. Future studies using multichannel or source-informed detection strategies, with separate treatment of slow and fast spindles, will be better suited to capture the spatial dynamics of these sleep oscillations. Fourth, the use of large anatomical ROIs may mask subregional contributions of specific thalamic nuclei or hippocampal subfields. Finally, without a memory task, we cannot establish a direct behavioral link between sleep-rhythm-locked activation and memory consolidation. Future studies combining ultra-high-field fMRI or iEEG with cognitive tasks, as well as multichannel or source-informed detection strategies that separately characterize slow and fast spindles, will be better suited to refine our understanding of subregional network dynamics and the functional significance of sleep oscillations.”

      Methods, Page 25, Lines 556-566

      “It is worth noting that the primary aim of EEG rhythm detection was to identify reliable event times for EEG-informed fMRI modelling. Detection was performed on the F3 electrode because this channel provided stable signal quality during prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to this reference, and electrodes near the vertex that were in prolonged contact with the head coil during supine sleep, were more susceptible to reduced signal contrast, impedance drift, and pressure-related degradation of electrode-scalp contact. We therefore used F3 as a pragmatic choice to maximize reliable event timing in N2/3 sleep. This choice was not intended to characterise the full scalp topography of SOs or spindles, and it may underrepresent the centro-parietal fast spindle component. As a sensitivity analysis, we repeated the main EEG-informed fMRI GLM using Fz, a midline frontal electrode, with the same detection and modelling procedure.”

      Re 7: It is not clear to me why/how larger voxels would reduce susceptibility-related distortions and partial volume effects. Usually, the opposite is true. This should be elaborated.

      What we meant was that we chose a relatively large voxel size to preserve signal-to-noise ratio and whole-brain coverage within a feasible repetition time for a long overnight EEG-fMRI protocol. This choice is useful for maintaining BOLD sensitivity in sleep recordings, where head motion, physiological noise, and participant comfort are major practical constraints. We agree that it may not be accurate to describe it as reducing susceptibility-related distortion or partial volume effects.

      We have rewritten the Methods to state this trade-off directly. The voxel size of 3.5 × 3.5 × 4.2 mm<sup>3</sup> allowed whole-brain coverage with a TR of 2000 ms, which was important for modelling sleep-rhythm-related BOLD responses across the whole brain during prolonged nocturnal recordings. A smaller voxel size would have improved spatial specificity, but would also have required either a longer TR, reduced brain coverage, or lower SNR, none of which would have been ideal for the present EEG-fMRI sleep design. We now explicitly acknowledge the cost of this choice.

      Methods, Page 21 Lines 453-463

      “For the functional scans, whole-brain images were acquired using a T2*-weighted gradient echo-planar imaging (EPI) sequence sensitive to the BOLD contrast. The sequence parameters were as follows: 33 slices in interleaved ascending order, TR = 2000 ms, TE = 30 ms, voxel size = 3.5 × 3.5 × 4.2 mm<sup>3</sup>, FA = 90°, matrix = 64 × 64, gap = 0.7 mm. A relatively large voxel size was chosen to preserve signal-to-noise ratio while maintaining whole-brain coverage within a feasible repetition time. This compromise was important for the prolonged overnight EEG-fMRI sleep protocol, where head motion, physiological noise, participant comfort, and sustained acquisition stability are substantial practical constraints (Bodurka et al., 2007; Laufs et al., 2008). A smaller voxel size would have improved spatial specificity, but would have required either a longer repetition time, reduced brain coverage, or lower signal-to-noise ratio.”

      Reviewer #2 (Public review):

      In this study, Wang and colleagues aimed to explore brain-wide activation patterns associated with NREM sleep oscillations, including slow oscillations (SOs), spindles, and SO-spindle coupling events. Their findings reveal that SO-spindle events corresponded with increased activation in both the thalamus and hippocampus. Additionally, they observed that SO-spindle coupling was linked to heightened functional connectivity from the hippocampus to the thalamus, and from the thalamus to the medial prefrontal cortex-three key regions involved in memory consolidation and episodic memory processes.

      This study's findings are timely and highly relevant to the field. The authors' extensive data collection, involving 107 participants sleeping in an fMRI while undergoing simultaneous EEG recording, deserves special recognition. If shared, this unique dataset could lead to further valuable insights.

      Thank you for this encouraging assessment. We appreciate your recognition of the effort involved in collecting this simultaneous EEG-fMRI sleep dataset. Below, we respond directly to your remaining concern.

      Comments on revisions:

      The authors' efforts in revising the manuscript and addressing the reviewers' comments are certainly commendable. However, I remain concerned about potential issues in detecting sleep-related oscillations (SOs, spindles, and consequently coupled SO-spindle events), which may arise due to suboptimal parameter selection or inaccurate sleep staging, potentially impacting all subsequent analyses.

      A review of Supplementary Tables 1-4 reveals an unusually high number of detected SOs and spindles during sleep stage N1 and REM sleep. While the authors correctly note that a percentile-based detection approach will always identify a certain number of events across sleep stages, the particularly high counts in N1 and REM are concerning. To mitigate the limitations of this method, the authors could have performed event detection independently of sleep stages (i.e., across the entire dataset for each participant) and subsequently assigned the detected events to the corresponding sleep stages. If the event counts in N1 and REM remained disproportionately high, this would indicate a fundamental issue with the detection procedure.

      In the previous version, thresholds were estimated separately within each sleep stage. As you point out, this can force the detector to identify a relatively large number of SOs and spindles in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 events.

      We have therefore revised the detection procedure so that event detection is no longer based on separate stage-wise thresholds. Following the logic of your suggestion, we first defined a fixed threshold for each participant and then assigned the detected events to their corresponding sleep stages afterwards. We used N2/3 sleep to define the SO and spindle thresholds because this is the stage in which these events are physiologically expected and most reliably observed. These same N2/3-derived thresholds were then applied unchanged to N1 and REM. This avoids the circularity of forcing a percentile-defined number of detections within each sleep stage.

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We also agree with you that the remaining detections in N1 and REM should not be interpreted as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We now state this explicitly in the manuscript. They are reported only as descriptive detector outputs under the fixed N2/3-derived criterion. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      We also would like to clarify our sleep staging procedure. The sleep staging was first performed using an established automated algorithm, YASA toolkit (Vallat & Walker, 2021), and then manually reviewed by two sleep experts. More importantly for the central results, all EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep. Thus, N1 and REM detections did not enter the event regressors used for the main fMRI analyses and do not affect the interpretation of the hippocampal or thalamic findings.

      Finally, we agree that the absolute number of SO-spindle coupling events depends on the detection threshold. We therefore tested whether the main fMRI result depended on the specific SO threshold. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We have made this clearer in the revised manuscript.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Reviewer #3 (Public review):

      Summary:

      Wang et al., examined the brain activity patterns during sleep, especially when locked to those canonical sleep rhythms such as SO, spindle, and their coupling. Analyzing data from a large sample, the authors found significant coupling between spindles and SOs, particularly during the up-state of the SO. Moreover, the authors examined the patterns of whole-brain activity locked to these sleep rhythms. The authors next investigated the functional connectivity analyses, and found enhanced connectivity between the hippocampus and the thalamus and the medial PFC. These results reinforced the theoretical model of sleep-dependent memory consolidation, such that SO-spindle coupling is conducive for systems-level memory reactivation and consolidation.

      Strengths:

      There are obvious strengths in this work, including the large sample size, state-of-the-art neuroimaging and neural oscillation analyses, and the richness of results. The results now inform hemodynamic neural activity that coincided with SO-spindle couplings.

      Weaknesses:

      My earlier comments were about the inability to make inferences on memory given the lack of memory tasks, and the weakness in using the open-ended cognitive state decoding.

      Comments on revisions:

      The current revision has addressed these major concerns. The authors expanded discussions regarding the theoretical implications of the work in a more nuanced manner.

      Thank you for taking the time to re-evaluate the manuscript. We are pleased that the revised Discussion now reads as more nuanced, especially in relation to the limits of the memory-related interpretation. Your earlier comments helped us sharpen both the claims and the framing, and we are grateful for that.

      References:

      Bergmann, T. O., Mölle, M., Diedrichs, J., Born, J., & Siebner, H. R. (2012). Sleep spindle-related reactivation of category-specific cortical regions after learning face-scene associations. Neuroimage, 59(3), 2733-2742.

      Bodurka, J., Ye, F., Petridou, N., Murphy, K., & Bandettini, P. A. (2007). Mapping the MRI voxel volume in which thermal noise matches physiological noise—implications for fMRI. Neuroimage, 34(2), 542-549.

      Caporro, M., Haneef, Z., Yeh, H. J., Lenartowicz, A., Buttinelli, C., Parvizi, J., & Stern, J. M. (2012). Functional MRI of sleep spindles and K-complexes. Clinical neurophysiology, 123(2), 303-309.

      Czisch, M., Wehrle, R., Stiegler, A., Peters, H., Andrade, K., Holsboer, F., & Sämann, P. G. (2009). Acoustic oddball during NREM sleep: a combined EEG/fMRI study. PloS one, 4(8), e6749.

      Fogel, S., Albouy, G., King, B. R., Lungu, O., Vien, C., Bore, A., Pinsard, B., Benali, H., Carrier, J., & Doyon, J. (2017). Reactivation or transformation? Motor memory consolidation associated with cerebral activation time-locked to sleep spindles. PloS one, 12(4), e0174755.

      Hahn, M. A., Heib, D., Schabus, M., Hoedlmoser, K., & Helfrich, R. F. (2020). Slow oscillation-spindle coupling predicts enhanced memory formation from childhood to adolescence. Elife, 9, e53730.

      Hale, J. R., White, T. P., Mayhew, S. D., Wilson, R. S., Rollings, D. T., Khalsa, S., Arvanitis, T. N., & Bagshaw, A. P. (2016). Altered thalamocortical and intra-thalamic functional connectivity during light sleep compared with wake. Neuroimage, 125, 657-667.

      Helfrich, R. F., Lendner, J. D., Mander, B. A., Guillen, H., Paff, M., Mnatsakanyan, L., Vadera, S., Walker, M. P., Lin, J. J., & Knight, R. T. (2019). Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans. Nature Communications, 10(1), 3572.

      Helfrich, R. F., Mander, B. A., Jagust, W. J., Knight, R. T., & Walker, M. P. (2018). Old brains come uncoupled in sleep: slow wave-spindle synchrony, brain atrophy, and forgetting. Neuron, 97(1), 221-230. e224.

      Huang, Q., Xiao, Z., Yu, Q., Luo, Y., Xu, J., Qu, Y., Dolan, R., Behrens, T., & Liu, Y. (2024). Replay-triggered brain-wide activation in humans. Nature Communications, 15(1), 7185.

      Ilhan-Bayrakcı, M., Cabral-Calderin, Y., Bergmann, T. O., Tüscher, O., & Stroh, A. (2022). Individual slow wave events give rise to macroscopic fMRI signatures and drive the strength of the BOLD signal in human resting-state EEG-fMRI recordings. Cerebral Cortex, 32(21), 4782-4796.

      Laufs, H., Daunizeau, J., Carmichael, D. W., & Kleinschmidt, A. (2008). Recent advances in recording electrophysiological data simultaneously with magnetic resonance imaging. Neuroimage, 40(2), 515-528.

      Maingret, N., Girardeau, G., Todorova, R., Goutierre, M., & Zugaro, M. (2016). Hippocampo-cortical coupling mediates memory consolidation during sleep. Nature Neuroscience, 19(7), 959-964.

      Moehlman, T. M., de Zwart, J. A., Chappel-Farley, M. G., Liu, X., McClain, I. B., Chang, C., Mandelkow, H., Özbay, P. S., Johnson, N. L., & Bieber, R. E. (2019). All-night functional magnetic resonance imaging sleep studies. Journal of neuroscience methods, 316, 83-98.

      Ngo, H.-V., Fell, J., & Staresina, B. (2020). Sleep spindles mediate hippocampal-neocortical coupling during long-duration ripples. Elife, 9, e57011.

      Ngo, H. V., Martinetz, T., Born, J., & Molle, M. (2013). Auditory closed-loop stimulation of the sleep slow oscillation enhances memory. Neuron, 78(3), 545-553.

      Picchioni, D., Horovitz, S. G., Fukunaga, M., Carr, W. S., Meltzer, J. A., Balkin, T. J., Duyn, J. H., & Braun, A. R. (2011). Infraslow EEG oscillations organize large-scale cortical–subcortical interactions during sleep: a combined EEG/fMRI study. Brain research, 1374, 63-72.

      Schabus, M., Dang-Vu, T. T., Albouy, G., Balteau, E., Boly, M., Carrier, J., Darsaud, A., Degueldre, C., Desseilles, M., & Gais, S. (2007). Hemodynamic cerebral correlates of sleep spindles during human non-rapid eye movement sleep. Proceedings of the National Academy of Sciences, 104(32), 13164-13169.

      Schreiner, T., Kaufmann, E., Noachtar, S., Mehrkens, J.-H., & Staudigl, T. (2022). The human thalamus orchestrates neocortical oscillations during NREM sleep. Nature Communications, 13(1), 5231.

      Schreiner, T., Petzka, M., Staudigl, T., & Staresina, B. P. (2021). Endogenous memory reactivation during sleep in humans is clocked by slow oscillation-spindle complexes. Nature Communications, 12(1), 3112.

      Staresina, B. P., Bergmann, T. O., Bonnefond, M., van der Meij, R., Jensen, O., Deuker, L., Elger, C. E., Axmacher, N., & Fell, J. (2015). Hierarchical nesting of slow oscillations, spindles and ripples in the human hippocampus during sleep. Nature Neuroscience, 18(11), 1679-1686.

      Staresina, B. P., Niediek, J., Borger, V., Surges, R., & Mormann, F. (2023). How coupled slow oscillations, spindles and ripples coordinate neuronal processing and communication during human sleep. Nature Neuroscience, 1-9.

      Vallat, R., & Walker, M. P. (2021). An open-source, high-performance tool for automated sleep staging. Elife, 10.

    1. eLife Assessment

      By mining the Logan assemblage of the Sequence Read Archive to reveal substantial papillomavirus diversity, this valuable study establishes a framework for integrating viral discovery with host, geographic, and ecological metadata. The evidence for identifying novel papillomavirus sequences is convincing, supported by large-scale sequence searches, established L1-based typing criteria, phylogenetic analyses, protein annotation, and structural comparisons. The broader host-association and ecological conclusions are more tenuous due to heterogeneous public metadata and uneven sampling, and would benefit from clearer discussion of potential biases, contamination, endogenous viral elements, and alternative explanations. The work should be of interest to a broad range of colleagues in the areas of microbiome research, virology, and viral evolution.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript titled "Petabase-scale papillomavirus discovery" capitalizes on Logan assemblages to identify novel and known PVs.

      This is a brilliant use of Logan assemblages, and this paper highlights the use of this, plus also shows that one person's trash is another one's gold. Superb paper and I applaud the authors for starting with the PVs as easier to identify due to their set of genes coupled with the conserved L1 protein and associated typing for PVs (10% pairwise identity threshold for identification of new PV types).

      Strengths:

      This study highlights the hidden gems in public resources, especially if mined properly. Thanks to Logan assemblages, this is possible, and this manuscript highlights this with their data mining of papillomaviruses, identifying known and novel PVs in pangolins, lizards, fish, and white rhinos.

      Weaknesses:

      None identified

    3. Reviewer #2 (Public review):

      This study applies the Logan assemblage of the Sequence Read Archive to investigate papillomavirus diversity at a large scale. The authors combine sequence similarity searches, phylogenetic analyses, protein annotation, structural comparisons, and metadata integration to identify and characterize novel papillomavirus sequences. Beyond papillomavirus sequence discovery, the study develops a framework for associating viral sequences with host, geographic, and ecological metadata through the integration of both established and recently developed large-scale public sequence repositories. This framework may provide a foundation for similar investigations across other viral taxa.

      An important aspect of the work is the systematic processing and integration of metadata across a large number of sequencing libraries. The approaches developed to aggregate, curate, and interpret host and environmental information may be broadly applicable to future large-scale studies of other viral families. At the same time, interpretations based on metadata-derived ecological and host-association patterns should be considered in the context of the inherent limitations of public sequencing repositories, including uneven taxonomic representation, heterogeneous sampling strategies, laboratory-derived samples, and variable sequencing protocols. The comments below primarily address methodological details, interpretation of ecological patterns, and opportunities to further strengthen the robustness of the analyses and conclusions.

      Pages 136-137: The manuscript focuses primarily on L1-containing contigs. Could the authors provide a summary of libraries containing other PV hallmark genes (e.g., E1 or E2) but lacking full-length L1 sequences? This would help assess how much additional PV diversity may remain inaccessible under an L1-centered framework. Furthermore, the manuscript does not provide a detailed discussion of technical or biological explanations for libraries with detected PV sequences but lacking L1 sequences. Could such cases represent incomplete assemblies, low-abundance infections, highly divergent PVs, or endogenous papillomavirus-derived elements? Clarifying these possibilities would help readers interpret the biological significance of these detections.

      Page 141: The rationale for performing the search in two sequential Logan releases is unclear. Does Logan v1.1 fully supersede v1.0? If so, why was the initial search performed on v1.0 rather than directly on v1.1? Providing a clearer description of the differences between the two database versions and explaining how these differences motivated the two-round search strategy would improve reproducibility.

      The data could have been explored at the libraries' read level. The analysis is currently limited to presence/absence and diversity patterns derived from assembled PV contigs. However, the identified PV-positive libraries provide an opportunity to explore abundance-related metrics. Read counts, coverage estimates, or other measures of sequence representation could be used to characterize PV abundance within libraries, providing additional ecological context and helping distinguish low-level incidental detections from strongly represented infections. In addition, the curated PV sequence dataset generated in this study could serve as a reference for targeted read-mapping analyses. Aligning reads from a subset of libraries classified as PV-negative may help determine whether PV sequences are present below assembly detection thresholds. Such an analysis could provide valuable insights into the sensitivity of assembly-based virus discovery approaches and help establish practical coverage or read-count thresholds for detecting low-abundance papillomaviruses. These results could have important implications for future surveillance, clinical, and environmental studies aimed at PV detection.

      Pages 155-156: Clustering was performed using sequence identity, while host, geographic, and ecological annotations were assigned from representative centroids. How frequently did clusters contain sequences associated with conflicting metadata (e.g., distinct hosts or geographic regions), and how were such cases handled?

      Page 160: The authors state that the 70% query coverage threshold was selected based on an observed bimodal distribution. Could this analysis be shown explicitly (e.g., in a supplementary figure), and could the authors discuss how sensitive the number of novel PV calls could have been to alternative coverage thresholds?

      Pages 167-168: The rationale for clustering highly divergent sequences at 60% nucleotide identity should be explained. Is this threshold associated with established genus-level classifications in this viral family, or was it chosen empirically?

      Pages 167-168: For the 45 sequences lacking nucleotide-level matches, did the authors investigate amino-acid similarity to known PV L1 proteins? Such analyses would help determine whether these sequences represent deeply divergent PVs or potentially more distant viral lineages.

      Pages 181-183: The biological interpretation of host-associated PV diversity may depend on library type. Could the authors summarize the proportion of samples originating from field collections, laboratory animals, cell culture systems, or experimental infections?

      Pages 243-248: Could the observed geographic and ecological patterns be influenced by laboratory-derived samples? Distinguishing field-collected samples from laboratory, captive, or experimental material would strengthen the ecological interpretations.

      Pages 254-261: The biome analyses focus on PV occurrence. An analysis of host composition across biomes would be highly informative and could help disentangle whether observed patterns reflect PV ecology or underlying host distributions.

      Page 294: Statements regarding structural similarity appear to rely primarily on visual comparisons. Could the authors provide quantitative structural alignment metrics (e.g., RMSD, TM-score, DALI score, Foldseek score) to support these conclusions?

      Page 321-324: Given the scale and curation of the dataset, the final case-study section is limited. Broader comparative analyses of gene content, ORF architecture, and composition related to host associations, phylogenetic relationships, and/or ecological variables could provide additional evolutionary insights beyond a small number of illustrative examples.

      Page 333: Given the emphasis on the feasibility of petabase-scale sequence mining, the manuscript would benefit from a more detailed description of the computational resources required. The reported ~10-hour runtime is difficult to interpret without information regarding hardware specifications, CPU-hours, memory requirements, storage footprint, and cloud infrastructure (if used). Such information is important for evaluating the reproducibility and practical applicability of the approach.

      Pages 408-409: Were metadata available regarding viral enrichment procedures, particle purification, or size-selection protocols? Such information could influence the interpretation of PV detection frequencies across library types.

      Pages 415-416: The differentiation between viral and endogenous viral sequences is one of the biggest challenges in viral metagenomics and large-scale data mining for viral sequences. This issue is particularly relevant because the distinction between exogenous and endogenous viral sequences may directly affect estimates of novel PV diversity and inferred host associations. The manuscript acknowledges that papillomavirus sequences recovered from DNA-based libraries may derive from integrated viral DNA. However, there is no systematic analysis addressing the potential contribution of endogenous papillomavirus elements (EVEs) to the reported diversity estimates. Given the large number of host genome sequencing projects represented in the SRA, some detected PV-like sequences may correspond to integrated or fossil viral sequences rather than exogenous viruses. The authors could discuss this possibility more explicitly and provide analyses evaluating the prevalence of integration signatures, disrupted ORFs, host-genome flanking regions, or other indicators that would help distinguish endogenous viral elements from actively circulating PVs.

      Pages 532-533: The study is described as "petabase-scale"; however, the analyses were performed on a pre-assembled and compressed representation of the SRA rather than directly on petabase-scale raw sequencing data. The authors may wish to clarify this distinction and explicitly acknowledge that the computational burden is substantially reduced by the Logan framework and, from this perspective of computational power applied, this study is not in the same context as Serratus and Logan.

      Perspective comment: One of the strengths of this study is the generation of a highly curated papillomavirus protein dataset spanning a broad range of known and newly identified PV diversity. Given the increasing importance of structure-based homology detection in virology, the authors may wish to discuss the potential of this resource for future structure-guided discovery efforts. Recent studies have shown that protein structure prediction and comparison can reveal extremely distant evolutionary relationships that are undetectable at the sequence level. The curated PV dataset generated here could serve as a valuable reference for searching unannotated proteins from metagenomic "dark matter" datasets for structural homologs or convergent folds related to papillomavirus proteins. Such approaches may help identify highly divergent PV lineages or previously unrecognized viral proteins that retain structural similarity despite extensive sequence divergence.

    1. eLife Assessment

      This valuable study reports that ALDH-abundant cells exhibit stem cell properties and may play a key role in endometrial epithelial development in mice. The data support the main conclusion and are convincing. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

    2. Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and builds on established hypotheses in the field. The manuscript is well written and scientifically sound and the important experimental limitations and interpretation caveats are presented throughout.

      The majority of my initial comments have been adequately addressed within the text.

      Remaining limitations include:

      (1) The impact of tamoxifen injection directly on Aldh1a1 expression in the developing uterus.

      (2) It would be beneficial to demonstrate the degree of cell death following diphtheria toxin treatment 24-48 hours after injection in Tam-treated mice at PND 10. It is not clear as to why the 4-day timepoint was selected. Cells expressing the DTR should begin undergoing apoptosis within several hours after treatment.

    3. Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      Comments on revised version.

      In the revised manuscript, comments have largely been addressed and the manuscript is improved. The authors have tempered their inference of the contribution of ALDH1A1+ cells to endometrial regeneration, but the conclusions are still somewhat overstated. However, the overall work provides valuable insight and strengthens the growing body of literature characterizing endometrial stem/progenitor cells and their function.

    4. Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      Comments on revised version.

      The authors have answered the questions in the revised manuscript, no further comments.

    5. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This valuable study reports that the ALDH-abundant cells display stem cell properties and may play a key role in the endometrial epithelial development in the mouse. The data supporting the main conclusion are solid, although further improvements are needed to strengthen the conclusions. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

      We thank the reviewers and editor for their critical reading and assessment of our manuscript. We carefully considered each of the points raised by the reviewers. In this document and in the edited manuscript and figures, we have carefully addressed each of the comments and requested modifications. In light of these changes, we expect that you will find that the manuscript has improved.

      We indicate our responses to the reviewers below in blue font and highlight the changes in the manuscript using the line numbers corresponding to the tracked version of the revised document.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and will build on established hypotheses in the field. The manuscript is well written and scientifically sound; however, several experimental limitations and interpretation caveats should be addressed.

      We thank the reviewer for their comments and expert assessment of our paper.

      (1) The methods surrounding the passage number and duration of culture following sorting prior to transcriptomic profiling should be clarified in the figure legends. Related to this, the representative images in Figures 1D and 1E do not appear consistent with the quantification presented in Figures 1F-H and should be reconciled.

      Thanks for this comment. We have now clarified this in the Figure 1 legend as follows,

      LINES 1026-1029: “Organoid formation assay performed immediately after luminal epithelial cell isolation and by plating equal numbers of viable ALDH<sup>LO</sup> (D) and ALDH<sup>HI</sup> (E) epithelial cells. ALDH<sup>LO</sup> and ALDH<sup>HI</sup> organoids were cultured for two weeks and passaged once prior to the organoid formation assays and transcriptomic analyses.”

      Regarding the second comment, we recognize that the images we showed may not have been the most representative of our quantification. As such, we replaced them with the organoid images so that they better reflect the quantification outlined in Figure 1F-H.

      (2) The conclusion that ALDH1A1+ cells are enriched in populations with stem cell characteristics relies primarily on transcriptomic analysis. Protein-level co-localization should be performed to strengthen this claim.

      We thank the reviewer for this comment. Unfortunately, the antibodies for many of these stem cell markers (such as LGR5, AXIN2, and SUSD2) are not well-suited for immunostaining. Others that have been proposed in human and are amenable to immunostaining are not suitable markers for mouse endometrial stem cells (such as CDH2). We hope that by showing that ALDH1A1 is expressed in patterns that are similar to the previously published stem cell markers LGR5 and AXIN2 (i.e., throughout the epithelium in the developing uterus and subsequently enriched in the tips of the endometrial glands of adult mice), along with transcriptomic studies, we can demonstrate its utility as a marker for mouse endometrial stem cells.

      (3) The overlap of 19 genes between the data set here and AXIN2 HI data is presented as evidence of shared stemness identity, but no statistical assessment of this overlap is provided. A hypergeometric test should be performed to determine whether this overlap is greater than expected by chance.

      Thank you for this suggestion. We have performed a hypergeometric test and determined that the reported shared genes between the two datasets are greater than is expected by chance. We have updated the results section to state the following:

      Lines 137-140: "We determined that the overlap between ALDH<sup>HI</sup> and Axin2<sup>+</sup> stemness marker genes was significantly greater than expected by chance for both upregulated (21/346 genes, 1.81-fold enrichment, p = 0.0067) and downregulated (19/674 genes, 1.67-fold enrichment, p = 0.021) gene sets (hypergeometric test, universe = 23,182 genes)."

      (4) The impact of tamoxifen injection on Aldh1a1 expression should be characterized in the neonatal uterus, as tamoxifen itself has known estrogenic activity that could confound interpretation of the lineage tracing results at early postnatal timepoints.

      Although we took measures to control for this possibility by using multiple time-points and models to trace the impact of Aldh1a1<sup>+</sup> cells in development and adulthood, we recognize the importance of this comment and acknowledge that this is a limitation in the design of our study. We have included the following text to the Discussion acknowledging this point:

      Lines 433-441: “Given the well-documented impacts of tamoxifen for lineage tracing studies, it is imperative to use doses of tamoxifen that will minimize estrogenic impacts and result in off-target effects (Rios et al., 2016). This often requires administration at doses that will achieve maximal recombination of the desired gene, while ensuring that the potential deleterious impacts of tamoxifen are minimized (Chen et al., 2023; Pimeisl et al., 2013). The cre/ERT2 tamoxifen inducible model is widely used to study uterine biology where it serves as a useful tool to interrogate the spatiotemporal impact of key genes, either through inactivation or for lineage tracing. Despite its widely documented utility across many tissue types and developmental timepoints, the use of tamoxifen and its impacts on the endometrium remain a limitation of our study, which we tried to address by implementing multiple timepoints, doses, and orthogonal assays in our experimental design.”

      (4b) Related to this, while low-dose tamoxifen is shown to label individual cells within 24 hours of injection, the translation dynamics of the label following Cre-mediated recombination can require up to 72 hours. The presence of only a few labeled clones at PND8 but multiple separate clones per cross-section at later timepoints warrants discussion and may reflect labeling kinetics rather than clonal expansion.

      The reviewer raises an important point. We agree that the 72hr-translation kinetics of the cre-mediated recombination is a legitimate consideration for interpreting our data and we have added the text below to the Discussion section acknowledging this point.

      We have addressed this by adding the following text to the discussion:

      Lines 417-422: We hypothesized that the singly labeled cells observed from one day tracing experiments expanded in a clonal fashion during the various timepoints we measured. We note that the translation kinetics of the labeled cells following cre-mediated recombination may contribute to the limited labeling observed at PND8/PND15 and there is a potential for delayed labeling of cells between 24 and 72 hours of tamoxifen administration. However, the continuous increase in labeled cells at the subsequent timepoints favors our interpretation of clonal expansion as the primary explanation.

      (5) It would strengthen the in vivo ablation data to validate the degree of cell death following diphtheria toxin treatment directly. It is possible that a general decrease in cell number rather than specific loss of a stem cell population is responsible for the observed reduction in gland number and FOXA2 expression (Tongtong et al 2017).

      We agree that this is an important control to incorporate into our experimental design. To rule out this possibility, we performed immunohistochemistry of cleaved caspase 3 in the uterine tissues of DTR<sup>flox/flox</sup> and DTR<sup>flox/flox</sup>;Aldh1a1<sup>cre/ERT2</sup> mice 4 days after administration of diphtheria toxin. The results indicate similar levels of cleaved caspase 3 detection in both genotypes, suggesting that the decrease in FOXA2+ cells is not due to non-specific cell death, but rather the result of ALDH1A1<sup>+</sup> cells. These data and the following text have been added to the manuscript:

      Lines 320-324: “We determined that the decreased in FOXA2<sup>+</sup> cells in the experimental mice was not the result of non-specific DT-mediated cell death, as similar levels of cleaved caspase 3-positive cells were detected in the DT-treated control ROSA26<sup>DTR/DTR</sup> and ROSA26<sup>DTR/DTR</sup>;Aldh1a1<sup>cre/ERT2/+</sup> mice 4 days post-diphtheria toxin administration (Figure S3G-H’).”

      (6) The lineage tracing data in the postpartum endometrium demonstrate that Aldh1a1-marked cells are present during regeneration, but it remains unclear whether these cells are preferentially activated or expanded in response to tissue injury. Coupling these studies with diphtheria toxin-mediated ablation during active regeneration would more directly test the proposed regenerative role of this population.

      This is a great point and one that we would be very interested in pursuing as follow-up studies in our future work. Regretfully, due to the long generation time and experimental procedures associated with these proposed studies, we are not able to include these experiments in the current manuscript. Thus, we have changed our wording and conclusions throughout the manuscript to be less definitive in terms of the role of Aldh1a1 in regeneration, since this will be the focus of future studies.

      The contribution of stromal Aldh1a1 lineage-positive cells is underexplored in the discussion, given the lineage tracing data showing stromal labeling across multiple timepoints and its potential relevance to mesenchymal-to-epithelial transition.

      Thank you for the suggestion. We have now expanded this section in the Discussion to include the following:

      Lines 496-504: We also found ALDH1A1<sup>+</sup> stromal cells were more prevalent when tracing began in adult mice. Other studies have shown that mesenchymal cells contribute to endometrial regeneration in the postpartum phase or after induced menses through a process of MET (Cousins et al., 2014; Kirkwood et al., 2022; Li et al., 2025). Similarly, lineage tracing studies have shown that MET is an active process and contributes to epithelial cell regeneration in the post-partum phase (Huang et al., 2012; Patterson et al., 2013). Although this is an area of active investigation in the field, with some contradicting reports, it is plausible to hypothesize that endometrial tissue has the capacity to undergo wound-healing and regeneration via several mechanisms (Ang et al., 2023; Ghosh et al., 2020). The process of MET in wound healing is widely documented in other organs, such as the kidney, liver and lung, where MET is associated with depletion of the resident epithelial cell pool (Bi et al., 2012; Niayesh-Mehr et al., 2024; Zeisberg et al., 2005).

      Finally, the word 'control' may overstate the functional evidence presented. 'Contribute' may be more accurate given the partial and context-dependent nature of the phenotypes observed.

      We agree with the reviewer’s point that control may overstate the evidence that we provide in the manuscript. To reflect this, we have edited the manuscript title and text to address this suggestion.

      Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      This study addresses an important area of research in the field of endometrial stem/progenitor cell biology. The authors are commended for their use of multiple complementary methods, including lineage tracing, DTR-mediated cell ablation, organoid assays, and RNA-seq in mouse and human models to assess the stem-like nature of Aldh1a1+ cells. The data support the stem/progenitor phenotype of Aldh1a1+ epithelial cells during endometrial development; however, there are noted discrepancies between organoid formation assays and lineage tracing experiments regarding the stemness of Aldh1a1+ epithelial cells in adults. Specifically, organoids were generated from adult cells and demonstrated in vitro stem cell activity; however, in vivo lineage-tracing of adult cells either during the estrous cycle or postpartum repair does not show expansion of Aldh1a1+ cells, suggesting they do not have stem/progenitor activity. Additionally, the stem-ness of epithelial vs stromal Aldh1a1+ cells is confounded in the study because epithelial cells were not purified for organoid experiments, epithelial cells were not exclusively lineage-traced as stromal cells were also labeled, and mesenchymal-epithelial transition was suggested to occur during postpartum repair. The following specific comments are presented to detail these concerns:

      We thank the reviewer for their critical reading of our manuscript and constructive comments.

      (1) The statement in the brief summary, "...critical for lifelong endometrial regeneration," is not supported by the data provided.

      We have edited the brief summary to exclude this statement, it now reads as follows:

      Lines 4-5: “We uncover ALDH1A1<sup>+</sup> cells as a group of hormone sensitive stem cells contributing to endometrial development and regeneration.”

      (2) AlDH1A1 is not restricted to the endometrial epithelium, and epithelial cells were not purified by flow cytometry for experiments in Figure 1. Figure 2 clearly shows the presence of mesenchymal cells, even using the described method for enriching for epithelial cells. Therefore, contaminating mesenchymal cells with high ALDH activity may confound the experimental results in Figure 1, either through promoting epithelial cell growth or through MET. The authors should provide clear evidence of epithelial purity in organoid experiments or that mesenchymal cells are not contained in the ALDHhi population. These comments also apply to the human organoid experiments in Figure 7.

      We thank the reviewer for raising this important point. Our group has been using the enzymatic method to routinely separate epithelial from stromal cell populations from the mouse uterus (see references dating back to 2015, PMID 26721398, 28324064, 34099644). In these experiments we typically obtain >98% purity in the epithelial and stromal cell compartments, respectively. We can directly observe this purity in the immunofluorescence images shown I Author response image 1 and Author response image 2, where mouse endometrial epithelial cells and stromal cells were enzymatically separated and immunostained with E-cadherin and vimentin antibodies to detect epithelial and mesenchymal cells in both cell preparations. The images show very few contaminating epithelial and stromal cells in either cell preparation. We have observed similar results when preparing epithelial and stromal cell preparation from the human endometrium, where the epithelial cell organoids display high purity with ~100% epithelial cell expression when we perform immunostaining.

      Author response image 1.

      Purity of mouse endometrial epithelial cells obtained via enzymatic and mechanical dissociation. A-B) Shows the epithelial (A) and stromal (B) cells plated on glass coverslips and immunostained with an epithelial cell marker (cytokeratin 8, red), a stromal cell marker (vimentin, green), and DAPI.

      Author response image 2.

      Human endometrial epithelial organoids were fixed and immunostained with cytokeratin 8 (green) and DAPI. The images are typical for our epithelial cell cultures and demonstrate that all epithelial cells are CK8-positive.

      (3) Lines 186-187: Susd2 was increased in EpSC clusters, yet this is a mesenchymal stem/progenitor marker in humans. The authors should discuss the implications of this.

      We thank the reviewer for highlighting this. We have now included the following in our Discussion to address this point:

      Lines 527-532: Clustering with this population of EpSCs were Susd2<sup>+</sup> cells, which are well-characterized mesenchymal progenitors that are enriched in the perivascular regions of the human endometrium (Darzi et al., 2016; Khanmohammadi et al., 2021). The presence of Susd2<sup>+</sup> cells, while unexpected in an epithelial stem cell niche, could indicate the presence of a transitional mesenchymal or perivascular cell that is differentiating into epithelium. Evidence for both mesenchymal and Nestin2<sup>+</sup> pericytes have been recently described in the mouse endometrial epithelium (Kirkwood et al., 2022; Li et al., 2025).

      (4) In Figure 5, RFP+ epithelial cells should be quantified as in previous figures to substantiate the statement in lines 279-280, "At PPD5, the proportion of RFP+ epithelial cells had expanded relative to PPD1 and PPD3 (Figure 5E-E')." Especially because in the low mag images (C-E), RFP+ epithelial cells appear to be most abundant at PPD1 and decrease at PPD3 and PPD5, suggesting that they may not be involved in endometrial regeneration/repair (contradicting the interpretation in line 285). Further, if there is in fact a decrease over postpartum repair, then regeneration should be removed from the title of the manuscript. RFP+ stromal cells should also be quantified.

      We appreciate this reviewer’s comment and agree that as stated, the conclusion is not fully supported by the data. To address this comment, we have edited the results so that they clearly indicate the results and remove any ambiguity:

      As requested, we quantified the number of RFP+ stromal and epithelial cells during the postpartum phase and noted that RFP+ cells were prominent in the stromal compartment of the endometrium. While RFP+ epithelial were also observed during these timepoints, they were less abundant than RFP+ stromal cells. Because the number of RFP+ cells did not significantly change over the postpartum phases in neither the stromal nor epithelial compartment, we have modified our conclusion to state that ALDH1A1+ cells are transiently detected in the regenerating endometrium.

      Results:

      Lines 287-294: “By analyzing the uterine tissues near the placental detachment site, we observed that RFP positive cells were prominent in the endometrial stromal cells that were adjacent to the luminal epithelium (Figure 5C-C’, green arrows). RFP<sup>+</sup> cells were also observed in the stromal cells near the placental detachment sites at PPD1 and PPD3 (Figure 5D’-E’, red & blue arrows) and in limited luminal epithelial cells (Figure 5D”,E”). Quantification of RFP+ cells throughout these postpartum phases indicated that stromal cells had more frequent ALDH1A1<sup>+</sup> stromal cells (360 ± 103, PPD1, n=3; 217 ± 107, PPD3, n=3; 254 ± 32, PPD5, n=4) than ALDH1A1<sup>+</sup> epithelial cells in the regenerating endometrium (65 ± 65, PPD1, n=3; 20 ± 10, PPD3, n=3; 114.25 ± 39, PPD5, n=4) (Figure S4).”

      Discussion:

      Lines 512-520: “We also noted that a majority of ALDH1A1<sup>+</sup> cells were localized to the active areas of endometrial regeneration near the placental detachment sites at PPD1 with a pronounced expression in the sub-epithelial stromal cells. As regeneration progressed, we continued to observe ALDH1A1<sup>+</sup> cells in the stromal compartment within the placental detachment sites at PPD3 and PPD5, with a progressive, but not statistically significant, increase in ALDH1A1<sup>+</sup> epithelial cells. Collectively, our data demonstrate that ALDH1A1<sup>+</sup> lineage cells participate in the restoration of endometrial architecture and functional compartments in the postpartum phase, even if their direct contribution is transient. Future detailed and mechanistic studies will be necessary to fully characterize their role in this process and their long-term consequence in postpartum regeneration.”

      (5) For Figure 7F, it should be clearly stated in the main text that the results are from one patient sample and the data presented are experimental replicates, so as not to be confused with biological replicates (the same for Supplementary Figure S4). Were B and G in Figure 7 also from one patient?

      Thanks for pointing this out. We have edited the figure legends in the main text and supplemental figures to indicate this.

      Lines 336-337: “…main figures show representative results from one patient sample performed in technical replicates, with additional patient samples included in the supplement…”

      (6) Lines 425-427: "Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed a complete restriction of ALDH1A1 to the glandular crypts." In Figure 2 S' ALDH1A1+ cells are visible in the LE (the staining is lighter than in the GE but looks real), contradicting this statement.

      This is an important distinction. We have now edited this part of the manuscript to state:

      Lines 458-461: “Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed enriched ALDH1A1 in the glandular crypts with weak luminal epithelial staining, while the ovariectomized controls had strong ALDH1A1 expression throughout the luminal and glandular epithelium.”

      (7) Lines 466-467: "In cycling mice, we found sporadic cells that expressed both stromal and epithelial markers in the ALDHA1+ cells." These data are not presented.

      We apologize for the confusion, this sentence has been removed from the discussion.

      (8) These data support the role of Aldh1a1+ cells in endometrial epithelial development, but conclusions about their role in repair/regeneration should be tempered as the data are much weaker here.

      We thank the reviewer for their overall assessment. To address this point, we have thoroughly edited the appropriate areas to temper the conclusions and ensure that they are strongly supported by our data. We have also edited the manuscript’s title to reflect this.

      Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      We thank the reviewer for their assessment of our manuscript.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      We thank the reviewer for their expert assessment and positive comments regarding our manuscript.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As this study evaluated Aldh1a1+ cells in both the epithelium and stroma, it is recommended that the title and abstract be revised to reflect this.

      Thank you. Both the title and abstract title have been updated to reflect this comment.

      (2) Lines 46-47 in the abstract: "Aldh1a1+ cells expanded during postnatal development, estrus cycling, and following post-partum repair." It is recommended to clarify stromal vs epithelial Aldh1a1+ cell expansion, as only the stromal cells expanded during the estrous cycle. Also, RFP+ epithelial cells were not quantified during postpartum repair and visually appear to decrease (see comment below regarding Figure 5), so this statement is misleading.

      The abstract was edited following the suggestions so that it depicts our results. Similarly, we have addressed the comment regarding Figure 5 and the interpretation of the post-partum regeneration experiments (see comment above for the full explanation and edits to the manuscript).

      (3) Lines 65-69, 186-187: the authors should be clear when describing markers of putative epithelial vs mesenchymal stem/progenitor cells in the introduction (e.g., CDH2/SSEA1/SOX9 for epithelial and SUSD2 for mesenchymal).

      Thank you for this suggestion. The Introduction section now states the following:

      Lines 70-73: “These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).”

      (4) Lines 78-80: "Studies tracing the fate, ablation, and proliferative capacity of Lgr5+ cells in the uterus identified an Lgr5+ niche that is enriched in the crypts of the glandular epithelium and promotes endometrial regeneration (Seishima et al., 2019)." This statement is incorrect regarding regeneration, as Lgr5 marks stem/progenitor cells in the developing uterus but not the adult, during which endometrial regeneration occurs. Please revise.

      Thank for you for this clarification. The statement has been revised and now reads as follows:

      Lines 82-84: “Studies tracing the fate, ablation, and proliferative capacity of Lgr5<sup>+</sup> cells in the uterus identified an Lgr5<sup>+</sup> niche that marked stem/progenitor cells in the developing uterus (Seishima et al., 2019).”

      (5) Recommend using free-form shapes to outline GE, LE, and EpSC in Fig 2C and stating in the text which clusters correspond to LE and GE (lines 167-169).

      Thank you, this has been edited in Figure 2C and in the text, which now reads as follows:

      Lines 176-177: “Clusters 5, 7, 21 and 16 were classified as glandular epithelial cells, and clusters 0, 2, 3, 10, 14, and 24 were classified as luminal epithelial cells (Figure 2C).”

      (6) "Estrus" refers to the specific stage of the "estrous" cycle. Estrus cycle is incorrect and should be estrous cycle.

      Thank you for pointing this out. It has been corrected throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Suggest increasing the font size for some of the labels in the figures.

      Thank you, we have increased font size in the figures to improve the quality.

      (2) Need to include more details of the human endometrial tissue used in this study: pathology, age, and stage of menstrual cycle.

      We thank the reviewer for the helpful suggestion. The details that are available to us have been included in Supplementary Table S5.

      (3) Line 65-67 - Different endometrial stem cell subsets - SUSD2+ reside in perivascular regions, while other markers are located in glandular epithelium, need revision.

      Thank you, this has been revised in the Introduction. The area now reads as follows:

      Lines 70-73: These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).

      (4) Include a discussion about the interpretation of their current finding in relation to the dynamic regeneration observed in human endometrium due to menstrual bleeding/tissue breakdown, compared to the cycles of growth and regression that occur in mice.

      This is a great suggestion. We have added the following statement to our Discussion section:

      Lines 537-544: Additionally, our studies in human endometrium extend our characterization of ALDH1A1 as an adult endometrial stem cell marker and emphasize the importance of ALDH1A1+ in the regenerative potential of the endometrium. The conserved hormonal responses between human and mouse endometrium support the hypothesis that cycles of proliferation, differentiation, and regression, whether through resorption/autophagy in mice or menstrual breakdown in humans, are governed by concerted growth factor signaling and dedicated stem cell populations with the capacity to expand and differentiate across repeated cycles of repair. Our detailed studies in both mouse and human models indicate that ALDH1A1+ cells represent a dedicated cell type within the endometrium with the potential to drive repair during adulthood. Collectively, these findings advance our understanding of the mechanisms that control endometrial cycling and regeneration throughout the reproductive lifespan.

      Reference

      Ang, C.J., Skokan, T.D., and McKinley, K.L. (2023). Mechanisms of Regeneration and Fibrosis in the Endometrium. Annu Rev Cell Dev Biol 39, 197-221.

      Bi, W.R., Jin, C.X., Xu, G.T., and Yang, C.Q. (2012). Bone morphogenetic protein-7 regulates Snail signaling in carbon tetrachloride-induced fibrosis in the rat liver. Exp Ther Med 4, 1022-1026.

      Chen, M.Y., Zhao, F.L., Chu, W.L., Bai, M.R., and Zhang, D.M. (2023). A review of tamoxifen administration regimen optimization for Cre/loxp system in mouse bone study. Biomed Pharmacother 165, 115045.

      Cousins, F.L., Murray, A., Esnal, A., Gibson, D.A., Critchley, H.O., and Saunders, P.T. (2014). Evidence from a mouse model that epithelial cell migration and mesenchymal-epithelial transition contribute to rapid restoration of uterine tissue integrity during menstruation. PLoS One 9, e86378.

      Cousins, F.L., Pandoy, R., Jin, S., and Gargett, C.E. (2021). The Elusive Endometrial Epithelial Stem/Progenitor Cells. Front Cell Dev Biol 9, 640319.

      Darzi, S., Werkmeister, J.A., Deane, J.A., and Gargett, C.E. (2016). Identification and Characterization of Human Endometrial Mesenchymal Stem/Stromal Cells and Their Potential for Cellular Therapy. Stem Cells Transl Med 5, 1127-1132.

      Ghosh, A., Syed, S.M., Kumar, M., Carpenter, T.J., Teixeira, J.M., Houairia, N., Negi, S., and Tanwar, P.S. (2020). In Vivo Cell Fate Tracing Provides No Evidence for Mesenchymal to Epithelial Transition in Adult Fallopian Tube and Uterus. Cell Rep 31, 107631.

      Huang, C.C., Orvis, G.D., Wang, Y., and Behringer, R.R. (2012). Stromal-to-epithelial transition during postpartum endometrial regeneration. PLoS One 7, e44285.

      Khanmohammadi, M., Mukherjee, S., Darzi, S., Paul, K., Werkmeister, J.A., Cousins, F.L., and Gargett, C.E. (2021). Identification and characterisation of maternal perivascular SUSD2(+) placental mesenchymal stem/stromal cells. Cell Tissue Res 385, 803-815.

      Kirkwood, P.M., Gibson, D.A., Shaw, I., Dobie, R., Kelepouri, O., Henderson, N.C., and Saunders, P.T.K. (2022). Single-cell RNA sequencing and lineage tracing confirm mesenchyme to epithelial transformation (MET) contributes to repair of the endometrium at menstruation. Elife 11.

      Li, S.Y., Whiteside, S., Li, B., Sun, X., and DeFalco, T. (2025). Mesenchymal-to-epithelial transition of perivascular cells contributes to endometrial re-epithelialization. Nat Commun 16, 10174.

      Niayesh-Mehr, R., Kalantar, M., Bontempi, G., Montaldo, C., Ebrahimi, S., Allameh, A., Babaei, G., Seif, F., and Strippoli, R. (2024). The role of epithelial-mesenchymal transition in pulmonary fibrosis: lessons from idiopathic pulmonary fibrosis and COVID-19. Cell Commun Signal 22, 542.

      Patterson, A.L., Zhang, L., Arango, N.A., Teixeira, J., and Pru, J.K. (2013). Mesenchymal-to-epithelial transition contributes to endometrial regeneration following natural and artificial decidualization. Stem Cells Dev 22, 964-974.

      Pimeisl, I.M., Tanriver, Y., Daza, R.A., Vauti, F., Hevner, R.F., Arnold, H.H., and Arnold, S.J. (2013). Generation and characterization of a tamoxifen-inducible Eomes(CreER) mouse line. Genesis 51, 725-733.

      Rios, A.C., Fu, N.Y., Cursons, J., Lindeman, G.J., and Visvader, J.E. (2016). The complexities and caveats of lineage tracing in the mammary gland. Breast Cancer Res 18, 116.

      Seishima, R., Leung, C., Yada, S., Murad, K.B.A., Tan, L.T., Hajamohideen, A., Tan, S.H., Itoh, H., Murakami, K., Ishida, Y., et al. (2019). Neonatal Wnt-dependent Lgr5 positive stem cells are essential for uterine gland development. Nat Commun 10, 5378.

      Zeisberg, M., Shah, A.A., and Kalluri, R. (2005). Bone morphogenic protein-7 induces mesenchymal to epithelial transition in adult renal fibroblasts and facilitates regeneration of injured kidney. J Biol Chem 280, 8094-8100.

    1. eLife Assessment

      The authors use solid high resolution microscopy techniques to present a model in which SARS-CoV-2 attachment and endocytosis are mediated by heparan sulfate, whereas ACE2 only functions downstream of these early processes. These findings serve as a potentially valuable starting point to examine SARS-CoV-2 entry models that challenge current paradigms. However, examination of the model in a clean heparan sulfate deficient background and in the context of TMPRSS2-dependent plasma membrane fusion are lacking. Thus, the broader impact remains uncertain in the absence of additional controls and orthogonal approaches.

    2. Reviewer #1 (Public review):

      This revised paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. The authors used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions downstream of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. The authors conclude their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and use of primary airway cells. While some studies were performed in the revision to address the Reviewer concerns, which improved the paper clarity, others were cursorily addressed, which limit the impact of the studies. Particularly. experiments were not performed to account for TMPRSS2 expression and plasma membrane fusion. Moreover, addition of studies in which hACE2 is expressed in cells genetically lacking HS were not designed. Thus, it the picture remains unclear picture exactly where downstream hACE2 functions and how this might differ given new structural models of TMPRSS2 activation (PMID: 42050172), which occur after ACE2 recognition of spike on the cell surface.

    3. Reviewer #2 (Public review):

      In the manuscript by Han et al, the authors assess binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that SARS-CoV-2 spike (on the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well described but the authors here attempt to make the claim that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. Additional controls would be of great benefit.

      Further, it is this reviewers opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. A more balanced and nuanced view of their interesting data would be of value.

      Major:

      The authors need to rigorously define a "receptor." This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm) while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is not necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this definition. This is proven in Fig 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959 but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. This is inconsistent with the claim that HS is broadly important (beyond the BHK cells overexpressing ACE2 that are used here).

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide inadequate discussion into the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2 deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used as protease expression is a key determinant of the site of fusion.

      An alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays are TMPRSS2 vs Cathepsin L dependent (using E64d and camostat for instance) as mentioned above.

      Each figure legend should clearly state how many independent experiments and replicates per experiment were performed.

      All bar plots should show individual dots (i.e. Fig 1G) to better reveal the variance of each dataset.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. A prevailing model that has been developed over the last five years by work from many laboratories using a variety of biochemical, structural, and microscopic approaches is that HS acts a co-receptor for SARS-CoV-2; its binding to SARS-CoV-2 both concentrates virus on the surface of target cells and allosterically alters the spike protein to promote an "up/open" RBD conformation that enables engagement of the proteinaceous receptor human ACE2 on the cell surface (PMID: 32970989, 35926454, 38055954, 39401361, 40548749). These two events enable plasma membrane fusion (after a cleavage event promoted by plasma membrane TMPSS2) or endocytosis and subsequent pH-dependent fusion (which requires a cathepsin L-mediated cleavage of the spike).

      The authors in this study used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions "downstream" of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. Blocking HS binding with pixantrone, a drug under clinical evaluation for cancer (due to its anti-topoisomerase II activity), inhibited SARS-CoV-2 Omicron JN.1 variant from attaching to and infecting human airway cells. The authors conclude that their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the interesting and technically well-executed study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and the use of primary airway cells. Nonetheless, there are issues that need to be addressed to buttress the proposed model compared to earlier ones. These include: (a) the distinction between macropinocytosis and receptor-mediated endocytosis and what this might mean for productive SARS-CoV-2 infection; (b) the need to account for TMPRSS2 expression and plasma membrane fusion; (c) addition of genetic studies in which hACE2 is expressed in cells lacking HS; (d) an unclear picture of exactly where downstream hACE2 functions; and (e) and a need for comparative/additional study of earlier SARS-CoV-2 variants, which preferentially fuse at the plasma membrane.

      We thank the reviewer for the strong support of this manuscript. We addressed the reviewer’s concerns in the Recommendations to the authors. We did not distinguish whether the endocytic route is macropinocytosis or receptor-mediated endocytosis, because it is a separate study beyond the scope of the present work. We did not examine earlier SARS-CoV-2 variants because we considered it a study beyond the scope of the present work, but a good idea that we may work on in the future. For detail on how we address the remaining concerns, please see our response to the reviewer’s Recommendations for the authors.

      Reviewer #2 (Public review):

      In this manuscript by Han et al, the authors assess the binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that the SARS-CoV-2 spike (in the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface, which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well-described, but the authors attempt to make the claim here that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears to be of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. The manuscript is imprecise and far overstates the actual findings shown by the data. Additional controls would be of great benefit.

      Further, it is this reviewer's opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. While the manuscript provides new insight into the details of this process, the manuscript attempts to oversell this finding by applying new words rather than new molecular details. The authors would be better served by presenting a more balanced and nuanced view of their interesting data. In this reviewer's opinion, the salesmanship significantly detracts from the data and manuscript.

      We thank the reviewer for pointing out that our manuscript is of interest to the field. However, we do not think that we oversell our data. hACE2 has been widely considered the receptor (or the binding partner) that mediates SARS-CoV-2 cell-surface attachment, whereas HS is considered only an attachment factor that facilitates SARS-CoV-2 binding with hACE2 at the cell surface. In the present work, we found that HS, but not hACE2, is the cell-surface attachment receptor (or binding partner), whereas hACE2 is not essential for attachment, but acts downstream of virus endocytosis to facilitate viral genome expression. This finding suggests significant modification of the current model by replacing the attachment receptor (or binding partner) from hACE2 to HS, treating HS as a primary receptor rather than an attachment factor, and relocating the hACE2 action site from the cell surface to the endosome. For these reasons, we do not consider these statements overselling our data. However, as the reviewer suggested in his/her specific comments, we revised the manuscript to ensure that we did not overgeneralize our findings (see our responses to the reviewer’s Recommendations to the authors).

      Major Comments:

      The authors need to rigorously define a "receptor" vs an "attachment factor." They also should avoid ambiguous terms such as "receptor underlying ...attachment" and "attachment receptor" (or at least clearly define them). Much of their argument hinges on the specific definition of these terms. This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm), while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is neither necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this fairly textbook definition. This is proven in Figure 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959, but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. The authors need to discuss this important information and reconcile it with their data and model if they want to claim that HS is broadly important.

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction, which likely won't excite a drug company. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide little discussion of the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2-deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used, as protease expression is a key determinant of the site of fusion.

      The claim that "SARS-CoV2 JN.1 variant binds to heparan sulfate, not hACE2, in primary human airway cells" is extraordinary and thus requires extraordinary evidence.

      First, PIX reduces attachment by 5-fold, which is not the same as "nearly abolished." Also, anti-ACE2 "nearly abolished" entry in 7D, while PIX did not. If the authors want to make these claims, an alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays is TMPRSS2 vs Cathepsin L dependent (using E64d and camostat, for instance) as mentioned above.

      Each figure should clearly state how many independent experiments and replicates per experiment were performed. What does "3 experiments" mean? Are these three independent experiments or three wells on one day?

      In the well-accepted current model, hACE2 is considered the receptor mediating SARS-CoV-2 cell-surface attachment, entry into cells, and infection, whereas HS is an attachment factor that facilitates SARS-CoV-2 binding to hACE2 at the cell surface. The present work revises this view: HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, with ACE2 acting downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined.

      We made this point clearer throughout the newly revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docks at HS clusters.

      The cited CRISPR-screen literature supports context-dependent host-factor usage. However, the absence of HS biosynthesis genes from a given screen does not prove that HS is irrelevant in that cell type; it only indicates that HS biosynthesis was not detected as a genetic dependency under that assay’s conditions. Such negative results can reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency. In the revised manuscript, we included the following in the Discussion:

      “While some studies using genome-wide CRISPR screening to identify genes involved in SARS-CoV-2 reveal genes for HS biosynthesis, others do not (45-50). The negative result, which might reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency, needs to be verified with specific gene knockout.”

      The ~5-fold reduction is likely due to the inhibitor not completely abolishing HS-virus binding. We revised the Discussion to strengthen the suggestion that targeting the virus cell-surface attachment by interfering HS binding is a therapeutic strategy to prevent and treat COVID-19, as in the following:

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding. However, this strategy has not been the focus for developing methods to prevent and treat COVID-19, likely because HS is considered only a regulator that is not essential for SARS-COV-2 entry. Our finding that HS is the attachment receptor re-emphasizes the importance of perturbing virus-HS binding, the first step of the viral entry, to efficiently block SARS-CoV-2 infection. Further supporting this view, inhibition of HS binding with a clinically used HS-binding agent, pixantrone, inhibits authentic SARS-CoV-2 JN.1 subvariant binding with HS on the cell surface and infection in primary human airway cells (Figs. 6, 7). These results suggest a combinatorial anti-SARS-CoV-2 strategy: early HS blockade to prevent attachment combined with ACE2 targeting to inhibit post-attachment steps”

      We include a sentence in the Discussion that our suggestions are limited to the cells we examined as below.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      Three experiments refer to three independent experiments. We added “independent” accordingly throughout the manuscript.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors define a new paradigm for the attachment and endocytosis of SARS-CoV-2 in which cell surface heparan sulfate (HS) is the primary receptor, with ACE2 having a downstream role within endocytic vesicles. This has implications for the importance of targeting virion-HS interactions as a therapeutic strategy.

      Strengths:

      The authors show that viruses are internalized via dynamin-dependent endocytosis and that endocytic internalization is the major pathway for pseudotyped SARS-CoV-2 genome expression. They show that HS-mediated viral attachment is a critical step preceding viral endocytosis and also subsequent genome expression. Further, they show that hACE2 acts downstream of endocytosis to promote viral infection, and may be co-internalised with virions after HS attachment. Pseudotyped virus and authentic SARS-CoV-2 provide similar results. In addition, the authors demonstrate that remarkable clusters of multiple HS chains exist on the cell surface, visualised by a number of elegant microscopy methods, and that these represent the docking sites for virions. These visualisations are an important general contribution in themselves to understanding the nanoscale interactions of HS at the cell surface.

      The use of a complementary range of methods, virus constructs, and cell models is a strength, and the results clearly support the conclusions.

      Overall, the results convincingly demonstrate a different model to the currently accepted mechanism in which the ACE2 protein is regarded as the cell surface receptor for SARS-CoV-2. Here, the authors provide compelling evidence that cell surface clusters of HS are the primary docking site, with ACE2 interactions occurring later, after endocytosis (whilst still being essential for viral genome expression). This is an exciting and important landmark evidence which supports the view that HS-virion interactions should be viewed as a key site for anti-viral drug targeting, likely in strategies that also target the downstream ACE2-based mechanism of viral entry within endosomes.

      We thank the reviewer for the strong support of the present work.

      Weaknesses:

      This reviewer identified only minor points regarding citing and discussing other studies and typos, which can be corrected.

      We have addressed these points in the revised manuscript. For detail, please see our response to the reviewer’s Recommendations to the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Pathway of internalization.

      The authors show clearly that labeled SARS-CoV-2 (pseudovirus or authentic virus) can become internalized in cells lacking hACE2, and this process depends on HS. However, they also show that this pathway is non-productive with regard to infection. Are the entry vesicles mediated by HS alone, HS + hACE2, and hACE2 alone the same? Or does the combination of co-receptor (HS + hACE2) drive SARS-CoV-2 into endocytic vesicles, whereas HS alone promotes macro- or micropinocytosis (lines 361-362). If HS alone directed SARS-CoV-2 into a non-productive entry vesicle, then hACE2 likely would be acting concurrently with HS and not downstream. A more detailed analysis of the different entry vesicles/pathways that occur with HS alone, HS + hACE2, and hACE2 alone is needed.

      During endocytosis, we did not detect a difference in the size distribution of virus-containing vesicles between BHK (HS alone) and BHK<sub>hACE2</sub> cells (HS+hACE2) (Fig. 2D). The similarity in the vesicle size suggests a similar endocytic path with HS alone or with HS + hACE2. In the revised manuscript, we added the following sentence.

      “Third, 3D-STED imaging showed that A490-labeled vesicle’s full-width-at-half-maximum (W<sub>H</sub>) was 363 ± 17 nm (n = 55) in BHK cells, similar to that (333 ± 13 nm, n = 70) in BHK<sub>hACE2</sub> cells (Fig. 2B-D), supporting a similar endocytic path regardless of hACE2 presence or not.”

      (2) TMPRSS2 and plasma membrane fusion.

      Although the authors allude to membrane fusion as an alternate mechanism of entry, their mechanistic experiments do not address the roles of HS and hACE2 in this process, possibly because their BHK and other cells do not co-express significant levels of TMPRSS2. While many Omicron variants preferentially enter cells via endocytosis (relative to antecedent strains in the pandemic) because of spike mutations that reduce cleavage by TMPRSS2 (PMID: 35104837, 36625591, 35145066), plasma membrane fusion can still occur. The authors should add experiments with co-expression of TMPRSS2/hACE2 [with or without HS] and earlier SARS-CoV-2 variants to establish the role of HS in plasma membrane fusion. Also, are there differences in entry pathways if viruses are prepared in cells expressing TMPRSS2?

      We thank the reviewer for this important comment and agree that our mechanistic experiments were not designed to address TMPRSS2-supported plasma membrane fusion. The reviewer’s suggestion for direct testing of HS function in TMPRSS2-supported plasma membrane fusion, including hACE2/TMPRSS2 co-expression and comparison with earlier SARS-CoV-2 variants, will require dedicated experiments and detection of the fusion pathway that we have not yet designed. It is beyond the scope of the present work. In the revised manuscript, we clarify that the observed ACE2-independent uptake and the predominance of endocytic entry refer to the tested cell systems and do not exclude TMPRSS2-dependent plasma membrane fusion in other cell types, as in the following.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Experiments with hACE2 in cells lacking HS.

      Apart from drug treatment (heparinases or pixantrone) studies shown, the current studies do not directly address whether expression of hACE2 on human cells can allow for endocytosis and productive infection in the complete and genetic absence of HS. The only experiments that use genetically deficient cells are the CHO [hamster] cell studies, and these cells lack hACE2 expression. The authors should knock out a key HS biosynthesis gene (e.g., B4GALT7) in more relevant human cells (e.g., A549-hACE2; ideally with sorted subpopulations having different levels of surface hACE2 expression) and assess endocytosis and infection. This is important given studies in the literature by others suggesting that KO of HS expression reduces but does not abrogate SARS-CoV-2 infection.

      We thank the reviewer for these comments. We showed that virus endocytosis is independent of hACE2 (Fig. 1). The reviewer’s question is whether hACE2 alone can allow for endocytosis of viruses. We have shown that in either BHK (without hACE2) or BHK<sub>hACE2</sub> cells (BHK cells expressed with hACE2), heparinase I/II/III mixture (HPRase) nearly abolished cell-surface immunolabelled HS (Fig. 3D), reduced cell-surface virus attachment by ~83-85% (Fig. 3E), reduced viral uptake by ~80% (Fig. 3F, 3G). These results suggest that hACE2 is not essential for viral attachment and endocytosis. We did not test whether hACE2 alone (without HS) plays a minor role for viral attachment and endocytosis, because to our knowledge, HS is present in nearly every cell. Under this physiological condition, it is HS, not hACE2, that plays an essential role in viral cell-surface attachment and endocytosis. In the revised manuscript, we added a sentence admitting that we did not test whether hACE2 alone is sufficient to support viral uptake and productive infection, as in the following.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis.

      (4) hACE2 function in entry.

      In many places, the authors suggest that hACE2-spike functional interaction occurs "downstream" of HS-dependent binding and endocytosis (e.g., lines 25, 33, 48, 210, 309, 312, 318, 333, 346). However, in their model, it is not clear where exactly this interaction occurs. Are the authors suggesting that this spike binds hACE2 on the cell surface, but this has nothing to do with endocytosis, or that the interaction with hACE2 is occurring at a post-entry step? Can they experimentally demonstrate the stage at which hACE2 is functioning? Is it the same or different in cells lacking HS? What about when TMPRSS2 is present?

      We showed that viral attachment and endocytosis are independent of hACE2, whereas entry as determined by viral gene expression, depends on hACE2. We also showed that most virions bind to HS, not hACE on the cell surface. Based on these results, we propose a model that hACE2 functions downstream of virion endocytosis. We cannot rule out the possibility that a small subset of viruses can also bind to hACE2 after their binding with HS at the cell surface.

      We have not been able to design an experiment to visualize hACE2 mediated virion fusion in endosomes, where hACE2 may facilitate virus fusion and delivery of viral genomes to the cytosol. Productive infection requires only a limited number of successful virion–hACE2 engagement events. While many internalized virions can be visualized, the specific virion or vesicle that ultimately gives rise to productive infection cannot be identified from the present imaging data. This makes it difficult to trace the precise stage or compartment in which the functionally relevant spike–hACE2 interaction occurs. In the revised manuscript, we added a paragraph discussing this limitation as below.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis. Although our data suggest that hACE2 functions downstream of endocytosis to facilitate viral fusion at the endosome for genome delivery to the cytosol, we do not know whether hACE2 binding with the virus occurs at the cell surface or endosomes. The binding may occur in both places, but not essential for virus attachment and endocytosis.”

      (5) Other comments.

      (a) Figure 1A. "Antibody" is misspelled.

      Corrected. Thank you.

      (b) The imaging experiments with pseudoviruses and authentic viruses lack any information on the multiplicity of infection or the number of virions added per cell. If this is particularly high and non-physiological (e.g., >100), is it possible that such conditions might enable viruses to enter [dominantly] through secondary [non-infectious] pathways?

      To address the reviewer’s concern, we used flow cytometry to measure cell-associated VSV-S signal as we diluted the virus by ~600-fold. We found that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHK<sub>hACE2</sub> cells over a ~600-fold dilution of the virus (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations. In the revised manuscript, we included the following sentence and Fig. S7 (Supplementary Information).

      “Flow cytometry also showed that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHKhACE2 cells over a ~600-fold dilution of the virus concentration (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations.”

      (c) Figure 1C and elsewhere. Most of the internalization studies rely on various imaging modalities to demonstrate the pseudovirus or virus on or in the cell. The experiments would be strengthened by inclusion of data from orthogonal binding/internalization assays that measuring virion-associated viral RNA on the surface [4oC binding assay] or inside the cell [after a 37oC temperature shift and exogenous proteinase K and RNAse A treatment]) - such assays can be performed at much lower MOI (e.g., <1, addressed comment #2 above) an also allow more objective quantitation and kinetic analyses of virus internalization (e.g., 0, 5, 15, 30 min at 37oC).

      We demonstrate virion attachment and uptake using multiple approaches, including confocal, STED, and EM analysis, showing virions with the expected morphology at the cell surface and in the cytosol. Furthermore, flow cytometric analysis provides population-level quantitation supporting the same overall conclusion. Thus, while we appreciate and agree that an RNA-based binding/internalization assay would provide additional information, we do not consider it essential to the main conclusion of this work.

      (d) Figure 2. (i) Is there any indication of which vesicles the bath dye is in? Is most of this fluid taken up by micropinocytosis? Are these the same vesicles where the virus that is destined for productive infection (HS/hACE2 engaging) transits? (ii) In all panels, can the authors clearly indicate/label which cells are being used (BHK or BHK-hACE2)? (iii) For the studies with dynasore or dominant-negative dynamin-2-K44A, the readout is at 24 h, a late timepoint, which also could affect virus egress and spread. Can the studies be repeated at much earlier time points (e.g., 15 min to 2 h) to demonstrate that viruses are internalized via dynamin-dependent endocytosis in these cells?

      (i) The bath dye A490 was used as a fluid-phase marker for endocytic uptake, rather than as a marker for a specific vesicle class or intracellular compartment. In principle, any vesicle that takes up extracellular fluid could become labelled by this approach. Since nearly all viruses are in the A490-containing vesicles, productive virus infection must come from some of these vesicles.

      (ii) In the revised Fig. 2 legends, we explicitly indicate which cells are used for each panel.

      (iii) To address the reviewer’s concern, we examined earlier time points for dynasore treatment and found that the virus uptake and genome expression were already markedly reduced at 1 h and 8 h after virus incubation. In the revised manuscript, we described these results as below and in Fig. S5.

      “Fourth, dynasore or dominant-negative dynamin 2-K44A overexpression, which inhibits fission of dynamin-dependent endocytosis [28-30], substantially reduced V-A647 internalized 1-24 h after viral incubation (Figs. 2F-G, S5).

      In addition to inhibiting V-A647 endocytosis, dynasore or dynamin 2-K44A inhibited V-EGFP expression 8-24 h after virus incubation by ~66-77% (Figs. 2F-G, S5), suggesting that endocytosis is the main route for viral genome expression.”

      (e) Line 225. "Envelop" should be "envelope".

      Corrected, thank you.

      (f) Line 235. The authors should clarify that they conclude that the "Omicron variant" of SARS-CoV-2 enters "BHK" cells indistinguishably from VSV-S.

      Thank you for pointing this out. We have rephrased the conclusion as “…omicron variant of SARS-CoV-2 enters BHK cells indistinguishably to VSV-S.”

      (g) Line 278. What happens to virus binding if the authors ectopically express hACE2 in CHO-K1 WT and CHO-pgsA-745 cells?

      We did not perform this experiment (see also our response to major comment 3 above).

      (h) Lines 280-281 and elsewhere (line 635). The authors state "pixantrone (PIX), a drug under clinical trial that binds HS to inhibit HS binding with proteins...." The authors should clarify that the drug is under clinical evaluation for cancer treatment because of its DNA intercalating activity (and not its HS binding activity) and cite any relevant ongoing trials. Also, in line 635, is reference #46 correct?

      As suggested, we modified this sentence as “pixantrone (PIX), a drug under clinical trial for cancer treatment due to its DNA intercalating activity, which can bind HS to inhibit HS binding with proteins”

      (i) Line 281-282. The authors should confirm in a Supplementary Figure that the anti-hACE2 antibody used blocks SARS-CoV-2-JN.1 binding to ACE2.

      In Figure 7D, we showed that PIX and anti-hACE2 antibody block SARS-CoV-2-JN.1 infection, suggesting that anti-hACE2 blocks SARS-CoV-2-JN.1 binding with hACE2.

      (j) Figure 7B. Can hACE2 co-localization be added to this panel?

      We did not perform this experiment. We addressed the role of ACE2 in these airway cells in subsequent panels of Fig. 7.

      (k) Figure 7C. The quantitative data show a 50% reduction in binding signal with pixantrone, whereas the microscopy images appear to show a much greater effect. Can more representative images be shown so that the data better corresponds?

      A ~50% effect is not as visually obvious as the current Fig. 7C. Therefore, we chose not to change the images. However, the statistics in Fig. 7C (right) clearly indicate an average effect of about 50%, as the reviewer pointed out.

      (l) In the Discussion, it is not necessary to use Figure callouts (as done in the Results). Please remove, with the exception of reference to the model.

      We prefer to call out Figures in the Discussion so that we can remind the readers where to find the data. The readers may choose to neglect these callouts. But some readers may read most the discussion part without going through the results carefully. In this case, the figure callouts may help these readers.

      (m) Please delete all references to "new" or "novel" models. It is unnecessary.

      As the reviewer suggested, we deleted “new” and “novel” throughout the revised manuscript.

      (n) Figure legends. Please make sure each panel indicates the # of independent experiments performed. This is included for some but not all panels. Also, a few panels use an unpaired t-test where an ANOVA with multiple comparisons is required (e.g., Figure 1G and S1).

      We agree that, for the three-group sub-comparisons shown within Fig. 1G and Fig. S1, the relevant analyses should account for multiple comparisons. In the revised manuscript, we therefore analyzed these predefined three-group subsets using ordinary one-way ANOVA followed by Dunnett’s multiple-comparisons test, with BHK or Vero used as the reference group as appropriate. The two-group comparisons were analyzed using unpaired two-tailed t-tests.

      Reviewer #2 (Recommendations for the authors):

      (1) It is well established that ACE2 is the receptor for SARS-CoV-2. The authors should not downplay this by saying it is "widely assumed", "typically thought", etc. The specific molecular details at various stages of entry (i.e, the role of HS) remain a bit unclear, but it is disingenuous to imply ACE2 is not the bona fide receptor by any conventional definition.

      The present work does not challenge the well-established view that ACE2 is the receptor for SARS-CoV-2 entry/infection, but suggests that HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, whereas ACE2 acts downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined. We made this point clearer throughout the revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docked at the HS clusters.

      As the reviewer suggested, we removed “assumed” and “typical” and clarify that our findings do not challenge this concept. For example, we modified the abstract

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely assumed as the receptor.”

      as

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely considered the receptor for cell-surface attachment and subsequent cell entry.”

      (2) When the authors state pseudovirus internalization is independent of ACE2, they should clarify that this is the case in cells not expressing TMPRSS2. Most physiologically relevant cell types express TMPRSS2, which will facilitate entry at the plasma membrane.

      As the reviewer suggested, we included the following sentence in the Discussion section: “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Line 130: "Endocytic internalization is the main viral infection pathway" and Line 180-181 is not precise and should be rephrased to include the cell types described in the figure. This may be true in BHK-ACE2 cells, but the evidence in this section does not show that this is universally or broadly true.

      We agree and have revised these sentences to limit the conclusions to the experimental context directly supported by our data. Specifically, our results support endocytic uptake as the major route leading to pseudovirus genome expression in the pseudovirus assays and cell types examined here, rather than as a universal entry mechanism for SARS-CoV-2 across cell types. We have therefore modified the subsection title and the relevant sentence in the Results to explicitly refer to the tested cells/assays.

      Across the revised manuscript, we have accordingly revised the text to distinguish initial virion docking/attachment from productive entry, to acknowledge ACE2 as the established receptor for productive infection, and to limit our mechanistic conclusions to the cellular systems directly tested here.

      (4) All bar plots should show individual dots (i.e., Figure 1G) to better reveal the variance of each dataset.

      While we respect the reviewer’s suggestion, this is not required in the journal style. We prefer plotting bar graphs without individual data points, which often makes it difficult to see the mean values.

      (5) Line 57: This is not accurate. HIV uses a receptor and a co-receptor, for instance.

      We thank the reviewer for noting this inaccuracy. We agree that viral entry frequently involves coordinated engagement of multiple host factors rather than a single receptor, for example, HIV requires both a primary receptor and a co-receptor. We have revised the statement in the Introduction (Line 57–58) to reflect that entry can involve receptors together with co-receptors and/or attachment factors, which collectively facilitate membrane fusion or endocytic uptake.

      In the Introduction (Line 57), we replaced the sentence with “Viral entry is often initiated by engagement of host receptors and associated co-factors that together facilitate subsequent viral membrane penetration.”

      (6) Line 60: "most" --> "many"

      As suggested, we have changed “most” to “many”.

      (7) Remove "clinically relevant" in reference JN.1, as JN.1 is not circulating currently. A more appropriate term could be "full-length" or "authentic", or "wild-type".

      As suggested, we changed it to “authentic”.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors omit to mention the work of Zhang et al, 2023 Nature Comms. "Host heparan sulfate promotes ACE2 super-cluster assembly and enhances SARS-CoV-2-associated syncytium formation". These authors also use PIXN and MTN compounds and define different mechanisms based on ACE2 clustering for virus entry. The authors should mention this work in the Discussion and try to reconcile the different findings.

      As suggested, we include the following discussion in the revised manuscript.

      “Consistent with this possibility, HS may promote spike-dependent ACE2 super-cluster assembly at the cell surface and enhance SARS-CoV-2–associated syncytium formation, suggesting that HS may organize ACE2 nanoscale architecture in a cell–cell fusion context [43].”

      (2) The authors should strengthen their case for the validity of HS-virion interactions as a therapeutic target by mentioning studies showing effectiveness of interference with HS-Covid interactions by heparin and other investigational drugs eg. first study to demonstrate heparin inhibition of SARS CoV2 attachment, Mycroft-West et al, Thromb Haemostatis, 2020; recent report of successful clinical trials of nebulized heparin, The Lancet, Sept 2025; and the superior efficacy of HS mimetic Pixatimod/PG545 compared to heparin (Guimond et al 2022 ACS Chemical Sciences).

      We thank the reviewer for this suggestion and add the following paragraph with citations the reviewer mentioned in the Discussion section.

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding.”

      (3) Figure 1a: incorrect label for antibody.

      Corrected, thank you.

      (4) Some misspellings noted in the manuscript, e.g., MINFLLUX, so please recheck the manuscript for typos.

      We have rechecked the manuscript and corrected the typos.

    1. eLife Assessment

      This important study investigates the mechanisms by which Mycobacterium tuberculosis suppresses protective Th17 differentiation during infection. Using ESX-1- and PDIM-deficient mutants of M. tuberculosis, which lack functional eccC1 and fadD28 respectively, the authors demonstrate that these virulence factors actively restrict IL-17 responses via a T-bet-dependent pathway, independent of overall bacterial attenuation or infection duration. While the precise molecular interactions by which ESX-1 and PDIM modulate Th17 differentiation warrant further elucidation, the experiments are rigorously designed, the findings are clearly presented, and the evidence is solid for a specific role for these factors in limiting protective IL-17 immunity in M. tuberculosis infection.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript examines the factors that restrict the induction of IL-17-producing T cells during Mycobacterium tuberculosis (Mtb) infection. The authors show that neither infectious route, nor duration of infection are responsible. But they do show that mice that lack the Th1-defining transcription factor, a finding consistent with prior reports in the field of immunology. They also show that 2 highly attenuated Mtb mutants in ESX-1 and PDIM, two well-known Mtb virulence factors, do induce IL-17 producing T cells. In contrast, Mtb mutants in mmpl4 are also similarly attenuated, but do not induce IL-17-producing T cells, suggesting that this property is not simply a result of attenuation but due to specific properties of ESX-1 and PDIM-deficient mutants.

      Strengths:

      (1) It is interesting that mice infected with ESX-1 and PDIM mutants have increased induction of Th17 cells.

      (2) Data is solid and convincing throughout.

      Weaknesses:

      There are two main criticisms:

      (1) B6 mice, compared to humans are known to be very Th1 skewed and the Th1 transcription factor T-bet is known to be a strong inhibitor of Th17 responses. Thus, these Th17 inhibitory factors may be stronger in B6 mice than humans, as many humans do make Th17 responses to Mtb infection.

      (2) The molecular insights about how Th17 induction is somewhat limited. Tbet induction is known to restrict Th17 development and this is a t cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. It is not clear these factors related to each other in restricting Th17 induction.

      Additional points:

      (1) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." One consideration, however, is that ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but have different CFUs in the lung-draining LN where T cell priming occurs? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction. Thus, without examining the LN, some questions remain regarding the conclusion that the altered Th17 response in the attenuated strains is not due to the attenuation itself. However, I agree the CFU in the LN probably reflects that in the lung, and if so, the author's conclusions would be sound.

      (2) Do LN cDC1 and high levels of IL-12 p35 manifest in mice infected with the mmpl4 mutant? Likewise do LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb). If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than to those infected with ESX-1 and PDIM mutants. These experiments would help provide more convincing evidence that the identified mechanisms are due specifically to outcomes regarding Th17 induction. However, I agree the author's conclusions are the most likely explanation given the current data.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors tackle an important question of why IL-17 production and TH17 responses are lower than expected during Mtb infection. The authors identify an axis of cross-regulation between TH1 and TH17 cells and provide data to support roles for Mtb virulence factors ESX1 and PDIM in promoting TH1 responses and/or suppressing TH17 responses.

      Strengths:

      The strengths include the significance of the work, the combination of host and Mtb genetic models to dissect the mechanistic basis for regulation of IL-17 production from T cells during infection, and the rigor of the experiments. There are a number of exciting findings from the work, including the cross talk between T cell responses and the impact of ESX1 and PDIM on these responses. It is particularly striking that that IL17a deficient mice partially rescue the attenuation of ESX-1 and PDIM mutants.

      Comments on revised version.

      The revised manuscript has tempered a lot of the language in the original text to more accurately state (and not overstate) interpretations of the data. The claim that the effect is independent of route of infection seems a little too large of a claim when only two routes were tested (aerosol and intranasal). And although the authors revised the results section to acknowledge the contribution of an IFNg-dependent suppression of IL-17 production from T cells, the abstract has not been updated and still claims that all effects are independent of IFNg.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Zilinskas et al seeks to understand the mechanisms underlying the ability of Mtb to suppress Th17 differentiation. As Th17 responses are needed for protective immunity against TB, this is an important topic of investigation. They use Mtb mutants that lack eccC1 (from ESX-1 locus) and fadD28 (encoding PDIM) and implicate a Tbet-dependent pathway by which Mtb modulates Th17 differentiation. The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is, however, unclear, which limits the novelty of the results.

      Strengths:

      Understanding how Mtb limits Th17 differentiation has implications for vaccine development. Comparative study of KO mice and Mtb mutants is a strength.

      Weaknesses:

      (1) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may reflect a difference in timing, since only 3-week data are shown; earlier and later time points would provide a better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, neutrophil data are important for clarifying the mechanism underlying the CFU outcomes.

      (2) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Fig 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6 etc) in the Tbet KO experiments.

      (3) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Fig 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      There are two main criticisms:

      (1) It is not clear how much the factors uncovered here are true beyond B6 mice. B6 mice, compared to humans, are known to be very Th1-skewed, and Tbet is a strong inhibitor of Th17-specific T cells. Many people make IL-17-producing T cells in response to Mtb infection.

      We appreciate the point that not all findings in mice are directly translatable to humans. The B6 mouse is widely used as a model organism for tuberculosis due to its tractability and the wealth of genetic tools available for this strain. While it is true that many individuals do produce Th17 cells after infection with Mtb, humans are still very Th1-dominant, and not all infected individuals produce Th17 cells. We can speculate that the mechanisms outlined in this paper may contribute to the reasons that Th17 responses are not more robust in humans, a finding that may be useful in guiding vaccine design in the future.

      (2) Very few novel insights are mechanistically revealed about how Th17 induction is restricted by Mtb. Tbet induction is known to restrict Th17 development, and this is a T-cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. How, if at all, are these factors related to each other in restricting Th17 induction? Also, the conclusion that it is not a result of attenuation is not completely convincing.

      While it is established that Th1 differentiation can inhibit Th17 differentiation, we believe that rigorously demonstrating this genetically in the context of Mtb infection remains important. Moreover, it cannot be assumed that IL-17 elicited by dampening the Th1 response can lead to enhanced control of infection. We view addressing this as a significant contribution. Furthermore, we also show that the ESX-1 and PDIM virulence factors are functionally linked by suppression of IL-17 responses. The effect is unlikely to be simply due to attenuation of the strains as an equally attenuated control strain does not elicit Th17 cells. We believe that these insights are both novel and important for understanding immune responses to Mtb.

      Other points:

      (1) The authors show that mice infected with a deficiency in ESX-1 have more IL-17-producing CD4 T cells in response to stimulation with an ESAT-6 peptide pool (Figure 3B). Because ESAT-6 is encoded by ESX-1, why do mice infected with this Mtb mutant have any ESAT-6-specific T cells? Is it an incomplete knockdown?

      The ESX-1 knock-out M. tuberculosis Erdman strain is a ΔEccC1 mutant. This strain can produce Esat-6 but cannot secrete Esat-6 out of the bacterial cell. Thus Esat-6 protein is present and able to be processed for MHC-II presentation. We also use Ag85b peptide pool stimulation and report similar effects as Esat-6 peptide pool stimulation.

      (2) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." I don't think the last sentence is necessarily true. I can imagine a scenario in which the induction of the Th17s is, in fact, due to the attenuation, and the Th17 induction still contributes to protection.

      We tested another attenuated M. tuberculosis strain with no known relationship with ESX-1 or PDIM, ΔMmpL4. This attenuated mutant fails to induce IL-17A–producing CD4 T cells to the same extent as observed in mice infected with ESX-1-deficient or PDIM-deficient strains, which is strong evidence that simple attenuation of virulence does not result in higher numbers of Th17 cells being elicited.

      (3) ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but what about the LN? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction.

      We acknowledge that this is a formal possibility, however we maintain that the phenotype is specific to ESX and PDIM mutants, rather than MmpL4 mutants. Even if this phenotype arises from a tissue-specific attenuation of ESX/PDIM mutants, it remains a specific phenotype of these mutants, and not all attenuated mutants, albeit less directly. More importantly, the observation that these mutants induce higher levels of the Th17-polarizing cytokine IL-23 from infected cells ex vivo suggests that this is not an indirect phenomenon.

      (4) Do LN cDC1 and high levels of IL-12 p35 in mice infected with the mmpl4 mutant? Likewise, LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb)? If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than those infected with ESX-1 and PDIM mutants.

      Because ΔMmpL4 and complemented strains resulted in T cell profiles that were not different from the wild-type, we did not measure mediastinal lymph node dendritic cell expression of IL-12 p35 and IL-23 p19 in infections with these mutants.

      (5) ESX-1 and PDIM are very different virulence factors - a protein secretory pathway and cell wall lipid, respectively? Mechanistically, how would mutants in these pathways give very similar outcomes regarding Th17 cells unless it was simply as an aspect of their attenuation? Perhaps, mmpl4 mutants simply differ in some aspects of their attenuation, such as bacterial burdens in LNs, or their interaction with cDCs?

      We are not the first to link phenotypes of ESX-1 and PDIM. Both systems have both been shown to be important for M. tuberculosis permeabilization of the host cell phagosome after phagocytosis, and for suppression of type I IFN responses, among other responses. Thus, these seemingly different virulence factors clearly work together to support specific virulence traits during infection. The exact mechanism of how ESX-1 and PDIM interact is not completely understood and is an area for future investigation.

      Reviewer #2 (Public review):

      The following conclusions and interpretations should be revisited, rephrased, and re-evaluated:

      (1) The manuscript neglects to analyze T cell responses in the dLN, which is the critical site where these responses are initiated (only DC cytokine production is measured in the dLN). The differences in the lungs could reflect trafficking of T cells to the lungs, local lung T cell responses, or durability of the T cell responses in the lungs. The authors state in the last results section that "These results indicate that the ESX-1 and PDIM virulence factors impact naïve T cell differentiation at the draining mediastinal lymph node..." but T cell responses are never measured in the dLN.

      Due to the limited size of the mediastinal lymph node at 3 weeks post infection, we were unable to obtain enough cells for both myeloid cell analysis and T cell analysis, as we perform staining for these panels separately due to the decrease in viability of myeloid cells observed during T cell restimulation. In addition, because T cells in the lung are the population of cells most critical for mediating the outcome of infection, we believe analyzing the T cell response in the lymph nodes though interesting, is not crucial for this study. We have edited the manuscript to be clearer, as suggested by the reviewer.

      (2) Figure 2: The authors state that "Importantly, IFN-γ deficient mice did not exhibit elevated levels of IL-17A producing CD4 T cells demonstrating that IFN-γ production is not the mechanism by which Th1 T cells limit a Th17 response during Mtb infection", but the difference is significantly different and even more obvious in Panel B. In fact, if the Panel D y-axis was on a log scale, the Ifng-/- would likely look more like Tbet-/- than WT. Based on this data, it seems like IFNg is having an effect and should not be completely discounted. Does the deletion of Ifng affect the number of Tbet+ T cells?

      We agree that the IFN-γ<sup>-/-</sup> have only 5x more IL-17 producing CD4 T cells than WT mice while Tbet<sup>-/-</sup>mice exhibit a 25-fold increase compared to WT. We have added this information to the text, and now point out that IFN-γ production is not the sole mechanism by which Th1 T cells limit a Th17 response during Mtb infection.

      In addition, the deletion of Tbet results in an increased number of IFNg+IL-17+ double positive T cells (Figure 2B), in addition to a sizable IFNg single positive T cell population maintained in the Tbet-/- mice (10x the negative control of Ifng-/-). Is this why Tbet deletion is not as severe as Ifng deletion, because T cells are still making IFNg?

      It is possible that the residual IFN-γ produced by T-bet-deficient animals contributes to their relatively modest susceptibility to infection. However, our data show that deletion of IL-17 in this background renders T-bet–deficient mice nearly as susceptible as IFN-γ deficient mice, arguing that the remaining IFN-γ is not a major protective factor.

      Along these lines, the statement in the text that, "Tbet-/-Il17a-/- mice completely lacked both IFN-γ producing...." T cells is not supported by the data in Figure 2C. Tbet-/-Il17a-/- mice look to have more gamma-producing T cells than Tbet-/- mice (which is already 10x the negative control of Ifng-/- in panel 2B if one includes the gamma single positive and IFNg/IL-17 double positive).

      We have amended the language in the text to be more consistent with the data.

      (3) In the Results sections describing Figures 3, 4, and 5, the authors equate IL-17 production by T cells with TH17 responses and IFNg expression with TH1, but Tbet and RORgt expression in the T cells should be measured to make conclusions about TH1 and TH17. Or the authors can rephrase their findings to specifically state the observations as IFNg or IL-17 expressing CD4+ T cells.

      We believe that calling a CD4 T cell in the lung that is producing IFN-γ (and not IL-17) a Th1 cell is appropriate. Potentially confounding cells include those which also produce IL17, which we have ruled out, or T<sub>FH</sub> cells that may be common in lymph nodes but are not common in lungs at this time point and under these conditions.

      (4) Conceptually, do the authors think that ESX1/PDIM promotes TH1 responses and this blocks TH17 or are ESX1/PDIM blocking TH17 responses directly, allowing for increased TH1 responses? It would be helpful to clarify the model in this regard, describe how the data supports one model or the other, and then make sure the language is consistent throughout. Can these effects on T cell responses be tested and recapitulated in vitro using infected APC and T cell co-cultures?

      While it is possible that PDIM and ESAT-6 suppress Th17 through promotion of Th1 differentiation, we do not have data to support this model currently. However, we have added a comment making this point to the discussion.

      Reviewer #3 (Public review):

      Weaknesses:

      (1) The authors should acknowledge and reference key findings from the literature that have identified suppression of Th17 differentiation as an Mtb virulence mechanism, e.g., the role of the Hip1 protease and CD40 signaling (Madan-Lala JI 2014, Sia Plos Path 2017, Enriquez iScience 2022) and Khader JI 2005, showing the requirement of IL-23 for Th17 responses in vivo in a TB mouse model.

      We thank the reviewer for pointing these references out and have added them to the discussion section of the manuscript.

      (2) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may be due to a difference in timing, since only 3-week data are shown; earlier and later time points would provide better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, data on neutrophils are important for clarifying the mechanism for the CFU outcomes.

      We agree that it is surprising that, in the context of TB, Th17 responses are protective whereas excessive neutrophil recruitment is detrimental to the host. In IFN-γ–deficient mice, neutrophils are recruited and contribute to the increased susceptibility of this strain (PMID: 21967766). In separate work from our lab, we have shown that the phenotype of neutrophils recruited to the lungs during Mtb infection influences disease outcome (PMID: 40937719). It is possible that differences in the host environment and the timing of the response shape the effects of neutrophils on the host; these and the other questions raised by the reviewer will be the subject of future studies.

      (3) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Figure 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6, etc.) in the Tbet KO experiments.

      While these are interesting points, investigating mechanisms of Tbet-dependent suppression of IL-17 is beyond the scope of this study.

      (4) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Figure 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants. Additionally, what was the level of Type I IFNs elicited by these mutants?

      We included the MmpL4 knockout Mtb Erdman strain as a control to ensure that attenuation of mutants is not the cause of the increase in IL-17. We also showed that eliminating type I IFN signaling by deleting its receptor has minimal impact on Th17 differentiation, even in the context of a host that produces excess type I IFN. Therefore we do not believe that type I IFN elicited by these mutants is explanatory for the phenotype.

      (5) Since macrophages have been implicated in the reduced cytokines seen in the ESX-1 mutant, IL-23 and other cytokine data on lung macrophages would complement the DC data.

      Because dendritic cells are primarily responsible for priming CD4 T cell responses, we believe that this result in macrophages would not substantially alter our conclusions. That said, it was demonstrated previously that macrophages infected with ESX-1 mutants produce less IL-12p40, a subunit of IL-23.

      (6) Figure 5. There are many fewer DCs overall in the eccC1 and fadD28 mutant groups, which could account for the increased % IL-23p19 in DCs (5D). What were the levels of IL-23 in DC1s?

      The amount of IL-23 p19+ in type I conventional dendritic cells (cDC1s) was near zero as shown in supplementary figure 6A. cDC1s are known to not express IL-23 p19 in mice.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What do the authors mean by the alternative secretion part of "ESX1 Type VII alternative secretion system" that they refer to?

      Bacterial alternative secretion systems facilitate the export of proteins from the bacterial cell independent of the canonical Sec-dependent secretion system required for export of most secreted bacterial proteins across the inner membrane. However to avoid confusion, we have removed the word alternative.

      (2) Not sure naïve fits in this sentence at the end of the introduction: "Furthermore, we observe a strong Th17 response during infection with ΔESX-1 or PDIM lacking Mtb in naïve mice....".

      We have removed the word Naïve.

      (3) Figure legend for 1A-C says analysis performed at 21 dpi, but the figure shows the time course.

      We have corrected this error.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1 should show the non-stimulated flow plot.

      We have added the unstimulated samples.

      (2) The % IL-17 in the flow plots is not consistent across Figures 1, 2, and 3. Not sure why the scales for the Y-axis for IL-17 differ so much between Figures 1 and 2/3. IS there a technical issue with compensation?

      We did not experience any difficulties with compensations. These experiments were done over several years of work. For every experiment, new single-color controls were used and gating was done with the FMO gating strategy. Minor variation such as we see here is not surprising.

      (3) Discuss Yeh et al J Neuroimmunol 2014- show that IFNγ inhibits Th17 differentiation and function via Tbet-dependent and Tbet-independent mechanisms.

      We have added this reference to the manuscript.

    1. eLife Assessment

      This useful study reports findings that support the use of the Open Field Test in Drosophila as a model to study "emotion-like states", which are behavioral responses to several stressful or aversive treatments, and resilience upon their subsequent removal. Behavioral data, by employing established stress-causing treatments and genetic manipulations, are solid. The main advance of this work over previous Drosophila work using a similar experimental setup is systematic analysis of various stressors.

    2. Reviewer #1 (Public review):

      Summary:

      Animal behavior is continuously influenced by the internal state moment by moment, including emotion primitives as the authors pointed out. Although emotion is a more human-related state, evolutional conservation is undeniable, which can be inferred by the behavioral manifestation. To further elaborate the neuronal mechanisms of emotion primitives, the simplest behavioral parameter related to emotional primitives should be well characterized. In this study, the authors described in detail of wall-following behavior (WAFO) and the total walking distance (TOWA) using flies after subjecting them to various conditions or flies being genetically manipulated according to the previous reports that could affect emotion primitives. Overall, the study is well designed and structured. In addition, the discussion on emotion primitives will be of value to the field.

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      Conceptual concerns:

      My suggestion to strengthen the authors' conclusion that "TOWA can be interpreted as a behavioral proxy for exogenously induced arousal" was to show that an increase in TOWA after stress exposure can be observed in a small (1-cm) arena that acts as an exogenous arousing stimulus, but not in a larger arena (>6.6 cm) where such arousing effects are absent. This comparison would demonstrate that basal locomotor activity measured in larger arenas is not altered by stress, whereas the additional component observed in smaller arenas reflects stress-induced internal state. Therefore, the authors would be able to distinguish clearly the effects of stressors or experiences on either simple locomotion or an emotion-like internal state. Then the future works can follow this protocol using smaller and larger arena to assess emotion-like internal state.

      I appreciate the significant authors' efforts to monitor TOWA using arenas with different diameters up to 6 cm. However, the conclusion was unfortunately the same as that obtained using the 1-cm arena. As the authors commented, flies do not show persistent and quantifiable wall-following in arenas larger than 5.8 cm, which limits further examination of this question. I personally agree with the authors' interpretation, but I hope that the authors obtain more definitive experimental contrasts to support this claim in the future study.

    3. Reviewer #2 (Public review):

      Summary:

      In terms of data, the revised manuscript is by and large the same, though the authors added new experiments examining the effects of the Open Field Test (OFT) arena size (Fig. 1 Supplement 1) and sex and mating status (Fig. 6). The authors have also provided textual revisions, partially addressing my previous major criticism about novelty over the work of Mohammad et al., 2016, Curr Biol. They argue that the main advance is the systematic inclusion of Total Walking (TOWA) data (e.g. Introduction, page 6 top in the tracked changes document). While I am still not convinced that the findings represent a huge leap forward over that previous work, the authors' systematic analysis is very nice and may prove of use to those seeking to develop Drosophila as a model for studying emotion primitives.

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stress-like emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      Weaknesses:

      The authors have addressed some of my previous points with textual revisions and in their rebuttal. As stated above, the conceptual advance over Mohammad et al. remains in my opinion limited, but I appreciate that this point is now clearly discussed in the manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      We appreciate this assessment and fully agree that the simplicity of the assay and the demonstration of its scalability and reproducibility are a strength.

      Weaknesses:

      The weakness of the study is the lack of further experiments to support their assumption related to TOWA. The authors suggested that TOWA can be interpreted as a behavioral proxy for exogenously induced arousal. However, it could be interpreted as higher activity, although the authors argued that the circadian clock increasing locomotor activity around ZT0 and ZT12 does not affect TOWA, and therefore TOWA is not related to the locomotor activity per se. As the author cited, flies lose locomotor activity in the circular arena of 6.6 cm in diameter, whereas they continuously move during a 1-h recording in the authors' arena of 1 cm in diameter.

      I would agree that the arena of 1 cm in diameter, but not 6.6 cm in diameter, serves as an exogenous stimulus inducing arousal, and TOWA is manifested by arousal. However, TOWA would also be affected by other behavioral parameters, including the activity, motivation for exploration, or perception of the space. Therefore, it could be reasonable to re-examine some of the flies tested in this study in the circular arena of 6.6 cm in diameter. If arousal is biased by the components presented in Figure 6 and TOWA can assess mainly exogenously induced arousal, the treatment altering TOWA in the arena of 1 cm in diameter would not affect their behavior in the arena of 6.6 cm in diameter. My concern is that Figure 6 may demonstrate too simplistic a diagram to interpret the results. I would suggest adding the experiments using the arena of 6.6 cm diameter or softening the argument.

      We are grateful that you prompted us to investigate the relation between TOWA and arousal and different arena diameters in more depth. Based on your comments, we compared naïve and stressed behaviour between arena diameters of 1, 2.2 and 5.8 cm. The sizes were chosen as to optimally comply with the camera field of view in our setup. Naïve flies showed stable locomotor activity throughout the 60 min of recordings in the different arenas (new Figure 1 – figure supplement 1 A-B’’). Moreover, no significant difference in TOWA over the first 10 min was found between the different arena sizes (new Figure 1 – figure supplement 1 C). This suggests to us that arenas with a diameter smaller than 6.0 cm (and not only with a 1 cm diameter) induce some form of activity that resembles stimulated activity as defined by Meehan and Wilson (1987). A mechanical shock (shake) resulted in significantly increased WAFO in all arena sizes (new Figure 1 – figure supplement 1 D-D’’). We did not test the effects in larger arenas > 5.8 cm, as we found that flies do not longer show persistent and quantifiable wall following. We started to see that also in few flies in the 5.8 cm arena – these few flies were excluded from our analysis presented in new Figure 1 – figure supplement 1. Unlike the naïve response, the stress-induced response in WAFO and TOWA appears to be transient (new Figure 1 – figure supplement 2), which is in line with the definition of emotions as a transient state.

      Reviewer #2 (Public review):

      Summary:

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stresslike emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      We are glad about this assessment and share your opinion.

      Weaknesses:

      The conceptual advance of this research is unclear, as previous work (Mohammad et al., 2016, Curr Biol.) carried out similar treatments and manipulations and reached largely similar conclusions.

      Thank you very much for bringing this up. We rewrote respective parts of the introduction (second last paragraph) and discussion to more clearly outline the advances over the previous work by Mohammad et al. 2016. While our study builds upon Mohammad et al. 2016, the conceptual advance and novelty is that we constitute and treat TOWA as a second and independent dimension equal to WAFO in the OFT. Mohammad et al. had measured locomotor activity (reported as average speed in their paper (total distance walked/time of recording), but primarily to test the dependency of WAFO on locomotor activity. They found that WAFO metrics were poorly correlated with average walking speed, showing a significant degree of independence of both measures – a finding that our results confirm. However, unlike us, they did not consider average speed/TOWA further for their analysis, possibly because they focused on anxiety-like behaviour while our study looked broader on emotion-like behaviour in general. We further used a round (not square) arena to exclude “cornering” in order to reduce the complexity of the assay, which may explain differences of observed speed/TOWA between our studies.

      Moreover, while WAFO is a good proxy for 'stress', I am not convinced that TOWA necessarily represents an emotional state in all cases. Indeed, as the authors themselves acknowledge, changes in total walking may be associated with other factors, such as starvation-induced hyperactivity, physical exhaustion after sleep deprivation, increased sex drive after mating, alcohol sedation, etc.

      Your comment raises a question in comparative research on emotions which is very difficult if not impossible to conclusively answer. At first sight, the most conservative stance seems to be to completely disregard the idea of emotions and affective experiences in animals. This, however, would mean that we cannot use animal models to study the basics of emotions (= emotion primitives) and would ignore that by all likelihood emotions are a product of evolution and hence should exist at least in more basic forms in animals. Obviously, we have no means to ask flies or any other animal whether they connect “hunger” or “mating” to a feeling or an emotional state (which must not be conscious) but can only observe the behaviour. We here adopt the often-cited “Pankseppian” view (based on the book of Jaak Panksepp: “Affective Neuroscience”) and firmly believe that – in order to fully understand how the brain drives behaviour- we also need to take affective states into account that bias behaviour towards adaptive responses.

      In short, we are unfortunately unable to give a clear and definite answer to your comment whether TOWA represents an emotional state in all cases. Perhaps you are right. We believe, however, that a “Pankseppian” view is adequate, and we may ask what evidence exists that shows that starvation-induced hyperactivity or post-mating is not associated with an affective emotion-like state in the fly or any other animal.

      Another unclear point is the interpretation of some unexpected results, such as the finding that both serotonin transporter overexpression and its knockdown give the same phenotype.

      Thank you very much for this comment, which we also received by reviewer #1. As suggested by the other reviewer, a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT may be a differential effect on the different serotonin receptor subtypes expressed in the brain. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – both higher and lower than optimum levels result in behavioural impairments in adults (see e.g., Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Finally, there are some issues with the use of the OFT in rodent research (e.g., inconsistent effects of anxiolytic drugs; see Rosso et al., 2022, Neurosci Biobehav Rev., for a meta-analysis). These should be explained to place the Drosophila findings in their appropriate context.

      Thank you very much for bringing this systematic review to our attention which assessed the usefulness of various behavioural tests including the OFT to study the effect of anxiolytic drugs in rodents. Overall, the review casts “serious doubt on both construct and predictive validity” of behavioural tests for anxiolytics. While diazepam (the only drug used in our study) was the drug with the most consistent effects across the analysed behavioural assays, only 59% of the OFTs revealed significant effects. We were already aware of earlier findings in the same direction (Prut and Belzung 2003 10.1016/s0014-2999(03)01272-x), but as we only used one drug did not include a discussion in the manuscript. Unfortunately, the number of studies employing the OFT in flies is very small and does not yet allow for a similar comparison. We now changed the respective sentences in the discussion:

      “In rodents, diazepam mostly but not consistently leads to an anxiolytic response in the OFT behaviour which questions the usefulness of the OFT for testing anxiolytic drugs (see (Prut and Belzung 2003; Rosso et al. 2022)).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Overexpression of SerT suppressed increased WAFO after electric shocks. However, knockdown of SerT did similar. It is worth reporting this finding, but possible mechanisms could be mentioned in the manuscript. For example, given that flies carry five serotonin receptor genes, upregulation of overall serotonin level may affect the specific serotonin receptor, but downregulation of it may affect other receptors, leading to unexpected outcomes. Exploring any other possibilities could be better to add to guide future research.

      Thank you very much for this comment and the suggestion of a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – higher or lower levels result in behavioural impairments in adults (see e.g. Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Reviewer #2 (Recommendations for the authors):

      (1) The advance over Mohammad et al., 2016, Curr Biol. must be clearly and emphatically articulated in the Introduction and/or Discussion. It is otherwise impossible to appreciate what is conceptually novel about this work.

      We have now rewritten parts of the two last paragraphs in the introduction to make the advances clearer. As outlined above, we conceptually advanced the analysis of OFT behaviour by integrating TOWA as a second and independent dimension in our analysis. During the revision process, we have spent great effort to better characterise the nature of the locomotor activity encountered in the OFT (Figure 1 – supplementary figures 1 and 2). Also, this is an advancement over previous studies, including Mohammad et al. 2016. We further included new results on the general effect of neuropeptides (silver mutants, impaired in neuropeptide processing).

      (2) The most important metric for a stress-like emotional primitive is WAFO. Therefore, figures should be revised in a way that highlights WAFO differences. Figure 4 is a good example of this. In contrast, in Figures 1-3, the WAFO box-and-whisker plots are very small and obscured under the raw WAFO-TOWA plots, which are difficult to see (especially given the light blue background of all plots) and redundant. I strongly recommend just showing the WAFO box-and-whisker plots for the sake of visibility, clarity, and brevity.

      Although we understand your reasoning, we would like to stick with the old figures as we consider TOWA as important to characterise the OFT response as WAFO (see our comments above regarding the conceptual advances of our study over Mohammad et al. 2016). It is true that Figures 1-3 are small, but at least in our print-out well legible. Further, it appears that in the current version of the eLife system the resolution is downsampled. In addition, we anticipate that in the version of record figures can be enlarged online as in other eLife articles.

      (3) Given the criticisms against the OFT in rodent research, one of which is inconsistent effects of anxiolytic drugs (Rosso et al., 2022, Neurosci Biobehav Rev.), it may be useful to expand pharmacological treatments beyond diazepam.

      As our focus is not on the testing of anxiolytic drugs and since it was already very difficult to be granted access to diazepam (we are not at a medical institution), we refrained from testing further drugs. Moreover, as rightfully mentioned by you, the OFT may not be the best test for the efficacy of anxiolytic drugs. On the other hand, diazepam was the most consistent anxiolytic in the OFT in rodents (see Rosso et al. 2022).

      (4) The authors should show results of the effects of at least some stressors/punishments on WAFO/TOWA of female flies to understand if observed effects are sex-specific.

      Thank you very much for bringing this topic to our attention. To test whether the effects are sex-specific, we now performed several new experiments. First, we compared the naïve OFT response of mated and unmated males and females (see new Figure 6). This revealed that without prior stress treatment, the WAFO response is independent of sex and mating status. In contrast, the naïve TOWA response turned up to be sex- and mating state-specific (see new Figure 6). To test whether the mated females show a different stress-induced OFT response to males, we applied mechanical stress (shake) that we had also used to assess the effects of arena diameter (Figure 1 – supplementary Fig. 1 D-D’’). After a first round of shaking, females showed increased TOWA, but WAFO was unaffected. A second round of shaking, however, led to a significant increase in WAFO and TOWA. This suggests that the OFT response is qualitatively similar between the sexes and mating status, yet the threshold for elicited responses differs between males and females. We now added a respective paragraph to the main text in the results section plus a new figure (Fig. 6).

    1. eLife Assessment

      This important study addresses a timely issue at the intersection of mitochondrial and telomere biology by focusing on the relationship between naturally occurring variants in the mitochondrial genome and telomere length. This work thereby provides a conceptual and experimental framework for investigating communication between mitochondria and telomeres. Using an innovative transmitochondrial cybrid approach, the authors provide evidence that mitochondrial DNA variants influence telomere maintenance through effects on mitochondrial function, reactive oxygen species, and NAD⁺-dependent repair processes. The evidence supporting the central conclusion that mitochondrial genotype influences telomere-associated phenotypes is convincing and is strengthened by the use of complementary functional and rescue experiments. However, some of the mechanistic interpretations and broader conclusions regarding telomere length inheritance in humans would benefit from additional donors and longitudinal analyses following cybrid generation, or more cautious framing.

    2. Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result. Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes? In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

    4. Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R²=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R²=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      We appreciate the reviewer raising the important link between pathogenic POT1 variants and lymphoid malignancies driven by elongated telomeres. We would like to clarify, however, that our introduction focused not on the pathological telomere attrition characteristic of telomere biology disorders, but rather on the natural, non-pathological variations observed in individuals with baseline telomere lengths on the shorter end of the spectrum.

      Nevertheless, we will include this observation regarding patients with long telomeres in the introduction to underscore the complex relationship between telomere length and tumorigenesis.

      The two following publications by the Armanios lab will be added:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      Because this study investigated the potential maternal inheritance of TL via the mitochondrial genome, our cohort consisted predominantly of female donors (6/7). Donor #5, the husband of Donor #7, was the only male included. We will update Figure 1D to include the sex of each donor and clarify this rationale in the text.

      Notably, our analysis revealed minimal influence of sex on TL, which cannot account for the observed differences between the extreme groups.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      We will do so.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result.

      We thank the reviewer for this suggestion. Following their advice, we used our metaphase spread FISH analyses to quantify chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells. However, because Cybrid 2 metaphase spreads were of insufficient quality for adequate chromosome analysis, we will restrict our telomere fusion comments exclusively to Cybrid 1 and include the quantifications in our revised manuscript.

      Author response image 1.

      The number of fusions and total chromosomes analyzed are indicated on each bar.

      Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      While this is a strong argument, the telomeres within these cybrid models are critically short. Consequently, the FISH signal intensity falls below the threshold required for reliable quantification of telomere-free ends.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      The reviewer is correct that we have not formally demonstrated this in our experimental system. Instead, our assumption was based on the fact that the initial phase of Rho0 cell repopulation involves a temporary ROS burst, partly driven by incompletely assembled ETC supercomplexes that are known to elevate ROS levels (Maranzana et al, 2013). This is supported by our experimental observation that the NAC antioxidant, combined with NR, potently inhibits telomere shortening in cybrids with low CI activity. We propose to add this explanation in the revised manuscript.

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes?

      We cannot fully explain the discrepancy, but we indeed suspect in vivo oxidative stress differs significantly from cell cultures (21% O<sub>2</sub>). Additionally, early cybrid replenishment involves an unknown telomere elongation step that may offset mitochondrial ROS-induced shortening.

      To further investigate this question, we measured telomere length across varying population doublings (PDs) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include (as Supplementary figure) and discuss these data in the revised manuscript.

      Author response image 2.

      In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      We excluded Rho0 cells from our superoxide measurements because these cells were grown in a different culture medium. Instead, we focused on comparing mitochondrial ROS across cybrids to accurately correlate these values with the telomere length of the corresponding donors’ lymphocytes. Consequently, Rho0 cell measurements would not have contributed to this correlation analysis.

      We propose to discuss this in the revised manuscript.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

      We thank the reviewer for this constructive feedback and entirely agree that including more donors would have added depth to our findings. While we acknowledge this limitation, pursuing this further is currently impossible without acquiring new ethical approvals and establishing fresh collaborations with clinicians.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      We limited our TL measurements to the earliest viable time point after cybrid formation to avoid the confounding effects of cellular adaptation in culture. For instance, the early telomere shortening observed in Cybrid 1 and Cybrid 2 was later alleviated—likely due to the upregulation of the NAD+ salvage pathway genes NAMPT and NAPRT1.

      To accurately capture the effects specific to cybrid formation, we isolated 4 independent clones per donor, all of which showed highly consistent TL values, as shown in Figure 2A and S3D.

      While we acknowledge the reviewer's point about multi-passage longitudinal analyses, we did, in fact, measure telomere length across varying population doublings (PDs) in Cybrid 3 (high mito ROS levels), 6 and 7 (low mito ROS levels) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include a Supplementary figure and discuss these data in the revised manuscript.

      We further propose to carefully edit the manuscript so as to clarify the text and, whenever possible, add quantifications. Among others, as suggested by Reviewer #1, we will add the quantification of chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells in our revised manuscript (See Author response image 1).

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

      Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R<sup>2</sup>=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      We thank the reviewer for this insightful comment. As suggested, we will explicitly describe the sampling structure upon introducing the correlation analysis. Furthermore, we have conducted the requested leave-one-out analysis:

      - removing Cyb3: R<sup>2</sup>=0.730; p=0.0302

      - removing Cyb6: R<sup>2</sup>=0.782; p=0.0194

      - removing Cyb1: R<sup>2</sup>=0.740; p=0.0278

      While we agree that additional donors would enhance the study, further experiments are however currently impossible without new ethical clearances and additional clinical partnerships.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R<sup>2</sup>=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      We acknowledge that our study does not establish the in vivo role of CI activity in TL regulation. Our abstract specifically highlights an in vitro phenomenon: “Under the specific conditions of cybrid formation, which involve a metabolic shift from glycolysis to oxidative phosphorylation, mtDNA variants associated with reduced CI activity induced rapid telomere shortening, …”.

      We are nevertheless happy to revise the text to clearly separate our in vitro results from in vivo biology as requested.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

      We agree that the evidence for the AT6 A177T inheritance claim remains inconclusive. To clarify, we do not argue that this mitochondrial variant is solely responsible for longer telomeres; indeed, the mtDNA genome of donor #1 suggests otherwise. Furthermore, the phenotypic impact of such mtDNA variants likely depends on nuclear variants in other telomere-related genes (e.g., hTERT or hTR), meaning AT6 A177T may not consistently result in elongated telomeres. Unfortunately, our ethical protocol precludes screening additional unrelated donors. We will revise the text to soften our statements accordingly.

      New references:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      Maranzana E et al. Mitochondrial respiratory supercomplex association limits production of reactive oxygen species from complex I. 2013. Antioxid Redox Signal 19, 1469-1480.

    1. eLife Assessment

      This fundamental manuscript describes a key role for the integrated stress response-regulated transcription factor CHOP in regulating liver biology in response to endoplasmic reticulum stress through both the downregulation of transcription factors involved in regulating hepatic identity and altering the capacity for integrated stress response and unfolded protein response signaling to induce protective signaling. The data supporting this model is convincing, but including some additional discussion on the mechanism and importance of the work in the context of the published literature would be helpful to better define the complex importance of CHOP signaling. This work will be of interest to a wide range of biologists interested in liver biology, stress-responsive signaling, and ER stress.

    2. Reviewer #1 (Public review):

      Summary:

      The predominant view on CHOP's functions during ER stress is that it promotes cell death. This is in contrast to a handful of reports in the literature that claim that CHOP is a positive regulator of protein synthesis during chronic ER stress, and therefore is part of the adaptation program to ER stress. These previous studies were performed in tissue culture cells. Velarde and co-authors have used a mouse model of induction of mild ER stress to study the function of CHOP in hepatocytes.

      Major strengths and weaknesses of the methods and results:

      The authors use state-of-the-art mice to manipulate (i) CHOP and (ii) ATF6, a protective factor of ER proteostasis, and address the hepatocyte responses to mild ER stress in vivo and in cultures. Validated gene expression programs are well correlated to liver pathology in the mouse models. This is a very well-done study.

      The authors clearly show that CHOP transitions hepatocytes under mild ER stress to a chronic ISR state, which is phenocopied by ATF6-depleted hepatocytes. So the conclusion that CHOP exacerbates ER stress in hepatocytes during mild ER stress is correct. It is also clear that CHOP targets negatively the transcription of hepatocyte identity genes, which opens a new direction of studies on the function of CHOP in secretory cells in general.

      Conclusion:

      This is a significant study that will benefit different research fields, and specifically studies on proteostasis, as was recently highlighted in Nat. Str. Mol. Biol. by experts in the field.

      To this reviewer, the importance of the study is that it links the function of a transcription factor (CHOP) to stress intensity (mild versus severe) in a physiological experimental model (hepatocyte function and pathology).

    3. Reviewer #2 (Public review):

      The Unfolded protein response (UPR) and related integrated stress response (ISR) are critical signaling systems for cell survival in response to acute stresses. While the UPR directs critical adaptive gene expression, certain chronic stresses switch this pathway towards cell death and disease. An important question concerns the mechanisms by which the UPR switches from being adaptive to maladaptive. Prevailing models focus on the transcription factor CHOP (DDIT3 or GADD153), whose levels are enhanced via the UPR, and extended/amplified amounts of CHOP are suggested to boost death-related gene expression. However, the literature and this manuscript point out a number of observations that do not neatly fit with this model, suggesting that there are still unresolved processes by which CHOP adjusts cell outcomes via the UPR.

      This manuscript features a nice hepatocyte-targeted knockout of CHOP to discern the contribution of CHOP in the transition between adaptive and maladaptive outcomes. The key ideas presented in this study are that CHOP-directed gene expression is focused on protein synthesis, metabolism, and hepatocyte identity. In the progression of the UPR, CHOP expression can lead to resumption of protein synthesis, which can assist in the translation of the UPR-directed transcriptome, which includes ATF6/XBP1-directed genes that aid the processing capacity of the endoplasmic reticulum (ER). However, enhanced nascent protein can further stress the ER. CHOP directs gene expression in both the first phase- acute and second phase-chronic in the UPR, and the pivotal decision lies in the transition between the phases.

      Overall, the manuscript includes some new ideas as well as refinements of earlier ones for CHOP-determination of UPR-directed cell fate. The CHOP-hepatocyte knockout mouse model helps to delineate the different tissue functions of CHOP, which has been a problem for some earlier studies. The manuscript progression of experiments is solid, and experimental design and documentation are rigorous. The manuscript text is largely clear, but there are portions that would benefit from fuller explanations of ideas.

      There are three points of concern. First, the manuscript model (Figure 7) lays out a timeline for the progression of the UPR between two phases. The study is not always clear about the times assayed, and there appears to be a single time point for measurements. Second, there is emphasis on protein synthesis changes in the model. It is true that the literature argues that resumption of protein synthesis concurrent with stress damage (i.e., GADD34-directed gene expression) is a key reason for the potentially debilitating effects of CHOP (e.g., Marciniak et al 2004, Han et al 2013). However, the manuscript does not feature protein synthesis measurements. Inclusion of bulk protein synthesis measurements in the context of this model system would strengthen the study and support for the model. Finally, for this reviewer, some of the most interesting ideas center on CHOP-directed transcription of genes that regulate hepatocyte identity. There is solid evidence for direct CHOP regulation of these genes, but the manuscript does not really develop and test the ramifications of these networks on cell fate during ER stress.

      Reviewer Concerns:

      (1) The abstract packs in a lot of information. The ideas would not be clear to a general reader. Furthermore, the UPR and ISR are referred to in the second-to-last sentence, but not defined earlier in the abstract.

      (2) There are some typos/grammar concerns.

      (3) ATF4 diminished with CHOP-depletion (Figure S2A). What is the mechanism here? Does this complicate the analysis of CHOP-directed gene expression? How does this fit with Figure 6J? The timelines for TM treatment are critical. The authors should more fully explain the time courses in the experiments.

      (4) Figures 2 and 3: There is a discussion on enhanced protein synthesis with loss of CHOP (reduced GADD34 expression). What is the time point - 8 hours TM? Emphasize, explain, and justify time points of experiments here and in later panels. It would strengthen the model with direct measurements of protein synthesis. The authors could include GADD34 protein measurements in these panels. Figure 3 - panel D - some abbreviations are not standard.

      (5) Figure 4: One of the most interesting in the manuscript is the transcription factors downstream of CHOP that are linked with hepatocyte differentiation and metabolism. The manuscript would be bolstered by developing some of these target genes into the Figure 7 transition model.

      (6) Figure 6: The comparison of CHOP and ATF6 target genes is a highlight of the manuscript. The literature on this topic is complex, and there are some suggestions that CHOP can be downstream of ATF6. Furthermore, there were some earlier models by Walter and others about extended induction of Perk (death) vs induction of other UPR sensors (survival) (e.g. PMID: 17991856). It would be helpful in the Discussion to delineate between these models and their critical differences.

    4. Reviewer #3 (Public review):

      In this manuscript, the authors aim to understand the function of the transcription factor CHOP, which is known to promote cell death during severe stress in the ER. The authors note that CHOP is induced during less severe stress, but its functional output is not well understood in these cases. Here, they study the effects of conditional knockouts of CHOP in hepatocytes of mice challenged with chemical inducers of ER stress.

      Tunicamycin (an ER stress inducer) injection leads to the upregulation of CHOP and lipid accumulation in the liver, but no significant cell death in the experiments outlined here. Conditional knockout of CHOP results in a number of differences in the way hepatocytes respond to stress, notably resulting in lower steatosis.

      There are two main findings supported by the data presented here. First, the authors show that CHOP suppresses the expression of ONECUT, a master regulator of hepatocyte differentiation and metabolism, during ER stress. They show by ChIP-seq that CHOP binds to the promoter region of this gene, and by RNA-seq that ONECUT expression is suppressed by ER stress in a CHOP-dependent manner. Many predicted targets of ONECUT1 were also suppressed by ER stress in a CHOP-dependent manner, though they were not bound directly by CHOP. The data support a model where CHOP down-regulates hepatocyte metabolism and identity via regulation of ONECUT1. This is a new and interesting finding, perhaps explaining the steatosis phenotype of livers that accompanies ER stress, although this was not tested directly.

      The second main finding of this paper is that CHOP deletion leads to an interesting assortment of effects on genes related to the ER stress response and integrated stress response (ISR). As expected, based on prior work, CHOP deletion led to more phosphorylation of eIF2alpha (CHOP is known to upregulate the phosphatase for this translation factor). However, unexpectedly, this did not cause increased expression of ATF4 (a transcription factor whose upregulation during stress is dependent on eIF2alpha phosphorylation) and its downstream targets; in fact, CHOP deletion had the opposite effect on these. In other words, CHOP seems to both turn off the initiating signal for the ISR (namely, eIF2alpha phosphorylation) and also promote the downstream signaling events that rely on this initiating signal. It makes sense that cells would do this, as restoring translation would be important for realizing the effects of the massive changes in gene expression initiated by ER stress, and yet this would exacerbate stress in the short term, so it would be counterproductive to also turn off the entire stress-regulated program. Having a factor (perhaps CHOP) that coordinates these two events makes sense. It will be interesting in future work to understand the mechanisms behind this regulation.

      Finally, CHOP deletion led to less activity of other aspects of the ER stress response, notably IRE1 (determined through measurement of XBP1 splicing and RIDD of Bloc1s1). This is explained by the continued phosphorylation of eIF2alpha in these knockouts, as the continued attenuation of translation would lessen the burden of misfolded proteins in the ER. Somewhat confusingly, the same pattern is not seen in downstream targets of XBP1. Less splicing, coupled with perhaps less translation of the spliced mRNA, should result in less active transcription factor and lower expression of its target genes in the CHOP KO. This is not observed in Figure 2, although the more global gene expression analysis suggests that all stress-dependent gene expression changes were weaker in the CHOP KO livers.

      The authors characterize the effects of CHOP, promoting restoration of protein synthesis and the accompanying exacerbation of stress while preserving the signaling that should relieve ER stress, as a switch from an acute to chronic phase of ER stress. This is mirrored in their analysis of ATF6 in a similar series of experiments. Although this is an interesting framework for thinking about the stress response, whether CHOP is the key factor or a supporting actor in regulating this transition will require a better understanding of the mechanisms involved.

    5. Author response:

      We thank the reviewers for their assessment of our work and their comments. We are grateful for their evaluation of our findings as fundamental and convincingly supported, and for their appreciation of the relative scope of this manuscript and of future work. The most direct requests for new experimental data are from reviewer #2, who asks for direct assessment of the effects of CHOP deletion on expression of GADD34 and on protein synthesis. We agree that these are important experiments to conduct for the revision.

      The reviewers requested more clarity on the experimental logic of the paper and on the place of our findings in the broader context of ER stress signaling, which we will be happy to provide in a revised manuscript. These revisions will include a more explicit consideration of how the regulation of metabolic genes by CHOP contributes to its effects in the liver independently of its role in regulating eIF2a dephosphorylation.

      In particular, there were concerns about the logic of the time points chosen that we feel are important to also address here. For analysis of ChopHKO animals, all experiments were carried out 8 hours after ER stress challenge. This is because, as we show in Fig. 1B and also in our previous paper on CHOP (1), this is the time point at which CHOP expression is at its maximum. Thereafter, hepatocytes become heterogeneous with respect to whether they do or do not express CHOP. This is an interesting finding because it suggests that CHOP is part of a cellular switch, and potentially even an effector of that switch—a point currently raised in the Discussion but worth further highlighting in a revision. At the practical level, it means that discerning the contribution of CHOP to ER stress signaling and adaptation at subsequent time points will require sophisticated single cell analyses that can discriminate cells that express CHOP from cells that do not, which are an important future direction.

      In contrast, for Atf6aHKO animals, all experiments were carried out 48 hours after ER stress challenge. As we have previously shown (2), at short time points after a stress challenge, such as 8 hours, there is very little difference in ER stress signaling between wild-type animals and those lacking ATF6a. The reason for this lack of distinction is that the major targets of ATF6a are ER chaperones and the like. Because adaptation to ER stress in the early phases of the response depends more on non-transcriptional mechanisms such as inhibition of protein synthesis and IRE1-dependent mRNA decay (RIDD), the failure to fully upregulate ATF6a targets is initially of little consequence. It is only at later time points when wild-type animals restore ER homeostasis and largely silence ER stress signaling. In contrast, at these same later points, animals lacking ATF6a show evidence of persistent ER stress, most notably in the form of persistent Xbp1 mRNA splicing and profound suppression of metabolic genes. Although the 8 hour time point for experiments in ChopHKO animals differs from the 48 hour time point for Atf6aHKO animals, the two lines of experimentation are united by the persistence of ongoing ER stress and of ISR signaling despite diminished eIF2a phosphorylation at the points when the presence of CHOP or the absence of ATF6a are of the greatest impact. A revised manuscript will present this logic more clearly.

      References

      (1)  Liu K, et al., EMBO Reports 25, 228 (2024)

      (2)  Rutkowski DT, et al., Dev. Cell 15, 829 (2008)

    1. eLife Assessment

      This important study combines behavioural analysis, voltage imaging and electrophysiology to advance our understanding of muscle coordination at the cell-to-cell level, in Caenorhabditis elegans. The evidence supporting the conclusions is convincing; however, the use of correlation in some aspects of data interpretation is a relative weakness. The technically sophisticated optogenetic voltage clamp approach introduced here can be applied to other small, transparent animals, making these findings of broad interest to researchers studying electrical coupling between cells or utilising optical electrophysiology techniques.

    2. Reviewer #1 (Public review):

      Summary:

      This study aims to reveal the contribution of individual gap junction proteins to the signal transmission and connectivity of living C. elegans animals in a completely non-invasive way through all-optical electrophysiology. The authors achieve this by simultaneous expression of bipoles, an excitatory/inhibitory light-activated actuator and Quasar2, a genetically encoded voltage dye. With this study, the authors extend their previous efforts to leverage the strength of optogenetic neurophysiology and set a new standard in this domain. In addition, they adapted their established methods to perform cell-specific optogenetic voltage clamp and revealed changes in gap junction connectivity. They also find that increasing excitability in innexin mutants is indicative of a reduction in gap-junction connectivity and current leaks.

      Strengths:

      This is an extremely strong manuscript, a technical feat and tour de force to infer junctional coupling through all-optical electrophysiology. The establishment of the voltage clamp method is powerful and allows researchers to obtain not only tight control over voltage signals but also permits the investigation of gap junction function in response to positive and negative voltage steps in a completely non-invasive fashion. This will be a new paradigm for investigating muscle electrophysiology in future.

      Weaknesses:

      This is a strong pioneering study, and I found very few technical weaknesses. The correlation quantification is relatively weak to establish connective causality, as a shared upstream input may lead to a similar perceived correlation. This is especially concerning for an average lag time of ~0, and the authors may want to investigate if there is unchanged connectivity in an unc-31 or unc-13 mutant. Conceptually, the local connectivity is scaled to account for behaviour: future studies may wish to perform this method on moving animals, and in specific neuronal populations, where a non-invasive optogenetic voltage clamp method will truly shine.

    3. Reviewer #2 (Public review):

      Summary:

      This technically sophisticated study combines behavioral analysis, voltage imaging, electrophysiology, and a newly developed cell-specific optogenetic voltage clamp (cOVC) approach to investigate gap-junction (GJ)-mediated coupling in C. elegans body-wall muscle cells. The work explores the coordination of muscle cells and systems physiology and introduces a method with potential utility beyond the nematode system studied.

      Strengths:

      The main strength of the work is the development and application of the cOVC method. This approach enables minimally invasive in vivo assessment of cell-to-cell electrical coupling in intact animals. This technique represents a meaningful advance over traditional electrophysiological techniques that require dissection or cell isolation.

      With respect to the GJ biology and function, the authors support their conclusions by integrating additional independent experimental approaches. Findings from behavioural/locomotion assays, voltage imaging, patch-clamp recordings, and cOVC measurements are generally consistent, particularly for unc-9 mutants, which show reduced synchronization of muscle cells, impaired electrical coupling, and severe locomotor defects.

      The gain-of-function experiment using murine Cx36 suggests that more or less electrical coupling can disrupt (normal) locomotion.

      Weaknesses:

      The main issue of this otherwise excellent manuscript relates to interpretation rather than experimental quality. Throughout the manuscript, increased correlation is often interpreted as evidence of increased electrical coupling. Are correlation, synchrony, and conductance equivalent measures? If not, how would this affect these correlations? Furthermore, could broader action potentials and altered excitability also increase correlation values? This concern could be addressed through a discussion of this limitation.

      Similarly, the proposed mechanism that reduced GJ coupling increases excitability through reduced leak currents is plausible but not directly demonstrated. Are alternative explanations, e.g., compensatory changes in ion-channel expression or gap-junction composition, possible? These could also be considered to improve the balance of this work.

      The conclusions regarding Cx36 overexpression would also benefit from more cautious wording, as developmental or localization effects have not been excluded.

      However, the experimental dataset is very strong. In my opinion, no major additional studies are needed. Direct analysis of compensatory changes in innexin expression or localization could strengthen the interpretation of the proposed mechanism. Overall, the study is of high technical quality, contains a notable methodological advance, and provides important insights into muscle synchronization and GJ biology.

    1. eLife Assessment

      This important study describes how selective lesions of key cortical and subcortical motor areas affect reaching actions in macaque monkeys. The results will be of interest to both basic and clinical researchers studying the neural control of movement. Kinematic analysis of movement quality is solid but could be improved by considering other metrics, especially those that relate to grasping. Evidence for the general claims related to the role of specific motor areas is incomplete because the lesions did not fully eliminate any single area while simultaneously involving neighbouring areas.

    2. Reviewer #1 (Public review):

      Summary:

      This is a very interesting and well-done study of the effects of selective lesions to the sensorimotor cortex and the red nucleus on control of upper limb movements. The findings that the red nucleus may subserve recovery of upper limb motor function after cortical lesions in macaques and the different motor functions of different cortical sensorimotor areas are significant findings of considerable interest to sensorimotor neuroscientists, neurologists and neurosurgeons. The methods are mostly excellent, but there are some questions about the use of endothelin lesions in cortical areas and the use of trajectory variability as a marker of movement quality and fine motor control. Furthermore, it is questionable that increased trajectory variability in reaching a target reflects reduced movement quality, reduced ability to independently control muscles, and is a proximal analog of reduced dexterity.

      Strengths:

      The rationale that rubrospinal projections onto spinal neurons may subserve the good recovery of upper limb movements observed after lesions of sensorimotor cortex is compelling. The methods involving complete lesions of the red nucleus followed by recovery prior to lesions affecting various sensorimotor cortical areas are a strength. The excellent interpretations offered in the Discussion section are also a strength.

      Weaknesses:

      There are weaknesses in the Methods, including:

      (1) no information on dimensions of the cup containing the food reward or types of food rewards,

      (2) recording 3D hand movements with a single camera,

      (3) cortical endothelin lesions were not very precise,

      (4) the use of trajectory variability as a measure of movement quality and reduced ability to independently control muscles.

      Some interpretations presented in the Discussion are not well supported. The discussion related to movement quality should be modified to focus on trajectory variability. The suggestion that rubrospinal projections onto motor neurons are apparently irreplaceable is not well justified because one monkey receiving a complete red nucleus lesion showed nearly full recovery of maximum movement speed, while the other monkey did not. The nearly full recovery of one monkey was probably due to new corticospinal connections onto motor neurons, whereas it is possible that the other monkey would have recovered better given more time before the 2nd lesion to cortical areas.

    3. Reviewer #2 (Public review):

      Summary:

      This study made selective lesions in motor cortical subregions and the magnocellular red nucleus in nine macaque monkeys, and evaluated reaching and grasping movements using maximum speed and trajectory variability. The results suggest that damage to the posterior old primary motor cortex (M1) was mainly associated with reduced maximum speed, whereas damage to the new M1 was mainly associated with increased trajectory variability. Damage to the anterior old M1 did not clearly add further impairment. Lesions of the magnocellular red nucleus (RNm) alone mainly reduced reaching speed, but recovery after subsequent cortical lesions was worse than after cortical lesions alone, suggesting that the rubrospinal pathway may be important for compensation after cortical damage in monkeys. Overall, this is a valuable study that examines differences among M1 subregions using selective lesions in macaques.

      Strengths:

      (1) This study tackles an important question. It attempts to decompose the diverse upper-limb impairments after stroke into the effects of different primary motor cortex subregions.

      (2) Another strength is that the lesions and behavioral impairments were evaluated quantitatively. The use of nine macaque monkeys with different lesion patterns, together with quantitative behavioral evaluation, provides a rare and valuable dataset. The authors also followed recovery using quantitative behavioral measures such as maximum speed and trajectory variability.

      (3) The inclusion of RNm lesions is also valuable, as it revisits the classic question raised by Lawrence and Kuypers (1968) of how brainstem descending pathways contribute to recovery after cortical motor damage.

      Weaknesses:

      (1) The main limitation is that the contribution of each cortical lesion is sometimes interpreted from largely qualitative comparisons. Because the lesion extent was not always limited to the intended region, it is difficult to fully separate the independent contribution of each subregion. Some conclusions are also based on comparisons between a small number of animals. The dataset itself is valuable, and the manuscript would be strengthened by presenting these conclusions more cautiously and explicitly acknowledging this limitation.

      (2) Because the behavioral evaluation is quantitative, it would be helpful to show the relationship between lesion size and behavioral impairment more quantitatively. For example, rank correlations between the lesion size of each cortical region and behavioral measures could help readers evaluate whether the type and size of lesion are related to behavioral impairment.

      (3) The discussion of area 4s could be further developed. The authors suggest that this region may have a different role, but the specific hypothesis is not fully clear. There has also been skepticism in the previous literature about area 4s, for example, Meyers et al. (1954), and this broader background could be discussed. (Meyers R, Knott JR, Skultety FM, Imler R (1954) On the Question as to the Existence of a "4s" Suppressor Mechanism. Journal of Neurosurgery 11:7-23.)

    4. Reviewer #3 (Public review):

      Summary:

      In this article, the authors performed targeted lesions in cortical areas involved in forelimb control of rhesus macaques. Using a reaching task with kinematic tracking, they compared kinematic variability (as a proxy for dexterity) and reaching speed (as a proxy for strength) before and after cortical (n=7) and magnocellular red nucleus (RNm) (n=2) lesion. Changes in these movement metrics were related to the location and extent of the lesions, reconstructed from histology. The authors report that lesions with a large component in New M1 had a pronounced effect on kinematic variability, whereas lesions with a large component in posterior Old M1 primarily affected reaching speed. Lesions of the RNm were performed in two animals approximately seven to eight weeks before the cortical lesion. By themselves, RNm lesions produced a significant but small reduction in reach speed. They also magnified the effect of the subsequent cortical lesions.

      Strengths:

      (1) For non-human primate (NHP) research, this is a large cohort.

      (2) The behavioural analyses are clear and precise.

      (3) The additional red nucleus lesions in two monkeys provide unique complementary information.

      Weaknesses:

      (1) Description of injuries. As described and reported in the current result section and figures, readers do not have a clear understanding of the lesion extent and location.

      (2) Lack of formal correlative analyses between lesion characteristics and behaviour. Currently, it seems that the conclusions are based on impressions between some aspects of the lesions and precise kinematic measures.

      (3) The data will be of interest to a large community of researchers working on brain injury and stroke recovery, as well as cortical motor control. There are, however, some major methodological issues in the current version of the manuscript that prevent a clear evaluation of the findings and potential contribution of the work in the field. These issues need to be addressed to support the conclusions.

    1. eLife Assessment

      The framework of this potentially important study – with the integration of multiple levels of analysis, glymphatic MRI, transcriptomics, functional MRI, and public amyloid maps, in one framework – is clever. The assertion that regional amyloid vulnerability may depend not just on neural activity alone, but on whether clearance is appropriately matched to activity, is an interesting and novel concept. However, the chosen approach to imaging glymphatic clearance relies on indirect inferences from a small subgroup. In its current form, the main conclusions of this study are therefore incompletely supported.

    2. Reviewer #1 (Public review):

      Summary:

      Regional differences in the brain's waste-clearance system may interact with neural activity to influence where amyloid-B accumulates. Using intrathecal GBCA administration to produce "Glymphatic MRI" in 96 subjects, the authors mapped cortical glymphatic influx and clearance and found distinct spatial patterns, with transcriptomic analyses linking better glymphatic function to neuronal cell types (through genes). In a subgroup with resting-state fMRI, regions with stronger resting-state activation generally showed higher contrast clearance, indicating a positive coupling between these processes. Notably, cortical regions where neural activity and glymphatic clearance were mismatched showed greater amyloid-β burden in a separate, publicly available PiB-PET dataset, suggesting that activity-clearance decoupling may contribute to regional vulnerability and neurodegeneration.

      Strengths:

      This is a rare and valuable dataset. Intrathecal contrast injection in ~100 subjects is quite a remarkable accomplishment alone, but the addition of resting-state fMRI, a correlative PiB cohort, and gene-expression pattern data is impressive.

      Weaknesses:

      This is a cross-sectional study, and we can't determine whether neural activity drives glymphatic clearance, whether glymphatic dysfunction alters neural activity, or whether both are shaped by a third factor. Language describing "flow", "influx", and "clearance" could be made more specific so the reader can more easily follow the methodological approach.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Li et al. investigated the relationships among regional cortical tracer dynamics following intrathecal gadolinium administration, neural activity, and amyloid-β deposition in humans. Using serial MRI acquisitions after intrathecal gadodiamide administration in 96 participants, the authors characterized regional signal enhancement and clearance patterns across the human cortex. They integrated these imaging measures with transcriptomic data (Allen Human Brain Atlas), resting-state fMRI outcomes, and an external amyloid PET dataset. The authors report that regions with more efficient tracer clearance are enriched for genes related to synaptic organization and neuronal cell types, that tracer clearance patterns are in parts spatially coupled to spontaneous neural activity, and that regional mismatch between neural activity and tracer clearance is associated with increased amyloid burden according to the PET dataset.

      Strengths:

      The study addresses an important and very timely question about the interaction among neural activity, cerebrospinal fluid dynamics (waste clearance), and regional vulnerability to neurodegeneration. Integrating serial post-contrast MRI, transcriptomics, resting-state fMRI, and amyloid imaging is ambitious and conceptually very interesting. The spatial characterization of cortical tracer dynamics is potentially valuable for the field, particularly given the increasing interest in human glymphatic imaging approaches and intrathecal contrast MRI, which provides an opportunity to assess CSF tracer dynamics without confounding tracer signal from the blood. The imaging preprocessing pipeline includes normalization of regional cortical signal intensity to a reference region within each session before calculation of longitudinal percentage change, which helps reduce inter-session variability within individuals for conventional T1-weighted imaging. The transcriptomic analyses linking tracer dynamics to neuronal and synaptic gene expression patterns are also interesting. In addition, the manuscript addresses recent literature on neurovascular coupling, glymphatic function, and amyloid vulnerability.

      Weaknesses:

      Several issues limit the strength of the conclusions. One concern relates to the interpretation of repeated post-intrathecal contrast MRI measurements as direct indicators of glymphatic influx and clearance. The approach presented by the authors measures regional signal changes following intrathecal gadodiamide administration, but does not directly visualize paravascular flow or establish that the observed signal dynamics specifically reflect glymphatic transport mechanisms. Although it is widely accepted that CSF influx occurs primarily along periarterial spaces as part of the glymphatic system, and the terminology "glymphatic MRI" is increasingly used in the literature, the physiological processes contributing to delayed parenchymal enhancement, including CSF-interstitial exchange mediated by convective bulk flow and/or extracellular diffusion, as well as transient and, in the case of linear gadolinium agents, even long-term tracer retention remain incompletely resolved. Importantly, tracer kinetics may not directly reflect interstitial fluid kinetics, as solute transport may also be influenced by compartmental and extracellular barriers, diffusion constraints, and tissue retention effects. As currently written, several sections of the manuscript appear to overstate what can be directly inferred from the imaging data. This issue may be particularly relevant given the intrathecal use of gadodiamide (Omniscan), a linear gadolinium-based contrast agent with known long-lasting tissue retention due to lower kinetic stability compared to macrocyclic agents. Sustained signal at later imaging time points may therefore not only reflect impaired glymphatic clearance dynamics may also be influenced by tissue retention of contrast material, particularly in the context of neurological disease. In addition, the participant cohort is heterogeneous and includes individuals with neuroinflammatory and neurodegenerative diseases, peripheral neuropathy, and motor neuron disease. Although the authors argue that the spatial tracer patterns are relatively preserved across neurodegenerative groups, this heterogeneity complicates interpretation of imaging data and raises the possibility that disease-related factors and altered tracer-tissue interactions contribute to the observed effects. Thus, the rationale for interpreting a greater tracer signal at 39h as evidence of impaired glymphatic clearance should be explained more carefully, particularly given the highly heterogeneous patient population.

      In addition, the analyses linking spontaneous neural activity and tracer clearance are based on a very small rs-fMRI subgroup (n = 15), limiting the generalizability. The interpretation of the "mismatch" analysis also requires caution. The mismatch index was computed from z-scored fALFF and tracer clearance and is subsequently associated with amyloid burden derived from the external PET dataset rather than from the studied participants themselves. Therefore, the observed spatial associations should be interpreted with greater caution rather than as evidence for a direct mechanistic relationship. The cross-sectional nature of the analyses also limits conclusions regarding the directionality and temporal sequence of the relationships between neural activity, tracer dynamics, and amyloid burden. Several statements in the Discussion currently imply stronger causal or biological conclusions than are directly supported by the data.

      Despite these limitations, the study presents an interesting dataset and proposes a framework for understanding regional vulnerability to protein accumulation in neurodegeneration. This work hopefully motivates further investigation into the important relationships among neural activity, CSF dynamics, and neurodegeneration in humans.

    4. Reviewer #3 (Public review):

      This manuscript addresses an interesting and timely question: whether regional glymphatic clearance in the human cortex is spatially coupled to neural activity and whether a mismatch between activity and clearance may help explain regional vulnerability to amyloid-β deposition. The authors use intrathecal gadolinium-based glymphatic MRI in 96 participants, derive cortical influx and clearance maps, integrate these with Allen Human Brain Atlas transcriptomic data, and then relate regional clearance to resting-state fMRI measures in a smaller subgroup. They further compare the resulting activity-clearance mismatch map with an open-source ¹¹C-PiB amyloid PET dataset. The overall concept is attractive because it attempts to connect glymphatic physiology, neuronal activity, and proteopathy at the regional level of the human brain, an important and understudied area.

      The main strength of the study is the use of direct intrathecal contrast-enhanced MRI to generate cortical maps of glymphatic tracer dynamics. This is a technically demanding approach and provides a richer spatial readout than indirect MRI proxies of glymphatic function. The authors show that the cortical tracer signal increases from 4.5 h to 15 h and then decreases by 39 h, allowing them to interpret the early signal as reflecting influx and the persistent signal at 39 h as impaired clearance. They further identify regional patterns, with faster influx in medial prefrontal/insular areas and slower clearance in dorsal prefrontal and parietal surface regions. The analysis is visually clear, and the use of cortical gradients is a useful way to reduce complex regional data into interpretable spatial axes.

      The multimodal integration is also interesting. The transcriptomic analysis suggests that regions with faster glymphatic clearance are enriched for synaptic organisation and neuronal activity-related pathways, while regions with slower clearance show enrichment for metabolic and mitochondrial pathways. The cell-type enrichment analysis further implicates excitatory and inhibitory neurons, oligodendrocyte lineage cells, microglia and, to a lesser extent, astrocytes. This provides a plausible biological bridge between regional neural activity and clearance function, and the sensitivity analysis using ReHo in addition to fALFF is a useful robustness check.

      However, the manuscript should be more careful in its causal interpretation. The study is cross-sectional and largely correlative in space. The finding that regions with higher spontaneous neural activity tend to show better glymphatic clearance is intriguing, but it does not establish that neural activity drives clearance in these participants. Conversely, it remains possible that better tissue integrity, vascular function, CSF access, cortical geometry, vascular density, or disease composition jointly influence both fMRI measures and tracer clearance. The authors do acknowledge some of these limitations, but the abstract and discussion should more consistently frame the findings as associations rather than evidence of an activity-clearance mechanism in humans.

      The most important limitation is the small size of the fMRI subgroup. Although the whole glymphatic MRI cohort includes 96 participants, the key activity-clearance analysis is based on only 15 individuals, including 11 with peripheral neuropathy and 4 with motor neuron disease. This is a very small and clinically heterogeneous sample on which to build a central conclusion about regional neural activity and glymphatic clearance. The authors show that the 39 h PC map in the fMRI subgroup resembles the whole-cohort map, which is helpful, but this does not address whether the fALFF-clearance relationship is robust at the individual level. The paper would be strengthened by reporting subject-level stability, leave-one-out analyses, and whether the association persists after excluding the four motor neuron disease cases.

      A second major concern is the interpretation of the amyloid analysis. The ¹¹C-PiB map is derived from an external open-source Alzheimer's disease dataset, not from the same participants who underwent glymphatic MRI and fMRI. Therefore, the association between activity-clearance mismatch and amyloid burden is a spatial correspondence across group-average maps, not an individual-level relationship. This is valuable for hypothesis generation, but should not be presented as evidence that a mismatch in the present cohort predicts amyloid deposition. The authors should clearly state that this analysis tests whether mismatch regions overlap with known amyloid-prone cortical regions, rather than directly linking mismatch to amyloidosis in individual participants.

      The definition of "mismatch" also needs clarification. The text defines the mismatch index as the negative absolute difference between z-fALFF and z-39h PC, and states that higher scores indicate greater mismatch. Because the index is negative, values closer to zero would normally indicate a smaller absolute difference rather than a greater mismatch. This should be checked carefully and corrected if necessary. More broadly, because a higher 39 h PC indicates worse clearance, the interpretation of match and mismatch categories is not intuitive. The authors should provide a clearer schematic and ensure that the mathematical definition, biological interpretation and figure labelling are fully aligned.

      Several technical confounds require more attention. Intrathecal gadolinium MRI is influenced by CSF dynamics, posture, sleep, circadian timing, renal clearance, age, intracranial pathology, and potentially diagnosis-specific differences. The authors acquired scans at fixed time points and noted that patients slept as usual, but individual sleep duration, sleep quality, posture, and daytime activity were not objectively measured. Given that the central claim concerns glymphatic clearance, these are not minor confounders. The authors should consider adjusting for age, sex, diagnosis, vascular risk factors, and relevant clinical variables where possible, and be more explicit about how heterogeneous disease indications may influence cortical tracer kinetics.

      The statistics are generally good. However, many correlations are performed across 400 cortical parcels, which are not independent biological samples. The paper would benefit from clearer separation between participant-level inference and region-level spatial inference. For example, the fALFF-clearance and mismatch-amyloid analyses are regional map correlations, not correlations across individuals. This should be clearly stated throughout. The authors should also report effect sizes and confidence intervals more consistently, and explain how multiple comparisons were controlled across transcriptomic, cell-type, fMRI, ReHo and amyloid analyses.

      The transcriptomic analysis is useful but should be presented as indirect. AHBA data come from six post-mortem brains; only the left hemisphere was used, and the donors were healthy and younger than the clinical cohort. Therefore, these data capture intrinsic regional gene-expression patterns rather than disease-state expression in the same individuals. The authors should avoid implying that the transcriptomic findings directly explain glymphatic function in their participants. The current discussion partly acknowledges this, but the framing in the abstract and results could be more cautious.

      There are also several points of presentation that should be improved. The manuscript should consistently distinguish glymphatic influx, glymphatic clearance, CSF tracer retention, and waste clearance. A 39 h residual gadolinium signal is a useful proxy for delayed clearance, but it is not the same as direct measurement of amyloid or tau clearance. The language around "waste clearance" and "amyloidosis" should therefore be precise. The authors should also clarity whether "higher clearance" corresponds to lower 39 h PC across all analyses, as this inversion is easy for readers to misinterpret.

    5. Author response:

      We would like to express our deepest gratitude to the Editors and Reviewers for their highly rigorous and constructive evaluation of our manuscript. We are greatly encouraged by the recognition of our study’s ambition, the unique value of the in vivo intrathecal contrast MRI dataset, and the conceptual novelty of linking macroscopic glymphatic physiology with neural activity and regional proteopathy.

      We fully agree with the thoughtful limitations and methodological concerns raised in the eLife Assessment and the Public Reviews. In our upcoming revised manuscript, we are implementing a comprehensive set of revisions to address these points. Specifically, our planned revisions focus on the following key areas:

      - Tempering Causal Interpretations: We agree that our cross-sectional design precludes definitive causal inferences. We are systematically revising the manuscript to soften causal language (e.g., replacing "drives" with "is spatially associated with"). We will explicitly frame our findings as macroscopic spatial associations and discuss the potential influence of joint physiological confounders.

      - Tightening Terminology and Imaging Physics: We are refining our terminology to more accurately reflect our MRI measurements. We will replace assertive terms like "direct glymphatic flow" with precise descriptors such as "imaging proxies for tracer enhancement and retention." Furthermore, we are expanding the Limitations section to explicitly acknowledge the confounding effects of Partial Volume Averaging (PVE), systemic tracer redistribution, and renal clearance kinetics.

      - Conducting Supplementary Imaging & Robustness Analyses: To address concerns regarding cohort heterogeneity and the sample size of the rs-fMRI subgroup (n=15), we are performing a series of rigorous supplementary analyses. This includes conducting sensitivity analyses (e.g., excluding the motor neuron disease subgroup) and applying leave-one-out cross-validation to rigorously assess the subject-level stability and robustness of the spatial coupling between neural activity and tracer clearance.

      - Clarifying the Conceptual Model and "Mismatch" Index: To improve readability, we are moving the anatomical definitions of the cortical gradients directly into the Results section. Additionally, we are introducing schematic diagram to intuitively explain the mathematical formulation and biological interpretation of the "activity-clearance mismatch" index.

      - Re-framing External Dataset Analyses: We are carefully re-framing the interpretations of the Allen Human Brain Atlas (AHBA) transcriptomic data and the external PiB-PET amyloid dataset, emphasizing that these reflect spatial correspondences of intrinsic regional vulnerability across groups, rather than individual-level direct interactions.

      We believe these revisions will significantly enhance the scientific rigor, clarity, and precision of our study.

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      We thank the reviewer for the kind comments.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      We have added more details on the annotation process and have included the list of instructions sent to the annotators. We have also included mmpose configs with the code provided. The following files include the relevant details:

      File including the list of instructions sent to the annotators:

      OpenMonkeyWild Photograph Rubric.pdf

      Mmpose configs:

      i) TopDownOAPDataset.py

      ii) animal_oap_dataset.py

      iii) init.py

      iv) hrnet_w48_oap_256x192_full.py

      Anaconda environment files:

      i) OpenApePose.yml

      ii) requirements.txt

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths:

      Open source dataset with excellent annotations on the format, as well as example code provided for working with it.

      Properties of the dataset are mostly well described.

      Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones?

      Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out).

      The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      We thank the reviewer for the kind comments.

      Weaknesses:

      More details on training hyperparameters used (preferably full config if trained via mmpose).

      We have now included mmpose configs and anaconda environment files that allow researchers to use the dataset with specific versions of mmpose and other packages we trained our models with. The list of files is provided above.

      Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010).

      We have included a datasheet for our dataset in the appendix lines 621-764.

      Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here.

      We have included the list of instructions sent to the Hive annotators in the supplementary materials. File: OpenMonkeyWild Photograph Rubric.pdf

      Should include model cards, as described in Mitchell et al (arXiv:1810.03993).

      We have included a model card for the included model in the results section line 359. See Author response image 1:

      Author response image 1.

      It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year.

      We agree that the source could introduce structural biases. This is why we included images from so many different sources and captured images at different times from the same source—in hopes that a large variety of background and lighting conditions are represented. However, doing so limits our ability to document each source background and lighting condition separately.

      Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically.

      The focus of this work is on overall keypoint localization accuracy and hence we wanted a metric that is easy to interpret and implement, in this case we made use of PCK (Percentage of Correct Keypoints). PCK is a simple and widely used metric that measures the percentage of correctly localized keypoints within a certain distance threshold from their corresponding groundtruth keypoints.

      A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

      We have now included a histogram of unnormalized bounding boxes in the manuscript, see Author response image 2:

      Author response image 2.

      Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential.

      We thank the reviewer for the kind comments.

      However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      We thank the reviewer for pointing this out. Inter-observer reliability is important for ensuring the quality of the annotations. We first used Amazon MTurk to crowd source annotations and found that the inter-observer reliability and the annotation quality was poor. This was the reason for choosing a commercial service such as Hive AI. As the crowd sourcing and quality control are managed by Hive through their internal procedures, we do not have access to data that can allow us to assess inter-observer reliability. However, the annotation quality was assessed by first author ND through manual inspections of the annotations visualized on all of the images the database. Additionally, our ablation experiments with high out of sample performances further vaildate the quality of the annotations.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      Our goal was to obtain as many images as possible from the most commonly studied ape species. In order to ensure a large enough database, we focused only on the species and combined images from as many sources as possible to reach our goal of ~10,000 images per species. With the wide range of people involved in obtaining the images, we could not ensure that all the photographers had the necessary expertise to differentiate individuals and subspecies of the subjects they were photographing. We could only ensure that the right species was being photographed. Hence, we cannot include more detailed information.

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      We created the training set to reflect the species composition of the whole dataset, but used test sets balanced by species. This was done to give a sense of the performance of a model that could be trained with the entire dataset, that does not have the species fully balanced. We believe that researchers interested in training models using this dataset for behavior tracking applications would use the entire dataset to fully leverage the variation in the dataset. However, for those interested in training models with balanced species, we provide an annotation file with all the images included, which would allow researchers to create their own training and test sets that meet their specific needs. We have added this justification in the manuscript to guide the other users with different needs. Lines 530-534: “We did not balance our training set for the species as we wanted to utilize the full variation in the dataset and assess models trained with the proportion of species as reflected in the dataset. We provide annotations including the entire dataset to allow others to make create their own training/validation/test sets that suit their needs.”

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

      We thank the reviewer for this important suggestion. Since we can separate a video into its constituent frames, one can indeed use the provided model or other models trained using this dataset for inference on the frames, thus allowing video tracking applications. We now include a short video clip of a chimpanzee with inferences from the provided model visualized in the supplementary materials.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Please provide a more thorough description of the annotation procedure (i.e., the instructions given to crowd workers)! See public review for reference on dataset annotation reporting cards.

      We have included the list of instructions for Hive annotators in the supplementary materials.

      An estimate of the crowd worker accuracy and variability would be super valuable!

      While we agree that this is useful, we do not have access to Hive internal data on crowd worker IDs that could allow us to estimate these metrics. Furthermore, we assessed each image manually to ensure good annotation quality.

      In the methods section it is reported that images were discarded because they were either too blurry, small, or highly occluded. Further quantification could be provided. How many images were discarded per species?

      It’s not really clear to us why this is interesting or important. We used a large number of photographers and annotators, some of whom gave a high ratio of great images; some of whom gave a poor ratio. But it’s not clear what those ratios tell us.

      Placing the numerical values at the end of the bars would make the graphs more readable in Figures 4 and 5.

      We thank the reviewer for this suggestion. While we agree that this can help, we do not have space to include the number in a font size that would be readable. Smaller font sizes that are likely to fit may not be readable for all readers. We have included the numerical values in the main text in the results section for those interested and hope that the figures provide a qualitative sense of the results to the readers.

    2. Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential. However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

    3. Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths: - Open source dataset with excellent annotations on the format, as well as example code provided for working with it. - Properties of the dataset are mostly well described. - Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones? - Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out). - The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      Weaknesses: - More details on training hyperparameters used (preferably full config if trained via mmpose). - Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010). - Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here. - Should include model cards, as described in Mitchell et al (arXiv:1810.03993). - It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year. - Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically. - A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

    4. eLife Assessment

      The OpenApePose dataset presented in this manuscript represents an important contribution to the field of primate behaviour and computer-vision science with methodological applications that are sure to be applicable for a wide variety of taxa. The analysis supporting the utility of this database is solid and compelling but would benefit from some additional clarity, particularly with regards to the annotation of landmarks, model parameters and division of the dataset for training, validation and testing.

    1. eLife Assessment

      This study addresses an important question in aging biology by combining metabolomics, transcriptomics, molecular genetics, and functional analyses to examine how cytosolic acetyl-CoA metabolism influences late-life fitness in replicatively aging yeast. The evidence supporting the roles of AMPK activation, mitochondrial acetyl-CoA utilization, and fatty acid synthesis in preserving fitness during aging is convincing overall, and the engineered A2A strain provides an elegant demonstration that coordinated modulation of distinct acetyl-CoA metabolic branches can increase the proportion of aged cells with a low-senescence phenotype. The study provides significant insight into mechanisms that allow aging cells to maintain fitness without extending replicative lifespan.

    2. Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fuelled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. The reviewers have achieved their aims and the results support their conclusions. This work provides important insight into molecular mechanisms that allow aging without loss of fitness and will be of interest to scientists in the field of aging and metabolism.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells. Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metabolic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Comments on revised version.

      I am fine with the revisions.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fueled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. This work will be of interest to scientists in the field of aging and metabolism. Some clarifications in the text would address the following concerns, which would increase the impact of the study:

      (1) What does activation of AMPK (via PGDP-Sak1 expression) do to the replicative lifespan? How many bud scars, in general, do the subpopulations that are older - yet have less Tom70 (increased mitochondrial fitness) - have, after the 48 hrs timepoint that they are examining? How many divisions occurred in this 48hr time period - i.e. is it long enough to have all cells reach the end of their replicative lifespan? This information is important to rule out that a subset of the mutant cells just divided faster and hence had more divisions within 48 hrs (growing faster and living longer are different things). Having identical growth curves doesn't indicate per se that they all divide at the same rate, as there may be a subpopulation that divides faster and a subpopulation that doesn't grow so well.

      Increasing AMPK activity increases replicative lifespan [PMID: 25869125], but given our finding that AMPK activation splits the population, such replicative lifespan assays are hard to interpret. Bud scar counts have a similar issue. Hence we restricted the lifespan and bud scar analyses to wt and A2A which are more homogenous (Figures S2 B and E). A2A cells at 48 h have ~25% more bud scars than wt cells. Yes, by 48 h most of the cells have lost viability (Figure 2E). The reviewer is correct that you can't properly compare the lifespan curves if the cells divide at different rates, hence our follow-up test of wt at 48 h vs A2A at 40 h viability after we had confirmed that these time points captured cells at equivalent replicative ages (Figure 2D, E). This shows that viability of A2A is slightly lower than wt at matched age, indicating a slightly shorter lifespan. 

      (2) A2A cells do not have an extended replicative lifespan (RLS) but show an increase in the "low senescence" population (Figure 2). If the cells are not becoming senescent, why don't they have longer RLS? Not having a longer lifespan seems inconsistent with the statement that "bud scar counting confirmed that A2A cells reach a higher age than wild type", which comes back to how many times the cells can divide in the 48hr timepoint studied and their rate of cell division? Also, the lifespan curve shown is plotted against time, not cell division number, which does not take into account different division times of cells within the population (described above). It would be much more useful to show standard lifespan curves showing cell division numbers per lifespan per cell.

      Our observation that cells can reach the end of life without senescing is consistent with other studies that have studied the life course of individual cells by microscopy [PMID: 31291577, 32675375]. These studies always highlight some proportion of the cells that reach the end of life with no or minimal senescence, though this fraction varies with the experimental system. The question of why cells lose viability without senescing is a complete unknown in the field, and reflects a wider lack of consensus as to why yeast lose viability with replicative age.

      In liquid culture we can only assess viability over time, not cell division number, which we agree is not optimal and we are wary about making strong statements on lifespan for exactly the reasons the reviewer notes. Unfortunately, it is clear from the comparison of liquid and solid media lifespans performed by the Gottschling lab [PMID: 19652178] that culture system has a huge effect on lifespan, with cells in classical plate-based microdissection assays living far longer than the same strains do in liquid. This means that lifespans determined by microdissection-based assays are of questionable relevance to ageing studies performed in liquid culture. Senescence cannot be assayed on plates, while microfluidic systems lack the throughput necessary and preclude key techniques like RNA-seq, so liquid culture assays were the only option for this work. We agree that this leaves an unsatisfactory approximation for lifespan measurements, but we consider it critical that everything is measured in the same system. We therefore restricted our conclusion on lifespan to simply say that lifespan of A2A cells is not extended which our data in Figures 2D, E, S2B does support (see also answer to Q1), and therefore with the majority of A2A cells showing low senescence marks and high fitness at 48 h we can conclude that lifespan and fitness loss must be separable.

      We have added a note of these limitations of lifespan measurements in the materials and methods section of the manuscript.

      (3) Increased "fitness" of the old cells is implied from the increased size of the colonies that the old cells can make. However, this is a measure of the fitness of the daughters per se, not the old mother cells. Are the old mothers just passing on healthier mitochondria and more lipids to the daughters, such that they can divide more times? If the aged cells have an "increased fitness", why don't they divide more times themselves (i.e. live longer?).

      Yes, colony growth speed is defined by daughter cell replication, but as long as the daughters and subsequent generations divide at the same rate irrespective of whether they come from a young or old mothers then the size of the colony after 24 hours varies based on the time it took the initial mother to produce a daughter. This is what the assay really measures. We note that aged wildtype mothers often do not divide at all in the first 24 hours after being put on an agar plate (hence the tiny reported colony size), even though they do eventually produce a daughter which then forms a colony, whereas A2A cells tend to produce the first daughter rapidly whether young or old. It is known that daughters of aged wildtype mothers also divide slower, as to some extent do grand-daughters (PMID: 2644196), which will also contribute to differences in colony size, and this may well result from a lipid and/or mitochondrial contribution, but the primary driver of colony size in 24 hours is the time the mother took to initially divide. We have added this detail to the materials and methods section of the manuscript.

      As noted above, the mechanistic basis of lifespan is unknown, but although senescence can shorten lifespan, our work and that of others shows that lifespan is still limited in the absence of senescence.

      (4) The statement is made that "these experiments define two classes of aging cells with distinct metabolic needs, coherent with the model of two aging trajectories previously proposed (referencing Nan Hao's work)". However, the big difference here is that in Nan Hao's work, their two aging trajectories influenced the length of lifespan, but that does not appear to be the case here. That distinction should be made clear. Perhaps the authors could also speculate as to why the A2A yeast stops dividing after presumably the same number of cell divisions, even though they have an activated AMPK and activated fatty acid synthesis pathway.

      Yes, this is a good point and we have added this distinction to the Discussion:

      “Here we have characterised two classes of ageing cells seemingly differentiated by high and low availability of cytosolic Acetyl-CoA, consistent with a previous demonstration that ageing follows two trajectories in yeast though it should be noted that in this previous report, the two trajectories also differed in replicative lifespan (6).”

      We would love to speculate on why the A2A cells don't have an extended lifespan, but at this point we don't have a strong hypothesis. We have come up with many theories for this, but none that we haven’t managed to disprove experimentally. One thing worth considering is that many cells which lose replicative viability in liquid culture and probably in plate assays remain intact – for example, DNA and RNA integrity is not compromised over 24- 48 h – so those cells are probably not dead per se. But we also detect apoptosis-sized DNA fragments, which must come from dead cells, so there is clearly not a single mechanism defining the end of replicative lifespan.

      (5) I am a bit confused by the use of the word "senescence" by this lab here and in their previous growth on galactose studies. If yeast don't senesce, which is usually defined as an irreversible arrest of the cell cycle where cells stop dividing, shouldn't the yeast that do not senesce still be dividing and hence have a longer lifespan? Should a different term be used rather than senescence? Such as "fitness late in life". The authors giving their definition of senescence may help reduce this apparent contradiction.

      We completely agree, this is confusing and noted this distinction in the Introduction. Use of the term senescence to mean a loss of fitness late in life in yeast stems from the classical definition of senescence as applied to whole organisms. However, the term senescence as applied to cells has a more specific meaning in terms of the cell cycle as the reviewer notes. As an individual S. cerevisiae is both a cell and an organism, the terminology clashes. However, the marker we largely employ (Tom70-GFP) which in our hands is a very good proxy for fitness was originally defined as marking the senescence entry point (SEP), so overall we feel we can't avoid the term.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells.

      Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metab to olic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Weaknesses:

      (1) While the manuscript focuses on mitochondrial transport and fatty acid synthesis, cytosolic acetyl-CoA is also a key regulator of histone acetylation and chromatin silencing. It would strengthen the study to consider whether acetyl-CoA depletion contributes to improved fitness through enhanced rDNA silencing. Given the well-established role of rDNA instability in yeast aging, additional experiments examining rDNA silencing and stability would be valuable. For example, monitoring rDNA copy number changes (not necessarily ERCs) under AMPK activation, oleic acid supplementation, and in the A2A strain, similar to approaches used in the authors' prior work, would help clarify whether chromatin regulation contributes to the observed phenotypes.

      We have added data addressing these points to the manuscript and Supplemental Figures 2, 3 and 4, though the outcomes are complex. Histone acetylation changes chromatin accessibility and could therefore alter global gene expression; in accord with this, RNA-seq shows that P<sub>GPD</sub>-SAK1 reduces known age-linked gene expression dysregulation. However, A2A does not further reduce the effect, meaning either that another driver exists in addition to cytosolic acetyl-CoA, or that age-linked gene expression dysregulation is unrelated to cytosolic acetyl-CoA. Oleic acid has little effect on age-linked gene expression dysregulation despite rescuing fitness. With regard to rDNA silencing, transcription of the rDNA intergenic spacer non-coding RNAs promotes ERC formation; we have added data showing that ERC accumulation is not reduced in A2A but slightly higher coherent with the higher replicative age of A2A at 48 h, which suggests silencing is not better in A2A. By RNA-seq, these intergenic spacer transcripts are massively upregulated with age, but this will be a consequence of the increased genomic copy number on ERCs; the upregulation is less in A2A than other conditions, but this arises because the log phase spacer transcript levels are higher and so does not reflect better rDNA silencing. We have previously assayed for heritable changes in rDNA copy number arising during ageing and found (to our surprise) absolutely nothing, so we don't expect any changes under these conditions. The upregulation of transcripts from Sir2-repressed telomeric and MAT loci with age is decreased in P<sub>GPD</sub>-SAK1 and A2A, but the effect size is not different from any other low-expressed genes so we do not think there is a particular effect at loci subject to chromatin silencing (see our previous study Zylstra et al PMID 37643194 for evidence that Sir2-mediated gene silencing is not affected by age). We have added our conclusions from these experiments to the Discussion.

      (2) The current data do not fully distinguish whether AMPK activation and oleic acid supplementation act on distinct subpopulations of aging cells. An alternative explanation is that oleic acid supplementation enhances mitochondrial function and acts additively with AMPK activation, thereby increasing the fraction of cells in the "low senescence" state. Since this distinction is not central to the main conclusions, I suggest softening the language around subpopulation specificity. Emphasizing instead that the A2A strain coordinately modulates multiple branches of acetyl-CoA metabolism to improve late-life fitness would maintain the strength of the central message without over interpretation.

      We respectfully disagree with the reviewer on this point. We show that P<sub>GPD</sub>-SAK1 rescues senescence in ~half the population by a Cat2/Mls1 dependent mechanism (Figure 1F). We then show that in A2A, which rescues most cells, deletion of CAT2/MLS1 restores senescence in ~half the cells (Figure 3F/G). This cannot be explained by an additive mechanism as this would either result in all cells being partially rescued in the P<sub>GPD</sub>-SAK1 and in the A2A cat2Δ mls1Δ mutants, which is definitely not the case either by Tom70-GFP or fitness. Instead the population splits into high/low senescence and fit/unfit cells in the different assays.

      On the specific point of whether lipid synthesis additively increases mitochondrial function, we have added oxygen consumption rate data showing that A2A cells respire more than P<sub>GPD</sub>-SAK1 at 48h but only by a relatively small amount (Figure S3D), so there is indeed an additive improvement in mitochondrial function, but too little to explain the difference in population fitness in our opinion.

      We realise that the reviewer is asking more specifically about oleic acid, but again in the flow data, Figure 4C, what changes with oleic acid or P<sub>GPD</sub>-SAK1 is the proportion of cells in the low Tom70 / high WGA sector. Under an additive effect model, oleic acid or P<sub>GPD</sub>-SAK1 individually would partially reduce Tom70 and partially increase WGA, but the population in the low Tom70 / high WGA sector has the same average Tom70/WGA values in oleic acid, P<sub>GPD</sub>-SAK1 or P<sub>GPD</sub>-SAK1+oleic acid. It is the proportion of cells in this population that changes. Furthermore, under an additive model, wildtype cells aged with oleic acid would not have highest fitness than P<sub>GPD</sub>-SAK1 or A2A (Figure 4D) as these individual cells would lack the mitochondrial upregulation from P<sub>GPD</sub>-SAK1.

      (3) The manuscript proposes that lipid starvation and excess acetyl-CoA are major drivers of senescence in distinct subpopulations of wild-type aging cells. This conclusion is not yet fully supported by the presented data. Direct measurements of age-dependent divergence in acetyl-CoA and fatty acid levels at the single-cell level would be needed to substantiate this model. Based on the current evidence, a more conservative interpretation would be that aging cells exhibit differential sensitivity to perturbations in acetyl-CoA and lipid metabolism. Accordingly, I recommend revising the statement in the Abstract ("We further implicate lipid starvation and excess acetyl coenzyme A availability as major drivers of senescence...") and the corresponding discussion text to better align with the data.

      We agree and have adjusted the abstract to make it clearer that the lipid starvation / excess acetyl-coA interpretation is a model.

      “Our findings support a model in which lipid starvation and excess acetyl-coenzyme A availability are major drivers of senescence in replicatively aged wild-type yeast.”

      Reviewer #3 (Public review):

      Summary:

      These findings suggest that PGPD-SAK1 yeast show a subpopulation with lowered TOM70-GFP expression in high bud scar staining aged cells. Deletion of CAT2 or MLS1 reduces this effect. A PGPD-SAK1 acc1S1157A double mutant (called "A2A" here) shows an even larger effect of lowered tom70 expression in high bud scar staining aged cells. Utilization of various additional mutants involved in acetyl-CoA transport, carnitine shuttle, respiration, etc., leads the authors to conclude that these shifts in TOM70-GFP in aged cells are linked to the AMPK-fatty acid metabolic regulatory system.

      Strengths:

      These extensive and clearly described experiments reveal interesting changes in TOM70-GFP intensity in subsets of aged yeast in several mutants eventually identified as linked to the AMPK-fatty acid metabolic regulatory system.

      Weaknesses:

      (1) 3 biological replicates for mRNASeq is low.

      Thank you for pointing this out. We performed another replicate after posting the initial preprint to confirm the finding but didn’t update the figure in the eLife-reviewed version. We have added this to the scatter plots and analysis in Figure 1, there are minor changes but the set of genes we followed up are still highly significant. For ageing experiments, we sequence to n=3 as a first pass which is sufficient to detect widespread age-linked gene expression effects, and add more replicates if required to solidify findings for specific sets of genes. Hence, the additional RNAseq experiments we have added to the manuscript to Address Reviewer 2’s comments on widespread gene expression effects are also n=3-4.

      (2) While "Traditional conceptions of ageing implicate a progressive accumulation of damage leading to systemic degradation in performance until death, with evolutionary pressures acting to maximise early life fitness and fecundity at the expense of ageing health." is tangential perhaps to the data and conclusions of the study, both claims of this sentence are at best controversial, and the manuscript is no weaker for their omission.

      We would prefer not to remove this sentence, which we see as important to a major message of the manuscript: that ageing does not have to involve a loss of fitness before death. Outside the ageing biology field, ageing is often described as the progressive wearing out of components leading to decline and death (‘like an old car’ is a common analogy); in the ageing field this is certainly controversial, but outside the field it remains the normal understanding. This is what we mean by traditional conceptions, and it is important to consider the contradiction between this widely held viewpoint and our findings (and of course those of many others in the ageing field).

      The second part of the sentence about evolutionary pressures alludes to antagonistic pleiotropy, which we have now made explicit. Antagonistic pleiotropies as a driving mechanism for ageing, while not universally accepted, are as far as we can tell the most widely accepted type of theory in the ageing field. Our interpretation that yeast are bet-hedging as a population growth strategy and this drives ageing in the long term is a classic antagonistic pleiotropy and we need to raise this concept in the introduction.

      (3) The statement that "Here, we determine the basis of senescence and fitness loss in replicatively ageing yeast" is a bit strong as a summary of the present careful work presented here. If the authors had created yeast mutants that retained fitness indefinitely, this would be a more appropriate strength of claim to summarize the work.

      We agree and have moderated this sentence:

      “Here, we show that senescence and fitness loss in replicatively ageing yeast can be almost completely avoided without extension of lifespan by rewiring the conserved AMPK-fatty acid metabolic regulatory system.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The labelling of Figure 3G horizontal axis needs to be realigned with the data.

      Fixed – thank you.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 3G, the x-axis labels appear misaligned and should be corrected for clarity.

      Fixed – thank you.

      (2) Figures S3B and S3C appear to be mislabeled and should be revised.

      Fixed – thank you.

      (3) On page 6 (3rd paragraph), the statement that the beneficial impact arises from acetyl-CoA removal "rather than a benefit of respiration" may be overstated. The data support a role for acetyl-CoA removal but do not fully exclude a contribution from respiration. A more balanced phrasing would improve accuracy.

      We have revised this sentence and also added data:

      “Working in sip2Δ to avoid an increase in AMPK activity due to reduced Acetyl-CoA availability, we observed that ald6Δ increased the low senescence population through decreasing Tom70-GFP (S3C), and therefore the beneficial impact of PGPD-SAK1 on this pathway arises primarily through Acetyl-CoA removal. It is possible that respiration is adding to this benefit, and we detect a significant increase in Oxygen Consumption Rate in aged PGPD-SAK1 cells, but the further increase in A2A is smaller and we consider that this cannot fully explain the effect of acc1S1157A.”

      Reviewer #3 (Recommendations for the authors):

      This manuscript is clearly written, and the data are clearly presented. While 3 biological replicates is inadvisably low for mRNASeq, the subsequent experiments motivated by the genes identified there nevertheless stand on their own as presented.

      Thank you.

    1. eLife Assessment

      This study offers important insight into the pathogenic basis of intragenic frameshift deletions in the carboxy-terminal domain of MECP2, which account for some Rett syndrome cases, yet similar variants also appear in unaffected individuals. Using base editing and mouse models, the authors present convincing evidence supporting the pathogenicity of select deletion variants, with potential implications for therapeutic development.

    2. Reviewer #1 (Public review):

      Summary:

      The authors scrutinized difference in C terminal region variant profiles between Rett syndrome patients and healthy individuals and pinpointed that subtle genetic alternation can cause benign or pathogenic output, which harbors important implication in Rett syndrome diagnosis and proposing therapeutic strategy. This work will be beneficial to clinicians and basic scientists who work on Rett syndrome and carries potential to be applied to other Mendelian rare diseases.

      Strengths:

      Well-designed genetic and molecular experiments, translating genetic differences into functional and clinical changes. This is a unique study resolving subtle changes in sequences give rise to dramatic phenotypic consequences.

      Comments on revised version.

      Improvements were made during the revision.

    3. Reviewer #2 (Public review):

      Summary:

      This study by Guy and Bird and colleagues is a natural follow-up to their 2018 Human Molecular Genetics paper, further clarifying the molecular basis of C-terminal deletions (CTDs) in MECP2 and how they contribute to Rett syndrome. The authors combine human genetic data with well-designed experiments in embryonic stem cells, differentiated neurons, and knock-in mice to explain why some CTD mutations are disease-causing while others are harmless. They show that pathogenic mutations create a specific amino acid motif at the C-terminus, where +2 frameshifts produce a PPX ending that greatly reduces MeCP2 protein levels (likely due to translational stalling) whereas +1 frameshifts generating SPRTX endings are well tolerated.

      Strengths:

      This is a comprehensive and rigorous study that convincingly pinpoints the molecular mechanism behind CTD pathogenicity, with strong agreement between the cell-based and animal data. The authors also provide a proof of principle that modifying the PPX termination codon can restore MeCP2-CTD protein levels and rescue symptoms in mice. In addition, they demonstrate that adenine base editing can correct this defect in cultured cells and increase MeCP2-CTD protein levels. Overall, this is a well-executed study that provides important mechanistic and translational insight into a clinically important class of MECP2 mutations.

      Weaknesses:

      The adenine base editing to change the termination codon is shown feasible in generated cell lines, but yet to be shown in vivo in animal models.

      Comments on revised version.

      The authors have addressed all of my questions and comments.

    4. Author response:

      The following is the authors’ response to the original reviews.

      This is a summary of the changes that have been made to the Reviewed Preprint:

      (1) The data from RettBASE which was analysed in the manuscript has been added in the form of four supplementary tables. Supplementary Table 1 contains the download of all MeCP2 mutations contained in RettBASE. Supplementary Tables 2-4 contain subsets of this data which were used in Figure 2B and Figure 3 SF2. Supplementary Table 4 also has the HGVS nomenclature for both e1 and e2 isoforms and the ClinVar Variation ID for each allele. Wording has been changed to clarify that analysis in the manuscript used this data from RettBASE and not the information that was deposited in ClinVar.

      (2) Similarly, Supplementary Tables 5-8 contain the gnomAD data that was used in the preparation of Figure 2, Figure 2 SF1 and Figure 3 SF1. The “high confidence” alleles in Supplementary Table 8 have been annotated with their HGVS names.

      (3) The criteria for selecting “high confidence” RettBASE and gnomAD alleles have been more explicitly stated in both the Results and Materials and Methods sections.

      (4) An additional “high confidence” RTT mutation (c.1152_1195) has been added to Figures 2B and 3B.

      (5) Figure 3 Supplementary Figure 2 has been added to show reading frame data for all frameshifting deletions in the C-terminal deletion-prone region (CT-DPR), showing that +2 frameshifts predominate in this larger data set, not just in the “high confidence” set. This has necessitated changing the previous Fig. 3 SF2 to Fig. 3 SF3.

      (6) A summary of the genetic alterations described in the manuscript, and their outcomes, has been added as Figure 7.

      (7) A simple flow chart which assists in the classification of human CTDs as “likely benign” or “likely pathogenic” has been added as Figure 8. This will aid future assessment of novel mutations in this region.

      (8) Additions have been made to the Materials and Methods section to comply with reporting guidelines.

      (9) Minor changes have been made to the text to correct typographical errors and to clarify meaning.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors scrutinized differences in C-terminal region variant profiles between Rett syndrome patients and healthy individuals and pinpointed that subtle genetic alternation can cause benign or pathogenic output, which harbors important implications in Rett syndrome diagnosis and proposes a therapeutic strategy. This work will be beneficial to clinicians and basic scientists who work on Rett syndrome, and carries the potential to be applied to other Mendelian rare diseases.

      Strengths:

      Well-designed genetic and molecular experiments translate genetic differences into functional and clinical changes. This is a unique study resolving subtle changes in sequences that give rise to dramatic phenotypic consequences.

      Weaknesses:

      There are many base-editing and protein-expression changes throughout the manuscript, and they cause confusion. It would be helpful to readers if authors could provide a simple summary diagram at the end of the paper.

      We have added a summary diagram, as suggested (Figure 7). We have also provided a flowchart which shows how to classify human CTDs as “likely benign” or “likely pathogenic” based on their location.

      Reviewer #2 (Public review):

      Summary:

      This study by Guy and Bird and colleagues is a natural follow-up to their 2018 Human Molecular Genetics paper, further clarifying the molecular basis of C-terminal deletions (CTDs) in MECP2 and how they contribute to Rett syndrome. The authors combine human genetic data with well-designed experiments in embryonic stem cells, differentiated neurons, and knock-in mice to explain why some CTD mutations are disease-causing while others are harmless. They show that pathogenic mutations create a specific amino acid motif at the C-terminus, where +2 frameshifts produce a PPX ending that greatly reduces MeCP2 protein levels (likely due to translational stalling) whereas +1 frameshifts generating SPRTX endings are well tolerated.

      Strengths:

      This is a comprehensive and rigorous study that convincingly pinpoints the molecular mechanism behind CTD pathogenicity, with strong agreement between the cell-based and animal data. The authors also provide a proof of principle that modifying the PPX termination codon can restore MeCP2-CTD protein levels and rescue symptoms in mice. In addition, they demonstrate that adenine base editing can correct this defect in cultured cells and increase MeCP2-CTD protein levels. Overall, this is a well-executed study that provides important mechanistic and translational insight into a clinically important class of MECP2 mutations.

      Weaknesses:

      The adenine base editing to change the termination codon is shown to be feasible in generated cell lines, but has yet to be shown in vivo in animal models.

      This work is the obvious next step and is in progress. However, with the rise in pre- and neonatal genetic testing we felt it was important to disseminate our findings as soon as possible. The family pedigree in Figure 3C is a clear illustration of this need

      Reviewer #3 (Public review):

      Summary:

      Guy et al. explored the variation in the pathogenicity of carboxy-terminal frameshift deletions in the X-linked MECP2 gene. Loss-of-function variants in MECP2 are associated with Rett syndrome, a severe neurodevelopmental disorder. Although 100's of pathogenic MECP2 variants have been found in people with Rett syndrome, 8 recurrent point mutations are found in ~65% of disease cases, and frameshift insertions/deletions (indels) variants resulting in production of carboxy-terminal truncated (CTT) MeCP2 protein account for ~10% of cases. Many of these occur in a "deletion prone region" (DPR) between c.1110-1210, with common recurrent deletions c.1157-1197del (CTD1) and c.1164_1207del (CTD2). While two major protein functional domains have been defined in MeCP2, the methyl-binding domain (MBD) and the NCoR interacting domain (NID), the functional role of the carboxy-terminal domain (CTD, beyond the NID, predicted to have a disordered protein structure) has not been identified, and previous work by this group and others demonstrated that a Mecp2 "minigene" lacking the CTD retains MeCP2 function suggesting that the CTD is dispensable. This raises an important question: If the CTD is dispensable, what is the pathogenic basis of the various CTT frameshift variants? Prior work from this group demonstrated that genetically engineered mice expressing the CTD1 variant had decreased expression of Mecp2 RNA and MeCP2 protein and decreased survival, but those expressing the CTD2 variant had normal Mecp2 RNA and protein and survival. However, they noted that differences between the mouse and human coding sequences resulted in different terminal sequences between the two common CTD, with CTD1 ending in -PPX in both mouse and human, but CTD2 ending in -PPC in human but -SPX in mouse, and in the previous paper they demonstrated in humanized mouse ES cells (edited to have the same -PPX termination) containing the CTD2 deletion resulted in decreased Mecp2 RNA and protein levels. This previous work provides the underlying hypotheses that they sought to explore, which is that the pathological basis of disease causing CTD relates to the formation of truncated proteins that end with a specific amino acid sequence (-PPX), which leads to decreased mRNA and protein levels, whereas tolerated, non-pathogenic CTD do not lead to production of truncated proteins ending in this sequence and retain normal mRNA/protein expression.

      In this manuscript, they evaluate missense variants, in-frame deletions, and frame shift deletions within the DPR from the aggregated Genome Aggregated Database (gnomAD) and find that the "apparently" normal individuals within gnomAD had numerous tolerated missense variants and in-frame deletions within this region, as well as frameshift deletions (in hemizygous males) in the defined region. All of the gnomAD deletions within this region resulted in terminal amino acid sequences -SPRTX (due to +1 frameshift), whereas nearly all deletion variants in this region from people with Rett syndrome (from the Clinvar copy of the former RettBase database) had a terminal -PPX sequence, due to a +2 frameshift. They hypothesized that terminal proline codons causing ribosomal stalling and "nonsense mediated decay like" degradation of mRNA (with subsequent decreased protein expression) was the basis of the specific pathogenicity of the +2 frameshift variants, and that utilizing adenine base editors (ABE) to convert the termination codon to a tryptophan could correct this issue. They demonstrate this by engineering the change into mouse embryonic stem cell lines and mouse lines containing the CTD1 deletion and show that this change normalized Mecp2 mRNA and protein levels and mouse phenotypes. Finally, they performed an initial proof-of-concept in an inducible HEK cell line and showed the ability of targeted ABE to edit the correct adenine and cause production of the expected larger truncated Mecp2 protein from CTD1 constructs.

      The findings of this manuscript provide a level of support for their hypothesis about the pathogenicity versus non-pathogenicity of some MECP2 CTT intragenic deletions and provide preliminary evidence for a novel therapeutic approach for Rett syndrome; however, limitations in their analysis do not fully support the broader conclusions presented.

      Strengths:

      (1) Utilization of publicly available databases containing aggregated genetic sequencing data from adult cohorts (gnomAD) and people with Rett syndrome (Clinvar copy of RettBase) to compare differences in the composition of the resulting terminal amino acid sequences resulting from deletions presumed to be pathogenic (n+2) versus presumed to be tolerated (n+1).

      (2) Evaluation of a unique human pedigree containing an n+1 deletion in this region that was reported as pathogenic, with demonstration of inheritance of this from the unaffected father and presence within other unaffected family members.

      (3) Development of a novel engineered mouse model of a previously assumed n+1 pathogenic variant to demonstrate lack of detrimental effect, supporting that this is likely a benign variant and not causative of Rett syndrome.

      (4) Creation and evaluation of novel cell lines and mouse models to test the hypothesis that the pathogenicity of the n+2 deletion variants could be altered by a single base change in the frameshifted stop codon.

      (5) Initial proof-of-concept experiments demonstrating the potential of ABE to correct the pathogenicity of these n+2 deletion variants.

      Weaknesses:

      (1) While the use of the large aggregated gnomAD genetic data benefits from the overall size of the data, the presence of genetic variants within this collection does not inherently mean that they are "neutral" or benign. While gnomAD does not include children, it does include aggregated data from a variety of projects targeting neuropsychiatric (and other conditions), so there is information in gnomAD from people with various medical/neuropsychiatric conditions. The authors do make some acknowledgement of this and argue that the presence of intragenic deletion variants in their region of interest in hemizygous males indicates that it is highly likely that these are tolerated, non-pathogenic variants. Broadly, it is likely true that gnomAD MECP2 variants found in hemizygous males are unlikely to cause Rett syndrome in heterozygous females, it does not necessarily mean that these variants have no potential to cause other, milder, neuropsychiatric disorders. As a clear example, within gnomAD, there is a hemizygous male with the rs28934908 C>T variant that results in p.A140V (p.A152V in e1 transcript numbering convention). This pathogenic variant has been found in a number of pedigrees with an X-linked intellectual disability pattern, in which males have a clear neurodevelopmental disorder and heterozygous females have mild intellectual disability (see PMIDs 12325019, 24328834 as representative examples of a large number of publications describing this). Thus, while their claim that hemizygous deletion variants in gnomAD are unlikely to cause Rett syndrome, that cannot make the definitive statement that they are not pathogenic and completely benign, especially when only found in a very small number of individuals in gnomAD.

      We have included the possibility that mutations found in gnomAD may give rise to less severe neurological conditions in the discussion.

      (2) The authors focus exclusively on deletions within the "DPR", they define as between c.1110-1210 and say that these deletions account for 10% of Rett syndrome cases. However, the published studies that are the basis for this 10% estimate include all genetic variants (frameshift deletions, insertions, complex insertion/deletions, nonsense variants) resulting in truncations beyond the NID. For example, Bebbington 2010 (PMID: 19914908), which includes frameshift indels as early as c.905 and beyond c.1210. Further specific examples from RettBase are described below, but the important point is that their evaluation of only frameshift variants within c.1110-1210 is not truly representative of the totality of genetic variants that collectively are considered CTT and account for 10% of Rett cases.

      The vast majority of C-terminal truncating mutations do occur within the “CT-DPR”, likely due to its C-rich nucleotide sequence and the presence of microhomologies within the region. Looking at frameshifting deletions in RettBASE that start after the NID, a large proportion of these end within the CT-DPR and result in a -PPX ending. We decided to restrict our analysis to the c.1110-1210 region to avoid including the rarer examples that may have a different reason for their pathogenicity. We do not assert that all C-terminal truncations are pathogenic due to this mechanism, but current evidence suggests that most are.

      (3) The authors say that they evaluated the putative pathogenic variants contained within RettBase (which is no longer available, but the data were transferred to Clinvar) for all cases with Classic Rett syndrome and de novo deletion variants within their defined DPR domain. Looking at the data from the Clinvar copy of RettBase, there are a number (n=143) of c-terminal truncating variants (either frameshift or nonsense) present beyond the NID, but the authors only discuss 14 deletion frameshift variants in this manuscript. A number of these variants have molecular features that do not fall into the pathogenic classification proposed by the authors and are not addressed in the manuscript and do not support the generalization of the conclusions presented in this manuscript, especially the conclusion that the determination of pathogenicity of all c-terminal truncating variants can be determined according to their proposed n+2 rule, or that all of the 10% of people with Rett syndrome and c-terminal truncating variants could be treated by using a base editor to correct the -PPX termination codon.

      It is important to state here that we did not use the data in ClinVar for our analysis, but the original information that was held in RettBASE. We have clarified this in the manuscript and have now included supplementary tables containing the data we downloaded from both RettBASE and gnomAD. Table 1 contains all the RettBASE entries with MECP2 mutations, while Tables 2-4 contain subsets of this data pertaining to CTDs. We have extended our “high confidence” set of RTT alleles to contain one more that was previously overlooked (c.1152_1195) to bring the total number of alleles to 15. Taken together these alleles account for 158 individual entries in RettBASE. We have now included an analysis of all frameshifting mutations in the CT-DPR (Supplementary Table 3, Fig. 3 SF2) which covers 69 different mutations and 260 individual entries. Of these, +2 frameshifts make up the large majority, in contrast to the gnomAD data shown in Supplementary Table 8 and Fig. 3 SF1.

      (4) The HEK-based system utilized is convenient for doing the initial experiments testing ABE; however, it represents an artificial system expressing cDNA without splicing. Canonical NMD is dependent on splicing, and while non-canonical "NMD-like" processes are less well understood, a concern is whether the artificial system used can adequately predict efficacy in a native setting that includes introns and splicing.

      We disagree with this opinion. We show that the loss of protein and mRNA seen with knock in mouse and human alleles is recapitulated when using a cDNA-based transgene in the HEK system, demonstrating that the mechanism of loss does not involve factors bound at splice junctions etc. We also demonstrate the effect of the A to G change at the stop codon is the same whether we do this by base editing our cDNA transgene in T-REx cells or by making the CTD1 X>W knock-in mouse. Both result in increased levels of a slightly extended but still truncated protein.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The phrase in the title, "an alternative therapeutic approach" is only insinuated in the manuscript, making it rather inappropriate to be in the title.

      In this study we use adenine base editing to modify RTT-causing CTD mutations in T-REx cells which is clearly a precursor to developing a therapy, utilising the new findings in this study. We therefore feel that the use of “an alternative therapeutic approach” can be justified.

      Reviewer #2 (Recommendations for the authors):

      I have a few minor comments for the authors to consider:

      (1) Please double-check Figure 2 Supplementary Figure 1, as the allele count for E394K does not appear to be in the thousands; rather, E397K seems to be the variant shown in the graph.

      Yes, this was an error and has been corrected to E397K in the text. Thank you for spotting it.

      (2) On page 12, the phrase “common DNA sequence features shared by all CTDs that give rise to RTT” might be better described as “amino acid sequence features.”

      This has been altered in the text as suggested.

      (3) On page 3, the sentence "analysis of patient mutations and experimental data from mouse models support a role in transcriptional repression" cites Gabel et al. 2015 and Kinde et al. 2016, which focus on null alleles but not patient mutations. It would be appropriate to also cite Johnson et al. 2017, which analyzed MeCP2 T158M and R106W patient mutations.

      This has now been cited.

      (4) In the same sentence, Bajikar et al. 2025 are described as studying the "acute loss of MeCP2," but gene expression was analyzed after a week or longer period of time, not minutes to hours as in degron-mediated degradation systems.

      The word “acute” has been removed from the text.

      (5) On page 7, the statement that "E394K is common... who were later found to have additional pathogenic MECP2 mutations" should include a supporting reference.

      We have now included the references Moncla et al (2002) and Wan et al (1999) to address this.

      (6) Similarly, on page 10, the sentence "these findings question the validity of two cases where individuals presented with classical Rett..." is missing a reference to the case report mentioned.

      The references Bienvenu et al (2000) and Philippe et al (2006) have been added.

      (7) While the manuscript is well written and full of detail, if space is an issue, the authors might consider tightening sections that reiterate findings from their 2018 HMG paper.

      A section discussing the CTD2 allele from the 2018 HMG paper has been removed from the results section.

      Reviewer #3 (Recommendations for the authors):

      (1) Overall, the manuscript is rather dense and potentially challenging to follow easily, especially for a non-expert reader.

      Minor edits have been made to the text which will hopefully make it easier to follow. We have also added two new figures (Figures 7 and 8) to summarize the different alleles and edits which appear in the paper, and to show how to determine whether a CTD in the region is likely to be benign or pathogenic.

      (2) The introduction of data presented in Figure 1 within the manuscript introduction seems inappropriate and should be moved to the results section.

      We would say that this is unconventional rather than inappropriate, and is referred to in the introduction, so we would prefer to leave it as it is.

      (3) Providing specific, common nomenclature for genetic repository variants (rs numbers, gnomAD IDs, etc) somewhere would be beneficial. This is an issue because of the complexity of numbering (either coding or protein) for MECP2 due to the different transcript-based numbering systems.

      This nomenclature is now included in the supplementary tables of data from RettBASE and gnomAD.

      (4) As described in the public comments, there are a number of MECP2 genetic variants listed in the Clinvar copy of RettBase, resulting in c-terminal truncations that are not mentioned or discussed within the manuscript. Without the level of detail present in the original form of RettBase (number of events, de novo, etc) in the currently available Clinvar iteration, it is unclear why a number of variants, even within the limited DPR region, were not mentioned. A supplementary file including the more complete information from the RettBase version, with a complete listing of all c-terminal truncating variants, and an explicit rationale for the exclusion of variants would be helpful.

      We have now included supplementary tables with our download of all MECP2 mutations which were held in RettBASE. We have further added tables with the subset of mutations that we have analysed and have more explicitly stated our criteria for defining the “high confidence” sets of mutations. We did not download the data relating to “evidence of pathogenicity” (ie de novo?, absent from parents etc) from RettBASE, but annotated our list of CTDs with this information while RettBASE was still available. This was used in Supplementary Table 3.

      (5) A discussion of the limitations, notably that the fact that the focus exclusively on deletion variants within a restricted region (c.1110-1210) does not truly represent all genetic variants that cumulatively account for 10% of Rett cases, is needed. Furthermore, as pointed out, not all frameshift variants, even those that are n+2, result in the -PPX termination that is presented as the pathogenic basis of c-terminal truncations and amenable to correction by ABE. This should be noted in the discussion, as well as consideration of the late nonsense variants that cause c-terminal truncations (some of which would be very similar to the deletion variants discussed but without the proposed primary pathogenic driver, -PPX).

      We do not claim to explain the pathogenic mechanism of all C-terminal frameshift mutations found in cases of Rett syndrome. There will certainly be some that do not fit our explanation. However, we believe we have shown evidence that a large proportion of CTDs in RTT will be amenable to the therapy we propose.

      (6) Regarding point 3 in the public review, specifically:

      (a) n=7 nonsense variants (S360X, K363X, E397X, R453X, E455X) that do not carry the destabilizing -PPX sequence.

      (b) n=136 frameshift indel variants beyond the NID.

      (i) n=11 that have indels that extend past the native stop codon, n=4 of which start within the DPR domain (c.1110-1210) but would have a different terminal sequence than their proposed pathogenic -PPX sequence.

      (ii) n=125 frameshift indels with terminal breakpoint before the native stop codon

      (c) n=89 that have start or stop points within c.1110-1210

      (d) n=72 not mentioned within the manuscript.

      (e) n=27 are n+1, with 26/27 having what the authors term as the "tolerated" -SPRTX ending, but 1/27 having a frameshift beyond this region (c.1133_1361)

      (f) n=45 are n+2, with 32/45 ending in -PPX (supporting authors conclusion), but 10/45 will use the frameshift stop codon preceding the -PPX and have a different terminal sequence, and 3/45 result in a frameshift termination beyond the -PPX sequence.

      (g) n=36 have breakpoints either before c.1110 or after c.1210

      (h) n=16 start before c.1110, with 11/16 ending before c.1110. 4 of these 11 are n+2, but would use the earlier frameshift stop codon and not have -PPX terminal sequence. For 5/16, the indel extends past c.1210, with the n+2 leading to frameshift termination codons beyond the -PPX sequence.

      (i) n=20 indels start beyond c.1210, with 8/20 being n+1 and 12/20 being n+2, with neither leading to the -SPRTX or --PPX termination sequences characterized in this manuscript.

      As mentioned previously, we have used data taken from RettBASE, not from ClinVar. Both RettBASE and ClinVar will contain MECP2 mutations found in cases of RTT which are not the causative mutation. Databases of this kind contain sequencing errors and mutations that have been mistakenly assigned as causative. It is therefore imperative that the publicly available information is screened to only include mutations that meet stringent criteria. This is why we chose to start by looking at high confidence sets of mutations, with our conclusions supported by analysis of all such mutations in our region of study.

      As mentioned in response to point 5 above, we do not claim to explain every mutation in the region, but believe this study reveals an important disease mechanism for a large proportion of CTDs, leading to a potential therapy. It also contains significant information for predicting the likely prognosis of individuals with CTDs, who may remain healthy but are currently informed that their mutation is likely pathogenic. At present it seems that this is often based solely on the presence of a frameshifting mutation with similarities to bona fide RTT CTDs, without strong evidence of pathogenicity.

    1. eLife Assessment

      This valuable study presents a comprehensive exploration of c-di-GMP-associated protein interaction networks in Pseudomonas fluorescens, with a particular focus on biofilm-related phenotypes. The evidence is convincing, supported by a genome-wide yeast two-hybrid screen, phenotypic analyses, and experimental validation. The work identifies multiple interaction hubs and provides a resource that will be of use to the biofilm and c-di-GMP communities, while additional mechanistic exploration would further enhance its impact by clarifying the biological roles of many of the identified interactions.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript by Noirot-Gros et. al. presents a herculean effort to map the protein-protein interactome of the c-di-GMP signaling network in Pseudomonas fluorescens (Pf). C-di-GMP, the key driver of biofilm formation in bacteria, is controlled by a highly complex network of synthesis, degradation and effector proteins. Pf is no exception as it encodes dozens of such proteins. The authors use a Yeast Two-Hybrid approach genome-wide screen with 10 diguanylate cyclase (DGC) enzymes as bait to assess protein-protein interactions in this network. The results identify over one hundred such interactions with several different hubs, including c-di-GMP signaling, other signaling systems, membrane proteins, etc. The authors then explore the original bait proteins as well as identify interactors on biofilm formation-related phenotypes and swarming using a high-throughput CRISPRi expression knockdown approach. The amount of data generated is quite impressive. Much of the manuscript uses statistical-based network analysis to group different proteins based on their interactions or impact on phenotypes, which is a high-level analysis that can catalyze further study into this system. The authors chose three specific proteins to assess their impact on cell morphology, DNA repair, and protein localization. Overall, in my view, this is perhaps the best analysis of a c-di-GMP protein-protein interactome, and it provides a multitude of hypotheses to be tested. However, therein lies the weakness of the manuscript in that very few of these hypotheses are actually tested. But such is not the goal of this network analysis type of approach. Overall, I think the work will be highly impactful to those in the c-di-GMP field, and it provides a template for others attempting such analyses of protein-protein interactions.

      Strengths:

      The manuscript is impressive in the sheer scale of the protein-protein interactions identified, network analysis, and phenotypic analysis of specific proteins in the network. It is an impressive amount of work that could be very useful to the field. It is also statistically rigorous in its analysis of significant interactions or network nodes.

      Weaknesses:

      The weakness of the manuscript is that, with three exceptions, very few of the hypotheses are actually tested. For example, BifA is shown to be a network hub protein that interacts with many other diguanylate cyclases, and this is hypothesized to be through GGDEF heterodimerization. I appreciate that experimentally testing such a hypothesis is probably another entire manuscript, but some early forays into such ideas could be undertaken using AlphaFold structural modeling of protein-protein interactions compared with GGDEFs that don't form heterodimers. Also, an inherent weakness is that such detailed analyses of a c-di-GMP signaling network, in which each diguanylate cyclase and phosphodiesterase may respond to a unique cue, is that the network identified and the conclusions made are highly specific to the experimental conditions in which the work was done. Therefore, it is unclear how broadly these conclusions (i.e. BifA is the central regulator of c-di-GMP signaling) apply to other conditions. But it is impossible to get around such a limitation, and this work can lead to testing the robustness of the identified network in other environments.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Noirot-Gros and coworkers investigated the network of c-di-GMP associated protein complexes in Pseudomonas fluorescens. They did so by using a genome-wide yeast two-hybrid screen, and that was further probed by phenotypic screening that focused on biofilm and motility phenotypes. From this network map, they discovered that the phosphodiesterase DipA interacts with the GGDEF domains of many c-di-GMP-binding proteins.

      Strengths:

      (1) Broadness of screen led to identification of new interactions: The genome-wide yeast two-hybrid screening approach permitted broad investigation of c-di-GMP-associated protein-protein interactions. These interactions included some previously validated interactions as well as newly discovered interactions.

      (2) Complementary experimental validation: The proposed network was experimentally validated, including by using a CRISPRi-based approach in which the expression of genes encoding proteins identified in the network was systematically suppressed, and then the impact on the biofilm and motility phenotypes was assessed.

      Weaknesses:

      The findings would have been strengthened by further biochemical analysis, but this is likely beyond the scope of the paper.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Noirot-Gross et al take an open-ended approach to elucidate the c-diGMP-associated protein complexes in Pseudomonas fluorescens. Starting with 10 cyclic d-GMP putative proteins, they use a combination of genome-wide two-hybrid system followed by CRISPRi-mediated exploration of phenotypes to describe the cyclic di-GMP-associated regulation of biofilm formation, and how it relates to other functions. Overall, this work presents an excellent example of how genome annotations can be further confirmed with the use of integrated functional genomic approaches. Some areas of improvement can be applied to this manuscript to enhance readability and provide a clearer distinction between confirmatory results and new findings, which are provided below:

      Strengths:

      (1) The authors have explored their findings extensively and provide a comprehensive view of the topic.

      (2) The combination of genome-wide explorations of protein-protein interactions with the more focused phenotypic exploration of the interactions found provides a solid framework for the work presented.

      Weaknesses:

      (1) Overall goal of the work:

      While articles that describe open-ended approaches can be comprehensive and descriptive in nature, the authors should have a main overall goal, which can guide the reader through the main and most compelling findings at the end. As written, the overall goal is not clear. The network perspective is interesting, and the focus on biofilm formation appears in the title. Why P. fluorescens? How is cyclic di-GMP-mediated regulation of biofilm formation in P. fluorescens different from P. aeruginosa? Why would it be studied?

      (Positive or negative regulation of biofilm formation?)

      (2) Abstract:

      The abstract is very well written and guides the reader to the DipA as a hub protein in the network. From further reading, the article could clarify whether this finding is confirmatory or novel (does DipA play a similar role in P. aeruginosa?) It would be appropriate to mention the role of DipA in other Pseudomonas species from the beginning, and not only in the discussion session.

      (3) Introduction:

      The introduction is nicely written. An area of improvement could be giving more attention to protein interactions as relevant to c-di-GMP. The authors could consider an independent paragraph starting with line 84-85 "Protein-protein interactions involving DGCs, PDEs, and target effectors are crucial in establishing localized signalling through the generation of local pools of c-di-GMP", expanding on this particular aspect with an example of localized signal, after explaining that localization could help decipher specific function within the network of DGCs and PDEs. Then go into connecting biofilms with c-di-GMP and protein-protein interactions, using the example of GcbC and LapD.

      (4) The rationale of choosing 10 PDEs could be clarified. The nice diagrams shown in the supplementary table could be used as part of Figure 1, so the reader understands why these proteins were used, and what is known about them (for example, add them as Figure 1a).

      (5) Figures 1b and 2 convey the same information as in Figure 1a. They could be removed without affecting the understanding of the article.

      (6) CRISPRi and Figure 3. Figure 3 shows the methodology of CRICPR phenotypic screening. A diagram showing the CRISPRi system in P. fluorescens could help the non-expert reader. While the choice of 23 proteins related to the emerging hub DipA is clear, the choice of the other 33 genes could be better explained. Are these proteins already related to biofilm formation? Where are they part of the network detected? How about the other 14 SBW25 genes? The authors could clarify the rationale of the choices. Figure 4 could be combined with Figure 3 or moved to the supplementary material.

      (7) Figures 5, 6 and 7 represent solid network analysis of the findings. Still, they could be improved in clarity on the main findings. The authors conclude at the end of section 3.2.3 that there are networks that exert a "positive role" and a "negative role". The authors could show that in the figures, explaining what those roles are: more biofilm structural coding genes? positive or negative regulation of biofilm formation?)

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Noirot-Gros et. al. presents a herculean effort to map the protein-protein interactome of the c-di-GMP signaling network in Pseudomonas fluorescens (Pf). C-di-GMP, the key driver of biofilm formation in bacteria, is controlled by a highly complex network of synthesis, degradation and effector proteins. Pf is no exception as it encodes dozens of such proteins. The authors use a Yeast Two-Hybrid approach genome-wide screen with 10 diguanylate cyclase (DGC) enzymes as bait to assess protein-protein interactions in this network. The results identify over one hundred such interactions with several different hubs, including c-di-GMP signaling, other signaling systems, membrane proteins, etc. The authors then explore the original bait proteins as well as identify interactors on biofilm formation-related phenotypes and swarming using a high-throughput CRISPRi expression knockdown approach. The amount of data generated is quite impressive. Much of the manuscript uses statistical-based network analysis to group different proteins based on their interactions or impact on phenotypes, which is a high-level analysis that can catalyze further study into this system. The authors chose three specific proteins to assess their impact on cell morphology, DNA repair, and protein localization. Overall, in my view, this is perhaps the best analysis of a c-di-GMP protein-protein interactome, and it provides a multitude of hypotheses to be tested. However, therein lies the weakness of the manuscript in that very few of these hypotheses are actually tested. But such is not the goal of this network analysis type of approach. Overall, I think the work will be highly impactful to those in the c-di-GMP field, and it provides a template for others attempting such analyses of protein-protein interactions.

      Strengths:

      The manuscript is impressive in the sheer scale of the protein-protein interactions identified, network analysis, and phenotypic analysis of specific proteins in the network. It is an impressive amount of work that could be very useful to the field. It is also statistically rigorous in its analysis of significant interactions or network nodes.

      Weaknesses:

      The weakness of the manuscript is that, with three exceptions, very few of the hypotheses are actually tested. For example, BifA is shown to be a network hub protein that interacts with many other diguanylate cyclases, and this is hypothesized to be through GGDEF heterodimerization. I appreciate that experimentally testing such a hypothesis is probably another entire manuscript, but some early forays into such ideas could be undertaken using AlphaFold structural modeling of protein-protein interactions compared with GGDEFs that don't form heterodimers. Also, an inherent weakness is that such detailed analyses of a c-di-GMP signaling network, in which each diguanylate cyclase and phosphodiesterase may respond to a unique cue, is that the network identified and the conclusions made are highly specific to the experimental conditions in which the work was done. Therefore, it is unclear how broadly these conclusions (i.e. BifA is the central regulator of c-di-GMP signaling) apply to other conditions. But it is impossible to get around such a limitation, and this work can lead to testing the robustness of the identified network in other environments.

      We would like to thank the reviewer sincerely for their positive comments on our manuscript and for their constructive feedback. We recognize the limitations arising from the lack of extensive knowledge regarding the environmental cues that trigger the regulation of all CDG activities in P. fluorescens. We hypothesize that DipA acts as a central local hub that positively or negatively regulates the activity of its interacting CDG partners throughout the cell life cycle, lifestyle transitions and environmental signals. Testing this hypothesis would indeed require extensive biochemical and omics approaches. However, strengthening the significance of DipA complexes in silico using AlphaFold is a very appealing proposition and we are currently considering including this analysis in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Noirot-Gros and coworkers investigated the network of c-di-GMP associated protein complexes in Pseudomonas fluorescens. They did so by using a genome-wide yeast two-hybrid screen, and that was further probed by phenotypic screening that focused on biofilm and motility phenotypes. From this network map, they discovered that the phosphodiesterase DipA interacts with the GGDEF domains of many c-di-GMP-binding proteins.

      Strengths:

      (1) Broadness of screen led to identification of new interactions: The genome-wide yeast two-hybrid screening approach permitted broad investigation of c-di-GMP-associated protein-protein interactions. These interactions included some previously validated interactions as well as newly discovered interactions.

      (2) Complementary experimental validation: The proposed network was experimentally validated, including by using a CRISPRi-based approach in which the expression of genes encoding proteins identified in the network was systematically suppressed, and then the impact on the biofilm and motility phenotypes was assessed.

      Weaknesses:

      The findings would have been strengthened by further biochemical analysis, but this is likely beyond the scope of the paper.

      We would like to express our gratitude to the reviewer for their positive evaluation assessment, and for taking into account the limitations of the study's scope.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Noirot-Gross et al take an open-ended approach to elucidate the c-diGMP-associated protein complexes in Pseudomonas fluorescens. Starting with 10 cyclic d-GMP putative proteins, they use a combination of genome-wide two-hybrid system followed by CRISPRi-mediated exploration of phenotypes to describe the cyclic di-GMP-associated regulation of biofilm formation, and how it relates to other functions. Overall, this work presents an excellent example of how genome annotations can be further confirmed with the use of integrated functional genomic approaches. Some areas of improvement can be applied to this manuscript to enhance readability and provide a clearer distinction between confirmatory results and new findings, which are provided below:

      Strengths:

      (1) The authors have explored their findings extensively and provide a comprehensive view of the topic.

      (2) The combination of genome-wide explorations of protein-protein interactions with the more focused phenotypic exploration of the interactions found provides a solid framework for the work presented.

      Weaknesses:

      (1) Overall goal of the work:

      While articles that describe open-ended approaches can be comprehensive and descriptive in nature, the authors should have a main overall goal, which can guide the reader through the main and most compelling findings at the end. As written, the overall goal is not clear. The network perspective is interesting, and the focus on biofilm formation appears in the title. Why P. fluorescens? How is cyclic di-GMP-mediated regulation of biofilm formation in P. fluorescens different from P. aeruginosa? Why would it be studied? (Positive or negative regulation of biofilm formation?)

      We would like to express our appreciation to the reviewer for their thorough evaluation of our manuscript and for the constructive feedback they provided. The overall goal of this study will be further refined, and outlined in the introduction in the revised version of the manuscript.

      (2) Abstract:

      The abstract is very well written and guides the reader to the DipA as a hub protein in the network. From further reading, the article could clarify whether this finding is confirmatory or novel (does DipA play a similar role in P. aeruginosa?) It would be appropriate to mention the role of DipA in other Pseudomonas species from the beginning, and not only in the discussion session.

      (3) Introduction:

      The introduction is nicely written. An area of improvement could be giving more attention to protein interactions as relevant to c-di-GMP. The authors could consider an independent paragraph starting with line 84-85 "Protein-protein interactions involving DGCs, PDEs, and target effectors are crucial in establishing localized signalling through the generation of local pools of c-di-GMP", expanding on this particular aspect with an example of localized signal, after explaining that localization could help decipher specific function within the network of DGCs and PDEs. Then go into connecting biofilms with c-di-GMP and protein-protein interactions, using the example of GcbC and LapD.

      We propose highlighting the example to the local signalling cascade formed by the tripartite system YdaM, YciR and MlrA. This will be addressed in the revised version of the manuscript.

      (4) The rationale of choosing 10 PDEs could be clarified. The nice diagrams shown in the supplementary table could be used as part of Figure 1, so the reader understands why these proteins were used, and what is known about them (for example, add them as Figure 1a).

      We propose to include a specific section in the supplementary file to explain the whole rationale behind choosing these CDGs. These proteins were selected based on their involvement in different steps of biofilm formation in Pseudomonas, as well as their role in the ability of P. fluorescens strains to colonize plant roots.

      (5) Figures 1b and 2 convey the same information as in Figure 1a. They could be removed without affecting the understanding of the article.

      Figure 2 will be transferred in Supplementary as part of the Figure S1

      (6) CRISPRi and Figure 3. Figure 3 shows the methodology of CRICPR phenotypic screening. A diagram showing the CRISPRi system in P. fluorescens could help the non-expert reader. While the choice of 23 proteins related to the emerging hub DipA is clear, the choice of the other 33 genes could be better explained. Are these proteins already related to biofilm formation? Where are they part of the network detected? How about the other 14 SBW25 genes? The authors could clarify the rationale of the choices. Figure 4 could be combined with Figure 3 or moved to the supplementary material.

      A better description of the rationale behind the choice of tested interacting protein partners will be provided. We also agree to combine Figure 4 with Figure 3.

      (7) Figures 5, 6 and 7 represent solid network analysis of the findings. Still, they could be improved in clarity on the main findings. The authors conclude at the end of section 3.2.3 that there are networks that exert a "positive role" and a "negative role". The authors could show that in the figures, explaining what those roles are: more biofilm structural coding genes? positive or negative regulation of biofilm formation?)

  2. Jul 2026
    1. eLife Assessment

      This important study conducted by Hall and colleagues advances our understanding of host-pathogen co-evolution by providing an elegant multi-omic comparative framework that uncovers a profound, pathogen-directed epigenomic reprogramming of bovine alveolar macrophages uniquely driven by host-adapted Mycobacterium bovis. Convincing evidence supports these findings and draws on robust, multi-layered high-throughput sequencing data integrated with large-scale cattle GWAS metadata to prioritize actionable genetic variants linked to disease susceptibility and to highlight key immune response genes/pathways upregulated in response to infection and pathogen-specific host adaptation mechanisms.

    2. Reviewer #1 (Public review):

      This manuscript by Hall et al. uses a multi-omic approach to investigate how distinct members of the Mycobacterium tuberculosis complex (MTBC: M. bovis, M. tuberculosis, the attenuated M. bovis BCG vaccine strain, and gamma-irradiated M. bovis) affect bovine alveolar macrophage epigenetic and transcriptional responses after 24 hours. The investigators used RNA-Seq, ATAC-Seq, and ChIP-Seq to assess differential gene expression in each complex type and integrated gene transcription with chromatin accessibility and epigenetic modifications, highlighting key immune response genes/pathways upregulated in response to infection and pathogen-specific host adaptation mechanisms. The analysis also revealed that the most pronounced transcriptional and epigenetic responses were in the M. bovis-infected cells compared with the other complex types. Comparing top genes associated with M. bovis infection of macrophages to a GWAS data set revealed 4 key genes associated with increased susceptibility to infection.

      Overall, this is a technically sound manuscript that contains highly interesting and useful data on bovine innate immune responses to different types of Mycobacterium tuberculosis, which are important to the immunology and infectious disease community as well as the livestock industry. However, in its current format, the manuscript presents the data/figures in a way that is not particularly informative (despite the rich data set) and is too descriptive. We also have some general concerns and suggestions listed below.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Hall et al. present a rigorous and comprehensive multi-omic comparative analysis investigating how host-adapted and non-host-adapted mycobacteria reconfigure the bovine host immune response. By utilizing RNA-seq, ATAC-seq, and ChIP-seq across four histone marks (H3K4me3, H3K4me1, H3K27ac, and H3K27me3) plus CTCF binding, the authors track the regulatory dynamics of primary bovine alveolar macrophages (bAM) challenged with Mycobacterium bovis (MBO), Mycobacterium tuberculosis (MTU), M. bovis BCG, and gamma-irradiated (killed) M. bovis (IRR).

      The study highlights a profound, pathogen-driven epigenomic reprogramming that is largely unique to the host-adapted virulent pathogen (M. bovis). Crucially, the authors integrate these regulatory networks with existing Holstein-Friesian GWAS datasets to prioritize novel candidate genes (ERBB4, LRCH1, MRTFA, and RNPC3) associated with M. bovis infection susceptibility. This work represents a significant advancement in our understanding of host-pathogen interactions and animal resilience to bovine tuberculosis (bTB).

      Strengths:

      (1) The manuscript addresses a major socioeconomic problem in global livestock agriculture and human zoonotic health. By profiling host-adapted vs. non-host-adapted and live vs. dead bacilli, it provides fundamental insights into mycobacterial virulence mechanisms and evolutionary adaptation.

      (2) The multi-omic approach is robustly executed, and the sample size (n=6 for RNA-seq, n=3 for ChIP/ATAC-seq subgroups) is highly appropriate for primary livestock cell cultures.

      (3) The inclusion of live virulent, live attenuated, non-host-adapted, and killed strains allows the authors to dissect whether host responses are driven by passive PAMP recognition or active, pathogen-directed virulence factors.

      (4) The parallel mapping of four distinct histone marks alongside chromatin accessibility mapping (ATAC-seq) yields a highly refined picture of enhancer and promoter dynamics.

      Weaknesses:

      (1) The profound transcriptional response of the gamma-irradiated (IRR) M. bovis group (3,320 DEGs vs. 2,312 for MTU) is intriguing but lacks a deep biochemical and functional explanation.

      (2) While the paper provides a clear atlas of epigenetic alterations, the underlying mycobacterial effectors driving these specific chromatin alterations remain largely correlative.

    4. Reviewer #3 (Public review):

      Summary:

      Hall et al. use a multi-omics approach to investigate the responses of bovine alveolar macrophages to Mycobacterium tuberculosis infections and the underlying mechanisms that are shared between bacterial family members. The study is of particular importance for multiple reasons, including the impact of bovine infections on food supply (and resulting economic impacts) as well as the use of bovines as a large animal model to investigate the effects of M. tuberculosis infection. The authors isolated bovine alveolar macrophages and exposed them to infection with M. bovis, M. bovis BCG, irradiated M. bovis, or M. tuberculosis. 24 hours post-infection, samples were analyzed by RNAseq, ChIP-seq, and ATAC-seq to look at alterations in the transcriptome and epigenome. Through fluorescence imaging-based analysis, the authors show equivalent infection (bacterial uptake) in all groups, except for the irradiated M. bovis. Principal component analysis of the transcriptomic data demonstrated strong segregation for the M. bovine-infected cells compared to the other groups, which had some intermixing in the PCA. Among the differentially expressed genes, the cytokine IL36G was significantly upregulated across all four groups. This is significant as this cytokine enhanced autophagy in macrophages and subsequent M. tuberculosis killing activity. To further investigate the transcriptional changes, ChIP-seq and ATACseq were utilized to investigate chromatin changes in the form of differential affinity binding sites (DABS) and differential open chromatin regions (DOCR). TPMRSS2, a protease that plays an important role in multiple types of infections (tuberculosis, COVID, etc.), was found to be a significant DOCR in both the M. bovis and M. tuberculosis challenge groups, conferring enhanced TPMRSS2 expression in both of these groups. Using the integrative approach between all the omics data, the authors found that CD274, which encodes the PD-L1 protein, was upregulated in all four groups. PD-L1 is known to play an immunosuppressive role, and PD-L1+ macrophages have been shown to create "cold" microenvironments that would likely favor the mycobacterium. Lastly, SNP analysis found variants in four genes (ERBB4, LRCH1, MRTFA, RNPC3) that could serve as susceptibility genes.

      Strengths:

      Overall, this study demonstrates that challenge with M. bovis elicits a more extensive remodeling of chromatin and subsequent gene expression changes in macrophages compared to the other closely related strains. The work demonstrates the power of functional genomics and its utility in investigating the underlying changes that can affect responses to infections and subsequent outcomes. While the study lacks functional validation, the cohesive dataset is quite compelling, and from it, the authors draw conservative conclusions and are frank about their study limitations.

      Weaknesses:

      Lack of functional validation of some of the targets found.

    1. eLife Assessment

      This important study investigates how verbal and nonverbal working-memory processing is distributed across large-scale functional networks in the human brain using precision fMRI. By leveraging extensive within-subject data and individualized network mapping, the authors provide solid evidence that hemispheric specialization for verbal versus nonverbal information extends across multiple association networks and is reproducible across independent datasets. The use of state-of-the-art precision neuroimaging approaches reveals fine-grained laterality patterns that are likely obscured in conventional group analyses.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Sun et al. investigate the hemispheric lateralization of functional brain networks during verbal versus nonverbal working memory tasks. Utilizing state-of-the-art precision neuroimaging in highly sampled individuals, the authors define a set of distributed association networks and examine their task-evoked responses. The authors report a "generalized laterality effect," wherein multiple association networks appear to functionally split across the hemispheres, with left hemisphere components exhibiting a relative preference for verbal stimuli and right hemisphere components preferring nonverbal stimuli.

      Strength:

      The use of dense-sampling fMRI is a major strength of this study, allowing for a highly accurate, individual-specific mapping of network topologies that group-averaging typically obscures. Despite the interpretational concerns raised below, this high-quality, within-subject imaging dataset represents a valuable resource for the community. Furthermore, the inclusion of an independent prospective replication dataset provides valuable confidence in the robustness of the core imaging metrics. However, while the data quality is exceptionally high, the conceptual interpretations regarding "network splitting", "preferential recruitment", and the generalizability of the verbal/nonverbal dichotomy require significant refinement. Several methodological and statistical clarifications are needed to fully support the authors' claims.

      Weaknesses:

      Major:

      (1) The manuscript relies heavily on task contrast values (e.g., Face > Word) to conclude that networks functionally "split" their profiles, with specific hemispheres being "preferentially recruited" by either verbal or nonverbal materials. While the data clearly demonstrate relative hemispheric differences, claiming absolute bidirectional specialization and active recruitment appears to overstate the findings in two key ways:

      First, statistical evidence for true bidirectional "splitting" is scarce. A significant hemispheric difference confirms a relative shift in processing, but it does not permit claims about absolute preference. When examining the face>word effects against zero in the discovery dataset (Figure 3), the right hemisphere of the LANG, FPN-B, CG-OP, and SAL networks shows no significant preference for faces over words. In the replication dataset (Figure 6), two of the four targeted networks (FPN-A, CG-OP) similarly fail to show a significant right-hemisphere preference. Furthermore, it is unclear whether these tests against zero were corrected for multiple comparisons (e.g., 18 individual tests in Figure 3). Networks reporting significance at the uncorrected p < 0.05 level (such as the left hemisphere of LANG and FPN-B) might not survive standard correction, suggesting that even the left hemisphere's preference for words may be statistically marginal. Second, contrast differences in networks exhibiting negative signals may reflect relative deactivation rather than active recruitment. A mathematically positive contrast value derived from two negative activation states (e.g., Face [-6] > Word [-8]) does not indicate active "recruitment" for face processing. Instead, it merely reflects a relative difference in deactivation. Characterizing this dynamic as "preferential recruitment" misleads the reader regarding the actual physiological state of the network. Furthermore, such asymmetric suppression is frequently driven by generalized differences in task difficulty or cognitive effort, rather than true stimulus-specific processing.

      (2) Building on the previous point, it would be highly beneficial to include the behavioral data (such as accuracy and reaction times) for the N-Back conditions, which do not currently appear to be reported in the manuscript. This information is important because if one condition (e.g., the Face N-back) was significantly more challenging or required greater cognitive effort than the other (e.g., the Word N-back), the observed hemispheric dissociations might reflect differences in arousal, effort, or attentional deployment rather than stimulus-specific processing. ion

      (3) The manuscript claims a "generalized" hemispheric laterality effect. However, the supplemental figures suggest this effect may be highly sensitive to the specific stimuli used in the main text (unfamiliar Faces vs. rhyming Words), which represent extreme ends of visuospatial and phonological processing. As we talked about earlier, a true functional "split" implies that the hemispheres respond in opposite directions. However, visual inspection of the supplemental graphs reveals that for the vast majority of networks, both hemispheres are actually driven in the exact same direction (Figures S8 and S9).

      In the Face > Letter contrast (Figure S8), true bidirectional splits largely disappear. With the exception of FPN-A, the bars for both the left and right hemispheres point in the exact same visual direction for every network (e.g., both hemispheres are visually positive in DN-A and dATN-B, and visually negative in CG-OP and dATN-A). This indicates that the hemispheres actually share the same categorical preference and merely differ in magnitude. Strikingly, the LANG network shows no significant difference between Faces and Letters in either hemisphere, suggesting that the robust leftward shift observed in Figure 3 was possibly driven by the heavy semantic and phonological demands of the rhyming task, rather than a generic preference for "verbal" processing.

      Similarly, in the Scene > Word contrast (Figure S9), the hemispheres do not visually diverge in their response direction for most networks. For example, both hemispheres are visually negative (indicating a shared preference for Words) in the LANG, FPN-B, CG-OP, and SAL networks. Conversely, both hemispheres are visually positive (indicating a shared preference for Scenes) in the dATN-B and DN-A network.

      Because the hemispheres do not visually diverge in their response direction for most networks across these supplementary contrasts, the claim of a robust, generalized hemispheric "split" is unsupported.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Sun and colleagues use precision fMRI to investigate the spatial patterns of verbal versus nonverbal processing in the cortex. They first show that the left hemisphere tends to be more active for word than face processing, whereas the right hemisphere tends to be more active for face than word processing during a working memory task. They then show that this is not confined to a particular network specialized for verbal vs. nonverbal processing but is a pan-network hemispheric difference across 8/9 association networks. This was unexpected, as prior precision imaging work emphasized network specialization and the authors had hypothesized that verbal vs. nonverbal processing would show bilateral network-level effects in a single, right-lateralized network (FPN-B). They then replicated this pan-network effect in an independent dataset.

      Strengths:

      (1) This paper is neat, clear, and succinct. The authors convincingly show that there exists a hemispheric difference in face vs. word processing during a working memory paradigm across many association networks.

      (2) They replicate this result in a fully independent dataset. They do a particularly nice job setting up a prospective replication analysis over 4 distinct association networks.

      (3) They do a wonderful job framing the experiment - they provide context on how previous precision imaging findings would have suggested a specialized network for verbal vs. nonverbal processing, they replicate major findings from that work in this data (network laterality and bilateral network activation in multiple task contexts), and then show the surprising pan-network hemispheric laterality in two independent datasets.

      (4) This paper adds a unique perspective to precision imaging findings. Most precision imaging has shown that different networks seem to be specialized for different cognitive tasks - individual task activations tend to follow network boundaries and networks tend to be cohesively activated. These types of findings led to the idea that previously reported broad hemispheric effects for verbal vs. nonverbal processing could instead be due to a lateralized network specialized for verbal versus nonverbal processing. However, this paper accounts for individualized network topography and yet finds a hemispheric effect that transcends individual networks. I think this result will have a strong influence on how many readers think about network specialization.

      Weaknesses:

      (1) The evidence for the main result comes only from a single fMRI task (n-back) with a particular set of stimuli (faces, words) in which processing demands are not matched across verbal and nonverbal conditions (word blocks use rhyming, face blocks use an exact match). While the claim made is fairly broad (hemispheric laterality in verbal vs. nonverbal processing), it isn't clear how well this finding would generalize to other types of stimuli or processing. This was mentioned in the limitations section, but the paper would be strengthened by both additional evidence for a general verbal versus nonverbal processing effect and by a discussion of how these specific stimuli (words and faces) and processing demands (differences between rhyming and exact match) might affect the results. The authors did include supplemental post-hoc analyses of scene and letter conditions that are matched in processing, but did not discuss them in the main text.

      (2) While lateralization is seen in 8/9 networks, the effect is much larger in some networks (FP-A, DATN-A) versus others. Similarly, the face > word effect is much stronger in some regions of the cortex than others. For example, all the individuals shown exhibit a patch of RH mid-LPFC cortex with particularly strong face > word preference compared to anywhere else in association cortex and an analogous LH mid-LPFC region that shows a particularly strong word > face preference. While the authors are correct that lateralization is seen across many networks, there does seem to be a topography to it rather than a diffuse hemispheric difference. I wonder to what degree the general hemispheric laterality pattern could be driven by a subset of regions that happen to cross several association networks. Neither the differences between networks nor the finer-scale activation patterns are considered in this paper. The paper would be strengthened by considering them.

      (3) The paper is focused on association networks and provides an interesting report on face versus word processing in those networks specifically. In the brain maps, it is possible to see face versus word preference in non-association regions as well. The authors don't discuss these patterns, but the paper would be strengthened by describing the relationship of these patterns with known face and language processing systems in the sensory cortex.

    4. Reviewer #3 (Public review):

      Summary:

      This work takes a precision neuroscience approach to examining how different task demands elicit lateralized vs. bilateral activity in large-scale brain networks. Using ~35-150+ minutes of resting state data per person, the authors identify individualized networks and map their degree of laterality. They find that all association networks are bilateral, with some networks. such as the language network (left) and fronto-parietal B network (right), showing some marked degree of laterality. A sentence reading task evokes activity bilaterally in the language networks, while working memory load during an n-back task evokes bilateral activity largely in fronto-parietal, dorsal attention, and salience networks. Interestingly, though, the authors find that when digging into the N-back task and contrasting blocks that have rhyming words vs. faces, this elicits a more lateralized pattern of activity; left activation for rhyming and right activation for faces. This lateralization is seen across multiple association networks, not just the language or fronto-parietal networks. These findings are then replicated in another precision data set.

      Strengths:

      This work has several notable strengths. This study boasts a lot of data for each individual, allowing them to examine individualized functional networks and task activations. Given the marked individual differences in laterality (see Figures S4 and S5), a group-averaged network atlas and group-averaged activation maps would likely muddle some of the interesting effects.

      The authors also use an elegant approach for individualized network estimation, multisession hierarchical Bayesian modeling, or MSHBM. MSHBM allows vertex-based functional connectivity to steer individualized networks while also incorporating meaningful priors to identify comparable networks across individuals. This is a nice approach combining elements of a common atlas across individuals with data-driven approaches.

      The authors have a rich working memory N-back task with four different stimuli types with which they can examine processing different types of stimuli.

      A major strength of this work is the replication. The authors take their surprising findings and replicate the entire experiment in another sample. The replication sample notably has even more data per person than the discovery sample, allowing for robust and precise estimation of individual networks and task activity.

      Weaknesses:

      I'd like to frame this section as 'unsolved challenges' rather than 'weaknesses'.

      One of the strengths of this paper is that the working memory task that incorporates so many different stimuli and conditions also poses a caveat for interpreting the task stimuli effects. The 'word' condition requires multiple cognitive demands - working memory and rhyming/phonological processing. This makes it a little hard to draw complete parallels between the word and face conditions. It does, however, provide a useful backdrop to explain why there are such strong left laterality effects for the word condition. I would expect that explicit phonological processing that is required for a rhyming task would elicit more left lateralized activation than the sentence processing task, as fluent readers are likely not using explicit phonological processing for sentence processing. This could explain why, for the sentence reading task, they find more bilateral activation, but a task that specifically targets phonological processing (i.e., the 2-back word condition) would evoke more left lateralized activation above and beyond the activation supporting working memory, which is held constant across both the 'words' and 'faces' conditions. Ultimately, I believe this supports the broad idea of this work - that large-scale association networks can show lateralization, even beyond language network boundaries. However, the fact that there is a dual-task load that evokes phonological processing for the word condition is important for contextualizing these findings.

      When looking at plots of individual differences, it is clear that some people are more bilateral than others in certain networks (and sometimes overall). Because of the amount of data per person in this precision study, it naturally limits the overall sample size (N=29). This makes it difficult to interrogate individual differences in laterality and what might predict this effect across people.

    1. eLife Assessment

      Murrell et al. describe a high-throughput method, the Feeding Experimentation Device (FED3), to study food foraging strategies in mice. The authors provide solid evidence about key sex differences in foraging strategies, specifically the finding of greater male win-stay behavior. Further consideration of how sex differences could be uniquely influenced by FED3 testing conditions (e.g., single-housing, hormones) and task demands (e.g., 100-0 vs. 80-20) would be helpful. Given this open source FED3 platform, the authors provide valuable findings that have utility to the field of behavioral neuroscience, specifically to those interested in studying reward-related behavior.

    2. Reviewer #1 (Public review):

      Summary:

      This paper examines potential sex differences in the conflict between exploitation, pursuing food and rewards in previously-associated locations/paradigms, or exploration of new locations that might result in better outcomes. Dysregulation of this conflict may be an underlying behavioral modality of psychiatric diseases. They used four distinct tasks: a two-armed Bandit 100:0 task, a standard fixed ratio 1 task, a two-armed Bandit 80:20 task, and a closed-loop economy PR1 task that allows for the assessment of motivational breakpoint.

      Male mice show significantly higher accuracy under conditions of high probability known rewards, sticking with an action that just resulted in a reward or "win-staying". This was demonstrated in multiple paradigms, and there was a predictive nature of this behavior that could predict animal sex with modest accuracy. Under probabilistic environments, males were no longer more accurate than females but still used a higher win-stay strategy. A closed-loop PR1 task showed that there were no inherent differences in motivational breakpoint between sexes. Finally, the authors use simulations to determine an appropriate number of animals needed to detect these differences.

      Strengths:

      The manuscript attempts to resolve inconclusive sex differences that have heretofore been neglected or inconclusive due to insufficient power. The most impressive aspect of this paper is its scale, assaying 62 female mice and 74 male mice in identical exploration-exploitation tasks using high-throughput and noninvasive operant feeding via FED3. Very few labs can achieve this scale, which is necessary to detect sex differences with a small effect size.

      The authors use some sophisticated modeling approaches and analysis of data from the 136 mice to investigate the significance of these sex differences and interrogate other conditions. They also use simulations to model the likelihood of replicating these differences given a sample size. This is extremely helpful for other researchers as they consider sex as a biological variable.

      Weaknesses:

      The study is largely descriptive in nature and does not pursue any mechanism of the underlying differences, like hormones, neuromodulators, or circuits. The lack of estrous cycle tracking is acknowledged as a limitation.

    3. Reviewer #2 (Public review):

      Summary:

      Murrell and colleagues examine sex differences in mouse decision-making tasks, using the FED3 device, which allows for continuous data collection in the home-cage. Mice performed four tasks across two weeks, which provided all of their food. Across tasks, male mice were more likely to repeat a rewarded choice than females, which benefits decision accuracy in deterministic tasks. This work complements existing results for decision-making differences in males and females, affirming that this domain of cognition is particularly sensitive to sex differences. However, there are some specific features of the FED3 device, such as single housing, closed economy feeding, and 24-hour access that can uniquely influence decision-making in a (likely) sex-dependent manner that encourage considering these data as examining sex differences in a particular context, rather than as a generalized finding. At the same time, these data could offer new insights about nuances of behavior like circadian rhythms or bout analysis, uniquely enabled by the extended availability of the FED3 devices. The analyses in this paper also make an important point, encouraging researchers to use methods that allow for much larger N's to provide clearer and more robust results.

      The FED3 devices are an innovative new way to approach behavior, and have allowed the authors to test many dozens of mice in a battery of tasks, over which they see similar patterns of increased win-stay behavior in male C57b6 mice (wildtypes from several knockout lines). The authors point out that there are discrepancies in prior literature across tasks and species in terms of how sex differences influence decision making, but there are some particular ways that sex differences could interact with the FED3 devices that it would be interesting and important to consider further. In particular, the fact that the animals live with the device, singly housed, may be an underrecognized contributor to sex differences. Changes in social interaction and dominance arising from long-term single housing are very likely to impact males and females differently, for example.

      Continuous data collection is a fascinating way to look at learning and decision-making, but it also raises interesting questions about whether these dynamics are impacted as a function of continuous access to the device. In addition to summary metrics over the whole task, it might be valuable to look at learning across the task each day, and within circadian periods of each day. For example, it seems based on the example sessions for the 100-0 bandit task that animals take at least a few reversals to learn the task structure. How many trials does it take for a male or female mouse to reach some criteria of success? Do the sex differences exist at all time points? Does the light cycle affect the accuracy or trial counts? There are numerous such analyses that could particularly inform future use of the FEDs across laboratories, and identification of similar or distinct patterns of sex differences or behavior in other apparatuses, and would be a benefit to the field.

      The authors employed several computational techniques to identify parameters or features of behavior that might explain the sex differences they observed, and this is a strength of the manuscript. However, the win-stay lose-shift agent may not be an ideal match to make conclusions about exploitation, as it is unclear how win-stay and lose-shift strategies map onto explore/exploit tradeoffs. If an animal were exploiting an option, they may win-stay *and* lose-stay, if the task is probabilistic. Indeed, the model fit is weaker for the 80-20 bandit, suggesting this model may not reflect the actual strategies mice are engaging in, potentially in both bandit tasks. This point is particularly worth considering in light of the criteria for shifts in the tasks being based not on the number of trials completed, but on the number of pellets earned. When shifts are tied to reward collection rather than trials, it can amplify differences in behavior driven by reward consumption.

    4. Reviewer #3 (Public review):

      Summary:

      The manuscript by Murrell et al. describes a high-throughput approach for evaluating food foraging strategies in mice. Building on their prior publication describing the technical aspects of the Feeding Experimentation Device 3 (FED3), this study demonstrates the utility of the FED3 in evaluating decision-making in mice. The authors identify key differences in male and female foraging strategies that could not be accounted for by total food consumption or overall food motivation. Given the rapid adoption of the open source FED3 platform, this work is likely to be of broad interest and utility in the field.

      Strengths:

      (1) The use of cost-effective, open-source devices like FED3 provides substantial value to the scientific community. Validation of appropriate conditions for using this equipment is an important step toward broad adoption.

      (2) The authors implement a simple but elegant experimental design for studying food-motivated decision-making behavior. This approach could be applied to a wide range of preclinical disease models in future studies.

      (3) The study is well-powered to evaluate sex as the primary experimental factor (62 females, 74 males), allowing the authors to make convincing claims about differences in strategy. Additionally, the dataset provides a useful benchmark for power analyses in future studies involving more complex experimental designs.

      (4) The figures are clear and generally easy to interpret the primary findings.

      (5) The conclusions are appropriate and not overstated

      Weaknesses:

      (1) A major strength of this study is the potential utility for new investigators trying to implement cognitive behavioral tasks in mice. However, the present version provides limited background on the rationale for selecting the bandit task and on prior work applying it in similar contexts. Including additional background and discussion would better contextualize the approach for other groups considering adopting it in their own studies.

      (2) Some methodological details surrounding the initiation of the experiment could be clarified. Specifically, it is unclear if mice transitioned directly from standard housing conditions (group housed, standard chow) to the study conditions (single housed, FED3-based probabilistic learning), or intermediate acclimation/training steps were used, such as autoshaping, free access to new food pellets, or FR1 training. A more detailed experimental timeline (for example, see Figure 1 from PMID 39710132) would address this concern.

      (3) The authors evaluated multiple probabilistic conditions (100%, 90%, 80%, 70%, 60%), but ultimately focused on the 80% condition for this study. A more detailed explanation for how this conclusion was reached would be useful for future researchers working under different experimental conditions (i.e., age, strain, genotype, disease model) where other probabilistic conditions may be more appropriate.

    1. eLife Assessment

      This study sets out to show that the directionality of bacterial transport in confined environments is affected by a salt gradient. Using Pseudomonas putida in microfluidics devices, the authors hypothesize that salt-induced diffusiophoresis acts as a steering mechanism to guide microbial movement. However, the evidence to support this hypothesis is incomplete, because the underlying reorientation process is not directly demonstrated, and the necessary controls to rule out alternative hypotheses are not presented.

    2. Reviewer #1 (Public review):

      Summary:

      The authors claim that bacteria are guided by diffusiophoresis. They perform experiments of bacterial motility in microfluidic channels with salt gradients. The data show that P. pudita bacteria swim towards higher sodium chloride concentrations, but there is no evidence that this is due to a diffusiophoresis.

      Weaknesses:

      It is well known that bacteria perform chemotaxis in salt gradients (see e.g., PNAS 86, pp. 8358-8362, 1989). The underlying mechanism based on chemoreceptors is widely accepted, but the authors do not mention this possibility. I recommend a control experiment where the chemotaxis genes are knocked out. Even if this mechanism can be ruled out, the current data show no evidence for a mechanism based on diffusiophoresis.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigate how salt gradients influence the transport of Pseudomonas putida in confined microfluidic environments. They report that salt gradients enhance directional migration, increase run persistence, and promote transport toward contaminant-rich regions. To explain these observations, the authors propose a physical steering mechanism in which differential diffusiophoretic mobilities of the cell body and flagellar bundle generate an aligning torque that reorients cells along the salt gradient.

      Strengths:

      The study addresses an interesting question at the interface of microbiology, complex fluids, and active matter. Their experiments suggest that salt gradients influence bacterial transport behavior and lead to more persistent, directional motion. Once confirmed, the proposed mechanism would broaden our understanding of how environmental gradients can shape microbial migration through physical interactions in addition to more traditional sensing-based pathways.

      Weaknesses:

      The main limitation of the current study is that the proposed steering mechanism is not directly demonstrated. The evidence for the diffusiophoretic torque is largely inferred from trajectory statistics and theoretical modeling. While the observed transport behavior is convincing, the causal link between the observed migration patterns and the proposed reorientation mechanism remains less well established. In particular, the manuscript focuses primarily on cell trajectories and transport properties, whereas the proposed mechanism fundamentally involves changes in cell orientation. Additional evidence connecting orientation dynamics to the proposed torque mechanism would strengthen the conclusions.

      A related concern is whether alternative physical mechanisms associated with the imposed salt gradients have been fully excluded. For example, weak flow-mediated effects or other hydrodynamic influences could potentially contribute to the observed transport behavior. The manuscript would benefit from a more thorough discussion of such possibilities and a clearer justification for why the proposed diffusiophoretic mechanism should be regarded as the dominant explanation.

      The manuscript would also benefit from a clearer positioning within the broader literature on physically induced microbial transport and swimmer reorientation. Previous studies have demonstrated directed migration arising from rheotaxis (Marcos et al., 2012, PNAS) and viscosity-gradient-induced steering (Stehnach et al., 2021, Nature Physics). While the mechanism proposed here appears distinct, a more explicit discussion of how the present work relates to these earlier studies would help readers better understand the specific conceptual advance being made.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors claim that bacteria are guided by diffusiophoresis. They perform experiments of bacterial motility in microfluidic channels with salt gradients. The data show that P. putida bacteria swim towards higher sodium chloride concentrations, but there is no evidence that this is due to diffusiophoresis.

      Weaknesses:

      It is well known that bacteria perform chemotaxis in salt gradients (see e.g., PNAS 86, pp. 8358-8362, 1989). The underlying mechanism based on chemoreceptors is widely accepted, but the authors do not mention this possibility. I recommend a control experiment where the chemotaxis genes are knocked out. Even if this mechanism can be ruled out, the current data show no evidence for a mechanism based on diffusiophoresis.

      We thank the reviewer for raising this important comment. We agree that receptor-mediated salt taxis is a well-established mechanism in some bacteria, including the classic study by Qi and Adler (PNAS, 1989). We note, however, that the original manuscript did discuss this possibility and cited Qi and Adler in line 117: “We also note that we did not observe any significant difference in the tumble rates between the control and NaCl gradient cases (Figure 3h; Figure S2, SM), suggesting that NaCl gradients do not interfere with chemoreceptors (Qi and Adler, 1989).” In Figure S2, the run-time distributions show no significant difference between the no-gradient condition and the NaCl-gradient condition, with fitted tumble rates of λ = 0.41 s<sup>−1</sup> and λ = 0.38 s<sup>−1</sup>, respectively. We also do not observe a directional bias in run duration, namely longer runs up the salt gradient and shorter runs down the gradient, which would be expected for canonical chemoreceptor-mediated taxis.

      The basis for assigning the observed migration to diffusiophoresis is that the NaCl gradient produces a directional drift of the bacterial body without a measurable change in the run-and-tumble statistics. This behavior is consistent with our previous work [1], where non-motile bacteria were shown to undergo diffusiophoretic migration toward higher salt concentration. Because that migration occurred in non-motile cells and across different bacterial types and morphologies, it supports the interpretation that native bacterial surface charge can drive a non-specific diffusiophoretic response in salt gradients.

      That said, we agree with the reviewer that a genetic control would provide a stronger test against receptor-mediated chemotaxis. We will therefore perform additional experiments using a ∆cheA strain. Because CheA is required for canonical chemotactic signal transduction, observing the same directional migration in the ∆cheA mutant would directly test whether the NaCl gradient response persists in the absence of receptor-mediated chemotaxis. We will include these new data and revise the manuscript to more explicitly distinguish diffusiophoretic drift from chemoreceptormediated salt taxis.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate how salt gradients influence the transport of Pseudomonas putida in confined microfluidic environments. They report that salt gradients enhance directional migration, increase run persistence, and promote transport toward contaminant-rich regions. To explain these observations, the authors propose a physical steering mechanism in which differential diffusiophoretic mobilities of the cell body and flagellar bundle generate an aligning torque that reorients cells along the salt gradient.

      Strengths:

      The study addresses an interesting question at the interface of microbiology, complex fluids, and active matter. Their experiments suggest that salt gradients influence bacterial transport behavior and lead to more persistent, directional motion. Once confirmed, the proposed mechanism would broaden our understanding of how environmental gradients can shape microbial migration through physical interactions in addition to more traditional sensing-based pathways.

      Weaknesses:

      The main limitation of the current study is that the proposed steering mechanism is not directly demonstrated. The evidence for the diffusiophoretic torque is largely inferred from trajectory statistics and theoretical modeling. While the observed transport behavior is convincing, the causal link between the observed migration patterns and the proposed reorientation mechanism remains less well established. In particular, the manuscript focuses primarily on cell trajectories and transport properties, whereas the proposed mechanism fundamentally involves changes in cell orientation. Additional evidence connecting orientation dynamics to the proposed torque mechanism would strengthen the conclusions.

      We thank the reviewer for the constructive comments. We agree that the proposed steering mechanism should be supported by evidence that directly connects the salt gradient response to bacterial orientation dynamics, not only to trajectory-level transport statistics.

      We would like to clarify that the original manuscript already includes an orientation-based analysis in Figure 4f,g in the main text. The corresponding methodology and results are described in lines 173–183 and in the Supporting Information. Specifically, we quantified cell steering by measuring the change in body angle, ∆θ, along individual run trajectories as a function of arc length, s, using the orientation correlation ⟨cos(∆θ)⟩<sub>s</sub>. In the absence of salt gradients, the orientation correlation decays slowly with arc length, indicating persistent swimming along the initial

      Author response image 1.

      Instantaneous angular velocity as a function of heading angle relative to the salt gradient orientation. (a) Experimental and (b) simulated mean angular velocity of cells as a function of heading angle θ (measured relative to the gradient direction; θ = 0° points toward the gel/high-salt side) in the absence (blue) and presence (red) of a NaCl gradient. Positive and negative values indicate counterclockwise and clockwise rotation, respectively, with arrows showing rotation direction. Under the gradient, both experiments and simulations show a signed, angle-dependent rotation rate that is largest near θ = ±90° and approaches zero near θ = 0° and 180°, consistent with a restoring torque that steers cell heading toward the gradient direction. Simulations reproduce this behavior with comparable magnitude to the experimental measurements, and in the absence of a gradient, angular velocity remains relatively small in the no-gradient case with no consistent directional bias across heading angles.

      run direction. Under a salt gradient, the correlation decays more rapidly, indicating stronger directional reorientation during runs. Because the tumble statistics do not change significantly between the no-gradient and salt-gradient conditions, this enhanced orientational decorrelation is not attributed to increased tumbling or rotational noise. Instead, it is consistent with continuous deterministic steering during runs, as expected from a diffusiophoretic torque acting on the cell body–flagellar bundle system.

      To further address the reviewer’s concern, we performed an additional orientation-dynamics analysis using the same dataset shown in Figure 4. Following the approach used by Stehnach et al. [2], we calculated the instantaneous angular velocity during individual runs as a function of the cell heading angle relative to the salt gradient direction. The cell orientation was obtained from the run trajectories, and tumble events were excluded because they produce large transient angular velocity spikes that are not representative of continuous steering during runs.

      The new analysis is shown in Author response image 1. Under the no-gradient condition, the angular velocity remains small and nearly independent of heading angle. In contrast, under the salt gradient condition, the angular velocity becomes strongly heading-dependent. The angular velocity is largest when cells swim nearly perpendicular to the salt gradient, where a steering torque is expected to be maximal. Moreover, the sign of the angular velocity indicates rotation toward alignment with the gradient direction. This behavior is consistent with the proposed diffusiophoretic torque mechanism and provides a direct link between the observed transport behavior and salt-gradient-induced reorientation dynamics.

      We plan to add this angular velocity analysis to the revised manuscript and revise the relevant text to make the connection between trajectory statistics, orientation dynamics, and the proposed torque mechanism clearer.

      A related concern is whether alternative physical mechanisms associated with the imposed salt gradients have been fully excluded. For example, weak flow-mediated effects or other hydrodynamic influences could potentially contribute to the observed transport behavior. The manuscript would benefit from a more thorough discussion of such possibilities and a clearer justification for why the proposed diffusiophoretic mechanism should be regarded as the dominant explanation.

      We thank the reviewer for raising this important point. We agree that alternative physical mechanisms associated with the imposed salt gradient should be considered explicitly. In the revised manuscript, we will add a more detailed discussion explaining why flow-mediated or other hydrodynamic mechanisms are unlikely to account for the observed steering behavior.

      First, the characteristic diffusiophoretic drift velocity in our experiments is approximately u<sub>d</sub> ≈ 1 µm/s, corresponding to only about 1–5% of the typical swimming speed of P. putida. Thus, the proposed mechanism does not require externally driven advection of the cells. Instead, the salt gradient produces a weak but persistent differential diffusiophoretic slip on the cell body and flagellar bundle, which can generate a reorienting torque during active swimming.

      We also considered whether diffusio-osmotic flow along the channel walls could generate sufficient shear to induce rheotaxis. In a dead-end channel, the diffusio-osmotic velocity profile can be estimated as [3]

      which gives a wall shear rate (at z = h) of

      This value is below the shear rate threshold reported by Marcos et al. [2], where rheotactic drift becomes negligible for S < 0.1 s<sup>−1</sup>. Therefore, the shear generated by diffusio-osmotic wall flow in our experiments is too weak to explain the observed directional reorientation. This effect would be even smaller for P. putida, whose thin flagellar bundle is expected to experience weaker shear-induced alignment than organisms with larger flagellar structures.

      We further considered viscotaxis as a possible mechanism. However, the viscosity difference between 1 mM and 100 mM NaCl solutions is marginal. This contrast is far smaller than the approximately 4–5-fold viscosity difference reported to induce strong viscophobic turning in bacteria [4]. Thus, salt-gradient-induced viscosity variations are insufficient to account for the measured steering response.

      Taken together, these estimates indicate that hydrodynamic shear, rheotaxis, and viscotaxis are too weak under our experimental conditions to explain the observed migration and orientation dynamics. In contrast, the proposed nonuniform diffusiophoretic mechanism naturally accounts for the key observations: directional migration up the salt gradient, enhanced curvature during runs, heading-dependent angular velocity, and the absence of significant changes in tumble statistics. We will include this analysis in the revised manuscript to clarify why diffusiophoretic steering is the dominant mechanism under the present conditions.

      The manuscript would also benefit from a clearer positioning within the broader literature on physically induced microbial transport and swimmer reorientation. Previous studies have demonstrated directed migration arising from rheotaxis (Marcos et al., 2012, PNAS) and viscosity-gradient-induced steering (Stehnach et al., 2021, Nature Physics). While the mechanism proposed here appears distinct, a more explicit discussion of how the present work relates to these earlier studies would help readers better understand the specific conceptual advance being made.

      We thank the reviewer for this helpful suggestion. We agree that the manuscript should more clearly position the proposed mechanism within the broader literature on physically induced microbial transport and swimmer reorientation.

      In the revised manuscript, we have added a discussion comparing our results with prior studies on rheotaxis and viscosity-gradient-induced steering. Specifically, we now discuss the work of Marcos et al. [4], which showed that shear flow can generate a torque on the helical flagellum of Bacillus subtilis, producing rheotactic alignment independent of chemical sensing. We also discuss the work of Stehnach et al. [2], which showed that viscosity gradients can steer Chlamydomonas reinhardtii through asymmetric viscous drag on its two flagella, producing viscophobic turning down the viscosity gradient.

      The mechanism proposed in the present work is distinct from both of these cases. Unlike rheotaxis, it does not require externally imposed shear flow. Unlike viscophobic turning, it does not rely on a substantial viscosity contrast. Instead, we propose that a salt concentration gradient generates differential diffusiophoretic motion of the cell body and flagellar bundle, producing a torque that continuously reorients swimming cells along the gradient. This mechanism therefore identifies salt gradients as a distinct physical cue capable of steering bacteria through surface-mediated transport rather than through flow, viscosity contrast, or canonical chemoreceptor signaling.

      We have added the following text to the revised manuscript: ”Our findings also relate to other physical mechanisms of microbial reorientation. Bacterial rheotaxis, arising from a torque generated by shear flow acting on the helical flagellum, steers Bacillus subtilis independent of any chemical gradient [4]. Viscosity gradients similarly drive a viscophobic turning in the alga Chlamydomonas reinhardtii, where uneven viscous drag on its two flagella produces a torque that reorients cells down the gradient [2], a behavior confirmed by measuring angular velocity as a function of heading angle and revealing a sinusoidal form, ω(θ) = −ω <sub>visc</sub> sin(θ), which resembles what we report here for diffusiophoresis in (Figure S4). This identifies salt gradients, independent of flow or viscosity, as a distinct physical route by which swimming cells can be steered and guided, broadening the set of known nonchemoreceptor mechanisms for directed microbial transport.” References

      (1) V. S. Doan, P. Saingam, T. Yan, S. Shin, A trace amount of surfactants enables diffusiophoretic swimming of bacteria, ACS Nano 14 (10) (2020) 14219–14227.

      (2) M. R. Stehnach, N. Waisbord, D. M. Walkama, J. S. Guasto, Viscophobic turning dictates microalgae transport in viscosity gradients, Nature Physics 17 (8) (2021) 926–930.

      (3) S. Shin, E. Um, B. Sabass, J. T. Ault, M. Rahimi, P. B. Warren, H. A. Stone, Size-dependent control of colloid transport via solute gradients in dead-end channels, Proc. Natl. Acad. Sci. 113 (2) (2016) 257–261.

      (4) Marcos, H. C. Fu, T. R. Powers, R. Stocker, Bacterial rheotaxis, Proceedings of the National Academy of Sciences 109 (13) (2012) 4780–4785.

    1. eLife Assessment

      This important manuscript presents a way to understand allostery and allostery-inspired mechanical systems. It presents a mechanochemical model of an ATPase-like molecular machine in which geometric constraints and mechanically coupled linkages generate negative allostery, yielding a large stochastic reaction network that captures productive and futile cycling behavior. The work presents an interesting framework for bridging concepts from molecular motors and allostery. It provides a compelling, intuitive physical mechanism for mechanochemical transduction and reveals nontrivial tradeoffs that emerge from geometry-induced constraints.

    2. Reviewer #1 (Public review):

      Summary:

      This paper presents a creative physics-based model of an ATPase-like molecular machine using mechanically coupled linkages to mimic allosteric cycles. The authors construct the complete reaction network by enumerating all mechanically allowed binding configurations and transitions between them. The resulting system contains hundreds of microstates connected through node-level binding, dissociation, intramolecular rearrangements, cleavage, and ligation reactions. Stochastic simulations are then used to study how the machine cycles between ligand-bound and substrate-bound states. Overall, the manuscript presents an interesting and creative mechanochemical framework for modeling ATPase-like allosteric cycles integrating multivalent binding, geometric exclusion through rigidity arising from binding of substrates and ligands, with stochastic simulations.

      Strengths:

      (1) The manuscript presents a creative mechanochemical framework that combines multivalent binding, geometric exclusion, rigidity-based coupling, stochastic kinetics, and catalysis within a unified model.

      (2) The use of geometric exclusion and rigidity to generate negative allosteric coupling is elegant and provides an intuitive physical mechanism for coordinated molecular behavior.

      (3) The interpretation of catalysis as a transient release of mechanical constraints is conceptually interesting and offers a novel perspective on how energy-consuming reactions can regulate state transitions.

      (4) The distinction between productive and futile cycles is insightful and provides a useful framework for understanding pathway selection in molecular machines.

      (5) The explicit construction of a stochastic state network allows the authors to connect microscopic binding events with emergent cyclic behavior.

      (6) The work provides a conceptual platform for exploring how simple mechanical principles may give rise to allosteric regulation and mechanochemical transduction in synthetic molecular systems.

      Weaknesses:

      (1) The manuscript is very dense and difficult to follow. The notation and microstate labels (e.g., {S/L}:{10,5}) obscure the central ideas, and the stochastic model is not explained clearly enough. The authors should provide a simpler schematic, a mapping of state labels, and a step-by-step example of a productive cycle. The supplementary videos would also benefit from additional explanation.

      (2) The framework appears most applicable to mechanically gated motor proteins and may not generalize to allosteric enzymes that operate through conformational ensembles, dynamic coupling, or entropy-driven regulation. The scope of the model should be discussed more carefully.

      (3) The reported behavior appears highly dependent on specific parameter choices and rate hierarchies. A broader sensitivity analysis is needed to demonstrate robustness.

      (4) The binary treatment of states as either rigid or flexible oversimplifies the continuous energy landscapes and fluctuations observed in real biomolecules. The limitations of this approximation should be discussed.

      (5) The role of nonequilibrium thermodynamics is underdeveloped. The relationship between the model, ATP chemical potential, free-energy dissipation, entropy production, and the energetic cost of futile cycles should be discussed more explicitly.

      (6) Although a large number of microstates are enumerated, it remains unclear which states and pathways dominate the dynamics. A more coarse-grained analysis highlighting the key states and transitions would improve interpretability and facilitate comparison with experimental systems.

    3. Reviewer #2 (Public review):

      This is an interesting study aiming to capture the fundamental principles of ATPase-like machines with an elementary model of rigid bars. While molecular motors have been the subject of many studies in statistical physics, taking very simplified approaches, these past studies generally abstract away from geometrical constraints and do not account for allosteric mechanisms. In turn, several simple physical models of allostery are now available, but most consider only long-range effects without reference to reactivity or the conversion of chemical to mechanical energy. The essential ingredient of the present model is a form of negative allostery that stems from geometrical constraints. These constraints impose the presence of hidden microstates and the connections between states that form a reaction network. Nontrivial tradeoffs can then be derived, e.g., on the catalytic rate.

      One limitation of the approach is that energetic constraints, which are equally important, are not themselves derived from physical principles. This includes the different kinetic rates, the binding constants, and the mechanisms of reactivity, even though, in principle, they should follow from basic interaction energies in the context of thermal fluctuations. This is a legitimate choice of modeling level, handled notably by imposing a hierarchy between rates. It would be appreciated, however, if the overall logic of the derivation were clarified, presenting more clearly from the beginning what the fundamental hypotheses of the model are, which aspects are derived from these hypotheses, and which require additional assumptions. It seems indeed that the model involves both fundamental physical assumptions that are used to derive some emerging kinetic features (e.g., geometry imposes futile cycles) and global kinetic assumptions that are used to derive some microscopic features (e.g., a productive cycle imposes the relative values of the kinetic rates).

      The conclusion ends with a proposal to generalize to more elaborate models. I was wondering, however, if the opposite would not be desirable. The current model is already quite involved (as the 7 pages of SI listing transitions between states testify). Wouldn't it be possible to obtain some of the main results, e.g., the trade-offs defining an optimal cleavage rate, from an even simpler model?

    4. Reviewer #3 (Public review):

      Summary:

      In "Design of a minimal, allosteric, and ATPase-like machine using mechanical linkage," Omabegho introduces a simplified model of an allosterically-regulated molecular machine. The model machine takes inspiration from simple ATPase motors, specifically myosin or dynein, and is meant to capture the interactions between enzyme, ligand, and substrate that empower these enzymatic biomolecular machines. The model system attempts to specifically model the inhibitory allosteric regulation relevant to machine operation, following work on mechanical features of allostery in protein-analogue elastic network models. After the introduction of the model machine, the manuscript discusses the primary cycle by which the substrate serves to displace and then cyclically replace the ligand (roughly corresponding to a "recovery" and "power" stroke, respectively, in their biomolecular machine inspiration), as well as cataloguing other futile cycles or individual chemical pathways the machine may traverse. After an exhaustive enumeration of all possible states and transitions of the model machine, the actual mechano-chemical system is mapped to an abstracted stochastic chemical reaction network, which is finally studied numerically in more detail. We believe this model system is novel, interesting, and, importantly, interpretable and could prove to be useful in future modeling and rational design of simple bio-mimetic nanoscopic systems.

      Strengths:

      (1) Omabegho introduces a novel chemo-mechanical model system that captures the core behavior of a simple ATPase molecular machine. The model is relatively simple in construction, but, as the manuscript demonstrates, it can display very rich behavior depending on various competing chemical timescales and mechanisms. These features suggest that this system could, in fact, serve as a useful starting model if adopted by the wider community.

      (2) A key feature to highlight is the mechanistic interpretability of the model. By cutting to the seemingly core functional details of an ATPase-like machine, Omebegho is able to algorithmically produce an exhaustive listing of all allowed chemical states and all possible transitions between them, enabling the study of all possible mechanistic pathways and the relative frequency of each observed path. Further, Omebegho produced very clear visualizations of these states, transitions, pathways, and cycles that again facilitate the interpretability and utility of the model. This interpretability is key to our suggestion that this model could be more widely adopted as a clear ATPase analogue model for study by the broader biophysical community.

      Weaknesses:

      There are some issues to note.

      (1) First, although the model is inspired by ATPases, the cycle in question is not a model even for the significantly abstracted form presented in Supplemental Section 1.1, as noted in the manuscript. The allosteric regulation utilized in the model does not mandate that the ligand (a stand-in for actin, for example) displaces both products (stand-ins for ADP and P). Rather, one product (P) dissociates spontaneously, and the other, larger one (ADP) is allosterically displaced. We recognize mandating this small further detail would presumably complexify the model (i.e., it may require an additional tile in the enzyme construction), but it would also lead to a more accurate, while still simple and interpretable, picture of the biological molecular machines in question.

      (2) Much as we recognize and applaud the model's structure as simple and interpretable, when trying to actually study the chemical dynamics and observed dynamical behaviors, individual chemical rates of various binding and unbinding processes appear to be quite finely tuned. There is some attempt at studying the behavior of this model in various parameter tuning regimes, but the large number of model parameters makes a truly complete numerical study prohibitive. Given this difficulty, it would be instructive to know if a comparison could be made to the actual motivating biological systems and any observed chemical rates in the experiment. Further, given the desired cyclic behavior, it would be quite interesting to see to what degree this behavior is robust to various parameter choices. Again, preliminary work was done in the manuscript by biasing rates or numerical siloing experiments, but a far more exhaustive study is certainly worthwhile.

      (3) Perhaps our biggest critique of the manuscript is that, although it incorporates both mechanical and chemical aspects into the model system construction, all mechanical aspects of the model simply function to limit allowed state transitions. From our understanding, all mechanical aspects of the model are, in fact, abstracted away during any simulations. This modeling choice, of course, retains some mechanical inspiration while making the resulting system more tractable, but it is ultimately not an actual mechanical model. ATPases, especially motors such as myosin and dynein, are inherently mechanical, their fundamental features being the conversion of chemical fuel into mechanical motion. A fuller treatment, perhaps in future work, should include these physical degrees of freedom in simulation, thus truly tracking the interplay between physical enzyme mechanics and chemistry. The absence of actual mechanics as yet, beyond a strict allosteric restriction, is an inherent limitation of the model.

    1. eLife Assessment

      This study presents a useful application of trans-omic network analyses to existing human brain datasets, generating systems-level insights into metabolic dysregulation in Alzheimer's disease. Overall, the authors' analytical choices are solid, with appropriate use of existing data and methods, and with many of their results confirming previous findings. However, some of the authors' key claims, related to previously unknown details of regulatory relationships, are only partially supported due to limitations in dataset cell-type resolution and network robustness, as well as a lack of functional validation. This work will be of interest to cellular or systems neurobiologists studying Alzheimer's disease and could serve as a helpful starting point for future work.

      [Editors' note: this paper was reviewed by Review Commons.]

    2. Reviewer #1 (Public review):

      The authors aim to reconstruct a multi-layer metabolic regulatory network in Alzheimer's disease by integrating transcriptomic, proteomic, and metabolomic datasets from human brain tissue. By linking transcription factors, enzyme expression, metabolic reactions, and metabolites, the study seeks to provide a systems-level understanding of disease-associated metabolic dysregulation.

      The revised manuscript has improved in clarity and includes additional analyses in response to prior comments, particularly the incorporation of cell-type proportion estimates derived from matched single-cell datasets. This represents a meaningful step toward addressing concerns about the interpretation of bulk tissue and strengthens the study's descriptive rigor.

      However, several key limitations remain. First, although cell-type proportions are now estimated, this information is not integrated into downstream analyses. As a result, it remains unclear whether the observed metabolic changes reflect cell-intrinsic regulation or shifts in cellular composition. This distinction is critical for interpreting the inferred regulatory network and limits the strength of the conclusions.

      Second, the biological conclusions remain largely confirmatory. The reported downregulation of energy-related pathways, including the TCA cycle and oxidative phosphorylation, is consistent with prior literature. The trans-omic framework provides a structured representation of these changes, but the manuscript does not convincingly demonstrate that this approach yields new mechanistic insight beyond existing knowledge. In particular, the network is not sufficiently leveraged to identify novel regulatory relationships or generate testable hypotheses.

      Third, while the framework integrates multiple molecular layers, the contribution of upstream transcriptional regulation remains unclear. The most compelling findings appear to arise from protein and metabolite layers, and the manuscript does not clearly demonstrate how transcription factors or mRNA-level changes contribute to the interpretation of metabolic dysregulation. This weakens the claim that the study provides a fully integrated trans-omic perspective.

      Fourth, concerns regarding network robustness remain. The analysis relies on partially overlapping cohorts across omics modalities, and although this limitation is acknowledged, no formal sensitivity or robustness analyses are presented. This reduces confidence in the stability and generalizability of the inferred network structure.

      Finally, the handling of covariates remains limited. While the authors justify this based on data availability and prior studies, the lack of consistent adjustment for known confounders introduces uncertainty in attributing observed differences specifically to disease-related biology.

      Overall, the study presents a technically sound application of a trans-omic integration framework and provides a coherent overview of metabolic dysregulation in Alzheimer's disease. However, the findings are primarily descriptive and confirmatory, and the current analyses do not fully demonstrate that the approach yields novel biological insight. The work will be of interest to researchers in systems biology and multi-omics integration, but its impact on advancing understanding of disease mechanisms is likely to be moderate.

    3. Author response:

      General Statements:

      We appreciate the reviewers for the critical review of the manuscript and the valuable comments. We have carefully considered the reviewer’s comments and have revised our manuscript accordingly.

      Point-by-point description of the revisions:

      Reviewer #1 (Evidence, reproducibility and clarity):

      Major comments

      (1) This study leaves out lipid metabolism as a major energy metabolism pathway relevant to AD. The authors themselves cite the significance of acylcarnitines and CPT1A in AD (pg. 3, lines 32-33, pg. 4, lines 1-2). Lipid metabolism and homeostasis is known to be disrupted in AD1. Fatty acid oxidation is a known energy source in the prefrontal cortex2 and will also generate acetyl coA, which this study reveals is a significant decreased metabolite in AD. Furthermore, sphingomyelin emerges as one of the major decreased DEMs as well. Thus, lipid metabolism should be highlighted in Figure 3 and discussed throughout the manuscript; otherwise its omission should be clearly stated and justified.

      We appreciate the reviewer’s insightful comment regarding a critical role of lipid metabolism in AD. We recognize that lipid metabolism is a metabolic pathway deeply involved in AD pathology (Baloni et al., 2022, 2020; Varma et al., 2021). Accordingly, we have revised the Limitations section to more strongly emphasize its role as a vital energy source (pg. 13, lines 15-17). Regarding the visualization of lipid metabolism, we extracted lipid-related pathway from the trans-omic network but found that the regulatory relationships among DEPs and DEMs were excessively complex and interconnected. Thus, interpreting this regulatory network seemed to be more challenging compared to the other energy production pathways presented in our manuscript. Therefore, we have concluded that the pathway analysis in our trans-omic network may not be suitable for deeply elucidating the lipid dysregulation in AD. We have added a statement acknowledging this as a limitation of our current methodology in the revised manuscript (pg. 13, lines 13-22).

      (2) The covariates used for differential analysis should be discussed and justified. Notably, age is used as a covariate for transcriptomic analysis but not proteomic and metabolomic analysis, with no justification. Additionally, given the known importance of lipid metabolism in AD and the putative role of APOE in lipid homeostasis3, APOE genetic status should be considered as a covariate, or its omission should be justified.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, such as age at death and RIN, is that these data were not available for each sample. Thus, we referred to the original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, and years of education. Regarding the proteomic dataset, in the original article (Johnson et al., 2020), age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (3) The authors make a conclusion statement that suggests intervention: "Collectively, our data suggests that preserving or improving the ability to produce ATP and early intervention in the process of nitrogen metabolism are candidates for the prevention and treatment of dementia" (pg. 12, lines 12-14). This claim is not well-supported by the evidence provided in the study. There are a few limitations: (a) This was an observational, not interventional study; (b) The study did not establish whether the metabolic disruptions are causes or effects in AD; and (c) ATP or other bioenergetic indicators were not directly measured. Therefore, any statements about potential interventions should be removed or qualified as highly speculative.

      We agree with the reviewer that the statement regarding potential interventions was not sufficiently supported by our analyses. Accordingly, we have removed the sentence regarding prevention and treatment from the revised manuscript (e.g., we have deleted final paragraph of the previous manuscript).

      (4) In conjunction with the last point, the main conclusion of the study is that energy production is down in AD. The data presented in Figure 3 are consistent with this conclusion, but it is far from definitive due to limitations stated above in comments 3a and 3b. The authors should offer additional support for this conclusion: experimental follow-up, flux modeling, analysis of alternative datasets with ATP measurement, causal inference.

      We sincerely thank the reviewer for this valuable and constructive suggestion. Regarding flux modeling, we agree that metabolic flux analysis could provide important mechanistic insight. Indeed, previous studies have applied flux modeling in the context of lipid metabolism in Alzheimer’s disease (Baloni et al., 2022). We also attempted to perform flux modeling focusing on energy metabolism. However, we found it difficult to obtain biologically meaningful and robust results and therefore decided not to include these analyses in the current manuscript.

      With respect to ATP measurements, we fully agree that direct evidence of altered ATP levels would further strengthen our conclusion. However, to the best of our knowledge, there are currently no publicly available large-scale datasets that directly measure ATP levels in human postmortem brain tissues. This limitation makes it challenging to incorporate validation in the present study.

      Regarding experimental follow-up, we agree that functional validation is essential to confirm the mechanistic implications of our findings. We are actively considering follow-up experimental studies. However, we consider the present work to be a multi-omic integrative analysis aimed at identifying key molecular alterations and generating biologically important hypotheses. We have revised the Limitation section to more clearly position this manuscript as an observational systems-level analysis (pg. 13, lines 20-22).

      (5) The validation analysis did not sufficiently show the generalizability of this study's results. The authors demonstrated a correlation of 0.53 to the MSBB transcriptomics data and 0.60 to the AMP-AD DiverseCohorts proteomics data. Beyond these correlation coefficients, no meaningful comparison between the datasets is offered. How concordant are the differentially expressed features (or pathways) between the datasets? How robust would the trans-omic network be if incorporating the alternate datasets? Is the main conclusion (energy metabolism is down in AD) supported by the validation datasets? We think this analysis should be expanded and described in the main text.

      Although the results for external metabolomics datasets are reported in Fig S2C, correlation coefficients with the external data are not reported. The authors state, "Note that each study used different definitions for AD and CT groups, had variations in measurement methods and brain regions analyzed." We appreciate these limitations. However, the external data should be re-analyzed using the same definitions of AD and CT, if possible. The limitations and results (which DEMs are shared between datasets) should be discussed in the main text.

      We thank the reviewer for this important comment regarding the generalizability of our findings. In the revised manuscript, we have expanded the validation analyses and summarized the results in Figure S2. First, at the transcriptomic level, Figure S2B and S2C show the overlap between up- and downregulated genes in AD identified in our ROSMAP-derived analyses and those reported in a previously published large-scale meta-analysis of 2,114 postmortem samples across seven brain regions (Wan et al., 2020). A substantial proportion of DEGs were shared, supporting cross-cohort and cross-region robustness to some extent. At the proteomic level, Figure S2E shows a comparison between the ROSMAP and the AMP-AD DiverseCohorts datasets. We highlighted the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3 and calculated a separate correlation coefficient for this subset (Pearson coefficient = 0.86, p-value = 1.5e-7), further supporting our main conclusion. In addition, to assess the concordance between the two datasets in a threshold-independent manner, we additionally performed Rank-Rank Hypergeometric Overlap (RRHO) analysis (Figure S2E). RRHO analysis (Cahill et al., 2018; Plaisier et al., 2010) enables the comparison of ranked protein lists without relying on arbitrary differential expression cutoffs and has been used for cross-dataset comparison in several previous studies (Fröhlich et al., 2024; Maitra et al., 2023). The RRHO heatmaps demonstrated significant enrichment in the concordant quadrants, confirming systematic agreement between datasets beyond simple correlation coefficients. For metabolomics, Figure S2G shows RRHO analyses comparing the ROSMAP metabolomic data with other datasets measured by the same UPLC-MS/MS platform (Batra et al., 2024; Novotny et al., 2023), demonstrating significant concordance in ranked metabolite changes in AD.

      (6) The glycolysis analysis and discussion needs more development. Glycolysis and gluconeogenesis share many of the same enzymes, but they are not the same pathway and should not be discussed as such. To make a claim about the overall influence of enzyme and metabolite levels on glycolysis, the authors should focus on the energetically committing steps of glycolysis (hexokinase, phosphofructokinase, pyruvate kinase) in Figure 3A, and include the full/current version of the figure in the supplement. Gluconeogenesis-specific enzymes (pyruvate carboxylase, PEPCK) are not mentioned at all - are they among the DEPs/DEGs?

      We appreciate the reviewer’s comment regarding the distinction between glycolysis and gluconeogenesis pathway. Among the gluconeogenesis-specific enzyme proteins, G6PC1, FBP1, PC, and PCK2 were measured in our dataset, but none of them were identified as DEPs. In addition, gluconeogenesis is a process that occurs primarily in the liver and kidney rather than the brain. Given this biological context and the lack of significant changes in relevant enzymes, we have revised the terminology throughout the manuscript, replacing “glycolysis/gluconeogenesis pathway” with “glycolysis pathway” in the revised version.

      (7) Given that there wasn't good concordance between the DEGs and DEPs, did including the mRNA and transcription factor layers in the network really add anything useful? It seems like the main conclusions of the manuscript were driven by the protein and metabolite layers only. How many of the DE metabolic enzymes were coregulated at the transcript and protein level? It would be useful to include the 5-layer trans-omic network in the supplement to display these results. Given your network, at what level does it appear that energy metabolism is regulated?

      It is true that our primary conclusion regarding the regulation of energy metabolism is driven by the changes in protein and metabolite abundance. However, we consider the low concordance between mRNA and protein expression itself to be an important feature of AD pathology, as also reported in previous studies (Johnson et al., 2022; Tasaki et al., 2022). Although we did not perform a further analysis of this discordance, we believe that including the TF and mRNA layers into the metabolic trans-omic network strengthens a system-wide view of metabolic dysregulation in AD.

      Regarding the mRNA changes corresponding to the DEP enzymes, please refer to Figure S7A.

      (8) Comment further on the results from Figure 2D. What can be learned from identifying metabolites with the greatest degree centrality? What pathways other than energy metabolism are highlighted by the trans-omic network?

      We assume that some energetic indicators, including AMP and acetyl-CoA, and nitrogen metabolism-related metabolites, Glu, 2-oxoglutarate, and urea, can be potential key regulators of dysregulated metabolism in AD.

      (9) (Suggestion) We suggest the authors leverage their trans-omic network in additional ways beyond giving a snapshot of a few energy metabolism pathways. The analysis of top DEMs could go further. What pathways are impacted beyond energy metabolism? Among the metabolic reactions allosterically regulated by top DEMs, what metabolic pathways are enriched?

      We identified the enriched metabolic pathways that were allosterically regulated by DEMs in AD using Fisher’s exact test. Alanine, aspartate, and glutamate metabolism pathways were significantly enriched in 2-oxoglutarate, glutarate, alanine, and glutamate-regulating metabolic reactions. Arginine and proline metabolism pathway was enriched in N-methyl-L-arginine and putrescine-regulating metabolic reactions. Arginine biosynthesis pathway was enriched in arginine-regulating metabolic reactions. Glycerophospholipid metabolism pathway was enriched in CDP-ethanolamine-regulating metabolic reactions. Glycine, serine, and threonine metabolism pathway was enriched in serine-regulating metabolic reactions. Purine metabolism pathway was enriched in AMP-regulating metabolic reactions. Pyrimidine metabolism pathway was enriched in deoxyuridine and thymidine-regulating metabolic reactions. Sphingolipid metabolism pathway was enriched in sphingosine-regulating metabolic reactions. However, this analysis did not yield sufficiently valuable insights into the regulatory relationships among biomolecules in AD. Thus, we did not include these results in the revised manuscript.

      (10) (Suggestion) Figure 3 shows that most differential signal in AD points to lower energy production due to the combination of differentially expressed metabolites and enzymes, but we are not given much context about the strength of these among all the differential signals. We would suggest including volcano plots where the features of interest, i.e. DE enzymes and metabolites, are colored differently (or a similar figure).

      We thank the reviewer for this constructive suggestion. To provide better context regarding the importance of the differential signals, we have added volcano plots for mRNAs, proteins, and metabolites in Figure S4A, B, and C.

      (11) (Suggestion) The PPI network could be better leveraged to understand metabolic changes in AD. If nodes are grouped into subnetworks (e.g. by Louvain / Leiden clustering) and tested for pathway enrichment, could you find functional subnetworks of coordinately up- and down- regulated metabolic enzymes? This could yield some pathways of interest beyond the energy metabolism pathways already highlighted.

      We appreciate the reviewer’s suggestion to utilize the PPI network for subnetwork analysis. However, it is important to note that the proteomic dataset analyzed in this study is derived from the original work of (Johnson et al., 2020). In that paper, the authors already performed a Weighted Gene Co-expression Network Analysis (WGCNA) across several datasets to identify co-expressed modules and functional pathways.

      Given this, we assumed that applying additional clustering methods to the same dataset would be unlikely to yield significant biological insights beyond the established findings.

      Minor comments

      (1). "All genes" and "all metabolites" should not be the background for the proteomic and metabolic pathway enrichment analysis by Metascape and MetaboAnalyst. The background should be limited to the proteins and metabolites that were measured.

      We fully agree with the reviewer that using “all gene” or “all metabolites” as a background is not suitable for enrichment analyses. As suggested, we have revised the enrichment analyses using the measured proteins and metabolites as a background in both Metascape and MetaboAnalyst (Fig. S4D).

      (2) Highlight the metabolic enzymes in Fig S2B. Calculate a separate correlation coefficient for the enzymes extracted in the energy metabolism analysis from Fig 3.

      We appreciate the reviewer’s suggestion to refine the correlation analysis. As requested, we have revised Fig. S2D to explicitly highlight the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3. We calculated a separate correlation coefficient for the subset (Pearson coefficient = 0.86, p-value = 1.5e-7).

      (3) Use a multiple hypothesis adjusted p-value or q-value in Figure S3.

      We agree with the reviewer regarding the necessity of correcting for multiple comparisons. Accordingly, we have revised Fig. S4D using q-values.

      (4) Describe the methods used to calculate the logFC values from the validation dataset.

      We have revised the Methods to include a detailed description of the procedure used to calculate the log2FC values for the validation datasets (pg. 21, lines 13-15).

      (5) It is difficult to read Figure 3. We would recommend really emphasizing to the reader to refer to Fig S7B as a "key" to this figure. The description of the red/blue arrows and nodes in the methods section (pg. 24, lines 21-36, pg 25, lines 1-4) were also helpful, but very lengthy. We recommend putting an abridged version of this description into the Fig S7 figure legend.

      We appreciate the feedback regarding the readability of Fig. 3. As recommended, we have revised the manuscript to explicitly direct readers to Fig. S8B as an essential “key” for interpreting the network visualization (pg. 8, lines 28). Furthermore, we have added an abridged description of the network elements to the legend of Fig. S8B.

      (6) The S7 figure legend should refer to panels A and B, not E and F.

      We apologize for this oversight. We have corrected the legend of Fig. S8.

      (7) (Suggestion) Are any of the differentially expressed metabolites allosteric regulators of the DE transcription factors? This could be interesting to discuss.

      We appreciate the reviewer’s insightful suggestion about the potential allosteric regulation of the DETFs by DEMs. We conducted an extensive literature search to identify any reports related to this perspective. However, to the best of our knowledge, no such direct interactions have been reported to date.

      Reviewer #1 (Significance):

      The study's strength lies in leveraging three omics modalities across large patient cohorts (n ~ 150-240) to identify coherent signals between transcriptomics, proteomics, and metabolomics in postmortem DLPFC tissue. It was encouraging to see that the main result, showing downregulation for TCA, oxidative phosphorylation, and ketone body metabolism, emerged from consistent signals across both proteomics and metabolomics. This result was consistent with previous findings in other models cited by the author4,5 and other studies 6,7 demonstrating deficiency in energy-producing pathways in AD.

      Another strength of the study is the application of thoughtful methodology to connect differentially expressed proteins and metabolites via an intermediate data layer of metabolic reactions. The authors leverage the KEGG and BRENDA databases and apply sound logic to estimate the effects of enzyme level and metabolite level on pathway activity, with metabolites serving as substrate, product, or allosteric regulator for reactions. This trans-omic network methodology was developed in previous studies cited by the author8,9.

      However, as written, this study is limited in its contribution of new knowledge to the AD research field. The main conclusion (energy production is down in AD, due to regulatory disruption of energy metabolism) is not strongly supported (see comments 1, 3, and 4 for elaboration). The evidence could be improved by orthogonal approaches: further experimentation, further integration of external datasets, causal modeling, or flux modeling. Alternatively, even in the absence of new experimental and computational approaches, the story could be made more complete by further leveraging the trans-omic network to provide insights into (a) the regulation of energy metabolism; and (b) the impacts of key disrupted metabolites (see comments 7-9).

      The study is also limited in its demonstrating the power of these methodologies to provide integrative insights. As mentioned above, the integration of enzyme levels and metabolite levels is clearly useful (Figure 3). In contrast, the utility of the mRNA and transcription factor layers was not evident. The study did not appear to improve or expand upon trans-omic network methodology described in the previous works. Finally, the various analyses (analyzing the trans-omic network for nodes with the highest degree centrality, the PPI analysis, and viewing the energy metabolism pathways in the network) provided disparate results that were only tenuously connected in the discussion section.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary

      This manuscript integrates public transcriptomic, proteomic, and metabolomic datasets from ROSMAP DLPFC samples to construct a multi-layer metabolic trans-omic network in Alzheimer's disease. By linking transcription factors, enzyme mRNAs, proteins, metabolic reactions, and metabolites, the authors report coordinated downregulation of the TCA cycle, oxidative phosphorylation, and ketone body metabolism, along with mixed regulatory signals in glycolysis/gluconeogenesis. They interpret these patterns as indicative of broad energetic dysfunction and alterations in amino-acid/nitrogen metabolism in AD. While the framework is conceptually appealing, much of the analysis remains descriptive, and several biological interpretations extend beyond what the data can robustly support. The reliance on bulk tissue without accounting for cell-type composition, limited covariate adjustment, and the absence of validation or sensitivity analyses reduce confidence in the mechanistic conclusions. Overall, the study provides a preliminary systems-level overview, but additional rigor is needed before the proposed trans-omic regulatory insights can be considered convincing.

      Major Comments

      (1) Interpretation requires more cautious phrasing, and validation is essential. The manuscript frequently asserts that specific pathways are "inhibited" or that energetic deficits are "compensated," but these conclusions extend beyond what the descriptive, bulk-level data can support. Because no metabolic flux, causality, or direct functional measurements are included, the results should be framed as putative regulatory shifts, not confirmed impairments. Critically, key claims about pathway inhibition would require flux modeling, perturbation analyses, or experimental validation to be convincing. Without such validation, the mechanistic interpretations remain speculative.

      We thank the reviewer for this crucial comment. We fully agree that, given the descriptive and bulk-level nature of our analysis, mechanistic interpretations must be made with caution. In the absence of direct metabolic flux measurements or experimental validation, our findings should be interpreted as putative regulatory shifts rather than confirmed functional impairments. Accordingly, we have revised the manuscript to temper mechanistic claims. We have replaced definitive statements with more speculative phrasing (e.g., “Our analysis revealed a putative coordinated downregulation …” instead of “Our analysis revealed a coordinated downregulation …” in Abstract section; “we demonstrate the systems-level view of the potential dysregulated energy production …” instead of “we demonstrate the systems-level view of the dysregulated energy production …” in pg. 10, lines 25-26).

      (2) Although the authors acknowledge this in the limitations, bulk-level differences may primarily reflect altered proportions of neurons, astrocytes, microglia, and oligodendrocytes rather than true within-cell-type regulation. Incorporating a cell-type deconvolution or performing a sensitivity analysis would substantially improve interpretability. This issue also impacts the trans-omic network: if the molecules included originate from different cell types, the inferred regulatory relationships may not reflect true intracellular processes.

      We appreciate the reviewer’s point that bulk-level differences can reflect altered proportions of different brain cell types, subsequently affecting the inferred trans-omic network analysis. To assess the changes in cell type proportions of the samples that we used in our study, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglias, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two group. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      (3) Differential analysis covariates. For the differential expression analyses, only gender and PMI were included as covariates. Additional variables, such as age at death, RIN, neuropathological measures, and comorbidities, can strongly influence molecular profiles and should be considered to ensure that the observed differences reflect AD-related biology rather than confounding pathological or technical factors.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, including age at death and RIN, is that these data for each sample were not available. Thus, we referred to original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, or education. Regarding the proteomic dataset, in the original article, age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (4) Network stability and sample non-overlap. Proteomic, transcriptomic, and metabolomic data come from partially overlapping individuals. The authors should test whether the reconstructed network is robust to: different significance thresholds, restricting analyses to overlapping samples and alternative definitions of AD vs control.

      We appreciate the reviewer’s comment for the trans-omic network stability. In our study, the number of individuals for whom all omic modalities were measured was relatively small (n=25 in CT and n=35 in AD). This limited overlap reduces statistical power and can affect the downstream network construction. We have acknowledged this limitation in the revised manuscript and clarified that the reconstructed networks should be interpreted with caution regarding reproducibility and generalizability (pg. 13, lines 13-23).

      Minor Comments

      (1) Some TF enrichment and regulatory inferences lack explicit mention of multiple-testing correction.

      We apologize for the lack of clarity in our original description. We have corrected for multiple-testing for the TF inference. Thus, we have revised the Methods section to explicitly describe the correction method used and the threshold applied (pg. 23, lines 23-24).

      (2) The limitations section is strong but should explicitly discuss the influence of postmortem interval on metabolite levels.

      We appreciate the reviewer’s comment about the effect of postmortem interval on changes in metabolite levels. Accordingly, we have added the description of this perspective in our revised manuscript (pg. 13, lines 1-5).

      Reviewer #2 (Significance):

      The study extends a trans-omic integration framework, originally applied to metabolic disease, into the context of Alzheimer's pathology. Although the biological findings largely confirm known alterations in mitochondrial and energy metabolism, the network-based approach offers a structured way to view cross-layer regulatory changes. Its main advance is conceptual rather than biological, providing a unified framework rather than uncovering fundamentally new mechanisms. This work will primarily interest researchers in neurodegeneration and systems biology, as well as computational groups developing multi-omics integration methods.

      Reviewer #3 (Evidence, reproducibility and clarity):

      This study leverages existing transcriptomic, metabalomic and proteomic datasets from prefrontal cortex (PFC) to assess metabolic dysregulation in Alzheimer's disease (AD). They found a downregulation of multiple metabolic pathways, including TCA cycle, oxidative phosphorylation, and ketone metabolism, that may explain bioenergetic alterations in AD.

      The study used matching ROSMAP omics datasets from the DLPFC that have allowed more robust data integration. However, the datasets are all generated using bulk tissue, which makes data interpretation difficult. For example, the AD changes they observed may be due to shifts in cell type proportion with disease (e.g. cell death, neuron inflammation). Did the authors account for any potential shifts in cell type proportion in their analysis?

      If the assumption is that the changes in AD are cell intrinsic, which cell types are likely to be impacted? Can the authors integrate any existing single-cell analysis to infer which cell types may be driving the signals they detect, and whether this accounts for some of the antagonistic regulatory effects that were detected?

      We thank the reviewer for their insightful comments. We agree that the use of bulk tissue datasets cannot account for cell-type heterogeneity. As noted in our Limitations section (pg. 12, lines 24-27), we recognize that previous studies have found that the Braak stage is correlated positively with microglia and astrocyte proportions and negatively with oligodendrocyte proportion (Hannon et al., 2024; Shireby et al., 2022). Regarding the integration of single-cell analysis, we have referenced recent snRNA-seq findings (Mathys et al., 2024) in our Limitations section (pg. 12, lines 28-32) to deconvolve our bulk signatures.

      Furthermore, in our revised manuscript, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglia, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two groups. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      Reviewer #3 (Significance):

      The manuscript provides multimodal insight into metabolic dysregulation in AD in the PFC. Given that metabolic dysfunction is likely to play a major in disease pathogenesis, this is a study of importance. However, the findings lack granularity at the cell type level, which limits the impact of the study.

      Reference

      (1) Baloni, P., Arnold, M., Buitrago, L., Nho, K., Moreno, H., Huynh, K., Brauner, B., Louie, G., Kueider-Paisley, A., Suhre, K., Saykin, A. J., Ekroos, K., Meikle, P. J., Hood, L., Price, N. D., Alzheimer’s Disease Metabolomics Consortium, Doraiswamy, P. M., Funk, C. C., Hernández, A. I., … Kaddurah-Daouk, R. (2022). Multi-Omic analyses characterize the ceramide/sphingomyelin pathway as a therapeutic target in Alzheimer’s disease. Communications Biology, 5(1), 1074.

      (2) Baloni, P., Funk, C. C., Yan, J., Yurkovich, J. T., Kueider-Paisley, A., Nho, K., Heinken, A., Jia, W., Mahmoudiandehkordi, S., Louie, G., Saykin, A. J., Arnold, M., Kastenmüller, G., Griffiths, W. J., Thiele, I., Alzheimer’s Disease Metabolomics Consortium, Kaddurah-Daouk, R., & Price, N. D. (2020). Metabolic Network Analysis Reveals Altered Bile Acid Synthesis and Metabolism in Alzheimer’s Disease. Cell Reports. Medicine, 1(8), 100138.

      (3) Batra, R., Arnold, M., Wörheide, M. A., Allen, M., Wang, X., Blach, C., Levey, A. I., Seyfried, N. T., Ertekin-Taner, N., Bennett, D. A., Kastenmüller, G., Kaddurah-Daouk, R. F., Krumsiek, J., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2023). The landscape of metabolic brain alterations in Alzheimer’s disease. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(3), 980–998.

      (4) Batra, R., Krumsiek, J., Wang, X., Allen, M., Blach, C., Kastenmüller, G., Arnold, M., Ertekin-Taner, N., Kaddurah-Daouk, R., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2024). Comparative brain metabolomics reveals shared and distinct metabolic alterations in Alzheimer’s disease and progressive supranuclear palsy. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 20(12), 8294–8307.

      (5) Cahill, K. M., Huo, Z., Tseng, G. C., Logan, R. W., & Seney, M. L. (2018). Improved identification of concordant and discordant gene expression signatures using an updated rank-rank hypergeometric overlap approach. Scientific Reports, 8(1), 9588.

      (6) Fröhlich, A. S., Gerstner, N., Gagliardi, M., Ködel, M., Yusupov, N., Matosin, N., Czamara, D., Sauer, S., Roeh, S., Murek, V., Chatzinakos, C., Daskalakis, N. P., Knauer-Arloth, J., Ziller, M. J., & Binder, E. B. (2024). Single-nucleus transcriptomic profiling of human orbitofrontal cortex reveals convergent effects of aging and psychiatric disease. Nature Neuroscience, 27(10), 2021–2032.

      (7) Green, G. S., Fujita, M., Yang, H.-S., Taga, M., Cain, A., McCabe, C., Comandante-Lou, N., White, C. C., Schmidtner, A. K., Zeng, L., Sigalov, A., Wang, Y., Regev, A., Klein, H.-U., Menon, V., Bennett, D. A., Habib, N., & De Jager, P. L. (2024). Cellular communities reveal trajectories of brain ageing and Alzheimer’s disease. Nature, 633(8030), 634–645.

      (8) Hannon, E., Dempster, E. L., Davies, J. P., Chioza, B., Blake, G. E. T., Burrage, J., Policicchio, S., Franklin, A., Walker, E. M., Bamford, R. A., Schalkwyk, L. C., & Mill, J. (2024). Quantifying the proportion of different cell types in the human cortex using DNA methylation profiles. BMC Biology, 22(1), 17.

      (9) Johnson, E. C. B., Carter, E. K., Dammer, E. B., Duong, D. M., Gerasimov, E. S., Liu, Y., Liu, J., Betarbet, R., Ping, L., Yin, L., Serrano, G. E., Beach, T. G., Peng, J., De Jager, P. L., Haroutunian, V., Zhang, B., Gaiteri, C., Bennett, D. A., Gearing, M., … Seyfried, N. T. (2022). Large-scale deep multi-layer analysis of Alzheimer’s disease brain reveals strong proteomic disease-related changes not observed at the RNA level. Nature Neuroscience, 25(2), 213–225.

      (10) Johnson, E. C. B., Dammer, E. B., Duong, D. M., Ping, L., Zhou, M., Yin, L., Higginbotham, L. A., Guajardo, A., White, B., Troncoso, J. C., Thambisetty, M., Montine, T. J., Lee, E. B., Trojanowski, J. Q., Beach, T. G., Reiman, E. M., Haroutunian, V., Wang, M., Schadt, E., … Seyfried, N. T. (2020). Large-scale proteomic analysis of Alzheimer’s disease brain and cerebrospinal fluid reveals early changes in energy metabolism associated with microglia and astrocyte activation. Nature Medicine, 26(5), 769–780.

      (11) Maitra, M., Mitsuhashi, H., Rahimian, R., Chawla, A., Yang, J., Fiori, L. M., Davoli, M. A., Perlman, K., Aouabed, Z., Mash, D. C., Suderman, M., Mechawar, N., Turecki, G., & Nagy, C. (2023). Cell type specific transcriptomic differences in depression show similar patterns between males and females but implicate distinct cell types and genes. Nature Communications, 14(1), 2912.

      (12) Mathys, H., Boix, C. A., Akay, L. A., Xia, Z., Davila-Velderrain, J., Ng, A. P., Jiang, X., Abdelhady, G., Galani, K., Mantero, J., Band, N., James, B. T., Babu, S., Galiana-Melendez, F., Louderback, K., Prokopenko, D., Tanzi, R. E., Bennett, D. A., Tsai, L.-H., & Kellis, M. (2024). Single-cell multiregion dissection of Alzheimer’s disease. Nature, 632(8026), 858–868.

      (13) Novotny, B. C., Fernandez, M. V., Wang, C., Budde, J. P., Bergmann, K., Eteleeb, A. M., Bradley, J., Webster, C., Ebl, C., Norton, J., Gentsch, J., Dube, U., Wang, F., Morris, J. C., Bateman, R. J., Perrin, R. J., McDade, E., Xiong, C., Chhatwal, J., … Harari, O. (2023). Metabolomic and lipidomic signatures in autosomal dominant and late-onset Alzheimer’s disease brains. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(5), 1785–1799.

      (14) Plaisier, S. B., Taschereau, R., Wong, J. A., & Graeber, T. G. (2010). Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures. Nucleic Acids Research, 38(17), e169.

      (15) Shireby, G., Dempster, E. L., Policicchio, S., Smith, R. G., Pishva, E., Chioza, B., Davies, J. P., Burrage, J., Lunnon, K., Seiler Vellame, D., Love, S., Thomas, A., Brookes, K., Morgan, K., Francis, P., Hannon, E., & Mill, J. (2022). DNA methylation signatures of Alzheimer’s disease neuropathology in the cortex are primarily driven by variation in non-neuronal cell-types. Nature Communications, 13(1), 5620.

      (16) Tasaki, S., Xu, J., Avey, D. R., Johnson, L., Petyuk, V. A., Dawe, R. J., Bennett, D. A., Wang, Y., & Gaiteri, C. (2022). Inferring protein expression changes from mRNA in Alzheimer’s dementia using deep neural networks. Nature Communications, 13(1), 655.

      (17) Varma, V. R., Wang, Y., An, Y., Varma, S., Bilgel, M., Doshi, J., Legido-Quigley, C., Delgado, J. C., Oommen, A. M., Roberts, J. A., Wong, D. F., Davatzikos, C., Resnick, S. M., Troncoso, J. C., Pletnikova, O., O’Brien, R., Hak, E., Baak, B. N., Pfeiffer, R., … Thambisetty, M. (2021). Bile acid synthesis, modulation, and dementia: A metabolomic, transcriptomic, and pharmacoepidemiologic study. PLoS Medicine, 18(5), e1003615.

      (18) Wan, Y.-W., Al-Ouran, R., Mangleburg, C. G., Perumal, T. M., Lee, T. V., Allison, K., Swarup, V., Funk, C. C., Gaiteri, C., Allen, M., Wang, M., Neuner, S. M., Kaczorowski, C. C., Philip, V. M., Howell, G. R., Martini-Stoica, H., Zheng, H., Mei, H., Zhong, X., … Logsdon, B. A. (2020). Meta-Analysis of the Alzheimer’s Disease Human Brain Transcriptome and Functional Dissection in Mouse Models. Cell Reports, 32(2), 107908.

    1. eLife Assessment

      The authors propose a pipeline using large language models (LLMs) to benchmark experimentally well validated computational models of signaling networks. The findings are important, with available methods that guide testing future models and find new molecular interactions. Furthermore, the work shows how general-purpose LLMs can generate up to 91% of reactions of a bacterial metabolic network. The support is convincing and offers a number of performance metrics.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors did an excellent job with their resubmission, politely and elegantly answering the comments from the reviewers.]

      Summary:

      Large language models (LLMs) have been developed rapidly in recent years and are already contributing to progress across scientific fields. The manuscript tries to address a specific question: whether LLMs can accurately infer signaling networks from gene lists.

      Strengths:

      The manuscript raises a good question: whether current LLMs can accurately generate signaling networks from gene lists.

    3. Reviewer #2 (Public review):

      Summary:

      The authors evaluate whether commonly used LLMs (ChatGPT, Claude and Gemini) can reconstruct signalling networks and predict effects of network perturbations, and propose a pipeline for benchmarking future models. Across three phenotypes (hypertrophy, fibroblast signalling, and mechanosignalling), LLMs capture upstream ligand-receptor interactions and conserved crosstalk but fail to recover downstream transcriptional programmes. Logic-based simulations show that LLM-derived networks underperform compared to manually curated models. The authors also propose that their pipeline can be used for benchmarking future models aimed at reconstructing signalling networks.

      Strengths:

      The authors compare the outcomes from three LLMs with three manually curated and validated models. Additionally, they have investigated gene network reconstruction in the context of three distinct phenotypes. Using logic-based modelling, the authors assessed how LLM-derived networks predict perturbation effects, providing functional validation beyond network overlap.

      Weaknesses:

      The authors have used legacy models for all three LLMs, and the study would benefit from testing the current versions of the LLMs (ChatGPT 5.2, Claude 4.5 and Gemini 2.5). Additional metrics such as node coverage, node invention, direction accuracy and sign accuracy would be useful to make robust comparisons across models.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary

      Large language models (LLMs) have been developed rapidly in recent years and are already contributing to progress across scientific fields. The manuscript tries to address a specific question: whether LLMs can accurately infer signaling networks from gene lists. However, the evaluation is inadequate due to four major weaknesses described below. Despite these limitations, the authors conclude that current general-purpose LLMs lack adequate accuracy, which is already widely recognized. Its key contribution should instead be to provide concrete recommendations for the development of specialized LLMs for this task, which is completely absent. Developing such specific LLMs would be highly valuable, as they could substantially reduce the time required by researchers to analyze signaling networks.

      Strengths

      The manuscript raises a good question: whether current LLMs can accurately generate signaling networks from gene lists.

      Weaknesses:

      (1) The authors evaluate LLM performance using only three signaling networks: "hypertrophy", "fibroblast", and "mechanosignaling". Given the large number of well-established signaling pathways available, this is not a comprehensive assessment. Moreover, the analysis need not be restricted to signaling networks. Other network types, including metabolic and transcriptional regulatory networks, are already accessible in well-known databases such as KEGG, Reactome, BioCyc, WikiPathways, and Pathway Commons. Including these additional networks would substantially strengthen the evaluation.

      We agree with the reviewer that our evaluation of LLM performance is not comprehensive of all signaling networks, and that the benchmarking was previously limited to signaling networks. The purpose of this study is to benchmark LLMs against peer-reviewed computational models that make testable predictions and are highly validated experimentally, of which these three signaling networks are strong examples. KEGG, Reactome, Biocyc, WikiPathways, and Pathway Commons are databases that contain collections of individual reactions or pathways but are not computational models themselves and have not been experimentally validated in that sense.

      While this study focuses primarily on signaling networks, we agree that it would be useful to evaluate how well LLMs perform in generating networks of another type, for which predictive and validated computational models are available. Therefore, in new Figure 3 we now test the ability of LLMs to reconstruct the E. Coli core metabolic network, as well as test its ability to predict growth on metabolic substrates using flux balance analysis. We find that Claude Opus 4.6, GPT 5.2 Pro, and Gemini 3 Pro Preview perform well at reconstructing reactions from E. Coli core metabolism, but these reactions are not sufficient to predict growth on a variety of substrates. Given this expansion of scope, we replaced “signaling networks” in the title with “biochemical networks”.

      (2) In LLM evaluation, the authors use the gene lists that exactly match those in their "ground truth" networks, thereby fixing the set of nodes and evaluating only the predicted edges. However, in practical research, the relevant genes or nodes are not fully known. A more realistic assessment would therefore include gene lists with both genes present in the ground-truth network and additional genes absent from it, to evaluate the ability of the LLM to exclude irrelevant genes.

      We agree with the reviewer that evaluating the capacity of these LLMs to exclude additional genes is interesting. But because biological networks are always incompletely known, there is no “ground truth” of genes absent from a given network. Therefore, for the most rigorous benchmarking against a “ground truth”, we examine the positive predictive ability of LLMs. However, in response to this comment and point 3 below, we further examine additional measures of performance that include “false positives”.

      (3) The authors report only the recall/sensitivity of the LLM, without assessing specificity. In practical applications, if an LLM generates a large number of incorrect interactions that greatly exceed the correct ones, researchers may be misled or may lose confidence in the LLM output. Therefore, a comprehensive evaluation must include both sensitivity and specificity. Furthermore, it would be informative to check whether some of the "false positives" might in fact represent biologically plausible interactions that are absent from the manually curated "ground truth". Manually generated "ground truth" can overlook genuine interactions, and the ability of LLMs to recover such missing edges could be particularly valuable. This may even represent one of the most important potential contributions of LLMs.

      We agree with the reviewer that additional metrics could inform the evaluation of LLM performance. Therefore, as recommended, we calculated sensitivity, specificity, precision, negative predictive value (NPV), accuracy, and F1 score for each of the network models (hypertrophy, fibroblast, and mechanosignaling). These new results are summarized in confusion matrices shown in a new Supplementary Figures 3, 4, and 5. We performed this additional benchmarking across the 10 replicates for each LLM.

      One limitation of this approach is the substantial class imbalance within these confusion matrices. Because we interpret actual negatives as connections that are not found between any nodes of the ground truth models, there will be >10x more true negatives than any of the other classes. This makes specificity, accuracy, and the negative predictive value less informative.

      To illustrate this point, consider the ground-truth hypertrophy network which contains 191 connections between 106 nodes. The total number of possible connections between any two nodes is 106<sup>2</sup> = 11,236. Given that there are 191 actual positives, that leaves 11,045, actual negatives (as illustrated in the null predictor of Supplementary Figure 3A). As the number of node-to-node connections predicted by LLMs is on the order of a few hundred, the number of true negatives is always in the thousands, often outweighing the true positives in the specificity or accuracy calculations.

      The effects of this class imbalance are illustrated with the results of the “null predictors” in Supplementary Figure 4A, which are hypothetical models that fail to predict any connection between nodes (have predicted positive values of 0). These null predictors have sensitivities of 0 and specificities of 1 and high accuracies and NPVs because of the high true negative rates.

      The precision and F1 scores calculated using these confusion matrices are robust to these class imbalances. Indeed, there is substantial heterogeneity in the number of “false positives” connections generated by the LLMs as illustrated by the precision and F1 scores. We are hesitant to unequivocally label these novel, predicted connections as false positives because, as the reviewer points out, these connections could represent true molecular interactions that were undiscovered at the time of the ground truth models’ conception but have since been demonstrated experimentally and published.

      (4) It is widely known that applying differential equation models to highly complex biological networks, such as the three networks in the manuscript, is meaningless, because these systems involve a large number of parameters whose values can drastically alter the results. As Richard Feynman once said: "with four parameters I can fit an elephant, and with five I can make him wiggle his trunk." Thus, the evaluation of LLMs on "logic-based differential equation models" does not make much sense.

      Differential equation models have been the primary mathematical framework for studying complex systems for decades. We refer readers unfamiliar with differential equations to the Nobel prize-winning work of Hodgkin/Huxley (Physiology or Medicine1963, action potential of neurons), Prigogine (Chemistry 1977, non-equilibrium thermodynamics and pattern formation), John Nash (Economics 1994, game theory and Nash Equilibrium), Merton/Scholes (Economics 1997, dynamics of financial derivatives), and Manabe/Hasselmann (Physics 2021, dynamic modeling of atmosphere and oceans).

      We are confused by the quote of a joke by Richard Feynman about fitting equations to data in the shape of an elephant. While this famous joke is amusing, it is both misattributed (it was a recollection by Enrico Fermi in 1953 of a joke once made by John von Neumann) and deliberately hyperbolic (see https://en.wikipedia.org/wiki/Von_Neumann%27s_elephant). Regardless, the relevance of this joke to our study is unclear, because we are not fitting equations to data. As described in the text, in previous studies we validated the predictions of these three logic-based network models with experimental data that was not used to develop the models.

      Reviewer #1 (Recommendations for the authors):

      (1) All figures are in very poor resolution.

      Thank you for identifying this. We have fixed this issue, which was due to SVG embedding. We now embed as higher resolution PNG and provide full resolution files separately.

      (2) The manuscript does not include data availability or code availability.

      As described in the Methods, all code and data is now available via GitHub (https://github.com/saucermanlab/LLM-network-generation).

      Reviewer #2 (Public review):

      (1) Information on the accuracy of directionality of interaction would help understand if there is a bias towards either a positive or negative association.

      To assess if LLMs are biased in predicting either stimulatory or inhibitory connections, we examined the proportion of stimulatory and inhibitory connections for each ground truth model along with the prediction sets from the different LLMs (Author response table 1). These findings suggest that any directional bias is minimal and not conserved across the different ground truth models. These tables were not included in the revised manuscript.

      Author response table 1.

      Proportion of stimulatory and inhibitory connections present in each ground truth model and in the sets of connections predicted by each LLM (Claude, GPT, and Gemini).

      The primary subset of reaction types that the LLM’s tend to miss are often downstream, cell-type specific, and/or gene regulatory connection (e.g. Figure 1B). This observation is consistent with the fact that the ground truth models were constructed using experimental evidence from specific publications involving defined experimental models and cell types whereas the LLM’s presumably draw from the entire corpus of published literature.

      (2) Do all LLMs capture similar information, or are some LLMs better at capturing certain information than others? Further to this, it would be interesting to look into whether amalgamating information across all three LLMs results in a more accurate network.

      The reviewer asks interesting questions that can be qualitatively answered in Figure 1B and in the network visualizations (Supplementary Figures 1 and 2). These diagrams illustrate predicted connections that are shared between the nodes. To include a more quantitative, comprehensive assessment of this overlap, we have included Venn diagrams showing the extent to which LLMs capture shared information (Supplementary Figures 3-5). One such Venn diagram for the hypertrophy model is included in Supplementary Figure 3C. Indeed, it seems that there is substantial overlap in the “false positive” connections predicted by Claude and GPT. It could be interesting to evaluate if these connections are reflective of newly discovered molecular interactions, as referenced in our response to reviewer one point 3.

      While amalgamating information across all three LLMs might generate a more accurate network, these Venn diagrams illustrate that there remain connections within the ground truth models that are not represented in any of the prediction sets from the LLMs. Indeed, the union of all predicted connections for the hypertrophy network made by any of the 10 replicates from the different LLMs would still lack 33 ground truth connections (Supplemental Figure 3C).

      (3) Would it be possible to retrieve a confidence value of the interactions from the LLMs and conduct Precision, recall rate, AUPR and calibration analyses? These metrics would also help with the benchmarking process.

      We agree with the reviewer that including additional evaluative metrics is instructive. Therefore, for each reference network, we have included precision, specificity, recall, accuracy, and F1 score calculations (Supplementary Figures 3-5). Additionally, we conducted calibration analyses to illustrate the LLM’s reliability. While this latter analysis is interesting, we note that the calibration curves were generated using the frequency of predictions among the 10 replicates as a proxy for confidence. This frequency is highly sensitive to the temperature parameter of the LLMs. Different temperature settings could substantially influence the trajectories of these confidence curves and the distributions of the associated histograms. See Supplemental Figure 3B for calibration analyses of the hypertrophy model.

      We considered performing analyses resembling AUPR as suggested by the reviewer, but we did see a defensible way to vary a “threshold” for the classification of a connection as positive or negative. In this study, a connection predicted by the LLM either does or does not exist within the ground truth mode.

      Typos:

      (1) Generated networks to predict THE CLASSIC "fetal gene program" gene expression.

      Thank you for catching this error. We have corrected it.

      (2) A manually curated network has A functional accuracy of

      Thank you for catching this error. We have corrected it.

    1. eLife Assessment

      The study presents important findings revealing previously unresolved conformational dynamics of the heterodimeric type IV ABC transporter TmrAB using single-molecule FRET. The evidence presented is convincing, integrating careful experimental design with computational approaches to uncover states that are typically masked and difficult to detect. The work will be of interest to scientists studying the molecular mechanisms of primary active transport processes.

    2. Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

    3. Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATP-bound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATPbound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I mentioned that the final section of the Results part seems like an afterthought, especially since the heading suggests a broader scope.

      Reply: We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      The changes made to the section do not align with the wording of the reply. Please consider modifying it further.

      We appreciate the positive feedback and this final comment. We have revised the final section of the Results to better reflect the scope indicated by the heading. In addition to clarifying the kinetic analysis, we now explicitly relate our kinetic observations to previously determined thermodynamic measurements, showing that the rapid interconversion of ATP-bound conformations observed during steady-state turnover is consistent with a thermodynamic landscape characterized by a near-zero free-energy difference and entropy–enthalpy compensation. This revision more clearly integrates the kinetic and thermodynamic aspects of the transporter cycle.

    1. eLife Assessment

      This important study addresses a classic debate in visual processing, using a strong method applied to an impressive dataset obtained from a rare clinical population to evaluate hierarchical models of visual object perception. The paper provides compelling evidence that the hierarchical model is only partly supported: as expected, neural responses in ventral visual cortex show increased representational selectivity for faces along the posterior-anterior axes, but the onsets of the signals do not show a temporal hierarchy, indicating more parallel processing.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      Comments on revised version:

      The authors have performed additional analyses and checks and I would now rate the findings as important and compelling.

    3. Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The work is worth publishing for the data alone, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways, and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided, using resampling that preserve dependencies in a hierarchical manner, which is rare. Two types of analyses are also provided to support evidence in favour of the lack of practical significance for some of the comparisons. Collectively, the various analyses and rich illustrations provide convincing evidence in favour of the conclusions.

      Comments on revised version:

      The authors have addressed all my previous comments and the work is mostly limited by the lack of pre-registration and the exploratory nature of some of the analyses. However, with data and code available, other teams can assess the impact of researchers' degrees of freedom on the main outcomes.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      We would like to thank this reviewer for their positive comments and evaluation of our manuscript.

      Weaknesses:

      (1) A control analysis where areas have known differences in response onset should be performed to increase confidence that the proposed analyses would reveal expected results when a difference in response onset was present across areas. From Figure 3, it can be seen that many electrodes are placed in earlier visual areas (V1-V3) that have previously been shown to have earlier broadband responses to visual images compared to VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018). The same analyses as in Figures 4 and 5 should be used comparing VOTC to early visual areas to confirm that the analyses would detect that V1-V3 have earlier onsets compared to VOTC.

      First, we would like to mention that the analyses performed in our paper are commonly accepted analyses to extract time-domain information.

      Yet, the reviewer is right that, considering our claim, evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies would provide further support for our claims. The solution proposed by the reviewer is interesting but the number of face-selective recording contacts/sites in posterior occipital cortex (colored disks in figure 3) is too small for any meaningful comparison. Moreover, while absolute responses to visual images should indeed emerge earlier in early visual cortex than association cortex, it may not be the case for category-selective responses to visual images (which is what our claim is about).

      To address the reviewer’s concern, we used face-selective responses from regions known to have different response onset latencies: occipital and posterior temporal lobe vs. medial temporal lobe structures (e.g., Mormann et al., 2008, https://pmc.ncbi.nlm.nih.gov/articles/PMC2676868/). Waveforms and onset latencies for these regions (OCC, PTL, MTL) are shown in Author response image 1. Despite the small number of contacts showing significant face-selective activity in the MTL (N=20) and the lower SNR in this region, the onset latency differences between OCC/PTL and MTL are significant using all 4 methods of latency estimation (see methods in the main manuscript), with medium to large effect sizes. This was despite noisy latency estimates for the MTL (in particular, the ‘delta slope’ method could not be used to get meaningful Cohen’s d when comparing MTL to OCC). Latency estimates are also slightly higher for the ‘% of peak’ method than in the manuscript, as we estimated the latency at 25% of the peak (instead of 20% in the manuscript), again to allow meaningful estimations for the MTL.

      Author response image 1.

      In addition to this, we performed a simulation analysis where we statistically compared the measured PTL signals to ATL signals that have been artificially, incrementally, shifted forward in time. Author response image 2 shows (top row) the measured onset latencies differences between the 2 regions (PTL minus ATL) estimated using 4 different approaches as a function of the temporal shift applied to ATL, as well as the associated p-values (bottom row). As shown in Author response image 2, the ‘original’ unshifted data yields no significant difference between the 2 regions. The difference however becomes significant when ATL signals is shifted forward in time 10 to 30 ms, depending on the method used.

      Author response image 2.

      These 2 observations provide evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies, further supporting our claims.

      Last, as also suggested by reviewer 2, we conducted a thorough equivalence testing using ROPE and Bayesian factor to support the lack of differences between regions. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from early visual cortex than OCC). These analyses, now reported in the result section of the revised manuscript (Table 1), largely confirm the hypothesis of concurrent onset latencies across VOTC.

      (2) It is unclear why correlating mean timeseries helps understand how much variance is shared between regions (Figure 4). Any variance between images is lost when averaging time series across all images, and this metric thus overestimates the variance shared between areas. Moreover, the finding that correlating time domain signals across VOTC areas does not differ from correlating signals within an area could be driven by this averaging. For example, if the same analysis was done on electrodes in left and right V1 when half of the images had contrast in the left hemifield and the other half had contrast in the right hemifield, the average signals may correlate extremely well, while this correlation falls apart on a trial-by-trial basis. These analyses therefore need to be evaluated on a trial-by-trial basis.

      This is an interesting comment. We agree that variance between images is lost when averaging time-series across all images. However, to use the reviewer’s analogy, in order to support the claim that left and right V1 would show the same onset times and time-course (i.e., no hemispheric lateralization) for lateralized presentations (= the same kind of claim that we make in our paper), it’s the average response across images that should be compared, not a correlation run on a trial-by-trial basis (which would indeed falls apart because of a lack of response in the ipsilateral V1).

      Moreover, we would like to emphasize that the goal of this analysis in our paper is not to make claims about the variance shared between regions. In fact, this is not a key analysis in our paper, the outcome of which is not strictly necessary for the main argument made. Finally, if we were to perform a (time-consuming) image-by-image analysis in our study, correlations would be weak due to low signal-to-noise ratio (each face image appears only 1.6 times per stimulation sequence on average) and the fact that each face image appears after a different non-face image across presentations.

      (3) Previous studies on visual processing in VOTC have shown that evoked potentials are more predictive of the onset of visual stimuli than broadband activity (e.g. Miller et al., 2016, PLOS CB, https://doi.org/10.1371/journal.pcbi.1004660). Testing the prediction from a hierarchical representation that signals along the VOTC increase in onset time should therefore include an evaluation of evoked potential onsets in addition to broadband signals.

      We have used HFB responses in our study as these signals tend to be easier to characterize in the time domain than evoked potentials, and they are more local given their reduced SNR compared to evoked potentials (Jacques et al., 2022; https://pmc.ncbi.nlm.nih.gov/articles/PMC9457683/). Moreover we have previously shown highly correlated time courses across HFB and low frequency evoked potentials in the same paradigm (Jacques et al., 2022, eLife).

      Yet, to address this reviewer’s concern, we replicated the main analyses on low-frequency event-related potential signal, identifying contacts exhibiting significant face-selective responses in the same manner as in Jacques et al (2022). Namely, we start from bipolar-referenced sequences of recording corresponding to the full visual stimulation sequences (~70 s). For each recorded intracerebral contact, we average sequences in the time-domain, crop the average to contain an integer number of face frequency (1.2 Hz) cycles, run an FFT on the cropped sequences and identify the significant contacts with a Z-score procedure identical to that used for HFB signals. We then notch-filter out the visual response (6 Hz and harmonics) from the full length sequences, extract short epochs from the filtered sequences around the onset of each face image, average across epoch for each recording contact, subtract the mean signal measured in the baseline (-0.166 to 0 s relative to face onset) and take the absolute value (to be able to average across contacts despite differences of morphology and polarity). Significant contacts are then subjected to the same analyses as for the HFB signal.

      Results from these analyses are presented as supplementary material (Figure S9, Table S4) in the revised manuscript (referenced in lines of the main manuscript). While we were not able to obtain reliable latency estimates using the z-score method with the same parameters as for HFB signal, these analyses with ERP signal indicate similar onset latencies for ERPs as for HFB activity and largely replicate observations made with HFB. In particular, onset latencies were in a very similar range (~100 to 140 ms) with similar patterns across regions or along VOTC and between-region signal correlations. There were also a few significant face-selective responses over posterior ventro-medial occipital cortex, likely overlapping ‘early visual cortex’ (V1,V2v,V3v,hV4), probably due to limited low-level contributions in this paradigm (see Or et al., 2019, JOV; https://jov.arvojournals.org/article.aspx?articleid=2734585). Over these regions, onset latency was systematically earlier (up to 40ms) than in slightly more anterior regions, (i.e. anterior to -80 mm) where very little variability in onset latency was found up to the ATL region. We have acknowledged this in the revised manuscript.

      (4) Testing the second prediction, that the onset time of processing increases along the VOTC posterior to anterior path, is difficult using the iEEG broadband signal, because from a signal processing perspective, broadband signals are inherently temporally inaccurate, given that they are filtered. Any filtering in the signal introduces a certain level of temporal smoothing. The manuscript should clearly describe the level of temporal smoothing for the filter settings used.

      The reviewer is right that HFB signals are temporally smoothed, potentially yielding slightly underestimated onset latencies. However, our time-frequency analyses parameters ensured a minimal degree of smoothing. In fact, the original submission already contained a description of the expected temporal smoothing resulting from the wavelet transform. This is what we wrote in the original submission: “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had 20 ms of full width at half maximum (FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin), ensuring that onset timing information is accurate up to 10 ms (i.e half of the FWHM).”

      In the revised manuscript we further elaborate as follows:

      “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had a temporal spread of 20 ms (full width at half maximum - FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin). A simulation of HFB signals with a constant abrupt onset time and realistic signal-to-noise ratio indicates that the potential underestimation of onset latency due to the wavelet analysis is around 5-10 ms, which is on par with the value of the half width at half maximum (= FWHM/2 = 20/2 ms).”

      Author response image 3 displays simulated HFB signal (using identical wavelet parameters than in our manuscript) in an ideal scenario with a response starting at 150 ms in all trials (N=150 trials), reaching maximum 10 ms later. This provides a theoretical estimate of the slight underestimation of onset latency due to the wavelet transform. It shows onset latency estimates are at most 12 ms underestimation of true onset time.

      Author response image 3.

      That being said, given the physiological noise in the data, the fact that the response onset likely varies slightly from trial to trial, with a variable slope in activity increase, these wavelet parameters (within a certain margin) have likely little influence on the actual latency estimation.

      (5) The onsets of neural activity in VOTC are surprisingly early: around 80-100 ms. This is earlier than what has previously been reported. For example, the cited Quian Quiroga et al. (2023) found single neuron responses to have the earlier onset around 125 ms (their Figure 3). Similarly, the cited Jacques et al., 2016b and Kadipasaoglu et al., 2017 papers also observe broadband onsets in VOTC after 100 ms. Understanding the temporal smoothing in the broadband signal, as well as showing that typical evoked potentials have latencies compared to other work, would increase confidence that latencies are not underestimated due to factors in the analysis pipeline.

      In the revised manuscript, as suggested by reviewer 2, we have modified the data resampling strategy (using hierarchical bootstrap and permutation test that respects the nested structure of the data) to estimate onset latencies, confidence interval and permutation tests. Moreover, since the absolute onset latency estimates depend on the methods used, we now provide estimates using 4 different methods. The overall absolute onset latencies differ slightly across the 4 methods but all median onset latencies vary between 95 ms and 130 ms, which is similar to what was reported in some of the participants in Kadipasaoglu et al., 2017 (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0188834; note that latencies reported at the level of single sites or individual are usually higher due to lower signal-to-noise ratio). Moreover, these latencies for face-selective responses are actually very similar to those measured for non-selective/absolute responses to visual stimuli (e.g. Jacques et al., 2016: around 90-100 ms for faces in FG; Yoshor et al., 2007: ~100 ms in posterior Fusiform Gyrus; Regev et al. 2018: 90-120ms in posterior to middle FG). Other studies, measuring non-selective responses using more ‘conservative’ onset detection methods, find slightly later onset latencies (e.g. Cao et al. 2025: 139 ms in posterior FG; Martin et al., 2019: ~150 ms in ventral occipital -VO- regions).

      Cao R, Zhang J, Zheng J, Wang Y, Brunner P, Willie JT, Wang S. 2025. A neural computational framework for face processing in the human temporal lobe. Current Biology. DOI: https://doi.org/10.1016/j.cub.2025.02.063

      Martin AB, Yang X, Saalmann YB, Wang L, Shestyuk A, Lin JJ, Parvizi J, Knight RT, Kastner S. 2019. Temporal dynamics and response modulation across the human visual system in a spatial attention task: An ECoG study. Journal of Neuroscience 39:333–352. DOI: https://doi.org/10.1523/JNEUROSCI.1889-18.2018, PMID: 30459219

      Regev TI, Winawer J, Gerber EM, Knight RT, Deouell LY. 2018. Human posterior parietal cortex responds to visual stimuli as early as peristriate occipital cortex. European Journal of Neuroscience 48:3567–3582. DOI: https://doi.org/10.1111/ejn.14164, PMID: 30240547

      Yoshor D, Bosking WH, Ghose GM, Maunsell JHR. 2007. Receptive fields in human visual cortex mapped with surface electrodes. Cerebral Cortex 17:2293–2302. DOI: https://doi.org/10.1093/cercor/bhl138, PMID: 17172632

      As an important note, in the revised manuscript, we have removed data from 3 recording contacts from 1 participant that were located in the upper bank of the Calcarine Sulcus, which is actually outside of our VOTC region of interest. The 3 contacts being located in dorsal V1 or V2 were showing very early responses and were biasing our latency estimates for the OCC region.

      (6) Understanding the extent to which neural processing in the VOTC is hierarchical is essential for building models of vision that capture processing in the human brain, and the data provides novel insight into these processes.

      For additional context, a schematic figure of the hierarchical view and a more parallel system described in the paragraph on models of visual recognition (lines 553) would help the reader interpret and understand the implications of the paper.

      Our observations in the current study clearly indicates concurrent face-selective processing in the VOTC, which is incompatible with a serial hierarchical model. While we discuss how such concurrent activity could be implemented in the cortex (e.g., via direct input from ‘early visual cortex’ to different VOTC face-selective clusters), our data do not allow to provide more evidence in that respect to what already exists in the literature. Moreover, we are not providing data regarding connectivity (feedforward or re-entrant) either between face-selective regions or between these regions and ‘early visual cortex’.

      Author response image 4 shows a very simplified versions of standard hierarchical/serial versus concurrent/parallel models.

      Author response image 4.

      Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The data alone are valuable, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important, and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well-suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided. Collectively, the various analyses and rich illustrations provide superficially convincing evidence in favour of the conclusions.

      We thank the reviewer for their positive comments on our manuscript.

      Weaknesses:

      (1) The work was not pre-registered, and there is no sample size justification, whether for participants or trials/sequences. So a statistical reviewer should assess the sensitivity of the analyses to different approaches.

      Pre-registration of fundamental research in a clinical context is quite uncommon for intracranial data, owing, for instance to the time needed to accumulate data, or to the uncertainty of cortical sampling location in a given participant. Nevertheless, in the current study, sample size is much higher than in typical intracranial studies (usually 5-20 participants), in fact much higher than most typical Cognitive Neuroscience research. The same is true for the number of recording contacts (>10000 site here), and the number of trials considered for analysis. Each participant had a minimum of 164 face trials and an average of 262 trials (i.e. an average of 3.2 stimulation sequences of 82 trials), which is higher than most standard human electrophysiological studies.

      In the revised manuscript, to unsure that we have the maximum available power, and because our hypothesis is independent of hemisphere, we collapsed data across hemispheres for all analyses. We nevertheless provide analyses split by hemispheres as supplementary material.

      In addition, since onset latency estimations depend on the methods used, we now report onset latencies from 4 different methods (2 statistical and 2 non-statistical).

      (2) Frequentist NHST is used to claim lack of effects, which is inappropriate, see for instance:

      Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3

      Rouder, J. N., Morey, R. D., Verhagen, J., Province, J. M., & Wagenmakers, E.-J. (2016). Is There a Free Lunch in Inference? Topics in Cognitive Science, 8(3), 520-547. https://doi.org/10.1111/tops.12214

      Please see reply to the next comment (3).

      (3) In the frequentist realm, demonstrating similar effects between groups requires equivalence testing, with bounds (minimum effect sizes of interest) that should be pre-registered:

      Campbell, H., & Gustafson, P. (2024). The Bayes factor, HDI-ROPE, and frequentist equivalence tests can all be reverse engineered-Almost exactly-From one another: Reply to Linde et al. (2021). Psychological Methods, 29(3), 613-623. https://doi.org/10.1037/met0000507

      Riesthuis, P. (2024). Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science, 7(2), 25152459241240722. https://doi.org/10.1177/25152459241240722

      We thank the reviewer for pointing this out. In the revised manuscript we conduct and report a thorough examination of equivalence using ROPE and Bayesian factor to support the lack of differences between regions. We did not use TOST procedures as these have low power and require huge samples be meaningful (Riesthuis, 2024). Instead we relied on Bayes factor and descriptive proportion in ROPE. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from ‘early visual cortex’ than OCC). These analyses, now reported in the result section of the revised manuscript, along with effect sizes, confirm the hypothesis of concurrent onset latencies across VOTC.

      Riesthuis P. 2024. Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science 7.

      Detailed methods are reported as well:

      “In addition, to statistically assert whether onset latencies measured across main VOTC regions (OCC, PTL, ATL) were consistent with a concurrent (parallel) face-selective activation, we used Bayesian equivalence testing, relying on two separate metrics: (1) the percentage of differences in region of practical equivalence (ROPE), and (2) the Bayes factor using Cauchy prior. Equivalence bounds for ROPE were defined by combining two components: (1) a component of physiological variability and (2) a component reflecting expected delays in response onset between regions attributed to neural conduction delay, given the differential distances separating early visual cortex (EVC) from posterior face-selective regions (e.g. IOG) vs. anterior regions (ATL) and assuming signal mainly travels between VOTC regions through major postero-anterior axis fiber bundles of the Inferior longitudinal fasciculus (ILF) or the inferior fronto-occipital fasciculus (IFOF). Physiological variability corresponded to expected measurement noise and between-subject variability. Physiological equivalence bound was obtained by multiplying a Cohen’s d of 0.2 (conventional threshold for a negligible effect) with the (pooled) between-subjects variability in onset latency (computed for each region using a jackknife procedure and a correction factor of [N-participants – 1] to the jackknife standard deviation). For conduction delay bounds, the expected latency difference between regions under parallel activation depends on: (1) the distance between each region (considering the estimated origin of the signal is the same for all regions - EVC), and (2) neural conduction velocity. For inter-region distance we determined, for each region, the 5 and 95 percentile of the Talairach y-coordinate distribution and defined the maximal distance bounds as the distance between the y-coordinate corresponding to 5% of region 1 (e.g. OCC) to the coordinate corresponding 95% of region 2 (e.g. PTL). This resulted in the following maximum distance values: OCC-PTL: [54]mm; PTL-ATL: [54]mm; OCC-ATL: [83] mm. For conduction velocity, we used a constant value of 3.5 m/s, based on median axonal conduction velocity for cortico-cortical connections (Lemarachal et al., 2022; Van Blooijs et al., 2023), which was more conservative than using a range of values (e.g. 1.7 to 5.3 m/s based on Lemarachal et al., 2022). Maximum expected conduction delay was computed as [maximum distance / conduction speed] (e.g. for OCC to PTL: 0.054 / 3.5 = 15 ms). Resulting equivalence bounds (ROPE) were asymmetrical given than one region (e.g. OCC) is always closer to the source (EVC) than the other region (e.g. PTL) and were defined as [-1*physiologial_bound +1*physiologial_bound+max_conduction_delay]. For instance, using the z-score method to measured onset latencies, ROPE was [-12 to 27] ms for OCC to PTL, meaning that under equivalence, OCC can be activated up to 27 ms earlier than PTL (maximum conduction delay + noise), while allowing for some instances where OCC activates later (up to 12ms, due to noise only).

      For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (4) The lack of consideration for sample sizes, the lack of pre-registration, and the lack of a method to support the null (a cornerstone of this project to demonstrate equivalence onsets between areas), suggest that the work is exploratory. This is a strength: we need rich datasets to explore, test tools and generate new hypotheses. I strongly recommend embracing the exploration philosophy, and removing all inferential statistics: instead, provide even more detailed graphical representations (include onset distributions) and share the data immediately with all the pre-processing and analysis code.

      Data will be shared upon publication of the manuscript (see OSF repository in https://osf.io/2qzym). While we agree the dataset is large and could be explored in many ways, we do not consider the current study to be exploratory in nature. While our measurements could have turned out to clearly support hierarchical processing in human VOTC, our point in this manuscript is that the evidence derived from this large dataset unequivocally points instead toward concurrent activation of face-selective regions along the VOTC from IOG to antFG+ (i.e. a ~90 mm portion of cortex), with potential small variability accounted for by variability in axonal conduction velocity, signal-to-noise ratio or simple physiological variability. Other likely sources of variability such as type and density/size of fiber bundles across regions cannot easily be modeled with the current data set.

      (5) Even if the work was pre-registered, it would be very difficult to calculate p-values conditional on all the uncertainty around the number of participants, the number of contacts and the number of trials, as they are random variables, and sampling distributions of key inferences should be integrated over these unknown sources of variability. The difficulty of calculating/interpreting p-values that are conditional on so many pre-processing stages and sources of uncertainty is traditionally swept under the rug, but nevertheless well documented:

      Kruschke, J.K. (2013) Bayesian estimation supersedes the t test. J Exp Psychol Gen, 142, 573-603. https://pubmed.ncbi.nlm.nih.gov/22774788/

      Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin & Review, 14(5), 779-804. https://doi.org/10.3758/BF03194105 https://link.springer.com/article/10.3758/BF03194105

      All analyses and preprocessing stages are identical between regions and the number of trials is large enough not to be a constraining factor. As indicated above and below, we now report detailed equivalence testing and effect sizes and recomputed all statistics, taking into account the structure of the data as suggested by the reviewer.

      (6) Currently, there is no convincing evidence in the article to clearly support the main claims.

      Bootstrap confidence intervals were used to provide measures of uncertainty. However, the bootstrapping did not take the structure of the data into account, collapsing across important dependencies in that nested structure: participants > hemispheres > contacts > conditions > trials.

      Ignoring data dependencies and the uncertainty from trials could lead to a distorted CI. Sampling contacts with replacement is inappropriate because it breaks the structure of the data, mixing degrees of freedom across different levels of analysis. The key rule of the bootstrap is to follow the data acquisition process, and therefore, sampling participants with replacement should come first. In a hierarchical bootstrap, the process can be repeated at nested levels, so that for each resampled participant, then contacts are resampled (if treated as a random variable), then trials/sequences are resampled, keeping paired measurements together (hemispheres, and typically contacts in a standard EEG experiment with fixed montage). The same hierarchical resampling should be applied to all measurements and inferences to capture all sources of variability. Selectivity and timing should be quantified at each contact after resampling of trials/sequences before integrating across hemispheres and participants using appropriate and justified summary measures.

      The authors already recognise part of the problem, as they provide within-participant analyses. This is a very good step, inasmuch as it addresses the issue of mixing-up degrees of freedom across levels, but unfortunately these analyses are plagued with small sample sizes, making claims about the lack of differences even more problematic--classic lack of evidence == evidence of absence fallacy. In addition, there seem to be discrepancies between the mean and CI in some cases: 15 [-20, 20]; 8 [-24, 24].

      In light of the reviewer’s comment, we recomputed all timing analyses using a stratified hierarchical approach to evaluate confidence intervals (using bootstrapping), statistical comparisons (using permutation tests) and equivalence testing.

      This is what we wrote in the revised methods:

      “The first two timing parameters of face-selective response, onset and offset latencies, were quantified per main VOTC region using a hierarchical bootstrapping approach to respect the nested structure of the data (region > participants > contacts > trials). For each bootstrap iteration and each region, we first sampled participants with replacement. Within each sampled participant, we then sampled contacts and then trials within sampled contacts, with replacement. Resampled trials, then contacts within participants, then participants within a region, were successively averaged to obtain a bootstrapped region-level response from which we derived onset latency (4 different methods) and offset latency. We obtained bootstrap distributions of onsets/offsets using 2000 bootstrap iterations per region, allowing to compute the median and 95% confidence interval for these 2 parameters.”

      Then, later about permutation tests:

      “Statistical significance of latency differences between main VOTC regions was assessed using a hierarchical permutation test. We use a stratification approach to partition participants into a paired set (i.e. participants that had recording contacts in the two regions compared) and unpaired set (participants with contacts in a single region). For paired participants, the region labels were randomly swapped within subject (i.e., exchanging the signals from the two regions), thereby preserving all participant-, contact-, and trial-level structure while breaking the association between region and latency estimates. For unpaired participants, participants were randomly reassigned between regions while preserving the original group sizes, to generate pseudo-groups under the null hypothesis of no regional difference. In each permutation, signals were averaged across trials, then contacts, then participants within each permuted group and latency was computed and stored from the resulting region-level signals. We performed 10000 permutations to obtain a distribution of regional differences of latencies under the null hypothesis and determine the p-value as the fraction of the null distribution larger or smaller than the observed (non-permuted) difference.”

      And then about equivalence testing:

      “For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (7) Three other issues related to onsets:

      (a) FDR correction typically doesn't allow localisation claims, similarly to cluster inferences: Winkler, A. M., Taylor, P. A., Nichols, T. E., & Rorden, C. (2024). False Discovery Rate and Localizing Power (No. arXiv:2401.03554). arXiv. https://doi.org/10.48550/arXiv.2401.03554

      Rousselet, G. A. (2025). Using cluster-based permutation tests to estimate MEG/EEG onsets: How bad is it? European Journal of Neuroscience, 61(1), e16618. https://doi.org/10.1111/ejn.16618

      In fairness, we do not understand or share the reviewers’ concern here. Hundreds of fMRI or EEG studies use FDR or cluster tests to make inference about spatial or temporal location. We use FDR correction in one of the onset latency estimation method and only consider one-sided differences. Other methods in the revised manuscript do not use FDR correction.

      (b) Percentile bootstrap confidence intervals are inaccurate when applied to means. Alternatively, use a bootstrap-t method, or use the pb in conjunction with a robust measure of central tendency, such as a trimmed mean.

      Rousselet, G. A., Pernet, C. R., & Wilcox, R. R. (2021). The Percentile Bootstrap: A Primer With Step-by-Step Instructions in R. Advances in Methods and Practices in Psychological Science, 4(1), 2515245920911881.

      Again, we are not sure what the reviewer’s is referring to. The confidence intervals are computed on latency estimates from bootstrapped waveforms. In the revised manuscript, these waveforms are obtained by averaging (i.e. mean) resampled trials, resampled channels, resampled participants. A trimmed mean could not be applied in this condition, except perhaps when averaging across trials. But then the trimmed mean would have to be applied separately at each time sample which would disturbed within-, or between-trial, variability.

      (c) Defining onsets based on an arbitrary "at least 30 ms" rule is not recommended:

      Piai, V., Dahlslätt, K., & Maris, E. (2015). Statistically comparing EEG/MEG waveforms through successive significant univariate tests: How bad can it be? Psychophysiology, 52(3), 440-443. https://doi.org/10.1111/psyp.12335

      The rule of contiguous significant points is a heuristic that many researchers have used successfully to avoid spurious detection due to temporal autocorrelation. While we are aware that more sophisticated methods exist to correct for autocorrelation, such as cluster-based approaches, it is not directly usable since it requires comparing 2 conditions. The approach described in Piai et al., 2015 is interesting but incorrect as well since it relies on split-half simulations, which reduced signal-to-noise ratio, resulting in over estimated correction to be applied. In our revised manuscript, we rely on multiple methods to estimate onset latency, some of which not relying on this heuristic. Moreover, we apply plausible physiological constrains to our latency estimates, such as rejecting any onset before 40 ms after stimulus onset.

      (8) Figure 5 and matching analyses: There are much better tools than correlations to estimate connectivity and directionality. See for instance:

      Ince, R. A. A., Giordano, B. L., Kayser, C., Rousselet, G. A., Gross, J., & Schyns, P. G. (2017). A statistical framework for neuroimaging data analysis based on mutual information estimated via a Gaussian copula. Human Brain Mapping, 38(3), 1541-1573. https://doi.org/10.1002/hbm.23471

      (9) Pearson correlation is sensitive to other features of the data than an association, and is maximally sensitive to linear associations. Interpretation is difficult without seeing matching scatterplots and getting confirmation from alternative robust methods.

      We rely on Pearson correlation because this replicates the method used in Kadipasaoglu et al., 2017. It is also a widely accepted measure of (linear) relationship (in our situation we did expect linear or near linear relationships) in the literature. To address the reviewers concern, in the revised manuscript we nevertheless report, as supplementary material (Figure S10), the same functional connectivity analyses performed using the methodology and code provided in Ince et al. (2017). The results of this analyses are extremely similar to the results using Pearson’s coefficients.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 6, the response onset latencies are rendered in a smoothed manner on the brain surface. However, with this smoothing, variability between electrodes cannot be seen, and it would be better visualized in color, rendered in each electrode.

      Latency estimates computed at individual channels are noisy, which is why we do not report individual channels latencies but rather rely on averaging signals across contiguous channels, either across whole regions (Figure 4) or across smaller volumes as in Figure 6.

      (2) Onset latencies of 60 seconds seem extremely early compared to literature typically citing evoked responses with a latency of ~170ms. It would help if some additional sanity checks were shown, such as showing the latency of early visual responses. This would help with relative comparisons.

      In the revised manuscript the earliest median latency is 95 ms, which is in line with previous intracranial electrophysiology literature (e.g. Jacques et al., 2016; Jacques et al., 2022 ; https://pubmed.ncbi.nlm.nih.gov/26212070/; https://pubmed.ncbi.nlm.nih.gov/36074548/). The 99% confidence intervals can result in earlier latencies both due to some participants showing early responses and noise in latency estimates. Also please keep in mind that latency estimates are usually earlier when combining data across channels/participants compared to individual channels simply due to differences in SNR or across participants (see e.g. Kadipasaoglou et al., 2017).

      The reviewer indicates “…to literature typically citing evoked responses with a latency of ~170ms.”. We are assuming that they refer to the face-selective N170 ERP component measured on the scalp in EEG. Even with this ERP component, the face-selective response usually starts around 120-130 ms after stimulus onset (e.g. Rousselet et al., 2008; Jacques, Retter and Rossion, 2016; https://pubmed.ncbi.nlm.nih.gov/18831616/; https://pubmed.ncbi.nlm.nih.gov/27138205/) at scalp level. With the same highly sensitive paradigm as used here in EEG, we have systematically shown latency onsets of face-selective activity shortly after 100 ms (e.g., Retter et al., 2020; also Quek & Rossion, 2017) not accountable for by low-level visual cues (i.e., not present for phase-scrambled stimuli; Rossion et al., 2015; Or et al., 2019). Our latency onsets are also in line with spiking activity recorded with the same approach in the LatFG (Laurent et al., 2026) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5955677

      (3) Figure 6B shows that the variability of latencies in the ATL is larger than the variability of latencies in the PTL. It would be helpful to evaluate whether, rather than in mean onset latencies, there is a change in variability in onset latency along the VOTC.

      This is an interesting point. However, it is difficult to evaluate since SNR is reduced in the ATL compared to OCC or PTL (Jacques et al., 2022). As a result, any measured modulations in the variability of onset latencies along VOTC may simply reflect changes in the precision of latency estimation driven by SNR variability.

      (4) Line 415 typo: 'there appears to be no delay' instead of 'there appear to be no delay'.

      We thank the reviewer for their careful reading of our manuscript. This has been corrected.

      (5) The discussion states in lines 455-457 that "a large proportion of neuronal populations in anterior VOTC regions exhibiting similar activity to different face images independently of the context in which they appear". However, this claim about similar activity to different faces should be evaluated and tested at the single-trial level. In addition, it is not clear how context was varied in the experimental design.

      Context is variable because each face image appears directly after a different object image (or object images) in the sequence. We have shown also in previous studies with this paradigm in EEG that the time-course of face-selective responses is similar across base frequencies (3-15 Hz) unless the rate is too fast, and whether an orthogonal or explicit face categorization task is used (Retter et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32119982/). Note that we do not claim that activity is identical across images but similar – if it was not (largely) similar, the averaged response would be jittered and low.

      (6) Line 516-517: DTI does not provide evidence for whether connectivity is direct or not, and what the directionality of connectivity is between two areas. This sentence should therefore state "..., suggest independent connections between early visual cortex and face-selective regions...".

      This has been rephrased.

      (7) Line 555 in the discussion, the definition of low-level visual should be expanded to include other early visual areas that have been demonstrated to respond earlier than VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018), to avoid the suggestion that V1 directly projects synaptically to all of VOTC (e.g. Markov et al., 2014, Cerebral Cortex, https://doi.org/10.1093/cercor/bhs270).

      We are not proposing that V1 directly projects directly/synaptically to all of VOTC, i.e., without other low-level retinotoptic areas involved; only that face-selectivity in the association cortex is not organized hierarchically. We have revised this sentence.

      (8) Line 585, for the sentence: "with temporal synchrony strengthening their connections", evidence or citations should be provided.

      Citations have been provided.

      (9) It is not clear what is meant in the paragraph starting in line 571: do the authors suggest that top-down signals are not necessary for fast recognition of clear views of faces, or additionally argue that these top-down signals are not necessary for detecting ambiguous or degraded inputs as faces?

      Exactly: That top-down (i.e., descending) signals may contribute but would not be necessary for fast recognition of clear views of faces AND for detecting ambiguous or degraded inputs as faces.

      (10) No statement was provided on data or code availability.

      Data will be made available on a repository upon publication (see https://osf.io/2qzym).

      Reviewer #2 (Recommendations for the authors):

      (1) FDR correction: which one? Please provide a reference.

      We now provide a reference, both in the results and methods: Benjamini and Hochberg, 1995.

      (2) In the introduction, this statement is too strong: "arguably the most familiar and ecologically valid stimulus". It is unclear how static 2D representations of faces are the most familiar and valid stimuli. Could you rephrase this? What about other very familiar stimuli like letters, words and biological motion?

      This statement is not about static 2D images of faces, but faces in general (in their natural environment). We do consider human faces (in general, not restricted to laboratory context) to be indeed the most familiar and ecologically important stimulus, both from an ontogenetic and phylogenetic perspective, unlike written material.

      (3) About the questioning of a strict temporal hierarchy, this EEG reference comes to mind: Foxe, J. J., & Simpson, G. V. (2002). Flow of activation from V1 to the frontal cortex in humans. Experimental Brain Research, 142(1), 139-150. https://doi.org/10.1007/s00221-001-0906-7

      As confirmed by Foxe et al. ’s (2002) paper to which the reviewer is referring to, there is indeed ample evidence that areas in the dorsal stream or frontal cortex (e.g. FEF) are activated very soon after V1 and before many ventral stream regions (e.g. Lamme and Roelfsema, 2000; https://pubmed.ncbi.nlm.nih.gov/11074267/). While Foxe et al.’s 2002 is highly valuable, it can hardly be compared with our current study which looks specifically into ventral stream areas which are largely indistinguishable using scalp EEG as in Foxe et al. ’s paper.

      (4) Regarding statistical significance, there is no such thing as a "trend". The threshold for a trend should have been pre-registered and applied to both sides of the magical boundary, for instance, with matching conclusions for a "trend toward non-significance (p=0.04)". P values near 0.05 provide weak support against the null. I would suggest leaving it at that. Nothing special happens at 0.05.

      This no longer appears in the revised manuscript.

    1. eLife Assessment

      This important study provides convincing evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial. By combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia.

      Comments on revised version:

      Thank you for the opportunity to re-review this interesting paper.

      This reviewer struggled to follow along with the author's response letter. It was challenging because responses were brief and did not list the specific changes made, leaving the reviewer to search for changes in the text. As best I could tell, there were no changes made that matched some of the highest-priority suggestions.

      Critique #1 - This reviewer did not find the presentation of confidence intervals in the Abstract and other sections, which were suggested. Please note the format that was suggested in the original critique from Reviewer 1.

      Critique #2 - regarding removal of unsubstantiated claims and use of a p-value to compare analysis to a null hypothesis - it seems the authors skipped over this critique and did not address it.

      Regarding bias: the authors answered a different question than the one asked. The reviewer asked what proportion of transmissions were sampled; the authors stated that only communities from which phylo data was acquired were modeled. Was sampling 100% in those communities? Please provide the percentage and provide analysis that shed light on how sampling bias could impact the analysis.

      Regarding "cherries" - the reviewer did not understand the author's response. The query was regarding what percent of the total number of phylogenetic pairs (denominator) were the 355 that had high confidence in directionality (numerator). The response could be expressed be a proportion.

      The expectation of ART reducing the age of sources of transmission seems unrealistic to this reviewer. People on ART are not always adherent and can still transmit during gaps in adherence. ART dramatically increases life expectancy with HIV, which would have the opposite effect.

    3. Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex specific transmission dynamics in Zambia with a goal of allocation of resources.

      Strengths:

      Important analysis to hone in on key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches used, and while the phylogenetic data was markedly more limited, it mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses and providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce.

      Comments on revised version.

      The revised manuscript clarifies the impact and utility of this work and better allows the comparability of the two methods. Highlighting the differences (or lack thereof) between the undiagnosed and diagnosed population) simplifies the public health approach.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial; by combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength. However, the evidence supporting key conclusions is incomplete, with several claims insufficiently substantiated by the data presented. Improvements in data presentation (e.g., quantification of qualitative statements, statistical estimates, and clearer description of results) would substantially strengthen the paper.

      We thank the editor and reviewers for their positive comments. We have revised the manuscript in response to the points raised, as described below.

      First of all, we would like to summarise what we have changed regarding the presentation of summary statistics throughout the text. We agree that many of the statements in the original submission tended towards being qualitative. This was the result of shying away from presenting two separate estimates, with different ways of quantifying uncertainty, in the text. The phylogenetics could be presented as mean and confidence interval, while the IBM would need some measure of centrality (mean or median) and the highest density interval for a summary statistic (e.g. the mean age gap) as it varies over the posterior. These are not directly comparable. We have now changed this to present both where appropriate, with cautionary note about the difference between the CIs and HDIs (lines 257-260).

      We also were somewhat arbitrary regarding where we chose to summarise the posterior in the IBM or look at the best-fitting single simulation, and where we presented the mean as opposed to the median. We have done a considerable overhaul of what is presented in this revision:

      (1) We always present the posterior summary unless the level of detail is such that summarising uncertainty over the posterior is not feasible (e.g. in figures 3, 4 and 5). In the latter case we still use the best-fitting IBM replicate.

      (2) In the main text we always present the mean. For the phylogenetics the summary statistics are mean and confidence interval. For the IBM this is the posterior mean, and 95% HDI, of the mean of a particular statistic as calculated in each of the 1000 IBM replicates. For example, each replicate will have its own distribution of male source ages which have a mean value. These means also vary over the posterior, and a mean of them is calculated, as well as the HDI interval to represent posterior uncertainty. This “mean of means” may be a slightly confusing piece of terminology at first glance, but it allows us to properly capture posterior uncertainty in a way we mostly avoided in the first submission.

      One result of 1) above is a change to figure 6. It is now summarised over the posterior, with the result that time trends that were previously not evident become clear. This changes our conclusions slightly (lines 529-537) but it should be noted that the magnitudes of the trends remain small.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia. The current manuscript draft is methodologically strong, but needs revision to strengthen the take-home messages. As written, there are many possible take-away conclusions. For example, the agreement between IBM and phylogenetic analysis is noteworthy and provides a methodological focus. The revealed age patterns of transmission could be a focus. The effects of the PopART intervention and the consequences of a 1-year disruption could be a focus. It is important, though, that any main messages summarized by the authors are substantiated by the evidence provided and do not extrapolate beyond the data that have been generated. I recommend that the authors think deeply about what the most important, well-supported messages are and reframe the discussion and abstract accordingly.

      We have rewritten the abstract, and also made changes to the discussion in order to centre our message around the contribution of particular of demographic groups to transmission, and how, with that contribution revealed, such groups can be selected for specialised interventions.

      Strengths/weaknesses by section:

      (1) ABSTRACT

      The Abstract summarizes qualitative findings nicely, but the authors should incorporate quantitative results for all of the qualitative findings statements.

      The abstract in the revision is extensively revised, and contains quantitative estimates throughout, from both methodologies where appropriate.

      The ending claim is not substantiated by the modeling scenarios that have been run: "targeted interventions for demographic groups such as under-35 men may be the key to finally ending HIV." It is straightforward to run this specific scenario in the model to determine whether or not this is true.

      Our modelling framework is not set up to model the “last mile” of HIV elimination, notably as it has no component for MSM or FSW transmission, and we do not feel that we could confidently present results regarding it. As a result, this statement has been greatly softened in the new abstract (lines 75-78).

      The authors should add confidence intervals to the quantitative metrics, such as the 93.8% and 62.1% incidence reduction.

      These have been added.

      (2) RESULTS

      The authors should check the Results section for any qualitative claims not substantiated by the analyses performed, and ensure the corresponding analyses are presented to support the claims.

      The Results and Methods describe the model's implementation of the PopART intervention differently. The Methods describes it as including VMMC, TB, and STI services, while the Results only mentions intensified HIV testing and linkage.

      This is a slight misreading of the text. That paragraph in the Methods is describing the trial itself, not the modelling framework.

      A limitation of the model is that HIV disease progression is based on the ATHENA cohort in the Netherlands, which is a different HIV subtype (B) than the one in the research setting (C). The model should be configured using subtype C progression data, which have been published, or at least a sensitivity analysis should be conducted with respect to disease progression assumptions.

      The available literature does not suggest a significant difference in progression between subtypes B and C, and we have added text and citations to this effect (lines 699-701).

      In Table 2, the authors should consider adding a p-value to establish whether or not IBM and phylogenetics estimates are different.

      We have done this; the appropriate test was a posterior predictive check. See lines 261-263, 575-579 and 805-814.

      (3) DISCUSSION

      The literature review and comparison of study results to previously published phylogenetic studies is very nice. The authors could strengthen this by providing quantitative estimates with CIs for a more scientific comparison of the study results vs. prior studies, perhaps as a table or figure.

      We have expanded the discussion on this point (lines 504-527). We considered adding a table, but the existing literature that directly answers the questions we ask is quite limited and fragmentary. For example, Monod et al do not present a complete treatment of age gaps. The literature using regression analyses to identify predictors of HIV prevalence or incidence related to partner age is extensive, but those results are not directly comparable to ours.

      The authors state that due to "the narrow geographical catchment area... The results should not be automatically extrapolated to apply to other SSA settings." The authors should exercise this caution when comparing the results to studies in South Africa and elsewhere.

      We have made more explicit acknowledgements of these limitations (lines 598-600).

      There are many other limitations to the analysis, including some mentioned above, that are not acknowledged. The authors should think carefully about what the most important limitations are and acknowledge them honestly at the end of the Discussion section.

      The limitations paragraph has been revised (lines 598-605).

      Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex-specific heterosexual HIV transmission dynamics in Zambia, with the goal of allocating resources.

      Strengths:

      Important analysis to hone in on the key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches were used, and while the phylogenetic data were markedly more limited, they mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses. The authors did a nice job of providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce

      Weaknesses:

      To increase the impact and utility of this work, it would be helpful to parse the analysis just a bit further to estimate the roles of undiagnosed vs diagnosed and untreated subpopulations on this transmission. PopART is a multifaceted intervention, but the cost, effort, and approach to reengagement in care vs testing/treatment can be quite different.

      We have now provided stratified results by diagnosed and non-diagnosed status of the source, as well as an overall summary of the proportion of undiagnosed sources by age and sex. See lines 305-310, 539-547, and table 3.

      Recommendations for the authors:

      Reviewing Editor:

      We commend you for conducting a rigorous and comprehensive study titled "The age and sex dynamics of heterosexual HIV transmission in Zambia: an HPTN 071 (PopART) phylogenetic and modelling study" that significantly advances the understanding of HIV transmission dynamics in sub-Saharan Africa. The study utilizes an innovative dual-methodology approach integrating individual-based mathematical modelling (IBM) and pathogen phylogenetics to characterize heterosexual HIV transmission patterns by age and sex during the PopART trial in Zambia.

      This manuscript reports on HIV transmission dynamics in Zambia using data from the PopART study, combining individual-based modelling and phylogenetic analysis. The use of two independent methodologies enhances confidence in the consistency of the findings and enables robust cross-validation. The work addresses an important topic in HIV prevention, particularly in settings where resources may become more constrained, and offers insight into potential demographic targets for intervention.

      However, several aspects of the manuscript limit its current impact. The main take-home messages are diffuse and not clearly presented. Some conclusions in the abstract and discussion appear to go beyond the scope of the presented data. For instance, the claim that targeting under-35 men may be key to ending HIV is not directly tested in the modelling scenarios and should be reframed or removed unless supported by new analyses. Furthermore, important quantitative details, such as confidence intervals, p-values, and precise age group estimates, are lacking in key sections (e.g., the Abstract and Results).

      The authors are encouraged to clearly identify and communicate their central findings, ensure all claims are fully supported by their analyses, and make the data more accessible to readers by adding detailed, quantitative summaries where needed.

      The following are our recommendations to the Authors:

      (1) Clarify Study Objectives and Central Messages

      Reframe the abstract and discussion to highlight a clear, well-supported set of main findings.

      Avoid overgeneralized or unsubstantiated claims, especially those not directly tested by your model (e.g., the effectiveness of targeting under-35 men).

      As stated above, we have revised this text accordingly.

      (2) Support Qualitative Claims with Quantitative Data

      Provide numerical results, including effect sizes and confidence intervals, wherever qualitative trends are mentioned.

      For example, restate: "The largest gaps for female recipients were among the youngest" as "... in the age group XX-YY with OR = Z.Z (95% CI: A.A-B. B)."

      As mentioned at the top of the review, we have overhauled the treatment of summary statistics extensively, and now give confidence or highest density intervals throughout the text.

      (3) Improve the Results Section

      Check that all claims are supported by the analyses, and ensure figure references are accurate.

      The statements that went beyond what was supported, notably about ending the epidemic by targeting young men, have been removed. The typo in table references has been fixed.

      Annotate Figure 6 with trendline coefficients and p-values where applicable.

      The takeaway message of figure 6 has now changed and we no longer see no trend, just a minor one.

      Revise Figure 4 for clarity or consider replacing it with a tabular format.

      We would prefer to keep the current figure 4, as we have not found any clearer way to illustrate the patterns, which are the consequence of the phenomenon observed in figure 5. We have put more explicit descriptive text in the discussion, linking the two figures (lines 470-476).

      (4) Address Potential Bias and Model Assumptions More Rigorously

      Explain sampling bias in IBM and phylogenetics (e.g., how the 355 high-confidence phylogenetic pairs were selected).

      The reviewer comment regarding the 355 pairs was based on a misapprehension; we used all the pairs we found using the phyloscanner pipeline. There are no sampling bias issues involved in the IBM as every individual in the simulations is considered. Appendix 2 includes some sensitivity analysis results if the procedure used to find the 355 is changed.

      Discuss how the use of subtype B disease progression data from the ATHENA cohort may impact results in a subtype C setting. A sensitivity analysis would strengthen this.

      Subtype B progression data was used in the absence of any appropriate data from subtype C, but the literature does not suggest any major difference between the two (lines 699-701).

      (5) Include More Detail on Undiagnosed Populations and ART Effects

      Estimate the roles of undiagnosed and untreated subpopulations in driving transmission.

      As mentioned above, this analysis has been added.

      Clarify mechanistically how ART might influence age gaps in transmission dynamics.

      This now is clarified in the introduction (lines 127-129).

      (6) General Improvements

      Provide p-values where comparisons are made (e.g., in Table 2).

      Use consistent terminology and definitions across Methods and Results.

      Add more discussion on limitations, especially regarding generalizability to other SSA settings.

      All of these have been inserted as previously mentioned.

      By addressing these points, the manuscript would present a more coherent narrative and a stronger, evidence-based contribution to the field. We appreciate you all for your fantastic effort and hope you will reflect the feedback in your final paper.

      Reviewer #1 (Recommendations for the authors):

      Thank you for the opportunity to review this interesting manuscript.

      In the public review, I have recommended that the authors should incorporate quantitative results for all of the qualitative findings statements. As one example, I would recommend that "We found the largest gaps for female recipients were among the youngest of those recipients" is re-written as "The largest gaps for female recipients were in the age group XXX-YYY with OR=ZZZ (XXX-YYY)." such as odds ratios, and specific outcome definitions including ages. To give one more example: "immediate increase in the average age at transmission of both sources and recipients" could be rephrased as "increase in the average age at transmission by XXX (YYY-ZZZ) years for sources and XXX (YYY-ZZZ) for recipients over [TIME PERIOD]."

      We hope the revisions we have made to the statistical presentation are satisfactory as a response to this request.

      Again in the public review, I recommended checking the Results section for any qualitative claims not substantiated by the analyses performed, and ensuring the corresponding analyses are presented to support the claims. An example is: "Trends are minor or non-existent in the former two variables." - please annotate Figure 6 (assuming the authors meant to reference Figure 6 and not 7 here?) to show over what period trendlines were fit and provide the coefficient and CI. To support the stated claim even more strongly, a p-value might be apt with a null hypothesis of a slope of zero.

      Please check the numbering on all figure references in the text, as some appear to be misnumbered. E.g., where the text refers to Figure 7, I believe the authors meant to reference Figure 6.

      The change to how we handled the statistics has changed the message of figure 6 (which is now figure 7) and rendered this somewhat moot. We have checked that all figure and table references are now correct.

      Figure 3 is very nice, but if the axes were flipped on one panel, it would make them easier to compare, and then adding some statistics to assess whether the patterns are the same or different when a man vs woman is the source.

      We have flipped the axes here.

      Figure 4 was too complicated for me. I could not follow the Sankey flows because there is too much going on and overlapping. Consider revising to make it easier to digest... perhaps to table format?

      As mentioned above, we would prefer to keep this figure, but we have situated it better in the text.

      Reviewer #2 (Recommendations for the authors):

      A few points that would improve the clarity and the strength of the manuscript

      (1) There is a need to clarify more about how the IBM and phylogenetic data does not suffer from sampling bias. For e.g.,

      Line 205: What proportion of the transmissions modeled in the IBM from Zambia?

      All of them. We confined the analysis of the IBM to the Zambian communities from which phylogenetic data was acquired (lines 755-758).

      Line 217: What proportion of the phylogenetic pairs (cherries) suggesting transmission were the 355 that had high confidence in directionality. How do these pairs compare to the others

      There was no identification of “cherries” involved in picking these pairs; the phyloscanner procedure does not use that step. We confined our analysis solely to the pairs for which we did identify a direction of transmission; that is the 355. Appendix 2 includes a sensitivity analysis involving varying the parameters by which these were identified.

      (2) I appreciate the authors noting that MSM transmissions are unlikely to be playing a role in this cohort, as noted in previous work by the group. However, systematic undersampling of men is common in other study cohorts of HIV. While the MSM and heterosexual networks may be relatively distinct, undersampled men who are bridging the networks could impact the estimates. Can the authors use the time to diagnosis analysis (HIV phyloTSI) to estimate rates of undiagnosed men and women?

      We feel that this is beyond the scope of this work. The phylogenetics dataset in its totality could be used for this purpose (although it is probably highly biased towards undiagnosed individuals due to the considerable majority of samples coming from the healthcare facilities). However, we concentrate here solely on the subset involved in our probable transmission pairs, which is fairly small. Extending the scope to an exploration of the full dataset would seem like a separate study, which we do have plans to do.

      We have used the IBM for this question instead (lines 303-321), however, as MSM transmission was not modelled, it is also not ideal for answering this question. Ultimately we feel that the way these studies were implemented makes it an unsatisfactory tool for answering the MSM question, important as it is.

      (3) Expanding on the point above, in other settings, transmission to young men has been associated with partnerships with older men, and if these young men then transmitted to young women, would we see a similar effect as noted in these models (assuming the young men were less well sampled).

      Our previous work (Hall et al., 2024) suggested no excess of identified male-male pairs in the phylogenetics dataset which might suggest cryptic male-to-male transmission. The age disparities would be worth exploring had this been found, but is curtailed by the lack of it.

      (4) Related to the point above, is there an estimate of the populations (age and sex) that are undiagnosed in the IBM model? Can this be teased out... is transmission from men to women more likely 2/2 lack of diagnosis... or lack of engagement in care?

      We have explored results by diagnostic status as it pertains to age and sex, but we feel that moving on to a more general exploration of the role of diagnosis and lack of engagement in care is again going beyond the scope of what is already a long paper.

      (5) I'm still not fully clear as to why ART might affect age gaps. Can this be explained in more detail?

      See lines 127-129.

    1. eLife Assessment

      This valuable study describes PXGS, a poly-transgene expression system that exploits the mutually exclusive splicing of Dscam variable exon 4 to enable conditional, simultaneous expression of up to 12 transgenes in Drosophila, addressing a longstanding limitation in which conditional co-expression has been restricted to a handful of genes. The approach is conceptually elegant and technically accessible, with potential applications spanning neuroscience, synthetic biology, and biomanufacturing across arthropod species. The evidence that Dscam exon 4 splicing is preserved in a UAS vector and that individual alternates can be replaced with functional transgenes is solid, and the in vivo axonal re-wiring application provides a convincing proof of principle. Quantitative characterization of expression levels, a direct demonstration of expression across all twelve positions, and additional imaging controls would further substantiate the system's utility and scope.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes the development of an expression system enabling up to 12 transgenes using the alternatively spliced fourth exon of *Drosophila* *Dscam* gene under the control of a UAS element. This will be a useful tool if expression is needed in *Drosophila* cells (in culture or in vivo). where *Dscam* splicing machinery is active, which limits its use.

      Strengths:

      The tool developed is based on a well-established genomic element. The underlying idea is relatively simple yet effective.

      Weaknesses:

      The authors describe the weaknesses of their system well, most importantly, depending on the presence of adequate levels of Dscam splicing factors in targeted cells. This likely limits effective use of the methodology to some cell lines (e.g., S2) and certain tissues (nervous system and innate immune system). The manuscript could do a better job in showing protein expression levels more quantitatively, either in comparison to other methods or as absolute values (transcript numbers, protein molarity, etc.).

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Yu et al seek to develop a Drosophila genetic tool to simultaneously co-express up to 12 transgenes. They leverage the native Dscam exon 4 alternative splicing to generate a UAS to enable cell- and temporal-specific expression of transgenes. This tool is called the poly-transgene expression system (PXGS). Previous approaches to co-express transgenes have been limited to four to five genes, so PXGS would be a significant advancement, especially when examining processes that require robust expression of many genes to confer function. The authors showed that PXGS can drive expression of multiple (1) fluorescent reporters and (2) cell surface receptors in different cell types (neurons, glia, and muscles). However, there are major proof-of-principle experiments missing to demonstrate the utility of PXGS and its potential limitations. Additionally, some of the data is just not interpretable, and experimental rigor is significantly lacking.

      Strengths:

      Developing a genetic tool to co-express transgenes beyond what is currently available would be significant.

      Weaknesses:

      (1) While the authors stated that each PXGS construct can express 12 transgenes, this was not directly tested - the largest number of genes tested was in the PXGS_fluorophores, which has 4 genes inserted in 10 alternates (and therefore it can be determined if genes from all the alternates are spliced in at a meaningful level).

      a. First, the authors state that they tested the expression of the fluorophores in S2 cells using RT-PCR before generating the fly line. However, this data is not shown. Also, it is possible to test expression and localization of UAS transgenes in S2 cells with a ubiquitous GAL4, similar to what they did for GFP expression in Supplemental Figure 1.

      b. In the corresponding figure for this experiment (Figure 2), there are some concerning expression patterns and potential channel bleed-through/crosstalk. The nSyb-Gal4 is a pan-neuronal driver, yet expression of three fluorophores was extremely minimal. This could potentially be explained by the deterministic vs random alternative splicing. However, it is more concerning that in the GFP, RFP, and iRFP channels, the exact same tiny cluster of neurons is observed, suggesting potential bleed-through of the channels. Appropriate controls are required, including expression of a traditional fluorescent reporter with the nSyb-Gal4 they are using. And replicates would help with the experimental rigor.

      c. Additionally, it would be helpful to know if genes inserted at each alternate exon are expressed at a similar efficiency (vs. some alternate exons have higher levels of expression)

      (2) PXGS expression in non-neuronal cells: The authors attempt to show that fluorophores targeted to different cellular compartments can be expressed in neurons and non-neuronal cells (glia and ubiquitously).

      a. In Supplemental Figure 2A, they first use S2 cells to confirm expression, which does confirm. However, they use no markers to show that the fluorophores localize to the corresponding compartments (e.g., mitochondria and nucleus). In Supplementary Figure 2B, with that magnification and resolution, it is impossible to determine if the fluorophores localize properly.

      b. In Figure 3, it is impossible to know if there is any glial expression based on those images. They state that "subcellular localization of fluorophores was observed in the flight muscle", but again, that cannot be concluded from the images. Also, the schematic of the construct is the same one used in Supplementary Figure 2. Why show it again?

      (3) In the functional expression of PXGS transgenes section, while the authors used RT-PCR to show that each receptor gene is transcribed in S2 cells, it was not tested if they are correctly expressed, translated, and localized in the fly. The functional outcome observed (Figure 4 b-c) could be the result of misexpression of one or multiple genes.

      a. Supp Figure 3: Why does the Sli lane have so many bands?

      b. Why were these genes chosen for misexpression? Is there any evidence that they are required for the wiring of the mechanosensory neuron? Does the co-expression lead to an additive effect?

      c. The RT-PCR result (Supplemental Figure 3) showed variations (e.g., kek and kir being significantly dimmer than tutl; multiple products for Sli). Is there any explanation behind this, and could this be the outcome of some alternate exons being more efficiently spliced than others?

      d. Figure 4: These images seem to be taken with a widefield scope and only one plane. Is it possible that some of the pSC neurons are in a different Z plane, and they are not being captured here? There definitely is part of the axon terminal out of focus in some of the images. Also, most of the figure graph axes (e.g., 4b) are extremely difficult to read. And the figure overall is not easy to interpret.

      (4) Supplemental Figure 5: This figure is quickly mentioned in the Discussion without much explanation. First, this must be in the Results section since it is an experiment. Second, this needs more context because, as is, it seems like it was just thrown into the manuscript.

      (5) The authors mentioned that the size of the inserted genes could be a limitation for this technique and tested cell surface receptors of different sizes. However, there was no explicit discussion in the main text.

    4. Reviewer #3 (Public review):

      This paper, by Brian Chen and collaborators, adapts the highly alternatively spliced Dscam1 gene locus for use in a system for simultaneous multi-transgene expression in a variety of insect species. Specifically, they show that the hypervariable Dscam1 exon 4 region maintains its alternative splicing when placed in a UAS expression vector, and that each of the twelve exon 4 alternates can be replaced with an exogenous gene such that co-expression of up to twelve proteins can be achieved. Since the co-expression of more than a few proteins simultaneously is difficult, this represents a significant advance with multiple use cases. The authors validated the technique by assessing expression in vitro and in vivo, and by rewiring Drosophila sensory neuron axons by simultaneously expressing several cell surface receptors within the neuron. Overall, this is a clearly written paper that describes a potentially important new system. I have no major criticisms.

    5. Reviewer #4 (Public review):

      From the Reviewing Editor:

      All three reviewers recognized the conceptual originality of PXGS and its value to the Drosophila community and the broader multi-gene expression field. The core demonstration - that Dscam exon 4 mutually exclusive splicing is maintained in an exogenous UAS vector, and that individual exon alternates can be replaced with genes of interest for conditional in vivo expression - was viewed as solid and creative. The in vivo application re-wiring pSc axonal arbors using PXGS constructs loaded with cell surface receptors was noted as an encouraging functional validation.

      The reviewers differed in their overall enthusiasm. Reviewer 3 found the experiments straightforward, the results clear, and the system a significant advance with broad use cases, with no major criticisms. Reviewer 1 viewed the evidence as broadly solid, with the principal limitation being a lack of quantitative expression data and a dependence on adequate Dscam splicing-factor levels that constrain the system's applicable cell types - a limitation the authors themselves describe well. Reviewer 2 was the most critical, finding the underlying concept significant but the supporting data insufficiently rigorous to conclusively establish the tool's utility. The points below reflect the areas where reviewers - principally Reviewers 1 and 2 - felt the manuscript could be strengthened, should the authors choose to revise.

      (1) Quantitative characterization of expression. Reviewers 1 and 2 both noted the absence of a quantitative comparison of PXGS-driven expression - in absolute terms (transcript numbers, protein amounts) or relative to standard UAS constructs. Given that signal is inherently divided across 12 alternates per transcription event, characterizing expression efficiency and whether all alternates are spliced and expressed at comparable levels would substantially strengthen the manuscript. Reviewer 2 specifically asked whether genes at different exon 4 positions are expressed with similar efficiency.

      (2) Direct demonstration of 12-transgene expression. Reviewer 2 noted that, although the manuscript claims expression of up to 12 transgenes, this was not directly tested - the largest construct placed 4 distinct genes across 10 alternates, and the largest functional test used 3 genes per construct. Either a direct demonstration with more positions occupied or a more carefully bounded claim supported by the probabilistic framework would address this.

      (3) Controls and interpretability of fluorophore expression. Reviewer 2 raised concerns about Figure 2, where expression of three fluorophores under nSyb-Gal4 was minimal, and the same small neuronal cluster appeared across the GFP, RFP, and iRFP channels - raising the possibility of channel bleed-through. Appropriate controls (including a conventional single UAS-fluorophore driven by the same nSyb-Gal4) and replicates were requested. Reviewer 2 also noted that the S2 cell validation data for the fluorophore constructs, described as having been performed prior to fly line generation, are not shown.

      (4) Non-neuronal expression and subcellular localization. Reviewer 2 noted that the compartment-specific localization claims (mitochondria, nucleus) in Supplemental Figure 2 and Figure 3 are not supported by co-markers, and that the magnification and resolution in key panels are insufficient to confirm proper localization, particularly in glia.

      (5) Functional expression of receptor constructs. Reviewer 2 noted that, while RT-PCR confirms transcription of each receptor in S2 cells, correct translation and localization in the fly were not directly tested, so the observed phenotypes could reflect mis-expression of one or a subset of the genes. Clarification of why the specific genes were chosen, whether co-expression produces additive effects, and the cause of the variable RT-PCR band patterns (e.g., the multiple Sli products, dimmer kek and kir signals) was requested.

      (6) Supplemental Figure 5 (synthetic biology / RNAi). Reviewer 2 noted that this figure is mentioned only briefly in the Discussion despite representing an experiment, and recommended moving it to the Results with appropriate context. Reviewer 1 separately queried the meaning of the two white boxes in this figure and whether the RT-PCR convincingly supports expression of all genes shown.

      (7) Figure quality and labeling. Both reviewers flagged that the labels in Figure 4 (particularly panel d) and the axes in Figure 4b are too small to read. Additional labeling points were noted (alignment of "Repo-GAL4" and "brain" in Figure 3; unlabeled images in Supplemental Figures 2 and 4; clarification of whether the two lanes per group in Figure 1c are replicates or use different primers). Reviewer 1 also suggested that some figure legends describe conclusions rather than what is shown, and recommended that legends describe the data with interpretation kept to the text.

      (8) Additional points. Reviewer 1 suggested showing more of the gel in Figure 1 to demonstrate the absence of non-spliced fragments; clarifying the 1/12 probability argument (or moving it to the Discussion with transcript-number context); providing sequences for the fluorophore variants and fusion tags; and minor prose corrections ("Regardless if" → "Regardless of whether"; "dependent on three things" → "dependent on three factors"). Reviewer 2 noted that the gene-size limitation, though tested, is not explicitly discussed in the main text.

    6. Author response:

      We are pleased that the reviewers viewed the core demonstration (that Dscam mutually exclusive splicing is preserved in a vector and that exon alternates can be replaced with genes of interest) as a solid foundation for the system. We agree that the manuscript would be strengthened by clearer quantitative characterization of expression, additional controls for fluorophore imaging, improved presentation of the figures, and more precise wording about the current scope of evidence. In a revised manuscript, we plan to address these points by adding or clarifying quantitative expression analyses, including S2 cell validation data, adding appropriate imaging controls where available, revising claims about 12-transgene expression to distinguish design capacity from direct experimental demonstration, and improving figure labels and legends throughout.

      We also plan to expand the discussion of PXGS limitations, including cell-type dependence on Dscam splicing machinery, possible position effects, and gene size considerations. Finally, we will improve Methods reporting by adding resource identifiers, cell culture quality control information, statistical design details, and data/code availability statements where appropriate.

      We appreciate the opportunity to revise the manuscript and believe these changes will make the strengths and limitations of PXGS clearer to readers.

    1. eLife Assessment

      This study presents a valuable finding on linking the frequency of neural activity to cortical depths of blood flow in a naturalistic setting of participants listening to music. The presentation of evidence in the version of the original submission is incomplete, as further clarifications in methods and results, as well as performing additional analyses, would strengthen the study. The work will be of interest to cognitive neuroscientists working on multimodal recordings, auditory perception and music.

    2. Reviewer #1 (Public review):

      Summary:

      In their submitted work, Lee and colleagues examine the correlation between electrophysiological activity as measured by SEEG, and layer-specific activation patterns, as measured through 7T fMRI. This analysis was performed using patients undergoing monitoring for epilepsy surgery guidance, as well as healthy controls, as they both listened to music.

      They find that, in general, higher-frequency SEEG activity correlated positively with the fMRI signal, while lower frequencies correlated negatively. Across cortical depth, higher-frequency activity correlated positively with middle-to-upper layers, whereas lower frequencies showed their strongest negative correlations in superficial layers.

      Strengths:

      This is an interesting physiological study in that, to the best of this reviewer's knowledge, it has not been done before with auditory stimuli using the combination of iEEG (as opposed to scalp EEG) and fMRI. The framework fits well with models of layer-specific feedforward versus feedback processing (e.g., Bastos et al., 2012).

      Weaknesses:

      Its main limitations are a lack of specificity to the acoustic stimuli, the absence of correction for venous draining, and the fact that it is largely a replication/port of prior work.

    3. Reviewer #2 (Public review):

      Summary:

      The authors present an investigation of the relationship between the iEEG frequency bands signal and hemodynamic responses at different cortical depths. Based on this, the authors aim to uncover the layered origin of iEEG signals at different frequencies. The authors then interpret their results in terms of feedforward and feedback processing, arguing that the correlations between fMRI and iEEG signals reflect the interaction between both processes. In addition, the authors aim to infer the extent to which these processes are involved during naturalistic music processing.

      Strengths:

      This study combines the neural recording methodologies yielding the highest spatio-temporal precision achievable in humans, while using naturalistic auditory stimuli. This combination of recording methods and experimental design offers key insights regarding the precise origin of iEEG signals, which is necessary to improve the interpretability of future iEEG studies.

      Weaknesses:

      (1) The current framing of the paper leads the authors to interpret their findings in ways that are not warranted by the data. The main analysis of the paper consists of correlating the hemodynamic responses from different layers with iEEG signals from different frequency bands, which enables us to infer the relationship between the two signals. It does not, however, enable us to draw inferences regarding the extent of feedforward and feedback processing and the interaction between the two during naturalistic auditory processing. This would require comparing hemodynamic responses in different cortical layers or frequency bands activation against some baseline condition. Based on the presented analysis, statements such as "our frequency-specific results demonstrate that naturalistic music perception seamlessly integrates both feedforward and feedback processing streams" should be removed.

      (2) The presentation of existing literature omits key details and findings, making it difficult to fully understand the research question the authors are trying to address. For example, the author mentions studies showing that feed-forward and feedback processing are segregated across cortical layers and that feed-forward and feedback processing have distinct time-frequency signatures (lines 47-58). However, the authors do not mention which cortical layer or which frequency band is associated with which kind of processing. As a result, it is difficult for the reader to determine what exact hypothesis the author is trying to test in the study. This might also relate to the confusion raised in (1).

      (3) The method section omits key details. When describing the paradigm, the authors do not describe how the tones were presented, nor how the signals were synchronized. Similarly, there is no mention of the pipeline used for iEEG electrodes localization. In addition, the exact regressors that entered the generalized linear model of hemodynamic responses are not clearly stated: were all regressors (frequency bands + acoustic signal + HFA) entered together in a single model or in separate models? The mention of a cubic spline is also not sufficient for the reader to understand what was done and for which purpose. Finally, the exact tests used for some comparisons are omitted (in Figure 2, for example, no mention of the exact test used to compare betas between A1 and A2). The current structure of the method section is also quite difficult to follow: the authors switch back and forth between describing acquisition protocols and participant counts, for example.

      (4) The lack of methodological details (as described in point 3 above) casts doubts about the validity of some of the statistical tests reported. Throughout the paper, the authors present quantitative statements and statistical tests comparing the fitted beta parameters between brain regions (A1 and A2) and cortical depths. However, the authors do not mention any normalization procedure taken to ensure that the scale of the signals being compared was equated. If the overall magnitude of the signals in A1 differs from that of A2, the mean of the beta distribution is expected to differ as well. Similarly, if the signal-to-noise ratio differs between brain regions or cortical layers, so should the variance of the beta parameters across subjects, which might break the homoscedasticity assumption of some tests, which might or might not be a problem depending on the exact test the authors used (hence the importance of reporting them).

    1. eLife Assessment

      This study provides a useful anatomical resource by mapping the expression of four putative chemoreceptors in spinal cerebrospinal fluid-contacting neurons (CSF-cNs) of larval zebrafish. These descriptive findings offer an interesting entry point to explore how the nervous system senses signals within the spinal fluid microenvironment. The evidence supporting the spatial expression patterns of these receptors is convincing, utilizing high-resolution hybridization chain reaction (HCR) to validate previous transcriptomic data. However, the evidence remains incomplete regarding the actual functional roles of these receptors, as the study lacks protein-level validation, evidence of ligand availability in the CSF, or functional assays to demonstrate active chemoreception.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript examines the expression of putative chemoreceptors in CSF-contacting neurons of the larval zebrafish spinal cord. Using in situ hybridization, the authors show that sstr2a is preferentially expressed in ventral CSF-cNs, whereas grm2a, ptprna, and ldlrad2 are detected in both ventral and dorsolateral CSF-cNs, with additional expression in neighboring cells around the central canal.

      Strengths:

      The study provides useful anatomical information on the expression of putative chemoreceptors in CSF-contacting neurons. The experiments appear to be carefully performed, and the results are clearly presented with high-quality illustrations and informative schematics.

      Weaknesses:

      This work remains largely descriptive and based on mRNA expression. Therefore, the proposed roles in chemoreception, ligand sensing, lipid capture, or long-range CSF signaling remain speculative without protein-level or functional validation.

    3. Reviewer #2 (Public review):

      Summary:

      Verran et al. leverage a previously published RNAseq dataset of zebrafish cerebrospinal fluid contacting neurons (CSF-cNs) to identify potential receptors involved in chemosensory signalling in these neurons. They then validate expression of the identified receptors by hybridization chain reaction (HCR) in zebrafish larvae. This way they uncover potential roles for the somatostatin receptor Sstr2a, metabotropic glutamtate receptor Grm2a, LDL receptor Ldlrad2 and the Phosphatase receptor Ptprna, suggesting the existence of numerous chemo-sensory pathways in CSF-cNs and providing a potential entry point for further investigation.

      Strengths:

      This is a useful resource; the provided HCR data that demonstrates expression of these receptors in CSF-cNs is convincing, and the finding that CSF-cNs express these receptors is interesting.

      Weaknesses:

      The overall insight provided by this manuscript is rather limited, essentially just demonstrating the expression of 4 receptors in CSF-cNs, whose expression was predicted to be enriched in these neurons anyway by a previously published dataset.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed to identify new molecular pathways that could enable long-range signaling through the cerebrospinal fluid (CSF), focusing on a specialized class of neurons called CSF-contacting neurons (CSF-cNs) in larval zebrafish.

      Strengths:

      Anatomical validation of transcriptomic candidates using HCR, providing high-resolution spatial mapping of chemoreceptor expression in CSF-contacting neurons and neighboring spinal cord cells. The work broadens the potential understanding of CSF-cNs and offers a resource for future functional investigations of CSF-mediated signaling.

      Weaknesses:

      The principal limitation of the study is that the conclusions remain largely transcriptomic and inferential. Although HCR convincingly validates mRNA expression, no protein-level evidence is provided to demonstrate receptor translation or subcellular localization, leaving uncertainty regarding functional receptor availability at the CSF interface. Moreover, the study does not establish whether the proposed ligands are present in the relevant CSF microenvironment or engage the identified receptors in vivo. As such, the functional significance of the proposed chemosensory pathways remains speculative.

    1. eLife Assessment

      This work provides a reassessment of VBIT-4, a compound previously proposed to inhibit oligomerization of the crucial protein known as the mitochondrial voltage-dependent anion channel. Combining complementary experimental approaches with molecular dynamics simulations, the authors provide compelling evidence that VBIT-4 primarily disrupts lipid membranes and induces channel-independent cytotoxicity. The study has important implications for interpreting previous work using VBIT-4 as a probe of channel function and highlights the need to consider membrane-disruptive effects when evaluating drug mechanisms.

    2. Reviewer #1 (Public review):

      Summary:

      The Voltage-Dependent Anion Channel 1 (VDAC1) is the most abundant β-barrel protein in the outer mitochondrial membrane and the main conduit for metabolite and ion exchange between the cytosol and mitochondria. Its oligomerization has been proposed to control mitochondrion-mediated apoptosis, making it a prime target for therapeutic intervention in diseases associated with excessive cell death, such as neurodegenerative disorders and autoimmunity. VBIT-4 is a small molecule developed to inhibit VDAC oligomerization and has shown therapeutic potential in various preclinical models. Despite its widespread use, the mechanism of action of VBIT-4 has not yet been fully elucidated. In this paper, Ravishankar et al. combine a suite of biophysical approaches with computer simulations to demonstrate that VBIT-4 forms water-permeable defects in membrane bilayers without any detectable effects on VDAC1 channel properties or oligomerization. Furthermore, cytotoxicity assays revealed identical VBIT-4 IC50 values in wild-type and VDAC1-KO cells, indicating that its activity does not depend on VDAC1. Collectively, these findings cast significant doubt on the widely held assumption that VBIT-4 is a specific inhibitor of VDAC1 oligomerization. Instead, it appears that VBIT-4 functions as a membrane-active compound.

      Strengths:

      This is a carefully conducted and well-written study that highlights potential side effects of VBIT-4, a compound that has been used to study the role of VDAC1 in a range of physiological and pathological conditions. The work is of interest to a broad readership by showcasing the importance of a systematic assessment of drug-membrane interactions to identify potential off-target membrane-driven effects of small molecules that may be mistakenly attributed to the inhibition of specific proteins. Its strength lies in the variety of complementary approaches the authors used to rigorously challenge the effect of VBIT-4 on VDAC1 organization and function. Overall, the experimental data are compelling and of high quality.

      Weaknesses:

      The authors used high-speed atomic force microscopy (HS-AFM) to study the impact of VBIT-4 on VDAC1 oligomerization in real time at nanoscale resolution. Toward this end, they adsorbed POPC:POPE:cholesterol membranes reconstituted with or without VDAC1 on mica. This revealed that the addition of VBIT-4 produced small perforations in the bilayer that were independent of VDAC1. In the absence of VBIT-4, VDAC1 showed the characteristic honeycomb topography that the authors described in a previous study (Reference 17). To quantitatively assess whether VBIT-4 affects VDAC1 organization, they analyzed protein compaction within clusters using inter-protein distance measurements. This analysis revealed no significant difference in VDAC1 organization between control conditions, 1 uM and 10 uM VBIT-4, supporting a model in which VBIT-4 primarily perturbs the lipid matrix rather than VDAC1 assemblies. This conclusion is based on the assumption that VDAC channels retain some lateral mobility in bilayers adsorbed onto mica. Do the authors have evidence that this is indeed the case? Did they also perform HS-AFM on VDAC1-containing membranes treated with VBIT-4 prior to adsorption onto mica?