10,000 Matching Annotations
  1. May 2026
    1. Reviewer #1 (Public review):

      Summary:

      This study investigated how visuospatial attention influences the way people build simplified mental representations to support planning and decision-making. Using computational modeling and virtual maze navigation, the authors examined whether spatial proximity and the spatial arrangement of obstacles determine which elements are included in participants' internal models of a task. The study developed and tested an extension of the value-guided construal (VGC) model that incorporates features of spatial attention for selecting simpler task mental representation.

      Strengths:

      (1) Original Perspective: The study introduces an explicit attentional component to established models of planning, offering an approach that bridges perception, attention, and decision-making.

      (2) Methodological Approach: The combination of computational modeling, behavioral data, and eye-tracking provides converging measures to assess the relationship between attention and planning representations.

      (3) Cross-validated data: The study relies on the analysis of three separate datasets, two already published and an additional novel one. This allows for cross-validation of the findings and enhances the robustness of the evidence.

      (4) Focus on Individual Differences: Reports of how individual variability in attentional "spillover" correlates with the sparsity of task representations and spatial proximity add depth to the analysis.

      Appraisal of Aims and Results:

      The study sets out to determine how spatial attention shapes the construction of task representations in planning contexts. The authors provide evidence that spatial proximity and arrangement influence which environmental features are incorporated into internal models used for navigation, and that accounting for these effects improves model predictions. There is clear documentation of individual variation, with some participants showing greater attentional spillover and more sparse awareness profiles.

      Comments on revised version:

      The authors did a great job and I am very happy with the revised manuscript.

    2. Reviewer #2 (Public review):

      Summary:

      Castanheira et al. investigate the role of spatial attention for planning during three maze navigation experiments (one new experiment and two existing datasets). Effective planning in complex situations requires the construction of simplified representations of the task at hand. The authors find that these mental representations (as assessed by conscious awareness) of a given stimulus are influenced by (spatially) surrounding stimuli. Individual participants varied in the degree to which attention influenced their task representations, and this attentional effect correlated with the sparsity of representations (as measured by the range of awareness reports across all stimuli). Spatially grouping task-relevant information on either the left or right side of the maze led to mental representations more similar to optimal representations predicted by the value-guided construal (VGC) model - a normative model describing a theoretical approach to simplifying complex task information. Finally, the authors propose an update to this model, incorporating an attentional spotlight component; the revised descriptive model predicts empirical task representations better than the original (normative) VGC model.

      Strengths:

      The novelty of this study lies in the proposal and investigation of a cognitive mechanism through which a normative model like value-guided construal can enable human planning. After proposing attention as this mechanism, the authors make concrete hypotheses about mismatches between the VGC predictions and real human behavior, which are experimentally validated. Thus, not only does this study describe a possible mechanism for simplification of task information for planning, but the authors also propose a descriptive model, revising VGC to incorporate this attentional component.

      A strength of this paper is the variety of investigative approaches: analysis of existing data, novel experiment, and a computational approach to predict experimental findings from a theoretical model. Analyzing pre-existing datasets increases the size of the participant cohort and strengthens the authors' conclusions. Meanwhile, comparing the predictions of the existing normative model and the authors' own refined model is a clever approach to substantiate their claims. In addition, the authors describe several crucial controls, which are key to the interpretability of their results. In particular, the eye tracking results were critical.

      In summary, this paper constitutes an important step toward a more complete understanding of the human ability to plan.

      Comments on revised version:

      I am overall happy with the revision and agree that the authors have addressed most of the comments.

    3. Reviewer #3 (Public review):

      Summary:

      The authors build on a recent computational model of planning, the "value-guided construal" framework by Ho et al. (2022), which proposes that people plan by constructing simple models of a task, such as by attending to a subset of obstacles in a maze. They analyze both published experimental data and new experimental data from a task in which participants report attention to objects in mazes. The authors find that attention to objects is affected by spatial proximity to other objects (i.e., attentional overspill) as well as whether relevant objects are lateralized to the same hemifield. To account for these results, the authors propose a "spotlight-VGC" model, in which, after calculating attention scores based on the original VGC model, attention to objects is enhanced based on distance. They find that this model better explains participant responses when objects are lateralized to different hemifields. These results demonstrate complex interactions between filtering of task-relevant information and more classical signatures of attentional selection.

      Strengths:

      (1) The paper builds on existing modeling work in a novel manner and integrates classic results on attention into the computational framework.

      (2) The authors report new and extensive analyses of existing data that shed light on additional sources of systematic variability in responses related to attentional spillover effects

      (3) They collect new data using new stimuli in the original paradigm that directly test predictions related to the lateralization of task-relevant information, including eye tracking data that allows them to control for possible confounds.

      (4) The extended model (spotlight-VGC) provides a formal account of these new results.

      Comments on revised version:

      I also agree that the authors addressed our comments and the manuscript is much stronger now.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      This study investigated how visuospatial attention influences the way people build simplified mental representations to support planning and decision-making. Using computational modeling and virtual maze navigation, the authors examined whether spatial proximity and the spatial arrangement of obstacles determine which elements are included in participants' internal models of a task. The study developed and tested an extension of the value-guided construal (VGC) model that incorporates features of spatial attention for selecting simpler task mental representation.

      Strengths:

      (1) Original Perspective:

      The study introduces an explicit attentional component to established models of planning, offering an approach that bridges perception, attention, and decisionmaking.

      (2) Methodological Approach:

      The combination of computational modeling, behavioral data, and eye-tracking provides converging measures to assess the relationship between attention and planning representations.

      (3) Cross-validated data:

      The study relies on the analysis of three separate datasets, two already published and an additional novel one. This allows for cross-validation of the findings and enhances the robustness of the evidence.

      (4) Focus on Individual Differences:

      Reports of how individual variability in attentional "spillover" correlates with the sparsity of task representations and spatial proximity add depth to the analysis.

      We thank the Reviewer for their overall positive assessment of our work and their helpful comments. We have addressed each point below.

      Weaknesses:

      (1) Clarity of the VGC model and behavioral task:

      The exposition of the VGC model lacks sufficient detail for non-expert readers. It is not clear how this model infers which maze obstacles are relevant or irrelevant for planning, nor how the maze tasks specifically operationalize "planning" versus other cognitive processes.

      The method for classifying obstacles as relevant or irrelevant to the task and connecting metacognitive awareness (i.e., participants' reports of noticing obstacles) to attentional capture is not well justified. The rationale for why awareness serves as a valid attention proxy, as opposed to behavioral or neurophysiological markers, should be clearer.

      We thank the reviewer for urging further clarity here. Our work builds closely on the previous maze navigation paradigm and VGC model developed and reported by Ho et al. Nature (2022). We directly adopted variants of their maze stimuli, computational model and obstacle awareness measures, and married these with an investigation of the role of visuospatial attention. We agree that it would be useful for the reader to have a more in-depth description of the paradigm and model, and how it operationalises planning, without needing to refer back to the original Ho et al. paper. We have now added additional explanatory sections to the Introduction and Methods as follows:

      On page 4:

      “One elegant approach to forming such a simplified representation is to adaptively select the granularity of information required to complete the task (Ho et al., 2022), known as value-guided construal (VGC). Unlike previous accounts, which model human planning as a search over all items (e.g.., tube lines), the VGC model predicts that a cognitively limited decision-maker selects a manageable subset of information over which to plan— i.e., a task representation—balancing utility and complexity (Ho et al., 2022). In our example, the VGC algorithm would plan over a few relevant tube lines rather than planning over all possible stations. To select the representation that achieves the best balance between utility and complexity, the model searches across all possible combinations of tube lines, computing the value (i.e., the plan’s utility minus its cost) of each representation for planning a specific journey. The algorithm then selects the representation with the highest value, which ensures that an ideal observer selects a representation which only includes the items (i.e., tube lines) that lead to successful planning while excluding as many items as possible to keep the plan as simple as possible. For our purposes, items included in the representation are considered taskrelevant, while items that are not represented are considered task-irrelevant. This algorithm, therefore, provides a normative standard of an efficient plan to which we can compare people’s actual plans.”

      On page 6:

      “We operationalized planning using a maze navigation paradigm, akin to our tube-related example, where participants were required to plan a route through the maze, avoiding obstacles that blocked their path. Obstacles predicted by the sVGC model to be included in the representation were considered task-relevant.”

      “At the end of every trial, participants reported their awareness of specific obstacles (see Methods for details). The level of awareness reported for different obstacles provides a read-out of what features of the environment individuals were subjectively representing while solving a particular maze. While other markers of attention and awareness (for instance, behavioural or neurophysiological variables) could also be used, here we focused on direct awareness reports in order to relate our findings both to those of Ho and colleagues and to the subjective awareness reports used in consciousness science (e.g. the Perceptual Awareness Scale (Barnett et al., 2024; Overgaard & Sandberg, 2021; Ramsøy & Overgaard, 2004; Samaha et al., 2015)). Participants were instructed to maintain central fixation while planning (see dataset dSC 1), in line with previous empirical work using this task (Ho et al., 2022).”

      To visualize our effects, we binarized the predictions of the sVGC model such that obstacles with a marginalized probability greater than 0.5 were considered taskrelevant, while other obstacles were considered task-irrelevant (e.g., Figure 2b). We now clarify this point in the caption of Figure 2.

      (2) Attention framework:

      The account of attention is largely limited to the "spotlight" model. When solving a maze, participants trace the correct trail, following it mentally with their overt or covert attention. In this perspective, relevant concepts are also rooted in attention literature pertaining to object-based attention using tasks like curve tracing (e.g., Pooresmaeili & Roelfsema, 2014) and to mental maze solving (e.g., Wong & Scholl, 2024), which may be highly relevant and add nuance to the current work. This view of attention may be more pertinent to the task than models of simultaneously tracking multiple objects cited here. Prior work (notably from the Roelfsema group) indicates that attentional engagement in curve-tracing tasks may be a continuous, bottom-up process that progressively spreads along a trajectory, in time and space, rather than a "spotlight" that simply travels along the path. The spread of attention depends on the spatial proximity to distractors - a point that could also be pertinent to the findings here.

      Moreover, the tracing of a "solution" trail in a maze may be spontaneous and not only a top-down voluntary operation (Wong & Scholl, 2024), a finding that requires a more careful framing of the link to conscious perception discussed in the manuscript.

      Conceptualizing attention as a spatial spotlight may therefore oversimplify its role in navigation and planning. Perhaps the observed attentional modulation reflects a perceptual stage of building the trail in the maze rather than a filter for a later representation for more efficient decision making and planning. A fuller discussion of whether the current model and data can distinguish between these frameworks would benefit readers.

      We thank the reviewer for highlighting relevant findings in the attention literature that were missing from our discussion. We fully agree that a complete account of the interplay of planning, navigation, and attention is likely to recruit the kind of curvetracing processes highlighted by the reviewer. However, we emphasise that our current focus is not on the process of navigation through a maze, but on the process of construing the maze itself. In other words, we are focused not on how people represent their path from A to B, but how they represent the maze itself, which they then use as a basis for planning between A and B. The VGC model predicts that a subset of obstacles will be included in this construal. We think that a spotlight model is a good starting point for this work, because attention is being deployed across the whole maze stimulus, and then becomes attached to particular objects located in particular positions. This is a distinct process from that involved in navigating the path itself. Accordingly, our stimuli were designed such that task-relevant obstacles could be presented either proximally or distally to the optimal path (e.g., Figure 1a and Supplemental Figures S1-6). An obstacle that blocks any possible path on one side of the maze is task-relevant but located a long way from the optimal path. The results of Ho and colleagues’ (2022) third experiment demonstrate how task-relevant yet distal obstacles are better remembered than task-irrelevant proximal obstacles (see Figure 4 of Ho et al., 2022). We also observed that obstacles further away from the navigation path were often represented by participants (see Figures S1-6), which cannot be explained by curve tracing alone.

      While these results cannot definitively rule out the possibility that participants automatically trace the path while also construing the maze, they suggest that the value-guided construal process is an independent predictor of participants’ representations beyond proximity to the navigated path. To make this distinction clearer, we now cite the papers alluded to by the reviewer, in the Discussion on pages 28-29, while also acknowledging the potential for investigating attention during the navigation process itself:

      “Future work may also wish to examine the relevance of visuospatial attention for the navigation process itself in this task. While our present findings speak to how individuals perceive the maze while planning, it remains unclear how attention is deployed during navigation along a path, such as how object-based attention progressively spreads along trajectories in time and space(Pooresmaeili & Roelfsema, 2014; Wong & Scholl, 2024).”

      There is also one additional nuance to the current spotlight model that we were inspired to consider by the reviewer’s comment. This is the idea that attentional effects may spread within or along the obstacles themselves. We cannot explore this in the current data because we asked for awareness of the entire obstacles, not parts of obstacles, but it may be possible to explore this in future work, for instance, with eye tracking measures.

      More generally, the growth-cone (i.e., zoom lens) model of attention for curve tracing proposed by Roelfsema and colleagues shares considerable similarities with the spotlight of attention model. Both models argue for the grouping of spatially proximal items based on attention. While the growth-cone model argues for varying sizes of zoom lenses (i.e., receptive fields of neurons) that facilitate the tracing of proximal items, both models predict that spatially proximal items are preferentially processed together because of attention. Indeed, the spotlight model could model these varying zoom lenses by altering the width of the attentional spotlight dynamically across the visual scene based on the spatial proximity of obstacles. Following related comments by Reviewer 2, we now investigate inter-individual differences in the attentional spotlight of participants and observed that these differences significantly predict participants’ mental representations (see Attentional spotlight model of task representations). We have now updated the Discussion to include consideration of these alternative model frameworks:

      On page 27:

      “Second, in the current work we were unable to distinguish whether these attentional effects are driven by a fixed spotlight of attention, or whether attention operates akin to a zoom lens, shifting the ‘width’ of the focus of attention according to the task demands (Eriksen & St. James, 1986; Müller et al., 2003; Schad & Engbert, 2012). The latter view would be consistent with growth-cone models of attention in which the focus of attention expands and contracts in accordance with task demands, mirroring the various receptive field sizes in the visual hierarchy (Pooresmaeili et al., 2014; Pooresmaeili & Roelfsema, 2014). In partial support of this idea, we found significant inter-individual differences in the width of participants’ attentional spotlight (Figure S11). It is also possible that attention is deployed within or along parts of obstacles, rather than on entire obstacles. Future work using naturalistic measures of eye movements may be able to address these questions.”

      (3) Lateralization of attention:

      The analysis considers whether relevant information is distributed bilaterally or unilaterally across the visual display, but does not sufficiently address evidence for attentional asymmetries across the left and right visual fields due to hemispheric specialization (e.g., Bartolomeo & Seidel Malkinson, 2019). Whether effects differ for left versus right hemifield arrangements is not made explicit in the presented findings.

      We thank the reviewer for this suggestion. To address this point, we fitted a three-way interaction model between VGC model prediction, lateralization index, and side (left vs right hemifield). We did not find evidence for the three-way effect (β= 0.01, SE= 0.02, 95% CI [-0.03, 0.04], p = 0.738; ΔBIC = 58.30 in favour of the null effect; see table below), suggesting that the side to which participants lateralized their attention did not influence their task representations. This result is now reported on page 12:

      “This effect did not vary significantly as a function of the specific hemifield (i.e., left vs right) in which task-relevant information was presented (β= 0.01, SE= 0.02, 95% CI [-0.03, 0.04], p = 0.738; ΔBIC = 58.30 in favour of the null effect; see table S14).”

      We also explored inter-individual differences in participants’ tendency to lateralize their attention (see also the next point). We observed that participants tended to lateralize their attention slightly more to the right-hand side for non-lateralized maze stimuli, despite the normative sVGC model predicting that participants should not lateralize their attention for these stimuli (Figure 3c). These results may speak to potential asymmetries in lateralization, but given the exploratory nature of these analyses, they should be verified and replicated in future work.

      (4) Individual differences:

      Individual differences in attentional modulation are a strength of the work, but similar analyses exploring individual variation in lateralization effects could provide further insight, and the lack of such analyses may mask important effects.

      Thank you for this suggestion. In new analyses, we explored whether i) participants exhibited differences in their tendency to lateralize their awareness reports, and ii) whether the degree to which they tended to lateralize their awareness predicted their performance on a separate set of maze stimuli. In short, we observed substantial variation in participants’ tendency to lateralize their awareness (Figure S11) and found that this tendency reflected an inter-individual difference which was stable across maze types. We report these new findings on pages 14-16.

      “Inter-individual variation in lateralization of attention

      Next, we investigated participants’ tendency to pay attention to obstacles within a single hemifield (left vs right) regardless of the sVGC model predictions. To do so, we computed an awareness lateralization index (ALI) based on participants’ self-reported awareness reports of obstacles on each trial (Figure 3a). Large positive values indicate that participants were preferentially aware of the right hemifield, whereas negative values indicate preferential awareness of the left hemifield. Values close to zero indicate that participants paid attention to both hemifields equally (see Methods for details). We observed that participants’ tendency to lateralize their awareness varied greatly across the Ho datasets 1 and 2 (Figure 3b); some participants preferentially paid attention to a single hemifield, regardless of whether the sVGC model predictions were lateralized. For the dSC1 dataset, we observed that on some trials, participants significantly lateralized their awareness (|ALI| > 0.5; Figure 3c) even though the sVGC model predictions were non-lateralized. These findings suggest that participants’ tendency to pay attention to a single hemifield may represent an observable inter-individual difference in how they allocate their awareness to form task construals.”

      “To further explore these inter-individual differences, we tested whether participants’ tendencies to lateralize their attention to a single hemifield was consistent across trials and maze stimuli. We observed that participants’ tendency to lateralize their attention to a single hemifield was similar for left and right lateralized maze stimuli (Spearman ⍴= 0.72, Figure 3d). This suggests that participants who preferentially attended to a single hemifield did so regardless of which hemifield they should attend to. More consequentially, the tendency for participants to lateralize their awareness on maze stimuli whose model predictions were also lateralized linearly correlated with participants’ tendency to lateralize their attention on non-lateralized maze stimuli (Spearman ⍴= 0.88, Figure 3d). Taken together, these findings emphasize that some individuals tend to preferentially attend to a single hemifield when planning. This tendency, importantly, represents an inter-individual difference in how participants allocate their attention across various maze types.”

      (5) Distinction between overt and covert attention:

      The current report at times equates eye movement patterns with the locus of attention. However, attention can be covertly shifted without corresponding gaze changes (see, for example, Pooresmaeili & Roelfsema, 2014).

      We fully agree, and thank the reviewer for prompting further reflection on this distinction. In the online experiments run by Ho and colleagues (i.e., datasets Ho1 and Ho2), participants’ eye movements were not tracked, and therefore, they could not disambiguate whether participants were engaging in covert or overt attention to sample maze obstacles. In our third experiment (i.e., dataset dSC1), we both recorded eye movements and explicitly instructed participants to fixate centrally while viewing the maze. This ensured that participants oriented their attention only covertly during planning (see Figure S13-14).

      We now elaborate on this important distinction in the Results section of the manuscript, page 12:

      “In addition, we monitored participants’ eye movements in dataset dSC 1 to ensure that attention shifts would be covert as opposed to overt—a distinction which could not be determined in the online samples of datasets Ho 1 and 2.”

      On page 28:

      “Importantly, while the visuospatial attention effects observed in the Ho 1 and 2 datasets are likely driven by both covert and overt shifts in attention, the findings presented in experiment 3 (i.e., dSC1 dataset) rule out the contribution of overt shifts in attention through the use of eye tracking (see Figure S13-14)(Carrasco, 2011; Pooresmaeili & Roelfsema, 2014).”

      The implications for interpreting the relationship between eye movement, memory, and attention in this setting are not fully addressed. The potential dynamics of attention along a maze trajectory and their impact on lateralization analysis would benefit from further clarification.

      We thank the reviewer for urging more clarity here. The attentional dynamics we document in our study concern how people perceive / construe the maze itself, rather than how they deploy their attention to guide active navigation. We have now sought to make this distinction clear at a number of points in the paper. The core idea is that attention acts as an early filter to select which obstacles are part of a task construal, which then affects both awareness and memory.

      We have now clarified the focus of our study in the introduction on pages 5-7:

      “Our focus in this study was to examine how participants perceive and represent their environment (the maze stimulus). This is a distinct process to how participants orient their attention during navigation itself, which is not part of our current study. To do so, we harness classical signatures of attentional selection to characterise how visuospatial attention shapes awareness of maze obstacles during planning.” … “Our focus in the present study was to examine attentional effects on participants’ perception of the maze stimulus. We did not quantify how individuals deploy their attention in the phase in which they were navigating through the maze.”

      We did not explicitly test for memory effects in our new experiments, but Ho and colleagues demonstrated that the sVGC model predicted not only awareness reports, but also participants’ memory of obstacles (see Ho et al., 2022). Indeed, task representations computed from memory or awareness reports were strikingly similar in their experiments (Spearman ⍴ = 0.86 between memory accuracy and awareness; ⍴ = 0.86 between confidence in memory and awareness). In relation to eye movements, we refer the reviewer back to our previous response, which details how eye movements were measured and controlled during maze construal.

      Figure 1 legend (b) --> (c)

      We have corrected this typo in the figure caption.

      Reviewer #2 (Public review):

      Summary:

      Castanheira et al. investigate the role of spatial attention for planning during three maze navigation experiments (one new experiment and two existing datasets). Effective planning in complex situations requires the construction of simplified representations of the task at hand. The authors find that these mental representations (as assessed by conscious awareness) of a given stimulus are influenced by (spatially) surrounding stimuli. Individual participants varied in the degree to which attention influenced their task representations, and this attentional effect correlated with the sparsity of representations (as measured by the range of awareness reports across all stimuli). Spatially grouping taskrelevant information on either the left or right side of the maze led to mental representations more similar to optimal representations predicted by the valueguided construal (VGC) model - a normative model describing a theoretical approach to simplifying complex task information. Finally, the authors propose an update to this model, incorporating an attentional spotlight component; the revised descriptive model predicts empirical task representations better than the original (normative) VGC model.

      Strengths:

      The novelty of this study lies in the proposal and investigation of a cognitive mechanism through which a normative model like value-guided construal can enable human planning. After proposing attention as this mechanism, the authors make concrete hypotheses about mismatches between the VGC predictions and real human behavior, which are experimentally validated. Thus, not only does this study describe a possible mechanism for simplification of task information for planning, but the authors also propose a descriptive model, revising VGC to incorporate this attentional component.

      A strength of this paper is the variety of investigative approaches: analysis of existing data, novel experiment, and a computational approach to predict experimental findings from a theoretical model. Analyzing pre-existing datasets increases the size of the participant cohort and strengthens the authors' conclusions. Meanwhile, comparing the predictions of the existing normative model and the authors' own refined model is a clever approach to substantiate their claims. In addition, the authors describe several crucial controls, which are key to the interpretability of their results. In particular, the eye tracking results were critical.

      In summary, this paper constitutes an important step toward a more complete understanding of the human ability to plan.

      We thank the Reviewer for their thoughtful and positive assessment of our findings. We also appreciate the constructive feedback on our methodology, which we believe has substantially improved our manuscript.

      Weaknesses:

      (1) There is a critical conceptual gap in the study and its interpretation, mainly due to the reliance on a self-report metric of awareness (rather than an objective measure of behavioral performance).

      a. Awareness is tested by a 9-point self-report scale. It is currently unclear why awareness of task-irrelevant obstacles in this task would necessarily compromise optimal planning. There is no indication of whether self-reported awareness affects performance (e.g., navigation path distance, time to complete the maze, number of errors). Such behavioral evidence of planning would be more compelling.

      We thank the reviewer for prompting further reflection on the connection between construal and navigation performance. We wish to emphasise that the primary focus of our study was on measuring and modeling participants’ task construals using perceptual awareness judgments, building on the methods developed by Ho and colleagues, rather than on navigation performance itself. However, as the reviewer points out, there is a natural relationship between construal and performance – if you represent the wrong obstacles, plans may be disrupted.

      To explore the relationship between task construals and performance on the navigation task we first regressed out the effects of the sVGC model on participants’ awareness reports and computed the mean squared residuals for each trial. We then used these values to predict participants’ navigation response times on each trial. We observed a significant negative relationship, suggesting that on trials where participants’ representations showed greater deviations from the normative model, they were in fact faster at navigating the mazes. This relationship was surprising, and at odds with the initial idea that adhering to normative VGC aids in task performance. However, we think that this direction of effect may make sense if one considers that a large part of the actual construal (rather than the normative prediction) in our data was in fact driven by effects such as lateralisation which are not accounted for by the sVGC model. If one is faster at harnessing inductive biases such as lateralisation, then one may be faster to complete the maze but also show a greater deviation from the predictions of the original model.

      To further explore these effects, we next focused on the distinction between lateralised and non-lateralised mazes. Here, we reasoned that the initial phase of lateralised attentional selection would lead to lateralised mazes being easier to navigate than nonlateralised ones. We conducted new analyses to determine whether participants navigated lateralized maze stimuli faster and with fewer moves than maze stimuli with non-lateralized model predictions. As detailed in Methods, we excluded trials in which participants significantly deviated from the optimal number of moves (9 or more moves) and took longer than 20 seconds to solve the maze. In line with our interpretation that attention operates as an inductive bias, participants were faster and deviated less from the optimal path on lateralized compared to non-lateralized mazes.

      We now report these new results on navigation performance on pages 20-21:

      “Maze navigation performance

      The previous analyses focused on participants’ task representations during planning. We next sought to explore links between participants’ task representations and maze navigation performance. Participants performed the maze navigation task near-ceiling: they solved 95% of maze stimuli in under 20 seconds, with minimal deviation from the optimal path (i.e., 9 moves or fewer). Notwithstanding this limited variance in task performance, we explored whether participants’ task construals may have impacted their navigation speed. To do so, we first regressed out the effects of the sVGC model from participants’ awareness reports and used the mean squared residuals for each trial to predict response times (see Methods for details). Surprisingly, we observed a negative relationship between mean squared residual variance and response times (β = -0.31, SE = 0.05, 95% CI [-0.41, -0.21], p < 0.001), indicating that participants were faster on trials where the sVGC model explained less variance in their awareness reports. In other words, trials in which participants deviated more from the sVGC model predictions were solved faster. We note that one reason for this may be the strong influence of the lateralisation effect on navigation performance (see paragraph below), which itself is not part of the sVGC model prediction.”

      “We then explored whether participant performance differed between lateralised and nonlateralised mazes. Here, we reasoned that the initial phase of lateralised attentional selection would lead to lateralised mazes being easier to navigate than non-lateralised ones. Consistent with this hypothesis, participants were faster (β = -0.04, SE = 5.91*10<sup>3</sup>, 95% CI [-0.06, -0.03], p< 0.001) and followed the optimal path more closely (β = -0.59, SE = 0.09, 95% CI [-0.78, -0.40], p< 0.001) when maze stimuli were more lateralized.”

      And in the Discussion section, on page 23:

      “Mental representations and task performance

      We observed that participants were faster and deviated less from the optimal path on maze stimuli that were lateralized. This effect is not predicted by the original sVGC model but dovetails with the interpretation that early visuospatial attention operates as an inductive bias to guide the formation of simplified task representations. Surprisingly, we also observed that participants were faster to navigate mazes on trials where their simplified task representation deviated from the sVGC model prediction. We interpret this seemingly contradictory finding in the following way: there are several factors beyond the sVGC model – including, for instance, maze lateralisation – that predict both construal and performance on the maze navigation task. Further work is needed to understand how inductive biases such as lateralisation shape both construal and performance, and the real-world benefits that such strategies might afford for naturalistic stimuli.”

      b. Relatedly, it would have been more convincing to have an objective measure of awareness, for instance, how the presence or absence of a "task-irrelevant" obstacle affects performance (e.g., change navigation path distance or time to complete the maze), or whether participants can accurately recall the location of obstacles.

      We thank the reviewer for prompting further reflection on the validity and robustness of our awareness measures. We emphasise however that our focus is not (primarily) on maze navigation performance, but on task construal, which as noted in our previous response may come apart from navigation performance for a variety of reasons. Our primary goal is to measure participants’ subjective awareness of the maze as a marker of their idiosyncratic (conscious) mental representation on each trial. In doing so, we build on a rich tradition of measuring subjective awareness in consciousness and perception science (for instance, work using the Perceptual Awareness Scale, or detection judgments). In this sense, we think our awareness scale (following Ho et al.) represents a valid and straightforward way of assessing our target psychological construct. However, we also agree with the Reviewer that convergent evidence from other measures is always valuable. In Ho and colleagues’ original paper, they developed a variant of the maze task where participants had to recall the location of obstacles, as well as rate their awareness (Exp 3) and a variant in which participants could hover their mouse over hidden obstacles in the maze to reveal their location – an online metric of attentional deployment (Exp 4). These data afforded us the opportunity to validate the awareness reports against an objective measure of recall, as suggested by the Reviewer. In reanalysing these data, we observed that the obstacle awareness and memory/hover measures were strikingly correlated within two independent samples of participants (Spearman ⍴ = 0.86 between memory accuracy and awareness; ⍴ = 0.86 between confidence in memory and awareness; ⍴ = 0.76 between the probability of hovering over the obstacle and awareness; ⍴ = 0.65 between the duration of the mouse hovering and awareness). These re-analyses are now reported on page 22 of our manuscript, to highlight the convergent validity of the awareness metric:

      “Finally, we examined the convergent validity of participants’ awareness reports by reanalyzing the memory recall data reported in Ho and colleagues’ experiment(Ho et al., 2022). We reasoned that participants should demonstrate similar task representations regardless of the measure used to probe the construal. In line with this prediction, we observed that the obstacle awareness reports and memory/hover measures were strikingly correlated within three independent samples of participants (Spearman ⍴ = 0.86 between memory accuracy and awareness; ⍴ = 0.86 between confidence in memory and awareness; ⍴ = 0.76 between the probability of hovering over the obstacle and awareness; ⍴ = 0.65 between the duration of the mouse hovering and awareness; see Tables S18 and S19).”

      c. Consequently, I'm not sure that we can conclude that the spatial context does impact participants' ability to plan spatial navigation or to "incorporate taskrelevant information into their construal". We know that the spatial context affects subjective (self-reported) awareness, but the authors do not present evidence that spatial context affects behavioral performance.

      Following the line of argument above, we think it’s important to separate out task construal (the simplified representation of the maze, measured by awareness reports), and the impact of this on navigation and other aspects of behaviour. The awareness reports (and other convergent measures) show that task-relevant information (as predicted by the VGC) is incorporated into the construal, a process which is modulated by spatial context. These are the key targets of our modeling. Whether this impacts performance is a distinct question, and one that we now address in our response to point a above.

      d. Another concern that may complicate interpretation is the following: Figure 3c shows improved VGC model predictions (steeper slope) for mazes with greater lateralization. However, there are notable outliers in these plots, where a high lateralization index does not correspond to good model performance. There is currently no discussion/explanation of these cases.

      The Reviewer astutely points out some outliers in our analysis. While on average lateralized maze stimuli are represented more closely to the sVGC model, there are indeed some noticeable outlier mazes. These mazes represent stimuli in which participants tended to lateralize their attention to the ‘wrong hemifield’—e.g., participants were more aware of obstacles in the right hemifield despite sVGC model predicting that obstacles on the left hemifield were task-relevant. We believe this explains the poor sVGC model fits on these trials. We note, however, that on average participants were capable of attending to the correct hemifield without explicit instructions (i.e., 9 out of 12 mazes).

      We have now included a discussion of these outliers in the results section of the paper on page 12:

      “We note that for three maze stimuli whose model predictions were lateralized there was nevertheless a poor fit to the sVGC model (see Figure 2c, right panel). These outliers correspond to maze stimuli where participants, on average, lateralized their attention to the incorrect hemifield (i.e., the opposite hemifield to that predicted by the sVGC model).”

      (2) I noticed an issue with clarity regarding task-relevance. It is currently not fully clear which obstacles are "task irrelevant". Also, the term is used inconsistently, sometimes conflating with "awareness". For example, in the "Attentional spotlight model of task representations" section, the authors state that "taskrelevant information becomes less relevant when surrounded by task-irrelevant information". But they really mean that participants become less aware of those task-relevant obstacles. I assume task-relevance is an objective characteristic related to maze organization, not to a participant's construal. Indeed, the following paragraph provides evidence of model predictions of awareness.

      We apologize for any confusion regarding the terminology of our manuscript. We indeed use the terms task-relevant and task-irrelevant to refer to obstacles that are objectively predicted by the normative sVGC model or the attentional spotlight model to be included in (>0.5) or excluded from (<0.5) task construals, respectively. This designation reflects the predictions from the computational model and does not reflect participants’ reported awareness. We then ran linear hierarchical models to predict participants’ awareness reports from these model predictions. The Reviewer is correct that the task-relevance of obstacles is indeed related to the maze’s organization, and not related to participants’ subjective reports of awareness. We have now clarified this point throughout the manuscript to better emphasize the difference between the model predictions of taskrelevance and participants’ subjective reports.

      On page 17:

      “To achieve this, we computed the predictions of the existing VGC model for each obstacle’s task relevance in a given maze, and averaged these predictions within an attentional spotlight of 3 squares (Figure 4a & S8, see Methods for details). This process yielded novel model predictions, whereby some obstacles which were once predicted as task-irrelevant by the normative sVGC are now predicted as task-relevant by the attentional spotlight model. We depict the effects of this spatial spotlight in Figure 4a: task-irrelevant stimuli (plotted in grey; see middle left obstacle) neighbouring taskrelevant obstacles (plotted in orange) become more task-relevant, whereas taskrelevant information becomes less relevant when surrounded by task-irrelevant information (see bottom right orange obstacle). This deviation in model predictions from the normative sVGC model was used to predict participants’ awareness reports. We hypothesized that this spotlight-VGC model would predict participants’ reports better than the original VGC model, which does not account for spatial attention.”

      (3) The behavioral paradigm has some distinct disadvantages, and the validity of the task is not backed up by behavioral data.

      a. I understand the need for central fixation, but it also makes the task less naturalistic.

      The fixation cross was required on every trial such that participants could maintain central fixation for our eye tracking experiment. While this design is less naturalistic, it allows us to examine the eye movements of participants. Requiring participants to fixate during the ‘planning’ phase of the experiment allowed us to isolate the effects of covert attention from changes in awareness due to overt shifts in attention. In other words, differences in participants’ awareness reports in the 3rd experiment cannot be explained by longer fixation times to specific obstacles.

      b. The task with its top-down grid view does not seem to mimic real human navigation. Though this grid may be similar to mental maps we form for navigation, the sensory stimuli corresponding to possible paths and to spatial context during real-life navigation are very different.

      We agree with the reviewer that while our task is engaging for participants and simple to follow, it does not mimic naturalistic navigation in humans. There is a natural tension in computational / experimental work in cognitive science in wanting to build closely on previous results and paradigms, while ensuring that results can generalise to real-world contexts. Here, our choice of paradigm and measures was closely built on previous papers using this task from Ho and colleagues (2022, 2023). While preparing this response, we learnt that the MIT group had also harnessed this same task to develop a novel dynamic variant of the VGC model (Chen et al., 2026) called the Just in Time model (JIT). The advantage of building on this prior work is that we are able to iteratively refine and expand the VGC approach, and (in our case) bring it into closer contact with work on modeling the deployment of spatial attention in human vision. The top-down aspect of the maze notably facilitated the study of the spatial deployment of attention. We now discuss the novel dynamic variant of the VGC model in our paper on page 27:

      “We close by reflecting on opportunities for further work in this area. First, an important next step is to explore the process by which task representations are formed, and how inductive biases might affect the process of task construal. The sVGC model is a normative model of the optimal task representation. Since it’s construction involves an exhaustive calculation over possible paths, it is not a plausible basis for a model of the psychological process by which participants actually construct task representations. More recently a process model of task construal has been proposed, the Just in Time model (JIT). The hypothesis of the JIT model is that participants’ task representations are built up over time by iteratively simulating possible paths through the maze, affording insight into the construal process (Chen et al., 2026). In future work, it would be of interest to ask whether the attentional effects we observe in our experiments could be meshed with a dynamic JIT account of construal. We speculate that visuospatial attention may operate as an early filter, limiting the space of potential construals based on coarse spatial features of the environment, constraining a dynamic selection of obstacles. Brain imaging techniques with high time resolution, such as M/EEG, may be able to shed further light on how task representations are formed as participants plan.”

      c. Behavioral performance is not reported, so it is unknown whether participants are able to properly complete the task. The task seems pretty difficult to navigate, especially when the obstacles disappear, and in combination with the central fixation.

      Behavioural performance is now reported in response to point 1a above.

      d. There is no discussion of whether/how this navigation task generalizes to other forms of planning.

      We fully agree that an important next step would be to generalise our results on construal to naturalistic forms of planning – for instance, using immersive VR mazes, and or investigating cognitive rather than perceptual construals. We have now added a line to this effect to the Discussion on page 28.

      “An important next step to further our understanding of task representations would be to extend the current paradigm to other forms of planning and more naturalistic tasks, such as navigating immersive virtual reality (VR) environments, planning over cognitive rather than perceptual representations (e.g. planning over an abstract space), or internallyguided planning based on working memory.”

      Reviewer #2 (Recommendations for the authors):

      (1) There are, of course, benefits to simple tasks like the ones described, but it would be interesting to compare the results to a possible experiment in which a top-down grid/map is used for planning, but then task execution is carried out in a simulated environment corresponding to the map. Also, perhaps beyond the scope of the questions addressed in this paper, but I am curious how unexpected obstacles affect representations. For instance, if participants plan based on a topdown map and then begin "real" navigation but encounter an unexpected obstacle that was not indicated on the map, does this modulate representations/awareness of future obstacles (near vs. far)?

      We fully agree that all of these lines of investigation would be super interesting to pursue in future studies, and we have added a line to the discussion to that effect on page 28:

      “An important next step to further our understanding of task representations would be to extend the current paradigm to other forms of planning and more naturalistic tasks, such as navigating immersive virtual reality (VR) environments, planning over cognitive rather than perceptual representations (e.g.. planning over an abstract space), or internallyguided planning based on working memory.”

      (2) Regarding self-reported awareness as a metric, an additional experiment could ask participants to recreate the maze (identify locations of obstacles after they disappear). This would be a more objective measure of awareness.

      Yes indeed, and as described above, this was a metric used by Ho and colleagues in their previous experiment. As we describe in more detail above, the task representations obtained via memory or awareness reports demonstrated striking similarity (⍴ = 0.86).

      (3) What is meant by "all possible orientations of the maze" in this Methods sentence: "For dataset dSC 1, participants solved each of these 24 mazes four times (i.e., all possible orientations of the maze)"?

      We thank the Reviewer for prompting more clarity here. We vertically and horizontally reversed mazes (i.e., left-right flipped) such that participants could not predict the location of the goal or start location. In this way, each maze stimulus had four potential orientations. This resulted in 96 trials of 24 unique mazes. We have clarified this point in the Methods section on page 30:

      Maze stimuli were vertically and horizontally reversed (i.e., left-right flipped) such that participants could not predict the location of the start or goal location. This resulted in four potential orientations of each maze across all 24 mazes, 96 trials in total.

      (4) For lateralization, it was unclear until reading the Methods that the lateralization index was calculated using the VGC-predicted level of taskrelevance. From the main text and Figure 2, I assumed you were just counting the number of task-relevant obstacles on each side, rather than also quantifying relevance. I understood after reading the Methods, but this could be clarified further.

      We agree with the Reviewer that this was not evident from the text. We have now updated the Results section of the manuscript to clarify this point on page 11:

      “To test this hypothesis, we derived a measure of task-relevant lateralization inspired by the attention literature (Ghafari et al., 2024; Keefe & Störmer, 2021; Vollebregt et al., 2015) (Figure 2a). Specifically, we separated maze stimuli across the vertical meridian and computed the ratio of task-relevant information presented on the left versus right side derived from the sVGC model. For example, the maze shown in Figure 2a has twice the amount of task-relevant information presented in the left hemifield than in the right (lat. Index= 1/3). A lateralization index of 0.0 indicates that both hemifields contain equal amounts of task-relevant information (i.e., non-lateralized). The lateralization index was computed using the continuous VGC predictions for each obstacle (see Methods).”

      (5) The explanation in the Methods of how the width of the attentional spotlight was chosen references Figure 1b and Supplementary Figure S2, but it seems that Supplementary Figure S8 explains this more in the caption. Also, I don't see how Figure S2 supports this.

      We apologize for this typo. The explanation of how we selected the width of the attentional spotlight should indeed reference supplemental Figure 15 (previously Figure S8). We have now corrected this and elaborated on this choice in the Methods section on page 35:

      “We fixed the ‘width’ of the attentional spotlight to a distance of 3 squares based on the observation that the two neighbouring obstacles positively predicted the awareness of a probe. We observed that the mean and median distance between neighbouring obstacles of the 2nd rank (i.e., second closest) was 3 squares away for all mazes (Figure S15). We therefore opted to fix the value of the attention spotlight to 3 squares based on these observations. Future work utilizing this model should consider the statistics of their maze stimuli when deciding on the ‘width’ of the attentional spotlight.”

      (6) The attentional spotlight width was assumed to be 3 squares, based on the linear regression predictions of the effect of neighboring obstacles on stimulus awareness. Given the individual differences across participants, it would be interesting to choose a different attentional spotlight size for each participant. Would a participant-specific attentional spotlight width improve the predictions of the spotlight-VGC model?

      The Reviewer highlights a very interesting question: do individuals vary in terms of their attentional spotlight? To test this hypothesis, we first estimated the size of the attentional spotlight for each individual based on lateralized maze stimuli, and then used this to generate personalized attentional spotlight model predictions for each subject based on these values (Figure S11). We restricted this analysis to the dSC1 dataset, where we had substantially more trials (96 in total).

      In brief, we observed that indeed the personalized spotlight model fit participants’ awareness reports better than both a normative sVGC model and a group-level attentional spotlight model. We interpret these findings with some caution as i) a subset of individuals had flat attentional slopes and therefore were excluded from these analyses, and ii) we believe we require additional trials to ensure a robust model fit at the individual level. While our results are encouraging, we hope future investigations into inter-individual differences will extend these findings.

      We have included these additional analyses in the main text.

      On page 18:

      “To further explore inter-individual differences in task construal, we tested whether adjusting the attentional spotlight width to each participant’s awareness reports improved the predictions of the attentional spotlight model. To do so, we first determined the width attentional spotlight of each individual in the dSC1 dataset based on lateralized maze stimuli. We then generated person-specific attentional spotlight model predictions for the non-lateralized maze stimuli to avoid overfitting the data (Figure S11). We note that 7 participants had either flat attentional slopes or negative beta coefficients, which prevented the selection of an appropriate attentional spotlight width (see Methods for details). We observed a significant improvement in model fit for the person-specific attentional spotlight model relative to both the group-level attentional spotlight model (ΔBIC= -1487.39) and the normative sVGC model (ΔBIC= -1655.29). While the limited trial numbers per participant in our current dataset warrants caution in interpreting these findings, these findings do encourage further research on inter-individual differences in attentional deployment during planning.”

      On pages 23-24:

      “Inter-individual differences in attention

      We also observed considerable inter-individual differences in attentional effects across participants (Figure 1c). While some participants were strongly influenced by the spatial context of neighbouring stimuli, others showed more limited evidence for an attentional effect (Figure 1b). Inter-individual differences in attention predicted the sparsity of participants’ simplified representations: participants with larger attention effects exhibited sparser representations. Moreover, these inter-individual differences in effects of spatial proximity could be incorporated into the attentional spotlight model by varying the width of the spotlight, resulting in better model predictions.”

      “Beyond these spatial proximity effects, we also observed that participants varied in their tendency to lateralize their attention to a single hemifield (Figure 3). This tendency was observed across all three datasets, including on maze stimuli whose value-guided model predictions were not lateralized. This suggests that although a strategy of allocating attention is sub-optimal for these maze stimuli, some individuals preferentially attend to a single hemifield in a heuristic-like fashion. This tendency to attend to a single hemifield was a robust inter-individual difference across maze stimuli (Figure 3d), and dovetails with individual-level variation in spatial proximity effects. Taken together, these findings offer novel insights into how people vary in the ways they allocate spatial attention to solve complex problems. Future research could explore how these individual differences constrain performance on other tasks that require planning and search in highdimensional spaces.”

      On page 17 of the Supplemental Materials:

      (7) The supplementary text about lateralization effects, above Supplementary Table S8, references Table S6, but it is Table S6 does not seem to display lateralization results.

      We thank the Reviewer for pointing out this typo: we now refer to the correct supplementary table (S9).

      (8) Why does it matter that "the maze stimuli were not designed to test horizontalmeridian lateralization effects"? What is the effect on power? Is it because there is not a good enough range in lateralization indices? It would be good to clarify, or just remove that explanation, since the cortical retinotopy explanation seems more convincing.

      We did not specifically design the maze stimuli such that there is an equal number of obstacles above and below the horizontal meridian. As such, the lateralization index derived along the horizontal meridian does not control for the number of obstacles in each hemifield, which may influence participants’ awareness reports. In contrast, we designed maze stimuli such that this would not be a concern for the vertical meridian. We have clarified this point in the discussion on page 27.

      “Third, while we observed clear lateralization effects along the vertical meridian (i.e., left vs right hemifield), effects along the horizontal meridian were less clear (i.e., above vs below; see Table S15-16). One potential explanation of this asymmetry is the retinotopic organization of the cortex, in which spatially adjacent stimuli can be retinotopically distant if presented on the opposite side of the vertical (but not horizontal) meridian, facilitating distractor inhibition. Importantly, while the visuospatial attention effects observed in the Ho 1 and 2 datasets are likely driven by both covert and overt shifts in attention, the findings presented in experiment 3 (i.e., dSC1 dataset) rule out the contribution of overt shifts in attention through the use of eye tracking (see Figure S13-14)(Carrasco, 2011; Pooresmaeili & Roelfsema, 2014).”

      (9) For Figure 2c, it would be helpful to directly state what each dot and line mean.

      We updated the caption of Figure 2c to clarify what we are plotting: each point represents an obstacle, and each line the linear fit for a maze stimulus.

      “Each point represents an obstacle in a maze, and each line represents the model fit for that specific maze stimulus.”

      (10) Figures and wording imply there is only a single probe obstacle per trial, but methods and model imply that participants are asked to report awareness for every obstacle. This should be clarified.

      We apologize for any confusion regarding the methodology of our study. The Reviewer is correct that participants reported their awareness of every obstacle presented on a given trial. We have clarified this in the Results section of the manuscript on page 7:

      “Note, participants reported their awareness of every obstacle presented on a given trial.”

      We have also updated the caption of Figure 1 to clarify this point:

      “Once participants finished navigating the maze, they were asked to report their awareness of every obstacle presented on a given trial in a random order.”

      (11) What is the reason for the exclusion of participants (33 for experiment 1 and 26 for experiment 2)?

      Participants were excluded from the Ho et al. datasets 1 and 2 based on their preregistered exclusion criteria, as detailed in the Methods section of their paper. In short, trials were excluded if participants took longer than 20 seconds to complete the trial, or if they spent longer than 5 seconds in the initial state. Participants were excluded if less than 80% of trials remained after reaction time exclusions or if they failed 2 out of 3 comprehension checks. We have elaborated on this point in the Methods section on page 31.

      “Participants were excluded from analyses based on pre-registered exclusion criteria as detailed in (Ho et al., 2022). In short, participants were excluded if 20% or more of their trials were removed based on reaction times, or if they failed 2 out of 3 comprehension checks.”

      (12) The supplemental figures are not referenced in order, and some are not referenced at all; this should be fixed.

      We thank the Reviewer for pointing this out and have reorganized our Supplementary materials accordingly.

      Reviewer #3 (Public review):

      Summary:

      The authors build on a recent computational model of planning, the "value-guided construal" framework by Ho et al. (2022), which proposes that people plan by constructing simple models of a task, such as by attending to a subset of obstacles in a maze. They analyze both published experimental data and new experimental data from a task in which participants report attention to objects in mazes. The authors find that attention to objects is affected by spatial proximity to other objects (i.e., attentional overspill) as well as whether relevant objects are lateralized to the same hemifield. To account for these results, the authors propose a "spotlight-VGC" model, in which, after calculating attention scores based on the original VGC model, attention to objects is enhanced based on distance. They find that this model better explains participant responses when objects are lateralized to different hemifields. These results demonstrate complex interactions between filtering of task-relevant information and more classical signatures of attentional selection.

      Strengths:

      (1) The paper builds on existing modeling work in a novel manner and integrates classic results on attention into the computational framework.

      (2) The authors report new and extensive analyses of existing data that shed light on additional sources of systematic variability in responses related to attentional spillover effects

      (3) They collect new data using new stimuli in the original paradigm that directly test predictions related to the lateralization of task-relevant information, including eye tracking data that allows them to control for possible confounds.

      (4) The extended model (spotlight-VGC) provides a formal account of these new results.

      We thank the Reviewer for their positive assessment of our manuscript and their insightful comments, which has improved the clarity of our findings.

      Weaknesses:

      (1) The spotlight-VGC model has a free parameter - the "width" of the attentional spotlight. This seems to have been fixed to be 3 squares. It would be good if the authors could describe a more principled procedure for selecting the width so that others can use the model in other contexts.

      Our choice for this parameter was informed by the spatial effects reported in Figure 1b. We observed that the two closest neighbouring obstacles to a probe had similar awareness (i.e., positive beta weights). We therefore compute the mean and median distances between obstacle pairs that were the second closest obstacle to a probe. This distance was 3 squares away, as depicted in Figure S15. We fixed the width of the attentional spotlight across all studies based on this observation. We agree that future research utilizing this model may need to tune this hyperparameter depending on the mean distance between a probe and its neighbours.

      We have clarified this point in the methods section on page 35:

      “We fixed the ‘width’ of the attentional spotlight to a distance of 3 squares based on the observation that the two neighbouring obstacles positively predicted the awareness of a probe. We observed that the mean and median distance between neighbouring obstacles of the 2nd rank (i.e., second closest) was 3 squares away for all mazes (Figure S15). We therefore opted to fix the value of the attention spotlight to 3 squares based on these observations. Future work utilizing this model should consider the statistics of their maze stimuli when deciding on the ‘width’ of the attentional spotlight.”

      Following the suggestion of Reviewer 2 point 6, we now also explored inter-individual differences in this parameter. To do so, we first used the lateralized mazes in the dSC1 dataset to determine the optimal width of the attentional spotlight for each individual.

      Then, we used this spotlight to derive model predictions for each person. We observed that these personalized attentional spotlight model predictions fit participants’ awareness reports on non-lateralized mazes better than the fixed-width spotlight model. We believe this preliminary result suggests the importance of modelling inter-individual differences in attentional deployment during planning. We report these effects on page 17.

      (2) Have the authors considered other ways in which factors such as attentional spillover and lateralization could be incorporated into the model? The spotlightVGC model, as presented, involves first computing VGC predictions and only afterwards computing spillover. This seems psychologically implausible, since it supposes that the "optimal" representation is first formed and then it gets corrupted. Is there a way to integrate these biases directly into the VGC framework, perhaps as a prior on construals? The authors gesture towards this when they talk about "inductive biases", but this is not formalized.

      We thank the reviewer for bringing up this very important point. We think that a full computational treatment of the inductive bias would be a distinct project, but now seek to expand our discussion on the mechanisms by which representations could be formed. In this context, we specifically highlight novel computational work from the MIT group that was published as a preprint in the time since we submitted our paper, and which proposes a new process account of construal, the “Just in Time” (JIT) model. We also elaborate on a possible mechanism by which visuospatial attention may aid the dynamics of the construal process. In short, we agree with the reviewer that spatial attention may bias individuals to search over a subset of potential representations based on low-level spatial characteristics of the obstacles (e.g., their spatial spread in the visual field), prior to (or in concert with) a dynamic JIT-like selection process. We now elaborate on these possibilities on pages 27-28:

      “We close by reflecting on opportunities for further work in this area. First, an important next step is to explore the process by which task representations are formed, and how inductive biases might affect the process of task construal. The sVGC model is a normative model of the optimal task representation. Since it’s construction involves an exhaustive calculation over possible paths, it is not a plausible basis for a model of the psychological process by which participants actually construct task representations. More recently a process model of task construal has been proposed, the Just in Time model (JIT). The hypothesis of the JIT model is that participants’ task representations are built up over time by iteratively simulating possible paths through the maze, affording insight into the construal process (Chen et al., 2026). In future work, it would be of interest to ask whether the attentional effects we observe in our experiments could be meshed with a dynamic JIT account of construal. We speculate that visuospatial attention may operate as an early filter, limiting the space of potential construals based on coarse spatial features of the environment, constraining a dynamic selection of obstacles. Brain imaging techniques with high time resolution, such as M/EEG, may be able to shed further light on how task representations are formed as participants plan.”

      […]

      “Fourth, it will also be necessary to elaborate on how bottom-up and top-down aspects of attentional selection are combined to guide complex task representations and plans. Foundational questions remain unanswered, for instance: can multiple spatial locations be preferentially selected at once, i.e. are there multiple spotlights (Awh & Pashler, 2000; McMains & Somers, 2004; Pylyshyn & Storm, 1988; Shaw & Shaw, 1977)? There is also discourse on how spatial attention may move from one location to another: are the intervening visual regions between attended locations similarly selected (Dubois et al., 2009; Kr & Np, 1999; McMains & Somers, 2004, 2005)? Our findings tentatively suggest that individuals are able to attend to disparate spatial regions to form sparse task representations, yet there is substantial variability in how individuals orient their attention during the task. The present paradigm and computational modelling, in conjunction with carefully designed stimuli, may help resolve these outstanding questions.”

      (3) Can the authors rule out that the lateralization effects are the result of memory biases since the main measure used is a self-report of attention?

      We thank the reviewer for bringing up this important point. In our experiments, we sought to measure participants’ subjective awareness of the maze stimuli as a readout of their conscious task representation on each trial. This approach marries an extensive literature on measures of perceptual awareness in consciousness science (e.g., using the Perceptual Awareness Scale) with computational models of planning. Participants’ memory of (their awareness of) the obstacles is inherent to this approach, but just as with similar approaches in consciousness science (e.g. measures of iconic memory in the Sperling paradigm), we think it provides a reasonably “online” measure of awareness. It’s important of course to ensure that results obtained with awareness reports are not idiosyncratic, and generalise to other approaches to quantifying task representations.

      To further bolster the convergent validity of our awareness measure, we reanalyzed the data from Ho and colleagues. In their original paper, they developed a variant of the maze-navigation task where participants were asked to recall the location of obstacles as well as report their awareness (Exp 3) and a third variant of the task where participants could hover their cursors over hidden obstacles to reveal their locations (Exp 4). These data allowed us to validate the awareness reports against objective measures of recall and mouse-tracking data. We observed that the subjective awareness reports of participants were strikingly correlated with recall/hover measures across two independent samples of participants (Spearman ⍴ = 0.86 between memory accuracy and awareness; ⍴ = 0.86 between confidence in memory and awareness; ⍴ = 0.76 between the probability of hovering over the obstacle and awareness; ⍴ = 0.65 between the duration of the mouse hovering and awareness). We believe these findings validate participants’ awareness reports. These findings are now reported on page 22 of the manuscript.

      “Finally, we examined the convergent validity of participants’ awareness reports by reanalyzing the memory recall data reported in Ho and colleagues’ experiment (Ho et al., 2022). We reasoned that participants should demonstrate similar task representations regardless of the measure used to probe the construal. In line with this prediction, we observed that the obstacle awareness reports and memory/hover measures were strikingly correlated within three independent samples of participants (Spearman ⍴ = 0.86 between memory accuracy and awareness; ⍴ = 0.86 between confidence in memory and awareness; ⍴ = 0.76 between the probability of hovering over the obstacle and awareness; ⍴ = 0.65 between the duration of the mouse hovering and awareness; see Tables S18 and S19).”

    5. eLife Assessment

      This important study utilizes behavioral data and computational modeling to show that spatial properties of visual attention affect human planning. The methodology and statistical analyses are solid, though the way attention is conceptualized and modeled could be refined. The findings of this study will interest cognitive scientists studying attention, perception, and decision-making.

    6. Reviewer #1 (Public review):

      Summary: This study investigated how visuospatial attention influences the way people build simplified mental representations to support planning and decision-making. Using computational modeling and virtual maze navigation, the authors examined whether spatial proximity and the spatial arrangement of obstacles determine which elements are included in participants' internal models of a task. The study developed and tested an extension of the value-guided construal (VGC) model that incorporates features of spatial attention for selecting simpler task mental representation.

      Strengths:

      (1) Original Perspective: The study introduces an explicit attentional component to established models of planning, offering an approach that bridges perception, attention, and decision-making.

      (2) Methodological Approach: The combination of computational modeling, behavioral data, and eye-tracking provides converging measures to assess the relationship between attention and planning representations.

      (3) Cross-validated data: The study relies on the analysis of three separate datasets, two already published and an additional novel one. This allows for cross-validation of the findings and enhances the robustness of the evidence.

      (4) Focus on Individual Differences: Reports of how individual variability in attentional "spillover" correlates with the sparsity of task representations and spatial proximity add depth to the analysis.

      Weaknesses:

      (1) Clarity of the VGC model and behavioral task: The exposition of the VGC model lacks sufficient detail for non-expert readers. It is not clear how this model infers which maze obstacles are relevant or irrelevant for planning, nor how the maze tasks specifically operationalize "planning" versus other cognitive processes.

      The method for classifying obstacles as relevant or irrelevant to the task and connecting metacognitive awareness (i.e., participants' reports of noticing obstacles) to attentional capture is not well justified. The rationale for why awareness serves as a valid attention proxy, as opposed to behavioral or neurophysiological markers, should be clearer.

      (2) Attention framework: The account of attention is largely limited to the "spotlight" model. When solving a maze, participants trace the correct trail, following it mentally with their overt or covert attention. In this perspective, relevant concepts are also rooted in attention literature pertaining to object-based attention using tasks like curve tracing (e.g., Pooresmaeili & Roelfsema, 2014) and to mental maze solving (e.g., Wong & Scholl, 2024), which may be highly relevant and add nuance to the current work. This view of attention may be more pertinent to the task than models of simultaneously tracking multiple objects cited here. Prior work (notably from the Roelfsema group) indicates that attentional engagement in curve-tracing tasks may be a continuous, bottom-up process that progressively spreads along a trajectory, in time and space, rather than a "spotlight" that simply travels along the path. The spread of attention depends on the spatial proximity to distractors - a point that could also be pertinent to the findings here.

      Moreover, the tracing of a "solution" trail in a maze may be spontaneous and not only a top-down voluntary operation (Wong & Scholl, 2024), a finding that requires a more careful framing of the link to conscious perception discussed in the manuscript.

      Conceptualizing attention as a spatial spotlight may therefore oversimplify its role in navigation and planning. Perhaps the observed attentional modulation reflects a perceptual stage of building the trail in the maze rather than a filter for a later representation for more efficient decision making and planning. A fuller discussion of whether the current model and data can distinguish between these frameworks would benefit readers.

      (3) Lateralization of attention: The analysis considers whether relevant information is distributed bilaterally or unilaterally across the visual display, but does not sufficiently address evidence for attentional asymmetries across the left and right visual fields due to hemispheric specialization (e.g., Bartolomeo & Seidel Malkinson, 2019). Whether effects differ for left versus right hemifield arrangements is not made explicit in the presented findings.

      (4) Individual differences: Individual differences in attentional modulation are a strength of the work, but similar analyses exploring individual variation in lateralization effects could provide further insight, and the lack of such analyses may mask important effects.

      (5) Distinction between overt and covert attention: The current report at times equates eye movement patterns with the locus of attention. However, attention can be covertly shifted without corresponding gaze changes (see, for example, Pooresmaeili & Roelfsema, 2014).

      The implications for interpreting the relationship between eye movement, memory, and attention in this setting are not fully addressed. The potential dynamics of attention along a maze trajectory and their impact on lateralization analysis would benefit from further clarification.

      Appraisal of Aims and Results:

      The study sets out to determine how spatial attention shapes the construction of task representations in planning contexts. The authors provide evidence that spatial proximity and arrangement influence which environmental features are incorporated into internal models used for navigation, and that accounting for these effects improves model predictions. There is clear documentation of individual variation, with some participants showing greater attentional spillover and more sparse awareness profiles.

      However, some conceptual and methodological aspects would be clearer with greater engagement with the broader literature on attention dynamics, a more explicit justification of operational choices, and more targeted lateralization analyses.

    7. Reviewer #2 (Public review):

      Summary:

      Castanheira et al. investigate the role of spatial attention for planning during three maze navigation experiments (one new experiment and two existing datasets). Effective planning in complex situations requires the construction of simplified representations of the task at hand. The authors find that these mental representations (as assessed by conscious awareness) of a given stimulus are influenced by (spatially) surrounding stimuli. Individual participants varied in the degree to which attention influenced their task representations, and this attentional effect correlated with the sparsity of representations (as measured by the range of awareness reports across all stimuli). Spatially grouping task-relevant information on either the left or right side of the maze led to mental representations more similar to optimal representations predicted by the value-guided construal (VGC) model - a normative model describing a theoretical approach to simplifying complex task information. Finally, the authors propose an update to this model, incorporating an attentional spotlight component; the revised descriptive model predicts empirical task representations better than the original (normative) VGC model.

      Strengths:

      The novelty of this study lies in the proposal and investigation of a cognitive mechanism through which a normative model like value-guided construal can enable human planning. After proposing attention as this mechanism, the authors make concrete hypotheses about mismatches between the VGC predictions and real human behavior, which are experimentally validated. Thus, not only does this study describe a possible mechanism for simplification of task information for planning, but the authors also propose a descriptive model, revising VGC to incorporate this attentional component.

      A strength of this paper is the variety of investigative approaches: analysis of existing data, novel experiment, and a computational approach to predict experimental findings from a theoretical model. Analyzing pre-existing datasets increases the size of the participant cohort and strengthens the authors' conclusions. Meanwhile, comparing the predictions of the existing normative model and the authors' own refined model is a clever approach to substantiate their claims. In addition, the authors describe several crucial controls, which are key to the interpretability of their results. In particular, the eye tracking results were critical.

      In summary, this paper constitutes an important step toward a more complete understanding of the human ability to plan.

      Weaknesses:

      (1) There is a critical conceptual gap in the study and its interpretation, mainly due to the reliance on a self-report metric of awareness (rather than an objective measure of behavioral performance).

      a. Awareness is tested by a 9-point self-report scale. It is currently unclear why awareness of task-irrelevant obstacles in this task would necessarily compromise optimal planning. There is no indication of whether self-reported awareness affects performance (e.g., navigation path distance, time to complete the maze, number of errors). Such behavioral evidence of planning would be more compelling.

      b. Relatedly, it would have been more convincing to have an objective measure of awareness, for instance, how the presence or absence of a "task-irrelevant" obstacle affects performance (e.g., change navigation path distance or time to complete the maze), or whether participants can accurately recall the location of obstacles.

      c. Consequently, I'm not sure that we can conclude that the spatial context does impact participants' ability to plan spatial navigation or to "incorporate task-relevant information into their construal". We know that the spatial context affects subjective (self-reported) awareness, but the authors do not present evidence that spatial context affects behavioral performance.

      d. Another concern that may complicate interpretation is the following: Figure 3c shows improved VGC model predictions (steeper slope) for mazes with greater lateralization. However, there are notable outliers in these plots, where a high lateralization index does not correspond to good model performance. There is currently no discussion/explanation of these cases.

      (2) I noticed an issue with clarity regarding task-relevance. It is currently not fully clear which obstacles are "task irrelevant". Also, the term is used inconsistently, sometimes conflating with "awareness". For example, in the "Attentional spotlight model of task representations" section, the authors state that "task-relevant information becomes less relevant when surrounded by task-irrelevant information". But they really mean that participants become less aware of those task-relevant obstacles. I assume task-relevance is an objective characteristic related to maze organization, not to a participant's construal. Indeed, the following paragraph provides evidence of model predictions of awareness.

      (3) The behavioral paradigm has some distinct disadvantages, and the validity of the task is not backed up by behavioral data.

      a. I understand the need for central fixation, but it also makes the task less naturalistic.

      b. The task with its top-down grid view does not seem to mimic real human navigation. Though this grid may be similar to mental maps we form for navigation, the sensory stimuli corresponding to possible paths and to spatial context during real-life navigation are very different.

      c. Behavioral performance is not reported, so it is unknown whether participants are able to properly complete the task. The task seems pretty difficult to navigate, especially when the obstacles disappear, and in combination with the central fixation.

      d. There is no discussion of whether/how this navigation task generalizes to other forms of planning.

    8. Reviewer #3 (Public review):

      Summary:

      The authors build on a recent computational model of planning, the "value-guided construal" framework by Ho et al. (2022), which proposes that people plan by constructing simple models of a task, such as by attending to a subset of obstacles in a maze. They analyze both published experimental data and new experimental data from a task in which participants report attention to objects in mazes. The authors find that attention to objects is affected by spatial proximity to other objects (i.e., attentional overspill) as well as whether relevant objects are lateralized to the same hemifield. To account for these results, the authors propose a "spotlight-VGC" model, in which, after calculating attention scores based on the original VGC model, attention to objects is enhanced based on distance. They find that this model better explains participant responses when objects are lateralized to different hemifields. These results demonstrate complex interactions between filtering of task-relevant information and more classical signatures of attentional selection.

      Strengths:

      (1) The paper builds on existing modeling work in a novel manner and integrates classic results on attention into the computational framework.

      (2) The authors report new and extensive analyses of existing data that shed light on additional sources of systematic variability in responses related to attentional spillover effects

      (3) They collect new data using new stimuli in the original paradigm that directly test predictions related to the lateralization of task-relevant information, including eye tracking data that allows them to control for possible confounds.

      (4) The extended model (spotlight-VGC) provides a formal account of these new results.

      Weaknesses:

      (1) The spotlight-VGC model has a free parameter - the "width" of the attentional spotlight. This seems to have been fixed to be 3 squares. It would be good if the authors could describe a more principled procedure for selecting the width so that others can use the model in other contexts.

      (2) Have the authors considered other ways in which factors such as attentional spillover and lateralization could be incorporated into the model? The spotlight-VGC model, as presented, involves first computing VGC predictions and only afterwards computing spillover. This seems psychologically implausible, since it supposes that the "optimal" representation is first formed and then it gets corrupted. Is there a way to integrate these biases directly into the VGC framework, perhaps as a prior on construals? The authors gesture towards this when they talk about "inductive biases", but this is not formalized.

      (3) Can the authors rule out that the lateralization effects are the result of memory biases since the main measure used is a self-report of attention?

    1. eLife Assessment

      This manuscript presents a useful mean-field model for a network of Hodgkin-Huxley neurons retaining the equations for ion exchange between the intracellular and extracellular space. The mean-field model derived in this work relies on approximations and heuristic arguments that, on the one hand, allow a closed-form derivation of the mean-field equations, but also raise questions about their justifications and the degree to which the results agree with experiments as well as direct numerical simulations. While the revised manuscript is much improved, reviewers continue to question the methodology for reducing model dimensionality and therefore the evidence for the utility of this approach remains incomplete at present.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript the authors derive a mean-field model for a network of Hodgkin-Huxley neurons retaining the equations for ion exchange between the intracellular and extracellular space.<br /> The mean-field model derived in this work relies on approximations and heuristic arguments that, on the one hand, allow a closed-form derivation of the mean-field equations, and on the other hand restrict its validity to a limited regime of activity corresponding to quasi-synchronous neuronal populations. Therefore, rather than an exact mean-field representation, the model provides a description of a mesoscopic population of connected neurons driven by ion exchange dynamics.

      Strengths:

      The idea of deriving a mean-field model which relates the slow-timescale biophysical mechanism of ion exchange and transportation in the brain to the fast-timescale electrical activities of large neuronal ensembles.

      Weaknesses:

      The idea underlying this work is not completely implemented in practice.

      The derived mean field model do not show a one-to-one correspondence with the neural network simulations, except in strongly synchronous regimes. The agreement with the in vitro experiment is hardly evident, both for the mean-field model and for the network model. The assumptions made to derive the closed-form equations of the mean field model have not been justified by any biological reason, they just allow for the mathematical derivation. The final form of the mean-field equations do not clarify whether or not microscopic variables are used together with macroscopic variables in an inconsistent mixture.

      Comments on revisions:

      The main weaknesses I listed in the first report are still present, since the authors did not answer my questions on a solid basis. I report the list for completeness:

      (1) It seems that the reduction methodology that is employed is not the most suitable one for the single-neuron model they are considering.<br /> (2) The formulation of the mean-field derivation is unnecessarily complicated. It could be heavily simplified by following previously published approaches to derive biologically realistic neural masses.<br /> (3) The model seems to work only for highly synchronized situations and not for the standard asynchronous evolution usually observed in neural circuits.

      Therefore, my statement remains unchanged.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aiming in developing a neural mass model characterized by few collective variables mimicking the dynamics of a network of Hodgkin - Huxley neurons encompassing ion-exchange mechanisms. They describe in details the derivation of the mean-field model , then they compare experimental results obtained for the hippocampus of a mice with the neural network simulations and the mean-field results. Furthermore, they report a bifurcation analysis of the developed model and simulation of a small network containing various coupled neural masses, somehow moving towards the simulation of an entire connectome.

      Strengths:

      The author attempts to develop a mean-field model for a globally coupled network of heterogeneous Hodgkin-Huxley neurons with explicit ion exchange mechanism between the cell interior and exterior.

      Weaknesses:

      (1) They do not employ the reduction methodology more suited for the single neuron model they consider.<br /> (2) Their derivation of the neural mass model is based on several assumptions, and not all well justified.<br /> (3) Their formulation of the mean-field derivation is unnecessary complicated, it can be strongly simplified by following previously published approaches to derive biologically realistic neural masses.<br /> (4) Their model seems to work only for highly synchronized situations and not for the standard asynchronous evolution usually observed in neural circuits.

      General Statements:

      The authors honestly declared the many limitations of their approach, once assumed this the results of the mean-field are somehow inconsistent with the neural network simulations as expected.

      The authors suggest to employ this model for the simulations on the whole connectome to follow seizure propagation, however I believe that a simpler model, as the Epileptor, remains superior in this respect to this model. That indeed includes biophysical parameters but their correspondence with the ones employed in the network dynamics remain elusive, due to the many assumptions required to derive this mean field model. Furthermore it is more complicated than the Epileptor, I do not think that the present model will be largely employed by the community.

      Comments on revisions:

      The authors have corrected mistakes present in the manuscript and put a correct list of references.

      However, they refuse

      (1) To simplify the formulation of the model, the model contains unnecessary complications, as I have clearly written in my report, the authors agree, but they do not want to change the formulation;

      (2) To derive the mean field model in a simpler way, as possible, and as I asked many times in my Referee report, this would help the readers to understand the important aspect of the derivation, without not needed and confusing complicated formulations;

      (3) To compare direct simulations of the network with neural mass results in sub-section "Bifurcation analysis: emergent network states and multistability" to show bistability, as I asked.

      As a matter of fact the performed modifications do not solve my previous doubts on the validity of the results reported in the manuscript.

      Therefore, my previous assessments remain valid.

    4. Author response:

      The following is the authors’ response to the current reviews.

      Reviewer #1 (Public review)

      Summary:

      In this manuscript the authors derive a mean-field model for a network of Hodgkin-Huxley neurons retaining the equations for ion exchange between the intracellular and extracellular space.

      The mean-field model derived in this work relies on approximations and heuristic arguments that, on the one hand, allow a closed-form derivation of the mean-field equations, and on the other hand restrict its validity to a limited regime of activity corresponding to quasi-synchronous neuronal populations. Therefore, rather than an exact mean-field representation, the model provides a description of a mesoscopic population of connected neurons driven by ion exchange dynamics.

      Strengths:

      The idea of deriving a mean-field model which relates the slow-timescale biophysical mechanism of ion exchange and transportation in the brain to the fast-timescale electrical activities of large neuronal ensembles.

      Weaknesses:

      The idea underlying this work is not completely implemented in practice.

      The derived mean field model do not show a one-to-one correspondence with the neural network simulations, except in strongly synchronous regimes. The agreement with the in vitro experiment is hardly evident, both for the mean-field model and for the network model. The assumptions made to derive the closed-form equations of the mean field model have not been justified by any biological reason, they just allow for the mathematical derivation. The final form of the mean-field equations do not clarify whether or not microscopic variables are used together with macroscopic variables in an inconsistent mixture.

      Comments on revisions:

      The main weaknesses I listed in the first report are still present, since the authors did not answer my questions on a solid basis. I report the list for completeness:

      (1) It seems that the reduction methodology that is employed is not the most suitable one for the single-neuron model they are considering.

      (2) The formulation of the mean-field derivation is unnecessarily complicated. It could be heavily simplified by following previously published approaches to derive biologically realistic neural masses.

      (3) The model seems to work only for highly synchronized situations and not for the standard asynchronous evolution usually observed in neural circuits.

      Therefore, my statement remains unchanged.

      Reviewer #2 (Public review)

      Summary:

      The authors aiming in developing a neural mass model characterized by few collective variables mimicking the dynamics of a network of Hodgkin - Huxley neurons encompassing ion-exchange mechanisms. They describe in details the derivation of the mean-field model , then they compare experimental results obtained for the hippocampus of a mice with the neural network simulations and the mean-field results. Furthermore, they report a bifurcation analysis of the developed model and simulation of a small network containing various coupled neural masses, somehow moving towards the simulation of an entire connectome.

      Strengths:

      The author attempts to develop a mean-field model for a globally coupled network of heterogeneous Hodgkin-Huxley neurons with explicit ion exchange mechanism between the cell interior and exterior.

      Weaknesses:

      (1) They do not employ the reduction methodology more suited for the single neuron model they consider.

      (2) Their derivation of the neural mass model is based on several assumptions, and not all well justified.

      (3) Their formulation of the mean-field derivation is unnecessary complicated, it can be strongly simplified by following previously published approaches to derive biologically realistic neural masses.

      (4) Their model seems to work only for highly synchronized situations and not for the standard asynchronous evolution usually observed in neural circuits.

      General Statements:

      The authors honestly declared the many limitations of their approach, once assumed this the results of the mean-field are somehow inconsistent with the neural network simulations as expected.

      The authors suggest to employ this model for the simulations on the whole connectome to follow seizure propagation, however I believe that a simpler model, as the Epileptor, remains superior in this respect to this model. That indeed includes biophysical parameters but their correspondence with the ones employed in the network dynamics remain elusive, due to the many assumptions required to derive this mean field model. Furthermore it is more complicated than the Epileptor, I do not think that the present model will be largely employed by the community.

      Comments on revisions:

      The authors have corrected mistakes present in the manuscript and put a correct list of references.

      However, they refuse

      (1) To simplify the formulation of the model, the model contains unnecessary complications, as I have clearly written in my report, the authors agree, but they do not want to change the formulation;

      (2) To derive the mean field model in a simpler way, as possible, and as I asked many times in my Referee report, this would help the readers to understand the important aspect of the derivation, without not needed and confusing complicated formulations;

      (3) To compare direct simulations of the network with neural mass results in sub-section "Bifurcation analysis: emergent network states and multistability" to show bistability, as I asked.

      As a matter of fact the performed modifications do not solve my previous doubts on the validity of the results reported in the manuscript.

      Therefore, my previous assessments remain valid.

      We thank the editors and the two reviewers for their continued engagement with our manuscript. The three weaknesses retained from the first round are essentially identical between the two public reviews:

      (i) The reduction methodology is not the most suitable for the single-neuron model we consider;

      (ii) The mean-field derivation is unnecessarily complicated;

      (iii) The model works only in highly synchronous regimes and does not reproduce the asynchronous evolution typical of neural circuits.

      Both reviewers explicitly note that their assessments remain unchanged and we have decided not to alter the formulation of the model. We use this response to state—on the public record—exactly where we agree with the reviewers, where we disagree, and why.

      On point (i): the reduction methodology.

      We fully agree with the reviewers' technical observation: the Ott–Antonsen / Lorentzian-ansatz reduction in the form introduced by Montbrió, Pazó and Roxin (2015) is exact for canonical Type I neurons (QIF), whose membrane-potential equation is quadratic, and is not directly applicable to a Type II / Hodgkin–Huxley-type neuron whose voltage dynamics is cubic-like. On this point there is no disagreement.

      Where we differ is in the conclusion the reviewers draw from this observation. The reviewers read our work as applying an inappropriate reduction methodology to an inappropriate neuron model. We instead positioned our work, from the outset, as an extension of that methodology: we keep the biophysically detailed Hodgkin–Huxley substrate (because it is the only level at which extracellular ion concentrations, depolarization block, bursting and seizure-like events are biophysically grounded), and we adapt the reduction by approximating the cubic voltage nullcline as a piece-wise quadratic with two parabolas of opposite curvature. This is explicitly an approximate, not exact, mean-field. The Lorentzian ansatz is then applied on each branch of the piece-wise quadratic, with the limitations of this construction analyzed in the manuscript.

      The reviewers' alternative—starting from a Type I canonical model and grafting on biophysical features—would indeed yield an exact mean-field, but it would forfeit precisely what motivates our work: a tractable mesoscopic description in which the slow variables are physiologically interpretable ion concentrations rather than phenomenological parameters. The trade-off is that we give up exact rigour in order to construct a bridge between the Montbrió-style next-generation neural mass models on one side and the Epileptor on the other, with the additional benefit that the parameters of the resulting neural mass retain a biophysical correspondence (e.g., [K<sup>+</sup>]_bath, Δ[K<sup>+</sup>]_int, [K<sup>+</sup>]_g, the gating variable n) that the Epileptor does not afford.

      We therefore respectfully maintain our position: the methodology is not "the wrong reduction for a Type II neuron"; it is an extended reduction designed to be applicable beyond the Type I case, with explicitly characterized validity.

      On point (ii): the formulation is unnecessarily complicated.

      We agree with the reviewers that, given the assumptions we ultimately adopt, namely that the gating variable n and the potassium concentrations Δ[K<sup>+</sup>]_int and [K<sup>+</sup>]_g are treated as collective (mesoscopic) variables shared by the population, with n a function of the average membrane potential, the closed neural mass equations could be reached by the more direct path used by Guerreiro et al. (2022) and the related literature (R1–R7). In the revised manuscript we now state this explicitly, and we note that the same five-dimensional system arises under either derivation.

      Our choice to follow Chen and Campbell (2022) is motivated by the fact that it makes each approximation visible at the point where it is invoked. In particular, it exposes the moment-closure step (Eq. 19), the vanishing-flux boundary condition (Eq. 28), and the locations where microscopic and mesoscopic variables enter the description. We believe that for a reader trying to extend our framework, for instance to a setting with partial heterogeneity in the slow variables, or with stochastic gating, this is the more useful presentation. We have added a remark stating that the simpler Guerreiro-type derivation reaches the same equations under our assumptions, so that readers can take whichever route they find clearer.

      On point (iii): the model only works in highly synchronous regimes.

      Here we partially agree and partially disagree, and we would like the partial disagreement to appear on the public record.

      We agree that the Lorentzian ansatz is, strictly, valid in regimes where the population's membrane potential distribution is unimodal, that is, when essentially all neurons sit on the same side of the threshold V*. Where we disagree is with the implication that the mean-field model fails outside the strongly synchronous regime. The supplementary analysis in Fig. S2, added in the previous round, quantifies the error introduced by the first-moment approximation of n as a collective variable across the full range of [K<sup>+</sup>]_bath values, spanning quiescent, bursting, seizure-like, sustained ictal and depolarization-block dynamics. The fraction of neurons whose gating variable deviates from the population mean is below 2% for the parameters used throughout the manuscript, and the error becomes appreciable only during the brief transitions between sub- and supra-threshold states. These are precisely the moments at which the population is genuinely bimodal and the single-Lorentzian assumption is theoretically expected to leak. In other words, the error peaks coincide with the moments where our derivation tells us in advance that the assumption is locally invalid; the model "knows where it fails." Away from these transitions, the mean-field tracks the population average across all dynamical regimes shown in Fig. 3, not only in the most strongly synchronized ones.

      This is, in our view, the strongest argument we can make: we are not claiming exactness, and we are not unaware of the limitations. We have characterized them analytically (the construction of the piece-wise Lorentzian, and the theoretical reason a closed solution exists only when the two branches collapse onto one), and we have characterized them numerically (Fig. S2). The deviations are bounded, their location in parameter space is well identified, and they coincide with transitions where the underlying assumption is locally violated. We believe this constitutes a controlled approximation rather than an uncontrolled one, and we would like this distinction to be visible to readers of the Reviewed Preprint.

      We note, in this connection, that the reviewers' preferred reference point, the next-generation neural mass model of Montbrió et al. (2015), which is exact and one-to-one with its underlying network, is exact precisely because the underlying network is a network of QIF neurons. The corresponding statement for a network of Hodgkin–Huxley-type neurons with explicit ion exchange does not, to our knowledge, exist in closed form, and may not exist at all. The relevant question is therefore not whether our model matches the exactness of the QIF case, but whether the controlled approximation we provide is useful. Given the qualitative agreement with neural-network simulations across the full range of [K<sup>+</sup>]_bath, the qualitative agreement with the in vitro recordings, and the recovery of the expected bifurcation structure with new emergent regimes, we believe the answer is yes.

      Other outstanding points in the review.

      Reviewer 2 reiterates the view that the Epileptor remains superior for whole-connectome seizure-propagation simulations because it is simpler and better characterized. We do not dispute that the Epileptor is more thoroughly analyzed and more parsimonious. The complementarity we propose is not a replacement but a parameter-grounding, as the Epileptor's phenomenological parameters (excitability, slow permittivity) acquire, in the present framework, an interpretation in terms of measurable biophysical quantities (extracellular potassium, intracellular potassium variation, glial buffering).

      We thank the reviewers and editors once again for their careful reading, and we are grateful that the points of disagreement have been sharpened to a state where readers can judge them transparently.


      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors derive a mean-field model for a network of Hodgkin-Huxley neurons retaining the equations for ion exchange between the intracellular and extracellular space.

      The mean-field model derived in this work relies on approximations and heuristic arguments that, on the one hand, allow a closed-form derivation of the mean-field equations, and on the other hand restrict its validity to a limited regime of activity corresponding to quasi-synchronous neuronal populations. Therefore, rather than an exact mean-field representation, the model provides a description of a mesoscopic population of connected neurons driven by ion exchange dynamics.

      We agree with the reviewer's characterization. Our manuscript describes the derivation as relying on "approximations and heuristic arguments" and states that "the derivation is not exact"; what we provide is a controlled, approximate mesoscopic description in which the slow variables are physiologically interpretable ion concentrations rather than phenomenological parameters. An exact closed-form thermodynamic limit is, to our knowledge, available only for canonical Type I (QIF) networks (Montbrió, Pazó and Roxin, 2015) and a few of their extensions; it is not currently known for a Hodgkin–Huxley-type network with explicit ion-exchange dynamics. We acknowledge that the original description of the regime of validity may have caused confusion on this point, and in the revised manuscript we have therefore replaced the looser formulation "strongly synchronous regimes" by the more accurate "regimes where the membrane-potential distribution is unimodal and can be reasonably approximated by a Lorentzian" throughout the manuscript.

      Strengths:

      The idea of deriving a mean-field model that relates the slow-timescale biophysical mechanism of ion exchange and transportation in the brain to the fast-timescale electrical activities of large neuronal ensembles.

      We thank the reviewer for recognizing the motivation behind our work. This explicit coupling between slow biophysical ion dynamics and fast electrical activity is precisely the feature we tried to preserve in the reduction, even at the cost of giving up exactness.

      Weaknesses:

      The idea underlying this work is not completely implemented in practice.

      We address this general statement through the four specific sub-points the reviewer raises in the paragraph that follows.

      The derived mean field model does not show a one-to-one correspondence with the neural network simulations, except in strongly synchronous regimes.

      We partially agree and partially disagree. We agree that the Lorentzian ansatz is strictly valid where the membrane-potential distribution is unimodal, i.e. when essentially all neurons sit on the same side of the threshold V*. We disagree with the implication that the mean-field fails outside this regime. To make this claim quantitative, we added a new supplementary figure (Fig. S2) that quantifies the deviation of individual neurons' gating variables from the population mean across the full range of [K<sup>+</sup>]_bath values—quiescent, bursting, seizure-like, sustained ictal and depolarization-block dynamics. The fraction of deviating neurons is below 2% for the parameters used in the manuscript, with localized peaks only during the brief, genuinely bimodal transitions between sub- and supra-threshold states—precisely the moments at which the theory predicts the assumption to be locally invalid. Away from these transitions, the mean-field tracks the population average across all dynamical regimes shown in Fig. 3, not only in the strongly synchronized ones.

      The agreement with the in vitro experiment is hardly evident, both for the mean-field model and for the network model.

      We acknowledge that the experimental and simulated traces in the original Fig. 4 did not match quantitatively; this was never our intention. The figure and its caption have been reorganized in the revised manuscript to frame the comparison as qualitative: we aim to demonstrate the shared structure i.e., the slow modulation of fast population activity by extracellular potassium fluctuations, rather than to claim a quantitative fit.

      We also added two clarifications that account for the residual differences: (i) the network simulations were intentionally run with rescaled biophysical parameters (membrane capacitance, gating time constants) to keep the computational cost feasible, a standard practice when the goal is to validate dynamical mechanisms rather than absolute timescales; (ii) the in vitro LFP recordings were AC-coupled, so the slow DC components visible in the mean-field traces are filtered out at acquisition.

      The assumptions made to derive the closed-form equations of the mean-field model have not been justified by any biological reason, they just allow for the mathematical derivation.

      We agree that the modelling assumptions were scattered through the original derivation. In the revised manuscript, the three core assumptions are stated explicitly at the point of derivation: (i) the gating variable n is treated as a collective, population-averaged variable; (ii) the potassium concentrations Δ[K<sup>+</sup>]_int and [K<sup>+</sup>]_g are homogeneous across the population, biophysically justified by the rapid redistribution of ions through diffusion and electrochemical gradients, which enforces near-instantaneous equilibration at the mesoscopic scale; (iii) no heterogeneity is assumed at the level of ion dynamics. The meaning of "locally homogeneous" is now defined explicitly.

      On the biophysical motivation of the in vitro perturbation used in the experiment, we have added a new Methods subsection that explains how low extracellular Mg<sup>2+</sup> unblocks NMDARs and abolishes the divalent-cation stabilisation of the resting membrane potential, depolarising hippocampal neurons and increasing the driving force for outward K<sup>+</sup> currents. This provides a biophysical link between the experimental perturbation and the model's main control parameter, the extracellular potassium concentration. We also added a reference to the well-established model of epileptic discharges that underpins the experiment.

      The final form of the mean-field equations does not clarify whether or not microscopic variables are used together with macroscopic variables in an inconsistent mixture.

      We now explicitly acknowledge that in the spiking-network simulations the gating variable n is microscopic (each neuron has its own n_i), whereas in the mean-field derivation it is treated as mesoscopic and shared by the population. This asymmetry between modalities is discussed both in the Results and in the Limitations sections, and is identified as a likely source of some of the discrepancy between the two modalities.

      We have also made the notation in Eqs. (36)–(37) consistent (firing rate r used throughout, full current-based dV/dt̄ restored) and fixed the typos and broken equation/reference labels that contributed to the impression of inconsistency (Eqs. 18, 28, 29; the Fig. 2(c) [K<sup>+</sup>] bath label; the lost reference at line 696).

      Reviewer #2 (Public review):

      Summary:

      The authors aim to develop a neural mass model characterized by a few collective variables mimicking the dynamics of a network of Hodgkin – Huxley neurons encompassing ion-exchange mechanisms. They describe in detail the derivation of the mean-field model, then they compare experimental results obtained for the hippocampus of a mouse with the neural network simulations and the mean-field results. Furthermore, they report a bifurcation analysis of the developed model and simulation of a small network containing various coupled neural masses, somehow moving towards the simulation of an entire connectome.

      We thank the reviewer for the accurate summary of the manuscript's structure and aims.

      Strengths:

      The author attempts to develop a mean-field model for a globally coupled network of heterogeneous Hodgkin-Huxley neurons with an explicit ion exchange mechanism between the cell interior and exterior.

      We thank the reviewer for recognizing this objective. The retention of Hodgkin–Huxley dynamics with explicit ion exchange is precisely the feature that distinguishes our framework from QIF-based reductions, and it is what enables the slow variables of the resulting mean-field to retain a direct biophysical interpretation.

      Weaknesses:

      (1) It seems that the reduction methodology that is employed is not the most suitable one for the single-neuron model they are considering.

      We agree, on technical grounds, with the observation: the Ott–Antonsen / Lorentzian-ansatz reduction is exact for canonical Type I neurons (QIF) and is not directly applicable to a Type II Hodgkin–Huxley-type neuron with a cubic-like voltage nullcline. Where we differ is in the conclusion. We did not apply an inappropriate reduction to an inappropriate neuron; we deliberately extended the methodology by approximating the cubic nullcline as a piece-wise quadratic with two parabolas of opposite curvature, and then applying the Lorentzian ansatz on each branch. The result is an explicitly approximate, biophysically grounded mean-field, with its regime of validity stated and quantified (Fig. S2).

      To make this positioning explicit, we have added a paragraph to the Introduction that situates our work within the next-generation neural mass literature (Byrne et al. 2020; Montbrió, Pazó & Roxin 2015; Guerreiro et al. 2022; Forrester et al. 2024; Perl et al. 2023; Gerster et al. 2021; and works on short-term plasticity, adaptation, conductance-based reductions,

      spike-timing-dependent plasticity, random connectivity and noise) and clarifies that we see our contribution as complementary to these approaches, not as a competitor to the exact QIF reductions.

      (2) The authors' derivation of the neural mass model is based on several assumptions, and not all well justified.

      We agree that, in the original submission, the modelling assumptions were scattered through the derivation. In the revised manuscript, the three core assumptions are stated explicitly at the point of derivation: (i) the gating variable n is treated as a collective population-averaged variable; (ii) the potassium concentrations Δ[K<sup>+</sup>]_int and [K<sup>+</sup>]_g are homogeneous across the population, biophysically justified by the rapid redistribution of ions through diffusion and electrochemical gradients, which enforces near-instantaneous equilibration at the mesoscopic scale; (iii) no heterogeneity at the level of ion dynamics is assumed. The meaning of "locally homogeneous" is now defined explicitly. In addition, we have added Fig. S2, which quantifies numerically the error introduced by the moment-closure assumption (deviation below 2% for the parameters used in the manuscript).

      (3) The formulation of the mean-field derivation is unnecessarily complicated. It could be heavily simplified by following previously published approaches to derive biologically realistic neural masses.

      We agree that, under the assumptions ultimately adopted in our model—namely that n, Δ[K<sup>+</sup>]_int and [K<sup>+</sup>]_g are mesoscopic—the final five-dimensional system can be reached by the more direct path used by Guerreiro et al. (2022) and the related literature. We now state this explicitly in the revised manuscript and note that the same system arises under either derivation, so that the reader can take whichever route they find clearer. Our choice to retain the Chen and Campbell (2022) formalism is pedagogical: it exposes the moment-closure step (Eq. 19), the vanishing-flux boundary condition (Eq. 28), and the locations where microscopic versus mesoscopic variables enter the description, which is the more useful presentation for a reader wishing to extend the framework (e.g. to partial heterogeneity in the slow variables or to stochastic gating). We also made the notation in Eqs. (36)–(37) consistent (firing rate r used throughout, full current-based dV/dt̄ restored) and fixed a number of typos and broken equation/reference labels.

      (4) The model seems to work only for highly synchronized situations and not for the standard asynchronous evolution usually observed in neural circuits.

      We partially agree and partially disagree. We agree that the Lorentzian ansatz is strictly valid where the membrane-potential distribution is unimodal; we have replaced "strongly synchronous regimes" by this more accurate formulation throughout the manuscript. We disagree, however, with the implication that the mean-field is useful only in those regimes. Fig. S2, added in this revision, explicitly quantifies the deviation across all dynamical regimes (quiescent, bursting, seizure-like, sustained ictal and depolarization-block dynamics): it remains below 2% for the parameters used in the manuscript, with localized peaks only during the brief sub-to-supra-threshold transitions where the population is genuinely bimodal. Away from these transitions, the mean-field tracks the population average across all dynamical regimes shown in Fig. 3.

      General Statements:

      The authors honestly declared the many limitations of their approach. It is assumed that the results of the mean-field are somehow inconsistent with the neural network simulations as expected.

      We thank the reviewer for acknowledging that the limitations are honestly declared. As detailed above and quantified in Fig. S2, the deviation from the network simulations is bounded and well characterized; it is not assumed but measured.

      The authors suggest employing this model for the simulations on the whole connectome to follow seizure propagation, however, I believe that the Epileptor remains superior in this respect to this model. That indeed includes biophysical parameters but their correspondence with the ones employed in the network dynamics remains elusive, due to the many assumptions required to derive this mean-field model. Furthermore, it is more complicated than the Epileptor, I do not think that the present model will be largely employed by the community.

      We do not propose our model as a direct replacement for the Epileptor and we do not dispute that the Epileptor is more thoroughly analyzed and more parsimonious. The complementarity we propose is not a replacement but a parameter-grounding: the Epileptor's phenomenological parameters (excitability, slow permittivity) acquire, in our framework, a concrete interpretation in terms of measurable biophysical variables (extracellular potassium, intracellular potassium variation, glial buffering). Retaining the Hodgkin–Huxley substrate is essential to ground these variables biophysically.

      To make this complementarity more visible, the Limitations and Discussion section has been expanded to discuss the choice of a purely excitatory network as a first step (with excitatory–inhibitory generalizations available via the synaptic reversal potential) and to point to additional biological ingredients (calcium and other ions, plastic synapses, random connectivity and noise, adaptation, spike-timing-dependent plasticity) that the framework can accommodate, with reference to the next-generation neural mass literature.

      We thank the reviewers and editors for their careful reading. We hope this public response makes our reasoning, the limits of our approach, and the concrete revisions made in this round transparent.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In general, the writing is scattered. Every time a model is introduced, one starts from the general formulation only to find that a very simplified case is used with respect to that formulation, which is very confusing. Authors need to reduce unnecessary formulations that confuse the reader and make it clear which formulations are actually used.

      We thank the reviewer for this comment and understand the concern regarding the balance between general formulations and specific approximations. Our intention in including the more general equations and derivations (e.g., Eq. 7 and others) was pedagogical — to ensure completeness and transparency in the modeling steps, especially for readers less familiar with mean-field reductions of biophysically detailed models. These general forms also serve to clarify the assumptions underlying the simplifications we employ. In the latest version, we improved the clarity of core equations (e.g., Eq. 37), which form the basis of all simulations presented (see details below, in the answer to question 14).

      (2) The Introduction would benefit from a wider view of the literature. The literature on exact mean field models (i.e. derived from the Lorentzian Ansatz) has flourished in the last years. In particular, it would be worth considering the following papers, where exact neural mass models are applied to perform whole-brain and large-scale brain simulations:

      Forrester, M., Petros, S., Cattell, O., Lai, Y. M., O'Dea, R. D., Sotiropoulos, S., & Coombes, S. (2024). Whole brain functional connectivity: Insights from next generation neural mass modelling incorporating electrical synapses. PLOS Computational Biology, 20(12), e1012647.

      Perl, Y. S., Zamora-Lopez, G., Montbrio, E., Monge-Asensio, M., Vohryzek, J., Fittipaldi, S.,

      Campo, C. G., Moguilner, S., Ibanez, A., Tagliazucchi, E., Yeo, B. T. T., Kringelbach, M. L., & Deco, G. (2023). The impact of regional heterogeneity in whole-brain dynamics in the presence of oscillations. Network Neuroscience, 7(2), 632-660.

      Byrne, Aine, James Ross, Rachel Nicks, and Stephen Coombes. "Mean-field models for EEG/MEG: from oscillations to waves." Brain topography 35, no. 1 (2022): 36-53.

      Gerster, M., Taher, H., Skoch, A., Hlinka, J., Guye, M., Bartolomei, F.,... & Olmi, S. (2021). Patient-specific network connectivity combined with a next generation neural mass model to test clinical hypothesis of seizure propagation. Frontiers in Systems Neuroscience, 15, 675272.

      Byrne, Aine, Reuben D. O'Dea, Michael Forrester, James Ross, and Stephen Coombes. "Next-generation neural mass and field modeling." Journal of neurophysiology 123, no. 2 (2020): 726-742.

      Benitez-Stulz, Sophie, Samy Castro, Gregory Dumont, Boris Gutkin, and Demian Battaglia. "Compensating functional connectivity changes due to structural connectivity damage via modifications of local dynamics." bioRxiv (2024): 2024-05.

      We have added the following paragraph:

      “Recently, a class of these models, called next-generation neural mass models [42], has been developed based on an analytical approach introduced by [25] that allowed for the exact derivation of mean field parameters for a population of quadratic integrate-and-fire (QIF) neurons. These can be linked to EEG/MEG oscillations [43], including epipeltic seizures [43], and have been used to study various aspects of the whole-brain dynamics such as the low-dimensional manifold of the resting state [45,46], aging [47] and neural signatures of consciousness [48].”

      We have also modified the preceding paragraph of the introduction that now reads:

      “At the mesoscopic level, the observable properties of a neuronal ensemble are generally explained by statistical physics formalism of mean-field theory [19-22]. Mean-field models demonstrated a predictive value for studying the mesoscopic dynamics of neuronal populations [23], providing statistical descriptions of neuronal networks [2, 19, 24-29], which can be used to address questions related to network-level mechanisms [12, 24, 30].

      In general, neural mass models have a low enough number of parameters to be tractable and provide general intuitions regarding mechanisms underlying complex neuronal activity [31-36]. For example, statistical population measures, such as the firing rate, can be used to assess mesoscopic dynamics [1, 7, 31, 36-41].”

      (3) Moreover, conductance-based models have been already implemented in neural mass models not only in references [69, 71, 95], but also in:

      Guerreiro, I. C., Di Volo, M., & Gutkin, B. (2023). A new generation of reduction methods for networks of neurons with complex dynamic phenotypes.

      Capone, C., Di Volo, M., Romagnoni, A., Mattia, M., & Destexhe, A. (2019). State-dependent mean-field formalism to model different activity states in conductance-based networks of spiking neurons. Physical Review E, 100(6), 062413.

      We have added the following sentence:

      “Moreover, conductance-based couplings between the spiking neurons have been already implemented in neural mass models [58, 59, 91, 93, 121], but without an extracellular exchange mechanism.”

      (4) Sec. 1.1 As previously established in the literature, a system of all-to-all coupled neuronal equations can be solved exactly in the thermodynamic limit (i.e., infinite neurons limit) if the single neuron membrane potential equation is a quadratic function and if the instantaneous distribution of membrane potentials of neurons in a population is described by a Lorentzian [Montbrió, E., Pazó, D. & Roxin, A. Physical Review X 5 (2), 021028 (2015)]. This means that the thermodynamic limit can be performed for a Canonical Type I model like the quadratic integrate-and-fire.

      What is the biological justification and the reason to approximate a different neuron type (a type II neuron model), whose membrane potential equation resembles a cubic function, with a quadratic function? The fact that it can be solved in the quadratic approximation is not, in my opinion, a sufficient justification. It would be more correct to start from a type I neuron at the microscopic level with a quadratic function and then provide additional biological features.

      We thank the reviewer for raising this important point. We respectfully disagree with the notion that starting from a canonical Type I model (such as the quadratic integrate-and-fire neuron) would be a more biologically grounded approach. While the quadratic form is analytically convenient, it does not capture certain key features of neuronal excitability particularly those related to bursting, seizure-like events, and depolarization block which are closely tied to the cubic-like nullcline geometry arising in Hodgkin–Huxley-type models, especially in the presence of slow ion dynamics.

      Our work seeks to bridge biophysical realism with analytical tractability. The step-wise quadratic approximation we employ is specifically designed to mimic the cubic membrane potential profile that emerges from the full ion-exchange dynamics. While the Lorentzian Ansatz is not strictly justified in this case from first principles, we show that it yields a workable and biologically interpretable mean-field description, which aligns with single-neuron dynamics, population simulations, and even in vitro observations. To our knowledge, this is a novel contribution that extends mean-field modeling beyond currently available approaches, which are often restricted to simplified or phenomenological neuron models.

      In this context, using a quadratic approximation is not merely a mathematical convenience — it is a means to retain key dynamical features of more realistic (non-Type I) neurons within a tractable framework, enabling insights into complex behaviors like multistability and pathological bursting.

      (5) Sec. 1.2 As shown in Figure 3, the mean-field equations do not show a one-to-one correspondence with the neural network simulations, except in strongly synchronous regimes. This represents a strong limitation in the model, especially because exact neural mass models (as shown in Reference [23]) perfectly fit the dynamics of the underlying network model both in the asynchronous and in the synchronized regime.

      We appreciate the reviewer’s observation and acknowledge that our original description may have caused confusion. The model's validity is not strictly limited to strongly synchronous regimes, but rather to regimes where the distribution of membrane potentials across the neuronal population remains unimodal and can be reasonably approximated by a Lorentzian. This includes but is not restricted to—highly synchronized states.

      We agree that this distinction is important and have clarified it in the revised manuscript (e.g., “in strongly synchronous regimes” —> “in regimes where the membrane potentials' distribution is unimodal and can be reasonably approximated by a Lorentzian”).

      In contrast to exact mean-field reductions based on quadratic integrate-and-fire neurons (e.g., [23]), our model originates from a biophysically grounded HH-type neuron with ion exchange dynamics, and necessarily involves heuristic approximations to achieve a closed-form mean-field description. While this results in a less exact correspondence with network simulations in more heterogeneous or bimodal states, our goal was to retain biological interpretability and account for phenomena such as ion-driven bursting and seizure-like transitions, which are not captured by standard QIF-based neural masses.

      We see our contribution as complementary to existing exact reductions — offering a biophysically grounded alternative that remains tractable and informative in a relevant class of unimodal, mesoscopic dynamical regimes.

      (6) Sec. 1.3 In this section the authors show the comparison between in vitro experiments and simulations with both the network model and the neural mass model (Figure 4, panels a,b,c). The qualitative agreement that is supposed to be shown is hardly evident. The shape of the signals is different as is the type of bursting. The only agreement results in the fact that there are repeated spiking events at successive times in a periodic manner. However, the time scale of the simulations is different for neural network simulation and mean-field experiment, making it difficult to compare them. While the period of the bursting event is around 2 min for mean field simulation (in according with experiments), the time scale of the network simulation is 60 times smaller, thus meaning that we are considering completely different mechanisms and phenomena. The justification given by the authors, that "the parameters were modified to simulate shorter fluctuations (in the network of Hodgkin-Huxley neurons) for computational efficiency" is inappropriate.

      The poor agreement turns out to be even worse in the comparison between experiments and mean-field simulations shown in panels d and e of Figure 4. While the mean field simulation is characterized by a periodic behaviour both in the mean membrane potential and in the external potassium concentration, the in-vitro traces are not periodic and show an increasing irregular activity of the extracellular LFP in correspondence with increasing external potassium concentration.

      How it is possible to justify the implementation of this model if the working hypotheses are not supported by the results? The worst agreement of the network simulations with the experiments reinforces the doubt raised in the previous point: what is the reasoning underlying the choice of Hodgkin-Huxley as a single neuron model?

      We thank the reviewer for this detailed critique. We acknowledge that the comparisons in Figure 4 involve limitations and we now provide a clearer rationale and context in the revised manuscript. First, we emphasize that our intention is not to claim a quantitative match between the experimental and simulated traces, but rather to demonstrate that our model grounded in biophysical mechanisms such as ion exchange is capable of qualitatively reproducing a key feature observed experimentally: the slow modulation of neuronal activity by extracellular potassium concentration. For example, both in vitro (Fig. 4a, 4d) and in our simulations (Fig. 4b, 4e), bursts of activity ride on slower oscillations of potassium, and the interplay of fast and slow dynamics is central to both.

      Regarding the discrepancy in timescales between the neural network and mean-field simulations: the network simulations were intentionally run with accelerated dynamics by rescaling biophysical parameters (e.g., membrane capacitance and gating time constants) to keep the computational cost feasible. We now clarify in the manuscript that this choice is standard practice in computational modeling when the primary goal is to validate dynamical mechanisms rather than replicate absolute timescales.

      On the shape of LFP signals: the experimental recordings were AC-coupled, and the DC components associated with slower shifts in membrane potential such as those modeled in the mean-field simulations are not captured in those recordings. This limits the visibility of key features like the underlying potential jumps. Additionally, no claim is made regarding a specific bursting classification in either data or simulation.

      We agree that the experimental trace in Fig. 4d shows more complex, non-periodic dynamics (e.g., slowing burst frequency and irregularity), which are not captured by our current deterministic model. These differences could plausibly arise from additional physiological processes (e.g., stochastic transitions between metastable regimes or variability in ion regulation) that are not modeled here. In future work, such phenomena may be captured by introducing noise or parameter variability (see, e.g., Saggio et al., A taxonomy of seizure dynamotypes , elife 2020), or by allowing the parabola coefficients in the nullcline approximation to vary dynamically.

      Finally, regarding the choice of a Hodgkin–Huxley-type neuron: this model allows us to incorporate a biophysical description of ion exchange, which is central to the phenomena we study. While modeling the spiking mechanisms explicitly precludes certain mathematical simplifications available to very simplified neuron models with reset, it enables direct links between mesoscopic dynamics and measurable quantities such as extracellular potassium an essential objective of our work. To summarize, we rearranged Fig4:

      Potassium can have periodic behavior with V bursting riding on top (Fig.4 a). The model also shows this behavior at different timescales (Fig. b,c,e).

      AC LFP recording is filtered so we might not see the V jump during the bursts (because we do not have DC recordings). No claim about bursting class here.

      Potassium can also have more complex behavior (e.g., slowing down of burst frequency Fig.4.d), that the deterministic model do not show, but maybe exploring dynamical parameters (e.g., from parabolas or K_bath) or with added noise allowing to jump between regimes (reference Saggio et al. eLife 2020).

      (7) Sec. 1.5 Here six neural masses are coupled via long-range structural connections with random weights. Simulations of the system are shown for two different values of the global coupling parameter (G = 0 and G = 100). How many realisations of the network have been considered?

      We thank the reviewer for pointing this out. The presented simulation was intended as a proof-of-concept demonstration to illustrate the model’s capacity to support network-level propagation of pathological activity. For this purpose, we considered a single representative realization of the structural connectivity with random weights. Given the deterministic nature of the model and the qualitative focus of the demonstration, additional realizations do not qualitatively change the observed behavior — namely, the transition from localized to network-wide bursting as coupling strength increases. We have now clarified this in the revised manuscript.

      “This simulation serves as a proof of concept to illustrate how local pathological activity can propagate through a network depending on the strength of coupling. We used a single representative realization of randomly weighted structural connectivity. While we did not perform a systematic exploration of different realizations or coupling strengths, we observed that the qualitative behavior namely, the emergence of network-wide bursting beyond a critical coupling threshold remains robust across similar setups. The model is compatible with empirical connectome data and can be readily extended to simulations using realistic brain network architectures.”

      In future applications involving data-driven network architectures or variability analyses, we agree that exploring multiple realizations or empirical connectomes will be valuable.

      How do the results depend on the different choices of the random weights? What is the dependence of the emergent dynamics on G? What kind of dynamics can be observed varying smoothly the parameter G (e.g. from 0 to 100)?

      This section serves as a proof of concept to show that pathological activity in one node can propagate through the network when coupling is strong. We used a single random weight configuration and did not systematically explore variations in G or connectivity. While richer dynamics likely emerge across intermediate values of G, a full parameter sweep is beyond the scope of this study. We clarify this in the revised text (see answer above).

      (8) Sec. 2.1 In the description of the experiment it is mentioned that only Mg^{2+} is varied. What is the role played by Mg^{2+} variation in influencing the external potassium concentration variation? How the experiment can be linked to the model? How the hypothesis of introducing an equation for the potassium concentration current in the microscopic model is supported by the experiment and vice-versa?

      We thank the reviewer for this question. We have added a new subsection in the Methods explaining the.agnesium removal as a mean to influence the external potassium dynamics:

      “The membrane of hippocampal neurons is equipped with N-methyl-D-aspartate type glutamate receptors (NMDARs). These receptors have a very high affinity for glutamate and can, in principle, be activated by ambient glutamate present at low concentrations in the brain extracellular fluid (ECF). Under normal physiological conditions, this activation does not occur because extracellular magnesium ions (Mg<sup>2+</sup>) block the NMDAR channel at membrane potentials more negative than about –50 mV; this voltage-dependent block prevents receptor activation at rest. When extracellular magnesium is removed, the block is relieved, allowing NMDARs to be activated, leading to neuronal depolarization toward the action potential threshold [117].”

      “In addition, as a divalent cation, Mg<sup>2+</sup> interacts with the negatively charged neuronal membrane, contributing to the stabilization of the resting membrane potential. Lowering extracellular magnesium concentration disrupts this effect, resulting in membrane depolarization [118].”

      “Consequently, magnesium removal not only facilitates NMDAR-dependent depolarization, but also directly depolarizes neurons. This depolarization increases the driving force for outward potassium currents through K<sup>+</sup> channels, meaning that variations in Mg<sup>2+</sup> can indirectly influence external potassium dynamics during neuronal activity.”

      (9) Sec. 2.6 The modified version of the continuity equation has been derived following Reference [95], where the authors consider a network of Izhikevich neurons, and each neuron is modelled by a two-dimensional system consisting of a quadratic integrate and fire equation plus an equation that implements spike frequency adaptation. In particular, in [95] the authors achieve a closed set of mean-field equations with the inclusion of the mean-field dynamics of the adaptation variable by using a Lorentzian ansatz combined with the moment closure approach. The moment closure condition is also assumed in the present manuscript (Eq. 19). Under which assumptions is the implementation of the moment closure condition justified?

      We are thankful to the reviewer (and also to the R2) for pointing out to the validity of the justification of the assumptions that we have used in our formalism. We hence agree that the moment closure is not a sufficient justification for assuming that V depends on the mean n, which is neccessary for the derivation of Eq. 20, but in addition we need the assumption that n can be treated as a collective variable as it is done in the works mentioned by the reviewer 2. In addition we have performed numerical simulations of the full system to calculate the error term introduced by this approximation, and the results in the new Fig. S2 show that this is below 2% for each of the different dynamical regimes.

      We have hence modified the justification for Eq. (19) reading:

      “Next we assume a first-order moment closure condition for the variable n [59], justified by the numerical simulations of the full network (see Fig. S2) which show that for most of the neurons (close to 99 \% for the value of ∆ same as in the other simulations) the mean of the population is well capturing the behavior of the single neurons [122]. Finally, putting together these factors and assuming that n can be treated as a collective variable for each neuron (see Limitations of the model} section) we arrive to ” and also

      “The validity of the first moment closure, Eqs. (19), as in [59], is supported by the numerical simulations, which show that, both, during the silent regime and when seizure-like events occur, n<sub>i</sub> for most neurons track the network averaged ⟨n | V, η⟩. In particular, it is less than 2% of the neurons that fire while the mean is low, and vice-versa, Fig. S2. In less synchronized scenarios (larger ∆ or smaller J), however, this value would increase, but the mean would always capture the qualitative behaviour of the population.”

      This is also now explicitly mentioned in the following paragraph:

      “Unlike the mean membrane potential ⟨V⟩ and the firing rate (r), which can be explicitly derived from the continuity equation under the Lorentzian assumption, the expression for ⟨n(t)⟩ in Eq. (26) is formal. In our mean-field model, the gating variable (n) is treated as a global population variable, evolving deterministically as a function of the average membrane potential. Therefore, ⟨n(t)⟩ corresponds to the collective gating variable assumed to be shared by all neurons, and is not computed by averaging distinct microscopic (n<sub>i</sub>) values.”

      (10) Considering also the comments reported above, I think that it would make more sense to start from an Izhikevich neuron model as microscopic model and add the equations for the ionic currents as mesoscopic variables (i.e. written as population average variables), instead of starting from the Hodgkin-Huxley single neuron model and trying to make hardly justifiable approximations and simplifications.

      We respectfully disagree. While the Izhikevich model is computationally efficient, it lacks the biophysical detail required to capture key ion-driven mechanisms such as depolarization block, slow ion accumulation, and specific burst-initiation dynamics all of which are central to our study. The Hodgkin–Huxley framework, despite requiring approximation, provides the necessary physiological grounding to link microscopic ion exchange with emergent population behavior.

      (11) Sec. 2.7 What is the advantage of using six more parameters to fit, like R-,R+,c-,c+,I-,I+?

      This is in contradiction with the spirit of deriving a mean-field model, where the number of parameters should be reduced. What is the advantage of this mean-field derivation with respect to other mean-field derivations of Hodgkin-Huxley neurons, like the one in Reference [9]?

      The additional parameters (R±, c±, I±) are not arbitrary they compactly parametrize the cubic-like nonlinearity of the membrane potential dynamics in our stepwise-quadratic approximation. This trade-off allows us to preserve essential biophysical features of HH neurons (e.g., bursting regimes, depolarization block) within a tractable analytic framework. Compared to alternative approaches like in ref. [9], which focus on phenomenological reductions and do not yield an ODE system, our model offers more direct interpretability in terms of ion dynamics, providing a closer link between microscopic mechanisms and mesoscopic activity patterns.

      (12) Sec. 2.11 The derivation of the mean-field dynamics for the gating variable is rather heavy and difficult to follow. This section could be simplified, whilst also better explaining the underlying approximations and the validity of these approximations, which is currently missing.

      We agree that the derivation is technical, but we chose to retain it for transparency, as it follows the Chen and Campbell approach and makes key approximations such as moment closure explicit. We have now added a clarification that n is treated as a collective variable We hope that the current level of detail helps readers understand the assumptions underlying the gating variable dynamics.

      (13) Sec. 2.12 The derivation of Eqs. (36) is quite confusing and needs to be re-written in a clearer form. Why are both the variables x and r present in these equations, since they are proportional according to Eq. (25)?

      We thank the reviewer for pointing this out. We have adjusted the equations to improve clarity and now consistently express the firing rate in terms of a single variable. This removes the redundancy and simplifies the presentation.

      (14) Sec. 2.13 The derivation of Eqs. (37) is quite confusing and needs to be rewritten in a clearer form.

      Both the auxiliary variable x and the firing rate r are present in this equation, the same as in Eq. (36). Therefore it is presented as a set of equations for the auxiliary variable x and for the physical variable V. Moreover in the equation for dV/dt, the quadratic term in V has disappeared and it is not clear to me which are the variables corresponding to I- and I+. In particular, in Eqs. (36) there are two different current terms I-,I+ for the two equations related to dy/dt. In Eqs. (37) there is a single term (I_{cl} +I_{Na}+I_K+I_{pump})/C_m which is identical for both equations related to dV/dt. I was expecting two different terms also in Eqs. (37).

      We appreciate the reviewer’s close reading. To improve clarity, we now express the dynamics in terms of the firing rate r, replacing \dot{x} with \dot{r} in both Eq. (36) and Eq. (37) to avoid confusion.

      As for the current terms: in Eq. (37), we reverse the stepwise quadratic approximation and reintroduce the original ionic currents from Eq. (16). This is why the expressions involving I_{\text{cl}}, I_{\text{Na}}, I_K, and I_{\text{pump}} appear as a single summed term in \dot{V}, rather than the split I_-,I_+ terms used in the stepwise approximation. We now clarify this in the text.

      We also write V as \bar{V} to clarify that it refers to the average membrane potential for the neuronal population. Finally, we wrote the final equation in a more compact form to improve clarity (new Eq.38).

      (15) Moreover, while the equation for the gating variable n can be considered as a differential equation for a mesoscopic variable since n depends on average values only, it is not clear to me if the remaining variables 𝛥[K+]_{int}, [K+]_g can be considered mesoscopic or not. Since Eqs. (37) represent a mean-field model, I expect every variable to be a mean-field variable. This could be easily achievable for the extracellular potassium concentration, but I do not understand how a site-specific microscopic variable like the intracellular potassium concentration variation can be automatically inserted in a set of mean-field equations without any averaging or intermediate steps. This is a crucial point to be clarified for the validity of the neural mass equations.

      We thank the reviewer for raising this important point. In our model, we assume spatial homogeneity at the mesoscopic scale, meaning that ion concentrations — both intra- and extracellular — are uniformly distributed across the population. As a result, variables such as \Delta[K^+]_{\text{int}}, Δ[K+]int and [K+]g are treated as population-level averages, consistent with the mean-field framework.

      Moreover, the rate of change of intracellular potassium is tightly coupled to extracellular dynamics via ion exchange mechanisms, justifying its inclusion as a slow, mesoscopic variable. We now clarify this modeling assumption explicitly in the text.

      “By locally homogeneous, we mean that all neurons in the population are assumed to share the same extracellular and intracellular ionic environment and are connected with identical coupling rules, allowing us to treat the population as uniform with respect to ion dynamics and connectivity.”

      “These slow variables are in addition considered to be mesoscopic, meaning they are identical for every neuron in the population.”

      Minor points:

      (1) Figure 2, panel d. Please detail the variable on the y-axis, which is not reported in the figure.

      Done

      (2) Eq. (15) is cited in many parts of the manuscript, while it seems to me it would be more appropriate to reference Eq. (2). Is this a mistake or is there a reason to cite Eq. (15)?

      The reviewer is correct, we have had a wrong equation label, which we have now corrected.

      (3) Figure 4 Would it be possible to show enlargements of the mean membrane potential traces to directly compare the different bursting types shown by the simulation of the different models?

      The panel d already contains enlarged part of the membrane potential traces. For the rest, going back to the Q6, we want to stress again that our intention is not to claim a quantitative match between the experimental and simulated traces.

      (4) Figure 5 In the caption the author refers to "the generic model, single neuron model, and epileptor model". Could you please better explain the models referred to and why they are mentioned? Are the generic model and the single neuron model those that are presented in the Materials and Methods section? Or do you refer to completely different models, as for the epileptor?

      We have removed the reference to the generic model (we had in mind the canonical model for seizures by Saggio et al. 2017), since it is not mentioned in the paper, and we have clarified that the single neuron model and epileptor model, which were used to simulate seizure like events.

      (5) Sec 2.5 As already stated above, the authors need to reduce unnecessary formulations that confuse the reader. Here, for example, Eqs. (6) and (7) are unnecessary, in view of the fact that delta spikes are used (Eq. 8).

      We thank the reviewer for the suggestion, but we disagree, and we think it is better to start the derivations from the more general case, as done with Eqs. 6-7.

      (6) Sec. 2.6 Could you please better explain why in Eqs. (15) and (16), the variable V0 is introduced, while before and after this, the variable V is used?

      We thank the reviewer for the comment. In Eqs. (15) and (16), \dot{V}_0 denotes the free term of the membrane potential equation, i.e., the component driven solely by the intrinsic ionic currents and excluding the synaptic input I_syn. Only this \dot{V}_0 term (a function rather than an independent variable) is approximated by the piece-wise quadratic expression in Eq.(21). In contrast, the variable V represents the membrane–potential variable, which dynamics is obtained by combining \dot{V}_0 with the synaptic current contribution I_syn. In summary, there is no independent variable V_0; only the function \dot{V}_0 is introduced to represent the intrinsic (non-synaptic) component of the membrane–potential dynamics. We have now clarified this in the text.

      (7) In the square brackets of the r.h.s. of Eq. (18), for all the intermediate steps, it appears G^n(V,n) ϱ^V, while there should be G^n(V,n) ϱ^n.

      We thank the reviewer for catching this typo. We have corrected this in the revised manuscript.

      (8) Sec. 2.8 Here the authors affirm that "a double-Lorentzian (or a piece-wise Lorentzian) could be a suitable form for ρ^V (t, V | η). However, it is not clear under which conditions such an assumption would allow a solution to the continuity equation". What are the problems underlying the implementation of the double Lorentzian? It seems to be a more correct form than the single Lorentzian actually implemented.

      We thank the reviewer for this thoughtful question. In principle, a double-Lorentzian ansatz for \rho^V can indeed be implemented in several reasonable ways–for example, by enforcing that the combined area of the two Lorentzian components is normalized to one (to preserve the probabilistic interpretation) and by imposing smoothness constraints at their boundaries. However, despite exploring these implementations, we were unable to obtain non-trivial solutions of the continuity equation under this parametrization. The only solvable case we found is the degenerate one in which the two Lorentzians collapse onto each other (i.e., (x_- = x_+) and (y_- = y_+)), which reduces the ansatz to the single-Lorentzian form used in the manuscript. For this reason, although the double-Lorentzian is conceptually appealing, it did not yield practically useful solutions within our framework.

      (9) Eq. (28). The symbols used for the flux (especially those used in the second-to-last step once the inner integration is performed) are confusing and it is difficult to understand what they mean.

      We thank the reviewer for noting this issue. The problem was due to a LaTeX typo that prevented the vertical lines—indicating that the flux is evaluated at specific points—from rendering correctly. We have now corrected this.

      (10) Eq. (29) In the third step there are some misprints that impair comprehension.

      We thank the reviewer for noting this. We have corrected these misprints in the revised version.

      (11) Line 696. The reference is not displayed.

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      As a really general remark, this manuscript is written in a confusing manner, the authors present their model in a general formulation and their analysis in a complicated way that in the end is not needed, as I will explain in detail in the following.

      Another general question is why the authors want to employ the neural mass reduction methodology developed in [23] to obtain exact mean-field evolution for quadratic neurons (like the quadratic integrate and fire (QIF)) for a model that reveals a cubic dependence on the membrane potential, as the FizhHugh-Nagumo neuron (that indeed is a 2d reduction of the Hodkgin-Huxley model), to obtain an approximate neural mass model that somehow works qualitatively only for synchronized dynamics? Why not use another approach more suited to derive the neural mass model for cubic nonlinearity, as the one suggested in [33] and [69] by Di Volo and co-authors? What is the rationale behind the choice of the authors?

      We appreciate the reviewer’s critical feedback and the opportunity to clarify our methodological choices. Our decision to base the mean-field model on Hodgkin–Huxley-type neurons stems from the need to retain ion channel dynamics, which are essential to capture the coupling between membrane activity and extracellular ionic concentrations. This biophysical link is central to our study and cannot be achieved using more abstract neuron models such as QIF or FitzHugh-Nagumo alone.

      Regarding the mean-field reduction method: while the Ott-Antonsen/Lorentzian framework is indeed exact for QIF neurons, we adopted a stepwise quadratic approximation to apply a similar formalism to the cubic-like dynamics of the HH model. This choice enables us to analytically capture a rich set of behaviors, including bursting, depolarization block, and seizure-like dynamics, in a tractable mean-field system.

      We considered the approach of Di Volo and colleagues [33, 69], but their methodology is tailored to asynchronous irregular regimes, whereas our model is specifically designed to capture dynamics in quasi-synchronous or bursting regimes — including epileptiform activity — which are not covered by the assumptions of the Di Volo framework.

      We now clarify these modeling choices more explicitly in the revised manuscript.

      "Unlike phenomenological or reduced models, the Hodgkin–Huxley framework allows us to retain explicit ion exchange dynamics, which are essential for linking membrane behavior to extracellular potassium fluctuations. This level of biophysical detail is crucial for modeling pathological regimes such as seizure onset and propagation."

      Furthermore, the derivation of the neural mass equations is unnecessarily complicated, as a matter of fact, they approximate all the variables (except the membrane potentials of the single neurons) as collective variables (i.e. the gating variable and the potassium concentration) common to all the neurons. The neural network model for which they derive the neural mass model presents microscopic evolutions of the membrane potential cubic-like plus other global variables equal for all neurons, that depend on collective variables such as the mean membrane potential or the mean firing rate. Once clarified, the derivation of the neural mass model is much simpler, and it is not necessary to follow the approach reported in Reference [95] [Chen, L. & Campbell, S. A. Exact mean-field models for spiking neural networks with adaptation. Journal of Computational Neuroscience 50 (4), 445-469 (2022)] which is unnecessarily complicated. The authors can follow a much simpler methodology as explained by Guerriero et al in Reference [R6] (cited below) where the authors consider the same model studied in [95]. Such a methodology has been applied in many cases already, to introduce realistic aspects in the neural mass model [23] (see References [R1-R7] below). I strongly encourage the authors to reformulate their approach in a simpler and clearer manner, by following the approach reported in [R1-R7]. The manuscript will become more readable and it will gain in comprehension.

      We thank the reviewer for this helpful suggestion. We agree that, given the assumptions made in our derivation (i.e., shared gating and ion concentration variables across neurons), the mean-field equations could alternatively be obtained using the simpler methodology proposed by Guerriero et al. [R6] and related works [R1–R7]. However, we chose to follow the derivation presented by Chen and Campbell [95] because it makes the approximations (e.g., moment closure, flux boundary assumptions) explicit and generalizable to future extensions. However, we also acknowledge that the assumption of n to be treated as a collective variable is needed, and for clarity, we have now added a remark in the manuscript indicating that the same result could be recovered more directly using the approach of Guerriero et al.

      “We note that, under the assumption of globally shared gating and ion concentration variables across the neuronal population, the resulting mean-field equations can also be derived using simpler methods as proposed by Guerriero et al [58]. In this work, we follow the more general formalism of Chen and Campbell [59], which makes the role of key approximations (e.g., moment closure, vanishing flux at boundaries) explicit. This also facilitates potential generalizations to settings with partial heterogeneity or dynamic gating distributions.”

      “Finally, putting together these factors and assuming that n can be treated as a collective variable for each neuron”

      “Unlike the mean membrane potential ⟨V⟩ and the firing rate (r), which can be explicitly derived from the continuity equation under the Lorentzian assumption, the expression for ⟨n(t)⟩ in Eq. (26) is formal. In our mean-field model, the gating variable (n) is treated as a global population variable, evolving deterministically as a function of the average membrane potential. Therefore, ⟨n(t)⟩ corresponds to the collective gating variable assumed to be shared by all neurons, and is not computed by averaging distinct microscopic (n<sub>i</sub>) values.”

      Now I will examine in detail all the manuscript and report comments/remarks/suggestions numbered as (Q#) on how to improve the present manuscript to render it easier to read and more comprehensible, these are not minor remarks, just detailed ones.

      Introduction

      (Q1) The Introduction section needs a part devoted to the reduction methodology developed in [23] for QIF neurons and a presentation of previous works dealing with the introduction of biologically realistic aspects in the neural mass model derived in [23]. Here is a non exhaustive list of such papers concerning the introduction of the following realistic aspects in the neural mass developed in [23]:

      (I) short-term synaptic plasticity :

      [R1] Exact neural mass model for synaptic-based working memory H Taher, A Torcini, S Olmi, PLOS Computational Biology 16 (12), e1008533 (2020)

      [R2] Bursting in a next generation neural mass model with synaptic dynamics: a slow-fast approach H Taher, D Avitabile, M Desroches, Nonlinear Dynamics 108 (4), 4261-4285 (2022)

      [R3] Mean-field approximations of networks of spiking neurons with short-term synaptic plasticity R Gast, K Thomas R, H Schmidt, Physical Review E 104 (4), 044310 (2021)

      (II) spike frequency adaptation:

      [R4] Gast, Richard, Helmut Schmidt, Thomas R. Knösche. "A mean-field description of bursting dynamics in spiking neural networks with short-term adaptation." Neural computation 32.9 (2020): 1615-1634.

      [R5] Population spiking and bursting in next-generation neural masses with spike-frequency adaptation, A Ferrara, D Angulo-Garcia, A Torcini, S Olmi, Physical Review E 107 (2), 024311 (2023).

      (III) conductance-based neuron with a slow current (Izekievic model):

      [R6] A new generation of reduction methods for networks of neurons with complex dynamic phenotypes,IC Guerreiro, M Di Volo, B Gutkin, preprint arxiv: 2206.10370 (2022)

      (IV) spike timing-dependent plasticity:

      [R7] Mean-field approximations with adaptive coupling for networks with spike-timing-dependent plasticity, B Duchet, C Bick, Á Byrne, Neural computation 35 (9), 1481-1528 (2023).

      (V) random connectivity and noise:

      [R8] Mean-field models of populations of quadratic integrate-and-fire neurons with noise on the basis of the circular cumulant approach

      DS Goldobin Chaos: An Interdisciplinary Journal of Nonlinear Science 31 (8) (2021)

      [R9] A reduction methodology for fluctuation-driven population dynamics DS Goldobin, M Di Volo, A Torcini, Phys. Rev. Lett. 127, 038301 (2021)

      [R10] Shot noise in next-generation neural mass models for finite-size networks VV Klinshov, SY Kirillov Physical Review E 106 (6), L062302 (2022)

      I think the authors should refer in the introduction to these previous papers, where realistic biological aspects have been already introduced in the neural mass model developed in [23].

      We have added a whole pragaraph devoted to the next-generation neural mass models and in particular to the other works introducing biological realism in this class of models:

      “Recently, a class of these models, called next-generation neural mass models [42], has been developed based on an analytical approach introduced by [25] that allowed for the exact derivation of mean field parameters for a population of quadratic integrate-and-fire (QIF) neurons. These can be linked to EEG/MEG oscillations [43], including epipeltic seizures [44], and have been used to study various aspects of the whole-brain dynamics such as the low-dimensional manifold of the resting state [45, 46], aging [47] and neural sig natures of consciousness [48]. Number of works dealt with the introduction of biologically realistic aspects in the mostly phenomenological neural mass model derived in [25]. These included short-term synaptic plasticity [49–51], spike frequency adaptation [52, 53], spike timing-dependent plasticity [54], synaptic delay [29], random connectivity and noise [55–57], as well as an extension of the conductance-based neurons with a recovery variable [58–60].”

      (Q2) Line 117 - Please specify what you mean by locally homogeneous, here.

      Thank you for allowing us the opportunity to clarify this. We now report:

      "By locally homogeneous, we mean that all neurons in the population are assumed to share the same extracellular and intracellular ionic environment and are connected with identical coupling rules, allowing us to treat the population as uniform with respect to ion dynamics and connectivity."

      (Q3) In this sub-section the authors should clarify all the hypotheses they employ to derive the neural mass models, not only the Lorentzian approximation they did for a cubic model, but also the fact that they assume that the gating variable n is a global variable as well as that the potassium concentration are assumed to be the same for all neurons, that they assume no heterogeneity at this level. This is a fundamental aspect that should be clarified at this stage already.

      We thank the reviewer for this important observation. We agree and have revised the text in the derivation section to explicitly state all key assumptions. Specifically, we now clarify that:

      (1) The gating variable n is treated as a population-average (global) variable;

      (2) The potassium concentrations Δ[K+]int and [K+]g are assumed to be homogeneous across the neuronal population; and (3) No heterogeneity is assumed at the level of the ion dynamics.

      This assumption is biophysically motivated: ion concentrations — particularly extracellular potassium — tend to redistribute rapidly due to diffusion and electrochemical forces, leading to an effectively well-mixed environment at the mesoscopic scale. As such, assigning separate compartments to individual neurons is not justified in this modeling context. We now explicitly note this in the manuscript to avoid ambiguity.

      “3) We assume that the potassium concentrations, both intracellular(\( \Delta[K^+]_{\text{int}} \)) and extracellular (through the buffering variable \( [K^+]_g \)), are homogeneous across the neuronal population. This is justified physiologically by the rapid redistribution of ions through diffusion and electrochemical gradients, which enforce near-instantaneous equilibration at the mesoscopic scale. As such, assigning separate compartments to each neuron is neither practical nor biologically meaningful in this context. We assume that the potassium concentrations, both intracellular (\( \Delta[K^+]_{\text{int}} \)) and extracellular (through the buffering variable \( [K^+]_g \)), are homogeneous across the neuronal population. This is justified physiologically by the rapid redistribution of ions through diffusion and electrochemical gradients, which enforce near-instantaneous equilibration at the mesoscopic scale. As such, assigning separate compartments to each neuron is neither practical nor biologically meaningful in this context; 4) We assume that the gating variable n, which governs potassium conductance, can be treated as a population-averaged variable. This allows us to describe the neuronal ensemble using a reduced set of collective (mean-field) variables.”

      Comparison with neural network simulations

      (Q4) The comparison the authors perform between the microscopic model and the neural mass is misleading, From what the authors wrote it seems that you are considering 4 variables for each neuron in the network model (this is unclear from how the model is written in Eq (9)), I guess one for the membrane potential, one for the gating variable and two for the potassium concentration. However, this is not the network model for which the neural mass has been developed, the neural mass has been obtained for a network made of N + 3 variables (N membrane potentials and 3 collective variables for gate, and potassium concentrations) this is a sort of mesoscopic network models, analogously to what done previously in references [R1,R3,R4] above and others. If the authors would compare their neural mass with this mesoscopic model the agreement among the two would be improved.

      We agree with reviewer’s observation and we now acknowledge this issue in the Results and in the Limitations. We have already modified the text to explicitly state that for the mean filed derivations n is treated as a collective variable and we have added the following statements:

      “Also note that the gating variable n is treated as microscopic in the neural network, while in the derivations for the mean-field it is considered as a mesoscopic and identical for the whole population. This is likely responsible for some of the discrepancies between the two modalities.”

      “Moreover, the discrepancy between the two modalities would have likely been smaller if for the neural network we also adopted a gating variable that is mesoscopic and identical across the spiking neurons, as in similar works [49–51]. However, here we demonstrate the validity of the mean-field approximation even for the more natural, microscopic representation of the gating variable in the neural network.”

      Comparison with in vitro experiments

      (Q5) Experiment -- The experiment is performed in vitro on the intact Hippocampus of mice between postnatal days P5-P7. It is known [R1] that neuronal activity at an early developmental stage is provided in the Hippocampus by a network primarily driven by synchronized GABA_A that provides an excitatory action and generates giant depolarizing potentials (GDPs) [R11]. However, GDPs have frequencies in the range of 1 Hz - 0.1 Hz, not matching the oscillation frequencies reported by the authors. I have several questions here:

      (E1) At this stage P5-P7 are the interactions among neurons essentially excitatory? Or not, please explain why, Are the oscillations reported by the authors somehow related to GDPs? The depolarizing action of GABAergic transmission and the presence of GDPs during early rodent brain development, as described by Ben-Ari and some others researchers, are characteristics commonly observed in ex vivo brain preparations, but are not evident under physiological in vivo conditions (see doi: 10.3389/fphar.2012.00065).

      In our preparation—intact mouse hippocampus—GABAergic synaptic transmission is not depolarizing. This is evidenced by the fact that inhibition of ionotropic GABA_A receptors with bicuculline triggers interictal-like discharges, which are routinely used as a model of epileptiform activity (see doi: 10.1016/j.nbd.2014.12.013). Therefore, in our experiments at P5–P7, neuronal interactions are not purely excitatory, and the observed low Mg2+ induced oscillations are not related to GDP.

      (E2) What is the nature of the oscillations reported by the authors in Figure 4 ? Which is their origin, please explain in the text of the paper clearly.

      The model of epileptic discharges presented in our study was first introduced over 20 years ago and has since become a well-established paradigm for screening potential antiepileptic drugs and research on the mechanism of epileptic seizure. A detailed description of this model can be found in doi: 10.1046/j.1460-9568.2002.02143.x, and its pharmacological properties are reviewed in doi: 10.1046/j.1528-1157.2003.19503.x. These references have now been added to the manuscript for clarity.

      We have added the following:

      “The model of epileptic discharges presented in our study was first introduced over 20 years ago [115] and has since become a well-established paradigm for screening potential antiepileptic drugs and research on the mechanism of epileptic seizure [116].”

      (E3) How exactly does the concentration of extracellular potassium ions change, this is not clear even in Methods, please clarify.

      [R11] Excitatory actions of GABA during development: the nature of the nurture Y Ben-Ari, Nature Reviews Neuroscience 3 (9), 728-739 (2002).

      We have now added a new Subsection in the methods explaining how we use Mg2+ variation to influence the external potasium variation.

      “The membrane of hippocampal neurons is equipped with N-methyl-D aspartate type glutamate receptors (NMDARs). These receptors have a very high affinity for glutamate and can, in principle, be activated by ambient glutamate present at low concentrations in the brain extracellular fluid (ECF).Under normal physiological conditions, this activation does not occur because extracellular magnesium ions (Mg<sup>2+</sup>) block the NMDAR channel at membrane potentials more negative than about –50 mV; this voltage-dependent block prevents receptor activation at rest. When extracellular magnesium is removed, the block is relieved, allowing NMDARs to be activated, leading to neuronal depolarization toward the action potential threshold [117]. In addition, as a divalent cation, Mg<sup>2+</sup> interacts with the negatively charged neuronal membrane, contributing to the stabilization of the resting membrane potential. Lowering extracellular magnesium concentration disrupts this effect, resulting in membrane depolarization [118]”

      “Consequently, magnesium removal not only facilitates NMDAR-dependent depolarization, but also directly depolarizes neurons. This depolarization increases the driving force for outward potassium currents through K<sup>+</sup> channels, meaning that variations in Mg<sup>2+</sup> can indirectly influence external potassium dynamics during neuronal activity.”

      (Q6) Lines 187-191 and Figure 4 -- The authors wrote : "In Figure 4.c we show the membrane potential and external potassium for a simulation of N = 3000 coupled HH-like neurons showing a similar behavior, although the parameters were modified to simulate shorter fluctuations for computational efficiency." This sentence is unclear. What is clear from Figure 4 is that the network simulations gave rise to collective oscillations on a completely different scale seconds with respect to minutes and also the profile of the potassium concentration has a clearly different evolution. From Figure 4 one can conclude that network simulations have nothing to do with the neural mass evolution and the experiment. I think the authors should better clarify and describe the results reported in Figure 4.

      We thank the reviewer for the observation. We have revised the relevant section of the manuscript to clarify the interpretation of Figure 4 and avoid any implication of quantitative matching. As stated in our response to Reviewer 1 (comment 6), the comparison is intended to highlight the shared qualitative structure across experimental data, the neural mass model, and the network simulation — specifically, the modulation of fast bursting by slow extracellular potassium fluctuations. The difference in timescale in the network simulation arises from rescaled parameters used for computational efficiency. We now explicitly state this and have updated the figure caption and accompanying text accordingly to reflect these points.

      (Q7) Why do the authors consider a purely excitatory network to describe the experimental results? What is the reason for this choice? Why they do not consider as usual balanced excitatory- inhibitory networks? Please clarify this point.

      We thank the reviewer for raising this point. We chose to model a purely excitatory network as a first step in isolating the role of extracellular potassium dynamics in generating population-level bursting. This allows us to focus on the ion-driven modulation mechanisms without introducing additional complexity from inhibitory feedback. Similar modeling choices have been made in previous studies of bursting and seizure-like dynamics (e.g., Gutkin et al.,), where inhibition is omitted to emphasize intrinsic or modulatory mechanisms. We acknowledge that incorporating inhibitory populations is an important next step for capturing a broader range of dynamics, but for the current study, the excitatory-only network provides a minimal and interpretable framework aligned with our focus.

      (Q8) By comparing Figures 4 (a) and (b) it seems that the bursting activity observed in the experiment and in the mean-field simulations seem quite different, originating from different mechanisms and bifurcations, Can the authors comment on this?

      We thank the reviewer for this important observation. We have reorganized the presentation of Figure 4 and revised the accompanying text to better clarify the nature of the comparison (see also our response to Reviewer 1, point 6). Our aim is not to claim that the experimental and simulated bursts arise from identical bifurcation mechanisms, but rather to highlight shared qualitative features — in particular, slow modulation of population activity by extracellular potassium. We now also comment on the potential role of more complex or noise-driven bifurcations (see Saggio et al. 2020) in shaping experimental bursting dynamics, which are not fully captured by the current deterministic model.

      Bifurcation analysis: emergent network states and multistability

      (Q9) This sub-section will gain interest by reporting simulations of the network and of the neural mass model presenting bistable dynamics.

      We agree with the reviewer that this would be an important addition, but we believe that it goes beyond the scope of this work (for the computational reasons among others) and it remains for future work. We have however updated the bifurcation analysis section.

      Limitations of the model

      (Q10) Lines 276- 280 -- I think that the parameters c+,c_,R+,R_ depend not only on the slow variables, potassium concentrations but also on the actual value of the gate variable n. This should be stressed.

      We thank the reviewer for this helpful observation. We agree and have clarified in the revised manuscript. This reflects the mean-field assumption that n is treated as a collective variable, and we now make this dependency explicit in the text.

      “Furthermore, the parabola coefficients c_-,c_+, R_-, R_+ were fixed as constants, however, these coefficients could be made functions of the slow variables and the gating variable, which might unveil new dynamical regimes and extend the validity of the thermodynamic limit beyond the regimes described in this work. Also, in the case of constant values, an in-depth exploration of the parameter space is required to fully characterize the model and its bifurcation structure.”

      (Q11) The authors wrote: " Other limiting assumptions are the moment closure condition (19) and the assumptions that the functions (3) averaged across the neuronal population can be expressed as functions of the average membrane potential V and gating variable n (which is only true in the cases where the functions (3) can be reasonably approximated as linear functions in a range of V and n." Apart from that a parenthesis is lacking, I think that this last aspect has been already taken into account when performing the fit with 2 parabolas to the sum of the currents, or not? In case, please specify.

      We thank the reviewer for catching the missing parenthesis — this has been corrected in the revised manuscript. Regarding the modeling point: the two-parabola fit applies specifically to the membrane potential dynamics and captures the nonlinear dependence of the total current on V (eq.16). In contrast, the moment closure assumption involves approximating averages of nonlinear functions of both V and n, such as those appearing in the gating dynamics (e.g., n∞(V)). This is not directly accounted for by the parabola approximation, but is handled separately via the mean-field approximation of G^n as a function of the average variables (eq.15).

      (Q12) A limitation that should be stressed is that the authors in the neural mass model consider the gate variable and the potassium concentrations, as global variable equal for all neurons, and where n depends on the mena membrane potential, to write that the moment closure (19) is a limiting assumption is honestly too clear, please be explicit here.

      We have now the following two statements:

      “These slow variables are in addition considered to be mesoscopic, meaning they are identical for every neuron in the population.”

      “In our mean-field model, the gating variable (n) is treated as a global population variable, evolving deterministically as a function of the average membrane potential. Therefore, ⟨n(t)⟩ corresponds to the collective gating variable assumed to be shared by all neurons, and is not computed by averaging distinct microscopic (n<sub>i</sub>) values.”

      Discussion

      (Q13) The authors could discuss in this section the further biological ingredients they can introduce in their neural mass based on the previous works [R1-R9] that have already shown how to include plastic synapses, random connectivity, noise, adaptation, spike-timing-dependent plasticity, etc and which of these ingredients they consider more relevant for the whole brain dynamics.

      In order not to repeat the same statements from the Introduction, we have now addded the following sentence:

      “This approach, taking into account key biophysical details, offers a first step in considering the role of the glia in neural tissue excitability. Following this direction, other ions, such as calcium should be taken into consideration, as well as other effects such as plastic synapses, random connectivity, noise, adaptation, spike-timing-dependent plasticity, as already discussed in the Introduction.”

      (Q14) The authors should also discuss why they limited their analysis to purely excitatory networks, and what would change by including excitatory-inhibitory interactions in each single mass and across neural masses, if this makes sense or not.

      As stated in our response to Q7, we chose to focus on purely excitatory networks as a first step to isolate and study the core role of extracellular potassium dynamics in driving bursting behavior. This modeling choice allows for a minimal system where the interaction between intrinsic ionic mechanisms and network coupling is most transparent.

      We also note that excitatory and inhibitory effects can be modeled within the same formalism by adjusting the synaptic reversal potential — for example, $E_{syn}=0$mV for excitatory, and $E_{syn}=-80$mV for inhibitory interactions. Including inhibitory populations would introduce additional complexity and richer dynamical regimes (e.g., oscillatory instabilities, balance states), which are certainly of interest but beyond the scope of this study.

      Materials and Methods

      (Q15) Fig.2 - I think a plus is lost in panel (c) where it should be [K+bath];

      Thank you. We corrected the figure.

      (Q16) Caption of Figure 2- the authors wrote: "In the case where the derivative of the membrane potential is zero for V > V ⋆ (e.g., if the cubic function is shifted up by adding a constant current to the membrane potential derivative), the population is described by the red distribution in the steady state, and the continuity equation is governed by the negative parabola equation." This sentence is unclear, the authors mean in the case where the derivative of the membrane potential crosses zero at V > V*? Please clarify.

      We thank the reviewer for pointing this out. Yes, we refer to the case where the membrane potential derivative crosses zero at a point V>V∗. We have clarified this in the revised figure caption.

      (Q17) Lines 558-562 -- Eqs (6) and (7) are examples of unnecessary complications of which this manuscript is full of. Since the authors do not consider any synaptic dynamics and homogenous (equal) couplings, these equations are not needed, I strongly recommend removing Eqs (6) and (7) and limiting to the expression reported in Eq (8), which indeed should also be corrected see next remark.

      We appreciate the reviewer’s concern regarding clarity. As mentioned in our response to Reviewer 1, the inclusion of Eqs. (6) and (7) was intentional and serves a pedagogical purpose — to present the general structure of the network interactions before introducing simplifying assumptions. While we agree that Eq. (8) suffices for the simulations considered in this manuscript, we believe that showing the more general form helps clarify the model’s extensibility, for instance to cases with heterogeneous coupling or synaptic dynamics.

      (Q18) Eq (8) - line 562 - Since the authors assume no synaptic evolution, i.e. instantaneous post-synaptic potentials, they can clarify that Eq (8) represents the population firing rate that later will be one of the fundamental variables of the neural mass model and call it r, as in the following. Furthermore, $s_i$ does not depend on the neuron index $i$ in a fully coupled network with homogenous coupling, as in the present case, this quantity is the same for all neurons. Please drop the index and call it r since it is the population firing rate.

      We thank the reviewer for this useful suggestion. We now clarify in the text that under the assumptions of all-to-all homogeneous coupling and no synaptic dynamics, s_i is identical for all neurons and can be interpreted as the population firing rate r. This connection is made explicit in the revised manuscript.

      “Under the assumption of instantaneous synaptic transmission and homogeneous all-to-all coupling, the synaptic activation variable (s<sub>i</sub>) is the same for all neurons and corresponds to the population firing rate, which we denote by (r)”

      (Q19) Line 564-567 - Here the network model is incomplete, it is not sufficient that the authors report the evolution equation for the membrane potential Eq (9). They should report the evolution equation for the gate variable n and for the potassium concentration as done in Eq (1). This request is fundamental because it is unclear from the present formulation which are the variables that are microscopic (associated with the single neuron evolution) and which are global (common to all the neurons). This is a fundamental aspect and it should be clarified. I guess that n will depend on the neuron index $i$, while the potassium concentration it is unclear how the authors will consider them, global or local. I guess that the internal density should depend on the neuron index $i$ or not ? Anyway, I would like to know exactly which network model has been simulated e.g. to obtain the results reported in Figure 3.

      We thank the reviewer for this essential clarification request. In the revised manuscript, we now explicitly state the full network model, including the evolution equations for the gating variable n_i and potassium variables. While in some simulations we consider the full microscopic model involving 4N variables (where each neuron has its own V_i ,n_i ,Δ[K+]int_i ,[K+]g_i), for the mean-field reduction and mesoscopic comparisons we assume that the gating and potassium variables are shared across neurons. This assumption is consistent with prior work (e.g., Chen & Campbell) and is biophysically justified in the case of potassium due to its fast spatial equilibration in extracellular space. We also now mention this explicitly in the Limitations.

      (Q20) Continuity equation - Lines 568 - 597 - This part can be largely simplified and rewritten, as a matter of fact, the authors consider the gate variable n, the potassium concentrations as global (collective variables) depending on mean field values of <V> they can directly start from eq 20, by stating that they assume that the other variables (n, $\Delta[K^+]_{int}$, $[K^+]_g$) are collective variables, common to all the neurons, and that depends only on mean field variables as <V> or r. This has been done in many previous cases since the Ott-Antonsen Ansatz can be applied whenever the potential evolution is driven by quadratic terms and in the presence of mean field variables, the first indication of this was reported in 1993 by Watanabe and Strogatz for phase oscillators :

      [R12] Watanabe, Shinya, and Steven H. Strogatz. "Integrability of a globally coupled oscillator array." Physical review letters 70.16 (1993): 2391.

      Anyway, this approach has been previously employed to derive a neural mass model for networks of QIF neurons in the presence of various further neuronal variables (ranging from slow currents to plastic evolution of the couplings) describing more biologically realistic situations, see references [R1-R7] above. I strongly encourage the authors to reformulate their approach in a simpler and clearer manner, particularly interesting is for them the article [R6] by Guerriero et al, the authors examine exactly the same model as in Ref [95] [Chen, L. & Campbell, S. A. Exact mean-field models for spiking neural networks with adaptation. Journal of Computational Neuroscience 50 (4), 445-469 (2022)]. However, they solve the problem in a much more simple way, I encourage the authors to follow this approach.

      We thank the reviewer for the constructive suggestion. We acknowledge that, under the assumption that n, Δ[K+]int , and [K+]g are collective variables shared across the neuronal population, one could directly begin from Eq. (20) and proceed using the simpler approaches found in Guerriero et al. [R6] or related works [R1–R7]. However, we chose to retain the Chen & Campbell formalism, with additional clarification regarding the mesoscopic nature of the gatin variable, as it explicitly highlights the key approximations used in the derivation, which may be beneficial for readers seeking to extend the method. See also general response to reviewer 2 at the beginning.

      (Q21) Eq (26) -- I do not think the authors can estimate explicitly <n(t)> from the equation (26), as they do for the mean membrane potential and the firing rate. This is just a formal expression representing a collective variable, I do not think that <n> will coincide with the average of the values of n_i for each neuron. Please discuss this point, and in this case show that <n> indeed coincides with the average of all of the values of the single neuron gate variable n_i.

      We thank the reviewer for raising this important point. We agree that Eq. (26) is more formal than operational, as ⟨n(t)⟩ is not directly derived from the continuity equation in the same way as ⟨V⟩ or the firing rate r. Rather, it reflects our mean-field assumption that the gating variable evolves as a collective population-averaged quantity, governed by the dynamics of the average membrane potential. In our formulation, n is treated as a global variable shared across neurons, and thus ⟨n(t)⟩ effectively is the gating variable in the neural mass model — rather than the result of averaging heterogeneous n_i. We have clarified this distinction in the text to avoid suggesting that Eq. (26) provides an explicit estimate of microscopic gating dynamics.

      “Unlike the mean membrane potential ⟨V⟩ and the firing rate (r)>, which can be explicitly derived from the continuity equation under the Lorentzian assumption, the expression for ⟨n(t)⟩ in Eq. (26) is formal. In our mean-field model, the gating variable (n) is treated as a global population variable, evolving deterministically as a function of the average membrane potential. Therefore ⟨n(t)⟩ corresponds to the collective gating variable assumed to be shared by all neurons, and is not computed by averaging distinct microscopic (n<sub>i</sub>) values.”

      (Q22) Mean-field dynamics for the gating variable - All this sub-section is in my opinion not useful, if the authors assume from the beginning that <n(t)> is a global variable. Indeed in the end they write for <n(t)> the evolution equation Eq (30) which is the same equation as for the single neuron gate variable (1) but for the mean values of n and <V>. I suggest removing this sub-section.

      We thank the reviewer for this suggestion. We agree that, under the assumption that n is a global collective variable, the resulting equation for ⟨n(t)⟩\langle n(t) \rangle⟨n(t)⟩ is equivalent in form to the single-neuron gating equation, driven by the average membrane potential. However, we chose to retain this subsection to explicitly demonstrate how the gating dynamics enter into the mean-field formulation, especially for readers less familiar with this type of reduction. This step also mirrors the structure of the derivation used for other state variables in the model and maintains clarity for potential extensions where n may not be strictly global.

      (Q23) Line 696 - here an equation reference is lost.

      Thank you for pointing this out. We have corrected the text and restored the missing equation reference in the revised manuscript.

      (Q24) Eqs (36) -(37) -- Since the variables r and x entered in Eq (36) are essentially the same as Eq (25), apart from a constant R/pi, the use of two different names complicated in a useless manner an already complicated expression, Please decide to use everywhere r or x and then proceed consequently this applies also to Eq (37). This will also allow us to rewrite the equation in x or r in a more compact form.

      As noted in our response to Reviewer 1, point 14, we have revised Eq. (37) to ensure consistency in notation by replacing x with r throughout.

      (Q25) Eq (37) - This equation is written in a manner that is not careful enough, apart from that the authors are passed now from (x,y) to (pi*r/R,V) , therefore they should substitute everywhere x with r. Furthermore, the equation for the derivative of V is confusing, the authors should use the same approximate expression employed in eq (36) that makes explicit the quadratic dependence on V itself, otherwise, I believe that the equation is incorrect.

      In the same response to Reviewer 1, point 14, we also clarified the expression for \dot{V} in Eq. (37), we reintroduced the full current-based formulation (as in Eq. 16), reversing the quadratic approximation used earlier. This is now explicitly stated in the text, and we have improved the equation presentation to avoid confusion.

      (Q26) Eq (37) below line 708 - From this expression, it is clear that the gate variable n and the potassium variables are ruled exactly by the same equations as for the single neuron Eq (1) and that the Lorentzian Ansatz enter only in the rewriting of the evolution of the membrane potentials of the neurons in the network. In the end, the authors are doing exactly the same approximation made by many other authors [R1-R7], that these variables are collective, i.e. they are the same for all neurons, and in particular n=n(V) is a function of the mean membrane potential V. The mean field model that the authors derive corresponds to a microscopic model where the single neurons are heterogenous only in the intrinsic currents $\eta_i$, but they are all driven by collective variables, like n(V) and the potassium variables that are identical for all neurons. This should be clarified.

      We agree with the conclusion by the reviewer, and as seen through the previous responses, we now explicitly acknowledge the fact that n and the two slow variables are considered as a mesoscopic variables for the mean-field derivation, while for the spiking network, n remains microscopic.

    1. eLife Assessment

      This work by Pyne and Pandey et al. addresses DNase X (DNase1L1) activity at the macrophage phagocytic cup, using an innovative imaging approach that couples visualization of cup formation to spatially resolve DNA degradation. The methodology is technically sound, and the central finding that DNA digestion begins prior to phagolysosomal maturation is considered well supported, though some mechanistic claims may benefit from further evidence and more cautious framing. Overall, the study is solid and provides a valuable framework for investigating early events at the phagocytic cup that may shape responses to pathogens and inflammatory disease.

    2. Reviewer #1 (Public review):

      Pyne and Pandey et al. report the observation of early DNA degradation at the phagocytic cup during macrophage engulfment. Using an elegant experimental system that combines actin staining to visualise cup formation with direct monitoring of DNA degradation, the authors identify rapid recruitment of the membrane-bound nuclease DNase X (DNase1L1) to nascent phagocytic cups. This recruitment occurs within minutes of cup formation, is independent of DNA presence at the substrate, and appears to originate from intracellular membrane structures rather than from the extracellular environment. The results support the conclusion that DNase X activity is present at the phagocytic cup and that DNA digestion can begin prior to phagolysosomal maturation.

      The study is technically strong. The experimental system is clean, specific, and allows precise spatial and temporal detection of DNA degradation. The imaging-based approaches are carefully executed and enable convincing visualisation of DNase X recruitment and activity. The use of an alternative substrate beyond the primary SNS system strengthens the core observation, and the data broadly support the authors' central claim.

      However, several limitations temper the physiological interpretation. The system relies largely on short, free DNA substrates, leaving open how efficiently DNase X processes more complex or physiologically relevant DNA structures, such as nucleosome-bound DNA or neutrophil extracellular traps (NETs). It remains unclear whether DNase X deficiency would alter macrophage responses to larger nucleic acid structures, influence engulfment efficiency, or modify downstream inflammatory signalling pathways such as TLR9 or STING activation. Moreover, the experimental setup prevents full phagocytic cup closure, potentially prolonging DNase activity compared with physiological phagocytosis, which typically proceeds rapidly to cargo internalisation. For example, the peak signal observed in Figure 5 occurs approximately 90 minutes after phagocytic cup formation, a time point at which many phagocytic cups would be expected to have already closed under physiological conditions. Additional work using fully engulfed cargo in more physiological contexts would clarify whether early DNase X activity meaningfully contributes to overall DNA clearance kinetics.

      Mechanistically, the signal that triggers DNase X recruitment remains unresolved. Although actin rearrangement was excluded as the primary driver, the upstream cues that direct DNase X-containing membrane structures to the forming cup are not yet defined.

      In the broader context, early DNase X activity at the phagocytic cup could represent an additional safeguard against inflammatory signalling by limiting extracellular or surface-associated DNA before phagolysosomal degradation by DNase II. This mechanism may be particularly relevant in settings where DNA fragmentation before engulfment is incomplete, such as necroptosis or NET formation. Determining whether DNase X deficiency exacerbates inflammatory responses, alters DNA clearance efficiency in vivo, or contributes to immune pathology will be critical for establishing its physiological and disease relevance.

      Overall, this is a compelling study that introduces a novel concept of pre-phagolysosomal DNA digestion. The conclusions are well supported within the in vitro system used, but further investigation using diverse DNA substrates and physiologically relevant models will be required to fully define the impact of this mechanism on immune regulation and disease.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript presents an elegant and innovative imaging approach to visualize DNase activity at the interface between macrophages and extracellular substrates. The platform is technically strong and enables the study of localized DNA degradation with high spatial resolution. The work is of clear interest and provides a useful framework to investigate how immune cells process extracellular DNA. However, several aspects of the mechanistic interpretation and conceptual framing would benefit from clarification.

      Strengths:

      (1) The study introduces a creative and well-designed imaging platform that allows visualization of localized DNase activity at cell-substrate interfaces.

      (2) The approach is technically robust and represents a valuable tool that could be broadly useful to the field.

      (3) The experiments are thoughtfully designed and address an important question regarding how immune cells interact with extracellular DNA.

      (4) The work opens interesting avenues for studying DNA processing in contexts such as infection and inflammation.

      Weaknesses:

      While the experimental approach is strong, several key conclusions rely on interpretations that would benefit from further clarification:

      (1) First, the conclusion that DNaseX is recruited to phagocytic cups from the "cytoplasm" appears conceptually imprecise. Given that DNaseX is a membrane-anchored protein, it is unlikely to exist as a freely soluble cytoplasmic pool. A more plausible interpretation is that DNaseX is supplied from intracellular membrane compartments. This interpretation would also be more consistent with the data showing dependence on a membrane anchor.

      (2) Second, the interpretation that actin polymerization is not required for DNaseX recruitment raises concerns. Phagocytic cup formation is known to depend strongly on actin dynamics, and it is therefore unclear whether the structures observed under actin inhibition represent fully formed functional cups or partial cell-substrate contacts. This distinction is important for interpreting recruitment versus activity, particularly since enzymatic activity is reduced under these conditions.

      (3) Third, the identification of DNaseX as the main nuclease responsible for the observed activity is not fully resolved. The conclusions rely primarily on gene silencing and staining approaches, but the specificity of these strategies relative to other nucleases is not addressed. It therefore remains possible that additional enzymes contribute to the observed activity.

      (4) Finally, the interpretation of the biofilm experiments may be overstated. While the data clearly show localized DNA degradation in contact with macrophages, it is not fully established that this process depends specifically on phagocytic cup structures. An alternative explanation is that membrane-associated DNase activity more generally mediates this effect. In addition, the physiological relevance of this mechanism would benefit from further discussion.

      Overall, the study is technically strong and introduces a valuable methodology, but several central conclusions are only partially supported by the current data and would benefit from more cautious interpretation and clearer conceptual framing.

    1. eLife Assessment

      This important study combines single-molecule imaging and CUT&TAG to address the molecular mechanism underlying the differentiation process that initiates the formation of red blood cells in the bone marrow. The authors provide evidence that the transcription factor GATA2 transiently associates with a new set of genomic loci early in the differentiation process before it is replaced by GATA1. Together, the experiments across three biological systems are solid, but they could benefit from additional details and controls to strengthen the conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      During erythroid differentiation, hematopoietic progenitors relinquish multipotency and activate lineage programs. The switch from GATA2 to GATA1 is particularly important in this process, yet GATA2 chromatin‑binding kinetics remain undefined. The authors investigated GATA2-chromatin interaction dynamics during erythroid differentiation in three different cell systems using single‑molecule live‑cell imaging, and they also used CUT&Tag to profile GATA2 chromatin occupancy.

      By single‑molecule imaging, the authors report two interaction modes for GATA2: short‑lived (<1 s) and long‑lived (>5 s) binding. The proportion of long‑lived molecules, the number of binding events, and the duration of long‑lived binding change (or are maintained) during differentiation. Notably, long‑lived chromatin engagement by GATA2 increases during early erythroid differentiation and decreases at the late stage. CUT&Tag identifies regulatory elements selectively occupied by GATA2 during the early transition stage. Together, these results support a model in which transcription factor kinetics form a dynamic chromatin‑engagement profile that characterizes the GATA2‑to‑GATA1 transition.

      Strengths:

      (1) Characterizing transcription‑factor binding kinetics during the GATA2->GATA1 transition addresses a fundamental mechanism in erythroid differentiation.

      (2) Combining single‑molecule live imaging with CUT&Tag provides both dynamic and locus‑specific perspectives.

      (3) Single-molecule analysis across three different cell systems strengthens the potential generalizability of the findings and highlights biological variability.

      Weaknesses:

      I agree that single‑molecule imaging is a powerful approach for investigating GATA2 kinetics, but the single‑molecule data are the most important part of the paper and need improvement. The analyses focus on three measures: (i) duration of long binding, (ii) proportion of short‑ and long‑binding molecules, and (iii) total binding events. However, several methodological and control issues limit confidence in the kinetic interpretations. The authors should address the following major concerns.

      (1) Two binding states: justification and controls

      The authors propose two states of GATA2 binding. Are there only two states? Studies that separate short‑ and long‑lived binding (e.g., Chen et al., 2014, PMID: 25342811) address two states of transcriptional factors very carefully. Some long‑binding duration distributions here are very long‑tailed (e.g., Figure 2D middle), suggesting a possible third state. The authors must explain how they determined that two states provide the "best fit" to the data and how they classified "short" versus "long" binding.

      Controls should be included for long‑lived and short‑lived binding (e.g., histone proteins, HaloTag‑NLS, or a binding‑deficient GATA2 mutant) as in other studies. These controls are essential to exclude alternative explanations (see points below).

      (2) Exclude photophysical and focal‑plane artifacts

      The authors should exclude contributions from (i) photobleaching, (ii) blinking, and (iii) Z‑axis motion (disappearance from the focal plane). Although photobleaching correction is mentioned in the Methods, no details are provided. Describe and quantify the photobleaching correction and demonstrate that it was applied across all cell types and conditions.

      Some spots in the supplementary movies appear to blink or to move substantially between frames. Provide analyses or controls that distinguish true dissociation events from photophysical blinking/bleaching or axial motion.

      (3) HILO illumination and nuclear region sampled

      HILO is powerful but sensitive to illumination angle: slight changes sample different nuclear regions (e.g., nuclear interior versus periphery). The nuclear periphery is enriched in heterochromatin and may bias binding statistics. Explain how the authors controlled the HILO angle and confirmed that comparable nuclear regions were imaged across cells and conditions.

      (4) Quantification of event counts and long‑binding durations

      The number of binding events and measured long‑binding durations are strongly affected by imaging conditions (labeling/staining, bleaching, nucleus size, cell cycle state, focal plane, spot detectability, etc.). Imaging clarity appears to differ among cells/conditions in the supplementary movie. Provide more careful analysis describing how these variables were controlled or corrected for, and assess the sensitivity of results to choices in detection and tracking parameters.

      (5) Evidence that spots are single molecules

      The authors state that spots represent single molecules but do not provide supporting evidence. Spot brightness varies considerably in the movies. Brightness differences may reflect axial position. Provide evidence supporting single‑molecule assignment (e.g., single‑step photobleaching traces, brightness distributions compared to a known single‑molecule control, or photon count analysis).

      (6) Description of spot‑analysis pipeline

      The manuscript lacks a sufficient description of the spot‑analysis method. I reviewed the STRAP pipeline paper cited (Haque and Coleman 2025 bioRxiv) and the GitHub code, but the Methods in the current manuscript should include a detailed STRAP pipeline. This would enable readers to evaluate and reproduce the analyses.

      (7) Differences among cell systems

      The three cell systems yield notably different results (e.g., Figure 2C vs 4C and Figure 2D/3D vs 4D). Provide a more detailed explanation for these differences and discuss how biological variability, technical differences, or imaging biases might account for the discrepancies.

    3. Reviewer #2 (Public review):

      In this study, the authors address the molecular mechanism underlying the transcriptional changes during erythroid differentiation from hematopoietic progenitor cells. The authors combine single-molecule live cell imaging and CUT&RUN to analyze the chromatin binding properties of the GATA2 transcription factor prior to and after initiation of differentiation into the erythroid cell lineage. Using three distinct cellular systems, the authors demonstrate that the chromatin binding of GATA2 is transiently increased early in the differentiation process, as evidenced by increased chromatin binding residence time and the emergence of new genomic binding sites identified by CUT&RUN. The strength of the study lies in the combination of single-molecule imaging, which reports on binding dynamics but is agnostic of the binding site, with CUT&RUN, which reports on the binding sites but does not provide dynamic information. The authors clearly demonstrate that chromatin binding of GATA2 is altered early in the differentiation process and is later displaced as cells switch to expression of GATA1, which has been previously observed. The use of three distinct cell lines, in particular the GATA2-SNAP mouse model, is a strength in principle; however, the results are not fully consistent between the different cell systems. A key difference is that the G1E-ER4 and HPC7 cell line models express HaloTagged GATA2 in addition to the endogenous GATA2 protein. The authors go through great lengths to control GATA2-HaloTag expression levels, but they use polyclonal cell lines and do not analyze expression levels of the GATA2-HaloTag transgene, which is a key variable in interpreting their experimental results. Finally, a key variable determined in their single-molecule analysis is the number of binding events observed during the distinct differentiation changes. The number of binding events observed is influenced by the expression level of the tagged protein, which in turn is controlled by the Shield-1 ligand, and the fraction of molecules labeled with the HaloTag ligand. Since transgene protein levels and the labeling efficiency were not determined, it is hard to assess how reliable the measurements of the number of binding events are across all cell lines.

      To address the weaknesses summarized above the authors could take the following steps:

      (1) Determine the expression levels of the GATA2-HaloTag transgene over the course of differentiation under the conditions used for single-molecule imaging. This will not only allow them to determine the expression of the transgene but also the endogenous untagged protein with which the GATA2-HaloTag fusion proteins compete for binding sites.

      (2) To determine the fraction of molecules labeled during imaging, the authors could carry out a titration of the HaloTag ligand and compare the amount of labeled protein under single-molecule imaging conditions to that of saturating labeling of the HaloTag. This approach will ensure that the number of labeled molecules per cell is comparable across experimental conditions and allow the authors to draw more solid conclusions regarding the number of binding events.

      (3) The analysis of residence times using single-molecule imaging requires robust single-particle tracking without gaps or interruptions of trajectories. The authors should show images of their particle trajectories to demonstrate that their tracking is robust. Or even better, movies superimposing the trajectories onto the imaging data.

    4. Reviewer #3 (Public review):

      Hobbs et al. use live-cell single-molecule tracking (SMT) of HaloTag- and SNAP-tagged GATA2 combined with CUT&Tag chromatin profiling to examine how GATA2 chromatin engagement evolves during erythroid differentiation. Across three complementary systems, G1E-ER4 cells, HPC7 cells, and primary bone marrow progenitors from a new Gata2-SNAP knock-in mouse, they report a transient strengthening of long-lived GATA2 chromatin binding at the "Early" (2 h) erythroid stage, manifested either as increased residence time (G1E-ER4) or expansion of the long-lived bound fraction (HPC7, primary cells). CUT&Tag identifies 1,167 Early-restricted GATA2 peaks partitioning into GATA2-only (promoter-proximal, GATA/RUNX motifs) and GATA2+GATA1 co-bound (distal, GATA/E-box motifs) subclasses. The authors propose that this kinetic phase represents a previously unappreciated dimension of the GATA switch.

      This is a strong study with a genuinely novel finding-the non-monotonic kinetic behavior of GATA2 during erythroid priming, supported by complementary measurements in three biological systems. The issues below are largely clarifications, additional analyses of existing data, and modest refinements to the discussion. With these addressed, the manuscript will make a valuable contribution. I recommend a minor revision.

      Specific points:

      (1) Clarify the photobleaching correction and report per-cell bleach lifetimes.

      The long-lived residence time claim in G1E-ER4 cells depends on careful accounting for photobleaching, which the Methods indicate was done via a right-censoring model. For reviewer and reader confidence, the authors should report the per-stage (or per-cell distribution of) photobleaching lifetimes and the photobleach-corrected residence time values alongside the apparent values in Figure 2D. If feasible, including a brief supplementary analysis with an H2B-Halo or similar long-lived control under matched conditions would further solidify the quantitative claims. This is an analysis of existing data and should not require new imaging.

      (2) Unify or explicitly discuss the mechanistic differences across systems.

      The three systems show qualitatively different signatures: residence time change in G1E-ER4, bound fraction expansion in HPC7, and primary cells. The authors currently group these under "enhanced engagement," but these signatures imply different underlying mechanisms (koff decrease vs. increased kon or increased specific-binding-competent pool). The Discussion partially addresses this by noting engineered vs. native differences, but a more explicit framing in both Results and Discussion would help readers. Specifically, reporting an on-rate proxy (for example, binding events per unit time normalized to detectable molecule number) alongside koff would let readers see how the mechanistic pieces fit together. I do not think this changes the central message; it sharpens it.

      (3) Per-cell GATA2 concentration would strengthen the "uncoupling" claim.

      A central claim of the Figure 6 model is that chromatin engagement is uncoupled from protein abundance. The ectopic Shield-1 stabilization system is a reasonable design choice, but quantifying total nuclear GATA2-Halo signal (for example, from the pre-bleach frame or a brief high-power acquisition) on a per-cell basis across stages would directly support the interpretation. For the primary cells, where the biological claim is strongest, a western blot or quantitative immunofluorescence on the flow-sorted populations would make the uncoupling argument much more defensible. I recognize this may be one additional experiment, but it is a high-value one.

      (4) Additional single-cell distribution analysis.

      Figure 1E and Figures 2 to 4 show substantial cell-to-cell heterogeneity, and the Early populations in particular look potentially bimodal. Given that the authors cite Wheat et al. and Palii et al. on probabilistic hematopoietic transitions, a brief supplementary analysis using distribution-based statistics (K-S test, or mixture model) rather than, or alongside, mean-based ANOVA would align the analysis with this conceptual framing and may reveal whether the Early state represents a subpopulation transition rather than a uniform shift. This is purely an analysis of existing data.

      (5) Quantitative integration of CUT&Tag with SMT.

      The manuscript presents SMT and CUT&Tag as complementary but does not attempt to quantitatively connect them. A back-of-the-envelope calculation of whether a 21% increase in residence time (G1E-ER4), or the fraction expansion in other systems, is consistent with the acquisition of the 1,167 Early-restricted sites, given plausible site affinities, would substantially strengthen integration. Even if the calculation is approximate, framing it explicitly would help readers appreciate that the two datasets reinforce each other.

      (6) Short-lived kinetic interpretation and tracking parameters.

      The 1.5 s gap allowance in tracking is long relative to the 0.55 to 0.73 s short-lived residence times reported in primary cells (Figure Supplement 1F), which could affect the interpretation of the "slowing of target search" claim. A brief sensitivity analysis with tighter gap parameters in the supplement would reassure readers that this effect is robust. Additionally, please clarify how the inferred slowing of search, which should reduce kon, is reconciled with the increased number of binding events per cell at the Early stage.

      (7) CUT&Tag peak definition.

      The Early-restricted peak set is defined by presence and absence at q less than 0.01, which can be sensitive to near-threshold peaks. Please report either (a) the CUT&Tag signal intensity distribution at the 1,167 sites across all three stages as a quantitative scatter or density plot, beyond the heatmap in Figure 5C, or (b) the result of a differential binding analysis (for example, DESeq2 on read counts in a union peak set) as a supplementary confirmation. Please also state the number of CUT&Tag replicates per stage and the overlap of Early-restricted sets across replicates.

      (8) Knock-in mouse validation.

      The Gata2-SNAP allele is a valuable new tool, and it would benefit from slightly more quantitative validation in the supplement. A brief characterization of basic hematopoietic parameters in homozygotes (CBC, LSK/HSPC frequencies, or colony assays) would confirm that the tagged allele is truly physiological and would serve the community that will want to use this mouse going forward. If this has been done, please include it; if not, a statement about what was checked would suffice.

    5. Author response:

      We are writing to provide our provisional response to the public reviews. We note that the reviewers’ comments focus primarily on strengthening technical rigor and quantitative interpretation. We have designed the planned revisions to directly address the reviewers’ major concerns and to strengthen the study’s evidentiary basis. We plan to submit a revised manuscript for the final Version of Record.

      For clarity, we summarize below the major new experiments and analyses that address the reviewers’ primary concerns:

      (1)Validation of Tracking Parameters (Reviewers 1 & 3): We will re-analyze our single molecule tracking data with tighter gap-time allowances (0 seconds) to demonstrate the robustness of our interpretations of short- and long-lived kinetics. We will also generate a supplementary movie with binding trajectories superimposed directly on detected molecules to visually confirm tracking robustness.

      (2) Photobleaching & Two-State Controls (Reviewers 1 & 3): We will report per-cell photobleaching lifetimes derived from our global fluorescence decay. To strengthen this analysis, we will include supplementary measurements using a H2B-HaloTag control under matched imaging conditions and perform single-molecule tracking of GATA2 zinc-finger deletion mutants (N-terminal, C-terminal, and double) as a binding-deficient functional control.

      (3) Protein Expression & Labeling Efficiency (Reviewers 1 & 2): To address concerns about transgene expression and competition with endogenous proteins, we will quantify Halo-GATA2 levels in G1E-ER4 and HPC7 cells and SNAP-GATA2 levels in primary cells using standardized titration methods with established Halo-CTCF and SNAP-RPB1 reference systems.

      (4) Integration of SMT and CUT&Tag (Reviewer 3): We have conducted a quantitative foldchange analysis of our existing CUT&Tag dataset to complement our single-molecule kinetics.

      However, as detailed in our specific response below (R3 point 5), we emphasize that directly integrating population-level genomic occupancy measurements with single-cell kinetic measurements is not straightforward. We will therefore frame the relationship between these datasets as a conceptual consistency check rather than a strict quantitative integration. This quantitative analysis supports and refines the Early-restricted peak set, identifying a high confidence strict subset consistent with the broader presence/absence-defined set described in Figure 5 of the manuscript (see Author response images 1–3 and our response to R3 point 7).

      (5) Characterization of the GATA2-SNAP Mouse (Reviewer 3): We have characterized hematopoietic populations in the homozygous knock-in mouse, including lymphoid (CD3<sup>+</sup>/CD4<sup>+</sup>/CD8<sup>+</sup>/B220<sup>+</sup>/CD19<sup>+</sup>), myeloid (CD11b<sup>+</sup>/Gr1<sup>+</sup>), and erythroid (Ter119<sup>+</sup>) compartments. These data, presented in Author response image 4, indicate that normal mature hematopoietic output is preserved across genotypes. Statistical caveats are described in the corresponding figure legend and in our response to R3 point 8.

      Public Reviews:

      Reviewer 1 (Public review):

      (1) Two binding states: justification and controls

      The authors propose two states of GATA2 binding. Are there only two states? Some longbinding duration distributions here are very long-tailed (e.g., Figure 2D middle), suggesting a possible third state. The authors must explain how they determined that two states provide the best fit and how they classified short versus long binding. Controls should be included for long-lived and short-lived binding (e.g., histone proteins, HaloTag-NLS, or a binding-deficient GATA2 mutant).

      Agreed in part; we will attempt the requested binding-deficient control using existing GATA2 deletion constructs, complemented by GRID and H2B-HaloTag controls.

      We will clarify that the two-state framework is an operational model rather than a claim that GATA2 can occupy only two physical states. This approach is widely used in SMT studies of chromatin-associated transcription factors and transcription machinery (Gebhardt et al., 2013; Liu et al., 2014; Hansen et al., 2017; Kenworthy et al., 2022). In particular, Ling et al. (Science, 2026) recently used two-exponential survival-probability fitting across 58 Halotagged transcription-associated proteins to distinguish transient and stable chromatin-binding populations, while explicitly noting that the simplified two-state model provides a tractable framework even when the underlying physical behavior may be more heterogeneous.

      We agree that our current two-state model may under-represent the diversity of GATA2 chromatin-binding populations in single cells. However, even within this simplified framework, the existing analysis already indicates increased upper-tail dispersion of kinetic measurements (e.g., residence time and/or percentage of stable events) at the single-cell level in early erythroid cells. To support the goodness-of-fit metrics from our two-state fitting, as Reviewer 3 recommends, we will provide a supplementary table containing confidence intervals for the rate parameters and an F-test metric describing the differences between one- and two-state fits.

      To determine whether additional binding states exist, we will perform GRID (Genuine Rate Identification from Distributions), which does not bias the model toward a particular number of states and, in our experience across multiple proteins, yields fits with 3-5 binding populations. However, we have found that in many cases, GRID requires aggregating binding events from multiple cells to achieve consistently robust fits for the populations of relatively rare, long-lived (>~30 sec) binding events. Therefore, GRID will assess whether additional populations exist, but we will lose the ability to analyze changes in the cell populations at the single-cell level.

      We will include the multi-state analysis as a new supplementary figure. We will additionally clarify in the Results and Methods exactly how short- and long-lived binding events are classified (1-second threshold consistent with prior single-molecule frameworks for transcription-factor chromatin interactions; Gebhardt et al., 2013; Liu et al., 2014; Kenworthy et al., 2022) and direct the reviewer to these passages.

      For the requested controls, we will include H2B-HaloTag imaging under matched conditions as a long-lived reference for both photobleaching correction and as a positive control for stable chromatin association, addressing R1 point 2 and R3 point 1 simultaneously.

      We will also attempt to address the reviewer’s request for a binding-deficient control. We have lentiviral constructs in hand that encode GATA2 with a C-terminal zinc-finger deletion (which removes the primary DNA-binding domain), an N-terminal zinc-finger deletion, and a double deletion. We will perform single-molecule tracking of these mutants in the engineered cell systems and test whether removing GATA2’s specific DNA-binding capacity produces the predicted reduction in long-lived chromatin engagement, providing a functional perturbation control. The interpretation of these experiments will depend on the mutants expressing and localizing appropriately, which we will validate before drawing kinetic conclusions. We note that an analogous binding-deficient mutant cannot be examined in the physiological context of the Gata2SNAP knock-in mouse, and we will frame the cell-line mutant analyses accordingly. Together with GRID and the H2B-HaloTag control, these mutants provide complementary lines of validation for the two-state kinetic framework.

      (2) Photophysical and focal-plane artifacts

      The authors should exclude contributions from (i) photobleaching, (ii) blinking, and (iii) Z-axis motion. Describe and quantify the photobleaching correction. Provide analyses or controls that distinguish true dissociation events from photophysical blinking/bleaching or axial motion.

      Agreed.

      We will substantially expand the methodological description and provide three new pieces of supplementary analysis:

      - Photobleaching: A per-cell photobleaching-rate distribution will be plotted for each cell type and differentiation stage, and photobleach-corrected residence-time values will be reported alongside apparent values in the relevant figures. We will also perform H2B-HaloTag imaging under matched illumination, exposure, and dye conditions in each cell line as a longlived chromatin-bound reference, establishing per-cell-type bleach lifetimes to which the GATA2 measurements can be referenced. This approach follows recent SMT precedent in which H2B decay was used to correct residence-time measurements for photobleaching, chromatin and nuclear motion, microscope drift, defocalization, and dye photophysics (Ling et al., Science 2026). The right-censoring photobleach-correction model used in our analysis will be described in detail in the revised Methods, including parameter values and per-cell handling.

      - Blinking: The STRAP single-particle tracking pipeline already accommodates fluorophore blinking when linking trajectories across successive frames, following the multiple-targettracing framework of Sergé et al. (Nature Methods, 2008). This use of short gap-frame allowances to avoid artificially splitting trajectories due to fluorophore blinking or transient defocalization is consistent with recent live-cell SMT studies of chromatin-associated factors (Ling et al., Science 2026). We will add an explicit statement to the Methods describing how blinking-tolerant linkage parameters are set, and we will reanalyze representative datasets

      with stricter maximum off-frame settings to ensure this parameter does not drive our conclusions (also addressing R3 point 6).

      - Z-axis motion: Given our 500-ms exposure and the ~500-nm axial detection range of the HiLo configuration, axial loss is expected to be a minor contributor. We will quantify this indirectly by plotting, as a supplementary analysis, the maximum in-plane 2D spatial exploration of each binding trajectory, defined as the long-axis diameter of the 2D trajectory envelope. Although this does not directly measure z-position, it serves as a control for large apparent displacements that could reflect molecules moving out of the HiLo detection volume and demonstrates that observed dissociation events are not dominated by axial drift.

      Representative photobleaching traces from individual cells (lowest, highest, and median bleach rates) will be included to support the single-molecule interpretation (also addresses R1 point 5).

      (3) HILO illumination and nuclear region sampled

      HiLo is sensitive to illumination angle: slight changes sample different nuclear regions. Explain how the HiLo angle was controlled and confirmed comparable across cells and conditions.

      Agreed.

      We will add a Methods subsection describing our HiLo illumination procedure. In brief, we started at a TIRF-supercritical angle and reduced it toward epifluorescence just enough to achieve high imaging depth while minimizing out-of-focus background signal. Within each biological system (cell line or primary cells), the TIRF angle was held constant across Basal, Early, and Late conditions to ensure direct comparability of kinetic measurements across stages.

      (4) Quantification of event counts and long-binding durations

      The number of binding events and the duration of long-binding events are influenced by imaging conditions. Provide a more detailed analysis of how these variables were controlled and assess the sensitivity of the results to detection and tracking parameters.

      Agreed.

      We will (i) normalize per-cell binding-event counts to nuclear cross-sectional area (extracted from the segmented nuclear masks already in the STRAP pipeline) to control for differences in nuclear size; (ii) report the tracking-parameter sensitivity sweep described above; and (iii) confirm in the revised Methods that all imaging conditions (laser power, exposure, dye concentration, sample preparation) were held constant across stages and cell types, consistent with the existing manuscript text. Per the Reviewing Editor’s guidance, the planned labeling-efficiency and absolute-molecule-quantification experiments will further constrain the interpretation of binding-event counts across conditions.

      (5) Evidence that spots are single molecules

      Provide evidence that spots represent single molecules.

      Agreed.

      We will include a small number of per-event intensity traces from our STRAP tracking output, selected to illustrate the single-step photobleaching behavior characteristic of single-molecule emission (intensity remains approximately constant during the binding event and then drops to background in a single step). The nuclear-fluorescence measurements from the planned labeling-titration experiment will also allow us to confirm that bound-spot densities are consistent with single-molecule occupancy at the labeled fraction used for tracking.

      (6) Description of the spot-analysis pipeline

      The Methods should include a detailed STRAP pipeline description.

      Partially agreed; the existing STRAP reference is appropriate, but the Methods will be expanded.

      STRAP (Haque & Coleman, 2025) is a consolidated, automated implementation of two well-established, previously published frameworks: SLIMfast / multipletarget tracing (Sergé et al., 2008) and evalSPT (Normanno et al., 2015), both of which are cited in the original manuscript. We will expand the Methods to describe the parameter set used in our analysis (detection thresholds, linking radii, gap-frame allowance, photobleaching correction model) so that readers can assess the analysis without referring exclusively to the STRAP manuscript and code repository, while preserving the cited STRAP reference for the full algorithmic description. We respectfully suggest that a complete pipeline description duplicating Haque & Coleman (2025) would not be appropriate in a primary research article.

      (7) Differences among cell systems

      The three cell systems yield notably different results. Provide a more detailed explanation for these differences.

      Agreed.

      We will also explicitly describe the caveats of the engineered systems versus the native GATA2-SNAP primary-cell system, in which endogenous GATA2-SNAP remains under physiological regulation. Specifically, we will discuss how variables such as the GATA1null background, ectopic forced nuclear import of GATA1-ERT, and ectopic GATA2-Halo in G1E-ER4 cells, as well as ectopic GATA2-Halo, endogenous GATA1, and cytokine signaling in HPC7 cells, likely contribute to the observed differences in signatures.

      Reviewer 2 (Public review):

      (1) Expression levels of the GATA2-HaloTag transgene

      Determine the expression levels of the GATA2-HaloTag transgene over the course of differentiation under the conditions used for single-molecule imaging.

      Agreed.

      This is the central concern flagged by the Reviewing Editor. For each cell line (G1E-ER4 and HPC7), we will (i) measure total nuclear GATA2-Halo fluorescence per cell under matched acquisition conditions and (ii) convert this fluorescence intensity to absolute molecules per cell using a Halo-CTCF/U2OS reference standard (Cattoglio et al., 2019; absolute CTCF abundance quantification applied previously by our group). This will provide per-cell GATA2Halo molecule counts at each differentiation stage (Basal, Early, Late). For the primary GATA2SNAP cells, we will perform the analogous comparison against a SNAP-RPB1/U2OS standard.

      (2) Fraction of molecules labeled

      Carry out a titration of the HaloTag ligand and compare the amount of labeled protein under single-molecule imaging conditions to that of saturating labeling.

      Agreed.

      We will perform HaloTag-ligand and SNAP-tag-ligand titrations in each cell type, comparing nuclear fluorescence under the limiting-label conditions used for single-molecule tracking with that under saturating labeling. This will yield a per-cell-type labeled fraction and allow us to confirm that comparisons of binding-event counts across conditions are not confounded by differences in labeling efficiency. The labeled-fraction values will be reported in a new supplementary figure and incorporated into our quantification of binding-event rates.

      (3) Robust single-particle tracking

      Show images of particle trajectories or movies superimposing trajectories on imaging data.

      Agreed.

      We will generate visualizations of selected long-lived binding events with single-particle trajectories overlaid on the imaging data — using a multi-frame color overlay (e.g., five sequential frames in distinct colors superimposed) so that linkage of the spot across frames is visually unambiguous — and include them as a new supplementary figure or movie. Examples will be drawn from each cell system to demonstrate consistent tracking quality.

      Reviewer 3 (Public review):

      (1) Photobleaching correction; per-cell bleach lifetimes

      Report the per-stage (or per-cell) photobleaching lifetimes and the photobleachcorrected residence time values alongside apparent values, ideally with an H2B-Halo control.

      Agreed.

      Addressed by the photobleach-rate distribution and H2B-HaloTag control analyses described under R1 point 2. The supplementary figure will explicitly compare per-cell bleach lifetimes across stages, report photobleach-corrected residence-time values alongside apparent values and include H2B-HaloTag controls under matched conditions in each cell line.

      (2) Mechanistic differences across systems

      The three systems show qualitatively different signatures: residence time change in G1EER4, bound fraction expansion in HPC7 and primary cells. Reporting an on-rate proxy alongside k_off would help.

      Agreed.

      Addressed by the cross-system kinetic framing described under R1 point 7 and by the GRID state-spectrum analysis described under R1 point 1. We will explicitly frame the three systems in terms of underlying kinetic mechanism in both Results and Discussion, following the conceptual distinction emphasized by Ling et al. (Science 2026) in which residence time reports binding stability once engaged, whereas changes in bound fraction or event frequency can indicate altered association/recruitment efficiency. In this framework, the G1E-ER4 residencetime signature is consistent with reduced dissociation (a longer-lived bound state), while the longlived-fraction expansion in HPC7 and primary cells is consistent with an increased target-search efficiency or specific-binding-competent pool. Alongside the GRID-derived state-spectrum analysis, we will report an apparent engagement-rate proxy calculated as binding events per unit imaging time normalized to detectable molecule number; this proxy is an approximation, not a direct k_on measurement, as accurate determination of k_on from single-molecule tracking requires concentration-dependent on-rate experiments that are outside the scope of the present study. We thank the reviewer for this suggestion, which we agree sharpens rather than alters the central message.

      (3) Per-cell GATA2 concentration and the uncoupling claim

      Quantify total nuclear GATA2-Halo signal per cell across stages; for primary cells, a western blot or quantitative immunofluorescence on flow-sorted populations would make the uncoupling argument more defensible.

      Agreed.

      For the cell lines, the per-cell nuclear GATA2-Halo quantification described in our response to R2 point 1 addresses this point.

      For primary cells, where the biological claim is strongest, we will exploit the endogenous Gata2SNAP knock-in itself as a quantitative reporter of total GATA2 protein. Specifically, we will label flow-sorted CD71/Ter119 populations from Gata2-SNAP mouse bone marrow with SNAP-Cell 647-SiR at saturating concentration in a parallel acquisition to the limiting-label single-molecule tracking experiment. Total nuclear SNAP-GATA2 fluorescence at saturating labeling provides a measure of endogenous GATA2 abundance per cell at each erythroid stage, in the same chemistry used for our single-molecule measurements, and will be benchmarked against a SNAPRPB1/U2OS reference standard for absolute molecule counting. This approach (i) measures the protein of interest in the labeling chemistry already established in this study; (ii) avoids reliance on quantitative immunofluorescence, which we have not been able to validate under our flowsorted-cell conditions; and (iii) extends the same analytical framework — saturating versus limiting labeling, with U2OS reference standards — across cell lines and primary cells. Quantitative western blotting on flow-sorted populations remains an alternative we will consider if specifically requested by the reviewers.

      (4) Single-cell distribution analysis

      Distribution-based statistics (K-S test, mixture model) rather than (or alongside) meanbased ANOVA, particularly for the Early populations, which look potentially bimodal.

      Agreed.

      We will perform Kolmogorov–Smirnov and Gaussian mixture model analyses of the single-cell long-lived fraction and residence-time distributions across stages, reporting these alongside the existing Welch ANOVA results in a new supplementary figure. This analysis is consistent with the conceptual framework cited in the manuscript (Wheat et al., 2020; Palii et al., 2019) for probabilistic hematopoietic transitions and may reveal subpopulation structure underlying the Early-stage signal. The GRID analysis further complements this by formally testing whether multi-state mixture models are statistically preferred at each stage. However, GRID analysis requires aggregating binding events across cells, which limits our ability to monitor changes in population dispersion at the single-cell level.

      (5) Quantitative integration of CUT&Tag with SMT

      Attempt a back-of-the-envelope calculation of whether the residence-time or fraction changes are quantitatively consistent with the acquisition of the 1,167 Early-restricted sites.

      Partially agreed; will attempt an order-of-magnitude framing.

      We thank the reviewer for this thoughtful suggestion. We agree that more explicit framing of the quantitative relationship between the two datasets will strengthen the integration. We will add a paragraph to the Discussion presenting an order-of-magnitude calculation linking the observed residence-time and long-lived-fraction changes to the steady-state occupancy increase predicted at competent regulatory sites, with explicit caveats regarding (i) the inherently semi-quantitative nature of CUT&Tag signal and (ii) the assumptions required to translate population-averaged occupancy into the genome-wide site count observed. For the G1EER4 cells, we observe relatively minor shifts in population-mean behavior as single-cell dispersion increases. Therefore, it may be difficult to directly link population-based measurements (e.g. CUT&Tag) with single-cell kinetic measurements (SPT). This distinction between occupancy and dynamics is consistent with recent systematic SMT analysis of the eukaryotic transcription machinery, in which factors appearing persistently associated in ensemble genomic assays were shown to exchange on second-scale timescales in living cells (Ling et al., Science 2026), emphasizing that population genomic occupancy and single-molecule residence time are complementary but not directly interchangeable measurements. Closing this gap rigorously is a major hurdle for the field and will require substantial technology development on quantitative single-cell CUT&Tag occupancy measurements. We will therefore frame our analysis as a consistency check rather than a strict quantitative integration. The reviewer notes that this analysis “does not change the central message; it sharpens it,” and we agree.

      (6) Short-lived kinetic interpretation and tracking parameters

      The 1.5 s gap allowance is long relative to the short-lived residence times in primary cells. A sensitivity analysis with tighter gap parameters would help. Also clarify how slowing of search reconciles with increased binding events at Early.

      Agreed.

      Addressed by the tracking-parameter sensitivity analysis described under R1 point 2. We apologize for the lack of clarity in our original description of the gap allowance. Our current maximum off-frame parameter is set to 2 frames, corresponding to a 0.5-s gap allowance. We will rerun the tracking analysis on representative datasets using a maximum off-frame parameter of 1, corresponding to no missed frames, and will report the resulting residence-time distributions alongside the original analysis to demonstrate robustness. We will also clarify in the Results and Discussion how changes in short-lived binding kinetics are reconciled with the increase in detectable binding events at the Early stage, drawing on the apparent engagement-rate proxy interpreted alongside the GRID-derived state-spectrum analysis.

      (7) CUT&Tag peak definition and quantitative analysis

      Report (a) signal intensity distribution at the 1,167 sites across stages (scatter or density plot beyond the heatmap) or (b) differential binding analysis (e.g., DESeq2). State replicate count and overlap of Early-restricted sets across replicates.

      Agreed; normalized fold-change analysis completed, with replicate-aware differential binding analysis planned if additional replicates are generated.

      We have performed a normalized count-based fold-change analysis of the union peak set from the existing GATA2 CUT&Tag dataset (14,468 peaks) using the goodpeaks framework previously used in our group, yielding per-peak log2 fold-change values and discrete dynamicstatus calls (Gained / Lost / Unchanged at |log2FC| ≥ 2) for each of the two transitions (Basal → Early at 0 vs 2 h, and Early → Late at 2 vs 24 h). This provides a conservative quantitative complement to the presence/absence peak-calling analysis presented in Figure 5; if additional replicate data are generated, we will perform replicate-aware differential binding analysis (DiffBind/DESeq2; Love et al., 2014; Stark & Brown, 2011) and report replicate overlap. This analysis addresses option (b) of the reviewer’s request and also enables the visualization requested in option (a) as a cross-stage scatter (Author response image 1). We present the quantitative analysis as a supplement to the presence/absence-defined Early-restricted set in Figure 5 of the manuscript, providing two orthogonal lines of evidence for the same biology. We note that the CUT&Tag experiments were initially performed as a validation step to confirm that the tagged GATA2-Halo constructs recapitulate endogenous chromatin-binding behavior, including appropriate genomic localization and expected GATA switch dynamics. This validation supports the conclusion that the observed single-molecule kinetics reflect physiologically relevant GATA2 engagement. Having established this, we subsequently extended the dataset to perform the quantitative analyses presented here.

      Quantitative findings.

      - 384 peaks were Gained (|log2FC| ≥ 2) at the Basal → Early transition.

      - 1,006 peaks were Lost over the same transition.

      - 178 peaks were Gained at Basal → Early and subsequently Lost at Early → Late, defining the strict differentially-restricted Early set (Author response image 1, red points). This set represents the higher-confidence subset of the manuscript’s broader presence/absence-defined Earlyrestricted set (n = 1,167; defined as MACS2 peaks at q < 0.01 present at Early but absent at Basal and Late).

      - 200 peaks were Gained at Early and retained at Late, indicating stable acquisition.

      - 49 peaks were acquired only at the Late stage.

      The discrepancy between the broader presence/absence set (1,167) and the strict differential set (178) reflects the analytical choice the reviewer raised: presence/absence calls based on a peaksignificance threshold are sensitive to near-threshold peaks, whereas differential analysis with a fold-change cutoff captures only sites with quantitatively pronounced stage-restricted enrichment. We interpret these as two complementary definitions: the broader set captures all peaks meeting a stage-specific peak-calling criterion, and the strict subset isolates the most quantitatively dynamic core of that population.

      Importantly, the three named example loci shown in Figure 5D of the manuscript — Nono (promoter-proximal), Nr3c1 (intron 2), and Gata3 (distal intergenic) — all survive the strict differential criterion (each shows |log<sup>2</sub>FC| ≥ 2 in both transitions, consistent with a clean Gainedthen-Lost signature). The published example panel therefore represents the high-confidence intersection of both definitions, supporting the robustness of the manuscript’s selected illustrative cases.

      We will explicitly state the number of CUT&Tag replicates per stage in the revised Methods and figure legends. Where the differential analysis is currently based on a single replicate per stage, we will explicitly note this and treat the strict subset as a conservative confirmatory analysis. An additional replicate is under consideration for the full revision, and if performed, overlap of Earlyrestricted calls across replicates will be reported.

      Motif cross-validation against a matched-GC background using HOMER and/or MEME-ChIP is planned for the strict differential subset and will be reported alongside the original SeqPos analysis in the revised Figure 5F or its supplement.

      Author response image 1.

      Cross-stage log<sub>2</sub> fold-change scatter for GATA2 CUT&Tag peaks. Each point represents a single peak in the union peak set (n = 14,468). The x-axis shows the log2 fold change from Basal (0 h) to Early (2 h); the y-axis shows the log2 fold change from Early (2 h) to Late (24 h). The sign convention follows the field-standard direction (positive log2FC = increased signal at the later time point). Peaks are colored by dynamic-status classification: unchanged/other (gray; n = 9,794); Lost at Early (blue; n = 109); Gained at Early and retained at Late (orange; n = 200); acquired only at Late (teal; n = 49); and Early-restricted, defined as Gained at Early and Lost at Late with |log2FC| ≥ 2 in both transitions (red; n = 178). The Early-restricted population occupies the lower-right quadrant, consistent with a transient kinetic peak of GATA2 binding.

      Author response image 2.

      Density representation of GATA2 CUT&Tag peak dynamics with Early-restricted peaks highlighted.

      Author response image 2 is shown for illustrative reference and is not annotated with a separate legend; it presents the same data as Author response image 1 in a hexbin density format to emphasize the bulk of unchanged peaks at the origin and the spatial separation of the Early-restricted set.

      Author response image 3.

      Genomic-annotation comparison of newly acquired GATA2 binding at Early. Stacked-bar comparison of genomic annotations (ChIPseeker classification) for two definitions of the newly acquired GATA2 peaks at the Early erythroid stage: all peaks Gained at Basal → Early (orange; n = 384) and the strict Early-restricted subset (Gained then Lost; red; n = 178). Annotation categories shown: Promoter (≤1 kb of TSS), Intron, Distal Intergenic, and Other (Exon, 5′/3′ UTR, Downstream). Both peak sets contain substantial promoter-proximal and distal/intronic components, consistent with the two-subclass model described in Figure 5E–G of the manuscript (GATA2-only promoter-proximal peaks with GATA/RUNX motifs, and GATA2/GATA1 cobound distal peaks with composite GATA/E-box motifs). The strict subset shows a higher proportion of intronic and distal-intergenic sites and a lower proportion of promoter-proximal sites than the full Gained set; this difference will be discussed transparently in the revised Results. Motif analysis (HOMER/MEME-ChIP, planned for the full revision) will be performed on both peak sets to confirm that the GATA/RUNX and GATA/E-box subclass signatures are preserved.

      (8) Knock-in mouse hematopoietic validation

      A brief characterization of basic hematopoietic parameters in homozygotes (CBC, LSK/HSPC frequencies, or colony assays) would confirm the tagged allele is physiological.

      Agreed; data acquired and analyzed.

      We have characterized mature trilineage hematopoietic populations in whole bone marrow from wild-type, heterozygous (Gata2Het), and homozygous (Gata2Homo) Gata2-SNAP knock-in mice (n = 5 per genotype). Bone marrow cells were stained for myeloid (CD11b<sup>+</sup> Gr1<sup>+</sup>), lymphoid (CD3<sup>+</sup>/CD4<sup>+</sup>/CD8<sup>+</sup>/B220<sup>+</sup>/CD19<sup>+</sup>), and erythroid (Ter119<sup>+</sup>) markers and analyzed by flow cytometry. Lineage frequencies are shown as percentages of live bone marrow cells in a new Figure Supplement in the revised manuscript.

      For myeloid and erythroid populations, omnibus one-way ANOVA detected no significant differences across genotypes (Myeloid: F(2,12) = 2.616, P = 0.1140; Erythroid: F(2,12) = 0.4943, P = 0.6219). Dunnett’s multiple-comparisons test against the WT control did not detect significant pairwise differences for either knock-in genotype (Myeloid: WT vs Het P = 0.1351, WT vs Homo P = 0.9926; Erythroid: WT vs Het P = 0.7017, WT vs Homo P = 0.9602).

      For the lymphoid compartment, although the omnibus ANOVA reached significance (F(2,12) = 6.690, P = 0.0112), no pairwise comparison against WT remained significant after multiplecomparisons correction (Dunnett’s adjusted P values: WT vs Het = 0.1217; WT vs Homo = 0.2078). We therefore interpret this result conservatively. Brown-Forsythe and Bartlett’s tests showed no significant differences in variance across genotypes (P = 0.1423 and P = 0.0908), so the result is not attributable to unequal variances. We do not interpret these data as indicating an unambiguous lymphoid phenotype in either heterozygous or homozygous Gata2-SNAP mice; this interpretation is consistent with the broader pattern across all three lineages, in which no pairwise comparison against WT survives multiple-comparisons correction. We will note in the figure legend and in the Results text that more granular HSPC-compartment analysis (LSK, MPP, lineage-restricted progenitor frequencies) and a complete blood count (CBC) remain valuable directions for future characterization of this resource and will be considered for the full revision if specifically requested.

      Author response image 4.

      Bone marrow trilineage frequencies in Gata2-SNAP knock-in mice. Bone marrow was harvested from the femurs and tibias of wild-type (WT), heterozygous (Gata2Het), and homozygous (Gata2Homo) Gata2-SNAP knock-in mice (n = 5 per genotype; mixed sex; 12–14 weeks). After ACK lysis, cells were stained for myeloid (CD11b<sup>+</sup> Gr1<sup>+</sup>), lymphoid (CD3<sup>+</sup>/CD4<sup>+</sup>/CD8<sup>+</sup>/B220<sup>+</sup>/CD19<sup>+</sup>), and erythroid (Ter119<sup>+</sup>) markers and analyzed by flow cytometry. Each dot represents one mouse, and horizontal bars indicate genotype means. Statistical results: Myeloid: ANOVA F(2,12) = 2.616, P = 0.1140; Dunnett’s adjusted P values WT vs Het = 0.1351, WT vs Homo = 0.9926. Lymphoid: ANOVA F(2,12) = 6.690, P = 0.0112 (omnibus); Dunnett’s adjusted P values WT vs Het = 0.1217, WT vs Homo = 0.2078. Erythroid: ANOVA F(2,12) = 0.4943, P = 0.6219; Dunnett’s adjusted P values WT vs Het = 0.7017, WT vs Homo = 0.9602. Brown-Forsythe and Bartlett’s tests for unequal variance were non-significant in all three lineages. Although the lymphoid omnibus ANOVA reached nominal significance, no pairwise comparison with WT remained significant after multiple-comparison correction; we therefore interpret this result conservatively (see response to R3 point 8).

      Summary

      We thank the editors and the three reviewers for the constructive and detailed assessment. The planned revisions consist of:

      - Four new experiments [planned] (HaloTag/SNAP labeling efficiency and absolute molecule counts via U2OS reference standards; H2B-HaloTag photobleaching reference; percell quantification of total endogenous GATA2 in flow-sorted primary Gata2-SNAP populations via saturating SNAP-tag labeling, benchmarked against a SNAP-RPB1/U2OS reference standard; single-molecule tracking of GATA2 N-terminal, C-terminal, and double zinc-finger deletion mutants in the engineered cell systems as a binding-deficient functional control).

      - Six analyses of existing data (GRID multi-state fitting [planned]; per-cell bleach-rate distributions and photobleach-corrected residence times [planned]; tracking-parameter sensitivity [planned]; nuclear-area normalization and total-displacement controls [planned]; normalized fold-change CUT&Tag analysis [completed; motif cross-validation planned], presented in Author response images 1–3; distribution-based single-cell statistics [planned]).

      - One previously-acquired dataset [completed] (trilineage hematopoietic flow cytometry of homozygous Gata2-SNAP knock-in mice; presented in Author response image 4 with full statistical detail).

      - Substantial revisions to text and figures [planned] to address statistical reporting, methodological description, mechanistic framing of cross-system differences, and refinement of the Figure 6 schematic.

      With respect to the requested binding-deficient single-molecule control, we will attempt to address this directly using sequence-validated lentiviral constructs in hand encoding GATA2 mutants lacking the C-terminal zinc finger, the N-terminal zinc finger, or both. These mutant analyses will be complemented by GRID multi-state analysis and H2B-HaloTag controls, providing converging lines of validation for the two-state kinetic framework. We note that an analogous mutant cannot be examined in the physiological context of the Gata2-SNAP knock-in mouse, and we will frame the cell-line mutant analyses accordingly.

      We believe these revisions directly address the editors’ specific guidance regarding labeling efficiency and methodological clarification. We thank the editors and reviewers for their time and look forward to submitting the revised manuscript.

      References cited in this response:

      References listed below are cited in this provisional response in support of the planned analyses and methodology.

      Cattoglio, C., Pustova, I., Walther, N., Ho, J. J., Hantsche-Grininger, M., Inouye, C. J., Hossain, M. J., Dailey, G. M., Ellenberg, J., Darzacq, X., Tjian, R., & Hansen, A. S. (2019). Determining cellular CTCF and cohesin abundances to constrain 3D genome models. eLife, 8, e40164. https://doi.org/10.7554/eLife.40164

      Gebhardt, J. C. M., Suter, D. M., Roy, R., Zhao, Z. W., Chapman, A. R., Basu, S., Maniatis, T., & Xie, X. S. (2013). Single-molecule imaging of transcription factor binding to DNA in live mammalian cells. Nature Methods, 10(5), 421–426. https://doi.org/10.1038/nmeth.2411

      Hansen, A. S., Pustova, I., Cattoglio, C., Tjian, R., & Darzacq, X. (2017). CTCF and cohesin regulate chromatin loop stability with distinct dynamics. eLife, 6, e25776. https://doi.org/10.7554/eLife.25776

      Haque, N., & Coleman, R. A. (2025). Dynamic transcription pre-initiation complex assembly governs initiation efficiency. bioRxiv. https://doi.org/10.1101/2025.05.07.652662

      Heinz, S., Benner, C., Spann, N., Bertolino, E., Lin, Y. C., Laslo, P., Cheng, J. X., Murre, C., Singh, H., & Glass, C. K. (2010). Simple combinations of lineage-determining transcription factors prime cis-regulatory elements required for macrophage and B cell identities. Molecular Cell, 38(4), 576–589. https://doi.org/10.1016/j.molcel.2010.05.004

      Kaya-Okur, H. S., Wu, S. J., Codomo, C. A., Pledger, E. S., Bryson, T. D., Henikoff, J. G., Ahmad, K., & Henikoff, S. (2019). CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nature Communications, 10(1), 1930. https://doi.org/10.1038/s41467-019-09982-5

      Kenworthy, C. A., Haque, N., Liou, S.-H., Chandris, P., Wong, V., Dziuba, P., Lavis, L. D., Liu, W.-L., Singer, R. H., & Coleman, R. A. (2022). Bromodomains regulate dynamic targeting of the PBAF chromatin-remodeling complex to chromatin hubs. Biophysical Journal, 121(9), 1738–1752. https://doi.org/10.1016/j.bpj.2022.03.027

      Ling, Y. H., Liang, C., Wang, S., & Wu, C. (2026). Live-cell single-molecule dynamics of eukaryotic RNA polymerase machineries. Science, 391, eads0960. https://doi.org/10.1126/science.ads0960

      Liu, Z., Legant, W. R., Chen, B.-C., Li, L., Grimm, J. B., Lavis, L. D., Betzig, E., & Tjian, R. (2014). 3D imaging of Sox2 enhancer clusters in embryonic stem cells. eLife, 3, e04236. https://doi.org/10.7554/eLife.04236

      Loeffler, D., Wang, W., Hopf, A., Hilsenbeck, O., Bourgine, P. E., Rudolf, F., Martin, I., & Schroeder, T. (2018). Mouse and human HSPC immobilization in liquid culture by CD43- or CD44-antibody coating. Blood, 131(13), 1425–1429. https://doi.org/10.1182/blood-2017-07-794131

      Love, M. I., Huber, W., & Anders, S. (2014). Moderated estimation of fold change and dispersion for RNAseq data with DESeq2. Genome Biology, 15(12), 550. https://doi.org/10.1186/s13059-014-0550-8

      Machanick, P., & Bailey, T. L. (2011). MEME-ChIP: motif analysis of large DNA datasets. Bioinformatics, 27(12), 1696–1697. https://doi.org/10.1093/bioinformatics/btr189

      Normanno, D., Boudarène, L., Dugast-Darzacq, C., Chen, J., Richter, C., Proux, F., Bénichou, O., Voituriez, R., Darzacq, X., & Dahan, M. (2015). Probing the target search of DNA-binding proteins in mammalian cells using TetR as model searcher. Nature Communications, 6, 7357. https://doi.org/10.1038/ncomms8357

      Palii, C. G., Cheng, Q., Gillespie, M. A., Shannon, P., Mazurczyk, M., Napolitani, G., Price, N. D., Ranish, J. A., Morrissey, E., Higgs, D. R., & Brand, M. (2019). Single-cell proteomics reveal that quantitative changes in co-expressed lineage-specific transcription factors determine cell fate. Cell Stem Cell, 24(5), 812–825.e5. https://doi.org/10.1016/j.stem.2019.02.016

      Sergé, A., Bertaux, N., Rigneault, H., & Marguet, D. (2008). Dynamic multiple-target tracing to probe spatiotemporal cartography of cell membranes. Nature Methods, 5(8), 687–694. https://doi.org/10.1038/nmeth.1233

      Stark, R., & Brown, G. D. (2011). DiffBind: Differential binding analysis of ChIP-Seq peak data. Bioconductor. http://bioconductor.org/packages/release/bioc/html/DiffBind.html

      Taylor, S. J., Stauber, J., Bohorquez, O., Tatsumi, G., Kumari, R., Chakraborty, J., Bartholdy, B. A., Schwenger, E., Sundaravel, S., Farahat, A. A., Dutta, A., Koche, R. P., Steidl, U., & Wheat, J. C. (2024). Pharmacological restriction of genomic binding sites redirects PU.1 pioneer transcription factor activity. Nature Genetics, 56(10), 2213–2227. https://doi.org/10.1038/s41588-024-01911-7

      Wheat, J. C., Salsman, J., Reekie, I., Mathhwala, A., Black, K. L., Tiedt, R., Shroff, H., & Steidl, U. (2020). Single-molecule imaging of transcription dynamics in somatic stem cells. Nature, 583(7816), 431– 436. https://doi.org/10.1038/s41586-020-2432-4

    1. eLife Assessment

      This manuscript presents a valuable and timely contribution by incorporating desolvation barriers into coarse-grained models of biomolecular condensates. The findings are convincing, supported by a clear physical model and systematic simulations showing effects on phase behavior, packing, and dynamics. Some clarification and broader context would improve the manuscript, but it provides a foundation that will be of use for developing more realistic coarse-grained interaction schemes.

    2. Reviewer #1 (Public review):

      This manuscript is very interesting and timely. By introducing the critical effects of desolvation barriers and solvent (water)-separated minima into the implicit-solvent potentials (of mean force, PMFs) for coarse-grained molecular dynamics simulations of biomolecular liquid-liquid phase separation (LLPS), this work fills a gap that should be apparent to researchers of protein folding in the past couple of decades but has so far escaped deserved attention such that these basic features of aqueous solvation have seldom, though not never, been invoked in recent studies of biomolecular condensates. Although the present paper deals almost exclusively with homopolymers, this work can be a foundation for the future development of a new, more physical coarse-grained interaction scheme for simulating amino acid sequence-dependent effects, which I presume is the authors' ongoing or next endeavor. The results presented in this manuscript are highly valuable.

      However, there is room for improvement in the authors' description of (i) the broader impact of effects of desolvation barrier and solvent-separated minimum in the thermodynamics of biomolecular condensates, especially with regard to the ramifications on hydrostatic pressure-dependent effects; (ii) the physical implication of using a 20-parameter hydropathy scale rather than a 210-parameter pairwise amino acid interaction scheme; and (iii) temperature-dependent effects, including the authors' discussion of "enthalpic" and "entropic" contributions. In all these aspects, the authors' discussion should be put in a more comprehensive context of the existing literature. At a few other places, the description of the methods and results should be clarified as well. Accordingly, the authors should revise the manuscript to address the following items thoroughly within the revised manuscript (not merely in the response letter) with the additional references mentioned below included in the revised discussion:

      (1) In several places, e.g., on line 77 (p.2), the authors appear to suggest that "implicit-solvent representation" is the origin of the deficiency in commonly utilized coarse-grained potentials that this study is aiming to rectify. But desolvation barriers and solvent-separated minima are also features of implicit-solvent representations; they are just features that should be incorporated in more accurate implicit-solvent potentials. This point is stated quite clearly and accurately in the Abstract (p.1) but not consistently in the rest of the text. The authors should check the entire text carefully to ensure that a coherent, accurate perspective is presented.

      (2) In the discussion of the importance of desolvation barriers and solvent-separated minima in the Introduction (pp.1-3), connections should be drawn to recent works that utilize these PMF features to rationalize hydrostatic pressure (P)-modulated effects on biomolecular LLPS, including the P-dependent reentrant phase separation of alpha elastin; see Cinar et al. (2019) Chem Eur J 25:13049 (https://chemistry-europe.onlinelibrary.wiley.com/doi/full/10.1002/chem.201902210) and references therein, especially discussions around Figures 10, 11 & 13 in this reference.

      (3) In the lower panels of Figures 2D, E (p.5), what do the differently colored small circles in the double-minimum free energy profiles represent? Does the color shading have the same meaning as that in the upper panels? If so, what do the positions of the circles on the free energy profile represent? The authors should clarify this.

      (4) The discussion regarding entropy and enthalpy around Figure 2 is quite confusing as it stands. What do the authors mean exactly by the association of entropy or enthalpy with the desolvation barrier of the solvent-separated minimum? Are they referring to conformational entropy?

      (5) Do the authors assume that the PMF (effective implicit-solvent potential) is a purely enthalpic term? It appears to be the authors' assumption. If so, the assumption has to be stated clearly in their discussion of "entropy" vs "enthalpy" around Figure 2.

      (6) Closely related to points 3-5 above, it should be stated clearly that the "temperature" used in the authors' simulations does not represent experimental temperature if the authors are using purely enthalpic effective potentials because PMFs are in fact temperature-dependent. This clarification is necessary to avoid misunderstanding. In this regard, it should be noted that temperature-dependent effective interactions have been used for modeling biomolecular condensates in analytical theory (Lin, Song, Forman-Kay & Chan, J Mol Liq 2017, already in the citation list) as well as in coarse-grained molecular dynamics simulations [Dignon et al. (2019) ACS Cent Sci 5:821-830 (https://pubs.acs.org/doi/10.1021/acscentsci.9b00102); Chakravarti & Joseph (2025) Protein Sci 34:e70284 (https://onlinelibrary.wiley.com/doi/10.1002/pro.70284)]. The latter two studies, not cited currently, are particularly relevant and thus should be cited because the authors may wish to incorporate temperature-dependent features in their ongoing or future effort in constructing a more comprehensive coarse-grained interaction scheme for biomolecular LLPS simulation.

      (7) In tackling "entropy" vs "enthalpy", it should be noted that the temperature dependence of the effective interactions entails an entropic contribution (which is itself temperature dependent) in addition to conformational entropy. As for the effective potential with desolvation barrier and solvent-separated minimum, it should be noted that the decomposition into entropic and enthalpic contributions at the direct contact, desolvation barrier, and solvent-separated minimum can be dramatically different, see, e.g., MaCallum et al. (2007) PNAS 104:6206-6210 (https://www.pnas.org/doi/full/10.1073/pnas.0605859104) and references therein.

      (8) P.7, line 340: The proportionality relation follows directly from the standard Flory-Huggins result T_c = T chi(T)/chi_c, thus the proportionality constant is exactly 1/chi_c. Is this the standard relation that the authors are invoking here? The authors should clarify this.

      (9) The study on dynamic consequences on pp.8-11 is interesting, but clarifications are necessary:

      (i) The vertical schematic in Figure 4A should be explained in detail in its entirety. As it stands, no explanation is provided either in the figure caption or in the text. In particular, what does "elasticity driven" refer to?

      (ii) The top snapshot in Figure 4A is labeled t_sim = 0 ns. Does it mean that the snapshot shown is the only chain configuration that the authors used to start the simulation, and that the snapshot does NOT represent the result of any time evolution, no matter how short the duration is? However, if that is the case, why is this snapshot identified with spinodal decomposition if it is not the product of a time evolution from a more homogeneous configuration?

      (iii) Related to (ii) - do the rectangular boxes shown represent the entire simulation box or just part of the box containing the polymer chains? One would imagine that if the top snapshot represents spinodal decomposition, the simulation would have been started at a more uniform distribution a short time prior? Why is this not the case?

      (iv) What precisely do the small yellow beads and black-colored springs in the zoom-in image of Figure 4E represent?

      (10) In discussing dynamic effects, it is useful to draw connections to related works on the effect of chain flexibility on "aging" of condensate [Biswas & Potoyan (2024) PRX 45:9222-9245 (https://journals.aps.org/prxlife/abstract/10.1103/PRXLife.2.023011)] and characterization of viscoelasticity in simulations of biomolecular condensates [Tejedor et al. (2023) J Phys Chem B 127:4441-4459 (https://pubs.acs.org/doi/10.1021/acs.jpcb.3c01292)], as the effects of desolvation can be explored further based on these prior works.

      (11) Much of the present study is based on the original HPS formulation of Dignon et al. (2018). In this regard and also in anticipation of future development of improved interaction schemes, several issues should be stated and discussed, even if briefly:

      (i) The original HPS model has a basic shortcoming in accounting for the relative interaction strengths of, among others, arginine vs lysine residues [Das et al. (2020) PNAS 117:28795-28805 (https://www.pnas.org/doi/10.1073/pnas.2008122117)].

      (ii) Compared to 210-parameter pairwise interaction schemes, such as KH in Dignon et al. (2018) and Joseph et al. (2021), the 20-parameter interaction scheme is likely too restrictive to account for pairwise amino acid residue interactions [Wessén et al. (2022) J Phys Chem B 45:9222-9245 (https://pubs.acs.org/doi/10.1021/acs.jpcb.2c06181)].

      (iii) The height of the desolvation barrier may vary significantly for different amino acid residue pairs, see, e.g., Figure 11 of Cinar et al. (2019) mentioned above (and references therein). The authors should discuss these nuances in the revised version. They may also wish to take them into consideration in future investigations.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in the molecular simulation of biomolecular condensates. Most residue-level coarse-grained models used for IDP phase separation employ implicit solvent and represent effective interactions through relatively simple pairwise potentials. While these models have been very useful, they usually do not explicitly distinguish direct contacts from solvent-separated interactions, nor do they include an energetic barrier associated with water removal. This manuscript attempts to address that limitation by introducing desolvation-inspired terms into coarse-grained models and examining their consequences for phase behavior, chain conformations, dense-phase packing, and dynamics.

      Strengths:

      The central idea is physically well motivated. Using a simple homopolymer model, the authors show that increasing the desolvation barrier suppresses phase separation, whereas stabilizing solvent-separated contacts enhances phase separation. They further show that solvent-separated interactions can reduce dense-phase over-compaction, which is a meaningful result given the known challenges in obtaining both accurate single-chain dimensions and realistic dense-phase properties from the same coarse-grained model. The finding that desolvation-like terms can reshape dense-phase packing without simply rescaling the overall interaction strength is interesting and could be useful for future model development. I also found the attempt to connect conformational changes across dilute and dense phases with thermal distance from the critical point to be intriguing. The dynamic analysis, including the FRAP-like simulations and the discussion of kinetic arrest during coarsening, adds another useful dimension to the work.

      Weaknesses:

      At the same time, there are several places where the manuscript would benefit from more careful framing. First, the desolvation terms are still effective coarse-grained parameters rather than a direct representation of water molecules. The language sometimes gives the impression that desolvation is being treated explicitly, whereas the model introduces desolvation-inspired effective interactions into an implicit-solvent framework. Second, the conformational analysis is interesting, but the broader context of prior work on dilute-to-dense phase conformational reorganization of IDPs could be more clearly discussed. This would help clarify what is new in the present work, whether it is the conformational change itself, its dependence on desolvation terms, or the proposed scaling with distance from the critical point. Third, the dynamic results are potentially useful, but the manuscript should more clearly articulate what is nontrivial beyond the expected slowing of local rearrangements by an added barrier in the potential.

      Overall, I think this is a useful and potentially important contribution.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      This manuscript is very interesting and timely. By introducing the critical effects of desolvation barriers and solvent (water)-separated minima into the implicit-solvent potentials (of mean force, PMFs) for coarse-grained molecular dynamics simulations of biomolecular liquid-liquid phase separation (LLPS), this work fills a gap that should be apparent to researchers of protein folding in the past couple of decades but has so far escaped deserved attention such that these basic features of aqueous solvation have seldom, though not never, been invoked in recent studies of biomolecular condensates. Although the present paper deals almost exclusively with homopolymers, this work can be a foundation for the future development of a new, more physical coarse-grained interaction scheme for simulating amino acid sequence-dependent effects, which I presume is the authors' ongoing or next endeavor. The results presented in this manuscript are highly valuable.

      We thank the reviewer for all the insightful comments.

      However, there is room for improvement in the authors' description of (i) the broader impact of effects of desolvation barrier and solvent-separated minimum in the thermodynamics of biomolecular condensates, especially with regard to the ramifications on hydrostatic pressure-dependent effects; (ii) the physical implication of using a 20-parameter hydropathy scale rather than a 210-parameter pairwise amino acid interaction scheme; and (iii) temperature-dependent effects, including the authors' discussion of "enthalpic" and "entropic" contributions. In all these aspects, the authors' discussion should be put in a more comprehensive context of the existing literature. At a few other places, the description of the methods and results should be clarified as well. Accordingly, the authors should revise the manuscript to address the following items thoroughly within the revised manuscript (not merely in the response letter) with the additional references mentioned below included in the revised discussion:

      (1) In several places, e.g., on line 77 (p.2), the authors appear to suggest that "implicit-solvent representation" is the origin of the deficiency in commonly utilized coarse-grained potentials that this study is aiming to rectify. But desolvation barriers and solvent-separated minima are also features of implicit-solvent representations; they are just features that should be incorporated in more accurate implicit-solvent potentials. This point is stated quite clearly and accurately in the Abstract (p.1) but not consistently in the rest of the text. The authors should check the entire text carefully to ensure that a coherent, accurate perspective is presented.

      We thank the reviewer for the insightful comment and suggestion. In this work, rather than departing from the implicit‑solvent modeling framework, our intention is to incorporate the desolvation effect within the implicit solvent model framework. In the revised manuscript, we will revise the text to ensure this point is presented clearly and consistently throughout the paper.

      (2) In the discussion of the importance of desolvation barriers and solvent-separated minima in the Introduction (pp.1-3), connections should be drawn to recent works that utilize these PMF features to rationalize hydrostatic pressure (P)-modulated effects on biomolecular LLPS, including the P-dependent reentrant phase separation of alpha elastin; see Cinar et al. (2019) Chem Eur J 25:13049 (https://chemistry-europe.onlinelibrary.wiley.com/doi/full/10.1002/chem.201902210) and references therein, especially discussions around Figures 10, 11 & 13 in this reference.

      We thank the reviewer for bringing these references to our attention. The hydrostatic pressure modulated effects on LLPS provide important context for understanding the physical significance of desolvation barriers and solvent‑separated minima. In the revised manuscript, we will expand the literature discussion by incorporating previous studies on pressure‑modulated phase separation.

      (3) In the lower panels of Figures 2D, E (p.5), what do the differently colored small circles in the double-minimum free energy profiles represent? Does the color shading have the same meaning as that in the upper panels? If so, what do the positions of the circles on the free energy profile represent? The authors should clarify this.

      We thank the reviewer for the suggestion to improve the clarity of the figure. In the lower panels of Figures 2D and 2E, the colored dots were intended solely as a qualitative illustration of the populations of residue‑pair configurations along the effective energy surface. Their colors are not related to the color scale used in the phase diagrams shown in the upper panels. We will modify the color scheme to improve clarity.

      (4) The discussion regarding entropy and enthalpy around Figure 2 is quite confusing as it stands. What do the authors mean exactly by the association of entropy or enthalpy with the desolvation barrier of the solvent-separated minimum? Are they referring to conformational entropy?

      We apologize for the confusion. When the desolvation barrier is high, configurations with inter‑residue distances corresponding to the barrier region become difficult to access, thereby reducing the conformational entropy of the condensed phase. This interpretation is supported by Figure 2—figure supplement 1C, where increasing the desolvation barrier decreases the population in the barrier region of the radial distribution function, indicating that fewer residue‑pair configurations are sampled there. In contrast, increasing the depth of the solvent‑separated minimum makes the condensed phase more energetically favorable. In the revised manuscript, we will incorporate this discussion to improve clarity.

      (5) Do the authors assume that the PMF (effective implicit-solvent potential) is a purely enthalpic term? It appears to be the authors' assumption. If so, the assumption has to be stated clearly in their discussion of "entropy" vs "enthalpy" around Figure 2.

      We thank the reviewer for highlighting this important point. In this work, the PMF profile is constructed from atomistic simulation results, and thus both entropic and enthalpic contributions shape the overall PMF. In the revised manuscript, we will clarify that the PMF represents a free‑energy profile along the intermolecular distance and therefore incorporates enthalpic and entropic contributions from the solute, solvent, and configurational degrees of freedom.

      (6) Closely related to points 3-5 above, it should be stated clearly that the "temperature" used in the authors' simulations does not represent experimental temperature if the authors are using purely enthalpic effective potentials because PMFs are in fact temperature-dependent. This clarification is necessary to avoid misunderstanding. In this regard, it should be noted that temperature-dependent effective interactions have been used for modeling biomolecular condensates in analytical theory (Lin, Song, Forman-Kay & Chan, J Mol Liq 2017, already in the citation list) as well as in coarse-grained molecular dynamics simulations [Dignon et al. (2019) ACS Cent Sci 5:821-830 (https://pubs.acs.org/doi/10.1021/acscentsci.9b00102); Chakravarti & Joseph (2025) Protein Sci 34:e70284 (https://onlinelibrary.wiley.com/doi/10.1002/pro.70284)]. The latter two studies, not cited currently, are particularly relevant and thus should be cited because the authors may wish to incorporate temperature-dependent features in their ongoing or future effort in constructing a more comprehensive coarse-grained interaction scheme for biomolecular LLPS simulation.

      We thank the reviewer for raising this important point. We agree that PMFs and the corresponding effective interactions should be temperature dependent, and therefore the simulation temperature in our current temperature-independent CG potential cannot be interpreted as a fully quantitative experimental temperature. In the revised manuscript, we will clarify the above point. We will also expand the discussion to include previous studies that introduced temperature-dependent effective interactions in analytical theories and coarse-grained simulations of biomolecular condensates.

      (7) In tackling "entropy" vs "enthalpy", it should be noted that the temperature dependence of the effective interactions entails an entropic contribution (which is itself temperature dependent) in addition to conformational entropy. As for the effective potential with desolvation barrier and solvent-separated minimum, it should be noted that the decomposition into entropic and enthalpic contributions at the direct contact, desolvation barrier, and solvent-separated minimum can be dramatically different, see, e.g., MaCallum et al. (2007) PNAS 104:6206-6210 (https://www.pnas.org/doi/full/10.1073/pnas.0605859104) and references therein.

      We agree that a temperature‑dependent PMF includes entropic contributions beyond the configurational entropy discussed around Figure 2. In the present manuscript, our discussion of entropy in that context refers specifically to the reduced accessible configurational space of residue‑pair states in the coarse‑grained ensemble, rather than to a full thermodynamic decomposition of the PMF. In the revised manuscript, we will make this distinction explicit. We will also note that the direct‑contact minimum, desolvation barrier, and solvent‑separated minimum may each have distinct enthalpic and entropic components, and that resolving these components would require additional temperature‑dependent PMF calculations. We will discuss this as a limitation of the current model and as a direction for future parameterization.

      (8) P.7, line 340: The proportionality relation follows directly from the standard Flory-Huggins result T_c = T chi(T)/chi_c, thus the proportionality constant is exactly 1/chi_c. Is this the standard relation that the authors are invoking here? The authors should clarify this.

      We thank the reviewer for pointing this out. Yes, our argument uses the condition that chi_c is fixed at the critical point for a given chain length. We will revise the text to explicitly state this relation and add the missing intermediate step, so that the proportionality used in the manuscript is clearer.

      (9) The study on dynamic consequences on pp.8-11 is interesting, but clarifications are necessary:

      (i) The vertical schematic in Figure 4A should be explained in detail in its entirety. As it stands, no explanation is provided either in the figure caption or in the text. In particular, what does "elasticity driven" refer to?

      (ii) The top snapshot in Figure 4A is labeled t_sim = 0 ns. Does it mean that the snapshot shown is the only chain configuration that the authors used to start the simulation, and that the snapshot does NOT represent the result of any time evolution, no matter how short the duration is? However, if that is the case, why is this snapshot identified with spinodal decomposition if it is not the product of a time evolution from a more homogeneous configuration?

      (iii) Related to (ii) - do the rectangular boxes shown represent the entire simulation box or just part of the box containing the polymer chains? One would imagine that if the top snapshot represents spinodal decomposition, the simulation would have been started at a more uniform distribution a short time prior? Why is this not the case?

      (iv) What precisely do the small yellow beads and black-colored springs in the zoom-in image of Figure 4E represent?

      We thank the reviewer for pointing out these unclear issues in Figure 4. In the revised manuscript, we will better explain the vertical schematic in Figure 4A, including the progression from the early growth of density fluctuations, to intermediate kinetic arrest, and finally to late-stage coarsening. We will also clarify that “elasticity driven” refers to the resistance to domain deformation caused by transient inter-chain network connectivity. We will clarify that t_sim = 0 denotes the time immediately after the temperature quench from the high-temperature homogeneous state to the low-temperature two-phase region. This snapshot is the post-quench initial configuration, while spinodal decomposition refers to the subsequent amplification of density fluctuations after the quench. The displayed snapshot is one representative trajectory, not the only initial configuration used in the simulations. The quantitative kinetic analysis was averaged over multiple independent trajectories. The rectangular box represents the entire simulation box. Although the system was equilibrated at high temperature before the quench, instantaneous density fluctuations remain, so the initial configuration is not perfectly uniform. In Figure 4E, the yellow beads represent interacting residue pairs. The black springs schematically represent the transient elastic network formed by these interactions, rather than a precise structural model.

      (10) In discussing dynamic effects, it is useful to draw connections to related works on the effect of chain flexibility on "aging" of condensate [Biswas & Potoyan (2024) PRX 45:9222-9245 (https://journals.aps.org/prxlife/abstract/10.1103/PRXLife.2.023011)] and characterization of viscoelasticity in simulations of biomolecular condensates [Tejedor et al. (2023) J Phys Chem B 127:4441-4459 (https://pubs.acs.org/doi/10.1021/acs.jpcb.3c01292)], as the effects of desolvation can be explored further based on these prior works.

      We thank the reviewer for this helpful suggestion. In the revised Discussion, we will cite and discuss the related studies on condensate aging and viscoelasticity, including the effects of chain flexibility, sticker lifetime, desolvation, and transient network formation on condensate material properties. These works provide an important context for interpreting our dynamic results. We will clarify that desolvation may influence condensate dynamics not only by slowing local rearrangements, but also by modulating transient network connectivity, kinetic arrest, and viscoelastic relaxation.

      (11) Much of the present study is based on the original HPS formulation of Dignon et al. (2018). In this regard and also in anticipation of future development of improved interaction schemes, several issues should be stated and discussed, even if briefly:

      (i) The original HPS model has a basic shortcoming in accounting for the relative interaction strengths of, among others, arginine vs lysine residues [Das et al. (2020) PNAS 117:28795-28805 (https://www.pnas.org/doi/10.1073/pnas.2008122117)].

      (ii) Compared to 210-parameter pairwise interaction schemes, such as KH in Dignon et al. (2018) and Joseph et al. (2021), the 20-parameter interaction scheme is likely too restrictive to account for pairwise amino acid residue interactions [Wessén et al. (2022) J Phys Chem B 45:9222-9245 (https://pubs.acs.org/doi/10.1021/acs.jpcb.2c06181)].

      (iii) The height of the desolvation barrier may vary significantly for different amino acid residue pairs, see, e.g., Figure 11 of Cinar et al. (2019) mentioned above (and references therein). The authors should discuss these nuances in the revised version. They may also wish to take them into consideration in future investigations.

      We thank the reviewer for highlighting these important limitations of the original HPS-based framework. We agree that a 20‑parameter hydropathy‑scale model has limitation in fully capturing residue‑pair‑specific interactions, including well‑established differences such as those between arginine and lysine. In the revised manuscript, we will explicitly discuss this limitation and cite the suggested studies on residue‑specific and pairwise interaction schemes. We also agree that desolvation barriers and solvent‑separated minima are likely to depend on amino‑acid pair identity. In the present work, we employ a simplified, residue‑independent desolvation parameterization to isolate the general thermodynamic and kinetic consequences of desolvation in coarse‑grained LLPS simulations. In the revised Discussion, we will clarify this scope and emphasize that developing residue‑pair‑specific desolvation parameters, potentially within a 210‑parameter interaction framework, is an important direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in the molecular simulation of biomolecular condensates. Most residue-level coarse-grained models used for IDP phase separation employ implicit solvent and represent effective interactions through relatively simple pairwise potentials. While these models have been very useful, they usually do not explicitly distinguish direct contacts from solvent-separated interactions, nor do they include an energetic barrier associated with water removal. This manuscript attempts to address that limitation by introducing desolvation-inspired terms into coarse-grained models and examining their consequences for phase behavior, chain conformations, dense-phase packing, and dynamics.

      Strengths:

      The central idea is physically well motivated. Using a simple homopolymer model, the authors show that increasing the desolvation barrier suppresses phase separation, whereas stabilizing solvent-separated contacts enhances phase separation. They further show that solvent-separated interactions can reduce dense-phase over-compaction, which is a meaningful result given the known challenges in obtaining both accurate single-chain dimensions and realistic dense-phase properties from the same coarse-grained model. The finding that desolvation-like terms can reshape dense-phase packing without simply rescaling the overall interaction strength is interesting and could be useful for future model development. I also found the attempt to connect conformational changes across dilute and dense phases with thermal distance from the critical point to be intriguing. The dynamic analysis, including the FRAP-like simulations and the discussion of kinetic arrest during coarsening, adds another useful dimension to the work.

      We thank the reviewer for all these positive and constructive assessment and comments. We are encouraged that the reviewer found the central idea physically well motivated and recognized the value of introducing desolvation-inspired terms to distinguish direct contacts, solvent-separated interactions, and the energetic barrier associated with water removal in coarse-grained models of biomolecular condensates.

      Weaknesses:

      At the same time, there are several places where the manuscript would benefit from more careful framing. First, the desolvation terms are still effective coarse-grained parameters rather than a direct representation of water molecules. The language sometimes gives the impression that desolvation is being treated explicitly, whereas the model introduces desolvation-inspired effective interactions into an implicit-solvent framework.

      We agree that the current wording should more clearly reflect the nature of our model. The desolvation terms introduced in this work are effective coarse-grained interaction terms rather than an explicit molecular representation of water. In the revised manuscript, we will carefully revise the language throughout the article to describe the model as incorporating desolvation-inspired effective interactions within an implicit-solvent coarse-grained framework.

      Second, the conformational analysis is interesting, but the broader context of prior work on dilute-to-dense phase conformational reorganization of IDPs could be more clearly discussed. This would help clarify what is new in the present work, whether it is the conformational change itself, its dependence on desolvation terms, or the proposed scaling with distance from the critical point.

      We appreciate this suggestion. In the revised manuscript, we will place the conformational analysis in the context of prior work and discuss the observed conformational changes more explicitly from the perspective of desolvation-inspired interactions. We will also clarify the assumptions behind the scaling relation between conformational change and thermal distance from the critical point.

      Third, the dynamic results are potentially useful, but the manuscript should more clearly articulate what is nontrivial beyond the expected slowing of local rearrangements by an added barrier in the potential.

      We thank the reviewer for the suggestion. In the revised manuscript, we will clarify which aspects of the observed dynamics can be directly expected from the added desolvation barrier and which trends arise from the combined effects of desolvation on packing density, chain mobility, kinetic arrest, and coarsening.

      We again thank the editors and reviewers for their constructive comments and suggestions. We believe that the planned revisions will improve the precision of the model description, clarify the physical interpretation of the desolvation-inspired terms, expand the relevant literature context, and better define the scope and limitations of the current framework.

    1. eLife Assessment

      This study presents a valuable framework that uses anticipatory eye movements to track how expectations are formed and revised during implicit probabilistic sequence learning. The evidence supporting a behavioural dissociation between errors arising from environmental noise and errors reflecting an inaccurate internal model is solid, but the oculomotor data describe behaviour rather than explain the underlying computational mechanisms, and the stronger mechanistic claims - that learning is more repetition-based than error-driven - remain incomplete without formal comparison against computational models of error-driven learning. The emerging reaction-time difference between conditions appears driven by slowing to low-probability stimuli rather than facilitation of high-probability ones, an asymmetry that requires decomposition and consideration of alternative explanations. The potential contamination of the anticipatory measure by starting gaze position should be addressed through control analyses, and the "process-pure" framing should be tempered, given that oculomotor behaviour is itself subject to motor learning.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript presents an original quantitative approach for tracking the online formation and updating of prior beliefs. In an Alternating Serial Reaction Time task, participants were exposed to probabilistic visual streams, and their pre-stimulus saccadic behavior (i.e., the first eye movement after the previous stimulus disappeared) was monitored via eye-tracking. Since the stimuli followed an alternating probabilistic sequence, upcoming events did not appear with full certainty: some stimuli had a higher, some a lower probability. By comparing anticipatory oculomotor behavior between high and low probability events, the authors dissociated between learning/belief updating and general oculomotor noise. Noise-driven errors were more frequent than learning-dependent errors, which nonetheless triggered more belief updating (i.e., a change in oculomotor behavior in a subsequent encounter of the same event). Interestingly, updating depended more strongly on whether a prior belief was consistent with the task's probabilistic structure than on prediction errors. These findings suggest that incidental, implicit statistical learning may rely on conservative updating with a relatively low learning rate, or on errorless algorithms, rather than prediction errors per se.

      Strengths:

      By applying a fine-grained analysis of anticipatory oculomotor behavior, this work establishes new continuous metrics to quantify the gradual learning and refinement of prior expectations during statistical learning. These metrics provide convincing evidence of the dynamics of anticipatory oculomotor behavior.

      The method is paradigm-independent, offering generalizable metrics for tracking the dynamic formation and refinement of predictive models in any task involving probabilistic stimulus streams. In the future, computational modeling may leverage these continuous metrics to better dissect the mechanisms underlying statistical learning.

      Weaknesses:

      The authors subscribe to the idea that statistical learning is not a unified concept but rather is implemented via multiple underlying mechanisms. However, it remains unspecified what these different mechanisms could be, and how eye movements could contribute to distinguishing between them.

      The authors claim that they developed a novel methodological approach to probe whether anticipatory eye movements directly reflect priors, thereby filling an outstanding gap. However, this claim ignores mounting relevant work on structure learning using eye-tracking in the developmental field.

      The authors claim that their framework quantifies trial-by-trial oculomotor dynamics, while in fact the analyses use epochs (i.e. groups of multiple trials) as predictors. Why not use trial number as a predictor to truly investigate trial-by-trial dynamics that directly reflect anticipation, surprisal, and revision?

    3. Reviewer #2 (Public review):

      Summary:

      Hann and colleagues introduce a gaze-based analytical framework designed to capture, on a trial-by-trial basis, how people form and revise their predictions during implicit probabilistic sequence learning. Using an eye-tracking adaptation of an alternating sequence task, they record the first anticipatory saccade during the response-stimulus interval and classify each such saccade along two dimensions: whether it was directed toward a high- or low-probability upcoming stimulus (the learning-dependent vs. not-learning-dependent distinction), and whether the anticipated location coincided with the stimulus that actually appeared. A complementary iterative-updating metric codes whether a participant's prediction for a given three-element context is repeated or revised on successive encounters of that context.

      On the basis of these measures, the authors report that errors congruent with the inferred regularity - which they interpret as reflecting environmental noise - become progressively more frequent than errors reflecting an inaccurate internal model; that participants show a pronounced tendency to repeat their previous prediction rather than revise it; and that updates depend more on whether a prior belief is congruent with the task's statistical structure than on whether the previous prediction was confirmed. They interpret these results as evidence that statistical learning is less error-driven and more repetition-based (Hebbian in character) than is typically assumed.

      Strengths:

      The methodological ambition of the work is considerable, and the paper makes several contributions that are likely to be useful to the implicit-learning and predictive-processing communities. Using the first anticipatory saccade as a pre-response behavioral readout of prediction is conceptually well-motivated: it provides a trial-by-trial index of predictive orienting at a temporal resolution that manual reaction times cannot deliver, and it does so before the outcome of the trial is known. The explicit distinction between errors arising because the task's outcome is stochastic - that is, predictions congruent with the statistical structure but unconfirmed by the stochastic sample - and errors arising because the internal model is inaccurate is a theoretically meaningful move: predictive-coding and Bayesian accounts have long argued that these two sources of surprise should carry different weight for model revision, and the authors offer a behavioral operationalization of that distinction. The analytical pipeline is not tied to the specific paradigm used here and could be applied to other probabilistic sequence-learning tasks, which gives it broader methodological utility than a single-paradigm report. Finally, the demonstration that learners maintain their prior across successive occurrences of the same context, even when it has been disconfirmed by the most recent outcome, is a robust behavioral observation that speaks directly to an unresolved debate about whether statistical learning is dominantly error-driven.

      Weaknesses:

      The framework and the core behavioral observations are valuable, but several inferential steps - from the gaze signal to the cognitive constructs the authors invoke - are not fully supported by the present design, and these gaps affect how readers should interpret the stronger theoretical conclusions.

      The "process-pure" framing conflates sensitivity with construct purity. The authors repeatedly describe the eye-tracking measure as providing a more process-pure index of statistical learning than manual-response paradigms. Anticipatory saccades are themselves a learned motor behavior - the oculomotor system is among the most plastic motor outputs the primate brain generates, and sequence learning in the saccadic system is well-documented. The present design does not dissociate learning of the statistical structure from learning of the oculomotor sequence that expresses it, so the measure is not, on its face, free from the motor-learning confound that the authors criticize in button-press paradigms. The framing should be read as aspirational rather than as demonstrated by the present data.

      The oculomotor reaction-time data do not show the canonical signature of statistical learning. Reaction times for low-probability trials rise across epochs while those for high-probability trials remain approximately flat (Figure 5). The emerging difference between the two trial types, therefore, appears to be driven by a slowing of responses to low-probability stimuli rather than by a facilitation of responses to high-probability ones, and the authors do not rule out the alternative interpretations that this pattern reflects fatigue, a motor floor effect, or inhibition of unexpected locations. Because no fixation constraint is imposed during the response-stimulus interval, pre-stimulus gaze drift toward the anticipated location will artifactually reduce reaction time on precisely those trials the authors wish to treat as learning-driven; the fact that measured reaction times remain well above zero even on trials classified as correct anticipations is itself evidence that this contamination is present. The oculomotor reaction-time data, therefore, do not provide as clean a verification of learning as the manuscript implies.

      The correct/error labeling of anticipatory saccades incorporates information that the participant did not have. Because the first saccade occurs during the response-stimulus interval - that is, before the upcoming stimulus is revealed - the participant's internal predictive state is identical whether the trial is subsequently classified as a learning-dependent correct response or a learning-dependent error. Any difference in the epochwise frequency of these two categories must therefore be driven, at least in part, by the external stochastic structure of the task rather than by a difference in the predictive process itself. In particular, the observation that learning-dependent errors are the most frequent saccade type (Figure 7) is predicted by the prior probabilities of the outcomes alone, given a high-probability prediction, without appeal to any difference in predictive state. Readers should recognize that the theoretically meaningful contrast is between learning-dependent and not-learning-dependent anticipations (two categories), and that the four-way split risks confounding predictive state with outcome stochasticity.

      The iterative-updating metric does not distinguish prior revision from alternative processes. The binary update / no-update code, computed across non-contiguous occurrences of the same three-element context, does not discriminate between a genuine update of the internal model, simple episodic retrieval of a previously encountered triplet, and oculomotor perseveration. Without a formal generative model to anchor the interpretation, the central theoretical claim - that statistical learning is less error-driven than commonly assumed - is underdetermined by the data. The repetition pattern the authors observe is equally consistent with an error-driven model equipped with a low learning rate in a stable environment, an interpretation the authors themselves acknowledge in the Discussion. Adjudicating between these possibilities requires comparison against explicit computational models, which the present manuscript does not provide.

      Data loss and the absence of fixation control. An interpretable saccade is detected on fewer than half of all trials (48.76%; line 889), and the manuscript does not report the distribution of saccade counts per interval, the per-condition trial counts after all exclusions, or the decomposition of the 20% missing-data threshold into its underlying causes. Given that the entire inferential apparatus rests on this subset of trials, the degree of data loss is a relevant context for the reader. Separately, no fixation constraint is imposed between trials: the participant's starting gaze position at the onset of each response-stimulus interval is whatever position was reached at the end of the preceding response, and this starting position carries trial-history information correlated with the upcoming stimulus. This leaves open the possibility that what is classified as predictive orienting partly reflects the mechanical consequences of where the eye happened to be at the end of the previous trial. The authors defend the absence of a fixation cross on the grounds that it would transform the transitional structure of the task, but this is an empirical claim presented without a supporting citation.

      Heterogeneity within the high-probability condition is not addressed. The two routes to a high-probability triplet in the design - pattern-random-pattern (50% of trials) and random-pattern-random (12.5%) - differ both in their base rate and in the reliability of the contextual cue they provide. Collapsing across these subtypes is an analytical choice that may conceal heterogeneity in the underlying learning process.

      Appraisal: Do the results support the authors' conclusions?

      The framework succeeds in providing a trial-by-trial behavioral readout of predictive orienting that is more fine-grained than conventional reaction-time measures, and the behavioral dissociation between errors congruent with the regularity and errors reflecting an inaccurate internal model is a genuine empirical contribution. The conclusions about the mechanistic nature of statistical learning should be read as motivating hypotheses for future modeling work rather than as settled empirical claims.

      Impact and utility:

      The analytical framework introduced here is likely to be useful to researchers working on implicit learning, predictive processing, and Bayesian models of perception and cognition. The measure of predictive orienting and the iterative-updating code could be adapted to a range of probabilistic learning paradigms, and the behavioral dissociation between noise-driven and model-mismatch errors fills a methodological gap that the field has long acknowledged. The authors share their data and code openly, which will facilitate reuse. The most durable contribution of the paper is methodological; the theoretical claims about the nature of statistical learning will require additional computational modeling before they can be regarded as established.

    4. Author response:

      We thank the Reviewers for their time and effort reviewing our manuscript, we are particularly thankful for the literature recommendations of Reviewer 1, and the analysis ideas of Reviewer 2.

      We are glad that both Reviewers agree that the method we developed provides value to the field. We furthermore agree that our theoretical claims and conclusions could be supported by further analyses. Thus, we primarily plan to focus on this.

      We plan to strengthen our statements by:

      - Comparing our metrics to those of alternative learning processes and hypotheses

      - Additional analyses, including ones using standardized learning scores, collapsed saccade likelihoods for learning-dependent and not-learning-dependent saccades, angular deviations instead of the binary update variable, and a breakdown of high-probability triplets into ones that end with a pattern element or a random one.

      - Adding further information regarding saccades, trials without saccades, and saccade starting points.

      Furthermore, we plan to strengthen our Methods section: some of the Reviewers’ points potentially stem from our unclear description of the ASRT task, thus, the Task & Procedure section needs deeper and clearer explanations. Lastly, we will extend the Introduction, citing the literature recommended in the reviews, which indeed could provide further depth.

    1. eLife Assessment

      This valuable study uses large-scale 7T naturalistic fMRI data and nonlinear pRF modeling to map the tonotopic organization of the human auditory cortex, linking spectral tuning to speech selectivity and cortical hierarchy. The evidence is solid, demonstrating that movie-based stimuli can recover robust population-level auditory maps and offering tools for leveraging existing datasets, although there is room for improvement in relating static tonotopy to dynamic speech processing and in presentation clarity. The study will be of interest to a broad audience working on auditory cortex organization and mapping.

    2. Reviewer #1 (Public review):

      This paper reports an auditory-directed analysis of the HCP 7T short movie dataset. It has the goal of using the film audio to create tonotopic (pRF) maps and combine these with other HCP-provided data (e.g., T1/T2 ratio) to improve understanding of auditory cortex organization and relative functional segregation, particularly in reference to speech processing.

      The paper is ambitious, uses well-founded existing tools for combining data across subjects, and in the Discussion in particular, makes a lot of careful points about interpretation. The paper shows that, at least for a very large dataset on 7T (and for at least a few individual participants) good quality cross-subject-average tonotopic maps can be extracted from fMRI movie datasets via basic spectral modelling of the movie soundtracks. It also suggests ways that these movie-based maps can be combined to come up with potential models of cortical organization. The PCA analysis is a creative way of combining maps (see below for comments)

      These are valuable tools for the field in exploiting/exploring existing data, and I look forward to trying them out myself. I want to emphasize that this is not 'damning with faint praise' - a concrete demonstration of this approach with freely available tools/examples is not only the product of a lot of effort (thank you!), but will be an impetus to research going forward.

      In terms of the contribution to our understanding of auditory cortex organization, using this large N cohort, they replicate a number of findings in the literature from the last couple of decades, including the overlap of low frequency preference with greater speech stimulus preference (e.g. Moerel, de Martino, & Formisano, 2012, J Neuro), patterns of BF width across cortex (Moerel et al., various; Thomas et al. 2015), use of shorter and longer natural sounds (Moerel et al., 2012, 2014; Dick et al., 2012), the importance/influence of sustained spectral attention for tonotopic mapping (da Costa et al., 2013; Dick et al., 2017; Riecke et al. 2017), the use of tonotopy and 'myelin' mapping to establish areal or regional boundaries (Dick et al., 2012; de Martino et al., 2015; Besle et al., 2018, etc) and the overall shape and consistency of tonotopic maps (e.g., Talavage et al., 2004, Humphries et al., 2010 and many others). To my knowledge/memory, this is the first tonotopy paper that has used the cross-subject cortical-surface-based averaging techniques that are driven by more than curvature/sulcal alignment.

      The paper focuses in particular on creating new sets of ROIs based on the various maps derived from the data. Despite being quite familiar with this body of work, I found it difficult to follow how the ROIs were derived, and how and why they were different and/or an improvement over existing parcellation schemes (see for instance Sereno, Sood, & Huang, 2022 for a comprehensive parcellation framework across modalities including auditory, based on combined receptive surface mapping, myelin estimates, and other metrics).

      Given the hour of fast(ish) fMRI data on a 7T with pretty big voxels (so high SNR), one aspect of the results that I found surprising - and potentially informative - was the lack of reliable tonotopic 'mappability' in the majority of participants. The authors' analytic approach to computing the pRFs seems completely reasonable (and shows good average maps), and yet individual maps seem unreliable except for the very best examples. I wondered if this might be due to problems in data collection with earbuds becoming slightly uncoupled and therefore delivering a lot less lower-frequency response and also not preventing scanner noise from getting to the ear; this is often a problem with any in-scanner earbud system (including the Sensimetrics). I wondered if the robustness of the 'speech maps' was associated with that of tonotopy; if they are highly associated, that would suggest that either there were huge individual differences in auditory attention, or perhaps that there was some variability in the acoustic signal delivered to each participant.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors leverage a high-powered 7T fMRI dataset of subjects viewing naturalistic audiovisual movies to elucidate the topographic organization of the human auditory cortex. By applying a nonlinear pRF model, they successfully map tonotopic gradients extending beyond the auditory core into the STG and STS areas. A primary finding is a medial-to-lateral gradient of increasing response compressivity, which the authors claim mirrors the hierarchical cascade architecture of the visual system. Furthermore, the modeling reveals that regions exhibiting high speech selectivity predominantly occupy the low-frequency portions of non-primary tonotopic maps. The authors argue that this architecture reflects an efficient coding mechanism where the cortex magnifies specific spectral features to facilitate the transition from acoustic encoding to flexible speech representation.

      Overall, the study presents concise analyses and compelling high-resolution results that advance our understanding of auditory cortical organization. However, the manuscript currently exhibits several significant theoretical and methodological gaps that temper its broader claims. Most notably, the authors' reliance on a spatial, retinotopic-like analogy overlooks the fundamentally temporal nature of audition. Decoding continuous, natural speech relies heavily on dynamic, full-spectrum temporal integration and contextual recurrent computations, which are difficult to reconcile with the purely static, low-frequency spatial tuning observed here.

      Strengths:

      (1) The utilization of ultra-high-field 7T functional imaging combined with large-scale, naturalistic continuous stimuli provides an excellent signal-to-noise ratio and captures cortical responses under ecologically valid conditions.

      (2) The application of a non-linear pRF encoding model provides a robust, quantitative method for parameterizing and mapping tonotopic features across the cortex, moving beyond simple contrast-based parcellations.

      (3) The manuscript effectively demonstrates the relationship between category selectivity (e.g., speech) and underlying tonotopy, drawing an elegant and structurally useful analogy to the well-established relationship between category selectivity and retinotopy in the visual cortex.

      Weaknesses:

      (1) While the PCA mapping of the functional and structural parameter space is visually compelling, the robustness of this representational geometry across varying acoustic contexts remains ambiguous. Because the model relies on the specific statistical regularities of a single naturalistic audiovisual stimulus set, it is unclear if this low-dimensional structure would hold when tested against isolated speech sounds, environmental noise, or spectrally matched non-speech control stimuli.

      (2) The methodological descriptions currently lack the computational precision required for replication and deep evaluation. I would suggest that the exact mathematical formulation of the encoding model be fully specified in the Methods section. This should include an explicit definition of the objective function, a clear accounting of all terms and hyperparameters utilized during the fitting process, and the exact dimensionalities of both the input feature space and the resulting parameter space.

      (3) There is a critical theoretical disconnect between the observed static, low-frequency tuning in the STG and the known acoustic requirements for continuous speech perception. Speech is a full-spectrum signal; while fundamental frequencies and formants dominate the lower spectrum (which is vital for processing dynamic pitch contours), high-frequency bands (>1 kHz) carry indispensable phonetic information, such as the rapid spectrotemporal dynamics of consonants, especially fricatives. If the speech-responsive cortex is primarily and statically tuned to a low-frequency spectrum, it is unclear how the dynamic, high-frequency spectral information required for semantic decoding is represented. A rich body of electrophysiological literature documents diverse spectrogram coding in the STG. For example, Mesgarani et al. (Science, 2014) demonstrated using spectrotemporal receptive field models that neural populations in the STG are tuned to both low and high-frequency spectrograms well above 1 kHz. The authors must address this discrepancy and attempt to reconcile their static tonotopic findings with the existing literature on dynamic speech encoding.

      (4) While drawing parallels between visual and auditory processing hierarchies is conceptually attractive, the modalities face fundamentally different computational challenges. Vision is largely resolved in space, making a retinotopic spatial coding strategy ecologically and computationally sound. Audition, however, evolves continuously in time. Complex temporal structure, continuous temporal integration, and contextual recurrent computations are paramount for auditory processing, particularly for speech comprehension. In this sense, a purely spatial or tonotopic coding framework is insufficient to fully explain the complex temporal processing dynamics required in the higher-order auditory domain.

    4. Reviewer #3 (Public review):

      Summary:

      The work has the potential to identify the topographical organization of the auditory cortex, which remains controversial with current unnaturalistic sound stimulation, using an elegant approach developed in the visual domain with population receptive field mapping to study the organization of the visual system with naturalistic stimulation conditions.

      Strengths:

      This work presents an analysis of the topographic study of auditory cortical organization, using a substantial Human Connectome Project 7-Tesla functional imaging dataset in which 174 participants viewed naturalistic movies.

      Weaknesses:

      The key issue for the paper is that even the authors seem undecided on what the topographical results are and whether these results are consistent with, refute, or expand our notion of human auditory cortical field organization using this massive dataset obtained under movie-watching conditions. Short of this clarity, and much of the discussion of the issues surrounding topographic mapping is buried in the Supplementary materials section, it is not clear what the authors think the advance of the current work is beyond the large datasets.

      On the flip side, there is little consideration of the challenges of mapping the auditory cortex using naturalistic stimuli that prevent dissociating visual from auditory stimulation conditions, contributing to this clarity or lack thereof in tonotopic mapping.

      As such, the current manuscript struggles to achieve its full potential.

    1. eLife Assessment

      This important study identifies two pairs of dopaminergic neurons (DA-WED) in Drosophila that coordinate cardiac deceleration and locomotor responses to a mechanical threat. The evidence is convincing, supported by comprehensive optogenetic, physiological, and behavioral experiments showing that these neurons are required for and sufficient to drive threat-associated cardiac slowing. The proposed role of cardiac deceleration as an interoceptive contributor to locomotion is intriguing, but should be presented more cautiously, as the causal relationship between heartbeat changes and locomotor output remains less directly established. The work will be of broad interest to those interested in neural circuits, neuromodulation, and the integration of physiological and behavioral responses.

    2. Reviewer #1 (Public review):

      Summary:

      This study by Tsuji et al. explores a mechanical threat model in Drosophila using air puffs as a stimulus. The authors first establish the paradigm and show that air puffs induce cardiac deceleration along with increased locomotion. They then identify dopamine as a key regulator of this response and go on to map the underlying circuit. In doing so, they pinpoint two pairs of DA-WED neurons as critical players. They carefully used intersectional strategies to achieve relatively clean labeling of these neurons, which helps ensure that the observed effects can be attributed specifically to DA-WED neurons. They further show that DA-WED neurons are both required and sufficient to drive cardiac deceleration, and that their activity increases in response to air puff stimulation. These neurons also contribute to the locomotor response. Directly inducing cardiac deceleration via optogenetic manipulation of cardiomyocytes also increases locomotion, suggesting a link between cardiac state and behavioral output.

      Strengths:

      Overall, the experiments are thoughtfully designed, well-controlled, and clearly presented. The figures are easy to follow, and the conclusions are generally well supported by the data. The manuscript is also clearly written, with a discussion that acknowledges potential caveats and outlines future directions. The genetic tools, behavioral paradigm, heart rate measurement approaches, and stimulation methods introduced here will be valuable resources for the community.

      Weaknesses:

      A few minor points to add to the clarity of the manuscript:

      (1) The DA-WED driver (R48A08-AD ∩ VT008692-DBD ∩ TH-FLP) appears quite clean in the brain. However, since the study focuses on cardiac function and locomotion, it would be helpful to check expression in cardiomyocytes and the ventral nerve cord. This would help rule out any off-target expression that might contribute to the phenotypes and further support the idea of a descending pathway from brain dopaminergic neurons.

      (2) Since DA-WED>Kir2.1 abolishes the puff-induced locomotor response (Figure 5b), suggesting that DA-WED neurons are directly involved in mediating locomotion. In the model (Figure 5L), it might make more sense for the pathway from mechanical threat to locomotion to pass through DA-WED neurons. The authors could consider adjusting the schematic if they agree.

      (3) In line 408, Figure 5K should be 5L as it's a discussion of the model.

      (4) In Figure 5j, the x-axis is missing time labels. Even if it matches Figure 5h, adding labels would make it easier to interpret at a glance.

      (5) In line 312, it would be helpful to briefly explain why a 28 ms light pulse was used, compared to other pulse durations elsewhere in the paper.

      (6) The cardiac deceleration seems to recover quickly after the air puff ends, whereas the locomotor response persists longer (around 10-15 seconds; see Figure 1 and Figure 5). This difference might suggest that DA-WED neurons influence locomotion through an additional or partially independent pathway, beyond their role in cardiac regulation. It could be worth briefly discussing this possibility.

    3. Reviewer #2 (Public review):

      Summary:

      The authors study cardiac deceleration during threat responses in Drosophila. Particularly, it focuses on identifying the neuronal control of this deceleration. Using behavioral and cardiac tracking and analysis, genetics, and calcium imaging, they identify two pairs of dopaminergic neurons involved in cardiac deceleration during air puff responses

      Strengths:

      The study is overall well done, and the paper is clearly written. Particularly, the work on identifying the two pairs of dopaminergic neurons involved in cardiac deceleration using a series of drivers and generating new ones is rigorous and extensive. Finally, the authors manipulate the heartbeat to investigate how it influences threat responses

      Weaknesses:

      There are, however, several points that need to be clarified, as some claims are not entirely supported by evidence.

      The authors, for example, claim that dopaminergic neurons are responsible for cardiac deceleration (during the air puff, lines 182-3, page 9). However, based on the work in this study, it seems that other neurons could be involved in this control as well. In addition to dopaminergic neurons, the authors test serotonergic and octopaminergic neurons, which, based on silencing experiments, also show an implication in heart-beat deceleration. Furthermore, because they find that dopaminergic neurons are the only ones that, upon thermogenetic activation, lead to lower heart beat frequency, they conclude that the dopaminergic neurons are responsible for air -puff induced cardiac deceleration.

      However, these activation experiments are done in a different context than the air puff experiments (at a higher temperature, which could have an effect on the heartbeat changes upon activation of different neuron groups), and because silencing of other monoaminergic neuron types during the air puff also resulted in less cardiac deceleration, one cannot exclude the implication of octopaminergic or serotonergic neurons in air-puff-induced deceleration.

      Activation experiments without high temperatures (using, for example, optogenetics) and/or in the presence of the air puff would be important to determine that the dopaminergic neurons are the main type of monoaminergic neurons involved in air-puff-induced cardiac deceleration. Otherwise, the related claims should be rephrased in a way that clearly doesn't exclude a possible implication of other monoaminergic neurons.

      Regarding the interactions between the cardiac deceleration and locomotion, the authors propose, based on the results, that the optogenetic cardiac deceleration is sufficient to induce an increase in locomotion, and that it is the decrease in heartbeat that would be responsible via interoceptive pathways to trigger an increase in locomotion. In the model they propose, the DA-WED neurons would induce a decrease in heartbeat that, in turn, would trigger an increase in locomotion. There is not enough proof that cardiac deceleration is the one that triggers an increase in locomotion during air puff responses. As the authors themselves state, the experiments that would demonstrate this would involve preventing cardiac deceleration while optogenetically activating DA-WED. It can therefore not be excluded that the DA-WED neurons trigger an increase in locomotion that is possibly modulated by the cardiac activity. Both alternatives should be considered (models in Figures 4 and 5).

    4. Reviewer #3 (Public review):

      Summary:

      In this elegant study, Tsuji et al. identify a relationship in Drosophila between cardiodynamics and threatening stimuli where mild air puffs elicit a brief bradycardia that coincides with locomotion increases. They then take advantage of the arsenal of genetic tools available in the fruit fly to reveal the indispensability of dopamine, through the action of Dop1R2, in this phenomenon. Further, they pinpoint the source of this dopamine to two specific pairs of neurons - DA-WED that are threat-activated. They then test and find a potential role for cardiac interoception from the heart in linking behavior and cardiodynamics.

      Strengths:

      This is an interesting and timely story that brings together the tools of fruit fly systems neuroscience and links it with physiology. The experiments are well done and tell a very nice story. In particular, the primary message of the paper - that the authors have identified specific dopaminergic neurons that regulate cardiac activity - is sound.

      Weaknesses:

      There are no important problems with the scientific approach. Rather, there are some interpretive changes I would consider.

      (1) The changes in heart rate are small (10% or so), and, as far as I can tell, are evident for a beat or two. So the data may be better interpreted not as a change in rate but as a lengthening of diastole for a beat or two. That may seem a petty difference, but it might point to particular stretch-activated systems or changes in blood flow as the determinant.

      Heart rate must be averaged over time, and so might be blurring the effects. It may be useful to produce figures centered on beat count and duration rather than time. Because the effect may even be just on a single beat, we suggest the authors try plotting the average beat duration for each beat that follows the air puff. If it's really just the first beat, using a quantification of the change of this duration relative to the average that precedes the puff may produce more striking figures.

      (2) The author's model that cardiac deceleration leads to walking data is only partially supported by their data. In the first figure, the relationship between cardiac deceleration and walking probability seems to be inverted relative to their model (weak stimulus -> strong cardiac effect and weak locomotor effect; strong stimulus-> weak cardiac effect and strong locomotor effect). It is possible that this discrepancy may disappear when the authors look at beat duration rather than heart rate (for instance, if following the strong stimulus, there is a very long beat that is followed by tachycardia, thus weakening their observed HR change). It would also be easier to relate this data in Figure 1 to their interoceptive model if some data were shown that illustrated the relative timing of the cardiac change and the locomotor start.

      (3) Also, since the locomotor and cardiac changes are probabilistic, it would be very useful to see how their respective probabilities change when conditioned on the other. According to their interoceptive model, locomotion should preferentially increase on trials where cardiac deceleration occurs. The authors should discuss this incongruity and also potential alternative interpretations of their cardiac manipulation experiments. Perhaps the bradycardia makes them more sensitive to threats - as suggested in the introduction? Control flies show a mild increase in locomotion following green light (Figure 5j), so perhaps by slowing the heart, they are more sensitive and thus respond more strongly to this stimulus?

      (4) Looking at the example shapes of the beats in Figure 5g versus Figure 1c, the optogenetically induced diastole has a very different shape from the naturally occurring long beat. Thus, the exact cardiac stimulus may be unnatural. If this is true across trials and animals, it may be worth considering that the funny beat (like an anxiogenic atrial fibrillation in mammals) is the source of the fear and, in turn, locomotor behavior (also interesting!) rather than being a true replication of the cardiac events seen following the puff stimulus.

    1. eLife Assessment

      This study proposes that fitness level influences exercise-induced hypoalgesia in women. However, the evidence to support this claim is incomplete: the conclusions rely on a small interaction that emerges only under specific conditions and are incongruent with the title, the findings are inconsistent across pain modalities and stimulus intensities, the analysis approach does not fully exploit the continuous pain ratings collected, and the absence of a baseline condition limits the interpretability of results as reflecting true hypoalgesia. Additionally, the methods by which fitness level was categorized across cohorts can be questioned, and the results and figures do not clearly illustrate how between-group comparisons were conducted. With a proper revision, it could be useful for sports medicine practitioners to consider how they administer exercise protocols to help those experiencing pain.

    2. Reviewer #1 (Public review):

      Summary:

      The current study is a follow-up to a previously published study by the same research group (Nold et al. 2025). In the previous study, the authors had included a set of exploratory analyses which assessed the effects of fitness level (denominated by a relative FTP), sex, and drug treatment (Naxolone versus placebo). In this previous study, the authors state that "exploratory analysis showed a significant main effect of fitness level on differences in pain ratings in the [saline] condition... suggesting increased hypoalgesia with increasing fitness levels, pooled across all stimulus intensities".

      In the current study, the authors have recruited an additional 22 female participants (21 included in analysis) from local cycling clubs to assess if fitness level does indeed impact exercise-induced hypoalgesia responses to experimental thermal and pressure pain models.

      Strengths:

      The current study has the potential to present a convincing argument about the effect of fitness level and potentially other factors (e.g., sex) on exercise-induced hypoalgesia responses. Combining data across two of their primary studies would be highly fruitful to the research community interested in this area. Specifically, it has the potential to inform sports medicine practitioners and how they administer exercise protocols to help those experiencing pain with a further consideration for the fitness level (and maybe sex) of their patients.

      Weaknesses:

      However, the current study makes several bold claims about the role of fitness level and sex on exercise-induced hypoalgesia, which I do not believe that this study on its own - or in conjunction with the previously published study by the same authors - can make at present. Namely, the current study does not appear to conduct any specific analyses between the cohorts from either study (current and present). The results mention a difference in the group mean values in "fitness level" between cohorts, but the analysis itself on pain responses/exercise-induced hypoalgesia is limited only to the cohort from the current study. If the authors wanted to provide a convincing argument that fitness level has an effect on exercise-induced hypoalgesia, then the analysis of this study would have to include an analysis between the groups considered to be of "high" and "low" fitness level. I do not think the current study does this. Instead, it makes an assumption from the previous study (Nold et al. 2025) which only states that "exploratory analysis showed a significant main effect of fitness level on differences in pain ratings in the [saline] condition... suggesting increased hypoalgesia with increasing fitness levels, pooled across all stimulus intensities". The analysis of this study would have to include fitness level "high fitness" versus "low fitness" of participants across both studies in its statistical model to properly discern if fitness level has an impact on exercise-induced hypoalgesia.

      A similar comment can be made with respect to sex differences, as these have not been assessed in the analysis of this study either.

      Another area of weakness in this study is how "fitness level" has been demarcated across participants. One issue is how authors have assumed that the current cohort is 'fit', whereas the previous cohort was 'less fit', meaning that the authors could be coming to false conclusions about fitness level. In detail, figures within the current study show a large overlap between the 'fit' and 'less fit' cohorts, where some participants have a higher relative functional threshold power (FTP) in the 'less fit' cohort than the 'fit' cohort and vice versa. Therefore, I believe the authors should better demarcate between those that are in the 'more fit' and 'less fit' groups according to a validated and well-established criterion from the kinesiology and sport science literature. That being said, I think this may be problematic in some ways as FTP is considered a relatively poor measure to denote fitness levels, a limitation highlighted in the previous study's review.

      Altogether, whilst I commend the researchers on their body of work across the two studies, the current methods and analysis provide an incomplete assessment of their primary research question, and therefore, I would urge the authors to reconsider some of their methods/analysis and the framing of their results to better reflect the main research question they have attempted to answer. Likewise, I would recommend that readers ensure they consider the current results with caution until the authors have addressed some areas of concern which currently limit their main conclusions.

    3. Reviewer #2 (Public review):

      This study addresses an important question regarding exercise-induced modulation of pain in women, but the conclusions appear to be based on relatively limited and selective evidence. The authors report an interaction between exercise intensity and stimulus intensity, which they interpret as evidence for exercise-induced hypoalgesia and conclude that fitness, but not sex, modulates this effect. However, this main result relies on a relatively small interaction that emerges only under specific conditions, with inconsistent findings across pain modalities and stimulus intensities, and an analysis approach that does not fully exploit the continuous pain ratings collected. The lack of a baseline condition further limits the interpretability of the findings as reflecting hypoalgesia, and overall, the data provide a rather constrained basis for drawing broader conclusions.

      Strengths:

      (1) The focus on women is important and timely, particularly given the ambiguity in prior findings and the historical bias toward male-dominated samples.

      (2) The attempt to revisit previous findings in a new cohort is valuable in principle.

      Weaknesses:

      (1) The core interpretation may not be fully supported by the data

      The central claim-that the results demonstrate exercise-induced hypoalgesia and its dependence on fitness but not sex-does not appear to be fully supported by the evidence presented.

      1.1 Lack of baseline condition

      The absence of a no-exercise baseline substantially limits interpretation. The study compares high- and low-intensity exercise, but without a baseline, it is not possible to determine whether either condition produces hypoalgesia or hyperalgesia relative to calibration. The observed HI-LI difference, therefore, reflects only a relative contrast between exercise intensities, not an absolute reduction in pain. As a result, attributing the findings to "hypoalgesia" may be difficult to justify fully.

      1.2 Lack of internal replication across conditions

      The reported effect is highly specific and does not clearly generalise across the experimental design. It emerges significantly only for heat pain at the highest stimulus intensity, with no clear effects for other intensities and for pressure pain. Moreover, the main statistical result is a relatively small interaction effect with a modest p value, which translates into a difference of approximately 6-8 VAS units on a 150 scale. This combination-a small effect size, limited statistical strength, and restriction to a single condition-substantially weakens the evidence for a robust or generalisable effect.

      1.3 Deviations from the original study and selective use of data

      Although framed as a follow-up to previous work, the current study introduces substantial methodological changes, particularly in the acquisition and scaling of pain ratings (continuous vs post-hoc ratings, modified VAS with sub-threshold range). Despite collecting rich continuous data, the analysis focuses on peak responses to approximate the previous study. While this may aid comparability, it results in a strong emphasis on a single data point (highest intensity), rather than leveraging the full dataset. This limits both interpretability and comparability.

      1.4 Over-reliance on null results regarding sex differences

      The conclusion that fitness, but not sex, modulates exercise-induced pain may not be directly supported by the data presented. The current study includes only highly fit women, and comparisons with men or less-fit women rely on non-significant differences in a previous cohort. The absence of a significant difference does not provide evidence for equivalence, and no formal statistical support for a null effect is provided. As such, conclusions about the absence of sex differences would unfortunately benefit from more cautious interpretation.

      (2) Limited sample and lack of diversity

      The dataset is narrow in scope, comprising a small sample (N = 21) of healthy, highly fit women. Key demographic characteristics (e.g. age range, BMI distribution) are not fully presented, explored or discussed. This limits generalisability and makes it difficult to draw broader conclusions about exercise-induced pain modulation in women, as the main focus of the study.

      (3) Methodological choices limit the interpretability of the data

      Several methodological decisions would benefit from stronger justification:

      3.1 The use of a non-standard VAS scale (0-150 with a fixed pain threshold at 50) is unconventional and may influence how participants report pain, while limiting comparability with related literature.

      3.2 Participants explicitly reported expecting exercise to reduce pain, introducing a potential confound that is not presently addressed.

      3.3 A more comprehensive use of the full time series of pain ratings would provide a stronger and more transparent basis for interpretation of the present findings.

    1. eLife Assessment

      This study presents important findings on the relationship between nutrient availability and NAD/NADH levels, which in turn regulate biomass production in cancer cells. The authors provide convincing evidence to support their claims, offering insight into why it is difficult to predict which nutrients limit cancer cell growth: both cell type and nutrient availability together determine the oxidative capacity that constrains the synthesis of various metabolic intermediates. The manuscript will be of broad interest to researchers working in cancer and cell metabolism.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how cellular NAD/NADH ratios are controlled in cancer cell lines in vitro. The authors build on previous work, which shows that serine synthesis is sensitive to NAD/NADH ratios and PHGDH expression. Here, the authors demonstrate that serine synthesis is variable across a panel of cell lines, even when controlling for expression of serine synthesis enzymes such as PHGDH. The authors show that cellular NAD/NADH ratios correlate with the ability to synthesize serine and grow in serine-deprived environments when PHGDH levels remain constant. Investigating this variability in NAD/NADH ratios, the authors find that the cells that can positively respond to serine deprivation are able to increase oxygen consumption and cellular NAD/NADH ratios. Cells that do not increase oxygen consumption in response to serine deprivation do not increase NAD/NADH ratios and cannot grow well without serine. The authors go on to show that in cells with the ability to increase oxygen consumption upon serine deprivation, PHGDH expression alone is sufficient to fully restore growth-serine; in cells that cannot increase oxygen consumption, both PHGDH expression and interventions to increase NAD/NADH ratios are required to increase growth. Thus, cells need both PHGDH and NAD/NADH increases to maximize serine synthesis in response to serine deprivation. The authors previously showed that lipid synthesis likewise requires NAD regeneration. Interestingly, one cell line that does not increase oxygen consumption in response to serine limitation tends to increase oxygen consumption in response to lipid deprivation; accordingly, depriving this cell line of lipids increases the synthesis of serine. Together, these findings show that how cells respond to nutrient deprivation is highly variable and that the response to nutrient deprivation (for example, whether or not oxygen consumption is increased) will determine how well cells tolerate depletion of nutrients with related biosynthetic constraints. This work sheds light on the complexity of cancer cell metabolism and helps to explain why it is difficult to predict which nutrients will be limiting to any cancer cell type or environment.

      Strengths:

      (1) The authors use multiple interventions to manipulate NAD/NADH ratios in cells.

      (2) Experiments are well controlled and appropriately interpreted.

      Comments on revised version:

      The authors thoughtfully and thoroughly responded to all reviewer comments. The revised manuscript addresses the critiques.

    3. Reviewer #2 (Public review):

      In the manuscript "Cancer cells differentially modulate mitochondrial respiration to alter redox state and enable biomass synthesis in nutrient-limited environments", Chang et al investigate how cancer cells respond to the limitation of certain environmental nutrients by regulating the cellular NAD+/NADH ratio. They focus on serine and lipid metabolism, pathways known to be controlled by the NAD+/NADH ratio, and propose that changes in mitochondrial respiration in response to deprivation of these nutrients can influence the NAD+/NADH ratio, thereby impacting biomass synthesis.

      While the study is descriptive in nature and does not investigate specific molecular mechanisms that explain the crosstalk between nutrient availability and mitochondrial redox changes, the experimental component is robust, and the conclusions are well supported by the results. Some suggestions could further refine the conclusions and enhance the quality of the manuscript.

      Comments on revised version:

      The authors have provided a very comprehensive response. Their updated paper has improved, and the critiques have been mitigated.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents valuable findings on the relationship between nutrient availability and NAD/NADH levels, which in turn regulate biomass production in cancer cells. The authors provide solid evidence to support their claims, offering insight into why it is difficult to predict which nutrients limit cancer cell growth: both cell type and nutrient availability together determine the oxidative capacity that constrains the synthesis of various metabolic intermediates. The manuscript will be of interest to researchers working in cancer and cell metabolism.

      We thank the eLife Editor for evaluating our manuscript and for the positive comments.

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how cellular NAD/NADH ratios are controlled in cancer cell lines in vitro. The authors build on previous work, which shows that serine synthesis is sensitive to NAD/NADH ratios and PHGDH expression. Here, the authors demonstrate that serine synthesis is variable across a panel of cell lines, even when controlling for expression of serine synthesis enzymes such as PHGDH. The authors show that cellular NAD/NADH ratios correlate with the ability to synthesize serine and grow in serine-deprived environments when PHGDH levels remain constant. Investigating this variability in NAD/NADH ratios, the authors find that the cells that can positively respond to serine deprivation are able to increase oxygen consumption and cellular NAD/NADH ratios. Cells that do not increase oxygen consumption in response to serine deprivation do not increase NAD/NADH ratios and cannot grow well without serine. The authors go on to show that in cells with the ability to increase oxygen consumption upon serine deprivation, PHGDH expression alone is sufficient to fully restore growth-serine; in cells that cannot increase oxygen consumption, both PHGDH expression and interventions to increase NAD/NADH ratios are required to increase growth. Thus, cells need both PHGDH and NAD/NADH increases to maximize serine synthesis in response to serine deprivation. The authors previously showed that lipid synthesis likewise requires NAD regeneration. Interestingly, one cell line that does not increase oxygen consumption in response to serine limitation tends to increase oxygen consumption in response to lipid deprivation; accordingly, depriving this cell line of lipids increases the synthesis of serine. Together, these findings show that how cells respond to nutrient deprivation is highly variable and that the response to nutrient deprivation (for example, whether or not oxygen consumption is increased) will determine how well cells tolerate depletion of nutrients with related biosynthetic constraints. This work sheds light on the complexity of cancer cell metabolism and helps to explain why it is difficult to predict which nutrients will be limiting to any cancer cell type or environment.

      Strengths:

      (1) The authors use multiple interventions to manipulate NAD/NADH ratios in cells.

      (2) Experiments are well controlled and appropriately interpreted.

      Weaknesses:

      Overall the data support the conclusions of the manuscript. I have only two minor comments and suggestions:

      We thank Reviewer 1 for their insightful comments, which have helped us improve the manuscript.

      (1) Figure 2B/C: data are presented as relative to +serine, which shows how some cells respond to -serine, but may also be of interest to see how absolute (not relative) NAD/NADH levels correlate with serine synthesis and serine-independent proliferation. In other words, is it the dynamic increase in the ratio that is most important, or the absolute level of the ratio?

      We thank Reviewer 1 for raising this point about whether it is the absolute NAD+/NADH ratio, or the change in NAD+/NADH ratio, that is important for increasing serine synthesis and allowing proliferation under serine depleted conditions. We reported relative ratios for accessibility to a general audience, but agree that this information is informative and should be presented. We assessed the NAD+/NADH ratio using an enzymatic assay, which does not directly measure absolute concentrations of NAD+ or NADH (PMID: 26232225). However, we previously confirmed the assay is in the same linear range for both NAD+ and NADH, and thus is valid for assessing the NAD+/NADH ratio. We now provide the unnormalized NAD+/NADH ratio data in Supplementary Figure 2G of the revised manuscript. This shows that the considered cells exhibit a range of NAD+/NADH ratios, and redox responsive cells do not cluster in having a higher or lower NAD+/NADH ratio.

      To more formally answer Reviewer 1’s question about whether the absolute ratio or change in ratio is important for increasing serine synthesis, we measured the correlation coefficient between the unnormalized NAD+/NADH ratios and the proliferation rate of all examined cancer cells cultured with or without serine. These data are presented in Author response image 1. Of note, we find that there is a significant positive correlation between the raw values of the measured NAD+/NADH ratio and proliferation rate in both serine-replete (r = .371) and serine depleted (r = .562) conditions. However, this correlation is not strong, and when examining the cancer cells whose proliferation in serine depleted conditions cannot be fully explained by serine synthesis enzyme expression (Calu6, 8988T, A549, MIA PaCa-2, H1299, and HCT116), there is no significant correlation between the raw NAD+/NADH ratio and proliferation rate in serine depleted conditions. The association between the relative change in the NAD+/NADH ratio and proliferation rate is much stronger upon serine deprivation (r = .571), as presented in Figure 2C of the revised manuscript. This suggests that the dynamic increase in the ratio is more tightly linked to the change in serine synthesis rate and proliferation in serine depleted environments, and we discuss this point in the revised manuscript with the following text:

      “Of note, whether the NAD+/NADH ratio of a cell was more or less oxidized in serine-replete conditions was not predictive of response to serine withdrawal (Supplementary Figure 2G).” (Lines 163-165)

      Author response image 1.

      Correlations between unnormalized NAD+/NADH ratios and cell proliferation rates between (A) all cancer cells examined (Calu6, MCF7, MDA-MB-231, A549, 8988T, MIA PaCa-2, A375, H1299, HCT116, MDA-MB-231 with PHGDH overexpression) in serine-replete conditions, (B) all cancer cells examined in serine depleted conditions, and (C) select cancer cells (labeled in gray) where serine synthesis enzyme protein expression does not fully explain proliferation in serine depleted conditions. Pearson correlation coefficient and P values were calculated by simple linear regression, *p<0.05, **p<0.01. Data shown are means of three biological replicates ± SD.

      (2) Line 177-178: the authors write, "We hypothesized that the elevated NAD+/NADH ratio represented a cellular response to make the NAD+/NADH ratio more oxidized to enable serine synthesis". I recommend modest edits to avoid anthropomorphizing. It is possible that the ratio responds for reasons yet to be determined and not necessarily because the cell is deliberately trying to enable serine synthesis.

      We thank Reviewer 1 for raising this point. We agree that our data do not show whether the ratio is elevated for the deliberate purpose of enabling serine synthesis and have edited the text accordingly with the following edit to that line of the revised manuscript:

      “We hypothesized that a more oxidized NAD+/NADH ratio could support greater serine synthesis and thus sought to identify the processes that increase the NAD+/NADH ratio in some but not all cancer cells.” (Lines 190-192)

      Reviewer #2 (Public review):

      In the manuscript "Cancer cells differentially modulate mitochondrial respiration to alter redox state and enable biomass synthesis in nutrient-limited environments", Chang et al investigate how cancer cells respond to the limitation of certain environmental nutrients by regulating the cellular NAD+/NADH ratio. They focus on serine and lipid metabolism, pathways known to be controlled by the NAD+/NADH ratio, and propose that changes in mitochondrial respiration in response to deprivation of these nutrients can influence the NAD+/NADH ratio, thereby impacting biomass synthesis.

      While the study is descriptive in nature and does not investigate specific molecular mechanisms that explain the crosstalk between nutrient availability and mitochondrial redox changes, the experimental component is robust, and the conclusions are well supported by the results. Some suggestions could further refine the conclusions and enhance the quality of the manuscript.

      We thank Reviewer 2 for their time and for their suggestions to improve the manuscript.

      Main critiques:

      (1) Throughout the manuscript, the authors utilise the number of cell doublings per day as an endpoint readout of cell proliferation. It would be advisable to include a quantification of the cell number and to display the proliferation rate over time. This would provide valuable insights into the timeline of cellular responses and avoid potential confounding effects associated with the use of Sulforhodamine B dye, an indirect measure of cell proliferation based on protein content, which may be influenced by some of the interventions. Furthermore, it will help determine whether specific treatments reduce cellular doublings resulting from cell death. This concern is particularly evident in treatments with rotenone, e.g., Fig. 1G, where the increase in doublings could be attributed to cell death.

      We thank the reviewer for this suggestion and agree that assessment of cell count provides additional information beyond Sulforhodamine B dye as an indirect measure of proliferation. To address this, we directly measured cell number over time using Incucyte Live-Cell imaging analysis applied to A549 and H1299 cells cultured with or without serine for 72 hours. Consistent with results using sulforhodamine B, A549 cells doubled at a rate of 0.874 per day and H1299 cells doubled at a rate of 1.034 per day in serine-replete conditions. In serine depleted conditions, A549 cells doubled at a rate of 0.205 per day while H1299 cells doubled at a rate of 0.544 per day. We have added the cell number measurements over time as well as the corresponding calculated doublings per day in Supplementary Figure 2D and Supplementary Figure 2E of the revised manuscript.

      We also agree with Reviewer 2 that serine deprivation and rotenone treatment could potentially impact cell viability, which might confound phenotypes, including NAD+/NADH ratio measurements. To assess whether serine deprivation and rotenone treatment cause cell death, we measured cell viability using Sytox Green after exposing cells to these conditions for 72 hours. We find that there is indeed more cell death in cells cultured without serine at most concentrations of rotenone. However, cell death did not exceed 4% in any of the conditions tested, suggesting this is not a major contributor to the cell doubling phenotypes. These data are now presented in Supplementary Figure 1C of the revised manuscript. However, in light of Reviewer 2’s comments, along with a comment from Reviewer 3 about whether rotenone induces ROS and cellular stress responses, we have decided to remove the proliferation data involving rotenone that were in Figure 1F and 1G of the original manuscript. The rationale is that the potential confounding impacts of rotenone on viability make interpreting the proliferation data difficult. Instead, we have focused Figure 1 of the revised manuscript on the observation that there is specifically a correlation between the cell NAD+/NADH ratio and serine synthesis.

      (2) The authors propose a model in which the deprivation of extracellular nutrients impacts mitochondrial respiration, which in turn increases the NAD+/NADH ratio and ultimately affects metabolic biosynthetic pathways that occur in the cytosol, such as serine biosynthesis. The mechanism by which nutrient availability is sensed and transmitted across different cellular compartments to regulate mitochondrial redox status remains unclear. This concern is particularly relevant for serine metabolism, as its synthesis occurs in the cytosol, but the authors connect it to mitochondrial respiration. Compartment-specific measurements of NAD+/NADH ratio would help to understand to what extent the redox state is affected by nutrients in the mitochondria and in the cytoplasm (see also minor critiques point 2). Moreover, the use of the genetic tool LbNox could be employed to manipulate the NAD+/NADH ratio in a compartmentspecific manner, while also avoiding the toxicity of certain compounds, such as rotenone. This set of experiments would add depth to the investigation, which might otherwise appear too descriptive.

      (A) Compartment-specific measurements of NAD+/NADH ratio would help to understand to what extent the redox state is affected by nutrients in the mitochondria and in the cytoplasm

      The question of how nutrient availability is sensed and transmitted across cellular compartments to impact mitochondrial respiration is important. However, rigorous assessment of compartment-specific metabolism is quite challenging, as we are not aware of tools to accurately measure redox ratios in a compartment-specific manner. Direct assessment of cofactor levels in subcellular compartments requires long isolation times and are unlikely to be accurate (PMID: 27565352). Rapid immunopurification of mitochondria has been used to estimate metabolite levels and ratios, but accurate measurements are hindered by rapid oxidation of NADH to NAD+. The use of fluorescence lifetime imaging (FLIM) to monitor NADH levels does not allow for accurate monitoring of the NAD+/NADH ratio as NAD+ cannot be visualized and NADH cannot be distinguished from NADPH. Additionally, the resolution of FLIM to interrogate compartment-specific signals is limited (PMID: 38594590). Fluorescent sensors, such as SoNar, have been used to image the NAD+/NADH ratio in compartments, though SoNar is sensitive to pH changes, which vary across compartments, and it has been argued that these sensors are more suitable for qualitative, not quantitative, changes in the NAD+/NADH ratio (PMIDs: 25955212, 29181426). It has also been argued that sensors are not amenable to measurement of mitochondrial ratios, as the predicted ratios are too reduced for the range of the sensors. Given these technical limitations, we opted to attempt a rapid subcellular fractionation (~25 second process to separate cytoplasm and mitochondria) followed by enzyme-based measurements of the NAD+/NADH ratio (PMID: 36883551), acknowledging the limitations of this approach. We find that across both A549 and H1299 cells, the mitochondrial NAD+/NADH ratio is lower than the cytosolic NAD+/NADH ratio, as expected. Using this approach, we find that in A549 cells, serine depletion leads to a decreased cytosolic NAD+/NADH ratio compared to serine-replete conditions while having no impact on the mitochondrial NAD+/NADH ratio. On the other hand, serine depletion leads to an elevated cytosolic NAD+/NADH ratio in H1299 cells while also having no impact on the mitochondrial NAD+/NADH ratio. In parallel, we used extracellular pyruvate exposure as a positive control, which should support cytosolic NAD+ regeneration, and rotenone as a negative control, which should suppress mitochondrial NAD+ regeneration. We show that pyruvate led to an elevated cytosolic NAD+/NADH ratio whereas rotenone treatment led to a decreased cytosolic NAD+/NADH ratio. Despite rotenone inhibiting complex I of the electron transport chain, we did not observe a change in the mitochondrial NAD+/NADH ratio (Author response image 2). This likely indicates that this assay is not sensitive enough to detect changes in mitochondrial NAD+/NADH, and we opted not to include these data in the revised manuscript given the limitations of the approach.

      Author response image 2.

      Rapid subcellular fractionation to examine compartment-specific NAD+/NADH ratios. (A) Cytosolic and mitochondrial NAD+/NADH ratios of A549 cells grown with or without serine for 24 hours, n=3. (B) Cytosolic and mitochondrial NAD+/NADH ratios of H1299 cells grown with or without serine for 24 hours, n=3. (C) Cytosolic and mitochondrial NAD+/NADH ratios of H1299 cells treated with either 1 mM pyruvate or 50 nM rotenone for 24 hours, n=3. P-values were calculated using a Student’s t-test, *p<0.05, **p<0.01. Data shown are means ± SD.

      We nevertheless draw the following conclusions from these data:

      (1) Changes to mitochondrial NAD+/NADH either do not occur or are not captured with this approach. Even rotenone treatment, which inhibits complex I and might be expected to change mitochondrial redox state, does not change the measured mitochondrial NAD+/NADH ratio.

      (2) The whole cell NAD+/NADH ratio most likely reflects changes in the cytosolic NAD+/NADH ratio. While observing no impact on the mitochondrial NAD+/NADH ratio after rotenone treatment, we still find the cytosolic NAD+/NADH ratio is decreased. Moreover, both pyruvate and serine depletion led to an elevated cytosolic NAD+/NADH ratio in H1299 cells, which we observe at the whole cell level.

      (3) H1299 cells depleted of serine elevate the cytosolic NAD+/NADH ratio, while rotenone treatment decreased the cytosolic NAD+/NADH ratio despite changes in mitochondrial respiration. This suggests that redox shuttles, such as the malate aspartate shuttle, play a role in communicating changes in mitochondrial redox dynamics to the cytoplasm. We test this hypothesis as described in response to Reviewer 2, point B, below.

      (B) The mechanism by which nutrient availability is sensed and transmitted across different cellular compartments to regulate mitochondrial redox status remains unclear

      Multiple known shuttles are involved in exchanging redox equivalents between the mitochondria and the cytosol. It is likely that multiple shuttles are involved, or could be involved in the right context, but one major shuttle is the malate aspartate shuttle (MAS), and the MAS has been shown previously to support de novo serine synthesis (PMID: 37647199). Thus, we hypothesized that the MAS is involved in the response involving elevated mitochondrial respiration in H1299 cells to increase the whole cell NAD+/NADH ratio upon serine deprivation. To test this, we used CRISPR/Cas9 to generate H1299 cells lacking MAS components GOT1, MDH1, or GOT2 and measured the cell NAD+/NADH ratio. We did not knock out MDH2 given its integral role in the TCA cycle. We find that when MDH1 and GOT2 are knocked-out, H1299 cells no longer exhibit elevated whole cell NAD+/NADH ratios upon serine deprivation. Consistently, removing MDH1 and GOT2 also blunted the increase in oxygen consumption as well as the increase in serine synthesis upon serine deprivation. This suggests that MDH1 and GOT2 activity though the MAS support the process by which mitochondrial NAD+ regeneration is transmitted to the cytoplasm to support serine synthesis. We have added these data as Supplementary Figure 7 in the revised manuscript.

      (C) Moreover, the use of the genetic tool LbNox could be employed to manipulate the NAD+/NADH ratio in a compartment-specific manner

      We thank Reviewer 2 for the suggestion to consider whether LbNOX might be used to manipulate the NAD+/NADH ratio in a compartment-specific manner. We expressed LbNOX in both the cytoplasm and the mitochondria of A549 (serine non-responsive) cells. We predicted that if LbNOX expression, either in the cytoplasm or the mitochondria, affected the NAD+/NADH ratio, proliferation in serine depleted conditions might be improved. However, we found that expressing LbNOX in the cytoplasm or the mitochondria of A549 cells had no effect on the NAD+/NADH ratio. Thus, LbNOX expression in either compartment also did not change proliferation in serine depleted conditions. These data are consistent with the known limitations of this genetic tool. While LbNOX can increase NADH oxidation in response to some interventions like rotenone, it does not necessarily change the NAD+/NADH ratio of unperturbed cells. This was reported in the original description of LbNOX (PMID: 27124460). We confirmed that LbNOX was successfully expressed via immunoblotting, and also confirmed that LbNOX functioned by showing either cytoplasmic or mitochondrial LbNOX expression improves cell proliferation following complex I inhibition. Thus, expressing LbNOX in A549 cells is not informative for understanding compartment specific metabolism following serine deprivation. Nevertheless, as this question is likely to come up for other readers, we have included these data as Supplementary Figure 6 in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor critiques:

      (1) It seems clear from the authors' data that the response to serine depletion in terms of cell proliferation is not determined exclusively by PHGDH levels. It would be useful to measure the levels of the other two enzymes in the serine synthesis pathway and also to measure serine uptake under normal conditions in the different groups of cells. This information could provide some insight into the different responses of cancer cell lines to serine deprivation.

      (A) It would be useful to measure the levels of the other two enzymes in the serine synthesis pathway

      Reviewer 2 raises a fair point, and we agree that measuring levels of other enzymes in the serine synthesis pathway is informative. Thus, we measured the expression of phosphoserine aminotransferase 1 (PSAT1) and phosphoserine phosphatase (PSPH) across all cancer cells examined and find that, similar to PHGDH protein expression, PSAT1 and PSPH protein expression is lower in many cancer cells that are more sensitive to serine withdrawal (e.g. MCF7). However, among the cancer cells where PHGDH protein expression did not explain the response to serine withdrawal, the protein expression of PSAT1 and PSPH also did not explain how well the cells proliferate without environmental serine. These data have been included in Supplementary Figure 2B of the revised manuscript.

      Of note, we measured serine synthesis enzyme expression for the six cancer cell lines whose proliferation in serine depleted conditions better correlated with a change in the NAD+/NADH ratio than it did with PHGDH expression: Calu6, 8988T, A549, MIA PaCa2, H1299, and HCT116. For these cells, we correlated proliferation upon serine depletion with PHGDH, PSAT1, and PSPH protein expression and found that interestingly, there was a significant negative correlation between PHGDH protein expression and proliferation upon serine deprivation. This was not observed for PSAT1 expression, and a statistically significant positive correlation between proliferation and PSPH protein expression was noted, though the variation in PSPH protein expression was large. We have added these correlation data to the revised manuscript as Supplementary Figure 2F.

      (B) It would be useful to measure…serine uptake under normal conditions in the different groups of cells

      Per the Reviewer’s request, we performed absolute quantification of serine uptake rates in serine-replete conditions for three serine “non-responder” cancer cells (Calu6, 8988T, A549) and three serine “responder” cancer cells (MIA PaCa-2, H1299, HCT116). We did not observe a notable difference in serine uptake rate and whether cells responded to serine deprivation. Additionally, with the exception of 8988T cells having a higher serine uptake rate than the other cells, there was no statistical difference in serine uptake across the cancer cells tested (Author response image 3).

      Author response image 3.

      Basal serine uptake rate of exponentially growing cells in serine replete conditions. Serine levels were measured using GC MS before and after 24 hours of serine depletion and normalized by area under the growth curve (PMID: 26954548). P-values were calculated using one-way ANOVA followed by a post-hoc Tukey HSD test, *p<0.05, **p<0.01

      (2) The authors experimentally demonstrated that some cancer cells respond to serine depletion with an increase in mitochondrial respiration, but the molecular mechanism behind this is not addressed. There is some evidence in the literature showing that serine acts as an activator of the glycolytic enzyme PKM, which is coherent with an increased mitochondrial respiration in the absence of serine (PMID: 23064226). The authors could discuss their findings in the context of this paper. Additionally, they could provide some insights about baseline mitochondrial activity in the different cell lines. Indeed, it seems that "redox responsive cells" might have an overall increased basal OCR.

      We appreciate the suggestion that pyruvate kinase M (PKM) may mediate the elevation in mitochondrial respiration in response to serine depletion. Given that serine is an allosteric activator of PKM, and PKM suppression can increase mitochondrial OCR, we discuss this possibility in the Discussion section of the revised manuscript using the following text:

      “Interestingly, serine is an allosteric activator of the glycolytic enzyme pyruvate kinase, which converts phosphoenolpyruvate to pyruvate and generates ATP (Chaneton, 2012). Thus, decreased environment serine availability in addition to differences in pyruvate kinase activity may yield lower glycolytic ATP, resulting in greater mitochondrial respiration in serine redox responder cancer cells.” (Lines 443-447)

      Additionally, we appreciate the reviewer’s observation that redox responsive cells may have an overall increased basal respiration rate. We directly measured mitochondrial dependent oxygen consumption in the same assay to test whether redox responsive cells exhibit higher mitochondrial respiration. We find that while the redox responsive H1299 and MIA-PaCa2 cells have higher mitochondrial respiration than non-responsive cells, HCT116 cells that are also redox responsive to serine deprivation, did not exhibit higher mitochondrial respiration compared to redox non-responsive Calu6, 8988T, and A549 cells (Author response image 4). However, when comparing redox non-responders versus responders as a whole, there was a statistically significant difference in basal OCR. Together, this suggests that basal mitochondrial respiration rate in serine-replete conditions may be related in some cases to whether cancer cells elevate mitochondrial respiration and the NAD+/NADH ratio upon serine deprivation, but this cannot be the full explanation given the HCT116 cell data. We also acknowledge the reviewer’s statement that we do not understand the molecular mechanism by which respiration responds to serine deprivation and explicitly state this in the revised manuscript.

      Author response image 4.

      Basal Oxygen consumption rate (OCR) of cancer cells in serine-replete conditions. (A) Kinetic OCR measurements of cancer cells before and after rotenone and anti-mycin injection, n=8. Data shown are means ± SD. (B) Quantified mitochondrial OCR (removing residual OCR), n=8. Values are averages obtained over three measurements. P-values were calculated via nested ANOVA, ****p<0.001

      (3) There is a discrepancy between the basal values of the OCR from the same cell lines in different experiments, i.e., Figure 3A and Supp. Figure 3C, or in different experiments, Figure 3A, Figure 5E, and Figure 6A. The authors need to comment on/clarify that. Moreover, authors are encouraged to show ECAR values to support the conclusion that lactate production is not differentially affected by serine depletion, and thus, does not contribute to the increase in the NAD+/NADH ratio.

      We recognize the differences in basal OCR values across different experiments. Given experiment-to-experiment variation and the need for different cartridges for each Seahorse experiment, we have found that measured OCR values using Seahorse assays vary across experiments despite the same conditions. Additionally, while we aim to seed the same number of cells per assay, cell seeding and cell quantification after each Seahorse assay can contribute to variation. Given this variability on a per-assay basis, we performed a singular experiment across all examined cancer cell lines considered to minimize variation in oxygen sensor calibration and address the reviewer question about whether absolute differences might contribute to response. These data are shown in Author response image 4.

      Regarding the reviewer’s request to present ECAR data, we note that measuring ECAR is dependent on using unbuffered media and for this reason do not routinely measure ECAR. Our concern is that removing serum from the culture conditions can impact OCR measurements, and we instead prioritized maintaining the same media composition across all sets of experiments (i.e., cell proliferation assays, NAD+/NADH assays, kinetic tracing assays, and OCR measurements). Additionally, we point out that ECAR does not directly measure lactate. We refer the Reviewer to data included in the manuscript where GC-MS was used to directly measure lactate secretion over time for cells cultured with or without serine. These data are presented as Supplementary Figure 3B in the revised manuscript.

      (4) There seems to be also a discrepancy between the levels of M+2 citrate and the fraction labelled (Figure 5C versus Supplementary Figure 6C) in the H1299 cell line upon serine depletion, whereby the M+2 fraction seems unexpectedly lower in serinedeprived cells. In those conditions, H1299 cells showed an increased mitochondrial respiration, which is consistent with increased total citrate levels. This could be explained by a faster TCA cycle activity and the presence of higher-order isotopologues of citrate upon serine starvation. Is this the case? Showing the abundance of the different citrate isotopologues and their contribution to the total pool would help to interpret the results.

      We thank Reviewer 2 for this thoughtful comment regarding the discrepancy between M+2 citrate produced (normalized ion counts per cell) versus fraction of the total intracellular citrate pool that is M+2 labeled in serine depleted H1299 cells. In our kinetic U-<sup>13</sup>C-glucose tracing experiments, where we performed isotope labeling for up to 15 minutes, we only see a greater presence of M+3 citrate from fully labeled glucose without robust changes in M+4, M+5, or M+6 citrate (Author response image 5). An elevated M+3 citrate could represent pyruvate carboxylase activity, where M+3 labeled pyruvate is converted to M+3 oxaloacetate that then reacts with unlabeled acetyl-CoA to generate M+3 citrate.

      We also find that the total citrate pool in H1299 cells is elevated upon serine depletion (see Supplementary Figure 6C in the original manuscript). Thus, the fractional contribution of an isotope to the citrate pool may decrease despite an increase in the amount of the particular isotope. In the original manuscript, we included data from kinetic U-<sup>13</sup>C-glutamine tracing in H1299 cells cultured with or without serine (Supplementary Figure 6I,J of the original manuscript). We find that H1299 cells depleted of serine exhibit greater M+4 citrate (via oxidative decarboxylation) and greater M+5 citrate (via reductive carboxylation) compared to serine-replete H1299 cells. Thus, one other potential explanation for why M+2 citrate from kinetic U-<sup>13</sup>C-glucose tracing represents a lower fraction of the total citrate pool in serine depleted H1299 cells is because there is a larger contribution from glutamine to the citrate pool. While there was no difference in the fraction of the citrate pool that consists of M+4 citrate, there was a greater fraction of the citrate pool labeled by M+5 citrate upon kinetic U-<sup>13</sup>C-glutamine tracing in serine depleted H1299 cells (see Author response image 6A, B). There was also a greater fraction of the citrate pool from M+6 citrate upon kinetic U-<sup>3</sup>C-glutamine tracing in serine depleted H1299 cells (Author response image 6C). This would require M+3 pyruvate labeling from glutamine, which may be due to malic enzyme, which converts M+4 malate to M+3 pyruvate. M+3 pyruvate may also be formed by PEPCK, which could convert M+4 oxaloacetate to M+3 phosphoenolpyruvate, leading to M+3 pyruvate. While understanding the source of M+6 citrate from glutamine is out of the scope of this study, it may highlight an interesting metabolic shift in H1299 cells depleted of serine that could elevate the total intracellular citrate pool.

      Author response image 5.

      Citrate isotopologues (A. M+3; B. M+4; C. M+5; D. M+6) from kinetic U-<sup>13</sup>C-glucose tracing in H1299 cells depleted of serine for 24 hours. For all measurements, citrate values were normalized to internal norvaline standard and cell number for each condition, n=3. Data shown are means ± SD.

      Author response image 6.

      Fraction of the citrate pool labeled by U-<sup>13</sup>C-glutamine in H1299 cells depleted of serine for 24 hours. (A) Fraction of the total citrate pool that is M+4 citrate (formed via oxidative decarboxylation), n=3. (B) Fraction of the total citrate pool that is M+5 citrate (formed via reductive carboxylation), n=3. (C) Fraction of the total citrate pool that is M+6 citrate, n=3. Data shown are means ± SD.

      (5) The lipid depletion part of the paper seems to be somewhat tangential. The effect of lipid depletion on the NAD+/NADH ratio in A549 cells is modest, and the effects of dual serine and lipid depletion on OCR and NAD+/NADH ratio are not consistent. Moreover, if the authors want to show that these different nutritional environments affect lipid synthesis, apart from glucose incorporation into citrate, they would need to show actual carbon incorporation into palmitate, probably at longer time points.

      We apologize for the lack of clarity for how mitochondrial respiration and the NAD+/NADH ratio play a role in governing glucose oxidation to citrate. To better highlight our logic and rationale for investigating alterations in NAD+/NADH homeostasis and citrate synthesis under lipid depletion, we have added the following text to the revised manuscript:

      “Oxidative biosynthetic reactions other than serine synthesis can also be constrained by the NAD+/NADH ratio. For example, cancer cells deprived of environmental lipids increase oxidative citrate production, and we have previously found that citrate synthesis, either through glucose oxidation or glutamine oxidation, is limited by NAD+ availability (Li, 2022) (Figure 5A, Supplementary Figure 8A). Thus, we sought to uncover whether the increase in the cell NAD+/NADH ratio by mitochondrial respiration in response to serine withdrawal specifically supports greater serine synthesis or also leads to greater oxidative citrate production.” (Lines 307-313)

      While we have previously shown that alterations to the NAD+/NADH ratio can modify both citrate production and palmitate synthesis under lipid depleted conditions (PMID: 35739397), we agree with Reviewer 2 that no conclusion can be made about lipid synthesis without direct measurements and have revised the manuscript accordingly.

      (6) In Figure 6C-6F, showing the results of the controls (+serine +lipids) will help to clarify the extent to which serine and citrate synthesis rates are affected by the different interventions.

      We thank the reviewer for the comment. Because we specifically asked how dual serine and lipid starvation impacted either serine or citrate synthesis compared to singular nutrient deprivation alone, we performed the experiments focusing on these conditions. We felt that conducting an experiment that specifically targeted our question would be make the findings more accessible as we had compared the +serine +lipid conditions to either serine or lipid depletion alone earlier in our manuscript (Figure 2D and Figure 5G,H of the revised manuscript).

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Chang and colleagues provides new insights into how cancer cells adapt their metabolism under nutrient-deprived conditions. They find cells respond differentially to serine and lipid deprivation via oxidising the cell redox state, which enables biomass synthesis and cell proliferation. They identified mitochondrial respiration as the major mechanism that dictates the endogenous NAD+/NADH ratio. By incorporating a dual stress paradigm, serine and lipid deprivation, the study further suggests that the NAD+/NADH ratio can serve as a link to orchestrate the complex interplay between multiple nutrient changes in the tumour microenvironment.

      Strengths:

      A novel aspect of this study is the idea that cancer cells are not uniformly passive victims of nutrient limitation; some can actively invoke endogenous NAD+ regeneration to combat nutrient stress. The conclusion is well-supported by comparing multiple cell lines from different tissues and genetic backgrounds, which improves generalizability. While most of the smaller conclusions align with common reasoning and expectations, the step-by-step deduction that leads to a novel 'big picture' is commendable. Another notable strength is the integration of dual stress (lipid and serine deprivation), which better mimics the complex tumor microenvironment with multiple nutrient fluctuations, raising the translational potential of these findings. The observation that lipid-deprived cells can stimulate serine synthesis and support proliferation in a subset of cancer cell lines offers a novel perspective on metabolic plasticity under starvation conditions.

      We thank Reviewer 3 for their time and for their comments to help us improve the manuscript. We also thank them for highlighting the strengths and significance of our findings.

      Weaknesses:

      (1) Although the authors derive a novel and valuable overarching concept, the presentation of this "big picture" is not clearly articulated, making it less accessible to readers outside the immediate field. It would greatly enhance the manuscript to include a clearer summary of the overarching model and its implications. Additionally, discussing the potential clinical significance and applications of the findings would increase the relevance and broader impact of the work. Finally, the manuscript's clarity and credibility are undermined by inconsistent figure labeling and the lack of statistical analysis, particularly for the Western blot data.

      (A) It would greatly enhance the manuscript to include a clearer summary of the overarching model and its implications. Additionally, discussing the potential clinical significance and applications of the findings would increase the relevance and broader impact of the work.

      We appreciate Reviewer 3’s suggestion to help clarify the findings of this study. To better articulate our overarching model, we have added the following text to the end of the Results section of the revised manuscript

      “Taken together, we propose a model where environmental nutrient availability can impact mitochondrial respiration based on the specific cancer. Because mitochondrial respiration is a major pathway that regenerates NAD<sup>+</sup>, changes to mitochondrial respiration can alter the cell NAD+/NADH ratio, influencing the activity of major NAD<sup>+</sup>-requiring metabolic reactions such as serine synthesis and citrate synthesis that can be important for proliferation. We further propose that changes to the cell NAD+/NADH ratio can impact all oxidative biosynthetic reactions if the enzyme machinery is present, but that specificity for how the cell NAD+/NADH ratio changes is dependent on both cell-intrinsic factors and cellextrinsic factors (Figure 7)." (Lines 396-404)

      Additionally, a new model figure was added as Figure 7 in the revised manuscript, which may help with understanding for a general audience.

      To better highlight the potential clinical significance of these findings, we have added the following at the end of the Discussion section of the revised manuscript:

      “Better understanding the mechanisms cells use to alter respiration and adjust the NAD+/NADH ratio in response to available nutrients could inform the complex interplay between cell-intrinsic and cell-extrinsic factors that determine cancer metabolic dependencies. This is particularly important to consider when targeting metabolism for cancer treatment. Many newer therapies targeting metabolism have not been successful in part because of metabolic plasticity to nutrient shifts (Amoedo, 2017; Fendt, 2020; Xiao, 2023). Co-targeting mitochondrial function limits metabolic adaptations and may also help predict the tissue nutrient conditions that result in pathway dependencies for specific cancers. Thus, better understanding how the cell NAD+/NADH ratio responds to nutrient levels in different cancers could improve selection of patients for cancer therapies that impact metabolism.” (Lines 483-492)

      (B) “…the manuscript's clarity and credibility are undermined by inconsistent figure labeling and the lack of statistical analysis, particularly for the Western blot data.”

      We apologize to the reviewer for any inconsistency in data presentation. To address the comment related to inconsistent figure labeling, we ensured all figures in the revised manuscript are labeled to allow readers to recognize what cell lines are used, what conditions are tested, what parameters are measured, and how the data may or may not be normalized. To address the reviewer’s comments about lack of statistical analysis, in the revised manuscript we ensured that statistical analyses are included for data presented in each figure, when appropriate. We also include a section titled “Statistics and Reproducibility” in the Methods section. In our revised manuscript, we have ensured that the p-value threshold is consistent throughout all figures, and have removed “ns” across the manuscript for consistency as suggested by Reviewer 3 in their minor comments. We also removed any explicit p-values included in figures where the p-values were close to reaching the threshold for significance (a=0.05). We have also performed additional statistical analyses where needed, including adding the pvalues for linear regression analyses, and ensured new data added to the revised manuscript also included appropriate statistical analyses.

      For western blot data, we show representative immunoblots. However, we measured PHGDH, PSAT1, and PSPH protein expression in three biological replicates across examined cancer cells and quantified the average serine synthesis protein expression from each replicate performed with error bars that denote standard deviation (see Author response image 7). We performed a nested ANOVA to examine whether there was a statistically significant difference in PHGDH, PSAT1, and PSPH protein expression between non-responder and responder cancer cells. Interestingly, as noted in our response to Reviewer 2, we find a significant negative association between PHGDH protein expression and response to serine deprivation among the six cancer cells where PHGDH protein expression did not explain proliferation upon serine depletion.

      Author response image 7.

      Serine synthesis enzyme protein expression in serine-replete and serine depleted cancer cells. (A) Immunoblots examining the expression of PHGDH, PSAT1, and PSPH in cancer cells as shown. HSP90 was used as a loading control. Data are from two separate biological replicates. (B) Mean levels of PHGDH, PSAT1, and PSPH normalized to loading control HSP90 across cancer cells from three separate biological replicates. Yellow denotes cancer cells that do not elevate mitochondrial respiration in response to serine depletion (non-responders). Blue denotes cancer cells that do elevate mitochondrial respiration in response to serine depletion (responders). P-values were calculated with nested ANOVA comparing non-responders and responders, **p<0.01

      (2) While this study identifies changes in serine synthesis, mitochondrial respiration, PHGDH protein levels, and NAD+/NADH ratio in different cell lines, some of these relationships appear correlative rather than causally established (Figure 2; Figure 5; Figure 6). Some claims are thus overinterpreted. For example, the co-occurrence of increased NAD+/NADH ratio and citrate levels under lipid deprivation in A549 cells does not establish causality (Figure 5). Direct perturbation experiments that manipulate NAD+/NADH and assess downstream effects on citrate synthesis would substantially strengthen the conclusions.

      We agree with Reviewer 3 that corresponding changes in proliferation, mitochondrial respiration, and serine synthesis are correlated to the NAD+/NADH ratio. As shown in Figure 4, we perturbed the NAD+/NADH ratio with FCCP and rotenone to measure downstream effects on serine synthesis. We also agree with the reviewer that doing similar experiments in the lipid depletion condition would highlight the relationship between the NAD+/NADH ratio and citrate synthesis. However, we point out that these experiments were already published in a manuscript from our group specifically showing that the NAD+/NADH ratio is limiting for citrate synthesis (PMID: 35739397). In that manuscript, the NAD+/NADH ratio was perturbed using electron transport chain inhibitors, including complex I inhibitors, which decreases the cell NAD+/NADH ratio. Exogenous electron acceptors were used to rescue the NAD+/NADH ratio, and under those conditions, cell proliferation, the NAD+/NADH ratio, and glucose and glutamine oxidation to citrate were measured with and without lipid depletion. We showed that decreasing the NAD+/NADH ratio decreases citrate synthesis through both glucose and glutamine oxidation and also affects palmitate synthesis. We could rescue citrate and palmitate synthesis by supplementing cells with exogenous electron acceptors. We also show that expressing cytosolic or mitochondrial NADH oxidase (LbNOX; PMID: 27124460) in mitochondrial complex I-inhibited cells rescues proliferation in lipid depleted conditions and that LbNOX expression raises oxidative citrate production at baseline. Given the extensive prior work showing the relationship between the NAD+/NADH ratio, oxidative citrate synthesis, and palmitate synthesis, efforts to repeat these same experiments for this manuscript were not warranted. We do show in the current manuscript that treating cells with AKB or FCCP, which raises the NAD+/NADH ratio, also increases glucose oxidation to citrate (Figure 5D of the original and revised manuscripts). We did this to confirm that the elevated M+2 citrate production from glucose in serine starved H1299 cells was related to an increase in the NAD+/NADH ratio as opposed to a specific response to serine depletion.

      The study focuses predominantly on mitochondrial respiration as a source of NAD+ regeneration. However, it will also be interesting to check other significant pathways, such as NAD+ salvage, which have been implicated in supporting serine biosynthesis. In addition, the subcellular distribution of NAD+ may distinguish whether some cells are truly redox-unresponsive. Mitochondrial NAD+ regeneration might counteract the cytosolic NAD+ consumption, rendering a relatively stable intracellular NAD+/NADH ratio. The malate-aspartate shuttle can be an interesting aspect.

      (A) The role of NAD+ salvage and serine biosynthesis

      Per the reviewer’s request, we investigated whether NAD+ salvage might be involved in supporting serine synthesis. Specifically, the reviewer comments highlight an interesting question about whether NAD+ salvage may differentially contribute to serine synthesis between cancer cells that elevate mitochondrial respiration in response to serine depletion and cancer cells that do not change mitochondrial respiration in response to serine depletion. Specifically, we wondered whether cancer cells that do not elevate mitochondrial respiration in response to serine depletion depend more on NAD+ salvage to support proliferation in serine depleted conditions. To test this, we treated A549 and H1299 cells in serine depleted conditions with increasing doses of the nicotinamide phosphoribosyltransferase (NAMPT) inhibitor FK866. However, we found no statistically significant difference in sensitivity to FK866 upon serine depletion in these cells based on ANCOVA analysis (p=0.9332). Interestingly, we observe that A549 cells are more sensitive to FK866 treatment than H1299 cells in serine-replete media conditions (ANCOVA analysis, p=0.0004). This suggests that A549 cells at baseline may have greater dependence on NAD+ salvage compared to H1299 cells, though this is not specific to the response to serine depletion. We then asked whether nicotinamide mononucleotide (NMN), the product of NAMPT and the immediate precursor to NAD+ in the salvage pathway, would rescue the proliferation of A549 cells cultured without serine. We find that adding 100 µM NMN, a concentration that can impact PHGDHdriven serine synthesis (PMID: 30157431), does not change proliferation of A549 cells cultured without serine, unlike supplementing cells with AKB or FCCP, which increase NADH oxidation to NAD+. Together, these data suggest that NAD+ salvage does not play a major role in differentiating the redox response to serine deprivation between responder and non-responder cells. We have added these data as Supplementary Figure 3C,D of the revised manuscript.

      (B) The role of the malate-aspartate shuttle and serine biosynthesis

      The MAS has been shown to play an important role in serine synthesis (PMID: 37647199) and may facilitate elevation in mitochondrial respiration in response to serine depletion. As stated in response to Reviewer 2, measuring subcellular compartmentspecific NAD+/NADH ratios accurately is not feasible, so we utilized a functional approach to interrogate the role of compartmentalization. Specifically, we tested a role for the malate-aspartate shuttle (MAS). Using CRISPR/Cas9, we generated GOT1, MDH1, and GOT2 deleted H1299 cells. We did not knock out MDH2 given its integral role in the TCA cycle. Using the knockout lines, we measured the whole cell NAD+/NADH ratio and found that MDH1 and GOT2 KO cells no longer exhibited an elevated cell NAD+/NADH ratio upon serine depletion compared to non-targeting controls (NTC). Consistently, MDH1 and GOT2 KO cells did not elevate OCR upon serine deprivation, nor did they exhibit greater serine synthesis rates compared to NTC cells. This suggests that MDH1 and GOT2 activity support the process by which mitochondrial NAD+ regeneration provides cytosolic NAD+ to support serine synthesis. We next asked whether MAS protein expression differed between cells that elevate respiration in response to serine depletion and cells that do not. While enzyme expression is not equivalent to activity, we wondered whether MAS protein expression would be lower in cells that do not increase their mitochondrial respiration upon serine depletion. However, we observed no major difference in GOT1, GOT2, MDH1, or MDH2 protein expression across the cancer cells examined (Author response image 8). Further experimentation is needed to measure MAS activity across lines and may reveal a mechanism by which mitochondrial respiration is governed by nutrient availability, such as levels of environmental serine.

      Author response image 8.

      Protein expression of the malate aspartate shuttle enzymes GOT1, MDH1, GOT2, and MDH2 in cancer cells cultured without serine for 24 hours. Membranes were first probed for GOT1 or GOT2 then stripped and re-probed for MDH1 or MDH2.

      (3) The authors should acknowledge the limitations of short-term isotope tracing in their experimental design. Differences in metabolic rates across cell lines can affect the kinetics of metabolite labeling, limiting the direct comparability of metabolic fluxes between them. As a result, observed changes may reflect transient adaptations rather than stable metabolic reprogramming. It is important to clarify that the study primarily captures short-term responses, and the conclusions may not extrapolate to longer-term adaptations or protein-level changes under sustained nutrient stress.

      We thank the reviewer for this comment. We apologize for any confusion around experimental approaches. We agree that in the case of acute changes in nutrient availability at the start of kinetic isotope tracing, the observed changes may reflect transient adaptations. However, cells are exposed to conditions for 24 hours prior to performing kinetic tracing. This approach allows us to examine changes that occurred in response to the nutrient condition, not acute changes. Additionally, we add fresh, prewarmed treatment media at least two hours prior to commencing kinetic isotope tracing. Upon analysis of kinetic isotope tracing, we examine whether cells were at metabolic steady state by monitoring metabolite levels over the course of tracing. For example, in the kinetic glucose tracing experiments in serine depleted cells, total serine levels are relatively stable throughout the experiment, and we find that total serine levels are greater in H1299 cells after 24 hours of serine starvation. Data showing total metabolite pools over the course of tracing are shown in the Supplementary Figures (for example, see Supplementary Figure 8C-H in the revised manuscript). The period of treatment prior to the start of kinetic isotope tracing is described in the figure legends and further detailed in the “Kinetic U-<sup>13</sup>C-Glucose Isotope Tracing Experiments” section of the Methods in the revised manuscript. To improve clarity, we added a kinetic graph showing total serine levels over time in Supplementary Figure 2I of the revised manuscript as this can address whether synthesis rates are captured while cells are at metabolic steady state. We also discuss these considerations better in the revised manuscript with the following text:

      “Importantly, we confirmed kinetic U-<sup>13</sup>C-glucose tracing was performed at metabolic steady state by ensuring metabolite levels were stable at each collected time point (Supplementary Figure 2I)” (Lines 178-180).

      Reviewer #3 (Recommendations for the authors):

      It is important to note that, in many cases, the data show only trends rather than statistically significant differences, or, if significance testing was performed, the results are not clearly labeled. For example, in Figure 1B, no p value was denoted in the figure, and the scale bar is quite high, precluding the conclusion that "AKB and rotenone dosedependently increased and decreased the cell NAD+/NADH ratio". In Figure 2E, no pvalue was shown to support the result that "H1299 cells had higher serine level than A549 cells". Inconsistencies in how significance is denoted across figures (e.g., asterisks vs. numerical values; "ns" vs. no label) make interpretation difficult. Marginal significance (e.g., p = 0.06 in Figure B) can be reported explicitly, but all figures should clearly denote whether comparisons are significant or not. Conclusions drawn from nonsignificant trends should be appropriately stated.

      We thank Reviewer 3 for this important comment and for highlighting specific instances where the manuscript could be improved. Please see response to Reviewer 3, Major Comment 1B. We also agree with Reviewer 3 that it is integral to ensure that conclusions made from non-significant trends are appropriately stated. For example, we explicitly mention that there was no statistically significant difference between the serine synthesis rate of A549 cells depleted of serine versus A549 cells depleted of both serine and lipids (Line 375). As another example, we changed the phrase “Moreover lipid depletion led to a greater fraction of total serine derived from glucose in serine depleted A549 cells” to “Moreover, lipid depletion appeared to lead to a greater fraction…” (Line 376).

      Western blot data supporting PHGDH expression variability across cell lines (e.g., Supplementary Figure 2B, 3E) appear to rely on single experiments. At least three biological replicates are required to substantiate claims about discordance between PHGDH levels and serine sensitivity. Supplementary Figure 4G presents overexpression validation based on a single Western blot without quantification. Including statistical validation from biological replicates would strengthen this point.

      We thank Reviewer 3 for this suggestion. Western blots were repeated 3 times, although data from a representative blot is shown. Please see response to Reviewer 3, Major Comment 1B.

      Certain data visualizations (e.g., Figure 2C) lack annotation indicating which data points correspond to which cell lines, limiting interpretability. All figures should include clear labels, consistent statistical notation, and complete legends. The author uses different color labels (redox-responsive (blue) and unresponsive (yellow) cell lines), which provides mechanistic clarity; however, this classification was not consistently used across the manuscript (e.g., Figures 2d and 2e). To further improve reader comprehension, consider adding conceptual schematic diagrams before each main result section to illustrate experimental logic, and a final diagram summarizing the proposed mechanism.

      We apologize for any unclear data presentation. In the revised manuscript we have added greater clarity around what cell lines are used in each experiment and have added explicit labeling to specify cancer cell lines in Figure 2C of the revised manuscript. Throughout, we have ensured that any serine redox non-responder cell lines are labeled in yellow while serine redox non-responder cell lines are labeled in blue. We have also ensured that any lipid redox responder cells are labeled in green while lipid redox non-responder cells are labeled in dark purple, a change from the original manuscript. Finally, we have also added a schematic to summarize the proposed model in Figure 7 of the revised manuscript.

      Although the authors provide justification for using H1299 and A549 as representative cell lines to study serine depletion, it remains unclear whether these two lines are equally suitable for investigating lipid depletion. Additional rationale or supporting data would help clarify their appropriateness for the lipid-related experiments.

      We thank Reviewer 3 for this suggestion. We opted to study H1299 and A549 cells under lipid deprivation to assess their responses in relation to the response to serine deprivation. We specifically wanted to know whether these findings related to serine deprivation applied to other nutrient depleted conditions. We clarify this logic in the revised manuscript by adding the following text:

      “Oxidative biosynthetic reactions other than serine synthesis can also be constrained by the NAD+/NADH ratio. For example, cancer cells deprived of environmental lipids increase oxidative citrate production, and we have previously found that citrate synthesis, either through glucose oxidation or glutamine oxidation, is limited by NAD+ availability (Li, 2022) (Figure 5A, Supplementary Figure 8A). Thus, we sought to uncover whether the increase in the cell NAD+/NADH ratio by mitochondrial respiration in response to serine withdrawal specifically supports greater serine synthesis or also leads to greater oxidative citrate production.” (Lines 307-313)

      We have also included more detailed justification for focusing our studies on A549 and H1299 to study serine depletion by adding the following statements to the manuscript:

      “We performed focused comparisons between A549 and H1299 cells because they exhibit differences in proliferation upon serine deprivation that are not explained by PHGDH protein expression, demonstrate differing responses of the cell NAD+/NADH ratio upon serine deprivation, and have similar basal proliferation rates.” (Lines 171-175)

      The concentration of serine in replete media should be explicitly stated and justified. If the intention is to mimic physiological conditions, alignment with human plasma levels would increase translational relevance.

      We agree that explicitly stating the concentration of serine in replete media is important. In the revised manuscript, we explicitly state that DMEM contains 400 uM of serine and that we use this concentration for serine-replete conditions (Line 102). While an important application of our manuscript is to better explain metabolic changes that can occur in physiologic conditions, we acknowledge that we did not test levels found in different tissues. Rather, by examining extreme conditions of high and low serine, we hoped to dissect how cells adapt to nutrient conditions, and testing the more subtle responses based on tissue serine levels will require a dedicated study.

      Rotenone may elevate ROS levels and trigger cellular stress responses, potentially confounding proliferation assays. The authors should validate that concentrations used do not induce cytotoxicity or excessive oxidative stress, and ideally measure ROS levels to support interpretation.

      We thank Reviewer 3 for raising this important point. We explicitly measured cell viability with the doses of rotenone used in this manuscript in cells cultured with or without serine. We find that rotenone dose-dependently increases cytotoxicity in A549 cells grown in serine-replete conditions in a statistically significant manner as calculated by simple linear regression. However, the cytotoxicity from rotenone is low (at most 4% in serine depleted conditions) and does not explain differences to rotenone sensitivity with respect to serine synthesis. These data have been added to Supplementary Figure 1C of the revised manuscript.

      Evidence for lipid depletion can enhance serine synthesis in A549 cells is inadequate, for the marginal difference in NAD+/NADH ratio and slight increase of M+3 serine levels. The statement "any perturbation that increases the NAD+/NADH ratio led to both elevated serine and citrate production, regardless of what nutrient was depleted from the environment" (introduction section) should be reworded.

      We thank Reviewer 3 for this suggestion. We have changed the above statement to the following:

      “Lastly, we find that any perturbation that increases the NAD+/NADH ratio, including lipid deprivation, could paradoxically improve the proliferation of cells in serine depleted conditions.” (Lines 90-92).

    1. eLife Assessment

      This study addresses an important gap in drug discovery by delivering a rigorous, large-scale evaluation of widely used co-folding methods for predicting ligand-bound protein complexes and virtual screening. A key strength is the comprehensive benchmarking framework, which leverages structures and chemical compounds that were absent from the AI models training set, thereby providing particularly compelling and unbiased evidence of co-folding performance. The findings clearly delineate the complementary roles of deep learning-based co-folding and physics-based docking, offering practical guidance for their rational integration into drug discovery workflows. Overall, the conclusions are well supported by thorough analyses across a representative set of cases and are highly convincing.

    2. Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      Strengths:

      Overall, this is a scientifically solid paper.

      The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Comments on revised version:

      The authors have adequately addressed my concerns.

    3. Reviewer #3 (Public review):

      Summary:

      Core conclusions are well-supported by data: co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods, provides an unbiased and rigorous benchmark dataset, which contains structures and compounds absent from the co-folding models training sets. Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment. The revised results clarify an intriguing finding: co-folding can predict correct ligand poses even when protein formations are mispredicted. The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      Weaknesses:

      The study identifies a major limitation of co-folding-failure to capture rare protein conformational changes, which deserve future investigation. The authors include uncalibrated Boltz-2 affinity data (addressing a prior comment) but note that large-scale free energy perturbation (FEP) comparisons are beyond their capabilities.

      Appraisal of Aims Achieved:

      The authors successfully achieved their primary aims and the results provide strong, well-supported evidence for their core conclusions. Key conclusions are grounded in the study's unbiased, training-set independent data, ensures the conclusions are not confounded by model memorization and are broadly applicable to the field's use of these co-folding models.

      Field Impact:

      This study provides a critical reality check for the field: co-folding models are powerful tools for pose prediction but are not yet standalone solutions for virtual screening, a key distinction that will prevent over-reliance on these models and guide more rational tool selection.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      We thank Reviewer 1 for this thoughtful summary of our work.

      Strengths:

      Overall, this is a scientifically solid paper. The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Weaknesses:

      My main concern is that the study's aim is a bit unclear. Modern benchmarking studies comparing physics-based docking with deep learning-based co-folding approaches (e.g., AF3, Boltz-2, Chai-1, and others) are increasingly expected to go beyond aggregate performance metrics.

      Indeed, we have gone into several examples of failures and successes for each of these methods. As we are not developing these methods ourselves, we also think this dataset will be a valuable contribution for improving them further.

      In addition to rigorous dataset construction, transparent methodology, and appropriate statistical evaluation, high-impact benchmarks typically provide actionable guidance on when each method class is most appropriate, reflecting their distinct inductive biases and practical constraints. Failure-mode analyses that link performance differences to protein flexibility, ligand chemistry, or binding-site characteristics are particularly valuable, as they move comparisons beyond "scoreboard" assessments toward mechanistic understanding.

      Right now, we do not observe meaningful trends that separate the failure modes for any individual method. This is covered in Supplementary Figures 6 and 7.

      While full biological validation is not expected, qualitative interpretation grounded in physical and biological principles strengthens conclusions. Providing reproducible workflows or reference pipelines is not mandatory, but it is increasingly viewed as a best practice because it facilitates adoption and helps contextualize results for practitioners.

      We note that our code is available (https://github.com/jongbin99/Cofolding/) and all structural data will be publicly accessible in the PDB alongside publication (we only held it back only for “blinding” during peer review to avoid contamination with any new deep learning methods).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Kim et al. evaluates the performance of three modern AI-based methods in predicting complex structures and binding affinities between proteins and chemical compounds. An honest 'prospective' evaluation is achieved by studying benchmark structures and chemical compounds that did not exist in the PDB at the time the AI structure prediction models (AlphaFold3, Chai-1, Boltz-2) were trained.

      Strengths:

      (1) The study addresses an important question in modern computational biology and drug discovery, and establishes the strengths and limitations of the three tools in solving various computational chemistry tasks, including compound pose prediction, active-inactive discrimination, and potency ranking.

      (2) The conclusions are based on examination of four separate targets and respective compound datasets, where for one of the targets, the authors also obtained numerous X-ray structures to serve as experimental answers for the binding pose prediction task.

      (3) The study reports relationships between structure prediction confidence, predicted energies (DOCK3.7), and affinity predictions (Boltz-2) with the geometric accuracy of compound pose prediction as well as the experimentally measured potency.

      (4) One of the key findings is the limited ability of co-folding methods to predict conformational rearrangements, which does not correlate with their ability to predict binding poses of the compounds inducing these rearrangements.

      (5) The findings could serve as useful guidelines for computational chemists in selecting appropriate software and scoring schemes for each task.

      We appreciate Reviewer 2’s summary of the novelty of the dataset and analysis.

      Weaknesses:

      While I consider this a solid study, several aspects would need to be addressed to make it really strong:

      (1) DOCK3.7 docking and scoring experiments were performed using one experimental structure of Mac1, selected from dozens of structures based on a criterion that is not sufficiently well justified. For sigma2 receptor, dopamine D4 receptor, and AmpC β-lactamase, it is not clear which structures or models were selected for docking at all. It is well known that geometry predictions, scoring, and active-inactive ROC AUCs are all strongly influenced by the selected structure. It would be important to attempt Mac1 docking using all available experimental Mac1 structures, or at least against representative structures in various conformations; it would also be quite insightful to compare results to docking of the same compound sets to AF3, Boltz-2 and Chai-1 predicted structures of Mac1. Same goes for the docking studies of sigma2, D4, and AmpC β-lactamase.

      In any program, a decision has to be made as to which template will be used for docking, we justified the choice in the methods:

      “We used this structure because the inhibitor (Z5014193706) was the most potent molecule with a structure determined around the same time as the ligands in this dataset were tested.”

      We stand by this as a reasonable assumption. Similarly, for sigma2, D4, and AmpC β-lactamase, the template was chosen in the respective papers:

      a) The σ2 receptor bound to cholesterol (PDB ID: 7MFI) was used in the docking calculations.

      - This structure was determined in the paper, the first structure of sigma2 and therefore a worthy template

      b) The D4 receptor campaign used PDB 5WIU

      - This was one of two D4 structures available and chosen because it was not bound to sodium

      c) For AmpC, the campaign used the structure in the Protein Data Bank (PDB) 1L2S

      - This maximizes comparisons to other docking studies that used the same receptor template.

      The major goal of this study is to compare different methods under reasonable (but perhaps as the reviewer points out, not optimal) conditions, not to optimize docking score.

      (2) For binding affinity predictions, as a control, authors should consider compound co-folding with an unrelated protein, or even with a pseudo-peptide that consists of a few random single amino acids - this would provide an honest baseline for such predictions.

      This suggestion would be valuable for understanding the performance for these methods from the perspective of ligand specificity (a valuable, but separate, goal). Surely this will generate some number or some prediction - but what would this baseline mean and how would it be relevant for drug discovery? Therefore, we do not think this suggestion is relevant for the issues being investigated in this manuscript.

      (3) ROC curves Figure 3 and elsewhere should be shown, and AUCs quantified/reported on a log or square-root scaled x-axis, to emphasize early enrichment, which is the area of practical significance for these predictions. For example, Figure 3A currently suggests that the pose prediction performance of AF3 exceeds that of Boltz-2 whereas the early enrichment is clearly better for Boltz-2.

      We agree with this, and added a semi-logAUC plot for Figure 3A. For Figure 5, we also generated a semi-logAUC plot to see early ligand enrichment clearly, added as Supplementary Figure 11. We added the text:

      “Considering its early enrichment performance, Boltz-2 Ligand ipTM was the strongest predictor of pose accuracy based on normalized logAUC (20.5% above random, Fig. 3a). In contrast, although Boltz-2 pIC50 showed poor overall discrimination, it overestimated its ability to enrich true positive poses at low false positive rates, despite having a weak early enrichment behavior”

      (4) 'Trained set' in figures and text should probably be 'training set'? Or otherwise explain this new term the first time it is introduced.

      Thank you for pointing out this for clarification. ‘Training set’ is the correct word, and we made changes appropriately across all figures and texts.

      (5) Figure 1 illustrates a projection onto the first two principal components of a space that apparently had only one (scalar) metric for each compound pair (% maximum common substructure or Tanimoto coefficient); the authors need to better explain the principle behind this analysis and visualization.

      This suggestion is valuable, since we often use PCA to reduce dimensionality for more complex features. For clarification, we actually have a full pairwise similarity matrix for all tested Mac1 compounds based on each of Tc and MCS%. PCA for each MCS% and Tc is a representation of each pairwise similarity matrix. We also made a change in Figure 1 caption to make this point clearer:

      “projection of compounds represented by their full pairwise similarity vectors (by ECFP-4 Tc and MCS%)”

      Reviewer #3 (Public review):

      Summary:

      This study's core conclusions are well-supported by data. It is shown that co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false-positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      (1) Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods.

      (2) Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment.

      (3) The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      We thank Reviewer 3 for pointing out the unprecedented and comprehensive nature of our study

      Weaknesses:

      (1) Limited generalization to diverse protein families (e.g., no ion channels/transporters).

      We agree - we have not explored the entire proteome and these are important target classes that will surely be investigated by future studies. We focused on targets here where we had large number of X-ray crystal structures (Mac1) and affinity/inhibition measurements from docking (the other three targets).

      (2) Ambiguity in the mechanism underlying co-folding's failure to predict rare conformational changes.

      Again, we agree. We are not the developers of these methods. We observe that these methods do not predict conformational changes with high fidelity and this weakness is an area that co-folding methods will surely prioritize in the future.

      (3) Virtual screen comparison is unbalanced (docking-prioritized hit lists bias results).

      We acknowledge this in the results: “An important caveat is that the hit-lists were composed of molecules prioritized by docking in the first place, giving it an advantage on these particular sets.” and discussion: “Finally, comparing co-folding to docking based on hit-lists themselves selected by docking is arguably unfair to co-folding. Counter-balancing this is the inclusion, in each of the three hit lists, of molecules that had mediocre and poor docking scores intentionally selected to test the correlation between docking score and hit-rate. Here too, the correlation between co-folding score and likelihood to bind, what we sometimes call a “dock-response-curve” was no better than docking’s, often worse (SFig.11).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are suggestions for revisions:

      (1) The writing is at times obtuse and hard to follow.

      This happens sometimes when multiple authors are writing together. We apologize and are happy to respond to specific areas that can be streamlined to be easier to follow.

      (2) In the Results section, "A set of 557 previously unreported Mac1 ligand complexes", the authors have compared the ligand poses across different metrics such as Tc - a standard, highly effective method in chemo-informatics and MCS (maximum common substructures); these are standard metrics for quantifying the structural similarity between pairs of small molecules. This part of the analysis checks whether this is memorization; it is critical to compare the two metrics, but it is not sufficient to draw a conclusion.

      Thank you for pointing out about the structural similarity of molecules co-folded to those present in the training set (resolved as Mac1 complexes and deposited in PDB before training dates). We have conducted an analysis where we do a pairwise similarity comparison for all ligands present in the PDB (regardless of the target), by both Tc and MCS, and overlay the cluster of ligands we tested (Mac1, AmpC, sigma2, D4). This should show where our tested benchmark datasets lie in the chemical space covered in the entire PDB. Each cluster (around 500 to 1300 compounds per target system) is overlaid on the cluster of all ligands deposited in PDB (over 50,000 compounds), and each cluster was relatively diverse by both Tc and MCS.

      (3) In the "Co folding can accurately reproduce poses of ligands dissimilar to those trained." Subsection under Results, the authors' conclusions are hard to follow; they state that the co-folding models often mispredict or miss the alternative conformation, but they also predict poses that are distinct from the training set. What does that imply?

      Our interpretation is actually a somewhat unsettling one: co-folding gets the ligand pose right even when it gets the protein wrong, and even when the ligand is novel. This suggests the models may be anchoring on conserved pharmacophoric interactions (like the adenosine-mimicking purine scaffold) rather than truly modeling the physics of the full complex. We added to the results section:

      This result suggests that co-folding reliably recapitulates dominant ligand-binding interactions even in the absence of accurate protein conformational modeling, providing further support to the idea that they are learning specific interaction patterns rather than a deeper physics-based representation (Masters et al. 2025).

      (4) The Discussion section connects the results and conclusions, but it can be challenging to grasp the study's overall message.

      We think the final paragraph hits on three major points:

      - Co-folding accurately predicts ligand poses for known binders, but fails to capture conformational changes

      - Co-folding does not reliably distinguish true binders from false positives in virtual screening hit lists

      - Docking and co-folding are complementary rather than competing tools

      (5) The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment. The value of the paper would be further enhanced by explaining how it differs from seemingly similar results reported in other studies, including the one cited in this manuscript (see https://www.biorxiv.org/content/10.64898/2025.12.04.692352v1).

      The Mac1 results are completely unique. However, the docking datasets are exactly the same as those analyzed in the Menon et al manuscript. We don’t think our results differs from conclusions of the Menon et al manuscript as we wrote: These observations are supported by a fascinating study on some of the same ligand sets as investigated here, using AlphaFold3, reaching similar conclusions (Menon et al. 2025).

      Reviewer #3 (Recommendations for the authors):

      (1) Expand target diversity to include ion channels, transporters, etc., beyond enzymes and GPCRs.

      (2) Investigate the cause of co-folding's failure in predicting rare conformational changes (e.g., adjust sampling, MSA inputs, or add experimental constraints).

      (3) Mitigate docking bias in virtual screens (e.g., re-analyze unbiased compound libraries).

      We addressed these three points in the public review above

      (4) Test Boltz-2's affinity predictions without linear calibration and compare with FEP.

      The data without linear calibration are included in the manuscript. Comparing such a large number of compounds with FEP is currently beyond our capabilities.

      (5) Conduct proof-of-concept to test co-folding-docking integration for better hit rates.

      We think this is well beyond the scope of this manuscript - but look forward to testing this idea in the future.

      We also got one community review that we respond to below:

      Summary

      This manuscript evaluates the performance of co-folding models when tasked with 1) the recapitulation of a large number of experimentally determined co-crystal structures of Mac1 with a series of Mac1 ligands and 2) the rescoring of hits to identify false positives originally derived from a set of large docking-based virtual screens. The evaluation leverages a dataset of crystal structures and affinity data from high-throughput crystallographic and biophysical screens, respectively. These data uniquely enable this report to focus on the ability of co-folding models to handle ligands, resulting in an analysis that is particularly timely given the wide adoption of co-folding models and the relative scarcity of such ligand-focused benchmarks among existing evaluations, which have primarily focused on protein structure prediction or binder design.

      Thank you for this thoughtful summary of our work

      Feedback

      The experiments and analyses in the manuscript are well thought-out and do not have any significant issues. There are a few high-level points that may improve the clarity and completeness of the results. Importantly, none of the suggested additional experiments will affect the conclusions of the paper, but rather help provide additional context for the results:

      The first section presents an exciting opportunity to frame the Mac1 ligands against ligands in the PDB more broadly. It would be informative to assess whether chemotypes that are easier or harder to predict accurately and confidently are over- or under-represented in the PDB as a whole. Note that this is not a recommendation that new scaffold similarity metrics be incorporated into the analysis, but rather that analyses similar to those already performed in the manuscript are performed using all ligands in the PDB. For example, PCA-based analyses similar to those in Fig. 1c could be used to examine Mac1 ligands in the context of all PDB ligands enabling questions such as whether similarity to a nearest PDB neighbor, cluster size in a Tc/MCS PCA space, or other frequency-based measures show any relationship with prediction vs. crystal structure RMSD. Such analyses could provide additional insight into how effectively models leverage ligand information present in the PDB overall, as opposed to biases arising specifically from scaffolds represented in Mac1 structures in the PDB, which are already well covered in the manuscript. The conclusion that Tc/MCS do not correlate with the ligand RMSDs for the ligands already associated with the Mac1 is well supported, and presumably suggests that a correlation would not exist against the backdrop of the PDB, but it would be interesting to see the data using analyses similar to those already done in the manuscript nonetheless.

      We are adding new figures in SFig.1 that consider how different clusters of ligands tested for our co-folding analysis are distributed across the chemical space in PDB. This is done by making a similarity comparison between every ligand in PDB and those tested in our analysis by Tc and MCS%, then plotting in PCA space for each metric. We are excited to see that each dataset covers a wide scope in PCA space, but at the same time, there are unexplored areas in the chemical space of PDB by co-folding.

      Similarly, even though the four proteins used in this manuscript are not themselves the primary focus of the analysis, it would be valuable to perform a high-level assessment of the precedent for each protein in the PDB (beyond the count of liganded structures in Table S6), either in protein sequence space (e.g., MSAs) or structural space (e.g., FoldSeek). An analysis like this would provide important context about whether any of the proteins in the study have close homologs with liganded structures in the PDB, or are generally overrepresented in the PDB. The fact that the AUC for L-pLDDT for AmpC is higher than σ2 and D4, for example, is notable given the relative abundance of liganded AmpC structures in the PDB (this raises potentially interesting questions related to where DOCK3.7 and AF3 actually place the ligands, given the orthosteric β-lactam binding pocket in AmpC, although this is outside of the scope of this manuscript).

      High-level assessment of the precedent for each protein in the PDB will definitely help to understand if proteins we used have close homologs with liganded structures in the PDB. Our Supplementary Table 6 covers the extent to which these liganded structures were available by cutoff dates for AF3, Chai-1 and Boltz-2. AmpC had more homologs than sigma2 and D4, and this may explain a better AUC for AF3 L-pLDDT specifically for this target.

      A discussion of the affinity probability results (`affinity_probability_binary`) from Boltz-2 is likely warranted in the second section in addition to the pIC50s that are already reported (`affinity_pred_value`). The former seems like it would be more applicable for section 2 of the manuscript, but both warrant inclusion—they should both be calculated by default when the affinity pipeline in Boltz-2 is turned on, so it wouldn't involve any more inference.

      As boltz-2 affinity module outputs both affinity probability binary output and affinity predicted value, we kept track of both metrics. So we tried re-ranking hit lists using both metrics. Where boltz-2 performed better (Sigma2, D4), binary probability values were more representative as a metric to differentiate true actives from non-binders. This was more clear in semi-logarithmic ROC plots. However, in AmpC, both Boltz-2 scoring metrics performed similarly. Such inconsistency in trend made it difficult to draw conclusions.

      Minor points

      A more detailed description of the experimental methods used to generate the ground-truth data in the introduction (even though these have been explained in prior works) would help orient the reader early on, and ground the benchmarking aspect of the story. In general, the abstract and introduction would benefit from a more cohesive through-line to tie the two complementary but orthogonal sections of the paper together.

      We will include a more thorough description alongside the PDB depositions. As for the two sections, we have tried to tie them together from the perspective of drug discovery workflows…

      The cutoffs in the "Co-folding can accurately reproduce..." section shift between 2.5 Å (from the ligand center of mass) and 2.0 Å. Is there a reason for this? Along similar lines, mentioning cutoffs for true positives/negatives when introducing the ROC analyses later on in the Mac1 section seems unnecessary since no cutoff should be necessary here.

      We used 2.5A distance to COM to just get at “broadly the correct binding site” for fast filtering and 2.0A RMSD because that is the broadly accepted standard in the field for “relatively correct binding pose”.

    1. eLife Assessment

      The nematode C. elegans is an ideal model in which to achieve the ambitious goal of a genome-wide atlas of protein expression and localization. In this paper, the authors develop a rational and useful strategy for at-scale tagging of all protein coding genes with fluorescent markers, providing solid evidence that it could be a feasible foundation for a large-scale, community-wide project.

    2. Reviewer #1 (Public review):

      Summary:

      Eroglu and Hobert demonstrate that injecting CRISPR guides and repair constructs to target three genes at a time, tagging each with a different fluorescent protein, and selecting which gene to tag with which fluorophore based on genes' expression levels, can improve efficiency of gene tagging.

      Strengths:

      This manuscript demonstrates that three genes can be targeted efficiently with three different fluorophores. It also presents some practical considerations, like using the fluorophore least complicated by agar/worm autofluorescence for genes with low expression levels, and cost calculations if the same methods were used on all genes.

      Weaknesses:

      Eroglu has demonstrated in a previous publication that single-stranded DNA injection can increase efficiency of CRISPR in C. elegans, while inserting two fluorescent proteins and a co-CRISPR marker into three loci, and Paix et al 2015 demonstrated simultaneous insertion of two fluorescent tags. The current work is valuable and incremental advance. In general, I applaud the authors' willingness to strategize about how whole proteome tagging might be accomplished. I predict that the advance here will be one of many small advances that will get the field to that goal. The title oversells the advance presented, in my view, since seems like one among many key advances, and the first sentence of the Discussion seems a more apt summary of the key advance here.

      Some injections targeted genes on the same chromosome together, which will create unnecessary issues when doing crossing that will be useful for some future experiments. This made me wonder if injecting 3 together really is helpful vs targeting each gene separately, since only 5 worms need to be injected. It cuts time down by 2/3, but perhaps avoiding targeting the same chromosome with two tags would be useful.

      The limited utility of current blue fluorescent proteins makes me wonder if it's worth using at this stage, before there are better blue fluorescent proteins, or better yet, far red, to avoid issues with live imaging under phototoxic UV or near-UV illumination.

    3. Reviewer #2 (Public review):

      Original Review:

      The manuscript by Eroglu and Hobert presents a set of strains each harboring up to three fluorescently tagged endogenous proteins. While there is technically nothing wrong with the method and the images are beautiful, we struggled to appreciate the advance of this work - who is this paper for?

      As a technical method, the advance is minimal since the first author had already demonstrated that three mutations (fluorophore insertion and co-CRISPR marker) could be introduced simultaneously.

      As a pilot for creating genome-scale resources, it is not clear whether three different fluorophores in one animal, while elegantly designed and implemented, will be desired by the broader community.

      Finally, the interpretation of the patterns observed in the created lines leaves much to be desired. A Table with all the observations must be included and can replace the tedious (and often wrong) descriptions of the observations with the different lines. It would be too much to point out every mistaken expectation of protein expression. Two examples include:

      The expectation that ACDH-10 is enriched in the intestine and epidermal tissues (hypodermis) is naïve - there are multiple paralogs of this protein (look at WormPaths or WormFlux) that may share functions in different tissues. There is also no reason to assume that fatty acid metabolism does not occur in other tissues (including the germline). Finally, there are no published studies about this enzyme, so we really don't know for sure what it's doing.

      The expectation that HXK-1 is ubiquitously expressed is similarly naïve. There are three paralogous enzymes that are all associated with the same reaction, and we have shown that these three function redundantly in vivo, perhaps in different tissues (PMID: 40011787). Moreover, single cell RNA-seq data (PMID: 38816550) also shows enrichment of hxk-1 in gonadal sheath cells.

      The table should have at least the following information: gene/protein name - Wormbase ID - TPM levels of single cell data assigned to tissues for L2, L4 and adult (all published) - tissues in which expression is observed in the lines presented by the authors.

      Other points:

      (1) We would encourage the authors to provide systematic validation of the reported insertions. The manuscript reports that 24 of 30 tags were isolated and visible but does not clearly state whether each isolated line was confirmed by sequence‑level validation to be correctly in‑frame and free of unintended mutations at the target locus.

      (2) The manuscript presents aggregated success counts (e.g., 8/10 mTagBFP2 tags, 9/10 mStayGold, 7/10 mScarlet3) and useful narrative descriptions of injection outcomes. We suggest also to include per‑locus success rates.

      (3) For pools that required re‑injection after initial failures, we would like to see a description of the specific changes that were made to the injection mixes or procedures (e.g., new repair template prep, different Cas9 reagent lot, guide redesign). This will be useful troubleshooting information for others.

      (4) The authors states that the fluorophore sequences are codon-optimized for C. elegans. We suggest they provide the exact donor/tag sequences used specifically state whether the fluorophore sequences contain any synthetic/artificial introns or other sequence modifications (e.g., silent PAM‑disrupting mutations) were included in the donor templates.

      (5) Page 3: Include a reference for "The C. elegans genome encodes around 20,000 genes"

      We hope these comments are useful.

      Comments on Revised Version:

      Overall, we found the responses to be quite recalcitrant.

      We have one remaining composite concern about the comparison between observed expression patterns with the new strains versus published data.

      First, the authors only report patterns for one stage while it should be not too much effort to image the different life stages. However, since this is a revision, we are not formally requesting they do this.

      Second, in the now provided Table (thank you) 'observed expression' (last column) is lacking for 9 of the 30 proteins, and for 6 of these the procedure was not successful. Why not report patterns for the other three? It is confusing also because on page 5, the authors say that "overall, 24 of 30 tags ...all of which were visible with fluorescence stereomicroscopy" - are we missing something? Also, they then said that they "obtained 6/9 of the originally failed tags"; why are the corresponding patterns not included in table 1, and are 9 proteins still labeled as "no" in the "success?" Column?

      Third, we strongly feel that the response to our comments about expression patterns is not adequate. On page 5 the authors say that "all proteins were expected to be ubiquitously expressed" and that "scRNA-seq indicated that transcript abundance was ubiquitous and without strong tissue-specific enrichment with few exceptions". However, in their rebuttal, the authors now argue for tissue-specific expression for proteins with paralogs, turning around their own argument! Moreover, their Table indicates that many genes show tissue-enriched expression by RNA-seq while many of their tagged proteins exhibit ubiquitous expression.

      Overall, this indicates that both the overall accomplishment of generating tagged protein strains and analyzing their expression is oversold.

    4. Reviewer #3 (Public review):

      Summary:

      The authors argue that establishing the expression pattern and sub-cellular localisation of an animal's proteome will highlight hypotheses for further study. This claim is probably accepted by many in the community. This manuscript seeks to confirm the feasibility of establishing such a resource, by using current transgenic methods to knock in DNA encoding different colored fluorescent tags into C. elegans genes.

      Strengths:

      The authors make the points above. For example, they provide evidence that the C. elegans germline harbors two populations of mitochondria that differ qualitatively in the proteins they express. They also confirm that labelling the whole proteome is an achievable goal with relatively limited resources and time.

      Weaknesses:

      The work is somewhat incremental in that it uses existing transgenic technology. Cell biology in C. elegans is challenging because of the small size of many of its cells, notably neurons. This can make establishing the sub-cellular localisation of a fluorescently tagged protein, or co-localizing it with another protein, tricky. The authors point out in their introduction that advances in light microscopy such as diSPIM, STED and ISM (a close relative of SIM), have increased the resolution of light microscopy. They also point out that recent advances in expansion microscopy can similarly help overcome the resolution limit. However, they do not use these technologies to characterize their transgenic strains.

    5. Reviewer #4 (Public review):

      Summary:

      Tagging the entire proteome of a metazoan would be a landmark achievement, providing a powerful complement and extension to existing "omic" catalogs in model systems. Here, Eroglu and Hobert argue that efficiently tagging multiple loci in a single "batch" would make the community-based achievement of this goal realistic. They provide rigorous evidence that such an approach is indeed feasible, exploring issues related to efficiency, design and screening strategies, disruption of gene function, and the potential for endogenously tagged alleles to reveal unexpected aspects of protein expression and localization. While the work has some minor gaps that are important to rigorously assess the feasibility of the proposed effort, the detailed and valuable insights that emerge should provide impetus to the community to coordinate efforts to make this ambitious goal a reality.

      Strengths:

      The work has numerous strengths. The authors provide compelling evidence that:

      - three distinct loci can be efficiently targeted with three distinct fluorescent tags in a single injection.

      - thoughtful targeting design can reduce the likelihood of disruption of function by the tag.

      - systematic design principles based on expression level and predicted localization/function can be used to optimize tagging strategies.

      - the resulting tags can provide unexpected insight into patterns of protein production and subcellular localization.

      Not all of these advances are novel in themselves, but taken together, they represent an important technical and conceptual advance. The most important strength comes from the exceptionally high value of the goal itself, in that the work is that it has the potential to spur a community-wide effort toward achieving the ambitious goal of proteome-wide tagging.

      Weaknesses:

      The work's shortcomings are minor.

      - One concern has to do with the feasibility of the proposed screening strategies. The experimental design cleverly coinjects tags for three loci in different gene expression 'zones'; this expression level determines which tag will be used. As the authors allude to, there is an important distinction between genes with the same overall FKPM value between those that are expressed broadly and those focally expressed in a specific tissue. The proposed strategy claims that there are a sufficient number of highly expressed genes "to be used as visible markers" for recovering successfully edited animals. It would be useful for the authors to discuss the issue of broad vs focused expression among this set of genes a bit more thoroughly, with an eye toward the issue of how likely it is that these genes could indeed consistently be used as visible markers, particularly for those at the low end of this limit.

      - What fraction of the proteome (on a per-gene basis) is secreted proteins? How difficult will it be to screen these for successful tags? Are there specific tags that would be more optimal for secreted proteins? (The authors mention the use of an SL2 or T2A cassette to label the cells in which these proteins are expressed but note that there are technical challenges associated with doing this at scale.)

      - For secreted and/or weakly expressed genes, it would be useful for the authors to estimate for what fraction of these would successful insertions need to be screened by PCR, and what resources (time and money) this would likely entail.

      - For how many genes would a single tag not capture all predicted isoforms?

      - Finally, some readers might object to the authors' assertion in the abstract that this work is "a first step in this direction" (presumably referring to designing a strategy for whole-proteome tagging). There is no concern that the authors are disregarding the extensive work of other groups, as they explicitly mention the contributions of other groups to the foundation that enables the present work. However, the spirit of the abstract could be misinterpreted by a well-intentioned reader.

    6. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      The nematode C. elegans is an ideal model in which to achieve the ambitious goal of a genome-wide atlas of protein expression and localization. In this paper, the authors explore the utility of a new and efficient method for labeling proteins with fluorescent tags, evaluating its potential to be the basis for a larger, genome-wide effort that is likely to be very useful for the community. While the evidence for the method itself is solid, carrying out this project at a large scale will require significant additional feasibility studies.

      We appreciate the editor’s recognition that the evidence for our method is solid and that a genome-wide protein atlas in C. elegans would be highly valuable to the community. However, we respectfully disagree that “significant additional feasibility studies” are required. Take the yeast proteome-wide GFP tagging project (Huh et al., Nature 2003). It achieved ~75% coverage of ~6,000 proteins directly from an established protocol without any prior significant feasibility studies, at least to our knowledge. While the C. elegans genome is 3 times in size, we would argue that our tagging protocol may even be less labor intensive as it does not involve any cloning and the screening is visual, requiring no molecular biology skills. Reviewer 3 notes: ‘They also provide convincing evidence that labelling the whole proteome is an achievable goal with relatively limited resources and time.’

      Our pilot study validates all key parameters for genome-wide scaling: editing efficiency at novel loci with untested reagents, viability of tagged worms, and detectability of multiple spectrally separated fluorophores across expression ranges. These address the core technical, biological, and practical challenges of large-scale endogenous tagging in a multicellular organism, leaving no fundamental barriers in our view.

      The proposed cost and timeline align quite favorably with established large-scale consortium projects: e.g., ENCODE pilot analyzed 1% of the human genome at ~$55 million over 4 years; Mouse Knockout Consortium scaled to ~20,000 genes over 20 years (ongoing) with ~$100 million; Human Protein Atlas mapped ~87% of proteins with antibodies in fixed cells (through much more labor intensive methods) over 20+ years at >$100 million. With ~8% of C. elegans genes already tagged (WormTagDB) and labs already tagging entire gene classes (PMID: 40463100), scaling our protocol to the proteome is feasible, potentially covering the genome in 5-6 years by a single lab or faster with distributed effort at a reagent cost of merely $2.2 million. The main barriers now are funding commitment and assembling collaborators, not further feasibility testing.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Eroglu and Hobert demonstrate that injecting CRISPR guides and repair constructs to target three genes at a time, tagging each with a different fluorescent protein, and selecting which gene to tag with which fluorophore based on genes' expression levels, can improve the efficiency of gene tagging.

      Strengths:

      This manuscript demonstrates that three genes can be targeted efficiently with three different fluorophores. It also presents some practical considerations, like using the fluorophore least complicated by agar/worm autofluorescence for genes with low expression levels, and cost calculations if the same methods were used on all genes.

      Weaknesses:

      Eroglu has demonstrated in a previous publication that single-stranded DNA injection can increase the efficiency of CRISPR in C. elegans while inserting two fluorescent proteins and a co-CRISPR marker into three loci. The current work is, therefore, an incremental advance. In general, I applaud the authors' willingness to think ahead to how whole proteome tagging might be accomplished, but I predict that the advance here will be one of many small advances that will get the field to that goal.

      Our manuscript indeed builds on prior multiplex editing (including our own co-CRISPR work), but the manuscript's primary contribution is not a novel technical breakthrough per se. Instead, our main goal was to pilot and strategize a feasible path to whole-proteome tagging in C. elegans and, most critically, test the following key parameters: (1) success rate of triple pools with prior untested reagents at novel targets; (2) utility of fluorophores across expression levels; (3) major effects on tagged protein function. In prior multiplexing, we used two targets which we already knew could be edited quite efficiently, with the 3rd target a point mutation with nearly 100% efficiency. Thus, it was not at all clear that picking 3 random genes and replacing the 3rd highly efficient locus with another less efficient large insertion would work or be sufficiently scalable for thousands of novel genes with unvalidated reagents at first pass.

      The title vastly oversells the advance in my view, and the first sentence of the Discussion seems a more apt summary of the key advance here.

      Some injections target genes on the same chromosome together, which will create unnecessary issues when doing necessary backcrossing, especially if the mutation rate is increased by CRISPR.

      We disagree with the reviewer’s assessment of the need for backcrossing, for two reasons: (1) Prior studies have shown that off-target mutations are not a serious concern in C. elegans (reviewed in PMID: 26336798). For instance, WGS of strains after CRISPR/Cas9 found negligible off-target effects (PMID: 25249454, PMID: 30420468 – using similar RNP/ssDNA method and multiple guides; PMID: 23979577, PMID: 27650892 using other methods). Targeted sequencing studies have reported similar findings, using various CRISPR/Cas9 methods, with essentially no mutations at sites other than the intended target (PMID: 23995389; PMID: 23817069). (2) If the goal is to tag the entire genome, the introduction of backcrossing should not reasonably be a routine part of the initial tagging.

      Lastly, if one really does want to backcross, the existence of tags on the same chromosome is actually an advantage because it permits selection for recombinants with wild-type chromosomes.

      Also, the need for backcrossing and perhaps sequencing made me wonder if injecting 3 together really is helpful vs targeting each gene separately, since only 5 worms need to be injected.

      Apart from our disagreement regarding backcrossing, we are puzzled by the reviewer’s comment. Why would one do single tagging at a time, rather than triple tagging if the whole point is to scale up tagging? It is important to keep in mind that the rate limiting step for tagging the whole genome is the number of injections that can be done per day. Since there is no cloning to generate the repair templates/guides and all other reagents are commercially available and not sample specific, these can be prepared quite rapidly. Being able to isolate multiple lines (together or independently) from the same injection increases throughput 3-fold and in our view does not provide any disadvantages as individual tags can be isolated independently if desired.

      Beyond the numerous technical advantages pooling provides (also lower cost and throughput for making injection mixes as well as imaging), our results show that it yields epistemic benefits as well: we would never have noted the subcellular pattern in Fig. 6B, C with different sets of mitochondria being marked by different mitochondrial proteins had we imaged them separately or even aligned to a pan-mitochondrial landmark. As we mentioned in the discussion, grouping proteins predicted to localize to the same compartment together can simultaneously test how uniform or differentiated such compartments are during the screen.

      The limited utility of current blue fluorescent proteins makes me wonder if it's worth using at all at this stage, before there are better blue (or far red) fluorescent proteins.

      We do not think that the utility of current BFPs is that limiting. At least the theoretical brightness of mTagBFP2 is comparable to that of EGFP (PMID: 30886412), which was useful for the bulk of currently tagged proteins. Due to modestly higher autofluorescence in the blue spectrum, the practical brightness is somewhat less ideal, but we have shown that many proteins are expressed high enough to be detected quite well with mTagBFP2 by eye at low magnification. We also note that many tags that are not visible by eye under a dissection scope become visible with long exposure cameras of widefield microscopes or modern confocal (GaAsP) detectors, so the list of genes detectable with mTagBFP2 is likely to be much higher. We routinely use mTagBFP2 to super-resolve subnuclear structures with endogenous tags (e.g., in the nucleolus), with some tags having lower annotated FPKMs than the genes tested here.

      Some literature reviews, particularly in the Introduction and Abstract, rely too much on recent examples from the authors' laboratory instead of presenting the state of the field. I'd like to have known what exactly has been done with simultaneous injection targeting multiple loci more thoroughly, comparing what has been accomplished to date by various laboratories' advances to date.

      We are not sure what the reviewer is referring to. In the Abstract, we do not refer to any literature. In the Introduction, we cite 28 papers, 6 of those from our lab (4 of which providing examples of protein tags). We do not believe that this can be fairly called an unbalanced presentation of the state of the field.

      This being said, we have gladly expanded our Introduction to provide more background on co-CRISPRing. Labs have routinely used co-conversion (“coCRISPR”) markers for picking out their intended edits (e.g., point mutations or insertions), as it has been shown by multiple groups that a CRISPR/Cas9 edit at one locus correlates with efficiency at other simultaneous targets (PMID: 25161212). Generally, making point mutations with the Cas9/RNP protocol is highly efficient, especially at specific loci such as dpy-10. However, multiple FP-sized insertions have not been routinely attempted. We and only one other group have successfully attempted it using previously working targets and reagents (e.g., 28% in PMID: 26187122). Importantly, the efficiency of such multiple insertions has never been assessed at scale and using entirely untested reagents at novel sites – critical parameters to determine for a whole genome approach. So, we test here (1) the efficiency of triple insertions and (2) the chance of getting them with new and untested guides and reagents.

      In our view, since we have to use some injection/coCRISPR marker anyway for those genes which are not expressed at dissecting-scope visible levels (likely most genes), using highly expressed intended targets as improvised markers in a pooled approach makes our approach much more efficient. It allows us to find the worms with the highest chance of yielding CRISPR insertions, which we can screen with higher power methods for the dimmer targets, while enabling us to co-isolate other intended targets. Insertions, being often heterozygous in F1, can be segregated independently if desired, or homozygosed together to facilitate maintenance then outcrossed individually by those interested in studying specific genes in more detail.

      In the revised version of this manuscript, we now discuss some of these points in the introduction section:

      “Currently, around 1554 proteins representing 8% of the proteome are estimated to have been endogenously tagged (Leyhr et al., 2025). However, at current rates, tagging the proteome is projected to take around 100 years and likely involve numerous duplicate attempts on a small number of commonly studied proteins (Leyhr et al., 2025). It will thus be crucial for the field to coordinate tagging efforts and scale up tagging protocols to enable coverage of the entire genome at a reasonable timescale and cost. Given the number of injections is a major time-limiting factor, pooling multiple injections into one would at minimum cut tagging time by a factor of 3. In C. elegans, screening for novel CRISPR/Cas9-induced genomic edits is already facilitated either by use of co-injection markers (i.e., plasmids that form extrachromosomal arrays) that yield phenotypes or fluorescence in progeny of successfully injected worms, or co-editing well characterized loci using established and highly efficient reagents which likewise yield visible phenotypes. In the latter approach, termed “co-CRISPR”, worms edited at the marker locus are most likely to also carry the intended edit (Arribere et al., 2014). Recent methods for CRISPR/Cas9 mediated genomic insertions have pushed efficiencies to sufficient levels to simultaneously insert multiple fluorophores (e.g., mNeonGreen and mScarlet) as well as a co-CRISPR marker (dpy-10) at three independent loci in a single injection (Eroglu et al., 2023; Paix et al., 2015). These attempts pooled reagents previously established to work efficiently and targeted genes that were known to yield functional fusion proteins when tagged. Thus, while in principle current methods could allow tagging of at least 3 independent loci in one injection if a co-CRISPR marker is omitted, it is not known to what extent such an approach could be generalized across the genome with previously unvalidated reagents (i.e., guides and repair template homology arms) at novel loci to yield functional tags”

      Reviewer #2 (Public review):

      The manuscript by Eroglu and Hobert presents a set of strains each harboring up to three fluorescently tagged endogenous proteins. While there is technically nothing wrong with the method and the images are beautiful, we struggled to appreciate the advance of this work - who is this paper for?

      We consider this paper to have two purposes: (1) motivate the community to come together to consider such genome-wide tagging approach; (2) provide a reference point for funding agencies that such an aim is not unreasonable and will provide novel interesting insights.

      As a technical method, the advance is minimal since the first author had already demonstrated that three mutations (fluorophore insertion and co-CRISPR marker) could be introduced simultaneously.

      We agree that the basic principle is similar. However, it was not clear that triple pooling three novel large edits would work, given the numbers in our original paper or that it would be scalable.

      The dpy-10 coCRISPR marker previously used is a highly efficient single site, with close to 100% hit rate. We also knew in the earlier study that the two pooled insertions already worked quite efficiently and did not disrupt the function of targeted proteins. Exchanging these plus dpy-10 for three novel tags was not guaranteed to succeed for many potential reasons, including both biological and technical. For instance, such a “marker free” approach necessitates that a significant number of targets in the genome should be expressed highly enough to be visible by fluorescence stereomicroscopy when tagged with current best fluorophores. The chance of disrupting gene function by tagging was also not explored in detail in C. elegans, nor whether one untested guide is generally sufficient. We think that establishing these parameters was meaningful and necessary for the goal of whole genome tagging. We have clarified some of these points in the text.

      As a pilot for creating genome-scale resources, it is not clear whether three different fluorophores in one animal, while elegantly designed and implemented, will be desired by the broader community. 

      The usage of three different fluorophores is largely driven by the ability to co-inject and therefore cut injection effort by a factor of three. Moreover, having all three fluorophores together facilitates imaging and maintenance. Lastly, co-labeling has the potential to reveal unexpected patterns of co-localization or lack thereof (example: two mitochondrial proteins that we found to not have overlapping distribution). We clarified this point in the revised text in both the results and discussion.

      Finally, the interpretation of the patterns observed in the created lines is somewhat lacking. A Table with all the observations must be included. This can replace the descriptions of the observations with the different lines, which could be somewhat laborious for the reader, and are often wrong. There are numerous mistaken expectations of protein expression here, but two examples include:

      We are not convinced that our expectations are mistaken. Below we respond to the reviewer’s specific examples, and we are open to hear from the reviewer about additional cases.

      (1) The expectation that ACDH-10 is enriched in the intestine and epidermal tissues (hypodermis).

      There are multiple paralogs of this protein (see WormPaths or WormFlux) that may share functions in different tissues. There is also no reason to assume that fatty acid metabolism does not occur in other tissues (including the germline). Finally, there are no published studies about this enzyme, so we really don't know for sure what it's doing.

      The expression of acdh-10 is annotated in multiple scRNA datasets as intestine and epidermal enriched (CeNGEN/Taylor et al. 2021, highest in epidermis; Ghaddar et al 2023 highest in intestine). We did not mean to imply that fatty acid metabolism does not occur in the gonad, nor that a paralog of acdh-10 could not be performing the same function in tissues where acdh-10 is not expressed.

      However, this raises an important question: why have different paralogs doing the same thing? Duplicate genes with the same function are generally not evolutionarily stable (PMID: 11073452, PMID: 24659815). That there are such striking tissue specific expression patterns of an essential or widely expressed protein class suggests that paralogs of the gene likely differ in some meaningful parameter that might align with tissue-specific functional needs or regulation. The reviewer’s statement that ‘there are no published studies about this enzyme, so we really don't know for sure what it's doing’ is in fact an excellent demonstration of our point; finding out where the duplicates are expressed can provide a starting point to uncover potential differences between the paralogs. At the very least it can delineate to what degree paralogs diverge in their expression across the proteome and identify which such cases merit further study. In a more ideal scenario, prior information of protein function could indicate that the involved pathway requires tissue specific regulation.

      (2) The expectation that HXK-1 is ubiquitously expressed.

      Three paralogous enzymes are all associated with the same reaction, and we have shown that these three function redundantly in vivo, perhaps in different tissues (PMID: 40011787).

      The cited paper (PMID: 40011787) does not show where they are expressed. We discussed redundancy/paralogs above in point 1, and in our view the same applies here. They may perform the same reaction but are likely to differ in some meaningful way, be it regulation or rate of activity, for them to be stably maintained as functional genes over evolution.

      Moreover, single-cell RNA-seq data (PMID: 38816550) also show enrichment of hxk-1 in gonadal sheath cells.

      The Ghaddar et al. and CeNGEN/Taylor et al. datasets do not show this. The scRNA paper cited (PMID: 38816550) also shows enrichment in neurons, pharynx, coelomocyte and germ cells which we did not note. In our view, these in fact further support our goals: often, transcript datasets alone (frequently used to infer tissue function) do not sufficiently predict protein expression. One can post hoc find an scRNA-seq dataset that aligns somewhat with our protein observations, but how does one know which to trust a priori? Disagreements between transcript datasets will ultimately require resolution at the protein level, in our view.

      To clarify these points, we added the following to the discussion section:

      “We also noted unexpected cell type dependent distributions of proteins involved in broadly important metabolic processes such as ACDH-10, which was depleted from the germline compared to other tissues, and HXK-1, which was highly enriched in the gonadal sheath. Notably, for these as well as other cases, scRNA-seq datasets were not sufficient to deduce a priori the observed cell type specific differences at the protein level. Importantly, many genes encoding metabolic enzymes including acdh-10 and hxk-1 have paralogs that likely perform similar catalytic functions. Yet, duplicate genes with identical functions are generally not evolutionarily stable (Adler et al., 2014; Lynch and Conery, 2000); thus such genes are likely to differ in some meaningful parameter (e.g., regulation or activity) that might align with tissue-specific functional needs. Fully annotating the expression patterns of paralogs at the protein level could indicate which tissues require unique metabolic needs and indicate which paralogous genes have undergone sub- versus neo-functionalization. For those proteins that are less functionally understood, unexpected distributions might indicate which merit further study.”

      The table should have at least the following information: gene/protein name - Wormbase ID - TPM levels of single cell data assigned to tissues for L2, L4, and adult (all published) - tissues in which expression is observed in the lines presented by the authors.

      We added some of this information such as annotated expression levels in young adults from various scRNA datasets (but not larval datasets as we did not image these). We note that each of these studies use different pipelines and report different metrics (scaled TPM/Z-score versus Seurat average expression versus TPM), so comparisons between them are not informative unless they are integrated and analyzed together.

      Reviewer #3 (Public review):

      Summary:

      The authors argue that establishing the expression pattern and subcellular localisation of an animal's proteome will highlight many hypotheses for further study. To make this point and show feasibility, they developed a pipeline to knock in DNA encoding fluorescent tags into C. elegans genes.

      Strengths:

      The authors effectively make the points above. For example, they provide evidence of two populations of mitochondria in the C. elegans germline that differ qualitatively in the proteins they express. They also provide convincing evidence that labelling the whole proteome is an achievable goal with relatively limited resources and time.

      We appreciate the referee’s recognition that whole proteome tagging is feasible.

      Weaknesses:

      Cell biology in C. elegans is challenging because of the small size of many of its cells, notably neurons. This can make establishing the sub-cellular localisation of a fluorescently tagged protein, or co-localizing it with another protein, tricky. The authors point out in their introduction that advances in light microscopy, such as diSPIM, STED, and ISM (a close relative of SIM), have increased the resolution of light microscopy. They also point out that recent advances in expansion microscopy can similarly help overcome the resolution limit.

      (1) Have the authors investigated if the three fluorescent tags they use are appropriate for super-resolution microscopy of C. elegans, e.g., STED or SIM? Would Elektra be better than mTAGBFP2? How does mScarlet3-S2 compare to mScarlet 3?

      All three tags work for ISM (i.e., Airyscan). We previously tried Electra (not for the genes tested here) but could not isolate positive tags. Given Electra is not that much brighter on paper than mTagBFP2 we did not pursue it further, though we recognize that these may simply have been unlucky injections. mScarlet3-S2 is quite a bit dimmer than mScarlet3 on paper – the advantage is that it has higher photostability. In our view, the limiting factor will be having FPs that are bright enough to screen, image and scale to the whole genome, so brightness will likely provide an advantage over photostability at this stage.

      (2) Have the authors investigated what tags could be used in expansion microscopy - that is, which retain antigenicity or even fluorescence after the protocol is applied? It may be useful to add different epitope tags to the knock-in cassettes for this purpose.

      mSG and mSc3 retain fluorescence after fixing with formaldehyde. We have not tested mTagBFP2 fluorescence in fixed worms. We agree that adding different epitope tags would be useful.

      The paper is fine as it stands. The experiments above could add value to it and future-proof it, but are not essential. If the experiments are not attempted, the authors could refer to the points above in the discussion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Merged figures appear saturated, and use colors that won't work for red-green colorblind viewers. 

      For all figures, we also show individual channels separately, which is common practice for making fluorescence images accessible to colorblind readers (PMID: 33788834). Figures highlighting non-overlap like 6B and C are already in accessible colors when merged (blue/green) and include a numerical quantification. 3-color RGB images preserve the greatest information for the highest number of individuals.

      (2) Targeting ubiquitously expressed genes as a proof of concept gives me some concern that this might underestimate the challenges that may be experienced with less widely expressed genes.

      While the genes were predicted to be ubiquitously expressed, many were not in practice, like HXK-1 and F54C8.1, which were also among the lower expressed genes on our list and highly cell type restricted. As discussed, the more tissue restricted a gene, the likelier that bulk RNA levels underestimate expression. Such genes are therefore more likely to be detected in a specific tissue. We routinely isolate tissue restricted endogenous tags, including those expressed in only a few neurons, with bulk FPKMs lower than the ranges tested in this manuscript.

      (3) Some results are not shown or referenced (autofluorescence, for example, is shown using a schematic in Figure 1C).

      We now provide representative images alongside what would be expected to be observed by eye during screening.

      (4) It would be useful to describe how to recover worms from what is shown in Figure 1A. 

      In the revised version, we added the following in the caption for Fig. 1A:

      “Selected worms expressing the brighter tag can be screened for dimmer tags by higher magnification and long exposure imaging. Worms can be recovered directly from slides if immobilized by levamisole as described (Ghanta et al., 2021). Alternatively, single hermaphrodite worms can be isolated, allowed to lay eggs, then screened.”

      (5) A blue bar of data must be missing from Figure 3B injection pool 5.

      As stated in the text, “All but one tag (cox-6B::mTagBFP2) was visible in the F1 generation of injected P0 animals, and these were subsequently isolated among F2 worms positive for the other tags in the pool.”

      To clarify that data points are not unintentionally omitted, we added the following text to the caption of Fig. 3B:

      “For group 5 including cox-6B::mTagBFP2, worms with detectable levels of mTagBFP2 fluorescence were not recovered in the F1 generation but were isolated among progeny of F1s positive for mStayGold and mScarlet3; we were thus unable to quantify efficiency for this locus at F1.”

      (6) Some expression or localization patterns were unexpected, but complications like germline silencing and protein mislocalization, with a small fraction localizing normally and rescuing function, were not presented as possibilities. Viability is used to confirm function, but without presenting whether this means 100% viability, less, or just the ability to maintain a strain.

      We already do discuss mislocalization and functionality issues in the Discussion, as well as tradeoffs of alternate methods. Any existing method to observe biological molecules, be it protein, RNA or DNA, has multiple drawbacks and sources of artifacts, which are unlikely to be fully eliminated in the foreseeable future.

      In regard to germline silencing of endogenously tagged genes in C. elegans, there is actually very little evidence for this. Collectively, various labs have now generated over 200 reporter alleles of germline-expressed genes (WormTagDB), with robust expression throughout the germline and retention of function. Likewise, numerous of our tags across fluorophores showed robust germline expressions including EEF-1A.1::mTagBFP2, Y22D7AL.10::mStayGold, and HAT-1::mScarlet3. In fact, overall transcript levels generally tended to underestimate germline enrichment at the protein level. We note that single-copy transgenes driven by eef-1A.1/eft-3 promoter by itself are frequently not expressed in the germline (PMID: 31064766); that we could detect EEF-1A.1 robustly in the germline when tagged endogenously is evidence that silencing is unlikely to be a widespread concern, and at the least less of a concern than single copy transgenes. We appreciate that for a transgene, presence/absence of specific sequence elements and genomic loci play a role in expression, but an endogenous tag captures all such information at a given locus.

      Indeed, we found only two reports of endogenous tags being silenced in the germline, the first being a novel tag (not fluorophore) which initially prevented expression at the tagged locus (PMID: 30109984), but after making changes to the sequence to avoid silencing signals the authors could rescue expression and thereafter saw robust expression in various novel contexts with this tag. The second example (PMID: 34547227) leaves open the possibility that germline repression of that particular gene might be a part of its endogenous regulation.

      Nevertheless, given it is probably rare if occurs at all, it will likely take a large scale tagging effort to uncover such cases at sufficient numbers to study. In our view, this further justifies tagging at large, ideally genomic, scales. If we do discover that there are numerous annotated germline proteins which we don’t observe by tagging, that would be interesting to study on its own.

      (7) Halotag is presented in the Discussion as a small tag, but it is bigger than GFP.

      Thank you for catching this. We have removed the discussion of Halotag. Given the comparable size to FPs, it would be unlikely to alleviate issues of tag functionality.

      (8) It would be useful to include FPKMs and viability percentages in Table 1.

      FPKM is included in column 6, but the title for this column is cut off. In the revised table FPKM values are now shown more clearly across stages.

      We did not quantify viability percentage. In our view it does not yield an informative metric when there is little information about the protein’s required dosage for function, which was the case for most proteins here. A haplosufficient gene might yield a full brood size even if 50% of protein function is lost; conversely, a highly dose sensitive protein could yield penetrant and severe inviability with mild perturbation of function. It also is not actionable information at this stage if there is no alternate tagging strategy as a baseline of comparison. The worms we picked to image all have viable embryos as adults, so in those individuals the genes were likely to be sufficiently expressed and functional.

      (9) Because establishing that a guide works well is a limiting step for many CRISPR experiments (once a guide works well, it's easy to inject 5 worms and get lines), I wondered if testing that for many genes is what is really needed in the field at this stage. 

      Guide quality is rarely an issue in C. elegans, as for all the genes here we tried only one guide, all of which were previously untested. We now clarified this in the discussion section:

      “Notably, we find that previously untested guide RNAs and homology arms perform exceptionally well at novel loci, as we only tested one set of reagents for each locus which yielded satisfactory tagging rates.”

      (10) For a manuscript where the injection is so central to what was done, I was surprised to read in the Acknowledgments that all of the injections were done by someone who is not included as an author.

      We are likewise surprised by such a comment but gladly clarify: Chi Chen has been with us as an expert microinjection specialist for more than 25 years and her very important technical contributions have been acknowledged in many dozen papers. Multiple authorship guidelines, including COPE’s and ICMJE’s, state that technical contributions alone do not qualify for authorship.

      Reviewer #2 (Recommendations for the authors):

      (1) We would encourage the authors to provide systematic validation of the reported insertions. The manuscript reports that 24 of 30 tags were isolated and visible, but does not clearly state whether each isolated line was confirmed by sequence‑level validation to be correctly in‑frame and free of unintended mutations at the target locus.

      We appreciate the reviewer’s concerns on fidelity. These parameters have been assessed in prior published work (e.g., PMID: 30504364, PMID: 34748534) and in our hands are in the range of 80% whenever we sequence non-fluorescent tags of similar sizes. The efficiencies we observed are high enough that one can expect to recover numerous worms with the exact intended sequence for each target, though we would argue mutations within the FP reporter are less likely to matter if it retains high fluorescence.

      (2) The manuscript presents aggregated success counts (e.g., 8/10 mTagBFP2 tags, 9/10 mStayGold, 7/10 mScarlet3) and useful narrative descriptions of injection outcomes. We also suggest including per‑locus success rates.

      Figure 3B shows per locus success rate and source data is provided for this figure. Each dot is an individual injection and the Y axis is per locus rate. We now worded this more clearly in the figure’s caption.

      “Total insertion efficiencies per locus for the indicated targets across injection pools.”

      (3) For pools that required re‑injection after initial failures, we would like to see a description of the specific changes that were made to the injection mixes or procedures (e.g., new repair template prep, different Cas9 reagent lot, guide redesign). This will be useful troubleshooting information for others.

      We re-made the exact same injection mix but with nanodrop to ensure the purity of the repair templates as assessed by absorbance ratios (A260/230 and A260/280) were sufficient after each purification step. No other changes were made. This is now specified in the methods section in the following way:

      “For re-runs of pools 4, 6 and 10 which failed initially, we regenerated the repair templates and ensured that after each column purification, the A260/230 ratio of the purified DNA was ≥2.2 and A260/280 was 1.8 ± 0.05 when measured with a Nanodrop spectrophotometer.”

      (4) The authors state that the fluorophore sequences are codon-optimized for C. elegans. We suggest they provide the exact donor/tag sequences, specifically state whether the fluorophore sequences contain any synthetic/artificial introns, or whether other sequence modifications (e.g., silent PAM‑disrupting mutations) were included in the donor templates. 

      This information is provided in Supplementary Table 1.

      (5) Page 3: Include a reference for "The C. elegans genome encodes around 20,000 genes" 

      We added a reference to the most recent release of the genome (WS237, May 2013). Spieth et al., 2014.

    1. eLife Assessment

      This important paper substantially advances our understanding of how Molidustat may work, beyond its canonical role, by identifying its therapeutic targets in cancer. This study presents a compelling and well-structured investigation into the therapeutic vulnerabilities of APC-mutant colorectal cancer. This work will be of broad interest to the cancer community in studying small molecules and their therapeutic targets.

    2. Reviewer #1 (Public review):

      [Editor's note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      The authors aimed to uncover novel therapeutic vulnerabilities in APC-mutant colorectal cancer (CRC), which constitutes the majority of CRC cases. They hypothesized that modulating oxygen-sensing pathways (via PHD inhibition) could disrupt adaptive stress responses in these tumours.

      Strengths:

      The study employs a powerful, two-pronged approach to identify Molidustat's targets. By using both Thermal Proteome Profiling (TPP) and an orthogonal chemical proteomic competition assay, the authors provide compelling evidence that GSTP1 is a genuine, direct off-target, effectively addressing the common limitation of indirect effects in proteomic screens.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aimed to determine Molidustat targets and the potential utility of these findings. They clearly demonstrate that Molidustat interferes with GSTP1 and some other proteins on top of PHD2. They also demonstrate that PHD2 deletion is not sufficient to recapitulate Molidustat effects in cells and proteomes. Finally, they demonstrate synthetic lethality in organoids for Molidustat and APC deletion.

      Strengths:

      The data on Molidustat proteomes, GSTP1 binding, inhibition and metabolic health of organoids is really clear. All biochemical, docking and omic data are really strong. The potential impact of these findings could be the use of Molidustat in APC null tumours and awareness of potential off-target effects.

    4. Reviewer #3 (Public review):

      In this paper, the authors revealed that Molidustat can induce a dose-dependent increase in Caspase-3/7 activity in the HT29 cell line, which is an APC-mutant colorectal cancer cell line. More importantly, they found that targeting PHD2 alone cannot cause cell death. By using thermal proteome profiling (TPP) and orthogonal chemical proteomic competition assays, they determined GTSP1 as a previously undiscovered off-target of Molidustat. They also revealed that combined PHD2 and GSTP1 loss leads to an increase in intracellular ROS and apoptosis. Moreover, they evaluated the effects of Molidustat in colonic organoids and showed that Molidustat has a high selectivity for colonic organoids with activated WNT signaling and/or KRAS pathway alterations, and this effect is not reproduced by hydroxylase inhibition alone, providing a new potential approach to targeting both PHD2 and GTSP1 for the treatment of APC-mutant CRC.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aimed to uncover novel therapeutic vulnerabilities in APC-mutant colorectal cancer (CRC), which constitutes the majority of CRC cases. They hypothesized that modulating oxygen-sensing pathways (via PHD inhibition) could disrupt adaptive stress responses in these tumours.

      Strengths:

      The study employs a powerful, two-pronged approach to identify Molidustat's targets. By using both Thermal Proteome Profiling (TPP) and an orthogonal chemical proteomic competition assay, the authors provide compelling evidence that GSTP1 is a genuine, direct off-target, effectively addressing the common limitation of indirect effects in proteomic screens.

      Weaknesses:

      (1) In Figure 1, the current data rely on a single guide RNA (sgRNA). To make the data solid, at least two independent sgRNAs targeting different regions of PHD2 should be used.

      We thank the reviewer for raising this. Clarity on the CRISPR strategy was missing from the original submission and we have now added the following to the Methods (Page 4). We did not use a single sgRNA. PHD2 was targeted with a pool of three chemically modified crRNAs:

      (IDT Alt-R; target sequences: 5'-TACAACCAGCATATGCTACA, 5'GTGGCTGCCGAAGCCGAGCC, 5'-GATAAGATCACCTGGATCGA)

      Delivered as in vitro assembled ribonucleoprotein complexes with high-fidelity Cas9. This format has been reported to achieve high on-target efficiency while minimising off-target cutting [1,2] such that any residual stochastic off-target events are distributed across the population and are not expected to manifest as a coherent phenotype at the population level. Working with pooled, unselected knockouts rather than single-cell clones also avoids the confounds of clonal heterogeneity that normally motivate the use of multiple independent guides and rescue experiments in single-clone workflows. We have previously validated this approach for GSTP1 knockout in a separate single-cell proteomics study [3], where loss of GSTP1 protein was observed in over 90% of single cells and GSTP1 was the most significantly altered protein between sgControl and sgGSTP1 populations.

      (2) Figure 3E: Asn205 site should be mutated to prove that whether Molidustat inhibits GSTP1 activity via Asn205 or not.

      This is a good suggestion, and we explored it in silico before concluding it was not tractable. We used PyMol mutagenesis to model Molidustat binding to GSTP1 variants at the predicted contact residues: Asn205 was mutated to Ala, Gly and Ser; Trp39 (predicted to hydrogen-bond Molidustat) was mutated to Ala, Phe and Thr; and a Tyr8Phe/Asn205Ser double mutant was also modelled. In every case, Molidustat reoriented within the active site and adopted an alternative hydrogen-bonding configuration (most commonly with Tyr8), yielding a docking score equal to or better than binding to native GSTP1 (Author response image 1– Author response image 4). The model therefore does not predict any single or double point mutant that would ablate Molidustat binding in a clean, interpretable way, and we could not design a rational loss-of-interaction mutant on this basis. Given this limitation, and that definitive mapping of the binding interface would require co-crystallography, which is beyond the scope of the present study, we have moved the docking model to the supplement and flagged it as predictive rather than definitive.

      Author response image 1.

      Molidustat in native GSTP1

      Author response image 2.

      Molidustat docking with mutated GSTP1, Asn205 mutated to Gln205

      Author response image 3.

      Molidustat docking with mutated GSTP1, Tyr39 mutated to Phe39

      Author response image 4.

      Molidustat docking with mutated GSTP1, Asn205 mutated to Ser205 and Tyr8 mutated to Phe8

      (3) Figure 5B and 5C: The metabolic imbalance phenotype observed upon dual knockout of PHD2 and GSTP1 requires rescue experiments to confirm on-target specificity.

      We thank the reviewer for this important point and agree that rescue experiments could represent the most direct demonstration of on-target specificity for the metabolic phenotype observed in Figures 5B and 5C. These rescue experiments are necessary when working with single clones, as they allow for comparing a knock-out clone with a reconstituted pool and sidestep the issue of clonal heterogeneity.

      In our case, we think that there is no advantage to doing so, as we work with pooled knockouts, so any clonal heterogeneity is diluted in the pool.

      One could even make the case that such a rescue experiment would introduce additional artefacts. Combined loss of PHD2 and GSTP1 leads to reduced cellular viability, with decreased proliferation and increased apoptosis, consistent with a synthetic lethal interaction. To devise a rescue experiment, we would have to isolate a single-cell clone (the pool is not a complete 100% knock out, WT cells would outgrow the knock out cells). The isolation of such a clone that has overcome the anti-proliferative insult of the double knockout is likely to have a phenotype distinct from the original, pooled population, as would the rescued have from the WT cells. For these reasons, we have not performed rescue experiments in the current study. We have added the absence of a rescue as a limitation to the study in the discussion

      “While genetic rescue experiments would provide definitive confirmation of on-target specificity, the pronounced loss-of-fitness and apoptotic phenotype observed upon combined PHD2 and GSTP1 loss limited the feasibility of establishing stable rescued double-knockout populations, and therefore represents a limitation of the current study.”

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to determine Molidustat targets and the potential utility of these findings. They clearly demonstrate that Molidustat interferes with GSTP1 and some other proteins on top of PHD2. They also demonstrate that PHD2 deletion is not sufficient to recapitulate Molidustat effects in cells and proteomes. Finally, they demonstrate synthetic lethality in organoids for Molidustat and APC deletion.

      Strengths:

      The data on Molidustat proteomes, GSTP1 binding, inhibition and metabolic health of organoids is really clear. All biochemical, docking and omic data are really strong. The potential impact of these findings could be the use of Molidustat in APC null tumours and awareness of potential off-target effects.

      Weaknesses:

      A main but minor weakness is that Molidustat also inhibits other PHDs, although these are less expressed. PHD1 has been shown to control the cell cycle and be expressed in the colon, where it is needed for viability. Although this does not explain the lack of effect of other PHD inhibitors, it does warrant some discussion. The use of MTT is not very good to detect viability when it measures metabolism; this also needs to be discussed and perhaps supplemented with colony or cell number measurements.

      Great point, for this reason, we have assayed apoptosis throughout. In addition, we have added a clonogenicity assay with APC organoids. Organoid cells were treated with an acute dose of Molidustat. We subsequently measured the level of Lgr5 (a stem cell marker) and of the ability of the cells to generate organoids (these data have been added as Figure 5 F-G.)

      Reviewer #3 (Public review):

      In this paper, the authors revealed that Molidustat can induce a dose-dependent increase in Caspase-3/7 activity in the HT29 cell line, which is an APC-mutant colorectal cancer cell line. More importantly, they found that targeting PHD2 alone cannot cause cell death. By using thermal proteome profiling (TPP) and orthogonal chemical proteomic competition assays, they determined GTSP1 as a previously undiscovered off-target of Molidustat. They also revealed that combined PHD2 and GSTP1 loss leads to an increase in intracellular ROS and apoptosis. Moreover, they evaluated the effects of Molidustat in colonic organoids and showed that

      Molidustat has a high selectivity for colonic organoids with activated WNT signaling and/or KRAS pathway alterations, and this effect is not reproduced by hydroxylase inhibition alone, providing a new potential approach to targeting both PHD2 and GTSP1 for the treatment of APC-mutant CRC.

      Specific comments:

      (1) What is the possible molecular mechanism of dual GSTP1/PHD2 loss, inducing cell death?

      This is an important question. Our data support a model in which combined loss of GSTP1 and PHD2 disrupts cellular redox homeostasis, leading to accumulation of reactive oxygen species, increased GSSG/GSH ratios, and depletion of antioxidant buffering capacity. This redox imbalance is accompanied by downregulation of pro-survival pathways. In this context, activation of apoptotic signalling, as evidenced by increased caspase-3/7 activity and proteomic enrichment of apoptosis-associated pathways, contributes to the observed cell death phenotype.

      While apoptosis is supported by our data, the magnitude of oxidative stress suggests that additional oxidative stress-associated cell death mechanisms may also contribute. We have clarified this point in the Discussion (Page 11).

      (2) Can the authors mutate the binding site of Molidustat on GTSP1 to verify the in silico docking results?

      This is a very important question. Currently, the model is of limited value. Reviewer 1 had a similar question. Can we refer you to Reviewer 1, question 2.

      (3) Evidence for Molidustat inhibiting PHD2 activity or stabilising HIF-1α should be provided.

      We thank the reviewer for this suggestion. Data showing HIF-1α stabilisation and evidence of downstream signalling is now added to Supplementary Figure 1.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I only have minor suggestions:

      Molidustat also inhibits other PHDs, although these are less expressed. PHD1 has been shown to control the cell cycle and be expressed in the colon, where it is needed for viability. Although this does not explain the lack of effect of other PHD inhibitors, it does warrant some discussion. The use of MTT is not very good to detect viability when it measures metabolism; this also needs to be discussed and perhaps supplemented with colony or cell number measurements.

      This is correct, PHD1 is of particular interest, given the effects inhibition/knock-out has on the inflamed colon. We have added a new paragraph to the Discussion (Page 13) that addresses the isoform selectivity of Molidustat. We note that, although developed as a PHD2 inhibitor, Molidustat retains appreciable activity against PHD1 and PHD3 [4], and we discuss the non-redundant and in some contexts opposing roles of PHD1 and PHD2 in the colon, PHD1 loss is protective in DSS colitis [5] and restrains colitis-associated tumour growth, whereas PHD2 loss in the tumour and stroma is reported to inhibit metastasis and treatment response [6]. We further note that this pattern of isoform engagement is shared with other pan-PHD inhibitors that did not phenocopy Molidustat in our screens, indicating that PHD isoform profile alone is insufficient to explain Molidustat’s distinctive activity and pointing to GSTP1 off-target engagement as the key distinguishing feature. We argue that localised colonic delivery (as discussed earlier in the Discussion) would concentrate drug at the APC-mutant epithelium while limiting systemic exposure.

      We fully agree with the reviewer, MTT measures metabolic activity/NADH levels rather than viability in the strict sense, and that this is particularly relevant for a compound that perturbs redox metabolism. We have added a clonogenicity assay in APC organoids (Fig. 5 F-G) to supplement the MTT and Cleaved Caspase 3 assays already present in the manuscript.

      (1) Lee, J. K. et al. Directed evolution of CRISPR-Cas9 to increase its specificity. Nat. Commun. 9, (2018).

      (2) Sakovina, L., Vokhtantsev, I., Vorobyeva, M., Vorobyev, P. & Novopashina, D. Improving Stability and Specificity of CRISPR/Cas9 System by Selective Modification of Guide RNAs with 2′-fluoro and Locked Nucleic Acid Nucleotides. Int. J. Mol. Sci. 23, (2022).

      (3) Makar, A. N., Holkham, J., Lilla, S., Wilkinson, S. & von Kriegsheim, A. Overcoming preservation challenges to enable single-cell proteomics of fixed cell and tissue samples with retained proteome integrity. Preprint at https://doi.org/10.1101/2025.03.10.642380 (2025).

      (4) Flamme, I. et al. Mimicking hypoxia to treat anemia: HIF-stabilizer BAY 85-3934 (molidustat) stimulates erythropoietin production without hypertensive effects. PLoS One 9, (2014).

      (5) Tambuwala, M. M. et al. Loss of prolyl hydroxylase-1 protects against colitis through reduced epithelial cell apoptosis and increased barrier function. Gastroenterology 139, (2010).

      (6) Leite de Oliveira, R. et al. Gene-Targeting of Phd2 Improves Tumor Response to Chemotherapy and Prevents Side-Toxicity. Cancer Cell 22, (2012).

    1. eLife Assessment

      This study used pupillometry to provide an objective assessment of a form of synesthesia in which people see additional color when reading numbers. It provides convincing evidence that subjective color ratings are matched by changes in pupil size that recapitulate brightness-mediated changes when exposed to the real color. The work provides a valuable contribution to the literature on both synesthetic perception and the use of pupillometry to probe perception and related psychological processes.

    2. Reviewer #1 (Public review):

      Summary:

      Knowing that small pupil-size variations accompany brightness variations (even when these are illusory), the authors asked whether pupil constrictions would accompany the synesthetic perception of a brighter color (compared with a darker one), induced by the presentation of a black-white character. This grapheme-colour synesthesia is only experienced by few participants, sixteen of whom were enrolled in this study. The results reliably showed that a relative pupil constriction would "betray" the perception of a brighter color in these participants, while no such effect would be observed in control participants who were asked to report a color in association with each grapheme, even though they did not perceive any.

      Strengths:

      The main strength of the study lays in its combination of psychophysics (brightness ratings) and pupillometry, which allowed for showing clear-cut results.

      Weaknesses:

      I only see the following relatively minor weaknesses, namely:

      - The pupil traces in Figure3 (main results) are heavily pre-processed (per-participant demeaned), loosing any feature besides the effect of interest. As I argued in my first review, I worry that this format gives unrealistic expectations about the effect (the perception of dark/bright colors do not generate a net dilation/constriction of the pupil; perception-related modulations of pupil size are always relative and generally small compared to the numerous other effects registered in pupil size; these include a pupil dilation that is more prominent in the controls and that gets analyzed later on in the manuscript; I do not think that eliminating one of the effects of interests from a main results figure helps the reader understand the results). In the revised manuscript, the authors addressed this concern by adding a Supplementary Figure 4, where a more complete representation of the results is shown (traces from individual trials are baseline corrected and averaged, resulting in more informative timecourses). I would strongly recommend that Supplementary Figure4 is brought to the main text (Figure3 could be presented in Supplementary).

      - Responses to physical brightness modulations were only measured in the synesthethes group, not in controls. The authors point out that pupillary light responses have been thoroughly characterized in previous studies, and conclude that synesthethes' responses were in line with the expectations both in terms of amplitude and latency. However, as we are not dealing with standardized measurements, subtle differences in pupil reactivity across the two populations remain a possibility. I recommend that this possibility is mentioned in the discussion.

      Impact:

      This work is likely to improve our understanding of synesthesia, providing a new tool to quantify the subjective sensations; an interesting potential extension would be using pupillometry for tracking changes over time of the synesthetic experiences, opening up the possibility to evaluate the importance of learning for this peculiar experience.

    3. Reviewer #2 (Public review):

      Synesthesia is a neurological condition where stimulation of one sensory channel leads to involuntary, automatic, and consistent experience of another, unrelated percept. For example, Sir Francis Galton (1880, Nature) famously described the robust tendency of some individual (synesthetes) to associate numerals with a distinct color. Ever since, synesthesia keeps attracting a broad interest in the cognitive neurosciences in light of its implications for the study of domains such as perception, consciousness, and brain connectivity, among others.

      Strauch, Leenaars, and Rouw measured pupil size in a group of 16 grapheme-color synesthetes and two matched control groups. The participants were presented with gray digits - that is, visual stimuli having identical physical properties in terms of brightness. Each participant subsequently rated the corresponding evoked color and brightness: unlike controls, synesthetes did so in a very consistent and reliable fashion. Accordingly, this was also shown in their pupils: despite the same objective luminance, digits associated with brighter percepts caused their pupils to constrict and digits associated with darker percepts caused their pupils to dilate more than controls. These results highlight how crossmodal correspondences are deeply rooted in synesthetes, and puts forward pupillometry as a particularly appealing biomarker for some phenomenological experience (at least those grounded in "brightness").

      Further strengths of the technique are its temporal resolution and its responsiveness to several constructs. Across several tasks, the authors show for example that responses to synesthetic light are somewhat slower than responses to real light (i.e., they are likely mediated), but at the same time faster than responses to mental imagery. The role of mental imagery can also be reasonably dismissed when considering the second feature of pupil size: its responsiveness to mental effort and cognitive load. The pupils tend to dilate with demanding, challenging tasks, and this was the case when control participants were asked to report the color of a digit for which they did not consistently experience a synesthetic association. The same task was, instead, seemingly effortless for synesthetes, again speaking in favor of the automaticity of number-color correspondences in their case.

      Overall, the findings by Strauch, Leenaars, and Rouw are highly significant for the field and likely to be impactful. The strength of their evidence, when accounting for the relatively small sample size and the inherent variability of both phenomenology (color perception and subjective reporting) and physiology (pupil size), is adequate and sufficiently convincing.

      Comments on revisions:

      I thank the authors for addressing all my comments in a satisfactory way. I think that the paper has improved, especially in terms of transparency of the reporting and clarity of the results.

    4. Reviewer #3 (Public review):

      Summary:

      In the present study, the authors examined pupillary responses to uncolored stimuli (number graphemes) among number-color synesthetes and non-synesthetes. After seeing a digit, the synesthetes and active control participants were asked to indicate which color they perceived using three dimensions of hue, saturation, and lightness. The lightness values were the primary independent variable for follow-up analyses. To see how the pupil responded to psychologically "bright" and "dark" digits, the authors split the reported lightness values at the median and plotted them. The synesthetes showed a pupillary constriction to digits they perceived as bright and dilation to digits they perceived as dark. Active control participants did not show that effect. In a subsequent block, only the synesthetes were shown the colors they reported perceiving as colored discs. Their pupillary responses were similar. The authors also found that the differences in pupillary responses between light and dark perceptions (with digits) were only slightly delayed in their onset to the perception of a colored disc, and therefore the color perception accompanying a digit is unlikely to be effortful or a retrieved association, but occurs rather automatically.

      Strengths:

      The authors employed a well-controlled and designed quasi-experiment comparing color-grapheme synesthetes to non-synesthetes and showed convincingly that the color perceptions accompanying graphemes alter the physical perception of brightness. They also made a reasoned attempt to ruled out the possibility that color associations are occurring effortful via retrieved associations.

      The follow are questions which I had asked in a first round of reviews, and which were answered adequately by the authors:

      (1) Are the pupillary responses among synesthetes, which objectively do not seem to match the degree of physical stimulation entering the retina, in any way maladaptive for eye functioning? I understand the constriction/dilation of the pupil to not only benefit visual acuity but also to protect the retina from damage. Are synesthetes at any risk of retinal damage due to over-dilation of the pupil to brighter stimuli? Or are these effects of a magnitude that is too small to matter? As reported in arbitrary units, it was hard to know how large these effects were in terms of measurable changes in dilation (e.g., millimeters).

      (2) Likewise, is the automatic synesthetic merging of two percepts something that could be learned such that natural synesthetes and "artificial" synesthetes would look similar? For example, if a group of non-synesthetic participants were to learn a color-grapheme association to automaticity, would you expect their pupillary responses to the graphemes look similar to the synesthetes? If so (or if not), what would this tell us anything about the phenomenology of synesthesia?

      (3) Do the synesthetic perceptions of digit graphemes merge in a sensible way? For example, if a synesthete sees a particular color with the digit 1, and a different color with the digit 9, what do they perceive when they see 19? or 1-9, or 1 9? Is there color blending, or an altogether different color perception?

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study used pupillometry to provide an objective assessment of a form of synesthesia in which people see additional color when reading numbers. It provides convincing evidence that subjective color ratings are matched by changes in pupil size that recapitulate brightnessmediated changes when exposed to the real color. The work provides a valuable contribution to the literature on both synesthetic perception and the use of pupillometry to probe perception and related psychological processes.

      We were pleased to learn that our manuscript was of interest to the reviewers and the editor. We thank the reviewers for their useful feedback and have addressed all their comments in the revised version. We here give the most prominent changes as quotes.

      We thank all reviewers and for their very helpful input.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Knowing that small pupil-size variations accompany brightness variations (even when these are illusory), the authors asked whether pupil constrictions would accompany the synesthetic perception of a brighter color (compared with a darker one), induced by the presentation of a blackwhite character. This grapheme-colour synesthesia is only experienced by a few participants, sixteen of whom were enrolled in this study. The results reliably showed that a relative pupil constriction would "betray" the perception of a brighter color in these participants, while no such effect would be observed in control participants who were asked to report a color in association with each grapheme, even though they did not perceive any.

      Strengths:

      The main strength of the study lies in its combination of psychophysics (brightness ratings) and pupillometry, which allowed for showing clear-cut results.

      Weaknesses:

      Some relatively minor weaknesses concern the ancillary analyses, which tackle secondary questions and are not entirely convincing.

      (1) The linear mixed model approach is a powerful way to identify important variables, but it does not clarify whether the key factors are between-subject or between-trial variations. Some variables are inherently defined at a subject level (e.g., PA scores), others are not. I would strongly recommend an alternative visualisation of the results to examine inter-individual variability.

      Visualizing the highly idiosyncratic effects is indeed challenging. Addressing R1’s point 4 and a point brought up by R2, we updated all figures to now visualize pupil size in millimeters instead of arbitrary units. Furthermore, we added a supplementary figure (supplementary figure 4) that visualizes pupil size change without demeaning (please see reply to point 4).

      To get a better grasp of the interaction between lightness and coupling strength, we further included the supplementary figure 5 that splits by lightness and coupling strength in synesthetes.

      Furthermore, as this review and response will be publicly available, Author response image 1 provides participant-mean traces per lightness bin in addition to the overall means and hopefully makes the stability/variability of effects visually clearer (in addition to the strip plots that attempt this for the average response).

      Author response image 1.

      We hope that these additional visualizations make the effects of interest more transparent. Ultimately, however, the LME figure likely provides the information best, albeit at the cost of complexity.

      (2) It is not clear why taking the first derivative of pupil size in Figure 5 would isolate the effect of arousal, eliminating those of luminance and contrast changes (in fact, one could argue for the opposite, since arousal effects are generally constant for extended periods of time while contrast effects are typically more local and transient).

      First, please note that the results in 2.3.1 cannot be explained by task or context effects such as luminance and contrast: the exact same active color reporting task (same task and context) was presented to synesthetes and non-synesthetes.

      Indeed, the reviewer is correct that the first derivative does not eliminate other concurrent pupil-driving effects, that was expressed wrongly in our original text. Indeed, any stimulus-locked effect, such as the luminance and contrast effects, but also the effort effect will reflect similarly in the derivative measure.

      We did take the derivative because pupil responses driven by other non-trial related activity, such as increasing tiredness or excitement over the course of trials differ almost by necessity between participants, thus creating variability. However, these effects are most likely happening at a slower timescale and thus show less in the derivative measure. Accordingly in past research, we previously found clearer response-locked effects in the past when using a derivative measure (Douze et al., 2025; Ten Brink et al., 2024). This way, we also hoped to get rid of such variability that happens between participants for this between participant analysis.

      Even if we were to use the same baseline corrected analysis, we would arrive at the same conclusion: we here directly compared baseline-corrected pupil sizes by taking individual differences into account (using a LME). In other words, we tested for the same question, but not relying on the derivative. We thus compared baseline-corrected pupil sizes using over-time LMEs. Group (active control vs. synesthete) gained significance between ~1.7s and 3s, aligning with the derivative-based result.

      Author response image 2.

      t-values of a per-time point LME predicting pupil response from group (synesthete/active control) Group reached significance.

      In sum, we deem the derivative more powerful/more appropriate in this context, but the interpretation of findings does not hinge on that analysis choice (as can be seen in the Author response image 2).

      We corrected the claims on the derivative as a measure cleaning out other effects that indeed was oversimplified as it stood. We now write:

      “Mental effort presents in task-evoked pupil dilations, yet other factors simultaneously affect the pupil, such as luminance and contrast changes at trial onset, as well as slower trends across the session (e.g., fatigue). To reduce the influence of these slower, non-trial-locked fluctuations while retaining the trial-evoked dynamics, we calculated the first derivative of the pupil time course to assess the velocity of pupillary changes (Butterworth filter, 18 Hz, order 3, 2.5 Hz lowpass, following our previous works [60, 61]).”

      Douze, B. T., Ten Brink, A. F., Dijkerman, H. C., & Strauch, C. (2025). Pupil responses objectively index pharmacologically altered tactile sensitivity. Cortex, 193, 90-104.

      Ten Brink, A. F., Heiner, I., Dijkerman, H. C., & Strauch, C. (2024). Pupil dilation reveals the intensity of touch. Psychophysiology, 61(6), e14538.

      (3) It is a pity that responses to physical brightness modulations were only measured in the synesthete group, not in controls, as this would have allowed for ruling out differences in pupil reactivity across the two populations.

      The reviewer is correct that this would allow additional comparisons, but argue that light responses in healthy control samples are very well documented and stereotypical. For instance, Bergamin & Kardon (2003) provide very systematic latency estimations, for low-luminance change stimuli in the realm of about 320ms that can accelerate to about 250ms for very strong luminance changes. Our relatively small luminance increments should thus be expected in this range. Indeed, this also well describes the response latencies we observed in synesthetes when exposed to the colored disks. While there is no detailed information about participants in Bergamin & Kardon (2003), data from previous studies shows very similar pupil light response profiles in a healthy student control population that matches our synesthetes well demographically (Strauch, Romein et al., 2022 Figure 2a, exact same lab as for the present study; Koevoet et al., 2025 Figure 3a). See also the further responses, baseline pupil size in millimeters across groups did not differ.

      Together, we can safely conclude that pupil light responses in synesthetes are not different from pupil light responses in controls. We agree with the reviewer that this is a sensible point to also make in the manuscript:

      “Specifically, pupil size first responded significantly to physical luminance after 330 ms (see Supplementary Figure 7 for per-timepoint LME; in line with response latencies of similar control populations, see Bergamin & Kardon [52], Koevoet et al. [40], and Strauch et al. [53]), but only responded significantly to synesthetic lightness at about 870 ms (see also Figure 3c vs e and Figure 4 for per-timepoint LME)”.

      Bergamin, O., & Kardon, R. H. (2003). Latency of the pupil light reflex: sample rate, stimulus intensity, and variation in normal subjects. Investigative Ophthalmology & Visual Science, 44(4), 1546-1554.

      Koevoet, D., Naber, M., Strauch, C. & Van der Stigchel, S. Presaccadic Attention Shifts Up-and Downwards: Evidence From the Pupil Light Response. Psychophysiology 62, e70047 (2025).

      Strauch, C., Romein, C., Naber, M., Van der Stigchel, S., & Ten Brink, A. F. (2022). The orienting response drives pseudoneglect—Evidence from an objective pupillometric method. Cortex, 151, 259-271.

      (4) Another concern is with the visualisation of the pupil traces in Figure 3 (main results); these were heavily pre-processed (per-participant demeaned), losing any feature besides the effect of interest and generating the unrealistic expectation that perception of dark/bright colors generate a net dilation/constriction of the pupil - whereas perception-related modulations of pupil size are always relative and generally small compared to the numerous other effects registered in pupil size. It would be far better to see the actual profiles, preserving the unfolding of dilations and constrictions over time, especially since these are further analysed in Figures 4 and 5.

      Indeed, the expectation that any dark synesthetic experience would lead to pupil dilation whereas any bright synesthetic experience would lead to constriction is not warranted – it would only do that relative to the counterfactual of not having that experience.

      Many factors affect the pupillary signal at the same time, and often differently across individuals (think of tiredness etc.), making merely baseline corrected traces seemingly noisy. Our visualization highlights that there is a systematic part to that variation that lies in the synesthetic brightness experience.

      Visualizing the effects of idiosyncratic experiences, varying within and between participants is challenging. For the theoretical insight brought about through our paper in Figure 4 (synesthesia being sensory in nature), demeaning is favorable in our opinion as it isolates the effect of interest in visualization. However, for methodological reasons and to better show effect sizes etc., there is certainly use in additional transparency. We now thus provide non-demeaned traces in the supplementary material as the reviewer suggested and also refer to these in the main manuscript. Furthermore, all figures are now provided in millimeters, with all pupil related analysis being rerun and updated to this end (without qualitative changes to the results). This should further rectify possibly inflated expectations about the absolute size of effects and allows to put effects into perspective across studies. We now added:

      “Pupillary data were transformed from arbitrary eyelink units to millimeters using a conversion factor obtained with an artificial eye (see Hayes & Petrov, 2016).”

      Hayes, T. R., & Petrov, A. A. (2016). Mapping and correcting the influence of gaze position on pupil size measurements. Behavior research methods, 48(2), 510-527.

      Impact:

      Despite these weaknesses, and especially if they are adequately addressed in the review, this work is likely to improve our understanding of synesthesia, providing a new tool to quantify the subjective sensations; an interesting potential extension would be using pupillometry for tracking changes over time of the synesthetic experiences, opening up the possibility to evaluate the importance of learning for this peculiar experience.

      We were happy to read our manuscript was evaluated this positively and hope that our replies can address the remaining smaller concerns and make findings more transparent to the readers.

      Reviewer #2 (Public review):

      Synesthesia is a neurological condition where stimulation of one sensory channel leads to involuntary, automatic, and consistent experience of another, unrelated percept. For example, Sir Francis Galton (1880, Nature) famously described the robust tendency of some individuals (synesthetes) to associate numerals with a distinct color. Ever since, synesthesia has continued to attract a broad interest in the cognitive neurosciences in light of its implications for the study of domains such as perception, consciousness, and brain connectivity, among others.

      Strauch, Leenaars, and Rouw measured pupil size in a group of 16 grapheme-color synesthetes and two matched control groups. The participants were presented with gray digits - that is, visual stimuli having identical physical properties in terms of brightness. Each participant subsequently rated the corresponding evoked color and brightness: unlike controls, synesthetes did so in a very consistent and reliable fashion. Accordingly, this was also shown in their pupils: despite the same objective luminance, digits associated with brighter percepts caused their pupils to constrict, and digits associated with darker percepts caused their pupils to dilate more than controls. These results highlight how crossmodal correspondences are deeply rooted in synesthetes, and put forward pupillometry as a particularly appealing biomarker for some phenomenological experience (at least those grounded in "brightness").

      Further strengths of the technique are its temporal resolution and its responsiveness to several constructs. Across several tasks, the authors show, for example, that responses to synesthetic light are somewhat slower than responses to real light (i.e., they are likely mediated), but at the same time faster than responses to mental imagery. The role of mental imagery can also be reasonably dismissed when considering the second feature of pupil size: its responsiveness to mental effort and cognitive load. The pupils tend to dilate with demanding, challenging tasks, and this was the case when control participants were asked to report the color of a digit for which they did not consistently experience a synesthetic association. The same task was, instead, seemingly effortless for synesthetes, again speaking in favor of the automaticity of number-color correspondences in their case.

      Overall, the findings by Strauch, Leenaars, and Rouw are highly significant for the field and likely to be impactful. The strength of their evidence, when accounting for the relatively small sample size and the inherent variability of both phenomenology (color perception and subjective reporting) and physiology (pupil size), is adequate and sufficiently convincing.

      We were glad to read this overall very positive assessment of our work and thank the reviewer for the additional non-public suggestions for improvements.

      Reviewer #3 (Public review):

      Summary:

      In the present study, the authors examined pupillary responses to uncolored stimuli (number graphemes) among number-color synesthetes and non-synesthetes. After seeing a digit, the synesthetes and active control participants were asked to indicate which color they perceived using three dimensions of hue, saturation, and lightness. The lightness values were the primary independent variable for follow-up analyses. To see how the pupil responded to psychologically "bright" and "dark" digits, the authors split the reported lightness values at the median and plotted them. The synesthetes showed a pupillary constriction to digits they perceived as bright and dilation to digits they perceived as dark. Active control participants did not show that effect. In a subsequent block, only the synesthetes were shown the colors they reported perceiving as colored discs. Their pupillary responses were similar. The authors also found that the differences in pupillary responses between light and dark perceptions (with digits) were only slightly delayed in their onset to the perception of a colored disc, and therefore, the color perception accompanying a digit is unlikely to be effortful or a retrieved association, but occurs rather automatically.

      Strengths:

      The authors employed a well-controlled and designed quasi-experiment comparing colorgrapheme synesthetes to non-synesthetes and showed convincingly that the color perceptions accompanying graphemes alter the physical perception of brightness. They also made a reasoned attempt to rule out the possibility that color associations are occurring effortfully via retrieved associations.

      We appreciate the positive assessment and useful suggestions for revision.

      Weaknesses:

      There are some areas in which the implications of these findings could be elaborated upon. I had the following questions:

      (1) Are the pupillary responses among synesthetes, which objectively do not seem to match the degree of physical stimulation entering the retina, in any way maladaptive for eye functioning? I understand the constriction/dilation of the pupil to not only benefit visual acuity but also to protect the retina from damage. Are synesthetes at any risk of retinal damage due to over-dilation of the pupil to brighter stimuli? Or are these effects of a magnitude that is too small to matter? As reported in arbitrary units, it was hard to know how large these effects were in terms of measurable changes in dilation (e.g., millimeters).

      This is an interesting point. Some argue that pupil size changes in a mid-range mildly affect optics thus affecting detection performance, contrast perception, and depth of field (Eberhardt et al., 2022, Mathôt & Ivanov 2019, Ruuskanen, Boehler, & Mathôt, 2025), rather than serving a protective role for the retina (Mathôt, 2018). Indeed, any effects reported here were quite small. We agree with the reviewer that this can be made more accessible by reporting effects in millimeters. We thus now adjusted all figures accordingly and write in the methods section:

      “Pupillary data were transformed from arbitrary eyelink units to millimeters using a conversion factor obtained with an artificial eye (see Hayes & Petrov, 2016).”

      Note that even the largest effects here (those elicited by physical luminance change in block 2 for the synesthetes) only caused differences in pupil size of about 0.3mm. This lies below the maximal pupil dilations observable in response maximal effort (about 0.5mm), for instance, and substantially below the full range of pupil size changes elicited through strong luminance stimulation (several millimeters). We therefore deem the changes in pupil size as obtained in our study too minor to be practically maladaptive for optics/perception.

      Eberhardt, L. V., Strauch, C., Hartmann, T. S., & Huckauf, A. (2022). Increasing pupil size is associated with improved detection performance in the periphery. Attention, perception, & psychophysics, 84(1), 138-149.

      Hayes, T. R., & Petrov, A. A. (2016). Mapping and correcting the influence of gaze position on pupil size measurements. Behavior research methods, 48(2), 510-527.

      Mathôt, S., & Ivanov, Y. (2019). The effect of pupil size and peripheral brightness on detection and discrimination performance. PeerJ, 7, e8220.

      Mathôt, S. (2018). Pupillometry: Psychology, physiology, and function. Journal of cognition, 1(1), 16.

      Ruuskanen, V., Boehler, C. N., & Mathôt, S. (2025). The Interplay of Spontaneous Pupil-Size Fluctuations and EEG Power in Near-Threshold Detection. Psychophysiology, 62(3), e70035.

      (2) Likewise, is the automatic synesthetic merging of two percepts something that could be learned such that natural synesthetes and "artificial" synesthetes would look similar? For example, if a group of non-synesthetic participants were to learn a color-grapheme association to automaticity, would you expect their pupillary responses to the graphemes look similar to the synesthetes'? If so (or if not), what would this tell us anything about the phenomenology of synesthesia?

      We find this question most interesting. Likely, different synesthesia researchers wouldn’t even fully agree on the most plausible answers to these questions. Training studies have shown that nonsynesthetes can be trained to associate particular colors to particular graphemes, as revealed in the synesthetic Stroop effect: interference effects of the learned color onto reporting the typeface color of the grapheme. The degree to which non-synesthetes can be trained to become similar to synesthetes is however still topic of debate.

      We now discuss as follows:

      “Future studies could examine to what degree training a non-synesthete to associate specific colors to particular inducers (e.g., digits), can provide similar patterns of results as genuine synesthesia (Bor et al., 2014, Colizoli et al., 2012, Rothen & Meier, 2014). Could learning produce similar brightness-related pupil effects in non-synesthetes? Similarly, would effort-linked responses diminish with increased training duration? The perhaps most interesting question relates to response latencies: Would a trained participant ever be able to produce brightnessrelated pupil effects as fast as a synesthete?”

      Bor, D., Rothen, N., Schwartzman, D. J., Clayton, S., & Seth, A. K. (2014). Adults can be trained to acquire synesthetic experiences. Scientific reports, 4(1), 7089.

      Colizoli, O., Murre, J. M., & Rouw, R. (2012). Pseudo-synesthesia through reading books with colored letters. PloS one, 7(6), e39799.

      Rothen, N., & Meier, B. (2014). Acquiring synaesthesia: insights from training studies. Frontiers in human neuroscience, 8, 109.

      (3) Do the synesthetic perceptions of digit graphemes merge in a sensible way? For example, if a synesthete sees a particular color with the digit 1, and a different color with the digit 9, what do they perceive when they see 19? or 1-9, or 1 9? Is there color blending, or an altogether different color perception?

      This is a very interesting question indeed. While each synesthete will have their own specific expression of synesthesia, there are regularities in how a combination of digits evokes synesthetic color. First, if asked about the color of a specific digit, each digit keeps its own color, as the color of a digit is linked to the identity of the digit (Dixon et al., 2006). Context effects are however possible, in particular when context alters the interpretation of the digit (Myles et al., 2003). A particularly common context in a multi-digit number is a dominant first digit, spreading its color to the subsequent digits in the number. However, as the digit color is linked to digit identity, what does ‘not’ happen is a mixing of colors into a qualitatively new color; for example, a yellow "1" and blue "9" do not merge into a green "19".

      Dixon, M. J., Smilek, D., Duffy, P. L., Zanna, M. P., & Merikle, P. M. (2006). The role of meaning in grapheme-colour synaesthesia. Cortex, 42(2), 243-252.

      Myles, K. M., Dixon, M. J., Smilek, D., & Merikle, P. M. (2003). Seeing double: The role of meaning in alphanumeric-colour synaesthesia. Brain and Cognition, 53(2), 342-345.

      Many thanks for the constructive assessment of our work.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) I am not sure I'd use the term 'cross-modal' given that the case considered here (graphemecolor) is purely visual.

      The reviewer is absolutely right: the term 'cross-modal' has a historical background rather than reflecting an exact factual accuracy. The term is still commonly used however, as it readily reflects how the induced additional experience is always of a different (sub)type than the inducing experience. There is a cross-over between experiences that might occur within the same sensory modality, or even induce awareness of a particular concept. But key to synesthesia is the crossover experience as the inducer and concurrent are different (sub)types of experiences. For example, seeing a letter can evoke a synesthetic experience of seeing a color, or evoke awareness of a particular gender or personality of that letter, but does not evoke another letter. To remain consistent with literature, we refer to 'cross-modality' when explaining the link to previous literature, but generally switched to using 'cross-over experience':

      “Therefore, synesthesia might provide a unique window into how the brain’s constructive processes can generate additional, conscious content, in cross-over experiences, often across modalities, going all the way down to the level of sensory phenomenology.”

      We adjusted throughout the manuscript accordingly.

      (2) I would not recommend focusing the introduction on the problem of qualia; this is a much more general and complex question than the one addressed in the study; the space of the introduction may be better used to present the actual object of study, giving a better picture of the synesthetic phenomenon and of previous work aimed at characterising it (behavioural, including PA scores and consistency measures, and neuroimaging). It is important to discuss how the pupillometric approach differs from the previously adopted neuroimaging techniques and what it can add to those.

      We agree that qualia is a very general and complex question. However, we respectfully disagree that this complex question is not the object of the study. What is remarkable about synesthesia is not the presence of an additional perceptual association per se, but the presence of a specific perceptual experience. As illustration, think of a test where an unconscious color association to the word 'banana' was tested. While a generic 'yellow' could semantically be linked and would likely be obtained in the (e.g. priming) experimental results, a follow-up question of picking on a color wheel the exact shade of yellow to this association, or describing the perceptual sensation of the color, would be non-sensical to the participants.

      This sharply contrasts with the current study: synesthetes, but not non-synesthetes, indicate a perceptual sensation of additional colors, and subsequently indeed the sensory properties of this percept (experienced brightness) affects the objective reflection of this sensation (pupil size) in synesthetes but not in non-synesthetes. In our view, the presence of additional qualia is key in understanding what sets synesthetic apart from non-synesthete associations, including so-called cross-modal correspondences (unconscious consistent associations across modalities, common to us all). We even believe that the reported qualia is what makes synesthesia so interesting in the first place. We now more clearly explain this link to qualia better in the introduction.

      "The most remarkable aspect of synesthesia is the subjective perceptual phenomenology of the induced colors, setting these sensations apart from color memory, thought, or amodal association. The contrast between synesthetes and non-synesthetes can thus offer an interesting doorway into examining qualia, the subjective perceptual phenomenology or first person (what's-it-like) perspective."

      We also improved the explanation of the synesthetic phenomenon, including a more detailed characterisation of behavioural measures (including consistency scores) and added neuroimaging studies. These changes have been incorporated into the text in response to previous comments (point 1- reviewer 1).

      Please note that we have chosen not to include more detailed discussion of PA scores. Our results show a trend but do not allow for a conclusive interpretation on PA scores, and we feel that placing greater emphasis on this topic might therefore be confusing or even misleading. Still, it would be a very interesting topic for follow-up research to examine how alterations in characteristics of the synesthetic experience influence pupil responses.

      The different synesthesia types all share the defining characteristics of an additional conscious and consistent experience. Synesthetes can verbally report their additional experience, and synesthetic sensations can be measured in behavioral paradigms such as the ’synesthetic Stroop’ effect, or brain activation patterns in sensory cortex [15]. Furthermore, test-retest paradigms show how synesthetic, but not non-synesthetic associations are highly specific and consistent [16-18]. Thus, over the past decades, research has established synesthesia as a ’real’ condition that can reliably be identified using behavior, neurophysiology, and neuroimaging [11, 13, 15–21]. The most remarkable aspect of synesthesia is the subjective perceptual phenomenology of the induced additional sensation, i.e., color in grapheme-color synesthesia. This sets synesthetic sensations apart from (color) memory, thought, or amodal association. Synesthesia can thus offer an interesting doorway into examining qualia, the subjective perceptual phenomenology or first person (what’s-it-like) perspective.

      We now discuss the pupillometric approach as it differs from the previously adopted neuroimaging techniques as follows:

      “Compared to neuroimaging studies [12,15,51], pupillometry may offer a more direct window into synesthetic phenomenology, as the directionality between pupil light reflex and perceived brightness is straightforward. Finally, improved understanding of the underlying processes can be obtained by contrasting responses to perceived versus actual (physical) brightness, given that the pupil light reflex is a well-characterised reflex arc involving few inferential steps.

      This adds to the explanation that was already present on how the current approach differs from previous techniques, and what it can add to those techniques:

      "Instead, current paradigms capturing synesthesia employ objective measures, but fail to capture its phenomenology [16, 17, 21, 23]."

      (3) There are a few typos and word repetitions.

      Many thanks – we identified typos and repetitions after another set of careful reads and hope to have eradicated them completely now.

      Reviewer #2 (Recommendations for the authors):

      I am overall very supportive of this work, but addressing the following points may enrich it further:

      (1) Paragraph 2.2.1. Here, models do not seem to compare synesthetes versus controls but rather assess the effects of interest separately in the two groups. The fact that experimental effects are significant in synesthetes, but not in controls, does not tell us much about differences between groups. Controls (e.g., Figure 3) do show a similar trend, albeit clearly smaller. There is one passage in which this issue appears to be tackled (page 10): "Critically, in an LME ran on synesthetes and controls and using only graphemes and the interaction of group and lightness as predictors, we found lightness to predict pupil size in synesthetes (t = -2.754, p = 0.006), but not controls (t = -1.134, p = 0.257)." But I am not sure that the reported statistics belong to the interaction - they seem to refer to the lightness effect within each group, not the difference.

      This is an important point, power for between-group comparisons is inherently limited for n = 16 per group (while still feasible for overall responses, things become trickier when less trials remain). A simple model of pupil ~ grapheme + group * lightness_scaled + (1 | participant) shows no significant interaction (despite one group showing the effect and the other not showing the effect significantly). The additional negative effect for group is in line with the effort-related effect reported later in the manuscript. Where does this leave us? Based on the lightness responses alone, the group difference can be characterized as a quantitative distinction, but the degree in which it is also a qualitative distinction cannot clearly be determined from current data. We revised the manuscript to make sure that such an interaction is not implied/ point to the absence of the significance of that interaction.

      The sensory nature of synesthetic color is supported by within-synesthete analyses, where coupling strength parametrically modulates the lightness-pupil relationship in a theoretically predicted manner. Importantly, the effort-related findings provide a complementary and statistically robust group comparison: synesthetes and controls performing the identical colorreporting task showed significantly different pupil dilation rates, directly demonstrating that the two groups differ in how they access color information. Together, these two independent pupillometric signatures, one tracking perceptual quality, one tracking effort, converge on the same conclusion and mutually reinforce the interpretation that synesthetic color constitutes genuine sensory phenomenology.

      Author response image 3.

      We now make this more explicit in the manuscript as follows:

      “We found significant modulations of pupil size by the lightness of the grapheme's synesthetic color - sustained and in the to-be-expected time window. Specifically, the pupil constricted more for brighter reported colors, and dilated more for darker reported colors, as predicted (Average pupil size 800-4000ms, t = -3.601, p < 0.001). In an LME ran for synesthetes and controls and using only graphemes and lightness as predictors, we found lightness to predict pupil size in synesthetes (t = 2.844, p = 0.004), but not controls (t = 0.606, p = 0.544). However, when taking group as interacting factor in a joint LME, there was no interaction of lightness and group (t = -0.949 p = 0.342).”

      and

      “For controls a separate model was run, now without the PA score as predictor (not assessed for controls). Neither lightness (t = -0.815, p = 0.415), coupling strength (t = 0.438, p = 0.661), nor their interaction gained significance (t = -1.058, p = 0.290; all for average pupil size between 800 ms and 4000 ms). Critically, we also ran a LME with the three-way interaction of coupling strength, group, and lightness (Wilkinson notation: pupil = grapheme + group + lightness * group + coupling strength * lightness * group + (1 | participant)). This analysis revealed a significant three-way interaction between lightness, coupling strength, and group (F = 3.86, p = .021), indicating that the lightness × coupling strength effect on pupil size was not equivalent across groups. Decomposing this interaction by group, the lightness × coupling strength slope was significant in synesthetes (t = 2.59, p = .010) but not in controls (t=-1.01, p=.311), suggesting that reported lightness and its coupling strength were more consistently related to pupil size in synesthetes than in controls. Note however, that this decomposition does not directly test whether the two slopes significantly differ from each other, however. Lastly, pupil size was marginally larger in controls than in synesthetes (t = 1.94, p = .062; see later sections for more in-depth analyses)”

      (2) The authors choose to analyze pupil size in arbitrary eye tracker units. This is fine, although I would recommend assessing and reporting whether the average pupil size (e.g., during the baseline) is roughly comparable between groups. The size of the effects may be difficult to compare between groups in the presence of very different baseline pupil size.

      Please see Author response image 4 for Baseline pupil sizes per group in millimeters. There were no differences between groups.

      Author response image 4.

      F2, 45) = 0.707, p = 0.499 (One-way Anova).

      We now write:

      “Baseline pupil sizes did not differ between groups (F(2, 45) = 0.707, p = 0.499).”

      We agree with the reviewer that millimeters are a more intuitive measure and updated all figures throughout manuscript and supplementary materials accordingly. We also briefly added to signal processing that this conversion was applied.

      “Pupillary data were transformed from arbitrary eyelink units to millimeters using a conversion factor obtained with an artificial eye (see Hayes & Petrov, 2016).”

      Hayes, T. R., & Petrov, A. A. (2016). Mapping and correcting the influence of gaze position on pupil size measurements. Behavior research methods, 48(2), 510-527.

      (3) If I understand correctly, the main task counted 120 trials overall (12 per digit). It seems, however, that only 3 and 4 participants remained with at least 50 trials (or 25 per median split by lightness) after preprocessing. This appears to be quite a massive data loss: is there a reason behind it? Please also clarify: the overall percentage of discarded trials; whether the median split by lightness was computed on all responses or only on those of the remaining, valid trials.

      This is an important point for clarification indeed. The exclusion of participants in Figure 3 applies only to that particular visualization, not to the statistical analyses. The linear mixed effects models (LMEs) used all available valid trials from all participants, with no participant-level exclusions. The figure-specific threshold (≥25 trials per median-split bin) was applied purely for display clarity, as plotting participants with very few trials per bin would produce unreliable/noisy and thus visually misleading traces (as we note in the figure caption and point readers to Supplementary Figure 1, which shows the same visualization without any exclusions).

      Since the paradigm required participants to repeat discarded trials until 120 valid trials were collected, all participants thus contributed exactly 120 valid trials to the analyses. There was therefore no data loss at the analysis level for the LME that is central to the claims of the manuscript (albeit more complex to grasp than the t-tests between bins).

      Why were there sometimes so little trials per brightness bin?

      First, participants differed in how dark or bright (synesthetic or forced-report) colors were overall, meaning that differing proportions thereof would fall above or below the 0.5 cutoff that overall, well represented the sample (but not necessarily every single participant). Note that this median split was not performed per individual but across all color reports to allow an apples-to-apples comparison.

      Second, participants often reported colors that differed in Hue and Saturation, but not Lightness. This is in line with synesthetes picking certain colors more often than others, as compared with non-synesthetes (Rouw & Root, 2019; Ward et al., 2025).

      We now include a new Supplementary Figure that visualizes responses on the Hue and Saturation dimensions of HSL space for both synesthetes and controls; fully saturated reports appear on the outer edge. We refer to the supplementary figure in the caption of Figure 2 as follows:

      "See Supplementary Figure 1 for color reports on the hue and saturation axes.”

      Rouw, R., & Root, N. B. (2019). Distinct colours in the ‘synaesthetic colour palette’. Philosophical Transactions of the Royal Society B: Biological Sciences, 374(1787).

      Ward, J., Maciel, S., Rouw, R., Simner, J., & Root, N. (2025). Synaesthesia is linked to differences in music preference and musical sophistication and a distinctive pattern of sound-color associations. Psychology of Music, 53(3), 453-473.

      Minor points:

      (1) "Building on this evidence, we hypothesized that the cross modal color phenomenology in synesthesia can, if truly sensory in nature, could likewise be (...)" -> may need rephrasing (can/could).

      Many thanks, fixed.

      (2) Caption of Figure 1: "Block 2 (synesthetes only): a colored disk and gray central patch, matching the average indicated color per digit, and the number and luminance of pixels of said digit were presented to assess externally triggered light responses." -> I find this sentence a bit hard to follow; perhaps consider rephrasing it.

      Agreed, we rephrased to:

      Block 2 (synesthetes only): a colored disk was presented, colored according to the synesthete's average indicated color for that digit. At its center sat a gray patch matching the luminance and pixel area of the original digit from Block 1, together allowing assessment of externally triggered light responses.

      (3) Figure 2 b: Consider truncating the y-axis to 1 if that improves the visualization.

      We adjusted the axis accordingly and added a bit more detail in the caption for the interpretation of the measure.

      (4) Caption of Figure 3 points to "see Supplementary Figure 1", but it should probably be SF2.

      Many thanks for spotting, all references to supplementary figures have been checked and are corrected now.

      Elvio Blini

      Reviewer #3 (Recommendations for the authors):

      (1) As a minor comment, there are some terms that felt overused in the manuscript. For example, the words "extraordinary" and "exceptional" were used multiple times throughout. I believe I understand the authors to mean them in their descriptive sense (i.e., outside the realm of typical experience), but in context, those words make it seem like they are touting their own experiment as "exceptional" or "extraordinary," which I don't believe was their intention.

      We agree. We removed words such as exceptional and extraordinary when they do not directly refer to the sensation throughout the manuscript (which is indeed how we intended to use it). We hope that this removes unnecessary and convoluting hyperbole.

      (2) It seemed counterintuitive to me that the color consistency score would be reverse-coded. In this case, the scores actually seem to indicate inconsistency, rather than consistency. Perhaps the raw scores can be inverted for a more intuitive interpretation that aligns with the terminology. I understand that they were following a previous publication in their method (Rothen et al., 2013).

      This manner of coding is counter-intuitive indeed. However, there are both logical and practical reasons to this approach. Importantly, this is indeed the standard way of reporting color consistency in synesthesia research (Carmichael et al., 2015; Eagleman et al., 2007; Root et al., 2025; Rothen et al., 2013). The calculation is based on a simple logic; a higher number reflects a larger distance in color space. An additional advantage is the clear and intuitive zero- reference: a score of zero implies choosing the exact same color. Finally, it intuitively reflects the distinction between synesthetes and non-synesthetes; there is by definition little variation across synesthetes (visualized at the bottom of the graph), then a 'cut-off line' (if consistency is used as diagnostic tool), and then the height of the range shows how large the range in consistency is, in that particular sample of non-synesthetes. In a way we therefore inherit a confusing definition/standard, but changing it would lead to new confusion instead. We now specifically clarify this in the caption as follows:

      “Note that higher consistency is reflected in lower color distance, hence lower values [17].”

      Carmichael, D.A., Down, M.P., Shillcock, R.C., Eagleman, D.M., Simner, J., 2015. Validating a standardised test battery for synesthesia: does the synesthesia battery reliably detect synesthesia? Conscious. Cogn. 33, 375–385

      Eagleman, D.M., Kagan, A.D., Nelson, S.S., Sagaram, D., Sarma, A.K., 2007. A standardized test battery for the study of synesthesia. J. Neurosci. Methods 159 (1), 139–145.

      Root, N., Chkhaidze, A., Melero, H., Sidoro -Dorso, A., Volberg, G., Zhang, Y., & Rouw, R. (2025). How “diagnostic” criteria interact to shape synesthetic behavior: The role of self-report and test–retest consistency in synesthesia research. Consciousness and Cognition, 129, 103819.

      Rothen, N., Seth, A.K., Witzel, C., Ward, J., 2013. Diagnosing synaesthesia with online colour pickers: maximising sensitivity and specificity. J. Neurosci. Methods 215 (1), 156–160.

    1. eLife Assessment

      This study presents a large, systematically curated catalog of non-canonical open reading frames (ncORFs) in human and mouse through the reanalysis of nearly 400 Ribo-seq datasets using a standardized pipeline; the resulting atlas consolidates ncORF annotations across tissues and provides a valuable resource for investigating non-canonical translation and ORF emergence. The main conclusions are supported by consistent data processing and multiple computational measures of translation and conservation. While the pipeline is transparent and technically robust, some analytical criteria and dataset limitations could be described more explicitly, and several downstream conclusions would benefit from more cautious interpretation, some evolutionary inferences are primarily correlative; dataset heterogeneity, uneven tissue representation, and limited experimental validation also constrain the strength of a subset of the findings. Overall, the evidence is solid, and the resource is likely to be broadly beneficial to the community.

    2. Reviewer #1 (Public review):

      This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

      Strengths:

      In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

      Weaknesses:

      Based on the ncORF catalog, some of the analyses were not properly done. Some of the results are descriptive.

      (1) Bias and representations of data source. Public ribo-seq datasets are unevenly distributed across tissues and cell lines, raising concerns about heterogeneity and underrepresentation of certain contexts. This may limit the generalizability of the catalog.

      (2) The discussion on modular domains of ncORFs is unclear, and the claim that they may originate via TE-related mechanisms is not well supported. Stronger evidence or clearer reasoning is needed.

      (3) The conservation comparisons are not fully convincing. Figure S7 shows only mild differences between ncORFs and CDS, and statistical significance is not clearly demonstrated. Comparisons with other non-coding RNAs should be added, and overlapping sequences between ncORFs and CDS should be excluded to avoid bias.

      (4) Figure 3 indicates that some ncORFs are subject to evolutionary constraints. This is not surprising. The authors should provide further analyses on more detailed features of these "conserved" ncORFs vs. the "non-conserved" ones. Some pretty informative works have been done in drosophila, worms, mouse, and human. Figure 3 suggests some ncORFs are under evolutionary constraint, but this is not unexpected. More granular analyses contrasting "conserved" versus "non-conserved" ncORFs would be informative. In fact, small ORFs, especially uORFs, have been extensively studied, for their functions and corss-species conservations. The authors should explicitly show what is new here in their analyses.

      (5) Translation levels are reported using RPF counts. However, translation efficiency (normalized by RNA expression) is a more appropriate measure to account for expression heterogeneity.

      (6) The correlation analyses between ncORF translation levels and PhyloCSF are confusing and largely descriptive. These sections need sharper framing and clearer conclusions.

      (7) Public ribo-seq datasets, generated by different research labs, are known for their strong batch effects. Representations of tissues and cells are also very unbalanced. Therefore, the co-translation analysis between ncORFs and canonical CDS is not well controlled. This should be done by referring to a recent large-scale ribo-seq meta-analysis (Nat Biotechnol. 2025. doi: 10.1038/s41587-025-02718-5).

      Comments on revisions:

      The authors have made efforts to address most of the previous concerns, and several points have been clarified or improved in the revision. However, in a number of cases, the responses rely more on acknowledgment and reframing rather than substantive analytical strengthening. Overall, the manuscript is improved, particularly in terms of clarity, transparency, and positioning of claims. I support its publication and look forward to seeing how the field engages with and discusses these claims.

    3. Reviewer #2 (Public review):

      Summary:

      Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

      Strengths:

      (1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

      (2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

      Weaknesses:

      (1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

      (2) Some analytical methods and standards were not clearly presented in the manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work compiles a comprehensive atlas of ncORFs across mammalian tissues and cell types, derived from reanalysis of ~400 public ribosome profiling datasets. The authors then evaluate cross-species conservation and functional signatures, proposing that evolutionarily ancient ncORFs tend to have higher translation potential, stronger expression, and closer relationships with canonical coding sequences.

      Strengths:

      In general, the study provides a large-scale and timely resource of annotated ncORFs, which could be broadly useful for the community. The authors collected ~400 public ribosome profiling datasets for annotations of ncORFs, which, to my best knowledge, is the largest collection of data for such a purpose. The catalog could facilitate future investigations into ncORF biology and broaden understanding of the coding potential of the "non-coding" genome.

      We thank the reviewer for the positive evaluation of our manuscript and for recognizing the significance of our contribution.

      Weaknesses:

      Based on the ncORF catalog, some of the analyses were not properly done. Some of the results are descriptive.

      (1) Bias and representations of the data source. Public ribo-seq datasets are unevenly distributed across tissues and cell lines, raising concerns about heterogeneity and underrepresentation of certain contexts. This may limit the generalizability of the catalog.

      We agree with the reviewer that the uneven distribution of public Ribo-seq datasets across tissues can inevitably introduce bias in the ncORF composition of our catalog. This bias is likely more pronounced in humans due to the narrower tissue coverage. We have addressed this point in the Discussion section of the revised manuscript.

      (2) The discussion on modular domains of ncORFs is unclear, and the claim that they may originate via TErelated mechanisms is not well supported. Stronger evidence or clearer reasoning is needed.

      We thank the reviewer for highlighting this point. We have revised the manuscript to more clearly explain the rationale behind our analysis of ncORF modular domains and have adopted more cautious language regarding their potential transposable element–related origins, limiting interpretations to what is directly supported by the data.

      (3) The conservation comparisons are not fully convincing. Figure S7 shows only mild differences between ncORFs and CDS, and statistical significance is not clearly demonstrated.

      Comparisons with other non-coding RNAs should be added, and overlapping sequences between ncORFs and CDS should be excluded to avoid bias.

      We thank the reviewer for this comment and apologize for the lack of clarity in the original figure. Both CDSs and ncORFs show significant deviation from zero Gnocchi scores (two-sided Wilcoxon signed-rank tests), which is now stated explicitly in the revised legend and text. CDS-overlapping ncORFs were already excluded in the original analysis; this has been clarified to avoid confusion.

      As suggested, we have added lncRNAs for comparison. ncORFs display modestly higher Gnocchi scores than lncRNAs, and this difference persists when restricting the analysis to lncRNA-derived ncORFs and their corresponding full-length lncRNAs (see revised Fig. S7). These additions strengthen the conservation comparison while controlling for transcript context.

      (4) Figure 3 indicates that some ncORFs are subject to evolutionary constraints. This is not surprising. The authors should provide further analyses on more detailed features of these "conserved" ncORFs vs. the "non-conserved" ones. Some pretty informative works have been done in Drosophila, worms, mice, and humans. Figure 3 suggests some ncORFs are under evolutionary constraint, but this is not unexpected. More granular analyses contrasting "conserved" versus "non-conserved" ncORFs would be informative. In fact, small ORFs, especially uORFs, have been extensively studied for their functions and cross-species conservation. The authors should explicitly show what is new here in their analyses.

      We thank the reviewer for this insightful comment. We agree that cross-species conservation of ncORFs (particularly uORFs) has been extensively investigated in prior studies, including our own.

      However, most prior analyses have focused on conservation of start codons or overall ORF integrity, which does not distinguish selection acting on translational activity from selection acting on the encoded peptide sequence itself. In contrast, our analysis leverages codon-level periodic PhyloP signals across the full ORF. The observed three-nucleotide periodicity is consistent with selective constraint at the amino acid level, rather than merely preservation of initiation sites or translational potential. Furthermore, our newly developed branch-length statistic uncovers lineage-restricted conservation patterns among ncORFs, enabling resolution of evolutionary dynamics not captured by conventional conservation metrics.

      Thus, while the existence of conserved ncORFs is not unexpected, the conceptual advance of our study lies in demonstrating that a subset exhibits coding-like evolutionary constraint consistent with selection on their peptide products, as well as revealing lineage-specific conservation patterns. We have clarified this distinction in the revised Discussion.

      (5) Translation levels are reported using RPF counts. However, translation efficiency (normalized by RNA expression) is a more appropriate measure to account for expression heterogeneity.

      We agree that translation efficiency (TE), which normalizes ribosome footprint counts by RNA abundance, is in principle an appropriate metric. We initially calculated TE and compared ncORFs with CDSs. However, we found that TE estimates for short ncORFs were substantially inflated by RPF enrichment near start and stop codons, leading to unstable and potentially misleading values.

      For CDSs, this bias is commonly addressed by excluding the first and last 10 to 20 codons when quantifying RPF density. This strategy is not feasible for ncORFs because of their short length. We therefore used RPF counts in the final analysis, applying stringent positional filtering. Only RPFs whose P sites fall within the ORF body, excluding start and stop codons, were counted. RPFs overlapping the ORF but with P sites outside the annotated frame, likely derived from adjacent ORFs or initiation or termination pausing, were excluded.

      TE and RPF counts both measure translation but capture different aspects. TE reflects ribosome density relative to transcript abundance, whereas RPF counts quantify overall ribosome engagement. Given the short lengths of ncORFs, count-based quantification provides a more robust and conservative estimate of their translational activity.

      (6) The correlation analyses between ncORF translation levels and PhyloCSF are confusing and largely descriptive. These sections need sharper framing and clearer conclusions.

      We thank the reviewer for this comment. We agree that the original presentation lacked clear framing. The relationship between PhyloCSF scores and mean ncORF translation levels across tissues is influenced by both evolutionary age and tissue specificity. Older ncORFs with higher coding potential tend to exhibit stronger tissue-restricted expression. As a result, their mean translation levels across all tissues appear lower, not because they are weakly translated, but because their translation is concentrated in specific tissues. This point is addressed in the revised manuscript.

      (7) Public ribo-seq datasets, generated by different research labs, are known for their strong batch effects. Representations of tissues and cells are also very unbalanced. Therefore, the co-translation analysis between ncORFs and canonical CDS is not well controlled. This should be done by referring to a recent large-scale ribo-seq meta-analysis (Nat Biotechnol. 2025. doi: 10.1038/s41587-025-02718-5).

      We thank the reviewer for highlighting this important study and for raising concerns regarding batch effects and tissue imbalance in public Ribo-seq datasets. We are aware that public Ribo-seq data generated by different laboratories are subject to substantial batch effects. During the ncORF annotation phase, we applied stringent quality-control criteria to minimize technical variability. For the co-translation analysis, inclusion criteria were relaxed to increase tissue and cell-type coverage. To partially mitigate representation bias, libraries derived from the same tissue or cell type were merged when quantifying ORF translation levels, thereby reducing overrepresentation from heavily sampled contexts.

      Nevertheless, we acknowledge that these measures cannot completely eliminate batch effects or imbalance inherent to public datasets. We agree that co-translation analysis would benefit from uniformly processed, high-quality datasets generated under standardized protocols with balanced tissue representation, representing a valuable direction for future research.

      Reviewer #2 (Public review):

      Summary:

      Chang et al. attempted to analyze a large number of ribo-seq datasets through a standardized pipeline, identifying novel non-canonical ORFs and elucidating their evolutionary and expression characteristics.

      Strengths:

      (1) The datasets analyzed by the authors are sufficiently comprehensive, and the use of standardized pipelines ensures excellent analytical consistency.

      (2) Their analyses of ORF evolution and co-expression further deepen our understanding of these ORFs.

      We thank the reviewer for the positive evaluation of our manuscript. It is encouraging to know that the analytical framework was found to be sound and appropriate.

      Weaknesses:

      (1) The authors primarily conducted analyses through bioinformatics, lacking sufficient wet-lab experimental evidence.

      We thank the reviewer for this comment and acknowledge this limitation. We agree that functional validation through wet-lab experiments would provide important mechanistic insight into individual ncORFs. However, this study was designed as a systematic, genome-wide computational analysis to characterize translated ncORFs across species and tissues. Our objective was to define global patterns of translation, conservation, and structural features using large-scale datasets. Given the breadth and scale of these analyses, experimental validation of specific ncORFs falls beyond the scope of the current study. We have clarified this point in the dicussion and noted that our results provide a framework for future targeted experimental investigation.

      (2) Regarding the evolution of non-canonical ORFs, a considerable amount of prior work already exists. The authors need to further clarify what new insights and discoveries they have made based on the analysis of such a large dataset.

      We thank the reviewer for this suggestion. Similar concerns were also raised by Reviewer #1. In response, we have revised the Discussion to more clearly delineate the conceptual advances enabled by our large-scale dataset.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Several aspects of the downstream analyses would benefit from additional refinement. The heterogeneity and tissue imbalance inherent in public Ribo-seq datasets introduce potential biases in ncORF detection and inferences about co-translation. Given the breadth of the dataset, it would also be informative to quantify how consistently the newly identified ncORFs are detected across samples-distinguishing those observed broadly across tissues, those enriched in specific contexts, and those detected in only a few datasets. Such stratification would help differentiate reproducibly translated ORFs from candidates requiring further validation.

      We thank the editor for the helpful comments. We agree that heterogeneity and tissue imbalance in public Ribo-seq datasets can influence ncORF detection and downstream interpretations. We have added discussion of this limitation in the revised manuscript.

      Detection of ncORF translation depends not only on biological activity but also on sequencing depth and data quality. Although all ncORFs reported here were reproducibly identified by multiple methods across independent libraries, we agree that those detected in a larger number of datasets represent stronger candidates for functional validation. Accordingly, we now report the number of methods and libraries in which each ncORF was detected in the final catalog (Supplementary Table 3). Overall, 22.3–26.3% of ncORFs were detected in more than 10 libraries, whereas more than half were observed in only two to five libraries (Fig. S1B), enabling clearer stratification of broadly translated versus more context-specific candidates.

      Some evolutionary and functional interpretations are largely descriptive or consistent with established findings for small ORFs, and the authors should more clearly articulate what is novel in their analyses. The criteria separating "young," "old," and "ancient" ORFs require clearer definition, and conservation analyses would be strengthened by improved statistical rigor and explicit exclusion of regions overlapping annotated coding sequences. Evidence for modular domain features or transposable element-related origins is limited and warrants either stronger support or more cautious framing. Proteomics validation is currently minimal and could be substantially reinforced using existing public MS resources.

      We thank the reviewer for these constructive comments. In the revised manuscript, we more clearly delineate the novel insights derived from our evolutionary analyses of ncORFs, distinguishing them from established findings on small ORFs.

      We have clarified the criteria used to classify ORFs by evolutionary age in figure 6E and refined the terminology describing “young,” “old,” and “ancient” categories to ensure precise definition. The conservation analyses have been strengthened through more rigorous statistical treatment and by explicitly excluding regions overlapping annotated coding sequences.

      With respect to modular domain features and potential transposable element–related origins, we have adopted more cautious language and limited our interpretations to what is directly supported by the data. Finally, we acknowledge that current proteomic validation remains limited and have clarified this point in the manuscript while outlining the potential for future integration of large-scale public mass spectrometry datasets in Discussion.

      The authors additionally report an interesting observation that many ncORFs on mRNA co-translate with the main CDS of the same gene. Because canonical models often posit that uORF translation suppresses downstream CDS translation, further analysis would be valuable. In particular, it would be useful to determine whether patterns of co-translation differ among ORF types or evolutionary categories and to discuss possible regulatory mechanisms underlying these relationships.

      We thank the editor for this thoughtful comment. As noted in our response to Reviewer #2, uORF–CDS co-translation does not contradict the canonical model in which uORFs repress downstream CDS translation. Co-translation reflects concurrent ribosome occupancy, whereas repression concerns the fraction of initiating ribosomes that ultimately reach and translate the CDS. Following the editor’s suggestion, we further examined whether co-translation patterns differ across ORF types or evolutionary categories. We found that ncORFs co-translating with their corresponding main CDSs are predominantly uORFs. However, these uORFs do not show statistically significant differences in conservation metrics or evolutionary age compared with other non-overlapping uORFs. Thus, we did not detect clear subtype- or age-specific distinctions among co-translating ncORFs. We have clarified these analyses in the revised manuscript.

      Addressing these points would enhance the precision, interpretability, and robustness of the study's conclusions.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors developed and refined a standardized pipeline to analyze nearly 400 ribo-seq datasets, identifying over 10,000 novel non-canonical ORFs in both human and mouse samples. Given the scale of this analysis, it is intriguing to consider how many of the newly identified non-canonical ORFs are consistently detected across multiple sample types (conservatively expressed ORFs), how many are restricted to specific tissues/ or tissue-specific ORFs), and how many were detected in only a single or very few samples (ORFs requiring further validation). Providing these data could offer new insights into understanding ORF translation.

      Thanks for this constructive suggestion. This information has been presented in the revised Supplementary Table 3 and in a newly added supplementary figure (Fig. S1B), which together provide a clearer overview of ncORF detection consistency and context specificity.

      (2) The authors' validation of MS data lacks specific details in the paper. Regarding the MS-supported ORF mentioned in Lane 117, which dataset's MS data is being referenced? Or does it refer to the content in Reference 20? At present, substantial research exists in both public general proteomics studies (e.g., CPTAC) and MS investigations targeting non-canonical ORFs. We recommend the authors incorporate additional MS data or public MS-based databases to strengthen validation in this area (PMID: 34129944, 39794466, 37823596,39413795).

      We thank the reviewer for this comment and for the helpful suggestions. The MS-supported ORFs mentioned in line 117 refer to the compilation reported in Reference 20, which integrates evidence from multiple independent proteomics studies. In addition, we examined MS-supported ORFs curated by GENCODE and PeptideAtlas, which are shown in Fig. 1E.

      We agree that incorporating additional MS datasets would further strengthen validation of ncORFs. Studies cited by the reviewer and recent community efforts such as the GENCODE and PeptideAtlas analyses (PMID: 39314370) provide valuable examples in this direction. However, performing a comprehensive reanalysis of more than 95,000 public human MS runs is computationally demanding and currently infeasible for our group given resource and funding constraints.

      To our knowledge, ongoing community-wide initiatives are working toward more comprehensive catalogs of translated human ncORFs. Large-scale, exhaustive MS searches will be particularly effective once a community consensus annotation framework for ncORFs is established. We have added discussion of these limitations and future directions in the revised manuscript.

      (3) The authors classified ncORFs into three groups-"Ancient," "Young," and "Old"-based on their origin nodes. However, both the "Young" and 'Old' groups appear to be "mammalian-specific," yet the specific criteria for their division remain unclear. It is recommended to more clearly define in the figure legend or main text how "Young" and "Old" are categorized (e.g., based on specific evolutionary nodes or distance thresholds from nodes to the end) to avoid reader confusion.

      In Fig. 5, “old” and “young” were intended as qualitative descriptors of relative evolutionary age based on the position of ncORF origination nodes along the phylogeny, as indicated on the x-axis. They were not meant to represent discrete categories. To avoid confusion, we have revised the manuscript to use “older” and “younger” throughout when referring to relative age differences. A binary classification is used only in Fig. 6E, where ncORFs are grouped into ancient (pre-mammalian) and younger (mammalian-specific) categories. This distinction is clearly defined in both the main text and the corresponding figure legend.

      (4) The authors observed an intriguing phenomenon: ncORFs on mRNA tend to co-translate with the main CDS of the same gene. However, the conventional view holds that uORF translation often inhibits the translation of the main CDS. I suggest the authors could refine their analysis in this section further. For instance, do different types of ORFs or ORFs at different evolutionary levels exhibit distinct levels of cotranslation with the main CDS? Additionally, while observing this phenomenon, the authors should also propose hypotheses regarding the regulatory mechanisms involved in these processes.

      We thank the reviewer for these constructive suggestions. After excluding CDS-overlapping ORFs, we identified 258 human and 128 mouse ncORFs that co-translate with their corresponding main CDSs. With the exception of 10 human dORFs, all remaining cases were uORFs. We compared these cotranslating ncORFs with other non-overlapping uORFs and dORFs but did not detect statistically significant differences in evolutionary age and conservation metrics. Because no clear distinguishing features emerged, we did not include these results in the manuscript.

      Importantly, the observation of uORF–CDS co-translation does not contradict the established repressive role of uORFs. Co-translation reflects concurrent ribosome occupancy, whereas repression concerns the proportion of initiating ribosomes that ultimately translate the CDS. For example, if two ribosomes initiate within a given interval and one translates the uORF while one translates the CDS, CDS output is reduced by 50% relative to a uORF-free transcript. If four ribosomes initiate under the same repressive regime, two may translate the uORF and two the CDS. In this case, absolute translation of both ORFs increases, while the fractional repression remains unchanged. Thus, co-translation is compatible with a regulatory model in which uORFs reduce CDS translation efficiency without abolishing it. This has been clarified in the revised manuscript.

    1. eLife Assessment

      This study offers an important contribution to our understanding of the role of layer 6b cortical neurons in sleep-wake regulation, providing new insight into how this understudied neural population may regulate cortical arousal via orexin signaling. The evidence supporting these findings is solid, although somewhat constrained by limitations in the specificity of the genetic targeting strategy. Nonetheless, the work introduces new avenues for uncovering how the classical wake-promoting peptide, orexin, exerts its effects on the cortex.

    2. Reviewer #1 (Public review):

      Summary:

      Meijer et al. sought to investigate the role of cortical layer 6b (L6b) neurons in modulating sleep-wake states and cortical oscillations under baseline and sleep deprived conditions and in response to orexin A and B. Using chronic EEG recordings in mice with silencing of Drd1a+ neurons (via constitutive Cre-dependent knockout of SNAP25), the authors report that while overall baseline sleep-wake architecture and response to sleep deprivation are minimal/unchanged, "L6b silencing leads" to a slowing of theta activity during wakefulness and REM sleep, and a reduction in EEG power during NREM sleep. The manuscript is well written with clarity and transparency. Although Drd1a+ neurons are not exclusive to L6b, the authors describe key future studies to identify a causal role for L6b neurons in brain state regulation. These studies contribute to a growing body of evidence that cortex-in addition to subcortical brain regions-plays a role in brain state regulation.

      Strengths:

      (1) The text is well written.

      (2) The authors are transparent about methodological details and study limitations.

      (3) The stated sleep, circadian, and orexin infusion experiments are well designed, executed, and analyzed.

      Weaknesses:

      (1) Outcomes are attributed to silencing cortical L6b neurons, but the genetic manipulation is not specific to L6b neurons or cortex. The authors acknowledge this as a limitation and offer targets for future studies to identify L6b neuron-specific contributions to stated outcomes that include spatially restricted manipulations.

      (2) Experiments use only male mice, which limits generalizability to females.

      Comments on revised version:

      The authors took great care in addressing my previous comments, and I do not have any additional concerns.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Meijer and colleagues investigated the effects of inactivation (conditional silencing) of cortical layer 6b neurons on sleep-wake states and EEG spectral power under the following three conditions: during natural sleep-wake states, after sleep deprivation, or after intracerebroventricular administration of orexin A and B. The authors report that silencing of L6b neurons did not have a significant effect on the total time spent in sleep-wake states, duration or number of state epochs, or the response to sleep deprivation. However, silencing of L6b neurons did slow down theta-frequency (6-9 Hz) during wake and REM sleep, and reduced the total EEG power during NREM sleep. Infusion of orexin A in the mice in which cortical layer 6b neurons were inactivated produced an increase in wakefulness. A similar effect was observed after infusion of orexin A in the mice in which these neurons were not silenced, but the effect (i.e., increase in wakefulness) was of a smaller magnitude. Silencing of cortical layer 6b neurons attenuated the effect of orexin B in increasing theta activity, as was observed in the control mice. The authors conclude that the cortical neurons in layer 6b play an essential role in state-dependent dynamics of brain activity, vigilance state control and sleep regulation.

      Strengths:

      - A focus on cortical layer 6b neurons, which is an understudied neuronal population, especially in the context of brain and behavioral state transitions.

      - The authors used a well-established mouse model to study the effect of inactivation of cortical layer 6b neurons.

      Weaknesses:

      - Although the authors used a highly selective approach to silence layer 6b neurons, the observed changes in EEG oscillations cannot be solely attributed to layer 6b neurons because of the ICV route for orexin administration.

      - The rationale for using only male rats is not provided.

      Comments on revised version:

      The authors have addressed my concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) All outcomes are attributed specifically to L6b neurons, but the genetic manipulation is not specific to L6b neurons. The authors acknowledge this as a limitation, but in my view, this global manipulation is more than a limitation - it affects the overall interpretations of the data. The Hoerder-Suabedissen et al., 2018 paper shows sparse, but also dense, expression of Drd1a+ neurons in brain regions outside of the L6b. Given this issue, the results are largely overstated throughout the paper.

      We appreciate the reviewer’s careful reading and concern that some of our statements may have overstated the implications of our data. The Drd1a Cre mouse model used (FK164) has a relatively selective expression of Drd1a Cre in cortex, but indeed some expression is seen subcortically. This is an acknowledged limitation which is now explicitly addressed in the revised manuscript.

      (2) It is not clear to me that the "silencing" of Drd1a+ neurons was verified.

      In our previous publications, we showed confirmation of the loss of regulated synaptic vesicle release from the Cre-positive neuronal population (Marques-Smith et al., 2016; Hoerder-Suabedissen et al., 2018; Messore et al., 2024). This has now been described in the revised manuscript.

      (3) There were various discrepancies (and potentially misattributions) between the stated significant differences in Supplementary Table T1 data and Figure 3a & S2 spectral plots. This issue makes it difficult to effectively evaluate the main text and stated outcomes.

      We thank the reviewer for their careful attention to the statistical analyses and for noting the inconsistencies in how the results of the spectral analysis were presented: in the text we described two-way ANOVAs with according posthoc tests but in the figures significance markers were positioned based on multiple t tests. We have now carefully revised the spectral results and implemented a consistent approach in statistical reporting and spectral plots. We have updated Supplementary Table T1, Figure 3a and S2 to ensure that all statistics are presented consistently throughout the manuscript, i.e. with two-way ANOVAs and accompanying posthoc tests. Please note that we performed all spectral analyses in the range between 0.5 and 128 Hz (excluding the range between 49-51.5 Hz due to electrical noise from the power grid) but only plot the range between 0.5-30 Hz as the spectral bands most relevant for sleep neurophysiology are contained in this range.

      Related, the authors stated that post hoc comparisons of EEG spectral frequency bins were not corrected for multiple testing. Instead, significance was only denoted if changes in at least two consecutive frequency bins were significant. However, there are multiple plots in which a single significance marker is placed over an isolated bin (i.e., 4c, 6, S5, S6). Unless each marker is equivalent to 2 consecutive frequency bins, these markers should be removed from the plots. Otherwise, please define the frequency and size of these markers in the main text.

      In line with the previous comment, we have adjusted markers to reflect the results from posthoc tests after two-way ANOVAs.Please note that Figure 6 and the related supplementary figures S5 and S6 have now been removed from the manuscript, as careful re-analysis indicated that the sample size was too low to support a strong conclusion regarding the comparison of orexin effects between genotypes. We stated in the text that we would only include posthoc significance when at least two consecutive bins were significant, but this was indeed not supported in our figure, where each marker reflects one 0.25 Hz bin. We have now adjusted our code to ensure that only markers are plotted when at least two consecutive bins are significant in bin-wise posthoc comparisons.

      (4) A rainbow color scale, as in Figure 3, we've now learned, can be misleading and difficult to interpret. The viridis color scale or a different diverging color scale are good alternatives.

      Thank you for pointing this out, we have adjusted the colour scale.

      (5) How much time elapsed between vehicle/orexin A & B infusions?

      There were 2-4 non-infusions days between infusions. We have added this information to methods.

      (6) For Figure 6, there are statistical discrepancies between the main text and the plots (pg. 10):

      (a) The text claims post hoc differences for relative ORXA frontal EEG, but there are no significance markers on the plot.

      (b) The text states that there were no post hoc differences for the relative ORXA occipital EEG, but significance markers are on the plot.

      (c) The main test for the relative ORXB frontal EEG was not significant, but there are post hoc significance markers on the plot.

      (d) For relative ORXB occipital EEG, there are significant markers on the plot outside of the stated range in the text.

      We agree with the reviewer, and we decided to exclude this figure from the manuscript as the sample size for some key comparisons was too low to support any strong conclusions and therefore presenting this analysis is potentially misleading. We explain the rationale for excluding this analyses in the revised manuscript.

      (7) Some important details are only available in figure captions, making it difficult to understand the main text. For example, when describing Figure 3c in the main text on page 7, it is not clear what type of transitions are being discussed without reading the figure caption. Likewise, a "decrease," "shift," and "change" are mentioned, but relative to what? Similar comment for the EEG theta activity description on pages 7 - 8. Please add relevant details to the main text.

      We have adjusted the wording in the main text to reflect more precisely which comparisons are shown in the figures.

      (8) Statistical comparisons for data in Figure 3e, post hoc analyses for data in Figure S7a-b REM data, and post hoc analyses for Figure S7c (not b) occipital EEG should be included to support differences claims. Please denote these differences on the respective plots.

      Please note that the previously named Supplementary Figures S5 and S6 have been removed from the manuscript, and that the Supplementary Figure S7 in this comment refers to the figure currently named Supplementary Figure S5.

      We have added the statistical comparisons for Figure 3e, Supplementary Figure S5A and Figure S5b to the results section. In Figure S5c, there was an overall genotype difference, but there was no significant time x genotype interaction, so we have not performed posthoc tests and did not plot posthoc significance markers for this figure. We have adjusted the wording in the results section to make this clearer. We have adjusted the reference to the figure S5c which was incorrect, thank you for your careful attention.

      (9) In the subsection titled "Layer 6b mediates effects of orexin on vigilance states (pg. 8)," there does not seem to be any stated differences between control and L6b silenced mice. A more accurate subtitle is needed.

      We agree with the reviewer and the title of this sub-section has now been changed accordingly.

      Reviewer #2 (Public review):

      Weaknesses:

      (1) Although the authors used a highly selective approach to silence layer 6b neurons, the observed changes in EEG oscillations cannot be solely attributed to layer 6b neurons because of the ICV route for orexin administration.

      We thank the reviewer for this important comment. The ICV route of orexin administration cannot guarantee that only cortical Drd1a-Cre–expressing neurons are reached by orexin, and the Drd1a-Cre driver line is highly selective but not entirely specific for layer 6b neurons (see also response to reviewer #1, comment 1). We have therefore changed the wording of the stated effects and addressed this consideration in the Limitations section of the manuscript. Please note that, as mentioned above, Figure 6 has now been excluded from the manuscript.

      (2) The rationale for using only male rats is not provided.

      We thank the reviewer for highlighting this omission. We now provide the rationale for using only male mice in the methods section as follows: “In the current study, only male mice were used, because our experimental protocol precluded the possibility of accurately monitoring the oestrous cycle, which has marked effects on brain activity, arousal and vigilance states. We therefore decided to use male mice only for the current study but are planning to use both sexes in future work.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Better descriptions of L6b connectivity will improve clarity in the second paragraph of the Introduction (pg. 3). For example, it is not explicitly stated that L6b projects to L5 before the authors describe L5. Therefore, the L5 description seems irrelevant.

      We thank the reviewer for this request for clarification. We mention the connectivity between L6b and L5 because L5 pyramidal neurons have recently been found to play a key role in sleep-wake regulation (Krone et al., Nat. Neurosci. 2021; Honjo et al., 2025; Wasilczuk et al, 2025; Krone et al., 2025). We have now amended the corresponding section of the introduction to emphasise the potential functional relevance of this connection as follows:

      “L5, the major output layer of the cortex, is also bidirectionally communicative with higher order thalamic nuclei (Hoerder-Suabedissen et al., 2018) as well as layer 5 pyramidal neurons (Zolnik et al., 2024). Since several subtypes of L5 pyramidal neurons have recently been shown to play important roles in distinct aspects of sleep-wake regulation (Krone et al., 2021, 2025; Hong et al. 2023; Wasilczuk et al. 2025; Honjo et al., 2025; Chouafeev et al., 2025); depth of anaesthesia (Wasilczuk et al. 2025), and the influence of stress on sleep (Chouafeev et al. 2025) the projections of orexin-sensitive L6b to L5 pyramidal neurons may be a key circuitry in the top-down regulation of brain states.”

      (2) There are plots where the y-axis tick label appears to be offset from the tick mark (4a, S5b, S6a).

      Thank you for spotting this graphical issue. We have removed the y-axis tick labels from Figure 4a to avoid confusion. Please note that we decided to remove Figure S5 and Figure S6, because after careful re-analysis we concluded that the group size was too small to draw conclusions on orexin spectra and that any results could be potentially misleading.

      (3) The 2-h time constant, I believe, is depicted in Figure 4H (not 4G).

      Thank you for spotting this. We have corrected the figure legends accordingly and double-checked that Figure 4G depicts the 2-h time constant and Figure 4H the 6-h time constant.

      (4) "...although there was an indication of a higher absolute theta-peak power in layer 6b silenced mice (Figure S6)," pg. 10. It is not clear to me how the data lead to this conclusion.

      Thank you for identifying this inconsistency, which resulted from a preliminary statistical analysis subsequently corrected. We have now improved the statistical analysis of spectral data (for more details see comments to both reviewers in public response) and removed this statement, which in fact is no longer supported by the data.

      (5) Exclusion of female mice is not listed as a limitation.

      We now discuss this limitation as follows:

      “In the current study, only male mice were used, because our experimental protocol precluded the possibility of accurately monitoring the oestrous cycle, which has marked effects on brain activity, arousal and vigilance states. We therefore decided to use male mice only for the current study but are planning to use both sexes in future work.”

      (6) A brief description of why Cplx3 and Tbr1 antibodies are being used will be helpful to include in the Methods (pg. 21) in addition to what is in the figure caption.

      We have added the following information to the methods section to clarify why we used these two antibodies: “rabbit α-Cplx3 to distinguish between L6a and L6b” “mouse α-Tbr1 to identify the L5-6 boundary”

      (7) Including a label/title for the Figure 2c spectral plots will be helpful. It is not immediately clear if these are light period & dark period data or frontal & occipital data.

      Thank you for pointing this out, we have updated the figure legend to clarify what is shown on this Figure

      Similar comments for S2 and S3a plots. Including a state label on the plots will be helpful in addition to the caption description.

      We have now added the state labels for Figure panels S2 and S3a for improved clarity.

      Reviewer #2 (Recommendations for the authors):

      This is a soundly conducted and well-written study that enhances our understanding of the cortical control of states of consciousness. I do not have any major concerns, but would like the authors to consider some alternate possibilities as suggested in my comments below:

      We thank the reviewer for this positive assessment of our manuscript and the helpful suggestions.

      (1) Given that the inactivation of layer6b neurons did not affect the time spent in sleep-wake states, to me it appears that these neurons likely have a role in creating the background neural conditions/oscillations supportive of an activated state rather than a direct role in behavioral state control.

      We completely agree with the reviewer and have made the wording more consistent throughout the manuscript, now using “brain state control” rather than “behavioural state control” to clarify that the main effect observed in the L6b-silenced mouse model is a change in spectral characteristics reflecting brain oscillations, rather than effects on vigilance states, which were modest.

      (2) Does the observed shift in REM sleep-related theta-peak frequency in the occipital derivation suggest changes in local neural processes, or could it be just a matter of better signal detection because theta is most prominent at or around the hippocampal region, which is approximately the location of occipital electrodes in this study.

      The source of the shift in REM sleep–related theta peak frequency in the occipital derivation cannot be established with EEG recordings alone. Additional intracortical or intrahippocampal recordings would be necessary to distinguish between the two possible explanations proposed by the reviewer. We have discussed this further in the revised manuscript.

      (3) Orexinergic system innervates multiple subcortical sites and widely covers the cortex too, because of which the effect of ICV orexins cannot be attributed to just layer6b neurons as described in the manuscript ("Layer 6b mediates effects of orexin on brain activity.").

      We agree with the reviewer that this is a limitation. We have now adjusted the subtitle of the paragraph describing the results from the ICV administration of orexin and further mention this important consideration in the ‘limitations’ section of the discussion.

      (4) While the current study is focused on sleep-wake mechanisms, the findings reported here have much broader implications for behavioral and/or brain state arousal and provide a mechanistic bridge between different states of consciousness, including general anesthesia. Therefore, the authors may consider tying these findings with the recent work on the role of the prefrontal cortex in arousal from general anesthesia and slow-wave sleep (PMID: 35436248, PMID: 29937348, PMID: 33328847).

      We thank the reviewer for this excellent recommendation. We are now citing these papers in the revised manuscript.

      (5) It's up to the authors, but I do not see the need for the section on Clinical Implications. It's very speculative, and it makes the entire discussion section heavy.<br />

      We have considerably shortened the discussion of potential clinical implications to make the manuscript more concise.

      (6) Figure 1: It's difficult to compare the EEG power the way figures are set up right now. I think it would enhance clarity if the authors separate the plots based on state and show power from the control and silenced neuronal group in the same plot. Also, the colors are too similar (essentially a shade of green/blue) to provide effective visual resolution. This is especially true in panel d. Please consider changing the color scheme.

      This comment seems to refer to Figure 2 and subsequent figures with analysis of vigilance states and EEG spectra (Figure 1 contains histological images). We have selected the colour scheme for colour-blind individuals. Therefore, the main difference is in the saturation, not the colour of the plots. We have tested the visibility of the colour scheme on a high-resolution screen with the original image files and can reassure the reviewer that the genotype differences, which are slightly blurred in the reduced-resolution figures provided within the combined text file for the review process, are easily distinguishable in the final figure quality.

      (7) I don't understand the y-axis scale in Figure 1. How can this be 500% and if it is, then 500% of what?

      This comment also seems to refer to the analysis of slow wave activity (SWA) in Figure 2 rather than to Figure 1 (histology figure). The percentage of SWA is normalised to the average SWA across the recording. Since NREM sleep is characterised by considerably higher SWA than wakefulness and REM sleep, the level of SWA during NREM sleep is in the range of 200-300%, and can be even higher after long wake episodes which are followed by a rebound of NREM sleep SWA. Hence, the upper limit of the y-axis in these (and subsequent) plots of SWA is 500% (of the average SWA). We have amended the figure legend to clarify that SWA is presented here as percentage of average SWA across the recording.

    1. eLife Assessment

      In this potentially valuable computational study, the authors conducted extensive atomistic and coarse-grained simulations to probe the temperature-dependent phase behaviors of ELF3, a disordered component of the evening complex in plant. The results aim to highlight the role of polyQ tracts in modulating temperature-responsive structural and condensation behavior. Despite considerable improvements in the revised manuscript, the level of evidence is considered incomplete, since several of the supplementary observables introduced to support the revised claim indicate that the variants studied are not statistically distinguishable within the reported replicate uncertainty.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript explores the role of the Evening Complex (EC), specifically focusing on ELF3, a disordered protein component of the EC, and its temperature-dependent phase behavior. The study highlights the role of polyQ tracts in modulating temperature-sensitive condensate formation and provides a combination of computational approaches, including REST2 simulations and coarse-grained Martini simulations, to investigate how polyQ tract length and sequence context influence this behavior.

      Strengths:

      The study addresses a key question in plant biology - how temperature influences circadian clock-mediated growth regulation through protein phase behavior. The manuscript introduces the novel finding that polyQ tract length modulates the temperature-dependent formation of helices and condensates.

      Weaknesses:

      (1) Coarse-Grained Simulation Results Not Supported by Data:

      The results presented in Figure 6A of the manuscript do not seem to show a clear trend in the number of clusters formed as a function of polyQ tract length. This is particularly evident in the comparison between 0Q and 7Q polyQ lengths, which display statistically similar values in terms of the number of clusters. The lack of distinction between these values raises questions about the sensitivity of the coarse-grained simulations to polyQ tract length, which the authors claim as a key modulator of condensate formation. This discrepancy weakens the argument that polyQ length directly impacts the clustering behavior in the simulations.

      Suggested Analysis:

      a) A more detailed statistical analysis should be performed to assess whether the observed differences between polyQ lengths are significant. This could involve hypothesis testing or the use of error bars in the graphs to better communicate the variability in the data.

      b) Additionally, the authors should examine whether there are other features, such as cluster shape or internal structure, that might differentiate between different polyQ lengths, even if the total number of clusters is similar.

      (2) Inconsistency in Cluster Size Across Temperatures (Figure 6B):

      The results in Figure 6B show a striking difference in the size of the largest cluster between temperatures of 290K and 300K. This abrupt shift in behavior lacks a clear mechanistic explanation. Typically, phase transitions driven by temperature are more gradual, unless there is some underlying structural or chemical shift that the authors have not accounted for. Without a clear explanation, this sudden change in behavior reduces confidence in the simulation results.

      Suggested Analysis:

      a) The authors should explore possible explanations for the dramatic difference in cluster size between 290K and 300K. For example, they could investigate whether specific interactions (such as the breaking or formation of hydrogen bonds or hydrophobic contacts) might explain the behavior at higher temperatures.

      b) It is important to check whether the coarse-grained simulation model has been adequately parameterized and scaled for accurate temperature dependence. Atomistic simulations of monomers and dimers with varying polyQ tract lengths could be used to fine-tune the coarse-grained model, ensuring it accurately reflects molecular behavior. The gross estimate of a 10% scaling factor might be insufficient and could lead to inaccurate representations of cluster formation.

      (3) Scaling of Coarse-Grained Model with Atomistic Simulations:

      As mentioned, the coarse-grained model used in the study may not have been properly scaled against atomistic data. A simple scaling factor of 10% may not be appropriate for accurately capturing the behavior of polyQ tracts across different lengths, especially considering their sensitivity to subtle changes in temperature. Without rigorous validation against atomistic simulations, the coarse-grained model's predictions could be skewed.

      Suggested Analysis:

      a) To address this, the authors should compare the coarse-grained model with atomistic simulations of monomeric and dimeric forms of ELF3 with different polyQ tract lengths. By comparing key structural parameters (e.g., radius of gyration, contact maps, and clustering propensity), the authors could adjust the coarse-grained model to more accurately reflect the atomistic behavior. The authors have wealth of atomistic simulation data that could afford such benchmarking and identification of scaling factor

      b) Additionally, the authors should investigate whether the assumed scaling factor of 10% is appropriate for each polyQ length or whether it needs to be refined based on specific properties, such as the number of hydrophobic interactions or secondary structure stability.

      (4) Lack of Analysis for Liquid-Like Behavior in Phase Separation:

      The simulations presented in the manuscript do not analyze the liquid-like behavior of ELF3 condensates, which is a key characteristic of liquid-liquid phase separation (LLPS). In LLPS systems, condensates are often dynamic, with chains exchanging between clusters, indicating liquid-like rather than solid-like behavior. The authors fail to probe this crucial aspect, which is necessary to support the claim that ELF3 undergoes phase separation.

      Suggested Analysis:

      a) The authors should conduct additional analyses to probe the liquid-like nature of the clusters formed by ELF3. One approach would be to analyze the dynamics of chain exchange between clusters, measuring how frequently chains leave one cluster and join another over time. This analysis would reveal whether the condensates behave as liquid-like, dynamic structures or more static, solid-like aggregates.

      b) Additionally, the temperature dependence of these exchange dynamics should be investigated. In true liquid-liquid phase separation, the rate of chain exchange is often sensitive to temperature. Observing how this rate changes between 290K and 300K, for instance, could help explain the abrupt shift in cluster size seen in Figure 6B.

      c) The authors should also analyze whether the internal structures of the condensates are consistent with a liquid-like phase. For example, radial distribution functions and contact lifetimes could be calculated to reveal whether the clusters exhibit liquid-like organization.

      (5) Lack of justification of polydispersity of polyQ:

      The authors don't provide any rationale for choice of different copies of polyQ used in the manuscript for their chain-growth simulation studies. It will be more apt if it can be motivated via some precedent experimental observations.

      (6) Lack of initiative to connect to Experiments:

      While the computational models and simulations provide robust theoretical insights, the absence of direct experimental validation weakens the overall impact of the manuscript. For example, experimental data on how specific mutations in the polyQ tract influence ELF3 behavior in vivo would significantly bolster the authors' claims. The manuscript would benefit from either citing existing experimental studies that corroborate these findings or from suggesting future experimental directions.

      Comments on revised version:

      The authors have now adequately addressed to the key concerns of manuscript. The manuscript in the present form looks significantly improved.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigate how ELF3, a disordered scaffolding protein in the plant circadian Evening Complex, responds to temperature by forming reversible nuclear condensates. They focus on the C-terminal prion-like domain and on a variable polyglutamine tract within it, asking how the tract length and surrounding sequence context tune temperature-responsive structural and condensation behavior. Using a tiered set of computational approaches, including sequence heuristics, hierarchical chain-growth ensembles, all-atom enhanced-sampling simulations, and coarse-grained condensate simulations of 100 monomers, they characterize wild-type, polyQ deletion, polyQ expansion, and an aromatic-disrupting F527A variant. In the revised manuscript, the central claim has been reframed so that polyQ length is now described as tuning condensate material properties rather than driving temperature-sensitive phase separation, with temperature-responsive condensation attributed primarily to a sticker-rich aromatic contact network.

      Strengths:

      The biological question is important and timely, and the multiscale computational strategy provides a fresh view of an intrinsically disordered protein and its variants. The all-atom enhanced sampling analyses identify a temperature-dependent long-range aromatic contact involving F527 and a methionine-tyrosine coordination motif, which are concrete and mechanistically interesting observations beyond what coarse-grained or sequence-only methods could provide. In response to the previous round of review the authors have added replicate averaged statistics with error bars on the new condensate analyses, introduced new dynamics observables including effective diffusivity, an anomalous diffusion exponent, the self van Hove function, shape anisotropy, per chain radius of gyration in the condensed phase, and a condensate lifetime, provided cluster size time series for transparency, justified the choice of polyQ tract lengths against published Arabidopsis polymorphisms, expanded the Methods with explicit formulas for the new analyses, and included a split half convergence check for the all atom ensembles. The reframing toward a sticker spacer interpretation is consistent with recent experimental work and represents a more cautious and defensible reading of the data.

      Weaknesses:

      Despite these substantive additions, several core concerns from the previous review remain only partially addressed, and, on close reading, the new supplementary analyses do not robustly support the reframed claim that polyQ length tunes condensate material properties. Error bars and replicate-averaged statistics were added to the new condensate panels, but the helical propensity and per-residue analyses throughout the rest of the manuscript still show only a single curve per temperature, so variability for these key observables remains unreported. Several of the newly added dynamics observables show that the variants are essentially indistinguishable within the reported uncertainty: the self van Hove distributions, the shape anisotropy distributions, and the per chain radius of gyration distributions in the condensed phase overlap almost entirely across variants, and the anomalous diffusion exponent has between replica spreads at low temperature that exceed the variant to variant differences, with variant orderings that change with temperature. The variant-dependent signal that does survive, namely a drop in condensate lifetime for the polyQ expansion and the aromatic mutant at the highest temperature studied, rests on a single temperature point, with replicate spreads spanning most of the metric's dynamic range.

      The cluster size time series at higher temperatures shows the dominant cluster oscillating over a wide range across replicas, indicating intermittent dissolution and incomplete convergence in the very temperature regime where the variant-specific claims are made. The only convergence test provided is a split-half radius-of-gyration analysis for the all-atom ensembles, with no slab-geometry or coexistence-density check for the coarse-grained condensate simulations. The polyQ deletion variant forms dominant clusters comparable in size to wild type at low and intermediate temperatures, which on its own argues that variable polyQ presence is not a primary determinant of clustering and supports the earlier concern that the temperature sensitive behavior is dominated by generic chain length and aromatic sticker effects rather than polyQ specific sequence effects, a concern that the reframing softens but does not resolve. Statistical significance is not assessed anywhere, and with three replicas and largely overlapping error bars, claims of variant-specific differences would benefit from explicit statistical tests. Minor quality control issues are also visible in the supplementary material, including a mislabeling of the aromatic mutant in two analysis panels and an inconsistent trajectory length for one variant at one temperature.

      Additional Context for Readers:

      Readers should interpret the molecular mechanism proposed here with caution. The reframing from polyQ length driving temperature-sensitive phase separation to polyQ length tuning of condensate material properties is more scientifically measured and aligns with recent experimental work, but several of the supplementary observables introduced to support this revised claim indicate that the variants studied are statistically indistinguishable within the reported replicate uncertainty. The most robust observation in the revised work is that the prion-like domain undergoes a temperature-responsive break of an aromatic contact in all-atom simulations and that aromatic sticker contacts dominate inter-protein interactions in coarse-grained condensate simulations. The mechanistic role of the polyQ tract, beyond generic chain length and hydration effects, remains, as in the original submission, not clearly established by the simulations presented. Independent experimental validation of the proposed aromatic contact and of the predicted material-state differences between polyQ variants will be needed to establish the molecular mechanism, and improved condensate convergence tests, uniformly reported error bars across all simulation-derived figures, and explicit statistical tests of variant-versus-variant differences would substantially strengthen confidence in the conclusions.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this potentially valuable computational study, the authors conducted atomistic and coarsegrained simulations to probe the temperature-dependent phase behaviors of ELF3, a disordered component of the evening complex in plant. The results aim to highlight the role of polyQ tracts in modulating the temperature sensitivity. The level of evidence is considered incomplete, due to the lack of systematic calibration of the coarse-grained model and limited statistical uncertainty analysis, especially considering the relatively subtle nature of the differences due to temperature change.

      We agree that the subtle temperature dependence of ELF3-PrD condensation requires rigorous uncertainty reporting and careful interpretation of CG predictions. In the revised manuscript we therefore (i) report mean ± SEM across independent replicas for all CG observables and provide full time series in the Supplementary Information, and (ii) expand our CG analysis beyond cluster counting to include condensate stability (size), lifetime, internal mobility (D, α), dynamic heterogeneity (van Hove), and structural descriptors (anisotropy, singlechain compaction/density). These additions strengthen the robustness of the conclusions and even enable physical explanations of recent experimental measurements on ELF3-PrD condensates.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript explores the role of the Evening Complex (EC), specifically focusing on ELF3, a disordered protein component of the EC, and its temperature-dependent phase behavior. The study highlights the role of polyQ tracts in modulating temperature-sensitive condensate formation and provides a combination of computational approaches, including REST2 simulations and coarse-grained Martini simulations, to investigate how polyQ tract length and sequence context influence this behavior.

      Strengths:

      The study addresses a key question in plant biology - how temperature influences circadian clock-mediated growth regulation through protein phase behavior. The manuscript introduces the novel finding that polyQ tract length modulates the temperature-dependent formation of helices and condensates.

      Weaknesses:

      (1) Coarse-Grained Simulation Results Not Supported by Data:

      The results presented in Figure 6A of the manuscript do not seem to show a clear trend in the number of clusters formed as a function of polyQ tract length. This is particularly evident in the comparison between 0Q and 7Q polyQ lengths, which display statistically similar values in terms of the number of clusters. The lack of distinction between these values raises questions about the sensitivity of the coarse-grained simulations to polyQ tract length, which the authors claim as a key modulator of condensate formation. This discrepancy weakens the argument that polyQ length directly impacts the clustering behavior in the simulations.

      Suggested Analysis:

      A more detailed statistical analysis should be performed to assess whether the observed differences between polyQ lengths are significant. This could involve hypothesis testing or the use of error bars in the graphs to better communicate the variability in the data.

      Additionally, the authors should examine whether there are other features, such as cluster shape or internal structure, that might differentiate between different polyQ lengths, even if the total number of clusters is similar.

      We agree that the number of clusters in Fig. 6A does not show a strong or monotonic dependence on polyQ length (e.g., 0Q vs 7Q can overlap within uncertainty). The cluster number is highly sensitive to coarsening kinetics and rapidly approaches a late-time plateau, and therefore is not our primary discriminator of variant-dependent condensation behavior.

      To address the reviewer’s request for statistical rigor and additional differentiating features, we have revised the analysis in two ways. First, we now report mean ± SEM across independent replicas for all key CG observables and provide full replicate time series in the Supplementary Information to make variability and convergence/coarsening explicit.

      Second, we shift our main CG conclusions away from “cluster number” and toward more diagnostic observables of condensate robustness and material state, including: (i) stability via the late-time mean largest-cluster size, (ii) persistence/lifetime via the fraction of frames with largest cluster size greater than 50, (iii) internal dynamics via MSD-derived DDD and anomalous exponent ααα, (iv) dynamic heterogeneity via self van Hove distributions relative to a Gaussian reference, and (v) morphology/internal structure via κ<sup>2</sup> and Rg distributions.

      Notably, the κ<sup>2</sup>/Rg distributions are broadly overlapping at 300 K, indicating that in our system variant differences are expressed more strongly in stability/persistence and internal dynamics (D/α/van Hove) than in a large shift in single-chain compaction at this temperature.

      This revised framing also aligns our interpretation with the experimental picture put forward by Huntin et al -- polyQ length modestly affects onset-like behavior but more strongly tunes condensed-phase regimes and dynamics.

      Relevant revisions have been made in the Results and the Discussion sections.

      (2) Inconsistency in Cluster Size Across Temperatures (Figure 6B):

      The results in Figure 6B show a striking difference in the size of the largest cluster between temperatures of 290K and 300K. This abrupt shift in behavior lacks a clear mechanistic explanation. Typically, phase transitions driven by temperature are more gradual, unless there is some underlying structural or chemical shift that the authors have not accounted for. Without a clear explanation, this sudden change in behavior reduces confidence in the simulation Results.

      Suggested Analysis:

      The authors should explore possible explanations for the dramatic difference in cluster size between 290K and 300K. For example, they could investigate whether specific interactions (such as the breaking or formation of hydrogen bonds or hydrophobic contacts) might explain the behavior at higher temperatures.

      It is important to check whether the coarse-grained simulation model has been adequately parameterized and scaled for accurate temperature dependence. Atomistic simulations of monomers and dimers with varying polyQ tract lengths could be used to fine-tune the coarsegrained model, ensuring it accurately reflects molecular behavior. The gross estimate of a 10% scaling factor might be insufficient and could lead to inaccurate representations of cluster formation.

      We agree that the apparently sharp change in largest-cluster size between 290 K and 300 K requires clearer interpretation. In the revised manuscript, we clarify that this behavior does not imply an abrupt thermodynamic phase transition; rather, in a finite (~100-chain) simulation box, the largest cluster size is sensitive to both (i) proximity to a coexistence boundary and (ii) coarsening kinetics. Consistent with this, all systems rapidly coarsen early and then approach a late-time plateau, so the dominant cluster size can change steeply when conditions shift the balance between one system-spanning droplet versus multiple long-lived subclusters.

      To distinguish “true loss of condensation” from “differences in coarsening state,” we added replica-averaged stability and persistence metrics (mean ± SEM) and full time series. Importantly, the condensate lifetime (fraction of frames with largest aggregate-population > 50) is ~1 at both 290 K and 300 K, indicating that both temperatures correspond to a persistently condensed regime, not intermittent nucleation/dissolution. We therefore interpret the smaller dominant cluster at 290 K as reflecting slower coarsening / stronger kinetic arrest, where reduced chain mobility delays merger/annealing into a single large droplet within the simulated time window, leaving a larger satellite/dispersed population despite sustained condensation.

      We further support this interpretation with mechanistic and dynamical analyses added in the revision. As temperature increases from 290 K to 300 K, we observe increased internal mobility (higher effective diffusivity, D) that would accelerate rearrangements and coalescence. In parallel, contact/desolvation analyses show progressive loss of protein-water contacts and gain of protein-protein contacts as clusters mature, and a residue-resolved comparison indicates net contact increases at 300 K relative to 290 K concentrated in aromatic-rich “sticker” regions, consistent with a strengthened intermolecular contact network that promotes more complete annealing at 300 K.

      (We address the reviewer’s points regarding Martini temperature scaling/parameterization together with points (3)-(4) below.)

      (3) Scaling of Coarse-Grained Model with Atomistic Simulations:

      As mentioned, the coarse-grained model used in the study may not have been properly scaled against atomistic data. A simple scaling factor of 10% may not be appropriate for accurately capturing the behavior of polyQ tracts across different lengths, especially considering their sensitivity to subtle changes in temperature. Without rigorous validation against atomistic simulations, the coarse-grained model's predictions could be skewed.

      (4) To address this, the authors should compare the coarse-grained model with atomistic simulations of monomeric and dimeric forms of ELF3 with different polyQ tract lengths. By comparing key structural parameters (e.g., radius of gyration, contact maps, and clustering propensity), the authors could adjust the coarse-grained model to more accurately reflect the atomistic behavior. The authors have wealth of atomistic simulation data that could afford such benchmarking and identification of scaling factor

      Additionally, the authors should investigate whether the assumed scaling factor of 10% is appropriate for each polyQ length or whether it needs to be refined based on specific properties, such as the number of hydrophobic interactions or secondary structure stability.

      We agree that temperature-dependent CG predictions must be interpreted carefully and that the interaction balance should be justified. In the revision, we therefore clarify both our calibration choice and the scope of inference.

      We use Martini 3 with a single, literature-motivated adjustment: protein-water Lennard-Jones interactions are strengthened by 10 percent, following an established strategy shown to improve IDP/multidomain protein behavior in Martini 3. This scaling is applied uniformly to all residues, polyQ lengths, and temperatures to avoid introducing construct-specific parameters and to preserve a controlled comparison across variants.

      We emphasize that our CG simulations are used in a comparative manner (how stability/dynamics/structure change with temperature and polyQ length under a fixed model), and we do not claim a quantitatively exact phase boundary or transition temperature for ELF3. In this spirit, and consistent with how Martini 3 has been used in prior work to probe thermally varying properties across temperature windows (while acknowledging documented limits to temperature transferability), we treat the temperature sweep as a comparative probe rather than an absolute calibration (https://doi.org/10.1063/5.0221199, 10.1021/acscentsci.5c00755, https://doi.org/10.1038/s41592-021-01098-3). Accordingly, we report replica uncertainty (mean ± SEM) for all CG observables and restrict conclusions to qualitative trends that are robust to replicate variability.

      Finally, while we do not undertake a full ELF3-specific reparameterization, we include qualitative checks linking atomistic and CG behavior: the CG model reproduces the same qualitative features of single-chain reorganization inferred from atomistic simulations — notably the radiusof-gyration (Fig. S8) and the rearrangement/exposure of aromatic “sticker” regions that correlate with strengthened intermolecular contacts in the condensate. We emphasize that these comparisons are intended as qualitative sanity checks on trend direction, not as a quantitative validation or calibration of an absolute phase boundary.

      (5) Lack of Analysis for Liquid-Like Behavior in Phase Separation:

      The simulations presented in the manuscript do not analyze the liquid-like behavior of ELF3 condensates, which is a key characteristic of liquid-liquid phase separation (LLPS). In LLPS systems, condensates are often dynamic, with chains exchanging between clusters, indicating liquid-like rather than solid-like behavior. The authors fail to probe this crucial aspect, which is necessary to support the claim that ELF3 undergoes phase separation.

      Suggested Analysis:

      The authors should conduct additional analyses to probe the liquid-like nature of the clusters formed by ELF3. One approach would be to analyze the dynamics of chain exchange between clusters, measuring how frequently chains leave one cluster and join another over time. This analysis would reveal whether the condensates behave as liquid- like, dynamic structures or more static, solid-like aggregates.

      Additionally, the temperature dependence of these exchange dynamics should be investigated. In true liquid-liquid phase separation, the rate of chain exchange is often sensitive to temperature. Observing how this rate changes between 290K and 300K, for instance, could help explain the abrupt shift in cluster size seen in Figure 6B.

      The authors should also analyze whether the internal structures of the condensates are consistent with a liquid-like phase. For example, radial distribution functions and contact lifetimes could be calculated to reveal whether the clusters exhibit liquid-like organization.

      We thank the reviewer for highlighting that liquid-like behavior is a key diagnostic for LLPS. We agree that our original manuscript did not explicitly quantify condensate material properties. In the revision, we therefore add several complementary analyses and figures that directly probe whether the condensed state in our simulations is liquid-like versus dynamically arrested, and how this depends on temperature and polyQ length.

      (i) Condensate persistence vs temperature (stability and lifetime).

      We now quantify two replica-averaged metrics with uncertainty (mean ± SEM): (a) stability, defined as the mean largest-cluster size over a late-time analysis window, and (b) lifetime, defined as the fraction of frames in which the dominant cluster exceeds a fixed size threshold. These results are shown in the new figures “Stability (Mean cluster size)” and “Lifetime (P [size > 50])”. In our system, both 290 K and 300 K correspond to a persistently condensed regime (lifetime ≈ 1 across variants), whereas at 340 K the lifetime drops substantially (≈0.3-0.5 depending on variant), indicating intermittent condensation / partial dissolution at high temperature. This directly demonstrates temperature-dependent persistence of the condensed state and clarifies that the key qualitative change at high temperature is loss of stability and intermittency, rather than a purely static cluster-size difference.

      (ii) Internal mobility and viscoelasticity (D and α).

      To probe liquid-like dynamics within the condensed state, we compute internal Mean squared displacement (MSD) and extract an effective internal diffusivity D(T) and anomalous exponent α(T) (new figures FIG X). In our system, D increases systematically with temperature for all variants, confirming that internal rearrangements accelerate at higher temperature. At the same time, α remains strongly subdiffusive (α ≈ 0.3-0.5), indicating constrained, non-Fickian motion rather than simple liquid diffusion. Importantly, we also observe variant-dependent mobility: around 300-320 K, 0Q exhibits markedly lower D than 19Q, consistent with stronger kinetic arrest in 0Q even when both variants are condensed. Together, these dynamics metrics show that our condensates are not ideal liquids, but instead occupy a viscoelastic / dynamically slowed regime with clear temperature dependence.

      (ii) Dynamic heterogeneity (self van Hove).

      We additionally compute the self van Hove displacement distributions (Fig. SX). In our system, the van Hove distributions deviate from a Gaussian reference matched to the MSD, with an excess of near-zero displacements relative to a simple Gaussian model. This non-Gaussian displacement statistics is consistent with heterogeneous/caging-like dynamics inside the condensed phase, further supporting a viscoelastic (gel-like) rather than purely liquid material state at the timescales accessible to simulation.

      (iv) Internal structure and morphology (Rg and anisotropy).

      Finally, we add structural descriptors as requested. The new Rg distribution and shape anisotropy (κ<sup>2</sup>) violin plots quantify single-chain compaction and heterogeneity in morphology within the condensed phase. In our system these structural distributions are broadly overlapping at 300 K, indicating that differences among variants are more strongly expressed in dynamics (D/α/van Hove) and stability/lifetime, rather than in a large change in single-chain compaction at this temperature. We report these results transparently and include them in the SI as additional mechanistic context.

      We now explicitly frame our CG condensed phases as viscoelastic/dynamically slowed condensates rather than assuming fully liquid droplets. This interpretation is consistent with experimental observations on ELF3 PrLD that report very slow recovery/gel-like behavior under some conditions (i.e., condensates can age into low-mobility hydrogel states).

      (6) Lack of justification of polydispersity of polyQ:

      The authors don't provide any rationale for choice of different copies of polyQ used in the manuscript for their chain- growth simulation studies. It will be more apt if it can be motivated via some precedent experimental observations.

      We agree and have clarified our rationale in the revised manuscript. ELF3’s polyQ tract is a naturally polymorphic short tandem repeat in Arabidopsis, reported to vary from roughly ~7 to ~29 glutamines in natural populations, and this variation has been linked to ELF3-dependent phenotypes and temperature-responsive growth (Undurraga et al.; Jung et al.). Importantly, recent ELF3 PrLD thermosensing/condensation experiments explicitly compare multiple polyQ lengths (including Q0, short/WT-like constructs such as Q7, and expanded tracts around ~Q20) and show that polyQ length tunes temperature-responsive phase behavior and condensate properties (Jung et al.; Hutin et al.).

      Accordingly, for our chain-growth ensembles we chose a small, experimentally motivated set that brackets this range - 0Q (deletion), 7Q (WT-like short), and expanded lengths 13Q and 19Q (with 19Q closely matching the ~Q20 construct used experimentally), so that our simulations map onto established constructs and naturally occurring variation rather than arbitrary copy numbers.

      The manuscript draft has been modified in the Results and Methods sections.

      Jung J-H. et al. A prion-like domain in ELF3 functions as a thermosensor in Arabidopsis. Nature (2020).

      Undurraga S. et al. Background-dependent effects of polyglutamine variation in the Arabidopsis thaliana gene ELF3. PNAS (2012), DOI: 10.1073/pnas.1211021109.

      Hutin S. et al. Phase separation and molecular ordering of the prion-like domain of the Arabidopsis thermosensory protein EARLY FLOWERING 3. PNAS (2023).

      (7) Lack of initiative to connect to Experiments:

      While the computational models and simulations provide robust theoretical insights, the absence of direct experimental validation weakens the overall impact of the manuscript. For example, experimental data on how specific mutations in the polyQ tract influence ELF3 behavior in vivo would significantly bolster the authors' claims. The manuscript would benefit from either citing existing experimental studies that corroborate these findings or from suggesting future experimental directions.

      We agree that our original submission did not make the experimental connections explicit enough, and we have strengthened this in the revision by (i) explicitly anchoring our results to published ELF3 thermosensing/condensation measurements and (ii) articulating concrete, experimentally testable mechanistic predictions that follow from the simulations.

      (i) Explicit connection to published experimental benchmarks: We now cite and discuss key experimental studies that directly probe ELF3 temperature responsiveness and polyQ effects. Jung et al. demonstrated temperature-triggered ELF3 condensation/speckle formation in vivo and showed that polyQ length modulates thermoresponsive behavior. More recently, Hutin et al. compared ELF3 PrLD constructs spanning polyQ lengths (e.g., Q0, Q7, and ~Q20) and reported temperature-triggered phase separation, condition-dependent condensed-phase regimes (droplet-like versus more arrested/gel-/hydrogel-like), and reduced mobility/immobile fractions quantified by FRAP in some regimes. In the revised manuscript we explicitly map these observations onto our results: our coarse-grained simulations capture temperature-dependent condensation propensity, while our added condensate dynamics analyses (MSD-derived internal mobility DDD, anomalous exponent α\alphaα, and self van Hove displacement statistics) indicate dynamically slowed/heterogeneous condensates rather than assuming ideal liquid droplets—consistent with experimentally observed slow FRAP recovery and arrested behavior under some conditions.

      (ii) Mechanistic Connections: While existing experiments establish that ELF3 condensation is temperature-triggered and tuned by polyQ length, they cannot directly resolve the molecular interaction changes that drive these macroscopic readouts. We therefore emphasize that our atomistic and coarse-grained analyses provide a mechanistic interpretation: temperature shifts reorganize and expose “sticker”-rich regions (notably aromatics), strengthening intermolecular contact networks that tune condensate stability and material properties. This framing aligns our conclusions with the experimental picture that polyQ length has modest effects on onset-like behavior but more strongly tunes condensed-phase robustness and dynamics (persistence, internal mobility, and arrest) across temperature

      The modifications relevant to this are in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to explore how a key protein in the circadian clock of plants, ELF3, responds to temperature changes by forming molecular condensates. They focused on understanding the role of a specific region of the protein, a polyQ tract, in promoting temperature-sensitive structural changes and regulating the formation of condensates. Through a series of computational simulations, they sought to uncover the molecular basis for ELF3's temperature responsiveness and its broader implications for plant growth and adaptation to environmental conditions.

      Strengths:

      The study's strength lies in its focus on an important biological question: how plants sense and respond to temperature changes at the molecular level. The authors employed a variety of computational techniques, including coarse-grained simulations, to explore the role of specific molecular features in this process. These methods provide a multi-scale view of protein behavior and offer valuable insights into how molecular structures may influence biological function.

      Weaknesses:

      However, there are notable weaknesses in the evidence provided. While the authors present trends in molecular changes, such as shifts in helical propensity and the formation of condensates, these results seem subtle and are not strongly substantiated by statistical analysis. The lack of error bars in the figures makes it difficult to distinguish between meaningful signals and potential noise in the data. Furthermore, the temperature-sensitive behavior appears to be influenced more by chain length than by sequence-specific effects of the polyQ region, raising questions about whether the findings truly capture the molecular mechanisms responsible for temperature sensing. Additionally, some simulations, particularly those related to the formation of condensates, do not appear fully converged, which casts further doubt on the robustness of the results.

      We appreciate the reviewer’s concerns regarding statistical support, sequence specificity, and convergence. In the revised manuscript we (i) report replicate-averaged means with uncertainty (mean ± SEM) for all key observables and add error bars/shaded bands to the relevant figures, (ii) provide the full time series plots in the Supplementary Information to make variability and equilibration transparent, and (iii) revise our interpretation to emphasize that polyQ length has only modest effects on onset-like metrics but more strongly tunes condensate stability and material state (lifetime, internal mobility (D), subdiffusion exponent (α), and non-Gaussian van Hove signatures). This revised framing is consistent with recent ELF3 PrLD experiments showing that polyQ variation can subtly affect onset while substantially modulating condensed-phase behavior and dynamics. Relevant changes to the main text have been made in the Results and Discussion section.

      Additional Context for Readers:

      Readers should interpret the results with caution, especially regarding the molecular mechanisms proposed for temperature sensing. While the study presents interesting trends, the evidence is not definitive, and the findings may be more reflective of general protein behavior (such as the effect of chain length on condensate formation) than specific sequence-driven responses to temperature. Further experimental studies and more converged simulations will be necessary to fully understand the role of ELF3 in temperature regulation.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I already have listed my possible recommendations for authors for revising their manuscript in the review. By addressing these issues, the authors could significantly improve the robustness of their conclusions and provide stronger evidence for ELF3's role in temperature-responsive phase separation.

    1. eLife Assessment

      This is an important study describing 'SPEx', a broadly accessible method that combines cell expansion, laser microdissection, and mass spectrometry to enable subcellular proteomic profiling. The authors provide convincing evidence that this flexible integration of established techniques provides a robust and practical approach for compartment-resolved spatial proteomics. The authors support their main claims with appropriate validation across multiple subcellular compartments and show that the method can recover known markers while also identifying previously uncharacterized components. Overall, the work is likely to be of broad interest to cell and molecular biologists, particularly those seeking scalable and cost-effective strategies for mapping organelle composition.

    2. Reviewer #1 (Public review):

      Summary:

      The authors present a novel approach to subcellular spatial proteomics by combining laser microdissection with expansion microscopy and LC-MS/MS analysis (SPEx). They implement two different workflows for LMD and LC-MS/MS quantification:

      (1)The standard approach, where an area of interest is cut out by LMD, subjected to proteomics analysis, and compared to the rest of the cell without the dissected ROI.

      (2) The subtraction approach, where ROIs are removed, and the remaining cellular material is compared to samples containing both the surrounding material and the ROI.

      The authors assess the technique by applying it to subcellular targets of various sizes, volumes, and protein compositions such as the nucleus, nucleoli, and Golgi. They demonstrate that SPEx can identify proteins enriched or reduced in ROIs.

      Strengths:

      The broad, relatively easy, and inexpensive applicability of this approach to potentially many cell types and subcellular areas of interest provides an exciting alternative to subcellular fractionation, native immunoprecipitation, or genetically encoded proximity labeling constructs. Moreover, by visually selecting ROIs for subsequent analysis, subcellular context or organelle morphology can be taken into account, as discussed by the authors in the discussion section.

      Weaknesses:

      While strongly supporting the sharing of this approach, we have a number of comments and questions that will improve the impact of the manuscript:

      (1) General:

      a) The manuscript would benefit from restructuring and language revision. In its current form, the writing is sometimes dense and verbose (in particular, the Results section). This makes it difficult to follow the authors' arguments.

      b) The authors mention the possibility of selecting organelles based on morphology. This is left for the discussion, but it seems like a missed opportunity - the authors could compare individual organelles in different morphological states, e.g., connected vs. fragmented mitochondria.

      (2) Technical:

      a) Why do the authors strive and optimize for a 10x expansion factor? Is SPEx compatible with a more standard 4x expansion, as e.g., used in the classic U-ExM approach (https://www.nature.com/articles/s41592-018-0238-1)? This could be added to the discussion.

      b) The U-ExM approach shows improved ultrastructural preservation when using 3%FA with 0.1% glutaraldehyde fixation (GA). Is SPEx compatible with the use of low amounts of GA for fixation?

      c) Related to the above, was the anchoring efficiency reduced only to achieve a 10x expansion factor or does this additionally affect the proteome coverage?

      d) Have the authors considered using alternative anchoring approaches, such as GMA (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0291506#pone.0291506.s001), which potentially increase the amount of sample retained in the hydrogel, thus allowing for better proteome coverage? This could be added to the discussion.

      e) The limitation of the approach to near-2D samples should be mentioned, and alternative approaches for more 3D samples could be discussed.

      f) How are peptides that are directly anchored to the hydrogel dealt with during LC-MS/MS analysis? Are they excluded, or can they be identified during the spectral search? The latter would allow us to get a deeper structural understanding of how proteins are actually anchored into hydrogels, which so far has not been assessed.

      An alternative approach to address this question would be to investigate if the peptide coverage of proteins detected by SPEx is enriched for peptides representing the folded core of proteins as opposed to the surface-exposed regions, which likely get more anchored into the hydrogel.

      g) Same question regarding peptides with NHS labeling. Can they be identified, or do they just compete for ionization and thus negatively affect coverage and dynamic range of the LC-MS/MS approach?

      h) How are the primary and secondary antibodies affecting the proteomics analysis identified as contaminants?

      i) Have the authors observed differences in proteomics coverage of only antibody vs NHS-labeling? Depending on the questions above, could pure antibody-based labeling increase proteomic coverage?

    3. Reviewer #2 (Public review):

      Summary:

      This study introduces a method that combines physical expansion of cells, imaging-guided isolation of defined regions, and protein identification to enable compartment-resolved analysis of protein composition at the subcellular scale. The authors aim to address a central limitation in existing approaches, namely the loss of spatial information during sample preparation or the indirect nature of proximity-based labeling methods. Using several cellular compartments as examples, they demonstrate that their approach can recover compartment-enriched protein sets and identify candidate proteins with previously unassigned localization.

      Strengths:

      A major strength of this work is the conceptual simplicity and accessibility of the approach. By combining established techniques in a modular way, the method avoids the need for genetic manipulation or specialized labeling strategies, making it broadly adaptable across experimental systems. The ability to directly select regions of interest based on imaging represents a clear advantage over indirect enrichment strategies and allows flexible targeting of both membrane-bound and non-membrane-bound compartments.

      The experimental design is also a strong aspect of the study. The use of complementary comparison strategies-analyzing isolated compartments alongside matched "subtracted" controls-provides an internal framework for assessing enrichment and depletion, increasing confidence in spatial assignment. The application of the method across multiple organelles of different sizes and properties demonstrates versatility, and the reported specificity for several compartments is encouraging. In particular, the ability to profile small and biochemically challenging structures highlights a potentially important niche for the approach.

      Weaknesses:

      Despite these strengths, several methodological limitations constrain the interpretation of the results. The most important relates to spatial accuracy in three dimensions. While lateral resolution is improved through physical expansion, the lack of depth resolution introduces uncertainty regarding contributions from structures above and below the selected region. Although the authors argue that this does not substantially affect specificity, the current evidence is largely indirect, and a more rigorous quantification of potential contamination would strengthen this conclusion.<br /> Quantitative interpretation also remains challenging. Because the measurements reflect total protein abundance rather than local concentration, differences in compartment size and protein density can influence enrichment values, particularly for small structures embedded within larger volumes. This issue is evident in the analysis of smaller compartments and complicates direct comparison across conditions. Additional normalization or modeling would help clarify how to interpret these measurements.

      Another limitation concerns variability in the expansion process and its downstream consequences. Differences in expansion factor across samples may affect the definition of regions of interest and introduce variability in sampling, yet the impact of this variability is not fully explored. Similarly, the use of a modified chemical treatment to preserve proteins for downstream analysis is central to the workflow but is not extensively validated with respect to preservation of spatial organization.

      While the identification of previously unannotated proteins is an appealing aspect of the study, validation is limited to a small number of examples, and broader support from independent datasets or literature context is lacking. In addition, the study primarily focuses on steady-state measurements in a single cell type, and therefore does not yet demonstrate the ability of the method to capture dynamic or condition-dependent changes in protein localization.

      Finally, the positioning of the method relative to existing approaches could be more clearly articulated. Although qualitative comparisons are provided, a more systematic and quantitative benchmarking against alternative strategies would help readers better understand the specific advantages and trade-offs.

    4. Reviewer #3 (Public review):

      Franziscus et al. describe an elegant approach for spatially specific proteome analysis. To achieve this, they expand fixed cells and subsequently use a laser to micro-dissect a region of interest, which is then analyzed by mass spectrometry.

      They demonstrate the effectiveness of their approach by analyzing the nucleus, nucleolus, and the Golgi, and benchmark their hits against previous datasets for these organelles.

      The manuscript is very well written and nicely guides the reader through the applied methods. The presented data is convincing, and I do not see the need for additional experimental verification of the protocol. The only minor concern is the novelty of the method and the presentation. A combination of expansion, laser microdissection, and proteomics has been applied in the past (PMID: 36450705, PMID: 39477916). In the manuscript, one of these studies is cited, though it does not become clear that this approach is already described. However, Franziscus et al. describe the approach better and make it more accessible to the reader, especially since the other studies described this methodology in combination with tissue expansion and not in combination with single cell expansion as it is done here. I would ask the authors to be clearer in the introduction about what others have already done and what their contribution is here. In general, I am convinced that the community will benefit from the presented protocol to analyze organelle proteomics in detail.

    5. Author Response:

      Reviewer #1 (Public review):

      Summary:

      The authors present a novel approach to subcellular spatial proteomics by combining laser microdissection with expansion microscopy and LC-MS/MS analysis (SPEx). They implement two different workflows for LMD and LC-MS/MS quantification:

      (1)The standard approach, where an area of interest is cut out by LMD, subjected to proteomics analysis, and compared to the rest of the cell without the dissected ROI.

      (2) The subtraction approach, where ROIs are removed, and the remaining cellular material is compared to samples containing both the surrounding material and the ROI.

      The authors assess the technique by applying it to subcellular targets of various sizes, volumes, and protein compositions such as the nucleus, nucleoli, and Golgi. They demonstrate that SPEx can identify proteins enriched or reduced in ROIs.

      Strengths:

      The broad, relatively easy, and inexpensive applicability of this approach to potentially many cell types and subcellular areas of interest provides an exciting alternative to subcellular fractionation, native immunoprecipitation, or genetically encoded proximity labeling constructs. Moreover, by visually selecting ROIs for subsequent analysis, subcellular context or organelle morphology can be taken into account, as discussed by the authors in the discussion section.

      Weaknesses:

      While strongly supporting the sharing of this approach, we have a number of comments and questions that will improve the impact of the manuscript:

      We thank the reviewer for the careful evaluation of our manuscript and the generally positive assessment. We plan on improving our manuscript based on the reviewers’ comments.

      (1) General:

      a) The manuscript would benefit from restructuring and language revision. In its current form, the writing is sometimes dense and verbose (in particular, the Results section). This makes it difficult to follow the authors' arguments.

      We will improve readability and clarity of the results section in the revised manuscript.

      b) The authors mention the possibility of selecting organelles based on morphology. This is left for the discussion, but it seems like a missed opportunity - the authors could compare individual organelles in different morphological states, e.g., connected vs. fragmented mitochondria.

      The authors agree with the reviewers’ assessment that investigating proteome of organelles based on morphology or cellular state is an exciting application of SPEx. While we plan experiments along this line in the future, we think that these experiments are beyond the scope of this manuscript, which is meant to describe the method and its general usefulness.

      (2) Technical:

      a) Why do the authors strive and optimize for a 10x expansion factor? Is SPEx compatible with a more standard 4x expansion, as e.g., used in the classic U-ExM approach (https://www.nature.com/articles/s41592-018-0238-1)? This could be added to the discussion.

      We aimed for 10x expansion solely because our ultimate goal is to cut out very small structures. Isolating structures as small as nucleoli would not be as reliable with a lower expansion factor (i.e. 4x) expansion. We did not assess the compatibility with U-ExM. We would assume that SPEx would also work with U-ExM as expansion method; omitting protease treatment, however. Still, we performed pilots with just 4x expansion (using TREx) in the early stages of optimization. We were able to isolate single cells and obtain similar protein coverage as with 10x expansion. We will further clarify our motivation to use 10x expansion in the discussion.

      We would also like to point out whether to U-ExM the standard method or not is rather subjective. Even though TREx was published three years later, it is also very widely used. The original expansion microscopy method was published three years prior to U-ExM.

      b) The U-ExM approach shows improved ultrastructural preservation when using 3%FA with 0.1% glutaraldehyde fixation (GA). Is SPEx compatible with the use of low amounts of GA for fixation?

      We tried different fixation methods in the early stages of this study (where expansion was not yet close to 10x). We saw a mild negative effect of GA on the expansion factor, so we avoided it in the later experiments since it also did not seem necessary to preserve the structure of our organelles of interest. However, the use of GA would generally be compatible with SPEx, potentially at the cost of a mild negative effect on expansion factor (see Author response image 1) and proteome coverage. We can add this information to the discussion.

      Author response image 1.

      Fixation methods mini-screen. Cells were fixed with the indicated reagents for 10 minutes at 37°C. After TREx expansion, the diameter of the nucleus was measured (A) and the resulting expansion factor compared to the non-expanded control was determined (B).

      Related to the above, was the anchoring efficiency reduced only to achieve a 10x expansion factor or does this additionally affect the proteome coverage?

      We solely lowered the anchoring in order to allow for higher expansion factors. In earlier pilots we performed proteomic analysis on samples that were just expanded 4x using standard TREx expansion (also using the original anchoring strategy from the TREx publication, consisting of 0.2 mg/ml AcX for overnight at RT). We presented the results of this pilot in Fig S1A. We still detected over 2,000 proteins from 10 cells, a coverage, which is highly similar to what we found in the final experiments (Figure 2F), in which the anchoring was lower yielding 10x expansion. Based on these data, we hypothesize that anchoring (and expansion factor!) has a negligible impact on protein coverage. We will clarify this in the manuscript.

      d) Have the authors considered using alternative anchoring approaches, such as GMA (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0291506#pone.0291506.s001), which potentially increase the amount of sample retained in the hydrogel, thus allowing for better proteome coverage? This could be added to the discussion.

      We did not use alternative anchoring approaches. We modified the TREx protocol to fit our purposes and since this was sufficient, we did not explore alternatives. However, using anchoring approaches, in which higher amounts of sample could be retained in the gel might be beneficial for the proteomics coverage. We will keep this suggestion in mind for future experiments. Thank you for the suggestion!

      e) The limitation of the approach to near-2D samples should be mentioned, and alternative approaches for more 3D samples could be discussed.

      The authors agree that SPEx is limited to near-2D samples at this point. We suggest that SPEx is applicable for 3D samples (e.g. in tissues) by performing cryosectioning. TREx has been shown to be compatible with sectioned tissue (Damstra et al., 2022). We will elaborate this in the discussion.

      f) How are peptides that are directly anchored to the hydrogel dealt with during LC-MS/MS analysis? Are they excluded, or can they be identified during the spectral search? The latter would allow us to get a deeper structural understanding of how proteins are actually anchored into hydrogels, which so far has not been assessed.

      The reviewer raises an interesting point. In general, peptides carrying the anchoring modification are analysed by LC-MS, but we did not include these specific modifications in the database search. Overall, we assumed that the labeling would be low and stochastic and hence should, if at all, only minimally affect the detection of peptides. Nevertheless, in response to the reviewers’ comment, we searched the MS data again for the crosslinking reagent linked to lysine residues. However, we could not get any confident hit for any peptide containing this modification. Since we cannot exclude that the modification precludes the identification of the corresponding peptides, we compared the number peptides generated by trypsin cleavage after arginine and lysine. As the human genome contains similar proportions of both amino acids, one would expect similar numbers of both peptide types being identified. Any modifications of lysine by the anchoring reagent used, would prevent tryptic cleavage and thus reduce the number of lysine peptides. As shown in Author response image 2, the number of lysine terminating is only slightly lower compared to arginine terminating peptides. Notably, the proteomics results of a different fixed human tissue sample directly extracted by laser capture micro dissection without expansion showed a very similar lysine to arginine peptide ratio. This indicates that the large majority of lysine residues is not modified and affected by the hydrogel anchoring.

      Author response image 2.

      Number of peptides identified either terminating with lysine (K) or arginine (R) across all samples shown in Figure 5F.

      An alternative approach to address this question would be to investigate if the peptide coverage of proteins detected by SPEx is enriched for peptides representing the folded core of proteins as opposed to the surface-exposed regions, which likely get more anchored into the hydrogel.

      Because of the negligible amounts of modified peptides, we did not investigate this potential bias of surface-exposed versus folded-core peptides.

      g) Same question regarding peptides with NHS labeling. Can they be identified, or do they just compete for ionization and thus negatively affect coverage and dynamic range of the LC-MS/MS approach?

      The reviewer raises a similar point as above for another lysine labeling used during the SPEx protocol. Again, we specifically looked for this modification by re-searching the raw MS data, but still could not identify any peptides, carrying this modification on a lysine residue. Even though we cannot exclude that this rather large modification prevents detection, considering the high number of lysine terminating peptides in our dataset (see Figure 2), we would expect that also this labeling step is stochastic and affects only a minor proportion of the proteins.

      h) How are the primary and secondary antibodies affecting the proteomics analysis identified as contaminants?

      We thank the reviewer for this comment. Since antibodies bind to proteins in a non-covalent manner, they will be released during the denaturing steps of the protocol. Of course, the antibodies will stay in the sample, be digested and analyzed and could, if very abundant, affect the analysis of the proteins from the samples. To check this possibility, we re-searched the MS data including the sequences of the antibodies used. To our surprise, we could not detect any peptides of these antibodies. This suggests that the concentrations of the antibodies used are much lower than those of the sample proteins and thus should not have any impact on the proteomics results.  We interpret this result also as a benefit of our method compared to organellar-IP.

      i) Have the authors observed differences in proteomics coverage of only antibody vs NHS-labeling? Depending on the questions above, could pure antibody-based labeling increase proteomic coverage?

      We did not perform this comparative analysis, since we always used NHS dyes. In the experiments presented in this manuscript, NHS dyes allowed easy visualization of the whole cell without the use of antibodies. This NHS staining was essential for this particular setup for sample acquisition. We cut out entire cells, cells lacking the nucleus and cells lacking the Golgi apparatus, which served as critical controls. However, other ways of detecting cell boundaries could be used to avoid NHS staining. As shown above, both, the anchor and NHS labeling are likewise sparse and stochastic. Moreover, we could not detect any impact of the antibody labeling to our results. Thus, we assume that both labeling procedures could be used.

      Reviewer #2 (Public review):

      Summary:

      This study introduces a method that combines physical expansion of cells, imaging-guided isolation of defined regions, and protein identification to enable compartment-resolved analysis of protein composition at the subcellular scale. The authors aim to address a central limitation in existing approaches, namely the loss of spatial information during sample preparation or the indirect nature of proximity-based labeling methods. Using several cellular compartments as examples, they demonstrate that their approach can recover compartment-enriched protein sets and identify candidate proteins with previously unassigned localization.

      Strengths:

      A major strength of this work is the conceptual simplicity and accessibility of the approach. By combining established techniques in a modular way, the method avoids the need for genetic manipulation or specialized labeling strategies, making it broadly adaptable across experimental systems. The ability to directly select regions of interest based on imaging represents a clear advantage over indirect enrichment strategies and allows flexible targeting of both membrane-bound and non-membrane-bound compartments.

      The experimental design is also a strong aspect of the study. The use of complementary comparison strategies-analyzing isolated compartments alongside matched "subtracted" controls-provides an internal framework for assessing enrichment and depletion, increasing confidence in spatial assignment. The application of the method across multiple organelles of different sizes and properties demonstrates versatility, and the reported specificity for several compartments is encouraging. In particular, the ability to profile small and biochemically challenging structures highlights a potentially important niche for the approach.

      Weaknesses:

      Despite these strengths, several methodological limitations constrain the interpretation of the results. The most important relates to spatial accuracy in three dimensions. While lateral resolution is improved through physical expansion, the lack of depth resolution introduces uncertainty regarding contributions from structures above and below the selected region. Although the authors argue that this does not substantially affect specificity, the current evidence is largely indirect, and a more rigorous quantification of potential contamination would strengthen this conclusion.

      Quantitative interpretation also remains challenging. Because the measurements reflect total protein abundance rather than local concentration, differences in compartment size and protein density can influence enrichment values, particularly for small structures embedded within larger volumes. This issue is evident in the analysis of smaller compartments and complicates direct comparison across conditions. Additional normalization or modeling would help clarify how to interpret these measurements.

      Another limitation concerns variability in the expansion process and its downstream consequences. Differences in expansion factor across samples may affect the definition of regions of interest and introduce variability in sampling, yet the impact of this variability is not fully explored. Similarly, the use of a modified chemical treatment to preserve proteins for downstream analysis is central to the workflow but is not extensively validated with respect to preservation of spatial organization.

      While the identification of previously unannotated proteins is an appealing aspect of the study, validation is limited to a small number of examples, and broader support from independent datasets or literature context is lacking. In addition, the study primarily focuses on steady-state measurements in a single cell type, and therefore does not yet demonstrate the ability of the method to capture dynamic or condition-dependent changes in protein localization.

      Finally, the positioning of the method relative to existing approaches could be more clearly articulated. Although qualitative comparisons are provided, a more systematic and quantitative benchmarking against alternative strategies would help readers better understand the specific advantages and trade-offs.

      We thank the reviewer for the careful evaluation of the manuscript and for the constructive feedback. We think the reviewer raises valid points and will address them in the revised manuscript.

      Reviewer #3 (Public review):

      Franziscus et al. describe an elegant approach for spatially specific proteome analysis. To achieve this, they expand fixed cells and subsequently use a laser to micro-dissect a region of interest, which is then analyzed by mass spectrometry.

      They demonstrate the effectiveness of their approach by analyzing the nucleus, nucleolus, and the Golgi, and benchmark their hits against previous datasets for these organelles.

      The manuscript is very well written and nicely guides the reader through the applied methods. The presented data is convincing, and I do not see the need for additional experimental verification of the protocol. The only minor concern is the novelty of the method and the presentation. A combination of expansion, laser microdissection, and proteomics has been applied in the past (PMID: 36450705, PMID: 39477916). In the manuscript, one of these studies is cited, though it does not become clear that this approach is already described. However, Franziscus et al. describe the approach better and make it more accessible to the reader, especially since the other studies described this methodology in combination with tissue expansion and not in combination with single cell expansion as it is done here. I would ask the authors to be clearer in the introduction about what others have already done and what their contribution is here. In general, I am convinced that the community will benefit from the presented protocol to analyze organelle proteomics in detail.

      We thank the reviewer for the careful evaluation of our manuscript and overwhelmingly positive assessment. We apologize for the omission of the mentioned citations, and will adjust the introduction to make it clearer what has already been done and what the advance our method provides.

      References

      Damstra HG, Mohar B, Eddison M, Akhmanova A, Kapitein LC, Tillberg PW. 2022. Visualizing cellular and tissue ultrastructure using Ten-fold Robust Expansion Microscopy (TREx). eLife 11:e73775. DOI: https://doi.org/10.7554/eLife.73775

      Gambarotto D, Hamel V, Guichard P. 2021. Ultrastructure expansion microscopy (U-ExM). Methods in Cell Biology 161:57–81. DOI: https://doi.org/10.1016/bs.mcb.2020.05.006, PMID: 33478697

      Liffner B, Silva TLA e., Vega-Rodriguez J, Absalon S. 2024. Mosquito Tissue Ultrastructure-Expansion Microscopy (MoTissU-ExM) enables ultrastructural and anatomical analysis of malaria parasites and their mosquito. BMC Methods 1:13. DOI: https://doi.org/10.1186/s44330-024-00013-4

    1. eLife Assessment

      This valuable study investigates whether high-level physical reasoning is grounded in real-time bodily and vestibular signals using an innovative combination of virtual tool-use tasks and galvanic vestibular stimulation. The evidence is incomplete, as the main claims rely on limited and partially exploratory effects, including uncorrected multiple comparisons and cross-study comparisons that weaken the strength of the conclusions. The work, if it can be supported by clearer statistical support and more cautious interpretation, will be of interest to researchers in embodied cognition and physical reasoning.

    2. Reviewer #1 (Public review):

      Summary:

      This study investigates a fundamental question in cognitive science: is our ability to reason about the physical world an abstract mental process, or is it "embodied"-directly rooted in our real-time physical interactions with the environment? The authors compared participants' performance in computerized reasoning games with and without Galvanic Vestibular Stimulation (GVS). They suggest that participants failed more often and utilized suboptimal strategies under GVS compared to a sham stimulation condition. Furthermore, they found that this detrimental effect of GVS was reduced when the games were governed by altered gravity (hyper- and hypo-gravity). Consequently, the authors conclude that the physical experience of the body modifies high-level cognitive skills, such as reasoning.

      Strengths:

      The manuscript is well-written, organized, and easy to follow, making complex concepts accessible. Also, combining a specialized physical reasoning task with real-time vestibular disruption (GVS) is an intriguing approach to testing the boundaries of embodied cognition.

      Weaknesses:

      (1) Lack of Overall Effects and Inflated Type I Error for Game-Level Effects

      The study utilizes a within-subject design. Taking Study 1 as an example, each subject participated in a familiarization session (4 games), a baseline session (12 games without stimulation), a GVS session (14 games), and a sham session (14 games). No game was repeated for any single subject. Performance was quantified using three primary measures (success rate, number of attempts, and time per attempt) and two strategy measures (tool switching and the distance between tool placements).

      For Study 1, to identify condition differences at the game level (i.e., Figure 2), the authors effectively conducted 70 independent t-tests (5 measures × 14 games). While 7 significant results were reported, this large number of independent tests invites an inflated Type I error rate, as no multiple-comparison correction appears to have been applied.

      A similar inflation is expected in Study 2, where 50 independent t-tests (5 measures × 10 games) yielded 5 significant comparisons (Figure 4). Although the authors might argue the direction of the differences is systematic, implying GVS generally impairs performance, at least one significant comparison shows the opposite effect: tool switching indicates that GVS led to better performance for the 'Table_A' game in Study 2 (Figure 4d), whereas the same variable indicated GVS led to worse performance in Study 1 (Figure 2d). I suspect that none of the significant game-level results would survive a proper statistical correction. If possible, the authors can redo statistical testing with corrections (FDR or Bonferroni) or with LMM using game as a random effect. Before proper statistical analyses, I strongly encourage the authors to refrain from drawing broad conclusions based on these isolated game-level results.

      Furthermore, when analyzing data across all games, the study found no significant effect of GVS on overall performance or strategy measures in either Study 1 or Study 2. This lack of an aggregate effect contradicts the authors' conclusion that participants failed more often or utilized suboptimal strategies under GVS.

      (2) Missing Rationale for Classification Analysis

      It is puzzling why the authors pursued two exploratory analyses on tool placement after revealing that the two related primary measures (tool positioning and switching) did not generate significant condition differences in Study 1. These additional analyses-the Dirichlet Process Gaussian Mixture Model and leave-one-out classification-were not pre-registered. In the absence of overall condition differences, the authors appear to be "doubling down" by applying sophisticated classification tools to the raw data without a clear prior rationale.

      (3) Insufficient Evidence for the Reduced Effect of GVS Under Altered Gravity

      To compare Study 1 and Study 2, the authors devised a "gravity-weighted index," but its definition is not sufficiently justified. The index assigns weights of 1, 2, and 3 to low-, medium-, and high-gravity-dependent games, respectively. The choice of these specific weights appears arbitrary, making the quantitative results difficult to interpret. More importantly, there is no citation or explanation regarding how these three levels of "gravity impact" were defined in the first place (Line 468). This index was also not pre-registered.

      The authors state that for the success rate index, a value close to -1 indicates a large negative difference for GVS, 0 indicates no difference, and 1 indicates a large positive difference. These are theoretical bounds; the actual distribution of each index should be examined to validate such claims. However, the paper lacks descriptive statistics for this composite index.

      Notably, the "reduction" of the GVS effect in altered gravity was only demonstrated in one of the five available indices (success rate, p = 0.046). In fact, the success rate in Study 2 was 66.7(sham) vs 67.3 (GVS) in Table 2. It is highly debatable whether this marginal result justifies the conclusion that GVS effects "were reduced when the games included reasoning about altered gravity".

      (4) Questionable Assumptions Regarding Strategy

      The authors assume that "big changes in tool positioning and frequent tool switching indicate poor evaluation of the failed outcome". This assumption is questionable. In solving this cognitive task, participants must explore and exploit solutions based on feedback. Large shifts in positioning or frequent tool switching might reflect active, adaptive exploration based on failed outcomes rather than a failure to evaluate them.

      (5) Confounding Factors in GVS Interpretation

      The central theoretical question is whether physical reasoning is grounded in physical experience. GVS is used here to manipulate that experience. However, GVS does not selectively target the vestibular nerve; it also activates distributed fronto-parietal attention networks and hippocampal circuits essential for any reasoning task. Additionally, the vestibular system is linked to the limbic system and the cerebellum, which regulate emotional reactivity and arousal. Because attention and emotion are likely affected by GVS, the authors should be much more cautious in attributing their behavioral findings solely to changes in the "physical experience of the body."

    3. Reviewer #2 (Public review):

      Summary

      The paper investigates whether the real-time physical experience of the body shapes high-level physical reasoning. Participants played a set of computerized tool-use reasoning games (the Virtual Tools paradigm) in which they must use knowledge of physical laws - including gravity, collisions, and inertia - to guide a ball into a target area. In Study 1, participants played the games under terrestrial gravity while receiving either Galvanic Vestibular Stimulation (GVS), which introduces noise into the vestibular organ and disrupts gravitational signalling, or a Sham condition with matched skin sensation. In Study 2, a separate cohort played the same games redesigned under hypogravity (0.5 g - half Earth g) or hypergravity (2 g - double Earth g), again with concurrent GVS or Sham stimulation. Performance was assessed through success rate, number of attempts, and time per attempt; strategy was assessed through the spatial distance between successive tool placements and the frequency of tool switching across attempts. A post-hoc gravity-weighted index (GWI) was computed to compare the effect of vestibular perturbation across the two studies. The main finding is that GVS impairs performance in gravity-dependent games under terrestrial gravity, yet the same perturbation appears to be neutral or even beneficial when the game environment involves non-terrestrial gravity - a result the authors interpret as evidence for an adaptable, body-grounded internal model of physics.

      Strengths

      One of the most notable strengths of this work is its conceptual positioning at the intersection of embodied cognition and physical reasoning. Rather than treating the human body either as an abstract information-processing device or as a purely biomechanical system, the authors take seriously the idea that cognition is scaffolded by ongoing sensorimotor state - and they test this idea with a paradigm that is both tractable and theoretically motivated. The use of the Virtual Tools paradigm is well-suited to this goal: the games vary systematically in their reliance on gravitational predictions, allowing selective impairment (rather than general disruption) to serve as a signature of embodied physical reasoning.

      The dual-study design is another strength. Testing the same vestibular perturbation under terrestrial and altered game-gravity conditions, and observing a reversal in its effect depending on context, provides a form of internal control that is conceptually compelling. The additional clustering analyses (Dirichlet Process Gaussian Mixture Model and leave-one-out kernel density classification) strengthen the strategy results beyond raw distance measures, confirming that GVS systematically shifts participants' spatial exploration strategies.

      The paper is also clearly written and engages meaningfully with relevant theoretical frameworks - predictive coding, embodied cognition, and stochastic resonance - making it accessible and stimulating for a broad audience.

      Weaknesses

      (1) Absence of multiple-comparisons correction. A large number of game-level pairwise t-tests are conducted in both studies (upward of twenty per study) without correction for familywise error rate. The game-level effects that anchor the main narrative - in Study 1 alone: Remove, GoalMove, Spiky, Falling_A, Shafts_B, Gap, and Chaining - arise from an uncorrected pool of comparisons. The probability that some of these constitute false positives is non-trivial. The authors should apply a correction (e.g., Benjamini-Hochberg) or at a minimum discuss this limitation explicitly.

      (2) The facilitation claim rests on a post-hoc and arbitrarily parameterized index. The gravity-weighted index (GWI), which drives the central cross-study comparison, uses integer coefficients (1, 2, 3) to weight games by gravity dependency level. These coefficients are entirely arbitrary and bear no principled relationship to the actual gravitational magnitudes used in the study. Why not use the gravity dependency ratings themselves, or the empirically estimated gravity impact scores from the computational modelling mentioned in the Methods? The choice of weights should be either principled or tested across a range of values to demonstrate robustness. Furthermore, the notation in equation (1) as currently typeset reads as "Gravity minus Weighted Index" rather than "Gravity-Weighted Index"; this should be corrected.

      (3) The "facilitation" interpretation exceeds what the data in Study 2 directly support. Across all games in Study 2, GVS versus Sham differences in absolute performance are non-significant in all directions. The facilitation claim derives entirely from the GWI being higher in Study 2 than in Study 1 - a between-subjects comparison involving different participant groups and a non-pre-registered metric. The language of "facilitation" should be tempered accordingly, or the authors should provide additional analyses to support this framing.

      (4) Gravitational manipulation is visual only, and the vestibular system is only one component of the gravity-sensing network. Gravity perception results, as the authors very well know, from a distributed multisensory integration process that involves, in addition to the vestibular system, visual, proprioceptive, and visceral inputs. The present paradigm manipulates gravitational context solely through visual cues and targets the vestibular system through GVS - a point the authors acknowledge but do not discuss in sufficient depth. It is important to distinguish clearly between real gravitational alterations (as achieved in parabolic flight or centrifuge environments, where the entire body is physically subjected to a different gravitational vector) and virtually altered gravity, where only one sensory modality is targeted while others remain anchored to 1 g. The scope of the conclusions should reflect this distinction.

      (5) The choice of 0.5 g and 2 g may lack sensitivity. Combining the two altered-gravity conditions in Study 2, because no significant effect of hypo versus hypergravity was found, is statistically pragmatic but conceptually unsatisfying. There is evidence in the space physiology literature that gravitational processing is not linearly symmetric around 1 g: threshold effects exist below and above terrestrial gravity that may not be captured by modest deviations (half and double g) - see refs below. It is worth discussing whether the absence of a hypo/hyper distinction in Study 2 reflects a genuine equivalence or a lack of sensitivity, and whether more extreme conditions (e.g., near-zero g or 4-5 g) might reveal different processing regimes. Whether 0.5 g and 2 g were sufficient to saturate the system or merely insufficient to perturb it remains an open question with direct implications for the interpretation of the null GWI effects on strategy measures.

      Lee SMC, Ribeiro LC, Martin DS, Zwart SR, Feiveson AH, Laurie SS, Macias BR, Crucian BE, Krieger S, Weber D, Grune T, Platts SH, Smith SM, and Stenger MB. Arterial structure and function during and after long-duration spaceflight. J Appl Physiol (1985) 129: 108-123, 2020.

      de Winkel KN, Clément G, Groen EL, and Werkhoven PJ. The perception of verticality in lunar and Martian gravity conditions. Neurosci Lett 529: 7-11, 2012.

      Clément G, Moore ST, Raphan T, and Cohen B. Perception of tilt (somatogravic illusion) in response to sustained linear acceleration during spaceflight. Exp Brain Res 138: 410-418, 2001.

      Benson AJ, Kass JR, and Vogel H. European vestibular experiments on the Spacelab-1 mission: 4. Thresholds of perception of whole-body linear oscillation. Exp Brain Res 64: 264-271, 1986.

      (6) High-level reasoning is not defined with sufficient precision. The term "high-level reasoning" appears from the title onward and in the heading of the Study 1 results section (line 138), but it is never formally defined. The reader needs a clearer account of what distinguishes high-level physical reasoning from low-level sensorimotor prediction, and where the games used here fall along that continuum. What specific physical competencies - ballistic trajectories, free-fall predictions, collision dynamics, frictional forces, inertial effects - are required across the game set? When describing the subset of games that drive key effects, this information is critical for evaluating whether effects are specific to gravity reasoning or to some other physical concept.

      (7) Performance measures are disconnected from underlying kinematics. The performance measures (success rate, number of attempts, time per attempt) are coarse, high-level summaries. Time per attempt is used as a proxy for performance efficiency, yet participants received no instructions regarding speed, and different individuals may have adopted systematically different speed-accuracy trade-offs. It would be valuable to know whether time per attempt correlates with attempt number within a given game (which would indicate within-game learning) and whether mouse movement data - trajectory, velocity, hesitation - were recorded and could be analysed to provide more mechanistic insight into strategy formation.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript investigates a theoretically important question in cognitive science: whether higher-level physical reasoning is an abstract, modular process or is grounded in real-time body-environment interactions. To address this question, the authors combine galvanic vestibular stimulation (GVS) with the Virtual Tools task to test whether perturbing vestibular gravity signals affects performance in physical reasoning. The study is conceptually innovative and has the potential to bridge embodied sensory processing and higher-level cognition. However, in its current form, the evidence only partially supports the main claims, and several aspects of the analysis and interpretation limit the strength of the conclusions.

      Strengths:

      A major strength of the manuscript is the originality of the experimental paradigm. The combination of galvanic vestibular stimulation (GVS), which perturbs gravity-related vestibular signals, with computerized game-based tasks that require physical reasoning provides a novel way to test whether ongoing bodily experience influences higher-level cognition. Conceptually, the study is highly original and meaningfully bridges two domains that are often studied separately: sensorimotor processing and higher-level cognition.

      Weaknesses:

      The main weakness of the manuscript is that its central conclusion is not strongly supported by the data. The key finding depends on a marginally significant cross-study comparison, whereas direct GVS-versus-Sham differences in Study 2 are minimal across aggregate measures. In addition, many game-level analyses involve a large number of uncorrected multiple comparisons, raising the possibility that some of the reported effects may reflect chance findings. The manuscript's most important metric, the Gravity-Weighted Index, was not preregistered and is exploratory in nature, yet it is treated as a primary basis for confirmatory conclusions. The cross-study comparison is also difficult to interpret because the two studies differ in participant samples, number of games, and partially in the stimulus set. Finally, the mechanistic claims in the Discussion-particularly those invoking predictive coding, stochastic resonance, or updating of internal gravity models-go well beyond what can be directly inferred from the present behavioral data. Overall, the study provides intriguing but limited evidence that vestibular signals may influence some physical reasoning tasks under specific conditions, rather than strong evidence for a broad account of physical reasoning as grounded in online vestibular processing

    1. eLife Assessment

      In this solid work, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role for two conserved acidic residues rather than a single one. This useful study used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors study two residues in the GHKL ATPase active site of Aq MutL and GyrB, and argue that the catalytic base function is shared between two conserved acidic residues that are 3 residues apart.

      They generated mutant versions in MutL and GyrB (both ala and the appropriate Asn/Gln version) and performed ATPase analysis. They also generated high-resolution crystal structures of the GyrB NTD with AMPPnP for WT and mutants of the two acidic residues. The data show that mutation in either of these residues does not fully kill activity (with the exception of the Alanine mutation of the first of the two, which interferes with ATP (or AMPPnP) binding). When the acidic residues are mutated to Asn/Gln, the catalytic water can still be positioned, and hence these mutants are more active than the Ala mutants. In both cases, the double mutation is catalytically dead.

      The authors then perform phylogenetic analysis and ancestral gene reconstruction, and based on this, they argue that HSP90 forms a different class of GHKL ATPases, and lost rather than gained this separate status.

      Strengths:

      The biochemical analysis seems solid.

      Weaknesses:

      (1) A major question that remains is why the mutations have so much more detrimental effect in MutL (100-fold lower kcat/KM) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      (2) The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the acidic residues and their mutants is not shown.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role of two conserved acidic residues rather than one. The authors have used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism.

      Through a detailed re-analysis of their previously published structure of the aqMutL NTD (ATPase domain) in complex with AMPPCP, they identified Glu29 and Glu32 as interacting with nucleophilic water for the catalysis. The authors carefully dissected the respective roles of these two acidic residues with a series of site-directed mutations. Mutations at Glu29 impaired ATPase activity without affecting protein secondary structure or ATP binding in the case of the E29Q mutant. Moreover, mutations at Glu32 did not affect secondary structure (except for E32G) but reduced ATPase activity. Activity was abolished when both residues (E29Q/E32Q) were mutated.

      The authors extended their study to another GHKL ATPase, aqGyrB. Their findings further supported the cooperative function of the corresponding acidic residues in aqGyrB (Glu48 and Asp51) during ATP hydrolysis. Mutation of these residues partially impaired ATP hydrolysis without affecting protein secondary structure. ATPase activity was completely lost in the double mutant E48Q/D51M. While the E48Q mutant retained the ability to bind ATP, the E48A mutant did not. High-resolution structures of the WT and E48A, E48Q, D51A, and D51N mutants of the aqGyrB NTD demonstrated that nucleophilic water positioning depended on these residues. E48 played a dominant role in water positioning and is critical for stabilising ATP lid formation and associated conformational changes, whereas D51 contributed cooperatively to catalysis.

      The authors investigated the functional impact of mutating the corresponding residues in the human MutL homologs PMS2 and MLH1. Clinical variants consistently exhibited reduced or abolished ATPase activity, providing a potential molecular basis for Lynch syndrome through impaired DNA mismatch repair.

      Lastly, through evolutionary analysis, the authors inferred that the second acidic residue was likely present in the common ancestor of MutL, GyrB, and MORC proteins, but was lost in the case of Hsp90.

      Strengths:

      (1) This study contains a detailed structural and biochemical analysis of a biologically important set of GHKL ATPases. The authors identify a second acidic residue that is conserved and contributes to catalysis in a large subset of GHKL ATPases. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed, which involves cooperative and partially overlapping roles for the catalytic residue pair. This revised mechanistic model is invaluable for the interpretation of clinical variants of GHKL ATPases such as PMS2 and MLH1.

      (2) The work described was performed to an excellent and rigorous technical standard. The structural and biochemical data are sound. The evidence supporting the claims is compelling.

      Weaknesses:

      (1) The identification in this study of a second acidic residue contributing to catalysis but not absolutely essential for catalysis is a useful finding. However, given that many structures of GHLK ATPases have been determined with different nucleotide analogs bound and that the essential role of the first acidic residue is well established, the importance and scope of the advances described here remain focused within the field of study of GHKL ATPases.

      (2) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of the effect (or likely effect) of disease-linked mutations in GHKL ATPases would have strengthened this study.

      (3) In MLH1, the E37K mutation completely abolishes ATPase activity, but the corresponding mutations in aqMutL, aqGyrB, and PMS2 do not. It remains unclear why E37K in MLH1 leads to complete loss of activity, as the authors propose that water molecule positioning via the first acidic residue, as well as ATP lid stabilisation and associated conformational changes, should still be possible.

      (4) The authors do not examine ATP binding in the E32 mutants of aqMutL NTD and the D51 mutants of aqGyrB, or AMPPNP binding of the NLH1 and PMS2 mutants. Hence, the relative contributions of the acidic residues to ATP binding and hydrolysis remain partially unclear.

      (5) The ATPase assays for PMS2 and MLH1 (Figure 7 and Table 1) were performed with purification/solubility tags still present. Hence, it cannot be ruled out that these tags influence the measured activities.

      (6) The authors suggest that the two-acidic-residue mechanism proposed in this study could be shared among several GHKL ATPase families, yet they also state that the hydrogen-bonding network was not observed in MutL and MORC family proteins. This raises doubt about how conserved the mechanism is, e.g., in MutL and MORC proteins.

    1. eLife Assessment

      This valuable study highlights the key role of NK cells and PD-L1+ neutrophils in worsening sepsis responses in the context of MASH (metabolic dysfunction-associated steatohepatitis). It focused on the role of neutrophils in mediating this effect, which is based on a choline-deficient high-fat diet model of various knockouts or selective ablation of immune cell types. While the data presented are of great interest, there are concerns around the reliability of the strength of the evidence provided, which is currently considered incomplete. The study may be of interest to researchers in immunopathological disease mechanisms once confirmatory studies have been completed.

      [Editors' note: the authors no longer have access to the original flow cytometry data and plan to compile new datasets in the future.]

    2. Reviewer #1 (Public review):

      Summary:

      By using an established NAFLD model, choline-deficient high-fat diet, Barros et al show that LPS challenge causes excessive IFN-γ production by hepatic NK cells which further induces recruitment and polarization of a PD-L1 positive neutrophil subset leading to massive TNFα production and increased host mortality. Genetic inhibition of IFN-γ or pharmacological blockade of PD-L1 decreases recruitment of these neutrophils and TNFα release, consequently preventing liver damage and decreasing host death.

      Since NAFLD is often accompanied by chronic, low-grade inflammation, it can lead to an overactive but dysfunctional immune response and increase the body's overall susceptibility to infections, therefore this is very important research question.

      Strengths:

      The biggest strength of the manuscript is vast number of mouse strains used.

      Weaknesses:

      After the review, there are still some open questions from my side:

      (1) I would like the authors to defend their choice of diet type since this has not been done in the review/response to authors. In case they cannot, we need additional proof (HFD or WD model).

      (2) Since the authors used same control groups (chow and HFCD), as required by the animal ethics committee, they must have power analysis test to show that the number of controls (but also in other groups) they used is enough to see the effect. Please provide it.

    3. Reviewer #2 (Public review):

      Summary:

      This is an extremely interesting mouse study, trying to understand how sepsis is tolerated during obesity/NAFLD. The researchers combine a well-established model of NASH (Choline-deficiency with High Fat Diet) with a sepsis model (IP injection of 10mg/kg LPS), leading to dramatic mortality in mice. Using this model, they characterize the complex contributions of immune cells. Specifically, they find that NK-cells and Neutrophils contribute the most to mortality in this model due to IFNG and PD-L1+ Neutrophils.

      Strengths:

      The biggest strength of the manuscript is how clear the primary phenotypes/endpoints of their model are. Within 6 hours of LPS injection, there is a stark elevation of liver inflammation and damage, which is exacerbated by a High Fat/CholineDeficient diet (HFCD). And after 1 day, almost all of the mice die. Using these endpoints, the authors were able to identify which cells were critical for mortality in the model and the specific mediators involved.

      Comments on revisions:

      I have no further comments.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We thank the editor and reviewers for their constructive questions, valuable feedback, and for approving our manuscript. We truly appreciate the opportunity to improve our work based on their insightful comments. Before addressing the editor’s and each referee’s remarks individually, we provide below a point-by-point response summarizing the revisions made.

      Duplication of control groups across experiments

      We appreciate the reviewers’ concern regarding the potential duplication of control groups. In the revised manuscript, we have explicitly clarified that independent groups of control mice were used for each experiment. These details are now clearly indicated in the Materials and Methods section to avoid any ambiguity and to reinforce the rigor of our experimental design (Page 15, Line 453-455): “Furthermore, knockout animals and those treated with pharmacological inhibitors or neutralizing antibodies shared the same control groups (chow and HFCD), as required by the animal ethics committee.”

      Validation of the MASLD model

      To strengthen the metabolic characterization of our MASLD model, we have now included additional parameters, including liver weight, Picrosirius staining and blood glucose measurements. These data are presented as new graphs in the revised manuscript and support the metabolic relevance of the HFCD diet model (Figure Suplementary S1). The corresponding description has been added to the Results section (Page 5, Lines 116-117) as follows: “Mice fed HFCD showed no increase in liver weight and collagen deposition as evidenced by Picrosirius staining (Fig. S1A and Fig. S1C)”

      Assessment of liver injury in RagKO and anti-NK1.1 mice

      We fully agree that assessment of liver injury is essential for these models. For mice treated with antiNK1.1, ALT levels are shown in Figure 4G, confirming increased liver injury after treatment. Regarding Rag⁻/⁻ mice, the animals exhibit exacerbation of liver injury when fed a HFCD diet and challenged with LPS (Page 7, Lines 183–184). The corresponding description has been added to the Results section (Page 7, Lines 175-176) as follows: “Interestingly, Rag1-deficient animals under the HFCD remained susceptible to the LPS challenge (Fig. 4C) with exacerbation of liver injury (Fig. 4D) ”

      Discussion of limitations

      We have expanded the Discussion section to provide a more comprehensive and balanced perspective on the limitations of our model and experimental approach (Page 13-14, Lines 401–414) “Our study presents several limitations that should be acknowledged and discussed. First, we cannot entirely rule out the possibility that our mice deficient in pro-inflammatory components exhibit reduced responsiveness to LPS. However, our ex vivo analyses using splenocytes from these animals revealed a preserved cytokine production following LPS stimulation. These results suggest that the in vivo differences observed are primarily driven by the MAFLD condition rather than by intrinsic defects in LPS sensitivity. Second, the absence of publicly available single-cell RNA-seq datasets from MAFLD subjects under endotoxemic or septic conditions limited our ability to perform direct translational comparisons. To overcome this, we analyzed existing MAFLD patients and experimental MAFLD datasets, which consistently demonstrated upregulation of IFN-y and TNF-α inflammatory pathways in MALFD. In line with these findings, our murine model revealed TNF-α⁺ myeloid and IFN-y⁺ NK cell populations, thereby reinforcing the validity and translational relevance of our results.”. This revision highlights the constraints of the MASLD model, the inherent variability among in vivo experiments, and the interpretative limitations related to immunodeficient mouse strains.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 4 the authors are showing the number of IFN+ positive CD4, CD8, and NK 1.1+ cells. Could they show from total IFNg production, how much it goes specifically on NK cells and how much on other cell populations since NK1.1 is NK but also NKT and gamma delta T cell marker? Also, in Figure 2E the authors see a substantial increase in IFNg signal in T cells.

      While we did not specifically assess IFNγ production in NKT cells or other minor populations, our data indicate that the NK1.1+CD3+ cells (NKT cells) cited in Page 7, Lines  188-192 were essentially absent in the liver tissue of LPS-challenged animals, as shown in Supplementary Figures 3C and S10. The corresponding description has been added to the Results section (Page 7, Lines 188-192) as follows: “We observed that the number of NK cells increased in the liver tissue of PBS-treated MAFLD mice compared with mice fed a control diet (Fig. 4E). LPS challenge increased the accumulation of NK1.1+CD3− NK cells in the liver tissue of MAFLD mice and the absence of NK1.1+CD3+ NKT cells (Fig. S3C and 4E)”.

      This absence was consistent across all experimental conditions, corroborating our focus on NK1.1+CD3− cells as the primary source of NK1.1-associated IFNγ production. Furthermore, data demonstrated in Figure 2E illustrate the presence of IFNγ primarily in NK cells. Therefore, the observed IFNγ signal, attributed to NK1.1+ cells, predominantly reflects conventional NK cells, with minimal contribution from NKT or γδ T cells.

      (2) In Figure 4C, the authors state that the results suggest that T and B cells do not contribute to susceptibility to LPS challenge. However, they observe a drop in survival compared to chow+LPS. Are the authors certain there is no statistical significance there?

      The observed decrease in survival is consistent with our expectations, as T and B cells are not the primary source of interferon-gamma (IFNγ) in this context. Even in their absence, animals remain susceptible to LPS challenge due to the presence of other IFNγ-producing cells that drive the observed lethality. We have carefully re-examined the statistical analysis and confirm that it was correctly performed.  

      (3) Since the survival curve and rate are exactly the same (60%) in Figures 3F, 3G, 4C, 4F, 5G, and 5H I would just like to double-check that the authors used different controls for each experiment.

      The number of mice used in each experiment was carefully determined to ensure sufficient statistical power while fully complying with the limits established by our institutional Animal Ethics Committee. To minimize animal use, the same control group was shared across multiple survival experiments. Despite using shared controls, the total number of animals per experimental group was adequate to produce robust and reproducible survival outcomes. All groups were properly randomized, and the shared control data were rigorously incorporated into statistical analyses. This strategy allowed us to maintain both ethical standards and the scientific rigor of our findings.

      (4) In Figure 5 the authors are saying that it is neutrophils but not monocytes mediate susceptibility of animals with NAFLD to endotoxemia. However, CXCR2i depletion and CCR2 knock out mice affect both monocytes/macrophages and neutrophils. And in Figures 5E, 5G, and 5H they see that a) LPS+CXCR2i decreases liver damage more than LPS+anti Ly6G, b) HFCD mice challenged with LPS and treated with anti-LY6G do not rescue survival to levels of CHOW LPS and c) anti Ly6G treatment helps less than CXCR2i. Therefore, from both knock out mice and depletion experiments the authors can conclude that most likely monocytes (but potentially also other cells) together with neutrophils are substantial for the development of endotoxemic shock in choline-deficient high-fat diet model.

      While neutrophils express CCR2, our data clearly show that CCR2 deficiency does not impair neutrophil migration, as demonstrated in Supplemental Figures 5A and 5B (added to the manuscript, page 8, lines 213–217). The corresponding description has been added to the Results section (Page 8, Lines 213217) as follows: ``Interestingly, animals deficient in monocyte migration (CCR2-/-) showed a high mortality rate compared to wild type after LPS challenge and neutrophil migration is not altered (Fig. 5SA and Fig. 5SB)``, In contrast, CCR2 deficiency primarily affects monocyte recruitment, yet in our experimental conditions, monocyte depletion or CCR2 knockout did not significantly alter the severity of endotoxemic shock, indicating that monocytes play a minimal role in mediating susceptibility in HFCD-fed mice.

      To specifically investigate neutrophils, we used pharmacological blockade of CXCR2 to inhibit migration and antibody-mediated neutrophil depletion. Both approaches have consistently demonstrated that neutrophils are critical factors in endotoxemic shock.

      These findings support our conclusion that neutrophils are the primary cellular contributors to susceptibility in HFCD-fed mice during endotoxemia, with monocytes making a negligible contribution under the tested conditions.

      (6) In Figure 6A (but also others with PD-L1) did the authors do isotype control? And can they show how much of PD1+ population goes on neutrophils, and how much on all the other populations?

      To address this issue, we performed additional analyses to assess the distribution of PD-L1 expression on CD45+CD11B+ leukocytes. These new results, detailed on Page 9, lines 245-250, and now presented in Supplemental Figure 6, demonstrate that PD-L1 expression is predominantly enriched in neutrophils compared to other immune subsets. This observation further reinforces our conclusion that neutrophils represent a major source of PD-L1 in our experimental model.

      To ensure the robustness of these findings, we also included FMO controls for PD-L1 staining in the newly added Supplemental Figure S6. These controls validate the specificity of our gating strategy and confirm the reliability of the detected PD-L1 signal. The corresponding description has been added to the Results section (Page 9, Lines 245-250) as follows: ``First, we observed that only the MAFLD diet caused a significant increase in PD-L1 expression in CD45+CD11b+ leukocytes after LPS challenge (Fig. S6C). We observed that within this population, neutrophils predominate in their expression when compared to monocytes (Fig. 6SA, Fig. 6SB, and Fig. 6SD). Furthermore, PD-L+1 neutrophils showed an exacerbated migration of PD-L1+ neutrophils towards the liver (Fig. 6A and 6B)”

      (7) In Figure 6D it is interesting that there is not an increase in PD-L1+ neutrophils in LPS HFCD IFNg+/+ mice in comparison to LPS chow IFNg+/+ mice, since those should be like WT mice (Figure 6A going from 50% to 97%) and so an increase should be seen?

      The apparent difference between Figures 6A and 6D likely reflects inter-experimental variability rather than a biological discrepancy. Although the absolute percentages of PD-L1⁺ neutrophils varied slightly among independent experiments, the overall phenotype and trend were consistently maintained namely, that PD-L1 expression on neutrophils is enhanced in response to LPS stimulation and modulated by IFNγ signaling. Thus, the data shown in Figure 6D are representative of this consistent phenotype despite minor quantitative variation.

      (8) In Figure 7 do the authors have isotype control for TNFa because gating seems a bit random so an isotype control graph would help a lot as supplementary information, in order to make the figure more persuasive

      To address the concern regarding gating in Figure 7, we have included the FMO showing TNFα as a histogram Supplementary Figure 8gG. These control reaffirm the accuracy and reliability of our gating strategy for TNFα, further supporting the robustness of our data. The corresponding description has been added to the Results section (Page 9, Lines 272-274) as follows:`` We observed an exacerbated TNF-α expression by PD-L1+ neutrophils from MAFLD when compared to control chow animals (Fig. 7A, Fig. 7B, Fig. 7D, and Fig8SG).

      (9) Figure 6C IFNg+/+ mice on CHOW +LPS is same as Figure 8E mice chow +LPS but just with different numbers. Can the authors explain this?

      Although the data points in Figures 6C and 8E may appear similar, we confirm that they originate from entirely independent experiments and represent distinct datasets. To enhance clarity and avoid any potential confusion, we have adjusted the figure presentation and sizing in the revised manuscript. These changes make it clear that the datasets, while comparable, are derived from separate experimental replicates.

      (10) Figure 1E chow B6+LPS is the same as Figure 5D B6+LPS but should they be different since those should be two different experiments?

      We confirm that Figures 1E and 5D correspond to data obtained from independent experiments. Although the experimental conditions were similar, each dataset was generated and analyzed separately to ensure the reproducibility and robustness of our results.

      Reviewer #2 (Recommendations for the authors):

      (1) Why did you look at kidney injury in Figure 1D? I think this should be explained a little.

      We assessed kidney injury alongside ALT, a marker of liver damage, because both the liver and kidneys are among the primary organs affected during sepsis and endotoxemia. This rationale has been added to the manuscript (page 5, lines 129–131): “Remarkably, compared to the Chow group, HFCD mice exposed to LPS did not show greater changes in other organs commonly affected by endotoxemia, such as the kidneys (Figure 1D).” By evaluating markers of injury in both organs, we aimed to determine whether our physiopathological condition was liver-specific or indicative of broader systemic injury.

      (2) I know Figure 2C isn't your data, but why are there so few NK cells, considering NK cells are a resident liver cell type? Doesn't that also bring into question some of your data if there are so few NK cells? And the IFNG expression (2E) looks to mostly come from T-cells (CD8?).

      The data shown in Figure 2C were reanalyzed from a separate NAFLD model based on a 60% high-fat diet. Although this model differs from ours, the observed low number of NK cells is consistent with expectations for animals subjected solely to a hyperlipidic diet, which primarily provides an inflammatory stimulus that promotes recruitment rather than maintaining high baseline NK cell numbers.

      In our experimental model, these observations align with published data. Specifically, liver tissue from NAFLD animals typically exhibits low baseline NK cell numbers, but upon LPS challenge, there is a marked increase in NK cell recruitment to the liver. This dynamic illustrates the interplay between dietinduced inflammation and immune cell recruitment in our experimental context and supports the interpretation of our IFNγ data.

      (3) In your methods, I think you didn't explain something. You said LPS was administered to 56 week old mice, but that HFCD diet was started in 5-6 week old mice and lasted 2 weeks, then LPS was administered. So LPS administration happened when the mice were 7-8 weeks old, right?

      We thank the reviewer for pointing out this inconsistency in our Methods section. The reviewer is correct: the HFCD diet was initiated in 5–6-week-old mice, and LPS was administered after 2 weeks on the diet, such that LPS challenge occurred when the mice were 7–8 weeks old.

      We have revised the Methods section (add page 15-16, lines 474–480).  to clarify this timeline and ensure it is accurately described in the manuscript. The corresponding description has been added to the Materials and Methods section (Page 14, Lines 436-442) as follows: “Lipopolysaccharide (LPS; Escherichia coli (O111:B4), L2630, Sigma-Aldrich, St. Louis, MO, USA) was administered intraperitoneally (i.p.; 10 mg/kg) in C57BL/6, CCR2 -/-, IFN-/-, and TNFR1R2 -/- mice. The HFCD was initiated in 5–6 week-old mice, and LPS was administered after 2 weeks on the diet, meaning that LPS administration occurred when the mice were 7–8 weeks old, with body weights ranging from 22 to 26 g. LPS was previously solubilized in sterile saline and frozen at -70°C. The animals were euthanized 6 hours after LPS administration”.

      (4) Throughout the manuscript, I would consider changing the term NAFLD to something else. I think HFCD diet is a closer model to NASH, so there needs to be some discussion on that. And the field is changing these terms, so NAFLD is now MASLD and NASH is now MASH.

      We appreciate the reviewer’s comment regarding the terminology and disease classification. In our experimental conditions, the animals were subjected to a high-fat, choline-deficient (HFCD) diet for only two weeks, a period considered very early in the progression of diet-induced liver disease. At this stage, histological analysis revealed lipid accumulation in hepatocytes without evidence of hepatocellular injury, inflammation, or fibrosis. Therefore, our model more closely resembles the metabolic-associated fatty liver disease (MAFLD, formerly NAFLD) stage rather than the more advanced metabolic-associated steatohepatitis (MASH, formerly NASH).

      Indeed, prolonged exposure to HFCD diets, typically 8 to 16 weeks, is required to induce the inflammatory and fibrotic features characteristic of MASH. Since our objective was to study the initial metabolic and immune alterations preceding overt liver injury, we believe that using the term MAFLD more accurately reflects the pathological stage represented in our model. Accordingly, we have revised the text to align with the updated nomenclature and disease context.

      (6) I am concerned about over interpretation of the publicly available RNA-seq data in Figure 2. This data comes from human NAFLD patients with unknown endotoxemia and mouse models using a traditional high-fat diet model. So it is hard to compare these very disparate datasets to yours. Also, if these datasets have elevated IFNG, why does your model require LPS injection?

      We thank the reviewer for their thoughtful comments regarding the interpretation of the RNA-seq data presented in Figure 2. We would like to clarify that the human NAFLD datasets referenced in our study do not specifically include patients with endotoxemia; rather, they focus on individuals with NAFLD alone.

      Comparing data from human and murine MAFLD models, we observed that NK cells, T cells, and neutrophils are present and contribute to the hepatic inflammatory environment. Our reanalysis indicates that the elevations of IFNγ and TNF in NAFLD are primarily derived from NK cells, T cells, and myeloid cells, respectively.

      In our experimental model, LPS administration was used to evaluate whether these immune populations particularly NK cells are further potentiated under a hyperinflammatory state, leading to exacerbated IFNγ production. This approach allows us to determine whether increased IFNγ contributes to worsening outcomes in NAFLD, providing mechanistic insights that cannot be obtained from static human or traditional mouse datasets alone.

      (7) The zoom-ins for the histology (for example, Figure 1E) don't look right compared to the dotted square. The shape and area expanded don't match. And the cells in the zoom-in don't look exactly the same either.

      We have thoroughly re-examined the histological sections and the corresponding zoom-ins, including the example in Figure 1E. Upon verification, we confirm that the zoom-ins accurately represent the highlighted areas indicated by the dotted squares. The apparent discrepancies in shape or cellular appearance are likely due to minor differences in orientation or cropping during figure preparation. Nevertheless, the content and regions depicted are consistent with the original sections.  

      (8) Did the authors measure myeloid infiltration in the CCR2-/- mice? Did you measure Neutrophil infiltration in the TNF-Receptor KO mice?

      Analysis of CD45+ cell migration in CCR2 knockout mice, as shown in Supplemental Figure 5C and 5D, demonstrates that the absence of CCR2 does not impair overall leukocyte migration. Similarly, assessment of neutrophil migration in TNF receptor (TNFR1/2) knockout mice, presented in Supplemental Figure 8A, shows that neutrophil trafficking is not affected in these animals. These results indicate that the respective knockouts do not compromise the migration of the analyzed immune populations, supporting the interpretations presented in our study.

      (9) Regarding Methods for RNA-seq Analysis. Was the Mitochondrial percentage cutoff 0.8%, because that seems low. And was there not a Padj or FDR cutoff for the differential expression?

      The mitochondrial percentage in our scRNA-seq analysis reflects the proportion of mitochondrial gene expression per cell, which serves as a quality control metric. A low mitochondrial gene expression percentage, such as the 0.8% cutoff used here, is indicative of highly viable cells.

      For differential gene expression analysis, we employed the FindMarkers function in Seurat with standard parameters: adjusted p-value (Padj) < 0.05 and log2 fold change > 0.25 for upregulated genes, and adjusted p-value < 0.05 with log2 fold change < -0.25 for downregulated genes. These thresholds ensure robust identification of differentially expressed genes while balancing sensitivity and specificity.

      (10) Regarding Methods for Flow Cytometry. How were IFNG and TNF staining performed? Was this an intracellular stain? Did you need to block secretion? TNF and IFNG antibodies have the same fluorophore (PE), so were these stainings and analyses performed separately?

      Six hours after LPS challenge, non-parenchymal liver cells were isolated using Percoll gradient centrifugation. Because the animals were in a hyperinflammatory state induced by LPS, no in vitro stimulation was performed; all staining was carried out immediately after cell isolation. Detection of IFNγ and TNF was performed via intracellular staining using the Foxp3 staining kit (eBioscience). Due to both antibodies being conjugated to PE, IFN-γ and TNF-α staining and analyses were conducted in separate experiments. These distinct staining protocols and analyses are detailed in Supplemental Figures 10 and 11. The corresponding description has been added to the Materials and Methods section (Page 16, Lines 490-493) as follows: ``As animals were already in a hyperinflammatory state, no additional in vitro stimulation was required. Intracellular detection of IFN-γ and TNF-α was conducted using the Foxp3 staining kit (eBioscience). Since both antibodies were conjugated to PE, staining and analyses were performed in separate experiments``

      Reviewer #3 (Recommendations for the authors):

      (1) Achieving an NAFLD model/disease is the starting point of this study. I understand that a two-week HFCD diet period was applied due to the decrease in lymphocyte numbers. Was it enough to initiate NAFLD then? Or is it a milder metabolic disease? Which parameters have been evaluated to accept this model as a NAFLD model?

      Indeed, the two-week HFCD diet induces an early-stage form of NAFLD, characterized by initial fat accumulation in the liver without significant hepatic injury. While this represents a milder metabolic phenotype, it is sufficient to study the inflammatory and immune responses associated with NAFLD. To validate this model, we assessed multiple parameters: liver weight, blood glucose levels, and collagen deposition. These measurements confirmed the presence of early-stage NAFLD features in the animals, providing a relevant and reliable context for investigating susceptibility to endotoxemia and immune cell dynamics. They are shown in Figure Suplementary 1 and the text was included in the manuscript (Page 5, Lines 116-117): “Mice fed HFCD showed no increase in liver weight and collagen deposition as evidenced by Picrosirius staining (Fig. S1A and Fig. S1C) ”.

      (2) It is true that the CD274 gene (encoding PD-L1) and the IFNGR2 gene, corresponding to the IFNγ receptor, are among the upregulated genes when authors analyzed the publicly available RNAseq data but they are not the most significantly elevated genes. What is the reasoning behind this cherrypicking? Why are other high DEGs not analyzed but these two are analyzed?

      We highlighted the expression of the IFN-γ receptor (IFNGR2) and CD274 (encoding PD-L1) in the publicly available RNA-seq data to align and corroborate these findings with the key results observed later in our study. To avoid redundancy, we chose to present these genes in the initial figures as they are directly relevant to the subsequent analyses. Regarding the broader analysis of human RNA-seq data, our primary objective was to identify enriched biological processes and pathways, which served as a foundation for the focus and direction of this study.

      (3) Figures 3C-3G: I understand that IFNg-/- and NFR1R2a-/- mice are not showing elevated liver damage but it may simply be because of the non-responsiveness to the LPS challenge. I suggest using a different challenge or recovery experiments with the cytokines to show that the challenge is successful and results are caused by NAFLD, truly. The same goes for Figure 6: Looking at Figure 6D one may think that IFNg deficiency alters the LPS response independent of the diet condition (or NAFLD condition).

      We appreciate the reviewer’s insightful comment and fully understand the concern regarding the potential non-responsiveness of IFN-γ⁻/⁻ and TNFR1R2a⁻/⁻ mice to the LPS challenge. To address this point and confirm that these knockout animals are indeed responsive to LPS stimulation, we conducted an additional set of ex vivo experiments.

      Specifically, WT and cytokine-deficient (IFN-γ⁻/⁻) mice were fed either Chow or HFCD for two weeks, after which spleens were collected, and splenocytes were challenged in vitro with LPS. We then quantified TNF, IFN, and IL-6 production to confirm that these mice are capable of mounting cytokine responses upon LPS stimulation.

      Due to current breeding limitations and a temporary issue in colony maintenance of TNF-deficient mice, we were unable to include TNFR1R2a⁻/⁻ animals in this additional experiment. Nevertheless, we prioritized performing the analysis with the available knockout line to avoid leaving this important point unaddressed.

      These additional data demonstrate that IFN-γ-deficient mice remain responsive to LPS, reinforcing that the differences observed in vivo are related to the NAFLD condition rather than a lack of LPS responsiveness.

      (4) Figure 1 vs Figure 4: Rag-/- mice seem more susceptible to LPS-derived death even after normal conditions. But If I compare the survival data between Figure 1 and Figure 4, Rag-/- HFCD diet mice seem to be doing better than wt mice after LPS treatment. (1 day survival vs 2 days survival). How do you explain these different outcomes?

      We thank the reviewer for this insightful question regarding the survival data in Figures 1 and 4. Although there is a one-day difference in survival outcomes, Rag-/- mice consistently exhibit increased susceptibility to LPS-induced mortality can influence the exact survival timing. Nonetheless, across all experiments, Rag-/- mice display a reproducible phenotype of heightened sensitivity to LPS challenge, which is supported by multiple independent observations in our study.

      (5) How do you explain Figure 4J in connection to the observation presented with Figure 7: TNFa tissue levels, even though significant, seem very similar between the conditions?

      We would like to clarify that the animals in this study are in a metabolic syndrome state, with early-stage NAFLD characterized by hepatic fat accumulation without significant tissue injury, as shown in Figure 1C.

      Under these conditions, the LPS challenge triggers an exacerbated inflammatory response, leading to increased secretion of IFN-γ and TNF-α, primarily from NK cells and neutrophils. While TNFα levels may appear visually similar across conditions, the HFCD mice exhibit a heightened predisposition for an amplified immune response compared to chow-fed mice. This difference is consistent with the functional outcomes observed in our study and highlights the diet-specific sensitization of the immune system.

    1. eLife Assessment

      By using a combination of patch clamp recordings, calcium imaging and computer modeling, the authors analyze the spatial distribution of voltage gated calcium channels at glutamatergic synapses formed between layer 5 pyramidal neurons (L5PNs) and between layer 2/3 and L5PNs in the prefrontal cortex (PFC) and primary somatosensory cortex (S1); they conclude that the calcium channel-vesicle coupling is looser in the PFC compared to S1, although additional experiments are needed to determine how the distinct functional characteristics of these synapses in different brain regions might affect data interpretation. Overall, these findings are important because they have implications for shaping synaptic plasticity and neural circuit function across brain regions. They are solid because they are based on the use of a multi-pronged approach, although the presentation would benefit from stronger integration of the current findings with the existing literature and a more explicit discussion of potential limitations and confounding factors for data interpretation.

    2. Reviewer #1 (Public review):

      Summary:

      This study asks whether synapses formed by the same broad neuronal class (excitatory pyramidal neurons, PN) adapt their presynaptic organization in a cortex-specific manner, comparing the prefrontal cortex (PFC) with the primary somatosensory cortex (S1). The authors combine sophisticated electrophysiology (paired recordings and extracellular minimal stimulation), pharmacological perturbations of presynaptic Ca²⁺-secretion coupling, bouton Ca²⁺ imaging, and mechanistic modeling. Across two prominent excitatory connections (Layer 5 (L5) PN-L5PN and L2/3-L5PN), they provide convergent evidence that mature PFC synapses operate with looser Ca²⁺ channel-release sensor coupling than their S1 counterparts.

      Overall, the study provides an appealing mechanistic link between synaptic nano/micro-architecture and cortical-area specialization. The idea that PFC synapses retain a more "plasticity-favoring" presynaptic state, while the primary sensory cortex emphasizes reliability and timing precision, is potentially impactful for how we think about circuit computation and plasticity across cortical hierarchies.

      Strengths:

      A major strength is the multi-pronged experimental strategy. The paper first establishes robust, area-dependent differences in synaptic efficacy, reliability, timing, and short-term plasticity (facilitation prevailing in PFC versus depression in S1), using both paired recordings and minimal extracellular stimulation paradigms. The coupling interpretation is then directly supported by differential sensitivity to EGTA (and appropriate positive-control effects of fast chelators). Finally, volume-averaged calcium signals are reported to be similar across areas, arguing against trivial explanations based on gross differences in calcium influx, and the modeling provides a quantitative framework for interpreting the observed chelator effects.

      Weaknesses:

      Limitations are minor and concern interpretation/clarity rather than core results. Some key inferences rely on indirect readouts (chelator sensitivity, fluctuation analysis-derived parameters, bouton-averaged calcium signals), each of which carries assumptions and potential confounds that should be discussed more explicitly. In particular, the repatching paradigm for the paired-recording EGTA experiment, though very impressive, and the limited number of extracellular calcium conditions used for fluctuation analysis (three concentrations), can influence quantitative estimates and the confidence intervals around them.

    3. Reviewer #2 (Public review):

      Schwarze et al. investigated whether synaptic efficacy is brain-region specific. To this end, they compared synaptic connections established by layer 5 (L5) neocortical pyramidal cells and between L5 and L2/3 pyramidal cells. In order to identify the mechanism of this brain region specificity, the authors employed several experimental approaches, including paired electrophysiological recordings, extracellular stimulation, low- and high-affinity intracellular calcium chelators (EGTA and BAPTA), multiple probability fluctuation analysis (MPFA), and intracellular measurements of calcium transients as well as computational modelling. The findings of the present study indicate that synaptic connections in the primary somatosensory cortex (S1) are significantly stronger and more reliable than those in the prefrontal cortex (PFC).

      The study is timely, and the topic is of significant interest to the neuroscience community. Despite the extensive research that has been carried out on the neuroanatomy and receptor distribution of different brain regions, comparatively little attention has been paid to differences in synaptic physiology. The authors' approach is characterised by its elegance and comprehensive nature, and the conclusions drawn are compelling. Nevertheless, there are a number of unresolved issues.

      Major points:

      (1) The authors state that data from the S1 cortex were obtained in a previous study. In the context of an explicitly comparative study (PFC vs. S1cortex), it would have been advantageous for the authors to perform a subset of experiments in which both cortices were obtained from a single animal. This is a feasible undertaking, given the spatial separation of the PFC and S1 cortex.

      (2) Figure 1A is somewhat misleading because it could suggest that the authors have performed dual recordings in identified PFC pyramidal cells.

      (3) PFC and S1 cortex in rodents differ markedly in their morphological organisation. For example, in all sensory cortices, layer 4 is very pronounced; however, in the PFC of rodent,s no clear layer 4 can be found. On the other hand, PFC shows a clear separation of layers 2 and 3, which is not visible inthe S1 cortex. Furthermore, PFC pyramidal cells in layers 2, 3, and 5 exhibit significant heterogeneity, diverging considerably from those found in layers 5a and 5b of S1 cortex. Thus, there is no clear correlation between L5 pyramidal cells in the PFC and the S1 cortex. In order to achieve a meaningful comparison of the data obtained in PFC and S1 cortex, it is necessary for the authors to determine whether the record is from similar pyramidal cell populations.

      (3) In addition, PFC pyramidal cells in layer 2, 3 and 5 are highly heterogeneous and differ markedly from those in layer 5a and 5b of S1 cortex. To achieve a meaningful comparison of the data obtained in the PFC and the S1 cortex, the authors need to determine whether the record from similar pyramidal cell populations.

      (4) For the S1 cortex, in rats it has been found that L5 synaptic connection between pairs of L5a pyramidal cells and pairs of L5b pyramidal cells differ markedly with respect to mean EPSP amplitude, latency and coefficient of variation (cv, a surrogate measure for the synaptic release probability) (cf. Markram et al., 1997; Frick et al., 2008). It is therefore likely that PFC and S1 pre- and postsynaptic pyramidal cells are not only morphologically and electrophysiological distinct but also with respect to their synaptic properties. At least, the authors need to discuss these confounding issues and preferentially address them experimentally. For example, it would be helpful to demonstrate that paired recordings were made from the same pyramidal cell types, perhaps by documenting their morphology and/or firing patterns. In addition, they should discuss the marked difference in EPSP amplitude and putative release probability between their data and the earlier studies.

      (5) In order to perform multiple probability fluctuation analysis (MPFA), a parabolic fit with a mere three points is inadequate, particularly because 2 mM and 5 mM Ca2+ are close to the peak of the variance-to-mean parabola, and only 1 mM Ca2+ is on its initial linear part. A more meaningful result would have been obtained with an additional Ca2+ concentration between 1.0 and 2.0 mM, as these are closer to the physiological range. In this context, the authors should have quoted the more recent and more detailed paper by the Silver group (Saviane and Silver, 2006; Lanore and Silver, 2016) and not just the Clements and Silver review paper.

      (6) Methods: The authors should clarify whether their paired recordings from L5 pyramidal cells involved whole-cell recordings from both pre- and postsynaptic neurons. From Figure 1B, it appears as if the presynaptic neurons were not recorded in whole cell mode but rather stimulated in cell-attached mode. This is also reflected in the artefact visible in the current trace recorded in the postsynaptic neuron. The authors should explicitly state their methodological approach and mention how reliable the timing of the presynaptic action potential was under these circumstances. The same holds true for the extracellular stimulation protocol. A significantly more detailed description of the experimental protocol is necessary here.

      (7) Methods: The authors use Student's t-test for data comparison. The authors should verify that the data distribution was indeed normal, e.g. by using a Shapiro-Wilk test. If this is not the case, non-parametric tests should be used.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Max Schwarze and colleagues examined the coupling distance between presynaptic Ca²⁺ channels and the vesicular release sensor at neocortical synapses in mice. They propose that Ca²⁺ channel-release sensor coupling differs across cortical areas, with relatively loose (microdomain) coupling in prefrontal cortex (PFC) and tighter (nanodomain) coupling in primary somatosensory cortex (S1) for comparable pyramidal-neuron synapse types. To test this, they combine paired recordings and minimal stimulation with chelator manipulations (EGTA/BAPTA), mean-variance/MPFA-style analyses, presynaptic Ca²⁺ imaging, and computational modeling. They conclude that presynaptic coupling organization is area-specific in the mature cortex and contributes to regional differences in synaptic timing, reliability, and short-term plasticity.

      Strengths:

      This study tackles an important question and is strengthened by a cohesive body of evidence assembled from multiple complementary approaches. A major asset is the inclusion of high-value datasets, particularly the paired recordings between L5 pyramidal neurons and the systematic assessment of EGTA sensitivity, which provide a solid functional foundation for the authors' central claims. The work is further distinguished by its genuinely multimodal design: combining electrophysiology with presynaptic calcium imaging (and integrating these observations with quantitative analyses and modeling) offers a more mechanistic view of neurotransmitter release than any single method could provide. Overall, the direct, within-framework comparison of presynaptic release-control mechanisms across cortical areas for comparable synapse types is compelling and gives the conclusions a level of robustness and interpretability that is often difficult to achieve in studies of cortical synaptic diversity.

      Weaknesses:

      Several aspects would benefit from clearer explanation, stronger integration with the existing literature, and a more explicit discussion of limitations and potential confounds. Without these additions, some conclusions remain speculative. Throughout the manuscript, the authors also often imply that different measurements reflect the same underlying synapse population. This is unlikely to be strictly true across all experiments and makes it difficult to integrate results from the various approaches into a single, unified set of functional synaptic properties. In addition, some statements-particularly those linking coupling mode to "higher-order neocortical functions"-appear broader than what is directly supported by the experiments and should be tempered or more precisely scoped.

      Below, I list several topics that could help better frame the main findings of the present study and clarify how it relates to previously published work.

      (1) The authors use EGTA sensitivity of EPSCs (together with additional metrics) to argue that S1 and PFC synapses differ in Ca²⁺ channel-release sensor coupling. While this is a plausible interpretation, EGTA effects are not uniquely determined by coupling distance and can also reflect differences in Ca²⁺ entry kinetics, action potential waveform, endogenous buffering/extrusion, or release-sensor/vesicle state. The authors use a constrained modeling approach, but the rationale for the different constraint sets is not fully clear from the current description. It would be helpful to expand and clarify the Methods section to explain how these constraints were defined, justified, and applied (and how alternative constraint choices would affect the results). In this context, the Abstract's broader claim that the study "reveals microdomain coupling as a presynaptic structure-function correlate of higher-order neocortical functions" appears overstated. Given the well-known diversity of cortical synapses even within a single region (e.g., synapses onto different interneuron subclasses or different PN cell types, extracortical sources like thalamus), the authors should clarify the intended scope: is the conclusion meant to apply broadly across synapse classes in S1 and PFC, or only to the specific connection type(s) examined here?

      (2) The chelator logic is sound in principle, but the Discussion should more explicitly acknowledge standard caveats and alternative explanations. The authors partly address this by including presynaptic Ca²⁺ imaging and modeling, yet it would help to explain more clearly how the combination of (i) chelator sensitivity, (ii) presynaptic Ca²⁺ signals, and (iii) model constraints rules out-or substantially reduces the likelihood of-changes in AP waveform, Ca²⁺ influx kinetics, buffering/extrusion, or sensor/vesicle state as the primary drivers. In addition, recent hypotheses emphasizing vesicle priming and/or release-site occupancy as contributors to apparent EGTA sensitivity should be discussed as a complementary or alternative interpretation.

      (3) A substantial portion of the S1 comparison appears to rely on previously published datasets. This should be made unambiguous in the Results and Methods, and it would be helpful to summarize this clearly (e.g., in a table indicating which figures/analyses use new data versus reanalysis of published data). If this information is already present, it should be highlighted more prominently.

      (4) The modeling is informative, but the choice of a specific VGCC-release-site geometry and channel arrangement is not sufficiently justified. The manuscript adopts a particular spatial configuration, yet the rationale for selecting this geometry, rather than other plausible architectures discussed in the literature, is not clearly explained, nor is it meaningfully revisited in the Discussion. The authors should justify why the same organization is assumed across two distinct cortical areas and, ideally, include (or at a minimum discuss) a sensitivity analysis showing how key inferences (e.g., coupling distance and channel number) depend on the assumed geometry.

      (5) The calcium imaging data are valuable, but given the diversity of synapses within each cortical layer, it is not clear that imaged boutons can be confidently assigned to the specific connection types being interrogated electrophysiologically. A substantial fraction of boutons likely corresponds to different postsynaptic targets (including interneurons and distinct pyramidal-cell classes), and this heterogeneity could complicate interpretation. This limitation should be discussed explicitly

      (6) In unitary connections, the authors assess EGTA effects alongside other functional parameters (strength, delay, short-term plasticity), which is a major strength. However, for L2/3 to L5 connections, it appears that EGTA sensitivity was tested primarily using extracellular stimulation. Given anatomical and circuit differences between PFC and S1, extracellular stimulation may recruit different synapse populations across regions, potentially confounding regional comparisons of EGTA sensitivity. This limitation should be acknowledged explicitly. While I am not requesting technically demanding L2/3↔L5 paired recordings in S1, the possibility that different synapse identities are being sampled should be treated as a meaningful source of uncertainty. The Discussion would also benefit from placing the magnitude of EGTA effects in the context of prior "loose coupling" literature, where comparatively large EGTA effects have been reported in some systems. In addition, the reported difference between adult PFC EGTA effects and S1 inhibition appears small (on the order of <10%) and should be interpreted cautiously, especially given that PFC and S1 mature on different timelines and P21-P26 is unlikely to reflect a mature PFC circuit state. The adult cohort (P90-P100) is therefore important, but the age mismatch complicates PFC-S1 comparisons; ideally, S1 should be assessed at matched ages, or this limitation should be discussed explicitly. Finally, for statistical robustness, in panel D of Figure 2, were the comparisons corrected for multiple testing to control Type I error?

      (7) Alterations in initial release probability are often associated with changes in short-term plasticity. In the present manuscript, the authors report similar initial release probability at PFC and S1 synapses, yet observe differences in short-term plasticity profiles. The mechanistic basis for this apparent dissociation is not addressed and should be discussed explicitly, including potential explanations.

      (8) There are multiple instances where the text appears to cite non-existent or misnumbered figure panels (e.g., references to "Figure 4G-I / 4J" when the relevant material appears elsewhere). These should be corrected throughout, as they currently reduce readability and confidence.

      (9) The Methods describe P21-P26 animals, whereas the Results include older cohorts (e.g., P90-P100) and additional regions (e.g., mPFC). The Methods should be updated so that all cohorts and regions analyzed in the Results are fully described.

    5. Author response:

      We will extend and clarify the text of the paper according to the suggestions of the reviewers. In particular we will extend the description and discussion of the calcium chelator approach, re-patching and multiple probability fluctuation analysis. We will also include in the Results section that volume-averaged calcium signals were measured and extend the description about measurement of the resting calcium and variability between boutons. Literature will be included and discussed as suggested.

      In order to avoid any misunderstandings, we will also make it clearer that recordings from

      L5PN – L5PN synapses in S1 were published in our preceding papers (Bornschein et al., 2019a, b), but that these data were partially reanalyzed for the comparison with recordings from L5PN – L5PN synapses in PFC (this paper). We will also emphasize that the recordings from L2/3 to L5PN synapses in S1 and PFC were made directly in the present study. We will include a supplementary table, which explicitly shows for each figure which data are from Bornschein et al. (2019a, b) and which data were obtained in the present study. 

      We will consider all points of the reviewers and the recommendations of the editors in detail in the revised manuscript and/or our pointwise response.

      We recognized one factual error in the public reviews:

      Reviewer 2, point 7: “Methods: The authors use Student's t-test for data comparison. The authors should verify that the data distribution was indeed normal, e.g. by using a Shapiro-Wilk test. If this is not the case, non-parametric tests should be used.”

      A detailed description of the statistics, including test for normality, is given in the Methods section. In particular we wrote in the Methods: “Normality was tested using the Shapiro-Wilk Test. (…) To compare pre- and post-treatment data the paired t-test or the Wilcoxon signed rank test (WSR) was used, depending on the distribution of the data. (…)”

      To further emphasize that the data was tested for normal distribution, we have also extended the description of the statistical tests in the figure legends.

      Bornschein G, Brachtendorf S, Schmidt H (2019a) Developmental increase of neocortical  presynaptic efficacy via maturation of vesicle replenishment. Front Synaptic Neurosci 11:36.

      Bornschein G, Eilers J, Schmidt H (2019b) Neocortical high probability release sites are formed by distinct Ca2+ channel-to-release sensor topographies during development. Cell Rep 28:1410-1418 e1414.

    1. eLife Assessment

      The findings of this study are important since they cover the repurposing of small molecules as metalloprotease and phospholipase inhibitors for early intervention in the treatment of bothropic envenoming in the Neotropics, and thus provide a strong rationale for the progression of these inhibitors into future preclinical and clinical evaluation for snakebite indications across various ecological zones, albeit the current evidence casts doubts on the viability of repurposing nafamostat. The strength of the evidence is solid; however, there are some weaknesses, such as a lack of translatability of the in vivo model and insufficient venom characterization. Thus, the strength of the evidence can be enhanced by using a mouse model in future studies. The paper remains of interest to ophiologists, biochemists and medicinal chemists.

    2. Reviewer #1 (Public review):

      Very nice and coherent body of work with appropriate in vitro to in vivo transition in methods.

      Lovely and easy to follow figures that can be understood even without the manuscript.

      My recommendation is that a sentence or two be added clearly stating the authors think nafamostat is off the table and suggest other approaches/drugs that might be considered instead of just making a general statement. I think all this can be done in a few sentences.

      Gabexate was administered to a snakebite victim in this case report from about 20 years ago and also a good example of the now better recognized threat to pregnancy.

      Nasu K, Ueda T, Miyakawa I. Intrauterine fetal death caused by pit viper venom poisoning in early pregnancy. Gynecol Obstet Invest. 2004;57(2):114-6. doi: 10.1159/000075676. Epub 2003 Dec 19. PMID: 14691344

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to test whether a defined set of small molecules can lessen damaging effects caused by venoms from several Bothrops species, and whether these effects are consistent enough to suggest a broadly applicable approach. They present a cross-venom dataset spanning in-vitro activity readouts and blood-based functional outcomes, and include a chicken embryo model to explore whether venom inhibition can translate into improved survival. The central message is that certain small molecules can reduce specific venom-driven effects across multiple samples, providing a comparative resource for the field and a basis for prioritizing future validation.

      Strengths:

      The main value of this work is the breadth and structure of the dataset, which places multiple venoms and multiple readouts into a single, comparable framework that should be useful for readers evaluating patterns across samples. The experimental flow is generally coherent, moving from activity measurements to functional outcomes and then to an in-vivo test, which helps the reader understand how the authors link mechanism-oriented assays to more integrated endpoints. The manuscript also provides practical information for the community by highlighting which readouts appear most consistently affected across venoms, which can help guide hypothesis generation and study design in follow-up work.

      Comments on revisions:

      I would like to thank the authors for answering my questions. The manuscript has gained in quality, knowing the limitations that are now better stated in the manuscript.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editors and reviewers for their assessment of this manuscript, and for the positive words highlighting the value of undertaking evaluation of small molecule drugs for snakebite in the neotropics, inclusive of the quality of this work and the value of the validated screening pipeline. We completely agree that the next steps for this work will be to evaluate the preclinical efficacy of the identified drugs in mouse models, though this considerable undertaking will form the basis of future work. Critically, the pipeline that we describe herein facilitates the selection of the most appropriate candidates to progress into such mouse studies, aligning with the 3Rs principles for minimising the need for animal research. The comment around insufficient venom characterisation seems somewhat misplaced – the objective of this project was not to characterise the venoms used, but to evaluate the in vitro inhibition of venom toxin family activities and identify the potential utility of specific repurposed drugs as therapeutics for snakebite in the neotropics. Venom characterisation of the diverse samples used in this project would represent an entire project and manuscript in its own right. We are pleased that the reviewers highlight the gap in research on serine protease inhibitors and the value this paper has in highlighting that more research is required in this area to identify a candidate that is more suitable for future clinical use than nafamostat.


      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Small molecule therapeutics for snakebite have received a lot of attention for their potential to close the gap between bite and treatment, where antivenom is not immediately available.

      Strengths:

      There has been a lot of focus on Africa, Asia, and India, but very little work related to neotropical regions. The authors seek to begin filling this gap in the preclinical literature. The authors use well-developed methods for preclinical assessment.

      Weaknesses:

      A clearer and more focused discussion of the limitations of the overall present work would be desirable (e.g. protection vs. rescue, why marimastat over prinomastat for in vivo assays when both have been through clinical trials for other indications; real-world feasibility of nafamostat, which has a half-life of 1-2 minutes compared to camostat, which has a half-life of hours). All of this could be improved in a revision.

      We thank the reviewer for their shared opinion of the potential value of small molecules as snakebite envenoming therapeutics and their insight on the gap in focus in the neotropics, which this manuscript aims to address.

      Our work in this manuscript included standard practice of pre-incubation between drug and venom for all in vitro studies, and sequential (i.e. not co-incubation) administration in the egg model. In our revised manuscript we will make these distinctions clearer. Use of a ‘rescue’ approach in the in vitro assays is not feasible due to the rapid destruction of the substrates used for assay readouts. The clearest rationale for the use of rescue models relates to their power within in vivo preclinical models (i.e. murine envenoming models) which, following the in vitro characterisations presented in this paper, are the logical next step for evaluating small molecule drugs for inhibiting neotropical snake venoms.

      Although both marimastat and prinomastat are repurposed drugs that have undergone clinical evaluation for other indications, marimastat has been more extensively characterised preclinically than prinomastat for snakebite, and will soon enter Phase II clinical trial evaluation for this indication (https://www.ddw-online.com/ophirex-to-produce-snake-venom-inhibitor-for-lstm-study-40669-202602/). Marimastat also has a longer half-life in humans of 8-10 hours (Millar et al. 1998), compared to prinomastat (2-5h, Hande et al. 2004). We will more clearly highlight the rationale for selecting marimastat in the revised manuscript.

      Although we appreciate the reviewer’s point regarding the short half-life of nafamostat (which is typically given by continuous iv infusion due to its short half-life), in the manuscript we have already stated that we do not recommend the progression of nafamostat as a snake venom serine protease (SVSP) inhibitor candidate due its low efficacy and off target effects. We highlight the need for the community to identify other serine protease inhibitors that might have utility for snakebite.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to test whether a defined set of small molecules can lessen damaging effects caused by venoms from several Bothrops species, and whether these effects are consistent enough to suggest a broadly applicable approach. They present a cross-venom dataset spanning in-vitro activity readouts and blood-based functional outcomes, and include a chicken embryo model to explore whether venom inhibition can translate into improved survival. The central message is that certain small molecules can reduce specific venom-driven effects across multiple samples, providing a comparative resource for the field and a basis for prioritizing future validation.

      Strengths:

      The main value of this work is the breadth and structure of the dataset, which places multiple venoms and multiple readouts into a single, comparable framework that should be useful for readers evaluating patterns across samples. The experimental flow is generally coherent, moving from activity measurements to functional outcomes and then to an in-vivo test, which helps the reader understand how the authors link mechanism-oriented assays to more integrated endpoints. The manuscript also provides practical information for the community by highlighting which readouts appear most consistently affected across venoms, which can help guide hypothesis generation and study design in follow-up work.

      Weaknesses:

      Several aspects of the study design and framing reduce the confidence with which readers can translate the findings beyond the specific experimental context presented. The evidence base is strongest in controlled in-vitro settings, while the bridge to real-world effectiveness remains limited, particularly for understanding performance under conditions that better reflect delayed treatment and systemic exposure. As a result, the manuscript is best interpreted as a well-organized comparative screening study with promising signals, rather than a definitive demonstration of a broadly effective, deployable intervention.

      We appreciate the reviewer’s opinion on the thorough and logical workflow we present in this manuscript and the value this pipeline providers the field for future and parallel work. We agree with the reviewer that this provides a well-organized comparative screening study applicable to different snake species or therapeutics. In relation to the comment on this manuscript being a definitive demonstration of a broadly effective, deployable intervention we agree with their opinion and are happy to clarify that while the evidence presented in this manuscript is promising, there is much work still to do before such molecules are ready for deployment for treating snakebite. Ultimately, this manuscript supports the growing evidence of the promising utility of marimastat and varespladib, and extends this evidence to neotropical snake venoms in a comparative manner. The next step will be to evaluate the efficacy of these molecules within in vivo murine preclinical models, which will be crucial for further supporting the evidence base for onward translation.

      Reviewer #3 (Public review):

      In this work, the authors wanted to evaluate repurposed small molecule inhibitors for the treatment of envenomation by snakes of the Bothrops genus; one of the most medically relevant in the Americas. I believe the objectives of the research were clearly achieved, and compelling evidence for the ability of these molecules to neutralize enzymatic and toxic activities of metalloproteinases and phospholipases in all the tested venoms is provided. Furthermore, the work highlights the limited efficacy of the tested serine protease inhibitor, suggesting a need for drug discovery campaigns to address toxicity caused by this protein family. The methods are well designed and performed, and the use of both in vitro and in vivo methodologies makes this a thorough and robust work.

      These results are extremely relevant, since they take us one step further to a potential orally administered snakebite treatment. The existence of such a treatment could improve the outcomes for thousands of snakebite victims worldwide. I have a few comments and questions that I hope will be useful to the authors:

      We thank the author for their high regard for the purpose and execution of this work. Their insight in relation to questions are supportive for an improved manuscript and discussion points for the field.

      During the introduction, the authors mention that small-molecule inhibitors can neutralize the localized tissue damage via cytotoxicity of some venoms, and cite PLA2s, SVMPs and/or cytotoxic 3FTxs as the main causing agents of this pathology. I am not aware of any direct effect described by small molecule inhibitors on cytotoxic 3FTxs alone. Has this been observed at all? Or is it more likely that the small molecule inhibitors act on the enzymatic toxins only, preventing synergistic effects with 3FTxs?

      We apologise for this error on our behalf. While inhibitory molecules have been described for cytotoxic 3FTxs, these are not small molecules as alluded to in the previous version of the manuscript. We have amended this text in our revised manuscript.

      I think it would be relevant to address the effects of non-enzymatic PLA2s, such as myotoxin II, which have been described in detail within Bothrops venoms. I believe there is some evidence of Varespladib also having a neutralizing effect on the myotoxicity caused by these non-enzymatic PLA2s. I suggest adding a comment about the contribution of these toxins in the discussion or in the section where PLA2 activity of the venoms is compared. In my opinion, right now it seems like these were overlooked.

      We thank the reviewer for highlighting this point. We agree that this is highly relevant and would benefit from discussion in the revised manuscript given the nature of our assays and the non-enzymatic mechanism of action of certain Bothrops PLA<sub>2</sub>s. We have added this to the discussion.

      Regarding Marimastat and the other MP inhibitors, are there any studies showing that they don't have an effect on endogenous MPs? I understand they have been approved for human use before, but is there any indication that they would not have an effect at the doses that would be required to treat envenomation?

      Most matrix metalloproteinases inhibitors will act on endogenous MPs to at least some extent (variable potency on different MMPs). Marimastat has demonstrated activity against endogenous metalloproteinases, including MMP1, which was hypothesised to cause severe joint pain when used chronically (i.e. frequent dosing over many weeks) for indications such as cancer, though this effect was reversible within 8 weeks of cessation of drug administration (Wojtowicz-Praga, 1998). Thus long-term use of matrix metalloproteinases inhibitors can cause safety concerns. However, the anticipated duration of dosing for snakebite, which is an acute life-threatening condition, is a few days. It is therefore unlikely that prior safety concerns observed following chronic dosing in cancer studies would apply to its potential use as a snakebite field therapy.

      Regarding the quenched fluorescence substrate used for enzymatic activity. Is there a possibility that some of the SVMPs would not act on this substrate, and therefore their activity or neutralization is not observed? Would it be relevant to test other substrates, such as gelatin, collagen, or even specific clotting factors?

      It has been observed that certain SVMPs (specifically several PI SVMPs) are not active against this ES010 substrate in vitro. The substrate used in the in vitro SVMP assay is reported by the manufacturer as a substrate for a wide range of MMPs which target the extracellular matrix components mentioned by the reviewer, i.e. collagenases and gelatinases as well as matrilysins, stromelysins and elastate. This in vitro assay combined with the coagulation assays are complementary in covering the main targets of SVMPs (ECM and clotting cascade), prior to haemorrhagic assessment in the egg model. Thus, we are confident that activity for the broad range of SVMP isoforms will be captured through the screening pipeline we have developed.

      Finally, could the authors comment or provide some bibliography regarding the translatability of the chicken embryo model in the context of envenomation?

      Our current model is based on an earlier egg embryo model (Sells et al. 1997, Sells et al. 1998 and Sells et al. 2000) which described good correlations (p<0.01) with the standard WHO murine preclinical envenoming model. These studies have assessed correlations for minimal haemorrhagic doses (MHDs), LD50s and ED50s in both models for a selection of viper venoms. As chicken embryos at day 6 of development have incomplete neural arcs, the model is not well suited for assessing neurotoxic effects, but can be effectively used for addressing venom-induced haemorrhage and lethality and for testing therapeutics. In addition, a more recent study (Yusuf et al. 2023) reported almost identical LD50s for the venom of Bitis arietans between the two in vivo approaches. The model is also being pursued as a preclinical testing model by an antivenom manufacturer with the focus of reducing the use of rodents in batch release testing (Verity et al. 2021). We will provide further clarification on the rationale for using the egg model, including the supportive references outlined above, in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The manuscript provides a useful comparative dataset across multiple Bothrops venoms and supports SVMP inhibition as a broadly effective lever in the authors in-vitro work. However, the strength of the 'pan-Bothrops' and translational claims is currently limited by insufficient characterization of the exact venom samples tested and by experimental designs that fall in clinically realistic rescue.

      Major comments:

      (1) The venoms used in this study are historical batches and are not formally characterized beyond SDS-PAGE and literature summaries, despite well-known intra- and inter-population venom variability; this weakens the generalization of the conclusions.

      To address this comment, we have increased clarity on our venom sources being historic, Due to the historic source locality is not available beyond country of origin, with the exception of B. lanceolatus which is endemic to Martinique. Figure 1 also makes clear that we agree with the reviewer that the variation is high within Bothrops species. We discuss this variation on the limitations in our sampling for making broad conclusions throughout the first paragraph of the discussion, with the final sentence stating Future proteomic characterisations of the specific venom samples used in this study, which were all sourced from a historical collection (except for B. lanceolatus), would be informative in this regard. Although venom composition of our samples has not been characterised, the focus of the manuscript is the characterisation of the whole venom functional activity through a wide ranging screening pipeline, and the generalisation of our findings is supported by the diversity of the venom samples (i.e. several species) despite them not being characterised (which is not critical for the focus of the study).

      (2) On a technical comment, the venom inhibition assays appear to rely on drug-first or preincubation conditions, which can easily overestimate efficacy compared with real snakebite envenomation, where toxins distribute and engage targets rapidly. Here, a translational gap is the clinical feasibility of the 'repurposed' inhibitors, as it is unclear whether the drugs central to the conclusions (especially marimastat, prinomastat and varespladib) are realistically available or stocked in hospitals or could be deployed in regions where Bothrops envenoming occurs. I think that the manuscript should clearly distinguish this from candidates with a plausible access and delivery pathway.

      Our work in this manuscript includes standard practice of pre-incubation between drug and venom for all in vitro studies, and sequential (i.e. not co-incubation) administration in the egg model. None of our methods administer drug-first. Throughout the methods and figure legends we have made these distinctions clearer. Use of a ‘rescue’ approach in the in vitro assays is not feasible due to the rapid destruction of the substrates used for assay readouts. The clearest rationale for the use of rescue models relates to their power within in vivo preclinical models (i.e. murine envenoming models), which would be the next step for this research programme.

      While the evidence presented in this manuscript is promising, there is much work still to do before such molecules are ready for deployment for treating snakebite, inclusive of the requirement to complete clinical trials, cost-benefit analysis and policy change and manufacturing/distribution feasibility assessments. Ultimately, this manuscript supports the growing evidence of the promising utility of marimastat and varespladib, and extends this evidence to neotropical snake venoms in a comparative manner. The next step will be to evaluate the efficacy of these molecules within rescue in vivo murine preclinical models, which will be crucial for further supporting the evidence base for onward translation. To further support this point we have included an additional section to the manuscript discussing the current preclinical and clinical progression of prinomastat and marimastat, which also incorporates the public comment on selection of marimastat over prinomastat.

      (3) In my opinion, the Nafamostat results and discussion need reframing, given weak SVSP inhibition and intrinsic anticoagulant behavior at 5 µM. Excluding it from certain analyses undermines interpretability, and it may be more appropriate to include it throughout as an explicit negative control condition (showing its baseline anticoagulant effect) rather than omitting it.

      Although we understand the reviewers opinion here, we disagree and believe that including nafomastat as a ‘negative control’ may present a negative reflection on the benefit that an efficacious serine protease inhibitor could provide. Furthermore, as the intrinsic anticoagulant effect of nafamostat cannot be de-coupled from direct SVSP toxin inhibition we were unable to interpret the activity which undermines the results. This can be seen in Figure 3b, which demonstrates that a false positive result would occur. For the serine protease assay, we do clearly discuss the lack of efficacy and justification of why EC<sub>50</sub> testing wasn’t appropriate within the guidance of our screening protocols.

      In the manuscript we have now further justified our approach in relation to the limitations of nafamostat as a snake venom serine protease (SVSP) inhibitor candidate due its low efficacy and off target effects. We highlight the need for the community to identify other serine protease inhibitors that might have utility for snakebite.

      (4) The data presentation needs consistent statistical analyses (currently absent for multiple key figures, including Figures 2, 3, 4, 6 and 7) and a clearer explanation for the dose of venom and drugs you choose. For example, Figure 3 relies on a fixed 5 µM drug concentration and very different venom amounts (50-100-250 ng), but it is not discussed whether such exposures are achievable in vivo, or how these concentrations map onto expected pharmacokinetics in patients. Likewise, Bothrops venoms can contain both pro- and anticoagulant activities, so the authors should justify how their framework accounts for anticoagulant components and why the observed plasma phenotypes are interpreted as they are

      In relation to the reviewers comment on the need for consistent analysis we thank the reviewer for flagging this and have now included these in figures 3, 4, 6 and 7. However, Figure 2 is presented to display the variation between all the venoms and ultimately used to select the most relevant doses for the latter inhibition experiments, therefore statistical analysis is not relevant for this figure. The updated statistical analysis now includes the following, which has been included in the relevant figure legends and results sections;

      Figure 3 - Bars indicate significant results (p = <0.05) identified through one-way ANOVA with Dunnett’s multiple comparisons test to the DMSO control

      Figure 4 - two-way ANOVA with Šídák's multiple comparisons test of each venom control compared to the matched venom treated with inhibitor

      Figure 6 – the CT and MCF data were analysed independently using one-way ANOVA with Tukey’s multiple comparisons test

      Figure 7 - Log-rank test (Mantel-Cox) with Holm- Šídák's multiple comparisons test against treatment vs venom-only control

      We have ensured that all figure legends clearly indicate the venom and drug dose to aid the clarity which the reviewer requested.

      The comment Figure 3 relies on a fixed 5 µM drug concentration and very different venom amounts (50-100-250 ng), but it is not discussed whether such exposures are achievable in vivo, or how these concentrations map onto expected pharmacokinetics in patients. is an understandable query however, in vitro assessment such as those carried out in this manuscript are not designed to directly inform pharmacokinetic/pharmacodynmanic interpretations, largely because they do not replicate real world envenoming (i.e. preincubation would not occur between a venom and treatment). This is why, as stated, follow on preclinical and clinical assessments are needed for onward progression of these inhibitors to inform dosing regimens that might achieve the necessary exposures required for in vivo venom neutralisation. That being said, PK/PD work has been initiated within Phase I trials, for example with DMPS Abouyannis et al. 2025 demonstrated a plasma exposure of >10 µg/mL for single doses of 1,200 mg and higher. This is equivalent to 80 µM, which although is lower than the EC<sub>50</sub> for some venoms in the clotting assay (Figure 3J), the venom dose (50 to 250 ng/ 50 µL, i.e. 1,000 to 5,000 ng/µL) is estimated to be >1000 times higher than a natural envenoming by Bothrops atrox at less than 1 ng/mL in serum (https://doi.org/10.1016/j.toxicon.2022.09.010). These extrapolations therefore indicate that the doses selected in our studies would have human clinical relevance.

      Finally, in terms of anticoagulant venom effects - these would be observed in our experimental approach either as reduced kinetic responses in the plasma clotting assay (as observed with nafamostat in Figure 3B) or as a prolonged clotting time in the thromboelastography assay (Figure 6). As stated in the results section Comparison of coagulation profiles, all of the venoms tested presented with a procoagulant effect. If underlying anticoagulant activity from PLA<sub>2</sub> toxins was to arise after inhibition of the procoagulant toxins (i.e. SVMPs by marimastat), as has been seen for certain other snake venoms previously, this would result in a percentage inhibition far greater than 100% in the plasma assay (Figure 3C to I) or as a prolonged clotting time in the thromboelastography assay. These described anticoagulant profiles were not observed with any venom tested in this study.

      (5) Finally, the in vivo evidence is limited to a chicken embryo model. To support your hypothesis, a conventional mouse model with delayed post-envenomation dosing (24-36 h monitoring) is needed to address both safety/toxicity and post-exposure efficacy, and to define a realistic therapeutic window, especially because venom toxins act very quickly and the timing of administration is central to the clinical utility of any small-molecule approach.

      We agree with the reviewer that the next important step for this research activity is utilising murine preclinical models to validate the in vitro and preliminary in vivo findings described in this manuscript. However, as stated above, this study provides the initial evidence base that the promising utility of marimastat, DMPS and varespladib as repurposed snakebite drugs extends to a range of neotropical viper venoms. Evaluating the safety, efficacy (both precincubation and rescue approaches) and PK/PD relationships to inform optimal dosing strategies of these molecules will be crucial next steps for the field. However, these activities are far from trivial and will take several years of additional research, and therefore fall outside the scope of this initial manuscript.

      To address the concern related to the evidence is limited to a chicken embryo model, we have included additional sentences to discuss the wider use of the egg model within snakebite research and related translation to murine studies.

      Minor comments:

      (1) Figure 2D: How do you discuss the fact that "no venom" has SVSP activity?

      The data for all in vitro assays in Figure 2 is presented as AUC from the raw data (absorbance or fluorescence), for consistency across assay. Therefore, all assays (B to D) have background signal in the absence of venom. The SVSP assay has a greater background signal.

      (2) For better understanding, I would suggest adding a dedicated column in Figure 4A with Nafamostat SVSP data reported as "N/D" where applicable.

      As stated in the results, due to the weak inhibitory activity EC<sub>50</sub> assessment was not justified, therefore adding this column would be redundant.

      (3) The introduction is too long relative to the experimental content and would benefit from tightening to sharpen the motivation and unmet need.

      We thank the reviewer for their opinion and we have reviewed the introductory section again. While we made minor edits throughout, we decided not to make substantial modifications to it.

      Reviewer #3 (Recommendations for the authors):

      I only have some minor comments:

      (1) In line 100, the word "that" is repeated.

      We thank the reviewer for spotting this error, which we have corrected.

      (2) Line 433. I believe the word "compromising" should be substituted by "comprising" here.

      We thank the reviewer for spotting this.

      (3) Figure 1 and supplementary: Bothrops asper venom has been very thoroughly studied, and using only one study from Costa Rica might underestimate the venom variation within the species. I suggest looking at the following study: https://doi.org/10.1016/j.toxicon.2022.106983. Maybe it is not necessary to change anything, but worth looking into.

      We appreciate the reviewer flagging this paper, it has been added to the manuscript (reference 48) and has provided additional data for Figure 1 and Supplementary table 1.

      (4) Methods: Given the intraspecies variation described for some of these species, I believe it is relevant to add the locality of origin of the venoms, and not only the country. I, of course, understand this is often unknown for historical samples.

      We have included the following sentence in the methods. Due to the historic nature of the venom samples, the source locality is not available beyond country of origin, with the exception of B. lanceolatus which is endemic to Martinique.

      (5) Figure 3: It is not very accurate to show an SD when the sample number is 2. I suggest, when possible, showing the mean and the two data points in the plots. This also applies to other figures where n=2. Also, in Figure 3D, does Marimastat seem to have an anticoagulant effect, or is this just within normal variation?

      We have removed the statement in the statistics paragraph of the methods Standard deviation (SD) for all kinetic reads and standard error for AUC is reported based on Prism v10 but kept the sentence. The sample sizes for HTS assays including the SVMP, PLA<sub>2</sub> and coagulation experiment are the average of the means from independent assays (n >2 within each independent assay). We understand the reviewer’s opinion on limited meaning of SD as well as SE for Fig 3 A to I, therefore we have changed the error bars to range, as we think that displaying the individual points would result in a lack of visual and analytic clarity.

      In relation to the query about marimastat anticoagulant effect in Fig 4D, as shown in 4B marimastat has no direct anticoagulant effect. The >100% inhibition for marimastat is likely to be normal variation as this is a biological assay which has high variability. However, it could also be that the strong inhibition of the SVMPs in B. asper along with limited SVSP activity has unmasked an anticoagulant effect of the remaining PLA<sub>2</sub> toxin which has high activity in this venom. That being said, as B. asper has a similar profile, we would have expected to see a similar profile in B. atrox in both the plasma and TEG assays. Therefore, assay variation seems the most likely reason for this observation.

    1. eLife Assessment

      This study introduces an important method to estimate the probability of malaria importation with a new Bayesian approach that integrates epidemiological, travel reports, and genetic data. The authors provide convincing evidence for the value of this model in identifying the main sources of malaria transmission and risk factors. This work will be of interest to the area of genomic epidemiology and public health strategies aiming to eliminate malaria.

    2. Reviewer #1 (Public review):

      Summary:

      This study presents a new Bayesian approach to estimate importation probabilities of malaria combining epidemiological data, travel history, and genetic data through pairwise IBD estimates. Importation is an important factor challenging malaria elimination, especially in low transmission settings. This paper focus on Magude and Matutuine, two districts in south Mozambique with very low malaria transmission. The results show isolation-by-distance in Mozambique, with genetic relatedness decreasing with distances larger than 100 km, and no spatial correlation for distances between 10 and 100 km. But again strong spatial correlation in distances smaller than 10 km. They report high genetic relatedness between Matutuine and Inhambane, higher than between Matutuine and Magude. Inhambane is the main source of importation in Matutuine, accounting for 63.5% of imported cases. Magude, on the other hand, shows smaller importation and travel rates than Matutuine, as it is a rural area with less mobility. Additionally, they report higher levels of importation and travel in the dry season, when transmission is lower. Also, no association with importation was found for occupation, sex and other factors. These data have practical implications for public health strategies aiming malaria elimination, for example, testing and treating travelers from Matutuine in the dry season.

      Strengths:

      The strength of this study relies in the combination of different sources of data - epidemiological, travel and genetic data - to estimate importation probabilities, the statistical analyses.

      Weaknesses:

      The authors recognize the limitations related to sample size and the biases of travel reports.

    3. Reviewer #2 (Public review):

      Summary:

      Based on a detailed dataset, the authors present a novel Bayesian approach to classify malaria cases as either imported or locally acquired.

      Strengths:

      The proposed Bayesian approach for case classification is simple, well justified, and allows the integration of parasite genomics, travel history, and epidemiological data.

      Weakness:

      While the authors aim to classify cases as imported or locally acquired, the method does not quantify the contribution of each case type to overall transmission, which the authors leave for future study.

    4. Reviewer #3 (Public review):

      This work provides a novel statistical model to identify imported malaria cases, which are an important challenge for elimination, particularly in low-transmission areas. This tool was applied in Plasmodium falciparum populations in Mozambique and determined differences in importation rates in two low-transmission districts in the South.

      Strengths:

      The study has several strengths, particularly the development of a novel Bayesian model integrating genomic, epidemiological, and travel data to estimate importation probabilities. The findings provided important insights into malaria transmission dynamics, including the identification of importation sources and regional differences in importation rates across Mozambique. These results highlight the potential value of targeted interventions among traveler populations to support malaria elimination efforts. Moreover, this approach could be adapted to other epidemiological settings.

      Weaknesses:

      The study has some limitations, including uneven sample representation across provinces, incomplete metadata for risk factor analysis and a proxy for transmission intensity. Future work will include a new sample collection effort and the incorporation of monthly malaria incidence estimates.

    5. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents a new Bayesian approach to estimate importation probabilities of malaria combining epidemiological data, travel history, and genetic data through pairwise IBD estimates. Importation is an important factor challenging malaria elimination, especially in low transmission settings. This paper focus on Magude and Matutuine, two districts in south Mozambique with very low malaria transmission. The results show isolation-by-distance in Mozambique, with genetic relatedness decreasing with distances larger than 100 km, and no spatial correlation for distances between 10 and 100 km. But again strong spatial correlation in distances smaller than 10 km. They report high genetic relatedness between Matutuine and Inhambane, higher than between Matutuine and Magude. Inhambane is the main source of importation in Matutuine, accounting for 63.5% of imported cases. Magude, on the other hand, shows smaller importation and travel rates than Matutuine, as it is a rural area with less mobility. Additionally, they report higher levels of importation and travel in the dry season, when transmission is lower. Also, no association with importation was found for occupation, sex and other factors. These data have practical implications for public health strategies aiming malaria elimination, for example, testing and treating travelers from Matutuine in the dry season.

      Strengths:

      The strength of this study relies in the combination of different sources of data - epidemiological, travel and genetic data - to estimate importation probabilities, the statistical analyses.

      Weaknesses:

      The authors recognize the limitations related to sample size and the biases of travel reports.

      We appreciate the review and comment about the manuscript.

      Reviewer #2 (Public review):

      Summary:

      Based on a detailed dataset, the authors present a novel Bayesian approach to classify malaria cases as either imported or locally acquired.

      Strengths:

      The proposed Bayesian approach for case classification is simple, well justified, and allows the integration of parasite genomics, travel history, and epidemiological data.

      Weakness:

      While the authors aim to classify cases as imported or locally acquired, the work lacks a quantification of the contribution of each case type to overall transmission.

      Comments on revisions:

      All my questions and concerns were satisfactorily addressed.

      We appreciate the review and comment about the manuscript. In fact, the approach does not pretend to quantify the contribution of each case to overall transmission. In the discussion we state it and refer to future work with this scope.

      Reviewer #3 (Public review):

      This work provides a novel statistical model to identify imported malaria cases, which are an important challenge for elimination, particularly in low-transmission areas. This tool was applied in Plasmodium falciparum populations in Mozambique and determined differences in importation rates in 2 low-transmission districts in the South.

      Strengths:

      The study has several strengths, mainly the development of a novel Bayesian model that integrates genomic, epidemiological, and travel data to estimate importation probabilities. The results showed insights into malaria transmission dynamics, particularly identifying importation sources and differences in importation rates in Mozambique. Finally, the relevance of the findings is to suggest interventions focusing on the traveler population to support efforts for malaria elimination.

      Weaknesses:

      The study also has some limitations, although the authors have plans to address them. The sample collection was not representative of some provinces, and not all samples had sufficient metadata for the risk factor analysis. Additionally, the authors used a proxy for transmission intensity and assumed some other conditions to calculate the importation probability for specific scenarios. They plan to conduct a new sample collection and include monthly malaria incidence estimates in the future.

      Comments on revisions:

      Delete "We added this text to the discussion" in line 302 (Discussion)

      I recommend adding the plans to address limitations indicated in the Response to Reviewers document in the Discussion. This would really strengthen the limitation section.

      Thank you for pointing to these aspects. We deleted the sentence mentioned. In the discussion section, we now finish the paragraph on limitations with the proposed future work to address them.

    1. eLife Assessment

      This important study presents a theoretically grounded framework for dimensionality reduction in single-cell RNA sequencing data, utilizing the principles of Riemannian manifolds. The proposed method addresses a critical challenge in bioinformatics-extracting highly informative latent dimensions without relying on the heuristic assumptions common in existing workflows. The evidence supporting the method's utility in estimating intrinsic dimensionality and identifying cell types is convincing, though the work would benefit from more rigorous validation against established ground truths and a clearer strategy for addressing prevalent batch effects.

    2. Reviewer #1 (Public review):

      Summary:

      Sidarta-Oliveira et al. present TopOMetry, a novel dimensionality reduction method based on the eigendecomposition of approximated Laplace-Beltrami Operator. Shortly, TopOMetry is an iterative version of the existing spectral methods (e.g., Laplacian Eigenmap or Diffusion map). It approximates the Laplacian operators twice, once in a "phenotypic space" and then once again in the eigenbases space. By doing this the approximated operator will contain more information of the manifold, which allows for more robust and accurate downstream analyses.

      Strengths:

      - Introduces operator-native fidelity scores and Riemannian diagnostics to single-cell analysis, enabling researchers to evaluate and trust embeddings - functionality absent in prior methods.<br /> - The approach was rigorously tested based on synthetic and real single-cell RNA-seq datasets.<br /> - The package is well-made and easily scalable to millions of cells.<br /> - The comprehensive documentation helps the end-users to run desired analyses.

      Weaknesses:

      - The method is an extension of the current state-of-art methods, not a fundamentally new one.

      Comments on revised version:

      The revised manuscript partially addresses the concerns raised in the prior review. The jargon weakness has been substantially mitigated by relocating mathematical derivations to the Methods section and simplifying language in the main text; this weakness has been updated accordingly.

      The introduction of operator-native fidelity scores and Riemannian diagnostics represents a meaningful addition and has been added to the Strengths. The benchmarking scope has also been notably expanded.

      The core weakness - that the method is an extension of existing spectral methods rather than a fundamentally new contribution - remains unchanged, as the authors' rebuttal did not provide a sufficiently precise mathematical argument to overturn it.

    3. Reviewer #2 (Public review):

      Summary:

      This work introduces a novel framework to systematically learn the latent dimensions of single-cell data, grounded in the theory of the Riemannian manifold. The authors demonstrate how this framework can be applied to various important tasks, such as estimating intrinsic dimensionalities, annotating cell types, etc. They did a great job of tackling an important but not yet established problem in the field and approaching it with a theoretically sound and novel approach. I think after a more rigorous and comprehensive validation, this work could be impactful.

      Strengths:

      - Dimensionality reduction is a routine step in analyzing many high-dimensional data, such as molecular data. While the downstream analysis results depend heavily on this step, existing methods rely on strong assumptions and are sometimes heuristic. The authors present a novel, theoretically grounded approach to address this important problem.

      - The authors demonstrated its usability in downstream analysis in a comprehensive manner. Especially, they show evidence suggesting novel T-cell subpopulations.

      - I commend the authors for releasing and maintaining their software well with comprehensive documentation. This significantly increases the usability and accessibility of the method.

      Weaknesses:

      - The paper lacks experiments that validate the results. It would be beneficial to see additional evaluation settings with better-established ground truths to more strongly demonstrate the method's effectiveness.

      - Batch effects are prevalent in single-cell data. The paper does not adequately address how the proposed method handles this issue.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) The method is an extension of the current state-of-art methods, not a fundamentally new one.

      We respectfully disagree with this characterization. While TopoMetry is inspired by the theory of spectral geometry, it is not a simple extension of existing dimensionality reduction methods such as Diffusion Maps. Instead, TopoMetry introduces a new framework for single-cell analysis that:

      Iteratively approximates manifold geometry by constructing refined diffusion operators on spectral scaffolds (“the geometry of the geometry”), a procedure not present in existing methods.

      Provides a unified workflow for dimensionality estimation, clustering, visualization, imputation, lineage inference, and diagnostics, all within the same geometric framework.

      Introduces operator-native fidelity scores and Riemannian diagnostics to single-cell analysis, enabling researchers to evaluate and trust embeddings—functionality absent in prior methods.

      Thus, TopoMetry represents a new paradigm for geometry-aware single-cell analysis, not merely a reimplementation of existing algorithms.

      (2) The paper contains a lot of jargon.

      We have thoroughly simplified the text throughout the manuscript. We now introduce geometric concepts in accessible terms, avoiding technical details where they are not essential for biological interpretation. For example, references to the Laplace–Beltrami operator and its eigenfunctions have been reduced and reframed in terms of “geometry,” “diffusion,” and “spectral scaffolds,” which are more intuitive for a general audience.

      Reviewer #1 (Recommendations for the authors):

      (1) What happens if the LBO is approximated more than twice? As the main idea of the method is an iterative approach to approximate LBO more precisely, then the authors would have already considered this. If so, this could be additionally discussed in the manuscript.

      We thank the reviewer for this important point. Indeed, TopoMetry’s design naturally supports iterating the Laplace–Beltrami operator (LBO) approximation beyond two steps. However, additional iterations (three or more) lead to only marginal improvements in final results while significantly increasing computational cost. In some tested cases, additional iterations could even over-smooth the data, reducing the resolution of fine-scale structure. The revised manuscript avoids an excessive focus on iterative LBO approximations and instead centers the narrative around representing and evaluating the underlying geometry of single-cell data.

      (2) As the paper describes the method in a very comprehensive way, as a result, it contains a lot of mathematical equations and jargon. This could hinder the visibility of the whole manuscript to biologists who do not have a background in mathematics. Thus, I strongly recommend that the authors consider moving a considerable amount of text to the supplementary material, and the main text should focus on the benchmarking results and the possible applications.

      We appreciate this recommendation and have substantially revised the manuscript to make it more accessible to a broad biological audience. In the revised version:

      We moved detailed mathematical derivations and operator definitions to the Methods section, keeping only the most essential concepts in the main text.

      We reframed technical terms (e.g., Laplace–Beltrami operator, eigenfunctions) in simpler and more intuitive language in the main text. 

      The Results section now emphasizes benchmarking outcomes and biological applications.

      Reviewer #2 (Public review):

      (1) To encourage the single-cell community to adopt this method, the authors should more clearly demonstrate its advantages over existing methods. There are many single cell analysis algorithms that are proposed in each task and some of them are widely used by biologists. However, the comparison in this work is somewhat limited. For example, Even methods mentioned in the relevant work paragraph (2nd paragraph) on page 2 are not all compared, or the reason why they are not included is not discussed. Also, I am curious how PC dimensions are determined. The choice of 300 PCs on page 11 seems arbitrary. Furthermore, the usefulness of dimension-reduced data also depends a lot on the preceding processing steps, such as highly variable gene selection. I understand it is hard to control all those factors, but I think there is room for improvement.

      We have substantially expanded the benchmarking and discussion of competing methods. These additions more clearly demonstrate TopoMetry’s advantages and robustness compared to widely adopted alternatives. In the revised manuscript:

      We now benchmark TopoMetry against 68 diverse single-cell datasets, far exceeding the scope of the original version.

      We explicitly compare TopoMetry with PCA→UMAP, standalone UMAP, and scVI. These workflows represent the de facto current standard in single-cell analysis. While numerous other approaches exist, a comprehensive benchmark of every possible workflow lies beyond the scope of this study and would itself warrant a dedicated report.

      We adopt the exact same preprocessing steps for all evaluated workflows to ensure a fair comparison, except for scVI, which requires gene counts data and performs its own internal preprocessing.

      We adjust the number of PCs used for each dataset based on the currently adopted “elbow point” ad hoc.

      (2) The paper lacks experiments that validate the results. It would be beneficial to see additional evaluation settings with better-established ground truths to more strongly demonstrate the method's effectiveness.

      We agree that validation is crucial and have strengthened this aspect:

      We introduce new geometry-preservation metrics and validate that TopoMetry outperforms current de facto standards.

      We demonstrate that TopoMetry resolves well-established ground-truth structures, such as the cell cycle in pancreas development and T cell proliferation, which PCA→UMAP fails to capture (Suppl. Fig. S3).

      We validate the biological relevance of novel T cell subpopulations by linking them to TCR clonotypes and clonal expansion patterns using datasets with paired VDJ information (ECCITE-TCR, TICA).

      We show that TopoMetry faithfully recovers expected lineage trajectories in atlas-scale datasets (MOCA).

      These analyses demonstrate that TopoMetry not only preserves geometry but also recovers biologically meaningful ground-truth structures. Further experimental investigation of biological insights obtained from the presented examples exceeds the scope of the presented methodological work.

      (3) The effect of various parameters, such as those involved in k-nearest neighbors (KNN) or choosing the appropriate Laplacian operator, is not comprehensively explored. How can we ensure the analysis is not overly sensitive to these parameters?

      We now explicitly address parameter robustness and show that results are stable across a wide range of k values (30–200) in the neighborhood graph (Suppl. Fig. S1e).

      The range of possible Laplacian operators was a design choice aimed at increasing user freedom, but we agree with the reviewer that this option could confuse readers and users. TopoMetry now only uses the appropriate operator (density-normalized graph Laplacian, a.k.a. diffusion operator), reducing variability and improving usability.

      (4) Batch effects are prevalent in single-cell data. The paper does not adequately address this issue.

      Several of the datasets we analyzed include cells from multiple donors and experimental batches, and TopoMetry successfully recovers consistent biological structure across these.

      TopoMetry’s spectral scaffolds can be integrated with data integration methods such as Harmony and Scanorama, which are employed to correct the latent PCA space in current practice.

      Reviewer #2 (Recommendations for the authors):

      (1) The paper introduces technical jargon without sufficient explanation abruptly many times. This makes it difficult for readers from a biological background to follow. Even I, with a more computational background, struggled to grasp some parts.

      We thank the reviewer for this feedback and have streamlined terminology throughout the manuscript, replacing jargon with more intuitive language and providing brief explanations when technical terms are first introduced. This makes the text more accessible to both computational and biological audiences.

      (2) There is no comparison of the computational cost of this method with existing approaches, which is an important factor for practical adoption. Including a benchmarking section on this would be useful.

      We thank the reviewer for this suggestion and have now included a runtime benchmark against PCA→UMAP, PHATE, and scVI (Suppl. Fig. 1f), showing that while TopoMetry is slightly slower than PCA→UMAP, it scales more favorably than alternative geometry-aware methods (PHATE) and neural networks (scVI).

      (3) TopOMetry allows users to obtain and evaluate dozens of possible representations. However, I wonder if this could introduce a user burden, increasing uncertainty and subjectivity, as users should examine them manually. I think this should be clarified.

      We appreciate this concern and have streamlined the workflow to minimize user burden. As shown in the original manuscript, representations learned with different TopoMetry kernels and Laplacian variants converge to highly similar results. Based on this, TopoMetry now defaults to the best-performing kernel and the most appropriate Laplacian operator, yielding only two scaffold representations (fixed-time and multiscale) and corresponding visualizations rather than dozens of alternatives. This removes the need for manual selection while retaining flexibility for advanced users. In addition, we introduced a single-line command that runs the entire analysis and generates a comprehensive PDF report, allowing users to evaluate results in a standardized and user-friendly way. Together, these changes eliminate unnecessary subjectivity and ensure consistent outputs across analyses.

      (4) Formatting. There are errors in figure numbering within the main text. For instance, it should be Figure 4 instead of Figure 3 on page 11. Some figures are not concise. For example, Figure 2 contains too much text, which detracts from its visibility. I recommend trimming the figures to improve clarity. A color map is missing in Figure 2, which could help better interpret the data.

      We have thoroughly adjusted the manuscript and figures for improved visibility and clarity.

      Broader Impact and Reception

      Since our preprint, TopoMetry has been used by Hale et al. (Science, 2024), where it helped reveal morphological T cell subpopulations, and in a recent preprint by Tedeschi et al. (2025). These independent applications highlight the utility and impact of TopoMetry beyond our group, supporting its relevance to diverse biological contexts. In addition, two independent studies performing multimodal integration of RNA and TCR data (Zhang et al., 2023 and Drost et al., 2024) have identified a diversity of T cell subpopulations that resembles the clusters identified by TopoMetry using only RNA data.

    1. eLife Assessment

      This study reports the relative importance of Tie1 and Tie2 signaling for atrial versus ventricular trabeculation. It is an important study and is one of the few works to date that have carefully and simultaneously analyzed these two processes. In line with a previous study in zebrafish, the authors demonstrate key differences between atrial and ventricular trabeculation. While the imaging and quantitative data were conducted with solid and validated methodology throughout the manuscript, the work would benefit from more rigourous approaches where Tie1/2 signaling is disrupted prior to the onset of atrial/ventricular trabeculation, to allow for a more direct comparison.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Ding et al. use genetic mouse models to demonstrate that atrial trabeculation is more dependent on Tie1/Tie2 signaling than ventricular trabeculation. With additional experimentation that would support the current claims, the results may hold significant value, as atrial trabeculation remains an understudied phenomenon in cardiac biology with potential implications for atrial cardiomyopathy and atrial fibrillation.

      Strengths:

      Detailed characterization of atrial versus ventricular trabeculation across different developmental timepoints, and the use of appropriate animal models to address the scientific question at hand.

      Weaknesses:

      The authors have consistently treated mice with tamoxifen after ventricular, but not atrial, trabeculation has already started. As such, the observed cardiac phenotypes - where predominantly atrial trabeculation is affected - might be a mere consequence of the precise time window in which Tie1/2 signaling was impaired, rather than a direct measurement of its relative importance for atrial versus ventricular trabeculation. The conclusions of the paper may thus be significantly strengthened by depleting Tie1/2 signaling prior to the onset of ventricular trabeculation, as is done for atrial trabeculation.

    3. Reviewer #2 (Public review):

      Summary:

      Ding et al. examine the role of TIE1 in cardiac chamber morphogenesis using genetic mouse models targeting Tie1, Tek, or both, and analyzing endocardial cell-mediated chamber formation across multiple embryonic developmental and postnatal stages, supported by analysis of published single-cell datasets and new bulk RNA seq analyses of murine cardiac tissue. The authors find that Tie1 and Tek expression is higher in atrial than ventricular endocardial cells. Notably, endothelial Tie1 is required for atrial trabeculation at E12.5, but is less critical in ventricular trabeculation. TIE1 also acts synergistically with TIE2 during atrial trabeculation. While Tie1 deficiency alone does not cause defects at E10.5, combined heterozygous deletion of Tek disrupts both atrial and ventricular development at E10.5. This synergy is further supported by analyses at later embryonic stages and in postnatal hearts.

      Strengths:

      The study is well-designed, clearly written, and supported by high-quality figures. The performed experiments demonstrate a previously unrecognized role for Tie1 in cardiac development and identify synergistic control of cardiac morphogenesis by Tie1 and Tie2. This synergy is consistent with the previously identified roles of Tie1 and Tek in venous development and with Tie1 involvement in angiopoietin-dependent postnatal vascular and lymphatic remodeling. Together, these findings support a role for Tie1 as a contributor to Ang1-Tie2 signaling during heart development.

      Weaknesses:

      The manuscript does not include direct mechanistic studies; however, RNA seq analysis of atria and ventricles showed reduced expression of Tek, Dll1, and Notch1 upon Tie1 deficiency in developing hearts. Although previously reported mechanisms, such as TIE1-TIE2 heterodimer formation and effects on endothelial junctions, migration, or survival are discussed, no direct mechanistic experiments are performed. Addressing some of these mechanisms would have clarified the basis of Tie1-Tie2 synergy. As two distinct Tie1 models are used, including one targeting the kinase domain, the authors should state whether phenotypes differed or were similar between models.

    4. Reviewer #3 (Public review):

      Summary:

      Ding et al. investigate the roles of TIE1 and TEK (Tie2) in mouse cardiac development, with a particular focus on atrial trabeculation. The authors employ multiple genetic models, including Tie1ICDflox/flox (with Cdh5-CreERT2), a knockout-first allele (EUCOMM, Tie1 tm1a/tm1a), and a Tek deletion model.

      Based on the dataset from Feng et al. 2022 Nat Commun, the authors report increased expression of Tie1 and Tek transcripts in atrial endocardial cells compared to ventricular cells at embryonic day (E) 14.5. Loss of Tie1 leads to early atrial trabeculation defects detectable at E12.5, whereas ventricular defects appear later and are less pronounced at E14.5. Chamber-specific RNA sequencing reveals stronger transcriptional changes in atrial tissue.

      Conditional deletion of Tek results in a similar phenotype, with more pronounced atrial defects. Combined deletion of Tie1 and Tek (Tie1 ΔICD/ΔICD; Tek+/-) leads to earlier and more severe defects in both atrial and ventricular trabeculation and results in embryonic lethality around E12.5, suggesting a synergistic interaction between the two genes.

      Conditional endothelial deletion of Tie1 combined with heterozygous global Tek at later embryonic stages allows analysis at later time points and again shows more severe defects in atrial trabeculation. Postnatal analysis of this model reveals reduced heart-to-body weight ratios and potential mild atrial abnormalities.

      Strengths:

      (1) The authors address chamber-specific signaling mechanisms underlying atrial versus ventricular trabeculation, an area of high developmental and clinical relevance.

      (2) The study provides a comprehensive temporal analysis across multiple embryonic stages.

      (3) The use of multiple genetic models strengthens the overall conclusions and allows comparative interpretation.

      (4) While focusing on trabeculation, the authors also include observations on coronary vessel development, increasing the broader relevance of the work. The findings are therefore of interest to the wider cardiovascular research community.

      Weaknesses:

      (1) Timing of recombination vs. trabeculation onset

      Ventricular trabeculation begins earlier than atrial trabeculation. Since tamoxifen (in contrast to 4-hydroxytamoxifen) requires metabolic activation, Cre-mediated recombination will occur with a delay. This suggests that atrial trabeculation may be targeted before its onset, whereas ventricular trabeculation may already be underway for 2-3 days at the time of effective gene deletion.

      How do the authors account for this discrepancy in their interpretation?

      Have earlier induction time points been tested to better capture the onset of ventricular trabeculation? This limitation should be explicitly discussed.

      (2) Clarity of genetic models and experimental design

      The study employs several genetic constructs. It would improve clarity if, for each experiment, the specific genetic model and tamoxifen regimen were clearly described before presenting the results.

      (3) Tie1 tm1a/tm1a phenotype vs. known global knockout

      Previous studies (PMID: 8846781, 7596437) show that complete Tie1 loss leads to severe edema, vascular rupture, and embryonic lethality around E13.5-E14.5.

      How does the Tie1 tm1a/tm1a allele differ, given that animals appear to survive longer? Is this allele hypomorphic rather than a full knockout?

      This point requires clarification.

      (4) Limited mechanistic insight

      While the authors aim to investigate underlying mechanisms, the current study is largely descriptive and based on mRNA expression and genetic interaction analyses (Tie1/Tek co-deletion). Direct mechanistic insights into signaling pathways remain limited. However, the dataset provides a valuable foundation for future mechanistic studies, which should be more clearly acknowledged in the discussion.

    5. Author response:

      eLife Assessment

      This study reports the relative importance of Tie1 and Tie2 signaling for atrial versus ventricular trabeculation. It is an important study and is one of the few works to date that have carefully and simultaneously analyzed these two processes. In line with a previous study in zebrafish, the authors demonstrate key differences between atrial and ventricular trabeculation. While the imaging and quantitative data were conducted with solid and validated methodology throughout the manuscript, the work would benefit from more rigourous approaches where Tie1/2 signaling is disrupted prior to the onset of atrial/ventricular trabeculation, to allow for a more direct comparison.

      We thank the editors for the eLife assessment. We would like to request that the following statement be modified: “…the work would benefit from more rigourous approaches where Tie1/2 signaling is disrupted prior to the onset of atrial/ventricular trabeculation, to allow for a more direct comparison”. We request this change for the following reasons:

      We utilized two distinct genetic mouse models in this study (as summarized in Fig. 7I), comprising conventional knockouts (Tie1<sup>tm1a/tm1a</sup>, Tie1<sup>ΔICD/ΔICD</sup> and Tie1<sup>ΔICD/ΔICD</sup>;Tek<sup>+/-</sup>) and inducible gene deletion models (Tek<sup>iECKO</sup>, Tie1ICD<sup>iECKO</sup>, and Tie1ICD<sup>iECKO</sup>;Tek<sup>+/-</sup>) [1-3]. The Tie1<sup>tm1a/tm1a</sup> line is equivalent to the previously published Tie1<sup>-/-</sup mouse line, as demonstrated in our prior work and by others [1, 2, 4-6]. Therefore, the Tie1 or Tek alleles were inactivated prior to the onset of atrial and ventricular trabeculations, as shown in Fig. 1, Fig. 2, Fig. 3, Fig. 5A-D, and Supplemental Fig. 3. Based on these findings, we propose that TIE1 is differentially required for atria versus ventricle morphogenesis, and acts synergistically with TIE2 during cardiac trabeculation.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Ding et al. use genetic mouse models to demonstrate that atrial trabeculation is more dependent on Tie1/Tie2 signaling than ventricular trabeculation. With additional experimentation that would support the current claims, the results may hold significant value, as atrial trabeculation remains an understudied phenomenon in cardiac biology with potential implications for atrial cardiomyopathy and atrial fibrillation.

      Strengths:

      Detailed characterization of atrial versus ventricular trabeculation across different developmental timepoints, and the use of appropriate animal models to address the scientific question at hand.

      Weaknesses:

      The authors have consistently treated mice with tamoxifen after ventricular, but not atrial, trabeculation has already started. As such, the observed cardiac phenotypes - where predominantly atrial trabeculation is affected - might be a mere consequence of the precise time window in which Tie1/2 signaling was impaired, rather than a direct measurement of its relative importance for atrial versus ventricular trabeculation. The conclusions of the paper may thus be significantly strengthened by depleting Tie1/2 signaling prior to the onset of ventricular trabeculation, as is done for atrial trabeculation.

      We thank the reviewer for the comments.

      Regarding the timeline of gene deletion and tamoxifen treatment, we would like to provide the following clarification.

      Fig. 1-3: As described in the Methods and Materials, Tie1<sup>tm1a/tm1a</sup> is a knockout first mouse model established from EUCOMM embryonic stem cells (EPD0735-3B07) targeting Tie1 gene. Therefore, the Tie1<sup>tm1a/tm1a</sup> line is equivalent to the previously published Tie1 null mice (Tie1<sup>-/-</sup>). The Tie1<sup>Flox/Flox</sup> mouse line (with exon 8 floxed) was generated when the lacZ reporter and neo-cassette were excised using the FLPeR mice.

      Fig. 5A-D: To investigate the synergy of TIE1 and TIE2 in cardiac trabeculation, we utilized the Tek<sup>+/-</sup> and Tie1<sup>ΔICD/+</sup> mouse lines and they were crossbred to generate double mutant mice harboring a homozygous Tie1 mutation and a single null Tek allele (Tie1<sup>ΔICD/ΔICD</sup>;Tek<sup>+/-</sup>). Although no obvious defects were observed in atrial or ventricular structures following Tie1 deficiency alone at E10.5, both atria and ventricle development were disrupted in Tie1<sup>ΔICD/ΔICD</sup>;Tek<sup>+/-</sup> mutants at the same stage (Fig. 5A-D).

      Supplemental Fig. 3: To verify the role of TIE1 in atrial development, we employed alternative knockout mouse line targeting the Tie1 intracellular domain by floxing exons 15 and exon 16 (Tie1ICD<sup>Flox/Flox</sup>). Mutants harboring these null alleles are designated as Tie1<sup>ΔICD/ ΔICD</sup>. As detailed in the previous publication [2], the line is also equivalent to the previously published Tie1 null mice (Tie1<sup>-/-</sup>). The cardiac phenotypes shown in Supplemental Fig. 3 are indeed similar to those of Tie1<sup>tm1a/tm1a</sup> mutant mice.

      For the inducible knockouts targeting Tie1, Tek and both, the results are shown in Fig. 4, Fig. 5E-H, Fig. 6, Fig. 7.

      Fig. 4: As mice homozygous for Tek mutation (Tek<sup>-/-</sup>) die before E10.5 [3, 7], we performed studies using the inducible knockout line targeting Tek (Tek<sup>Flox/-</sup>;Cdh5-Cre<sup>ERT2</sup> named as Tek<sup>iECKO</sup>), as shown in Fig. 4.

      Fig. 5-7: To investigate the synergy of TIE1 and TIE2 in the cardiac trabeculation at the later stages of embryogenesis (Fig. 5E-H, Fig. 6) and the postnatal stage (Fig. 7), we used the inducible knockout models targeting Tie1/Tek, including Tie1ICD<sup>iECKO</sup> (Tie1ICD<sup>Flox/-</sup>;Cdh5-Cre<sup>ERT2</sup>) and Tie1ICD<sup>iECKO</sup>;Tek<sup>+/-</sup> (Tie1ICD<sup>Flox/-</sup>;Cdh5-Cre<sup>ERT2</sup>;Tek<sup>+/-</sup>).

      Reviewer #2 (Public review):

      Summary:

      Ding et al. examine the role of TIE1 in cardiac chamber morphogenesis using genetic mouse models targeting Tie1, Tek, or both, and analyzing endocardial cell-mediated chamber formation across multiple embryonic developmental and postnatal stages, supported by analysis of published single-cell datasets and new bulk RNA seq analyses of murine cardiac tissue. The authors find that Tie1 and Tek expression is higher in atrial than ventricular endocardial cells. Notably, endothelial Tie1 is required for atrial trabeculation at E12.5, but is less critical in ventricular trabeculation. TIE1 also acts synergistically with TIE2 during atrial trabeculation. While Tie1 deficiency alone does not cause defects at E10.5, combined heterozygous deletion of Tek disrupts both atrial and ventricular development at E10.5. This synergy is further supported by analyses at later embryonic stages and in postnatal hearts.

      Strengths:

      The study is well-designed, clearly written, and supported by high-quality figures. The performed experiments demonstrate a previously unrecognized role for Tie1 in cardiac development and identify synergistic control of cardiac morphogenesis by Tie1 and Tie2. This synergy is consistent with the previously identified roles of Tie1 and Tek in venous development and with Tie1 involvement in angiopoietin-dependent postnatal vascular and lymphatic remodeling. Together, these findings support a role for Tie1 as a contributor to Ang1-Tie2 signaling during heart development.

      Weaknesses:

      The manuscript does not include direct mechanistic studies; however, RNA seq analysis of atria and ventricles showed reduced expression of Tek, Dll1, and Notch1 upon Tie1 deficiency in developing hearts. Although previously reported mechanisms, such as TIE1-TIE2 heterodimer formation and effects on endothelial junctions, migration, or survival are discussed, no direct mechanistic experiments are performed. Addressing some of these mechanisms would have clarified the basis of Tie1-Tie2 synergy. As two distinct Tie1 models are used, including one targeting the kinase domain, the authors should state whether phenotypes differed or were similar between models.

      We thank the reviewer for the comments. In this study, we have provided genetic evidence that TIE1 is differentially required for atrial versus ventricular trabeculation. Although the precise molecular mechanisms underlying TIE1 function require further investigation, we have provided compelling genetic evidence of its synergistic role with TIE2 during this process. The two genetic models targeting Tie1 (Tie1<sup>tm1a/tm1a</sup>, Tie1<sup>ΔICD/ΔICD</sup>) produced consistent cardiac and vascular phenotypes as shown in this study and our previous work [1, 2].

      Reviewer #3 (Public review):

      Summary:

      Ding et al. investigate the roles of TIE1 and TEK (Tie2) in mouse cardiac development, with a particular focus on atrial trabeculation. The authors employ multiple genetic models, including Tie1ICDflox/flox (with Cdh5-CreERT2), a knockout-first allele (EUCOMM, Tie1 tm1a/tm1a), and a Tek deletion model.

      Based on the dataset from Feng et al. 2022 Nat Commun, the authors report increased expression of Tie1 and Tek transcripts in atrial endocardial cells compared to ventricular cells at embryonic day (E) 14.5. Loss of Tie1 leads to early atrial trabeculation defects detectable at E12.5, whereas ventricular defects appear later and are less pronounced at E14.5. Chamber-specific RNA sequencing reveals stronger transcriptional changes in atrial tissue.

      Conditional deletion of Tek results in a similar phenotype, with more pronounced atrial defects. Combined deletion of Tie1 and Tek (Tie1 ΔICD/ΔICD; Tek+/-) leads to earlier and more severe defects in both atrial and ventricular trabeculation and results in embryonic lethality around E12.5, suggesting a synergistic interaction between the two genes.

      Conditional endothelial deletion of Tie1 combined with heterozygous global Tek at later embryonic stages allows analysis at later time points and again shows more severe defects in atrial trabeculation. Postnatal analysis of this model reveals reduced heart-to-body weight ratios and potential mild atrial abnormalities.

      Strengths:

      (1) The authors address chamber-specific signaling mechanisms underlying atrial versus ventricular trabeculation, an area of high developmental and clinical relevance.

      (2) The study provides a comprehensive temporal analysis across multiple embryonic stages.

      (3) The use of multiple genetic models strengthens the overall conclusions and allows comparative interpretation.

      (4) While focusing on trabeculation, the authors also include observations on coronary vessel development, increasing the broader relevance of the work. The findings are therefore of interest to the wider cardiovascular research community.

      Weaknesses:

      (1) Timing of recombination vs. trabeculation onset

      Ventricular trabeculation begins earlier than atrial trabeculation. Since tamoxifen (in contrast to 4-hydroxytamoxifen) requires metabolic activation, Cre-mediated recombination will occur with a delay. This suggests that atrial trabeculation may be targeted before its onset, whereas ventricular trabeculation may already be underway for 2-3 days at the time of effective gene deletion.

      How do the authors account for this discrepancy in their interpretation?

      Have earlier induction time points been tested to better capture the onset of ventricular trabeculation? This limitation should be explicitly discussed.

      (2) Clarity of genetic models and experimental design

      The study employs several genetic constructs. It would improve clarity if, for each experiment, the specific genetic model and tamoxifen regimen were clearly described before presenting the results.

      We thank the reviewer for the detailed and constructive comments. For studies employing the inducible gene deletion mouse models, the genetic models and tamoxifen treatment schemes have been provided in the related figures. For the rest of studies, we used the conventional knockouts targeting Tie1 and Tek (Tie1<sup>tm1a/tm1a</sup>, Tie1<sup>ΔICD/ΔICD</sup> and Tie1<sup>ΔICD/ΔICD</sup>;Tek<sup>+/-</sup>), as detailed above.

      (3) Tie1 tm1a/tm1a phenotype vs. known global knockout

      Previous studies (PMID: 8846781, 7596437) show that complete Tie1 loss leads to severe edema, vascular rupture, and embryonic lethality around E13.5-E14.5.

      How does the Tie1 tm1a/tm1a allele differ, given that animals appear to survive longer? Is this allele hypomorphic rather than a full knockout?

      This point requires clarification.

      Tie1<sup>tm1a/tm1a</sup> is equivalent to the full knockout (Tie1<sup>-/-</sup>). As demonstrated in our prior work, the Tie1<sup>ΔICD/ΔICD</sup> model produced lymphatic and blood vascular phenotypes similar to those of Tie1<sup>-/-</sup> mutants [1, 2, 5, 6].

      (4) Limited mechanistic insight

      While the authors aim to investigate underlying mechanisms, the current study is largely descriptive and based on mRNA expression and genetic interaction analyses (Tie1/Tek co-deletion). Direct mechanistic insights into signaling pathways remain limited. However, the dataset provides a valuable foundation for future mechanistic studies, which should be more clearly acknowledged in the discussion.

      We thank the reviewer for the comments. The manuscript will be revised accordingly, and a detailed response will be provided in our final submission.

      Reference

      (1) Cao, X., et al., Endothelial TIE1 Restricts Angiogenic Sprouting to Coordinate Vein Assembly in Synergy With Its Homologue TIE2. Arterioscler Thromb Vasc Biol, 2023. 43(8): p. e323-e338.

      (2) Shen, B., et al., Genetic dissection of tie pathway in mouse lymphatic maturation and valve development. Arterioscler Thromb Vasc Biol, 2014. 34(6): p. 1221-30.

      (3) Chu, M., et al., Angiopoietin receptor Tie2 is required for vein specification and maintenance via regulating COUP-TFII. Elife, 2016. 5:e21032.

      (4) Rodewald, H.R. and T.N. Sato, Tie1, a receptor tyrosine kinase essential for vascular endothelial cell integrity, is not critical for the development of hematopoietic cells. Oncogene, 1996. 12(2): p. 397-404.

      (5) D'Amico, G., et al., Loss of endothelial Tie1 receptor impairs lymphatic vessel development-brief report. Arterioscler Thromb Vasc Biol, 2010. 30(2): p. 207-9.

      (6) Qu, X., et al., Abnormal embryonic lymphatic vessel development in Tie1 hypomorphic mice. Development, 2010. 137(8): p. 1285-95.

      (7) Dumont, D.J., et al., Dominant-negative and targeted null mutations in the endothelial receptor tyrosine kinase, tek, reveal a critical role in vasculogenesis of the embryo. Genes Dev, 1994. 8(16): p. 1897-909.

    1. eLife Assessment

      This Review Article provides a scholarly, clear and well-structured review of intracranial research into the neural correlates of consciousness (NCCs). To our knowledge this is the first such review and is therefore likely to become a must-read for anyone working in the field of consciousness research. The authors discuss the difficulties that researchers must face when studying NCCs and how insights may emerge via intracranial recordings in humans. This no doubt reflects an in-depth, timely, and insightful contribution to the literature.

    2. Reviewer #1 (Public review):

      Summary:

      In this review paper, the authors describe the concept of neural correlates of consciousness (NCC) and explain how noninvasive neuroimaging methods fall short of being able to properly characterise an unconfounded NCC. They argue that intracranial research is a means to address this gap and provide a review of many intracranial neuroimaging studies that have sought to answer questions regarding the neural basis of perceptual consciousness.

      Strengths:

      The authors have provided an in-depth, timely, and scholarly contribution to the study of NCCs. First and foremost, the review surveys a vast array of literature. The authors synthesise findings such that a coherent narrative of what invasive electrophysiology studies have revealed about the neural basis of consciousness can be easily grasped by the reader. The authors also succeed in describing how single-cell recordings can interface with task-design to help mitigate the impact of confounded neural activity when searching for NCCs.

      The review is also, to the best of my knowledge, the first review to specifically target intracranial approaches to consciousness and to describe their results in a single article. This is a credit to the authors - as it becomes ever harder to apply strict tests to theories of consciousness using methods such as fMRI and M/EEG, it is important to have informative resources describing the results of human intracranial research so that theorists will have to constrain their theories further in accordance with such data. Additionally, the authors provide a compelling case for single-celled research in consciousness science, despite the dominance of theories situated at the system and circuit level of analysis. As far as the authors were aiming to provide a complete and coherent overview of intracranial approaches to the study of NCCs, I believe they have achieved their aim.

      Weaknesses:

      Overall, I feel positive about this paper. The authors have addressed my comments from my previous review and I see no significant weaknesses in the current version.

      Comment on revised version:

      No comments - congratulations to the authors!

    3. Reviewer #2 (Public review):

      Summary:

      In this work, the authors review the study of the neural correlates of consciousness (NCCs). They discuss several of the difficulties that researchers must face when studying NCCs, and argue that several of these difficulties can be alleviated by using intracranial recordings in humans.

      They describe what constitutes an NCC, and the difficulties to distinguish between an NCC proper from the prerequisites and consequences of conscious processing.

      They also describe the two main types of experimental designs used to study NCCs. These are the contrastive approach (with its report and non-report variants), and the supraliminal approach, each with their own merits and pitfalls.

      They discuss the limitations of non-invasive methods, such as fMRI, EEG and MEG, as well as the limitations of the use of invasive recordings in non-human animals.

      After setting the stage in this way, the authors provide an extensive review on the knowledge acquired by using invasive recordings in humans. This included population level measurements in vision and in other sensory modalities, as well as single neuron level studies. The authors also discuss studies of subcortical NCCs.

      The second half of this work discusses the theoretical insights gained through the use of intracranial recordings, as well as their limitations, and a perspective for future work.

      Strengths:

      This work offers an impressive review, which will serve as a useful reference document, both for newcomers to the study of NCC as for experienced researchers. The inclusion of non-visual and subcortical NCCs is of particular merit, as these have been understudied.

      Besides serving as a review, this work includes a perspective, exploring several directions to pursue for the progress of the field.

      Weaknesses:

      No major weaknesses.

      Appraisal of whether the authors achieved their aims:

      In this work, the authors have gathered an impressive review, and have discussed several important problems in the field of study of NCCs, as well as provided a perspective on how the field could move forward.

      Discussion of the likely impact of the work on the field:

      This work has the potential of becoming a must read for anyone working in the field of consciousness research.

      Comment on revised version:

      The authors have addressed all my concerns. Once again, my compliments for a nice piece of work.

    4. Reviewer #3 (Public review):

      Summary:

      This narrative review provides a clear, well-structured, and comprehensive synthesis of intracerebral recording work on the neural correlates of consciousness. It is written in an accessible manner that will be useful to a broad community of researchers, from those new to iEEG to specialists in the field.

      Strengths:

      The manuscript successfully integrates methodological and theoretical perspectives and offers a balanced overview of current sometimes contradicting evidence. As such, the manuscript is important as call for a concernted better exploration of NCCs using iEEG in the future.

      Weaknesses:

      The manuscript discusses extensively the use of "report" as a criterion for identifying conscious perception and its limitations for separating between correlates of consciousness and post consciousness processes, yet the term is not defined at the outset. The authors should specify what they mean by "report" (e.g., verbal report, nonverbal self-report, or any meta-cognitive indication of experience). Importantly, this definition should be explicitly linked to the theoretical landscape: whether the authors adopt an access-consciousness perspective in which (self) reportability is central, or whether the review also aims to address phenomenal consciousness. Making this conceptual grounding explicit at the beginning will help readers interpret the empirical work surveyed throughout the review.

      In addition, the review would benefit from an earlier introduction of the distinction between states and contents of consciousness. This distinction becomes important in the later section on anesthesia, sleep, and epileptic seizures, where the focus shifts from content-specific NCCs to alterations in global states. Presenting these definitions upfront, and briefly explaining how states and contents interact, would strengthen the coherence of the manuscript.

      Overall, this is an excellent and timely review. With clearer initial theoretical definitions of consciousness, the manuscript will offer an even stronger conceptual framework for interpreting intracerebral studies of consciousness.

      Comments on revised version:

      The current version of the manuscript is clear and complete. Kudos to the authors for their thorough revisions.

      My only remaining point concerns the definition of "report": "We define a report as any explicit behavioral response (whether verbal, manual, or otherwise) that communicates a participant's subjective state."

      It would be helpful to clarify whether this definition is intended to exclude purely internal, explicit self-reports that are not externally expressed. As currently formulated, the definition appears to require overt behavioral communication. However, this raises a conceptual issue in relation to the no-report paradigm literature, where the distinction between report, metacognitive access, and overt motor/verbal expression is precisely at stake.

      Could the authors specify whether "report" is meant to (i) be restricted to externally observable, behaviorally expressed reports, or (ii) extend to internally generated, explicit metacognitive judgments even when they are not communicated? Clarifying this point would help situate the manuscript more precisely within ongoing debates on the role of report in identifying neural correlates of consciousness.

    5. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      In this review paper, the authors describe the concept of neural correlates of consciousness (NCC) and explain how noninvasive neuroimaging methods fall short of being able to properly characterise an unconfounded NCC. They argue that intracranial research is a means to address this gap and provide a review of many intracranial neuroimaging studies that have sought to answer questions regarding the neural basis of perceptual consciousness.

      Strengths

      The authors have provided an in-depth, timely, and scholarly contribution to the study of NCCs. First and foremost, the review surveys a vast array of literature. The authors synthesise findings such that a coherent narrative of what invasive electrophysiology studies have revealed about the neural basis of consciousness can be easily grasped by the reader. The review is also, to the best of my knowledge, the first review to specifically target intracranial approaches to consciousness and to describe their results in a single article. This is a credit to the authors, as it becomes ever harder to apply strict tests to theories of consciousness using methods such as fMRI and M/EEG it is important to have informative resources describing the results of human intracranial research so that theorists will have to constrain their theories further in accordance with such data. As far as the authors were aiming to provide a complete and coherent overview of intracranial approaches to the study of NCCs, I believe they have achieved their aim.

      We appreciate the reviewer's positive feedback on our work.

      Weaknesses

      Overall, I feel positive about this paper. However, there are a couple of aspects to the manuscript that I think could be improved.

      (1) Distinguishing NCCs from their prerequisites or consequences

      This section in the introduction was particularly confusing to me. Namely, in this section, the authors' aim is to explain how intracranial recordings can help distinguish 'pure' NCCs from their antecedents and consequences. However, the authors almost exclusively describe different tasks (e.g., no-report tasks) that have been used to help solve this problem, rather than elaborating on how intracranial recordings may resolve this issue. The authors claim that no-report designs rely on null findings, and invasive recordings can be more sensitive to smaller effects, which can help in such cases. However, this motivation pertains to the previous sub-section (limits of noninvasive methods), since it is primarily concerned with the lack of temporal and spatial resolution of fMRI and M/EEG. It is not, in and of itself, a means to distinguish NCCs from their confounds.

      As such, in its current formulation, I do not find the argument that intracranial recordings are better suited to identifying pure NCCs (i.e. separating them from pre- or post-processing) convincing. To me, this is a problem solved through novel paradigms and better-developed theories. As it stands, the paper justifies my position by highlighting task developments that help to distinguish NCCs from prerequisites and consequences, rather than giving a novel argument as to why intracranial recordings outperform noninvasive methods beyond the reasons they explained in the previous section. Again, this position is justified when, from lines 505-506, the authors describe how none of the reported single-cell studies were able to dissociate NCCs from post-perceptual processing. As such, it seems as if, even with intracranial recording, NCCs and their confounds cannot be disentangled without appropriate tasks.

      The section 'Towards Better Behavioural Paradigms' is a clear attempt to address these issues and, as such, I am sure the authors share the same concerns as I am raising. Still, I remain unconvinced that the distinguishing of NCCs from pre-/post- processing is a fair motivation for using intracranial over noninvasive measures.

      We agree that distinguishing proper NCCs from their prerequisites or consequences is primarily a matter of experimental design and theoretical framework, not merely of recording modality. We did not mean to imply that intracranial recordings inherently solve this dissociation.This is now explicitly stated that at the beginning of this section. Instead, we argued that the high signal-to-noise ratio and spatiotemporal accuracy of sEEG offer a stronger "testing ground" for the null findings often relied on by no-report paradigms. This is now also further clarified in the revised section “Limits of noninvasive measures”.

      We also explicitly acknowledge, as the reviewer noted, that even the most precise recordings require careful task dissociations to distinguish NCCs from their prerequisites and consequences.

      (2) Drawing misleading conclusions from certain studies

      There are passages of the manuscript where the authors draw conclusions from studies that are not necessarily warranted by the studies they cite. For instance:

      Lines 265 - 271: "The results of these two studies revealed a complex pattern: on the one hand, HGA in the lateral occipitotemporal cortex and the ventral visual cortex correlated with stimulus strength. On the other hand, it also correlated with another factor that does not appear to play a role in visibility (repetition suppression), and did not correlate with a non-sensory factor that affects visibility reports (prior exposure). These results suggest that activity in occipitotemporal cortex regions reflecting higher-order visual processing may be a precursor to the NCC but not an NCC proper."

      It's possible to imagine a theory that would predict HGA could correlate with stimulus strength and repetition suppression, or that it would not correlate with prior exposure (e.g. prior exposure could impact response bias without affecting subjective visibility itself). The authors describe this exact ambiguity in interpretation later in the article (line 664), but in its current form, at least in line 270 (when the study is most extensively discussed), the manuscript heavily implies that HGA is not an NCC proper. This generates a false impression that intracranial recordings have conclusively determined that occipitotemporal HGA is not a pure NCC, which is certainly a premature conclusion.

      We agree that our interpretation of these studies (lines 265–271 of the previous version of the manuscript) was presented too definitively. We have modified the text (now lines 314-317) to soften this conclusion and align it with the more nuanced discussion later in the manuscript. Specifically, we now frame this as a "suggested dissociation" rather than a conclusive finding (line 730), and we explicitly acknowledge that alternative interpretations remain viable.

      Line 243: "Altogether, these early human intracranial studies indicate that early-latency visual processing steps, reflected in broadband and low gamma activity, occur irrespective of whether a stimulus is consciously perceived or not. They also identified a candidate NCC: later (>200 ms) activity in the occipitotemporal region responsible for higher-order visual processing."

      The authors claim in this section that later (>200ms) activity in occipitotemporal regions may be a candidate for an NCC. However, the Fisch et al. (2009) study they describe in support of this conclusion found that early (~150ms) activity could dissociate conscious and unconscious processing. This would suggest that it is early processing that lays claim to perceptual consciousness. The authors explicitly describe the Fisch et al results as showing evidence for early markers of consciousness (line 240: '...exhibited an early...response following recognized vs unrecognised stimuli.) Yet only a few lines later they use this to support the conclusion that a candidate NCC is 'later (>200ms) activity in the occipitotemporal region' (line 245). As such, I am not sure what conclusion the authors want me to make from these studies.

      This problem is repeated in lines 386-387: "Altogether, studies that investigated the cortical correlates of visual consciousness point to a role of neural responses starting ~250 ms after stimulus onset in the non-primary visual cortex and prefrontal cortex."

      This seems to be directly in conflict with the Fisch et al results, which show that correlates of consciousness can begin ~100ms earlier than the authors state in this passage.

      We thank the reviewer for pointing out this inconsistency. We agree that stating ">200 ms" conflicts with the findings of Fisch et al. (2009), who observed dissociations as early as ~150 ms. Our goal was to contrast the very early, stimulus-driven responses with the later responses that reflect consciousness. However, as the reviewer correctly notes, the exact "onset" of these signals varies across studies and paradigms. To address this, we have removed the specific ">200 ms" mentioned in line 245 of the previous version of the manuscript and updated the timing in line 284 to "starting 150 ms" to better reflect the results of Fisch et al. We also clarify that while the exact latency depends on the paradigm, a consistent finding is that activity representing conscious contents in higher-order visual cortex follows an initial wave of unconscious processes (lines 809-810).

      (3) Justifying single-neuron cortical correlates of consciousness

      The purpose of the present manuscript is to highlight why and how intracortical measures of neural activity can help reveal the neural correlates of perceptual consciousness. As such, in the section 'Single-neuron cortical correlates of perceptual consciousness', I think the paper is lacking an argument as to why single-neuron research is useful when searching for the NCC. Most theories of consciousness are based around circuit or system-level analyses (e.g., global ignition, recurrent feedback, prefrontal indexing, etc.) and usually do not make predictions about single cells. Without any elaboration or argument as to why single-cell research is necessary for a science of consciousness, the research described in this section, although excellent and valuable in its own right, seems out of place in the broader discussion of NCCs. A particularly strong interpretation here could be that intracranial recordings mislead researchers into studying single cells simply because it is the finest level of analysis, rather than because it offers helpful insight into the NCCs.

      It is true that many prominent theories of consciousness were developed based on macroscopic observations, largely due to the prevalence of non-invasive recordings in humans. However, we argue that recording single-unit activity is important for several reasons, and we made this clearer in the revised version. First, signals like fMRI, EEG (or even LFP) often conflate multiple distinct neural populations. SUA allows us to dissociate neurons representing the percept from neighboring neurons involved in task-related confounds (e.g., motor preparation or arousal) that would otherwise be blurred together. Therefore, some percepts might be represented by sparse coding involving a small, specific population of "concept" or "percept" cells. Electrophysiological studies in animal models reveal that various cognitive processes are encoded within neuronal subspaces that only emerge when single-unit activity is analyzed as lower-dimensional projections of the broader neural activity manifold (Mante et al., 2013; Ebitz & Hayden, 2021; Jayazeri & Afraz, 2017). Importantly, many neural computations are only discernible through the lens of population dynamics (i.e. with single neuron activity) (Vyas et al., 2021). We believe that providing high granularity through SUA recordings prevents over-aggregation of data, ensuring that even system-level theories can build on biologically accurate foundations.

      Moreover, some theories are defined at the cellular level. For instance, the Dendritic Integration Theory (Bachmann et al., 2020) posits that the integration of feedforward and feedback signals occurs at the level of individual pyramidal neurons. Without SUA, these cellular mechanisms remain untestable. Beyond spatial granularity, SUA also provides excellent temporal granularity, which is crucial for testing theories that rely on the precise timing of spikes (e.g., neural synchrony). As LFPs reflect average activity across populations, only SUA can confirm whether individual neurons lock their spikes to a specific phase, a mechanism hypothesized to bind features into a conscious whole.

      We added these points to a new section in the revised manuscript. References:

      Bachmann, T., Suzuki, M., & Aru, J. (2020). Dendritic integration theory: A thalamo-cortical theory of state and content of consciousness. Philosophy and the Mind Sciences, 1(II).

      Ebitz, R. B., & Hayden, B. Y. (2021). The population doctrine in cognitive neuroscience. Neuron, 109(19), 3055-3068.

      Jazayeri, M., & Afraz, A. (2017). Navigating the neural space in search of the neural code. Neuron, 93(5), 1003-1014.

      Mante, V., Sussillo, D., Shenoy, K. V., & Newsome, W. T. (2013). Context-dependent computation by recurrent dynamics in prefrontal cortex. nature, 503(7474), 78-84.

      Vyas, S., Golub, M. D., Sussillo, D., & Shenoy, K. V. (2020). Computation Through Neural Population Dynamics. Annual Review of Neuroscience, 43(1), 249-275.

      (4) No mention of combined fMRI-EEG research

      A minor point, but I was surprised that the authors did not mention any combined fMRI-EEG research when they were discussing the limits of noninvasive recordings. Intracortical recordings are one way to surpass the spatial and temporal resolution limits of M/EEG and fMRI respectively, but studies that combine fMRI and EEG are also an alternative means to solve this problem: by combining the spatial resolution of fMRI with the temporal resolution of EEG, researchers can - in theory - compare when and where certain activity patterns (be they univariate ERPs or multivariate patterns) arise. The authors do cite one paper (Dellert et al., 2021 JNeuro) that used this kind of setup, but they discuss it only with respect to the task and ignore the recording method. The argument for using intracranial recordings is weaker for not mentioning a viable, noninvasive alternative that resolves the same issues.

      We thank the reviewer for this point. We have added a discussion of fMRI-EEG to the "Limits of noninvasive measures" section (lines 167-171). While we acknowledge that fMRI-EEG is a powerful non-invasive tool for bridging spatial and temporal scales, we note that it relies on merging an indirect metabolic signal with a weak electrophysiological one filtered by the skull, which is computationally complex and often noisy. In contrast, intracranial recordings provide direct measures of both local field potentials and spiking activity within the same neural population, offering interpretability and signal-to-noise ratio that non-invasive combinations cannot match. In our view, this is not just an alternative to these methods, but a unique means of accessing the underlying neuronal ground truth.

      Reviewer #2 (Public review):

      Summary:

      In this work, the authors review the study of the neural correlates of consciousness (NCCs). They discuss several of the difficulties that researchers must face when studying NCCs, and argue that several of these difficulties can be alleviated by using intracranial recordings in humans.

      They describe what constitutes an NCC, and the difficulties to distinguish between an NCC proper from the prerequisites and consequences of conscious processing.

      They also describe the two main types of experimental designs used to study NCCs. These are the contrastive approach (with its report and non-report variants), and the supraliminal approach, each with its own merits and pitfalls.

      They discuss the limitations of non-invasive methods, such as fMRI, EEG and MEG, as well as the limitations of the use of invasive recordings in non-human animals.

      After setting the stage in this way, the authors provide an extensive review of the knowledge acquired by using invasive recordings in humans. This included population-level measurements in vision and in other sensory modalities, as well as single-neuron level studies. The authors also discuss studies of subcortical NCCs.

      The second half of this work discusses the theoretical insights gained through the use of intracranial recordings, as well as their limitations, and a perspective for future work.

      Strengths:

      This work offers an impressive review, which will serve as a useful reference document, both for newcomers to the study of NCC and for experienced researchers. The inclusion of non-visual and subcortical NCCs is of particular merit, as these have been understudied.

      Besides serving as a review, this work includes a perspective, exploring several directions to pursue for the progress of the field.

      We thank the reviewer for acknowledging the strength of our work.

      Weaknesses:

      The intention of the authors is to argue how some of the problems faced when studying NCCs are alleviated by the use of intracranial recordings in humans. But in some cases, the link between the problems related to the study of NCCs and the advantages of intracranial recordings over non-invasive methods is not clear.

      For example, the authors explain the difficulties in distinguishing between true NCCs from their prerequisites and consequences. This constitutes a difficult conceptual problems that plague all recording techniques. The authors don't provide a convincing explanation of how intracranial recordings offer advantages over EEG or MEG when dealing with these problems.

      We agree that the distinction between proper NCCs and their prerequisites or consequences is a fundamental challenge that affects all recording modalities. We did not intend to imply that intracranial recordings are a "silver bullet" for solving this conceptual problem in isolation, and we now explicitly state that at the beginning of this section (line 101).

      We have revised the section on "Distinguishing NCCs from their prerequisites or consequences" to clarify that intracranial recordings are a powerful tool when used in conjunction with appropriate experimental designs, rather than a standalone solution to these conceptual difficulties.

      For example, the authors explain how the use of non-report designs to rule out post-perceptual processing relies on null results, which, according to them, are harder to interpret given the low resolution of non-invasive methods. But the interpretation of null results is actually more complicated in the case of intracranial recordings. As the coverage achieved by the electrodes is sparse, if a null result is attested, it remains possible that a true effect was present in a nearby patch of cortex out of coverage.

      It is true that a null result in an intracranial study may simply reflect that the relevant neural population was not sampled by the specific electrode implantation scheme. However, we argue that interpreting null results is equally, if not more, complicated in non-invasive methods, albeit for different reasons. While M/EEG offers broader coverage, it is blind to many cortical sources because of their orientation (radial sources in MEG) or their location in deep sulci and subcortical structures. The signal-to-noise ratio of M/EEG is also much lower than that of intracranial EEG, making it more likely that null results obscure the existence of subtle effects (Parvizi & Kastner, 2018).

      To address this, we revised the manuscript to clarify that intracranial recordings provide high local certainty within the sampled regions (lines 224-227), whereas non-invasive methods provide broader coverage (lines 247-249). We now explicitly emphasize that drawing conclusions from null results based on intracranial recordings requires caution regarding electrode placement. We also point out that these approaches are complementary: M/EEG can identify large regions of interest, while sEEG can then provide high-resolution "ground truth" to confirm whether those regions are part of the NCC.

      Reference: Parvizi, J., & Kastner, S. (2018). Promises and limitations of human intracranial electroencephalography. Nature Neuroscience, 21(4), 474-483. https://doi.org/10.1038/s41593-018-0108-2

      The authors argue that the spatial resolution of intracranial recordings is better than that of EEG and MEG. While this is technically true (especially compared to EEG), the true spatial scale of the NCCs is unknown. If NCCs' span is in the mm range, then the additional spatial resolution of intracranial recordings might not be an advantage.

      We agree with the reviewer that the exact spatial scale of the NCC remains a topic of ongoing debate. However, we believe that the advantage of intracranial recordings holds true whether the NCC spans millimeters or centimeters. The main spatial limitation of non-invasive electrophysiology (M/EEG) is not just its spatial resolution but also the inverse problem. Since scalp sensors detect a mixture of signals from across the brain, different cortical configurations can produce identical scalp patterns. This makes it challenging to precisely locate the NCC or distinguish it from nearby activity (e.g., motor or attentional signals). When recording intracortically, a widespread NCC could be captured across multiple adjacent channels with high accuracy. Conversely, if the NCC is focal, it can be isolated with high spatial resolution. In either case, intracranial recordings eliminate the spatial ambiguity inherent in scalp recordings. We have revised the Introduction (lines 158-164) to clarify that the "spatial advantage" of intracranial recordings also pertains to the inverse problem, not merely to the ability to record from smaller cortical areas.

      Another factor that should be taken into consideration when assessing the spatial resolution of intracranial recordings is that while the listening zone of individual intracranial contacts is small, coverage is sparse and defined by clinical criteria (something that the authors discuss). In practice, the activity recorded by contacts is usually attributed to anatomically defined ROIs with a scale in the cm range. Given the sparse and uneven (across regions and patients) coverage afforded by intracranial recordings, the advantage of intracranial recordings in terms of spatial resolution is overstated.

      We thank the reviewer for raising this point regarding how intracranial data is often aggregated into regions of interest. We agree that if researchers generalize findings to large anatomical regions without accounting for single-channel recordings, some of the spatial benefits of intracranial recordings are indeed mitigated. We toned down some of the original claims accordingly, and acknowledged more clearly that clinical constraints of sEEG lead to sparse coverage (245-249).

      However, we maintain that even when using an ROI-based approach, intracranial recordings offer a clear advantage over non-invasive methods, in that they represent a direct measure from a specific patch of tissue, rather than a statistical estimate that may be contaminated by "leakage" from distant sources. To address the reviewer’s concern, we have updated the manuscript (lines 244-245) to emphasize the importance of relying on MNI coordinates and individual anatomy rather than solely on broad ROI labels.

      Appraisal of whether the authors achieved their aims:

      In this work, the authors have gathered an impressive review and have discussed several important problems in the field of study of NCCs, as well as provided a perspective on how the field could move forward.

      What is less clear is how the use of intracranial recordings per se holds potential to overcome problems such as the distinction between true NCCs and the prerequisites and consequences of conscious processing.

      Discussion of the likely impact of the work on the field:

      This work has the potential of becoming a must-read for anyone working in the field of consciousness research.

      Reviewer #3 (Public review):

      Summary:

      This narrative review provides a clear, well-structured, and comprehensive synthesis of intracerebral recording work on the neural correlates of consciousness. It is written in an accessible manner that will be useful to a broad community of researchers, from those new to iEEG to specialists in the field.

      Strengths:

      The manuscript successfully integrates methodological and theoretical perspectives and offers a balanced overview of current, sometimes contradicting evidence. As such, the manuscript is important as it calls for a concerted and better exploration of NCCs using iEEG in the future.

      We thank the reviewer for stating the importance of our work and its potential contribution to the field.

      Weaknesses:

      The manuscript extensively discusses the use of "report" as a criterion for identifying conscious perception and its limitations for separating between correlates of consciousness and post-consciousness processes, yet the term is not defined at the outset. The authors should specify what they mean by "report" (e.g., verbal report, nonverbal self-report, or any meta-cognitive indication of experience). Importantly, this definition should be explicitly linked to the theoretical landscape: whether the authors adopt an access-consciousness perspective in which (self) reportability is central, or whether the review also aims to address phenomenal consciousness. Making this conceptual grounding explicit at the beginning will help readers interpret the empirical work surveyed throughout the review.

      We agree that a clear definition of report is essential for the reader to interpret the empirical findings presented. We have added a definition to the Introduction (lines 108-111), specifying that we use "report" to refer to any explicit behavioral response (whether verbal, manual, or otherwise) that communicates a subject’s subjective state.

      Regarding the conceptual distinction between Phenomenal and Access consciousness, we refer to recent work from some of the co-authors (Mudrik et al., 2025), which suggests that P and A should not be seen as two types of consciousness, but rather as two necessary conditions for conscious experience. While a full discussion of this distinction is beyond the scope of this review, we now clearly state that our focus is on identifying neural activity that reflects the subjective experience itself, regardless of the downstream requirements of report.

      Reference: Mudrik, L., Faivre, N., Pitts, M., & Schurger, A. (2025). On a confusion about there being two types of consciousness. Trends in Cognitive Sciences. https://doi.org/10.1016/j.tics.2025.11.012

      In addition, the review would benefit from an earlier introduction of the distinction between states and contents of consciousness. This distinction becomes important in the later section on anaesthesia, sleep, and epileptic seizures, where the focus shifts from content-specific NCCs to alterations in global states. Presenting these definitions upfront and briefly explaining how states and contents interact would strengthen the coherence of the manuscript.

      We agree that clarifying the distinction between contents and levels of consciousness early on provides a stronger framework for the paper.

      We have added a brief clarification in the Introduction (lines 63-76): "It is also helpful to distinguish between levels of consciousness, defined as a global level of arousal or wakefulness (e.g., being awake vs. under anesthesia), and the contents of consciousness, defined as the specific subjective experiences one has while conscious (e.g., perceiving a visual stimulus; Bayne et al., 2016; Laureys, 2005). While the majority of this review focuses on 'content-specific' NCCs, the two dimensions are intrinsically linked, as global states typically set the conditions for the occurrence of specific conscious contents."

      Overall, this is an excellent and timely review. With clearer initial theoretical definitions of consciousness, the manuscript will offer an even stronger conceptual framework for interpreting intracerebral studies of consciousness.

      We thank the reviewer again for this highly positive assessment of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I would like to reiterate that I believe this is a very scholarly piece of writing, and I congratulate the authors on producing such a useful and timely manuscript. Below, I suggest just a few ways the authors may resolve some of the issues I raised in the public review. However, I would like to emphasise that these are merely suggestions - the authors may think of different and better ways to address these comments that are more in line with either their thinking or writing style, and I would certainly encourage the authors to follow their own preferences if they feel they are at odds with my suggestions.

      For the longer comment questioning whether intracranial recordings are really a way to isolate NCCs from their pre- and post-processing, there are two ways the authors could resolve this. One is that they collapse the section distinguishing NCCs from their prerequisites and consequences into the previous section regarding limits of noninvasive measures. For instance, they could make the point that null results are easier to interpret with intracranial recordings in this previous section. Then they could discuss how specific intracranial studies have been able to resolve questions of pre-/post- processing confounds when they introduce studies later in the manuscript. At the moment, the Distinguishing NCCs from their prerequisites and consequences section, at least to me, undermines the argument of why intracranial recordings are important because it spends too much time describing how tasks are the core component of isolating pure NCCs, and not the recording method.

      Alternatively, the authors could keep the structure as it is. In this case, I would urge the authors to emphasise the role of intracortical recordings here and to make the argument that this is a problem that intracortical recordings (rather than novel tasks) can solve more convincingly. Citing specific studies that combined intracortical recordings with no-report paradigms and emphasising how the invasive recording allowed the researchers to reach a conclusion that would not have been possible with noninvasive measures would also be helpful.

      We thank the reviewer for these useful suggestions and agree that we would not want readers to take from this paper that design issues can be fixed by using invasive recordings. Because confounding issues are crucial in research on the NCC, we believe it is important to include a section on this topic in the Introduction. However, as we explained in our response to the public review, we revised the section introducing Human intracranial electrophysiology to reflect that intracranial recordings are a complementary tool that improves the interpretability of no-report paradigms, rather than a “silver bullet” solution for confound issues. We also explicitly say now that this problem is relevant to all techniques in the study of consciousness, including intracranial recordings (line 101). Additionally, based on the reviewer’s suggestion, we have added a more detailed explanation of how studies that pair intracranial recordings with no-report paradigms provide a unique insight in the Temporal Insights section (lines 822-823).

      For my comment: Drawing misleading conclusions from certain studies, I think the public review speaks for itself. I would recommend that the authors make sure they are drawing correct conclusions from the studies they cite, and make clear from the outset where there is ambiguity in interpretation.

      We thank the reviewer for bringing these ambiguities to our attention. As explained in the response to the public review, we have modified the text accordingly.

      Finally, with regard to the single-cell analyses, I would imagine that most readers will share at least some scepticism around single neurons being the appropriate level of analysis for revealing the basis of perceptual experience. As such, I think it would strengthen the manuscript greatly if the authors could provide a brief argument as to how such work can either inform theories of consciousness or contribute more generally to the study of NCCs, given that the field and its theories are mostly biased towards studying system-level neural processes. I think single-cell analyses are extremely valuable to NCC research, and the authors have a good opportunity to frame these studies accordingly.

      We agree. As detailed in the response to the public review, we now specify (1) how a higher level of granularity in electrophysiological measurements can distinguish between awareness-related signals and confounds, (2) that these measurements provide an opportunity to study neuronal population dynamics where various cognitive processes have been shown to emerge in animals and (3) that single-neuron measurements are necessary to test predictions of theories that are defined at the cellular level

      Reviewer #2 (Recommendations for the authors):

      Recommendations for improving the writing and presentation:

      My compliments for having written an impressive review. Overall, I think that this is a beautiful piece of work that will be of great use to the community. My only concern is that the advantages of intracranial recordings over non-invasive methods in solving the difficulties faced in the study of NCCs are overstated.

      Here I provide more precise comments for your consideration.

      (1) On page 5, lines 100 to 102, you argue that "Scalp EEG and MEG have limitedanatomical resolution due to the overlap of deep and superficial brain signals at the scalp level and, in the case of EEG, the scattering of the adjacent electrical signals through the scalp". It would be good to provide precise estimates of the spatial resolutions of EEG, MEG and intracranial recordings, with accompanying references. Consider also that MEG is relatively insensitive to deep sources. I recommend this paper: Piastra et al. 2020 https://onlinelibrary.wiley.com/doi/10.1002/hbm.25272

      We thank the reviewer once again for their positive evaluation of our work. As detailed in the response to the public reviews, we now clarify that intracranial recordings provide high local certainty within the sampled regions (lines 224-227), whereas non-invasive methods provide broader coverage (lines 247-249). We thank the reviewer for their additional suggestions and have clarified our concern about the anatomical conclusions that can be drawn from scalp EEG and MEG data (lines 158-164).

      (2) On page 11, you describe work showing that activity in the occipitotemporal cortex mightreflect a precursor to consciousness, but not an NCC proper, except for the case of faces, in which the fusiform seems to behave like a true NCC. Could you discuss how these seemingly contradictory results could be reconciled?

      One possibility is that activity in some parts of the occipitotemporal cortex instantiates content-specific NCCs, i.e., correlates that are only specific to certain stimulus types (in this case: faces), while activity in other parts instantiates precursors of the NCCs. Because faces have been extensively studied, we might have uncovered the content-specific NCCs for these stimuli but not for others. This is now discussed in the text on lines 342-344. Based on reviewer 1’s suggestion, we have also toned down our claim about occipitotemporal activity being a precursor to the NCC.

      (3) From line 322, you start to discuss connectivity analyses. Adding a subheading mightimprove readability.

      We appreciate the suggestion; however, adding a subheading to a single paragraph would require restructuring the entire section, which could disrupt the flow. We believe the current format maintains clarity and cohesion.

      (4) In line 329, you write "It remains unclear to what extent these connectivity patterns reflectpost-perceptual processing and how the signals associated with perceptual consciousness in the occipitotemporal cortex interact with frontoparietal regions." But it's not clear why this is the case.

      We meant to make two separate points: (1) these studies did not control for report-related activity using no-report paradigms and (2) there has been no investigation so far of the interaction between occipitotemporal and frontoparietal signals associated with perceptual consciousness. These two points have been clarified in the text (lines 378-381).

      (5) In line 692, it would be good to clarify that Pereira 2021 is a single-neuron study.

      This has been clarified in the text.

      (6) The phrase "more research/work is needed" is repeated several times.

      Thank you for pointing this out. To avoid redundancy, we have deleted the second mention of this phrase.

    6. eLife Assessment

      The authors provide a scholarly review of intracranial research into the neural correlates of consciousness (NCCs). To our knowledge, this is the first such review, and it therefore may become a must-read for anyone working in the field of consciousness research. It is not so persuasive that intracranial recordings are better suited to identifying pure NCCs than other methods, which appears a problem instead solved through novel paradigms and better-developed theories - but this no doubt reflects an in-depth, timely, and insightful contribution to the literature.

    7. Reviewer #1 (Public review):

      Summary

      In this review paper, the authors describe the concept of neural correlates of consciousness (NCC) and explain how noninvasive neuroimaging methods fall short of being able to properly characterise an unconfounded NCC. They argue that intracranial research is a means to address this gap and provide a review of many intracranial neuroimaging studies that have sought to answer questions regarding the neural basis of perceptual consciousness.

      Strengths

      The authors have provided an in-depth, timely, and scholarly contribution to the study of NCCs. First and foremost, the review surveys a vast array of literature. The authors synthesise findings such that a coherent narrative of what invasive electrophysiology studies have revealed about the neural basis of consciousness can be easily grasped by the reader. The review is also, to the best of my knowledge, the first review to specifically target intracranial approaches to consciousness and to describe their results in a single article. This is a credit to the authors, as it becomes ever harder to apply strict tests to theories of consciousness using methods such as fMRI and M/EEG it is important to have informative resources describing the results of human intracranial research so that theorists will have to constrain their theories further in accordance with such data. As far as the authors were aiming to provide a complete and coherent overview of intracranial approaches to the study of NCCs, I believe they have achieved their aim.

      Weaknesses

      Overall, I feel positive about this paper. However, there are a couple of aspects to the manuscript that I think could be improved.

      (1) Distinguishing NCCs from their prerequisites or consequences

      This section in the introduction was particularly confusing to me. Namely, in this section, the authors' aim is to explain how intracranial recordings can help distinguish 'pure' NCCs from their antecedents and consequences. However, the authors almost exclusively describe different tasks (e.g., no-report tasks) that have been used to help solve this problem, rather than elaborating on how intracranial recordings may resolve this issue. The authors claim that no-report designs rely on null findings, and invasive recordings can be more sensitive to smaller effects, which can help in such cases. However, this motivation pertains to the previous sub-section (limits of noninvasive methods), since it is primarily concerned with the lack of temporal and spatial resolution of fMRI and M/EEG. It is not, in and of itself, a means to distinguish NCCs from their confounds.

      As such, in its current formulation, I do not find the argument that intracranial recordings are better suited to identifying pure NCCs (i.e. separating them from pre- or post-processing) convincing. To me, this is a problem solved through novel paradigms and better-developed theories. As it stands, the paper justifies my position by highlighting task developments that help to distinguish NCCs from prerequisites and consequences, rather than giving a novel argument as to why intracranial recordings outperform noninvasive methods beyond the reasons they explained in the previous section. Again, this position is justified when, from lines 505-506, the authors describe how none of the reported single-cell studies were able to dissociate NCCs from post-perceptual processing. As such, it seems as if, even with intracranial recording, NCCs and their confounds cannot be disentangled without appropriate tasks.

      The section 'Towards Better Behavioural Paradigms' is a clear attempt to address these issues and, as such, I am sure the authors share the same concerns as I am raising. Still, I remain unconvinced that the distinguishing of NCCs from pre-/post- processing is a fair motivation for using intracranial over noninvasive measures.

      (2) Drawing misleading conclusions from certain studies

      There are passages of the manuscript where the authors draw conclusions from studies that are not necessarily warranted by the studies they cite. For instance:

      Lines 265 - 271: "The results of these two studies revealed a complex pattern: on the one hand, HGA in the lateral occipitotemporal cortex and the ventral visual cortex correlated with stimulus strength. On the other hand, it also correlated with another factor that does not appear to play a role in visibility (repetition suppression), and did not correlate with a non-sensory factor that affects visibility reports (prior exposure). These results suggest that activity in occipitotemporal cortex regions reflecting higher-order visual processing may be a precursor to the NCC but not an NCC proper."

      It's possible to imagine a theory that would predict HGA could correlate with stimulus strength and repetition suppression, or that it would not correlate with prior exposure (e.g. prior exposure could impact response bias without affecting subjective visibility itself). The authors describe this exact ambiguity in interpretation later in the article (line 664), but in its current form, at least in line 270 (when the study is most extensively discussed), the manuscript heavily implies that HGA is not an NCC proper. This generates a false impression that intracranial recordings have conclusively determined that occipitotemporal HGA is not a pure NCC, which is certainly a premature conclusion.

      Line 243: "Altogether, these early human intracranial studies indicate that early-latency visual processing steps, reflected in broadband and low gamma activity, occur irrespective of whether a stimulus is consciously perceived or not. They also identified a candidate NCC: later (>200 ms) activity in the occipitotemporal region responsible for higher-order visual processing."

      The authors claim in this section that later (>200ms) activity in occipitotemporal regions may be a candidate for an NCC. However, the Fisch et al. (2009) study they describe in support of this conclusion found that early (~150ms) activity could dissociate conscious and unconscious processing. This would suggest that it is early processing that lays claim to perceptual consciousness. The authors explicitly describe the Fisch et al results as showing evidence for early markers of consciousness (line 240: '...exhibited an early...response following recognized vs unrecognised stimuli.) Yet only a few lines later they use this to support the conclusion that a candidate NCC is 'later (>200ms) activity in the occipitotemporal region' (line 245). As such, I am not sure what conclusion the authors want me to make from these studies.

      This problem is repeated in lines 386-387: "Altogether, studies that investigated the cortical correlates of visual consciousness point to a role of neural responses starting ~250 ms after stimulus onset in the non-primary visual cortex and prefrontal cortex."

      This seems to be directly in conflict with the Fisch et al results, which show that correlates of consciousness can begin ~100ms earlier than the authors state in this passage.

      (3) Justifying single-neuron cortical correlates of consciousness

      The purpose of the present manuscript is to highlight why and how intracortical measures of neural activity can help reveal the neural correlates of perceptual consciousness. As such, in the section 'Single-neuron cortical correlates of perceptual consciousness', I think the paper is lacking an argument as to why single-neuron research is useful when searching for the NCC. Most theories of consciousness are based around circuit or system-level analyses (e.g., global ignition, recurrent feedback, prefrontal indexing, etc.) and usually do not make predictions about single cells. Without any elaboration or argument as to why single-cell research is necessary for a science of consciousness, the research described in this section, although excellent and valuable in its own right, seems out of place in the broader discussion of NCCs. A particularly strong interpretation here could be that intracranial recordings mislead researchers into studying single cells simply because it is the finest level of analysis, rather than because it offers helpful insight into the NCCs.

      (4) No mention of combined fMRI-EEG research

      A minor point, but I was surprised that the authors did not mention any combined fMRI-EEG research when they were discussing the limits of noninvasive recordings. Intracortical recordings are one way to surpass the spatial and temporal resolution limits of M/EEG and fMRI respectively, but studies that combine fMRI and EEG are also an alternative means to solve this problem: by combining the spatial resolution of fMRI with the temporal resolution of EEG, researchers can - in theory - compare when and where certain activity patterns (be they univariate ERPs or multivariate patterns) arise. The authors do cite one paper (Dellert et al., 2021 JNeuro) that used this kind of setup, but they discuss it only with respect to the task and ignore the recording method. The argument for using intracranial recordings is weaker for not mentioning a viable, noninvasive alternative that resolves the same issues.

    8. Reviewer #2 (Public review):

      Summary:

      In this work, the authors review the study of the neural correlates of consciousness (NCCs). They discuss several of the difficulties that researchers must face when studying NCCs, and argue that several of these difficulties can be alleviated by using intracranial recordings in humans.

      They describe what constitutes an NCC, and the difficulties to distinguish between an NCC proper from the prerequisites and consequences of conscious processing.

      They also describe the two main types of experimental designs used to study NCCs. These are the contrastive approach (with its report and non-report variants), and the supraliminal approach, each with its own merits and pitfalls.

      They discuss the limitations of non-invasive methods, such as fMRI, EEG and MEG, as well as the limitations of the use of invasive recordings in non-human animals.

      After setting the stage in this way, the authors provide an extensive review of the knowledge acquired by using invasive recordings in humans. This included population-level measurements in vision and in other sensory modalities, as well as single-neuron level studies. The authors also discuss studies of subcortical NCCs.

      The second half of this work discusses the theoretical insights gained through the use of intracranial recordings, as well as their limitations, and a perspective for future work.

      Strengths:

      This work offers an impressive review, which will serve as a useful reference document, both for newcomers to the study of NCC and for experienced researchers. The inclusion of non-visual and subcortical NCCs is of particular merit, as these have been understudied.

      Besides serving as a review, this work includes a perspective, exploring several directions to pursue for the progress of the field.

      Weaknesses:

      The intention of the authors is to argue how some of the problems faced when studying NCCs are alleviated by the use of intracranial recordings in humans. But in some cases, the link between the problems related to the study of NCCs and the advantages of intracranial recordings over non-invasive methods is not clear.

      For example, the authors explain the difficulties in distinguishing between true NCCs from their prerequisites and consequences. This constitutes a difficult conceptual problems that plague all recording techniques. The authors don't provide a convincing explanation of how intracranial recordings offer advantages over EEG or MEG when dealing with these problems.

      For example, the authors explain how the use of non-report designs to rule out post-perceptual processing relies on null results, which, according to them, are harder to interpret given the low resolution of non-invasive methods. But the interpretation of null results is actually more complicated in the case of intracranial recordings. As the coverage achieved by the electrodes is sparse, if a null result is attested, it remains possible that a true effect was present in a nearby patch of cortex out of coverage.

      The authors argue that the spatial resolution of intracranial recordings is better than that of EEG and MEG. While this is technically true (especially compared to EEG), the true spatial scale of the NCCs is unknown. If NCCs' span is in the mm range, then the additional spatial resolution of intracranial recordings might not be an advantage.

      Another factor that should be taken into consideration when assessing the spatial resolution of intracranial recordings is that while the listening zone of individual intracranial contacts is small, coverage is sparse and defined by clinical criteria (something that the authors discuss). In practice, the activity recorded by contacts is usually attributed to anatomically defined ROIs with a scale in the cm range. Given the sparse and uneven (across regions and patients) coverage afforded by intracranial recordings, the advantage of intracranial recordings in terms of spatial resolution is overstated.

      Appraisal of whether the authors achieved their aims:

      In this work, the authors have gathered an impressive review and have discussed several important problems in the field of study of NCCs, as well as provided a perspective on how the field could move forward.

      What is less clear is how the use of intracranial recordings per se holds potential to overcome problems such as the distinction between true NCCs and the prerequisites and consequences of conscious processing.

      Discussion of the likely impact of the work on the field:

      This work has the potential of becoming a must-read for anyone working in the field of consciousness research.

    9. Reviewer #3 (Public review):

      Summary:

      This narrative review provides a clear, well-structured, and comprehensive synthesis of intracerebral recording work on the neural correlates of consciousness. It is written in an accessible manner that will be useful to a broad community of researchers, from those new to iEEG to specialists in the field.

      Strengths:

      The manuscript successfully integrates methodological and theoretical perspectives and offers a balanced overview of current, sometimes contradicting evidence. As such, the manuscript is important as it calls for a concerted and better exploration of NCCs using iEEG in the future.

      Weaknesses:

      The manuscript extensively discusses the use of "report" as a criterion for identifying conscious perception and its limitations for separating between correlates of consciousness and post-consciousness processes, yet the term is not defined at the outset. The authors should specify what they mean by "report" (e.g., verbal report, nonverbal self-report, or any meta-cognitive indication of experience). Importantly, this definition should be explicitly linked to the theoretical landscape: whether the authors adopt an access-consciousness perspective in which (self) reportability is central, or whether the review also aims to address phenomenal consciousness. Making this conceptual grounding explicit at the beginning will help readers interpret the empirical work surveyed throughout the review.

      In addition, the review would benefit from an earlier introduction of the distinction between states and contents of consciousness. This distinction becomes important in the later section on anaesthesia, sleep, and epileptic seizures, where the focus shifts from content-specific NCCs to alterations in global states. Presenting these definitions upfront and briefly explaining how states and contents interact would strengthen the coherence of the manuscript.

      Overall, this is an excellent and timely review. With clearer initial theoretical definitions of consciousness, the manuscript will offer an even stronger conceptual framework for interpreting intracerebral studies of consciousness.

    1. eLife Assessment

      This important study identifies a physiologically relevant interaction between LRRK2 and drebrin, an actin-binding protein crucial for neuronal structure. A solid body of evidence, including multiple cell models, highlights the complexities of how modifiers like BDNF intersect with LRRK2-kinase dependent function, and that many modifiers between AKT and LRRK2 in different cellular pathways are yet to be identified and understood.

    2. Reviewer #1 (Public review):

      Summary:

      LRRK2 protein is familially linked to Parkinson's disease by the presence of several gene variants that confer a gain-of-function effect on LRRK2 kinase activity.

      The authors examine the effects of BDNF stimulation in immortalized neuron-like cells, cultured mouse primary neurons, hIPSC-derived neurons, and brain tissue from genetically modified mice. They examine a LRRK2 regulatory phosphorylation residue, LRRK2 binding relationships, other kinase phosphorylation status, and measures of synaptic structure and function.

      Strengths:

      The study addresses an important research question: how does a PD-linked protein interact with other proteins, and contribute to responses to a well-characterized neuronal signalling pathway involved in the regulation of synaptic function and cell health.

      They employ a range of models and techniques to convincingly demonstrate that BDNF stimulation alters LRRK2 phosphorylation at pS935 and binding to many proteins. Several independent data sets lead to some exciting conclusions.<br /> In this re-revised manuscript, some aspects are very convincing and well validated e.g., drebrin binding to LRRK2, increased by BDNF, and reduced LRRK2 protein levels in young (but not mature) drebrin KO mice. A phosphoproteomic analysis of PD mutant Knock-in mouse brain is included. Overall, the links between LRRK2, LRRK2 activity, and the changes to synaptic molecules, structures, and activity are intriguing.

      Weaknesses:

      Enthusiasm for the title claim that "LRRK2 regulates synaptic function through BDNF signalling" is tempered by disconnected results across different model systems and inconsistent alterations upon kinase phosphorylation in SHSY5Y cell line and primary neurons. Exciting conclusions are sometimes not consistently supported by the data and/or only conducted in one of the models.

      BDNF increasing pS935 LRRK2 is quite well supported in cell lines, as is BDNF regulation of derbrin-LRRK2 binding. However, there is a lack of connection between this result and subsequent alterations to LRRK2 substrates e.g., phosphorylation of Rab GTPases, especially in neurons. Interesting omic data sets are provided, but with very little or no validation. For example, only drebrin protein was assessed in BDNF treatment omic, and the phosphoproteomic analysis of PD mutant Knock-in mouse is stand alone with no validation and G2019S is not explored elsewhere in the study.

      The major disconnect this reviewer struggles with is the conclusion that the quite clear data in SHSY5Y cells is the same as that from neurons regarding BDNF / LRRK2 and ERK / Akt. It seems they are not.

      ERK and Akt phosphorylation by BDNF is absent in CRISPR KO SHSY5Y cells.<br /> This conclusion is at odds with interpretation of neuronal data. To explain; in div14 neurons, BDNF's transient increase in pLRRK2 is seen and strongly prevented by MLi2. BDNF also increased pAkt & pERK1&2 in WT... but also in LRRK2 KO. Furthermore, this happened in the presence of MLi2 in WT despite no pLRRK2 increase. While the 5min BDNF induced increase to pAkt appears reduced in LKO, the same time BDNF in LKO with MLi2 is as high as WT (in these unquantified examples) and ERK is almost identical. This is described as "significantly reduced" but I see no replicates or quantification, and face value assessment of the blot argues against this.<br /> Thus, there is little or no evidence supporting that LRRK2 activity is involved in BDNF-stimulated increases in pAkt or pERK, upstream, in neurons as neither Mli2 nor KO prevented this.

      Synapse markers increased in WT neuron with BDNF treatment which did not happen in LKO neurons. So this process requires pLRRK2, but is unrelated to pAkt or pERK (which do still go up with BDNF in KO)? Similarly, an increase in synaptic activity in WT hiPSC neurons in response to BDNF seems lost in LRRK2 KO hiPSC neurons, although their activity is already increased and depending on the age of the cells the effects were different. Both of these experiments lack supporting evidence by other measures e.g., LRRK2 inhibition effects on BDNF-induced increases in WT and parallel biochemistry of p'd LRRK2, Akt, ERK in WT & KO.

      LRRK2 activating Akt1 has been published before (e.g., Ohta 2011 - not cited), but Ohta also conclude that LRRK2 gain of function mutations (more LRRK2 kinase activity) were associated with a reduced ability of LRRK2 to bind AND phosphorylate Akt at the same residue, in contradiction to the mechanism proposed here? This should be discussed. Here the authors also conclude Akt is Upstream of LRRK2. However, it appears from the data here in neurons that pLRRK2 increases in response to BDNF are separate from BDNF signalling to Akt.

      Of note, in comparison to bTubulin control, LKO total Akt levels appear consistently higher in this single example blot; a large increase in Akt would skew the ratio down, while absolute levels of pAkt (probably the most important matter for an active enzyme - what is the ratio against total protein stain) are similar or increased. These are major problems for the conclusions as presented.

      BDNF increased mEPSC frequency in hIPSC neurons; which didn't happen in LKO, which already had high frequency. Earlier in the manuscript BDNF is shown to alter synapse number in WT but not LKO mouse neurons, but no increase in synapse number was seen following BDNF treatment in any WT or LKO hiPSC neurons +/- BDFN.

      If we are to assume that the WT neurons have LRRK2 (not demonstrated), and that LRRK2 KO neurons have similar drebrin (not demonstrated) it is unclear how to interpret this result in the model of BDNF-LRRK2 being upstream of pERK/Akt. There is no evidence that the BDNF increase in WT is blocked by LRRK2 inhibition, nor has it been associated with changes (or not) to pAkt or ERK1, which would be expected in both WT and KO based on Figure 4C.

      There are many reports of acute and longer term BDNF application increasing event frequency in brain slices & primary neurons. Overexpression of BDNF in NPCs has also been shown to increase synapse function in hiPSC neurons derived from them. Here, BDNF has an effect on frequency in only one 6 comparisons (3 timepoints, two lines). Is it not concerning that expected BDNF effects occur at only one time point in WT, and that generally a lack of effect is more common both in WT and LKO... is this due to slow appearance of TrkB receptors and degeneration at 90 days?

      There are no other data provided to show that BDNF was having a consistent expected effect in human neurons (pAkt, pLRRK2 etc etc), and there is little to link between this data and that in previous figures of the study.

      The discussion of some of the weaknesses is mostly fair, asides the disparities noted above which are not.

    3. Reviewer #2 (Public review):

      The data show that BDNF regulates the PD-associated kinase LRRK2, they place LRRK2 within well-described BDNF pathways biochemically, and they show that LRRK2 can play a role mediating BDNF-driven synaptic outcomes at excitatory synapses. The chief strength is that the data provide a potential focal point for multiple observations that have been made across many labs. The findings will be of broad interest because LRRK2 has emerged as a protein that is likely to be part of Parkinson's pathology and its normal and pathological actions remain poorly understood.

      A major strength of the study is the multiple approaches that were used (biochemistry, bioinformatics, light and electron microscopy and electrophysiology) across different experimental models (cells, primary neurons, human neurons, mice) to identify and examine the impact of BDNF on LRRK2 signaling and functions. Noteworthy is also the employment of LRRK2KO preparations to validate outcomes and to place LRRK2 actions up or downstream.

      The demonstration that LRRK2 and drebrin interact directly is important and suggests that other interacting proteins identified biochemically and bioinformatically in the paper will be important to pursue.

    4. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This important work begins to understand how BDNF regulates the phosphorylation and activity of LRRK2. The overall strength of evidence has been assessed as compelling, though some claims are only partially supported. The work will be of interest for those that might pursue specific LRRK2 interactions and mutational effects on these pathways as the work continues to develop.

      We thank the editors and reviewers for the constructive feedback. We have revised the manuscript to improve clarity, strengthen statistical analysis and increase the western blot sample size in drebrin KO mice.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      LRRK2 protein is familially linked to Parkinson's disease by the presence of several gene variants that all confer a gain-of-function effect on LRRK2 kinase activity.

      The authors examine the effects of BDNF stimulation in immortalized neuron-like cells, cultured mouse primary neurons, hIPSC-derived neurons, and brain tissue from genetically modified mice. They examine a LRRK2 regulatory phosphorylation residue, LRRK2 binding relationships, and measures of synaptic structure and function.

      Strengths:

      The study addresses an important research question: how does a PD-linked protein interact with other proteins, and contribute to responses to a well-characterized neuronal signalling pathway involved in the regulation of synaptic function and cell health.

      They employ a range of good models and techniques to fairly convincingly demonstrate that BDNF stimulation alters LRRK2 phosphorylation and binding to many proteins. In this revised manuscript, aspects are well validated e.g., drebrin binding, but there is a disconnect between these findings and alterations to LRRK2 substrates. A convincing phosphoproteomic analysis of PD mutant Knock-in mouse brain is included. Overall the links between LRRK2, LRRK2 activity, and the changes to synaptic molecules, structures, and activity are intriguing.

      We thank this Reviewer for appreciating our work including the new experiments performed during the revisions.

      Weaknesses:

      The data sets remain disjointed, conclusions are sweeping, and not always in line with what the data is showing. Validation of 'omics' data is light. Some inconsistencies with the major conclusions are ignored. Several of the assays employed (western blotting especially) are underpowered, findings key to their interpretation are addressed in only one or other of the several models employed, and supporting observations are lacking.

      We understand the Reviewer’s points and agree that it is important to increase the sample size (animals) for western blot. In particular, we acknowledge that the initial experiments with the Dbn1 KO mice included only 3 mice, which was insufficient to draw any definitive conclusion on the effect, especially regarding pRab8 levels. In response to this, we have collected additional animals and repeated the experiment with N=7 wild-type and N=7 KO mice (2 months old). Despite a high degree of interindividual variability, we have confirmed that drebin KO mouse brains have reduced levels of pLRRK2 (Author response image 1). In the new figure 2H we included all the replicates (N=7+3 per genotype) for pLRRK2. However, we removed western blot for pRab8 because a new batch of pRab8 antibody did not yield specific results, making it impossible to reassess.

      Author response image 1.

      Western blot analysis of N=7 WT and N=7 drebrin KO whole brains.

      Main Conclusions of Abstract:

      (1) Increase in pLRRK2 Ser935 and pRAB after BDNF in SH-SY5Y & mouse neurons

      Well supported, but only for pLRRK2 in neurons, why not pERK pAkt & pRab?

      The response of pERK and pAKT in neurons is shown in figure 4C. We have repeatedly tried pRab (both pRab8 and pRab10) in primary neurons but with no success. In support of the difficulty in detecting pRab in primary neurons, we are not aware of studies in the literature of western blot analysis of pRabs in primary neuronal cultures. This is likely due to the high levels of PPM1H in neurons as discussed in Berndsen et at, eLife, 2019 (PMID: 31663853).

      (2) Omics Proteome remodelling of LRRK2 interactome with BDNF & different in G2019S mouse neurons.

      Supports that the phosphoproteome of G2019S is different. Drebrin interaction with LRRK2 very well supported. Link between drebrin and LRRK2 activity somewhat supported (pS935 site), but the consequence (non-specific pRab8) not supported, as there is no evidence of a change in LRRK2 substrate(s).

      As discussed above, we removed the pRab8 western blot in figure 2H as we could not confirm with the new set of mice and a new batch of pRab8 antibody.

      (3) Golgi 1 month LKO mouse altered dendritic spines, transient at 1m not older.

      Supported but very small transient change in spines, disconnected to other results (e.g., drebrin).

      We agree with the Reviewer that the observed effect is modest, still we believe it is important to report. As discussed in the discussion, one plausible explanation for the limited magnitude of the effect is functional compensation by LRRK1.

      (4) iPSC-derived neurons BDNF increases mEPSC frequency (transient at 70 not 50 or 90 days) in WT not KO "which appear to bypass this regulation through developmental compensation"

      Weak, not clear what is being bypassed.

      We reviewed the statistical analysis as described below.

      Main Conclusions Based on Old and New Figure / Data:

      (1) Increase in pLRRK2 Ser935 and pRAB after BDNF in SH-SY5Y & mouse neurons

      Well supported, but only for pLRRK2 in neurons, why not ERK Akt & Rab?

      The response of pERK and pAKT in neurons is shown in figure 4C. We have repeatedly tried pRab (both pRab8 and pRab10) in primary neurons but with no success. In support of the difficulty in detecting pRab in primary neurons, we are not aware of studies in the literature of western blot analysis of pRabs in primary neuronal cultures. This is likely due to the high levels of PPM1H in neurons as discussed in Berndsen et at, eLife, 2019 (PMID: 31663853).

      (2) BDNF promotes LRRK2 interaction with "post-synaptic actin cytoskeleton components"

      Tone down, only one postsynaptic validated - drebrin strong BUT CONTRADICTORY; link between drebrin and LRRK2 activity (pS935 site) supported, consequence (non-specific pRab8) broken, no evidence of change in LRRK2 substrate.

      As suggested we tone down the paragraph title and changed it as follow: “BDNF stimulates LRRK2 interaction with drebrin, an actin cytoskeletal-associated protein enriched at the postsynapse”. As mentioned above, pRab8 has not been incorporated.

      (3) LRRK2 G2019S striatal phosphoproteome is different from WT.

      It is different. Where is link to BDNF or Drebrin?

      We found that debrin S339 phosphorylation is 3.7 fold higher in G2019S KI mice as compared to WT, suggesting a potential functional connection between LRRK2 and drebrin. However, differences in phosphorylation do not necessarily translate into physiological effects so further validation is required. To test if BDNF can induce S339 drebrin phosphorylation in a LRRK2-dependent manner we plan an in vivo experiment where BDNF is acutely administered to WT vs G2019S-KI mice +/- MLi2 to control for LRRK2 dependency. This is an important experiment to establish the mechanistic link, though it will require sufficient time due to the necessary ethical authorization needed to administer BDNF in the mouse brain.

      (4) BDNF signaling is impaired in Lrrk2 knockout neurons

      TrkB changes seem higher in SHSY5Y. pAKT impaired, pERK not convincing. Primary neurons Akt slower but it and Erk mostly intact. MLi-2 did not block pAkt or pErk in WT or KO (higher in latter). Whatever is happening in KO, Mli-2 not really blocking effect in WT. If we are to assume that studying the KO was a means to understand LRRK2 function, the authors data should explain why we care if an effect is absent in LKO, if LRRK2 isn't doing the same job in WT?

      To further support the conclusion that this effect is reproducible and dependent on LRRK2 kinase activity acting upstream of AKT and ERK signaling, we probed the membranes shown in Figure 1H for phosphorylated and total AKT and ERK1/2. Consistent with our hypothesis, the inhibition of LRRK2 with MLi-2 significantly reduced BDNF-induced AKT and ERK1/2 phosphorylation (Author response image 2).

      Author response image 2.

      Western blot (same experiments as in figure 1) was performed using antibodies against phosphoThr202/185 ERK1/2, total ERK1/2 and phospho-Ser473 AKT, total AKT protein levels. Retinoic acid-differentiated SH-SY5Y cells stimulated with 100 ng/mL BDNF for 0, 5, 30, 60 mins. MLi-2 was used at 500 nM for 90 mins to inhibit LRRK2 kinase activity.

      BDNF increases synaptic puncta in WT not LKO (which start higher?). Is this BDNF increase blocked by LRRK2 inhibition?

      This is an important experiment that we plan to investigate in a future study.

      (5) Postsynaptic structural changes in Lrrk2 knockout neurons

      Golgi impregnation shows some very small spine changes at 1m. Not sustained over age. mRNA changes are very small (10% not even a fold... very weak and should be written as so). Derbrin levels reduced clearly at 1m, but probably also at 4 & 18. Underpowered, disconnected time course from the spine changes.

      While differences are small they have been observed in independent sets of mice through qPCR, histology, WB and TEM, supporting the consistency of the effect, although small. For clarity we rescaled the qPCR graphs to 0.

      (6) An effect on "spontaneous electrical activity" at Div70

      Weak. What is so special at 70 days that means we should be confident in the differences, or be satisfied that the other time points are legitimately ignored? These are 10-11 cells from 3 cultures assayed at 3 time points but only one is presented (rest in supplement). This should be a 2 (time) or 3 way (+culture RM) ANOVA. As it stands, in WT there is a little - no activity at 50 days, little to no at 70 days, and variable to lots or none at 90. BDNF did nothing at 50 or 90 but may have at 70. In KO low activity stable at 50 & 70, tanks at 90. BDNF would seem to have a similar effect on KO at 90 as WT at 70, but as there are only 7 cells it remains inconclusive. Thus the conclusion that BDNF signalling is broken in LKO is not well supported by the ephys data, nor is the BDNF effect in WT cells (even at the 70 day time point) shown to be susceptible to LRRK2 inhibition.

      We thank the Reviewer for suggesting a more comprehensive analysis of the data. Following this suggestion, we performed separate two-way ANOVAs (DIV × treatment) for WT and LRRK2 KO neurons. This analysis revealed significant main effects of DIV and BDNF treatment in WT neurons, indicating that synaptic activity increases with neuronal maturation and is globally enhanced by BDNF. In contrast, neither DIV nor BDNF treatment reached statistical significance in LRRK2KO neurons, and no DIV × treatment interaction was observed. These results indicate that BDNFdependent enhancement of synaptic activity is preserved in WT neurons but is lost in the absence of LRRK2. We have now incorporated this analysis into the main figure and removed the individual DIV50 and DIV90 plots from the supplementary material. We also revised the title of the last paragraph to reflect the outcome of this analysis and toned down our interpretation (page 12).

      Furthermore, we have added a paragraph to the Discussion section highlighting the limitations of this study. These include the variability observed in protein content and phosphorylation analyses by western blot, as well as the necessity to confirm the electrophysiological findings in larger datasets, including in dopaminergic neurons.

      Reviewer #2 (Public review):

      The data show that BDNF regulates the PD-associated kinase LRRK2, they place LRRK2 within welldescribed BDNF pathways biochemically, and they show that LRRK2 can play a role mediating BDNFdriven synaptic outcomes at excitatory synapses. The chief strength is that the data provide a potential focal point for multiple observations that have been made across many labs. The findings will be of broad interest because LRRK2 has emerged as a protein that is likely to be part of Parkinson's pathology and its normal and pathological actions remain poorly understood.

      We thank this Reviewer for appreciating our work and acknowledging that our findings will be of broad interest.

      A major strength of the study is the multiple approaches that were used (biochemistry, bioinformatics, light and electron microscopy and electrophysiology) across different experimental models (cells, primary neurons, human neurons, mice) to identify and examine the impact of BDNF on LRRK2 signaling and functions. Noteworthy is also the employment of LRRK2KO preparations to validate outcomes and to place LRRK2 actions up or downstream.

      Thank you to the Reviewer

      The demonstration that LRRK2 and drebrin interact directly is important and suggests that other interacting proteins identified biochemically and bioinformatically in the paper will be important to pursue.

      We agree with this statement

      Some data from different models do not fit well with one another (like mouse and human neurons). This is likely due to inherent differences in the preparations. Since different experiments were carried out on the different preps, however, it is not possible to cross compare. The lack of this information is viewed more as an open question than a cause for concern.

      We thank the Reviewer for raising this point. In response, we have added a new section to the Discussion explicitly addressing the limitations of the study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      MLi2 pretreatment experiment is nice. They state in legends BDNF treatment prior to MLi2, they mean prior MLI2 treatment. Or MLi2 pretreatment, prior to BDNF. However, this experiment is hard to interpret as it has no control (non BDNF treated) time course following MLi2, could this be (at least in part) a rebound effect produced by relief of inhibition? This should be discussed if not addressed directly by experiments.

      The non BDNF treated group represents the 0 time point. We have specified it in the figure legend. We have excluded the rephosphorylation kinetic after MLi-2 relief because pRabs increase significantly at 5 minutes, far exceeding the control levels. This observation gives us feel confidence that the effect if BDNF dependent.

      (1) "As suggested, we performed qPCR and observed that 1 month-old KO midbrain and cortex express lower levels of Dbn1 as compared to WT brains (Figure 5G). This result is in agreement with the western blot data (Figure 5H)."

      There is no Fig 5H? 5F? In 5F effect sizes are exaggerated with axes not crossing zero. There is a 10% reduction in mRNA (normally >1 or 2 fold changes would be considered biologically important?). This isn't much change, and should be presented as such. 1 month old WB in G are much more convincing of a reduction of drebrin levels, but what brain region is this from?

      We apologize for the error in the rebuttal, where we incorrectly referred to figure 5G (the correct is 5F), while what we called 5H is instead 5G. We checked the labeling in the manuscript text and it is correct.

      Following the Reviewer’s important suggestion, we rescaled all plots to start at zero. Although some differences appear relatively modest, they are statistically significant. Importantly, all brains used for qPCR analyses (N = 6 per genotype) were obtained from independent mice. In addition, independent cohorts of mice were used for spine morphology analyses (N = 3 per genotype), TEM analyses (N = 4), and western blot experiments (N = 3). Thus, the overall sample size across approaches is substantial.

      WB are from whole brain, now indicated in the figure legend.

      All blots are underpowered, especially given what appears to be an age dependent loss of drebrin in both genotypes beset by high variability

      (i) Western blots looking at pSer935 and pRab8 (pan Rab) in Dbn1 WT and knockout brains.

      "As reported and quantified in Figure 2I, we observed a significant decrease in pSer935 and a trend decrease in pRab8 in Dbn1 KO brains. This finding supports the notion that Drebrin forms a complex with LRRK2 that is important for its activity, e.g. upon BDNF stimulation."

      Non-sig data in Fig2I/H and especially Fig5G are important data but hard to interpret because the experiment is underpowered. I am surprised the authors designed studies on an n=3 western blot.

      For fig 2 this is a problem if they wish to correlate LRRK2 activity with drebrin. The KO have a clear 50% decrease in LRRK2 pS935 but no change to pRab8(pan).

      As discussed above, we increased the sample size by 7 additional mice per genotype (total of 10 brains analyzed).

      Why not look at Rab10, and certainly redo with a higher n than three. Of special confusion is the observation that the WT with the highest drebrin levels, is the animal with the lowest pS935 & pRab

      As discussed above neither pRab8 nor pRab10 returned convincing results in the new round of western blots. We acknowledge that future experiments should explore the phosphorylation levels of Rab12 which is emerging as a more reliable readout of LRRK2 kinase activity in the brain.

      (ii) "Reverse co-immunoprecipitation of YFP-drebrin full-length, N-terminal domain (1-256 aa) and Cterminal domain (256-649 aa) (plasmids kindly received from Professor Phillip R. Gordon-Weeks, Worth et al., J Cell Biol, 2013) with Flag-LRRK2 co-expressed in HEK293T cells. As shown in supplementary Fig. S2C, we confirm that YFP-drebrin binds LRRK2, with the N terminal region of drebrin appearing to be the major contributor to this interaction"

      CoIP with drebrin (and fragments) is very convincing.

      We thank the Reviewer for his/her comment/feedback

      Ephys data, presentation, and response to review is weak.

      We reanalyzed the data as suggested by the Reviewer and reviewed the text and interpretation.

      Reviewer #2 (Recommendations for the authors):

      p. 12, last paragraph. "sealing" should be "ceiling"

      We corrected the misspelled word

    1. eLife Assessment

      This important study provides solid evidence that early childhood malaria exposure affects the development of antibody responses to unrelated pathogens and vaccine-derived antigens in Kenyan children. The findings are of major public health importance and limitations of the observational study design are properly acknowledged.

    2. Reviewer #2 (Public review):

      Summary:

      The authors investigated whether early-life malaria exposure has long-term effects on immune responses to unrelated antigens. They leveraged a natural experiment in coastal Kenya where two adjacent communities (Junju and Ngerenya) experienced divergent malaria transmission patterns after 2004. Using 15 years of longitudinal data from 123 children with weekly malaria surveillance and annual serological sampling, they measured antibody responses to multiple pathogens using a protein microarray technology and ELISA.

      Strengths:

      (1) Extensive longitudinal data collection with weekly malaria surveillance, enabling precise exposure classification.

      (2) Use of a natural experiment design that allows for causal inference about malaria's immunological effects.

      (3) Broad panel of antigens tested, demonstrating generalized rather than antigen-specific effects.

      (4) Within-cohort analysis in Ngerenya controls for geographic and environmental factors.

      (5) Validation of key findings using both serologic microarray and ELISA.

      (6) Important public health implications for vaccine strategies in malaria-endemic regions.

      Weaknesses:

      (1) Due to its nature, the study lacks the ability to determine the direction of the associations found between malaria exposure and other IgG levels to unrelated pathogens.

      (2) No evaluation of the clinical Implications of the reduced IgG levels observed in the area with high malaria exposure.

      Assessment of Claims:

      The data appear to support the authors' primary claims. The strength of the evidence is limited by the observational nature of the study and the results should be interpreted in that light. Together with the currently available evidence of P. falciparum's impact on the host's immune function, this natural experiment design provides further evidence for a relationship between early malaria exposure and reduced antibody responses to other pathogens and vaccine-derived antigens. The within-Ngerenya analysis controls for geographic factors and thus enhances the quality of the evidence; there is limited physical, nutritional, and socio-economic information on factors that may have driven the observed changes.

      Impact and Utility:

      This work has fundamental implications for understanding vaccine effectiveness in malaria-endemic regions and may contribute to inform vaccination strategies. The findings, if confirmed, would suggest that children in areas of high malaria transmission may require modified immunization approaches. The dataset provides a valuable resource for future studies of malaria's immunological legacy.

      Context:

      This study builds on prior work showing acute immunosuppressive effects of malaria but uniquely attempts to demonstrate the durability of these effects years after exposure. The natural experiment design addresses limitations of previous observational studies by providing a more controlled comparison.

    3. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study sought to investigate the role that early childhood malaria exposure plays in the development of antibody responses to unrelated pathogens and vaccine-derived antigens in Kenyan children. In this natural experiment, the authors compare antibody levels among children who have been exposed to different levels of malaria transmission by using protein microarray technology. Although the findings are of importance, the evidence remains incomplete, and the analysis would benefit from a more in-depth evaluation of potential confounders. With the appropriate analysis, the findings will be of great interest for global health, immunology, and vaccine development.

      We thank the editors for highlighting the need for a more comprehensive evaluation of potential confounding. We agree that this is a critical aspect of the study and have now undertaken additional analyses to address this directly.

      The original longitudinal cohort was designed to investigate the acquisition of naturally acquired immunity to malaria and did not include systematic collection of anthropometric/nutritional, environmental or socioeconomic data, precluding direct adjustment for these factors within the primary dataset. However, to assess whether there were population-level differences in these factors, we leveraged contemporaneous hospital-based surveillance data from the same geographic regions, which includes measurements of anthropometry and nutritional status (muac, weight-for-age, and height-for-age) and detailed infection diagnostics.

      Using this independent dataset, we fitted mixed-effects regression models adjusting for age, calendar year, and concurrent infections (RSV, parainfluenza, influenza A, human metapneumovirus, OC43). For all three anthropometric indices, we found no evidence of systematic differences between children from Junju and Ngerenya. Adjusted differences were small and centred around zero (muac: −0.12, 95% CI −0.38 to 0.15, weight-for-age: −0.05, −0.28 to 0.19, height-for-age: 0.08, −0.17 to 0.33), with no consistent directional effect.

      As the longitudinal cohort was randomly selected from these underlying populations, these findings suggest that the groups were broadly comparable with respect to nutritional status and there were no differences in their exposure to the infections that were included in the analysis. We have incorporated these analyses into the revised manuscript, added a new figure focussed on this analysis -fig. 6, updated the statistical analysis and discussion sections), and believe they substantially strengthen the evidence by addressing a key source of potential confounding.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study shows that childhood malaria can weaken the antibody response to other vaccines and infections. This suggests that early exposure to P. falciparum may have a long-lasting effect on immunity, with implications for vaccine efficacy in endemic areas.

      Strengths:

      This study stands out for its longitudinal design, the use of robust immunological techniques, and the comparison between areas with different levels of malaria exposure. Its findings reveal that early malaria can weaken the response to childhood vaccines, with important implications for public health in endemic regions.

      We thank the reviewer for this comment

      Weaknesses:

      One of the study's main limitations is the lack of functional data confirming the clinical impact of the low antibody levels. Furthermore, although multiple immune responses were measured, other important components, such as cellular immunity, were not assessed. Furthermore, the results may not be generalizable to other regions.

      We thank the reviewer for this important comment and agree that the absence of functional immunological assays is a limitation of the current study. Our analysis was designed to determine whether early-life malaria exposure is associated with durable alterations in antibody responses to unrelated pathogens and vaccine antigens, rather than to establish the downstream functional consequences of these differences. As such, the study is able to demonstrate a broad and persistent attenuation of humoral responses but cannot directly determine whether the lower antibody levels observed translate into reduced neutralising capacity or diminished protection at the individual level.

      We have revised the manuscript to make this distinction more explicit. In the revised discussion, we now state that although reduced antibody titres to vaccine-preventable pathogens may have implications for long-term protection, the clinical significance of these differences remains to be established in future studies incorporating functional assays and clinical outcome data.

      Reviewer #2 (Public review):

      Summary:

      The authors investigated whether early-life malaria exposure has long-term effects on immune responses to unrelated antigens. They leveraged a natural experiment in coastal Kenya where two adjacent communities (Junju and Ngerenya) experienced divergent malaria transmission patterns after 2004. Using 15 years of longitudinal data from 123 children with weekly malaria surveillance and annual serological sampling, they measured antibody responses to multiple pathogens using a protein microarray technology and ELISA.

      Strengths:

      (1) Extensive longitudinal data collection with weekly malaria surveillance, enabling precise exposure classification.

      (2) Use of a natural experiment design that allows for causal inference about malaria's immunological effects.

      (3) Broad panel of antigens tested, demonstrating generalized rather than antigen-specific effects.

      (4) Within-cohort analysis in Ngerenya controls for geographic and environmental factors.

      (5) Validation of key findings using both serologic microarray and ELISA.

      (6) Important public health implications for vaccine strategies in malaria-endemic regions.

      We thank the reviewer for these comments

      Weaknesses:

      (1) Lack of participants' characteristics (socio-economic, nutritional, physical).

      We thank the reviewer for this important comment. We have now included a detailed summary of participant characteristics in Table 1to provide context for the study population. This includes key demographic and longitudinal variables stratified by cohort (Junju and Ngerenya), including sex distribution, age at study entry and exit, duration of follow-up, number of visits per participant, and total serum samples analysed. Detailed data on socio-economic status, nutritional status, and other environmental or physical characteristics were not consistently available across all participants and time points, and therefore could not be included. This has now been explicitly stated as a limitation in the discussion.

      (2) Somewhat limited sample size (longitudinal analysis of 123 children total), with further subdivision reducing statistical power for some analyses.

      We thank the reviewer for this important observation. The study is based on an intensively followed cohort with weekly malaria surveillance and repeated serological measurements throughout childhood, allowing detailed characterisation of early-life exposure and subsequent immune trajectories. This depth of longitudinal sampling provides resolution that is not achievable in larger cross-sectional studies. We acknowledge that subdivision of the cohort reduces statistical power for some analyses. Nevertheless, the key findings were consistent in several independent comparisons, including a reduction in antibody levels for broad panel of antigens in the malaria endemic setting, within-cohort analyses in Ngerenya that replicated this observation, and the confirmation of results generated on the protein microarray on the ELISA platform. The consistency of these findings across analytical approaches and measurement platforms reduces the likelihood that the observed effects are driven by small-sample variability. We have clarified this point in the revised discussion to emphasise that the strength of the study lies in the depth and longitudinal resolution of the data rather than the absolute sample size.

      (3) Potential confounding by unmeasured socioeconomic, nutritional, or environmental factors between communities.

      We thank the reviewer for this important point and agree that residual confounding between communities must be considered. As outlined in reponse to the editorial assesment, we have undertaken additional analyses using contemporaneous population-level data from the same regions and found no evidence of systematic differences in anthropometric indices between children from Junju and Ngerenya after accounting for age, calendar year, and concurrent infections, with effect estimates small and crossing zer. In addition, the within-Ngerenya analysis provides an internal comparison within a shared geographic and environmental setting, reducing the likelihood that unmeasured socioeconomic or environmental differences between communities account for the observed associations. The new analyses suggest that major population-level differences in nutritional status or infection burden are unlikely to explain the observed patterns. We have clarified this point in the revised discussion and explicitly acknowledge the possibility of residual confounding.

      (4) Lack of ability to determine the direction of the associations found between malaria exposure and other IgG levels to unrelated pathogens.

      We agree that, as an observational study, our analysis cannot definitively establish the direction of the association between malaria exposure and antibody responses to unrelated antigens. However, several features of the study design strengthen the inference that early-life malaria exposure contributes to the observed differences. First, malaria exposure was characterised prospectively through intensive weekly surveillance, allowing precise temporal definition of exposure in early childhood. Second, within the Ngerenya cohort, where children were exposed to different levels of malaria due to a rapid decline in transmission, those with even limited early-life exposure exhibited lower antibody responses at 10 years of age than malaria-naïve peers, despite residing in the same geographic and environmental context. In addition, we now show that these differences are not confined to a single timepoint but are evident across the full longitudinal follow-up after adjustment for age and repeated measurements. While we cannot exclude the possibility of residual confounding or bidirectional relationships, the convergence of evidence from the natural experiment design, within-cohort contrasts, and age-adjusted longitudinal analyses supports early-life malaria exposure as a key contributor to long-term differences in antibody responses. We have clarified this in the discussion.

      (5) Despite good longitudinal data, the main analysis was conducted as a cross-sectional analysis at age 10 for many comparisons, which limits the understanding of temporal dynamics.

      We thank the reviewer for highlighting this point. While age 10 was initially used as a standardised reference point for cross-sectional comparisons, the underlying dataset is longitudinal, with repeated antibody measurements across childhood. To address this more directly, we have now complemented these analyses with antigen-specific mixed-effects regression models incorporating all available longitudinal data, with adjustment for age and a random intercept for repeated measurements within individuals. These models demonstrate that the differences between cohorts are not confined to the age-10 cross-section but are evident in an age-adjusted longitudinal framework for multiple antigens. We have retained the age-10 comparisons for reference, but the primary inference is now based on the longitudinal mixed-effects analyses. These changes are reflected in the revised results and statistical analysis sections. We thank the reviewer for this astute point, which we think has substantially improved the manuscript.

      (6) Statistical analysis is limited to univariable comparisons without consideration for confounders or adjusting for multiple comparisons.

      We agree that the original analyses relied primarily on univariable comparisons. In the revised manuscript, we have extended the analytical framework to include mixed-effects regression models that account for age effects and repeated measurements within individuals. These models estimate the average age-adjusted difference in antibody responses between cohorts across the full follow-up period. We have also applied false discovery rate (FDR) correction to account for multiple antigen testing. For multiple antigens, the direction and magnitude of cohort differences remain consistent under this approach, strengthening the robustness of the findings beyond the original univariable comparisons. These analyses have been incorporated into the revised results and statistical analysis sections.

      (7) No mechanistic understanding of how early malaria exposure creates lasting immunosuppression.

      We agree that this study does not directly resolve the mechanistic basis underlying the observed long-term differences in antibody responses. The primary aim of this work was to identify and characterise durable alterations in humoral immune profiles associated with early-life malaria exposure, rather than to define the cellular or molecular pathways involved. However, our findings are consistent with a growing body of experimental and clinical literature suggesting that malaria infection can induce sustained perturbations in B cell and T cell compartments, including the expansion of atypical memory B cells, altered germinal centre responses, and increased regulatory immune activity. These mechanisms have been proposed to impair the generation and maintenance of effective humoral immunity. In the revised discussion, we have clarified that the mechanistic basis of this phenomenon remains to be fully defined and have expanded the discussion of plausible pathways in light of existing literature. We now explicitly position our findings as providing population-level evidence of a durable immunological phenotype that warrants further mechanistic investigation.

      (8) No understanding of the clinical Implications of the reduced IgG levels observed in the area with high malaria exposure.

      We agree that this study does not directly establish the clinical consequences of the reduced antibody levels observed in malaria-exposed children. The primary objective of this study was to characterise long-term differences in humoral immune profiles associated with early-life malaria exposure, rather than to assess downstream clinical outcomes or functional antibody activity. We have clarified this limitation in the revised discussion. Nevertheless, the breadth and consistency of the observed differences for multiple vaccine-preventable and infectious antigens raise the possibility that early-life malaria exposure may have implications for long-term immune protection. We now emphasise in the revised discussion that future studies incorporating functional assays and clinical outcome data will be required to determine whether these serological differences translate into altered susceptibility to infection or reduced vaccine effectiveness.

      Assessment of Claims:

      The data appear to support the authors' primary claims, but the strength of the evidence is limited, and the results should be interpreted with caution. Together with the currently available evidence of P. falciparum's impact on the host's immune function, this natural experiment design provides further evidence for a relationship between early malaria exposure and reduced antibody responses. The within-Ngerenya analysis controls for geographic factors and thus enhances the quality of the evidence, however, it still fails to account for the physical, nutritional, and socio-economic factors that may have driven the observed changes. Additionally, the mechanism underlying this effect remains unclear, and the clinical significance of reduced antibody levels is not established.

      We thank the reviewer for this assessment and for recognising the strengths of the natural experiment design and within-cohort analyses. We agree that, as an observational study, our findings should be interpreted appropriately. In the revised manuscript, we have undertaken additional analyses and clarifications to strengthen the evidential basis of our conclusions and to address the points raised. To address potential confounding by nutritional and related factors, we analysed contemporaneous hospital-based surveillance data from the same geographic populations since nutritional and socioeconomic variables were not consistently collected during the course of longitudinal follow up. For three independent anthropometric indices of nutrition status (muac, weight-for-age, and height-for-age), we found no evidence of systematic differences between children from Junju and Ngerenya after adjustment for age, calendar year, and concurrent infections. As the longitudinal cohort subjects were randomly drawn from these populations, these findings suggest that the two groups were broadly comparable with respect to early-life growth and nutritional status.

      We agree that the mechanistic basis of the observed differences is not resolved in this observational study. We have clarified this point in the revised discussion and expanded our consideration of plausible biological pathways based on existing literature, including perturbations in B cell and T cell compartments. Similarly, we now explicitly state that the clinical implications of reduced antibody levels remain to be determined and will require studies incorporating functional assays and clinical outcomes. We believe these revisions strengthen the manuscript by providing a more comprehensive interpretation of the data.

      Impact and Utility:

      This work has fundamental implications for understanding vaccine effectiveness in malaria-endemic regions and may contribute to informing vaccination strategies. The findings, if strengthened, would suggest that children in areas of high malaria transmission may require modified immunization approaches. The dataset provides a valuable resource for future studies of malaria's immunological legacy.

      We thank the reviewer for this comment

      Context:

      This study builds on prior work showing acute immunosuppressive effects of malaria but uniquely attempts to demonstrate the durability of these effects years after exposure. The natural experiment design addresses limitations of previous observational studies by providing a more controlled comparison.

      We thank the reviewer for this comment

      Recommendations for the authors:

      Reviewing Editor Comments:

      We suggest that further analyses of potential confounders such as anthropometric indices, socioeconomic status, and comorbidities would render the evidence more robust.

      We thank the Reviewing Editor for this important suggestion. We agree that careful consideration of potential confounding factors is critical to the interpretation of these findings, and have undertaken additional analyses to address this.

      Because anthropometric and related socioeconomic measurements were not collected systematically within the original longitudinal malaria cohort, we assessed potential population-level differences using hospital-based surveillance data from the same geographic regions. This dataset includes measurements of anthropometry (mid-upper arm circumference, weight-for-age, and height-for-age) as well as detailed infection diagnostics in childhood. Using these data, we fitted regression models adjusting for age, calendar year, and concurrent, clinically diagnosed infections. For all three anthropometric indices, we found no evidence of systematic differences between children from Junju and Ngerenya, with effect estimates small and crossing zero (fig. 6). As the longitudinal cohorts were randomly selected from these populations, these findings suggest that the groups were broadly comparable with respect to nutritional status and infection exposure. With respect to socioeconomic status and comorbidities, detailed individual-level data were not available within the longitudinal cohort. However, the within-Ngerenya analysis, where children with differing early-life malaria exposure were compared within the same geographic and environmental setting, provides a complementary control for these factors. We have incorporated these additional analyses and clarifications into the revised manuscript statistical analysis, discussion lines and believe they strengthen the robustness of the findings by addressing key sources of potential confounding.

      Reviewer #1 (Recommendations for the authors):

      The manuscript is well written, with clear and informative figures that effectively support the findings.

      We thank the reviewer for this comment

      Suggestions:

      (1) Although the study well controlled for malaria exposure, other environmental or infectious factors that influence immunity could be considered:

      Nutritional status in childhood (malnutrition impacts immune response), co-infections (helminths, respiratory viruses), socioeconomic differences, or differences in access to health services, even minimal, between Junju and Ngerenya.

      We thank the reviewer for highlighting the potential influence of environmental, infectious, and socioeconomic factors on immune responses. We agree that these are important considerations in the interpretation of observational data. To address nutritional status and concurrent infectious exposures, we analysed contemporaneous hospital-based surveillance data from the same geographic populations. This dataset includes measurements of anthropometric indices (mid-upper arm circumference, weight-for-age, and height-for-age) alongside detailed clinical diagnostics for common childhood infections. Using regression models adjusting for age, calendar year, and concurrent infections, we found no evidence of systematic differences in anthropometric profiles between children from Junju and Ngerenya (fig. 6). These findings suggest that the populations from which the longitudinal cohorts were randomly selected were comparable with regard to early-life growth and nutritional status. We agree that we cannot fully exclude the influence of unmeasured factors such as helminth infections, socioeconomic variation, or subtle differences in healthcare access. However, the within-Ngerenya analysis, where children with differing early-life malaria exposure were compared within the same geographic, environmental, and healthcare setting, provides an internal control for many of these factors. The persistence of similar patterns within this setting supports malaria exposure as a key contributor of the observed differences. We have clarified these considerations in the revised discussion and believe that, the additional analyses and within-cohort comparisons strengthen the robustness of our conclusions while acknowledging the limitations inherent to observational studies.

      (2) Measurement of other immunological markers:

      In addition to IgG, include: B cell subpopulations (naive, memory, atypical), cytokine levels (IL-10, IFN-γ) to characterize the immunological microenvironment.

      You could include these recommendations in the text for future studies.

      We thank the reviewer for this thoughtful suggestion. We agree that detailed immunophenotyping, including characterisation of B cell subpopulations and cytokine profiles, would provide important insight into the mechanisms underlying the observed differences in antibody responses. In the revised manuscript, we have expanded the discussion to highlight these important avenues for future work, including the potential role of altered B cell subsets (and regulatory or inflammatory cytokine environments in shaping long-term humoral responses).

      Reviewer #2 (Recommendations for the authors):

      The manuscript is well-written.

      We thank the reviewer for this comment

      (1) Methodological Clarifications:

      Do the authors have any information regarding the characteristics of these children that could be of use in understanding their immune responses better? (e.g., weight, height, BMI, socioeconomic status, HB level, access to health care, etc.).

      We thank the reviewer for highlighting the importance of participant characteristics in interpreting immune responses. Anthropometric and related clinical measures were not collected systematically within the original longitudinal malaria cohort, as the study was designed to investigate the acquisition of naturally acquired immunity to malaria.

      To address this, we analysed contemporaneous hospital-based surveillance data from the same geographic populations, which include measurements of anthropometric indices (mid-upper arm circumference, weight-for-age, and height-for-age) alongside detailed infection diagnostics. Using regression models adjusting for age, calendar year, and concurrent infections, we found no evidence of systematic differences in anthropometric profiles between children from Junju and Ngerenya (fig. 6) Detailed individual-level data on socioeconomic status, haemoglobin levels, and healthcare access were not available within the longitudinal cohort impeding direct adjustment in the longitudinal cohorts. However, the within-Ngerenya analysis, where children with differing early-life malaria exposure were compared within the same geographic and healthcare setting, provides an internal control for many of these factors. These considerations are now clarified in the revised discussion.

      Could the authors provide more detailed statistical analysis, including power calculations and multiple comparison corrections?

      In the revised manuscript, we have extended the statistical analysis and now include antigen-specific mixed-effects regression models incorporating all available longitudinal measurements, which is comprehensively described in the statistical analysis section. We have also applied false discovery rate (FDR) correction to account for multiple testing across antigens, and report both unadjusted and FDR-adjusted significance in the revised results. With respect to power, the sample size was determined by the number of children meeting inclusion criteria within the long-term surveillance cohorts in terms of availability of a sufficient number of longitudinal samples. We have clarified this in the revised manuscript.

      Clarify the criteria for selecting the 123-child subset from the larger surveillance cohorts.

      We thank the reviewer for this comment. The 123 children included in this analysis were selected from the larger surveillance cohorts based on the availability of sufficiently dense longitudinal serum sampling as described above. Specifically, children were required to have at least eight longitudinal samples available in the archive, enabling robust assessment of within-individual antibody trends over time. This criterion was applied to ensure adequate temporal resolution to examine the long-term stability of malaria-associated effects on antibody responses. Children with fewer available samples were therefore excluded, as limited sampling would not allow reliable characterisation of longitudinal patterns. We have clarified these inclusion criteria in the revised manuscript.

      (2) Additional Analyses and Data Presentation:

      The authors could consider dose-response analyses relating malaria episode frequency/timing to degree of immunosuppression or even AMA-1 IgG levels and degree of immunosuppression. How do they associate over time?

      We thank the reviewer for this suggestion. To address this, we examined the relationship between malaria exposure (using cumulative febrile malaria episode count derived from longitudinal surveillance data) and the magnitude of heterologous antibody responses. In mixed-effects models adjusting for age and repeated antibody measurements, higher malaria episode burden was associated with lower antibody responses against multiple antigens (fig 7).

      Analyze whether the effects vary by specific age at malaria exposure.

      We agree that age at exposure is an important consideration. We have now assessed how the relationship between malaria burden and antibody responses varies with age by including age as a non-linear term and modelling interactions between malaria exposure and age as described above. These analyses did not suggest substantial heterogeneity in the association over age, and therefore we have retained the simpler presentation for clarity.

      Provide correlation analyses between different antibody responses to assess whether suppression is generalized.

      We have addressed this by modelling responses jointly across a panel of heterologous antigens and by examining antigen-specific associations. The direction of effect was consistent for the majority of antigens, with no evidence of opposing trends, supporting a broad rather than antigen-specific effect.

      The authors could consider moving Figures 2a and b to the supplementary material.

      We thank the reviewer for this suggestion. We carefully considered whether panels 2a and 2b could be moved to the supplementary material. However, we have retained them in the main text because they provide a simple, intuitive illustration of how AMA1 antibody responses track with malaria exposure at the individual level, complementing the population-level analysis shown in fig. 2c. We feel that this helps establish the biological validity of the microarray platform in a way that is immediately interpretable to the reader, and therefore supports the interpretation of subsequent analyses.

      The authors could consider replacing Figures 3a and b with IgG levels from ALL vaccinated children and ALL non-vaccinated children.

      We thank the reviewer for this suggestion. We would like to retain these figures for the same reasons that have been articulated above for figures 2a and b.

      (3) Discussion Enhancements:

      The authors should consider expanding the discussion to address the limitations of the data more thoroughly, particularly regarding the potential differences between cohorts that could have contributed to the results.

      We have expanded the discussion to more explicitly address potential differences between cohorts that could contribute to the observed findings, including nutritional, socioeconomic, and environmental factors.

      The discussion needs to acknowledge the lack of directionality for the associations observed. As stated above, although I agree in general terms with the observations that the authors have made, it is not possible to distinguish between a suppressive effect of malaria on immune responses to infection-derived pathogens or a protective effect of malaria that leads to less exposure to infection-derived pathogens (and consequently lower IgG levels). The mechanisms behind these could include things like different health-seeking behaviors or social interactions from kids who have malaria versus those who don't, for example.

      We agree that, as an observational study, we cannot definitively establish the direction of the association between malaria exposure and antibody responses to unrelated antigens. We have now clarified this limitation explicitly in the discussion. We acknowledge the alternative interpretations raised by the reviewer, including the possibility that differences in exposure to other pathogens, potentially driven by behavioural, environmental or healthcare-related factors, could contribute to the observed patterns. At the same time, we note that the natural experiment design, prospective malaria exposure classification, and within-Ngerenya comparisons support early-life malaria exposure as a key contributing factor. We have revised the discussion to reflect this balance.

      Extend the discussion of potential biological mechanisms underlying durable immunosuppression.

      We thank the reviewer for this suggestion. We have expanded the discussion to more fully consider potential biological mechanisms that could underlie the observed long-term differences in antibody responses. Specifically, we now discuss evidence from prior studies indicating that malaria infection can induce sustained alterations in B cell and T cell compartments, including expansion of atypical memory B cells, disruption of germinal centre responses, and increased regulatory immune activity. We position our findings as providing population-level evidence of a durable immunological phenotype, while noting that targeted mechanistic studies will be required to define the underlying pathways.

      Extend the discussion around the clinical implications of the observed antibody level differences.

      In the revised discussion, we highlight that studies incorporating functional assays and clinical outcome data will be required to determine whether these serological differences translate into altered susceptibility to infection or reduced vaccine effectiveness.

      (4) Technical Issues:

      Could the authors please:

      (1) Clarify microarray data processing and quality control procedures.

      We thank the reviewer for this request. We have expanded the methods section to provide additional detail on microarray data processing and quality control procedures.

      (2) Provide information on inter-assay variability and batch effects.

      We have expanded the methods section to clarify how these were evaluated and addressed. Inter-assay variability was monitored using pooled adult serum included on every slide as a consistent positive control. This allowed us to assess slide-to-slide consistency in signal detection across the full antigen panel. In addition, fluorophore-conjugated IgG and IgA controls were printed directly onto each miniarray to confirm scanner performance independently of antigen–antibody interactions. At the sample level, each specimen was assayed on two independent miniarrays per slide, generating four spatially separated replicate measurements per antigen. Technical variability was quantified using the coefficient of variation (CV), and measurements with CV >20% were excluded from downstream analyses.

      (3) Include details on how missing data were handled in longitudinal analyses.

      We thank the reviewer for highlighting this point. We have added clarification in the statistical analysis section describing how missing data were handled. Specifically, mixed-effects models were used, which accommodate unbalanced longitudinal data without requiring imputation, allowing all available observations to contribute to the analysis.

      (4) Include details of the parameters of the LOWESS analysis shown in Figure 1.

      We have expanded the figure 1 legend to include the parameters used for the loess smoothing shown, including the smoothing span.

      (5) Include details of the samples used for Figure 3d (Negative and Pooled Adult Serum).

      We have clarified in the methods the nature and purpose of the samples used in Figure 3d. The negative control consisted of phosphate-buffered saline applied to a full miniarray in place of serum, allowing assessment of background and non-specific signal in the absence of antibody binding. The pooled adult serum comprised a composite of sera from multiple healthy adults from the same setting and was included as a positive reference sample, expected to contain a broad repertoire of antigen-specific antibodies. These controls were included on each slide to enable interpretation of assay performance, with the negative control defining baseline signal and the pooled adult serum providing a consistent reference for antigen recognition across the microarray.

    1. eLife Assessment

      This valuable study identifies a brown adipose tissue-specific heat shock factor 1 (HSF1)-alcohol dehydrogenase 5 (ADH5) axis that regulates oxidative stress and cellular senescence during aging. The authors show that ADH5 deficiency drives BAT dysfunction and contributes to organismal health decline in aged mice. The evidence is convincing, and the work will be of broad interest to the adipose tissue and aging research communities.

    2. Reviewer #1 (Public review):

      Sebag et al. addressed the role of ADH5 in BAT in the development of aging and metabolic disarrangements associated with it. This is a follow-up study after the authors' demonstration of the role of BAT ADH5 in glucose homeostasis, obesity, and cold tolerance. By ablating ADH5 specifically in brown adipocytes or pharmacologically modulating ADH5 through activation of its transcription factor, the authors conclude that preservation of BAT function is crucial for healthy aging and ADH5 is causally involved in this process. The topic is appealing given the rise in the aging population and the unclear role of BAT function in this process. Overall, the study uses several techniques and addresses several physiological and molecular manifestations of aging. Therefore, the findings contribute to the growing body of literature pointing to the biological role of BAT activity in aging.

      Comments on revised version:

      I have no further comments other than to congratulate the authors on the nice piece of work.

    3. Reviewer #2 (Public review):

      Summary:

      This study investigates the role of the enzyme Alcohol Dehydrogenase 5 (ADH5) in brown adipose tissue (BAT) during aging. BAT is crucial for thermogenesis and energy balance, but its function and mass diminish with age, contributing to metabolic dysfunction and age-related diseases. ADH5, also known as S-nitrosoglutathione reductase, regulates nitric oxide (NO) signaling by removing damaging S-nitrosylation modifications from proteins. The authors show that aging in mice leads to increased protein S-nitrosylation associated with a combination of increased Nos2 expression and reduced ADH5 expression in BAT, resulting in impaired metabolic and cognitive functions. Deletion of ADH5 in BAT accelerates tissue senescence and systemic metabolic decline. Mechanistically, aging suppresses ADH5 via downregulation of heat shock factor 1 (HSF1), a master regulator of protein homeostasis. Importantly, pharmacologically boosting HSF1 improves BAT function and mitigates both metabolic and cognitive declines in aged mice. The findings highlight a critical HSF1-ADH5 pathway in BAT that protects against aging-related dysfunction, suggesting that targeting this pathway may offer new therapeutic strategies for improving metabolic health and cognition during aging.

      Strengths:

      This research provides insight into the interplay between redox biology, proteostasis, and metabolic decline in aging. By showing that age regulates genes that control SNO status in BAT and further developing a therapy to target ADH5 in BAT to prevent age related decline, the authors have identified a putative mechanism to combat age related decline in BAT function.

      Weaknesses:

      None identified.

      Comments on revised version:

      Congratulations to the authors for this interesting manuscript. I don't want to pat myself on the back, but I found the increased Nos2 expression in Figure 1C of the revised manuscript very satisfying, as it reinforces the shift in the regulation of SNO status that happens in BAT with aging. I appreciate the authors addressing this suggestion.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers for their thoughtful and constructive comments. We fully agree that when two independent variables (genotype and age) are being evaluated, the statistical analysis must appropriately account for both factors and their potential interaction. We appreciate the reviewers’ guidance in strengthening the statistical rigor of our study.

      In response to this concern, we have carefully reanalyzed the relevant datasets using two-way ANOVA to properly assess the effects of genotype, age, and their interaction. The manuscript, figures, and figure legends have been revised accordingly. Specifically:

      Figure 1:

      The quantification of p16 expression in Fig. 1F has been reanalyzed using two-way ANOVA. The figure has been replotted, and the corresponding legend has been updated to reflect the revised statistical approach.

      Figure 2:

      The quantification of AUC in Fig. 2F has been reanalyzed using two-way ANOVA. The figure and legend have been updated accordingly.

      Figure 3:

      The quantification of F4/80 in Fig. 3C and 3D has been reanalyzed using two-way ANOVA. The figures and corresponding legends have been revised to reflect this updated analysis.

      Public Reviews:

      Reviewer #1 (Public review):

      Sebag et al. addressed the role of ADH5 in BAT in the development of aging and metabolic disarrangements associated with it. This is a follow-up study after the authors' demonstration of the role of BAT ADH5 in glucose homeostasis, obesity, and cold tolerance. By ablating ADH5 specifically in brown adipocytes or pharmacologically modulating ADH5 through activation of its transcription factor, the authors conclude that preservation of BAT function is crucial for healthy aging and ADH5 is causally involved in this process. The topic is appealing given the rise in the aging population and the unclear role of BAT function in this process. Overall, the study uses several techniques, is easy to follow, and addresses several physiological and molecular manifestations of aging. However, the study lacks an appropriate statistical analysis, which severely affects the conclusions of the work. Therefore, interpretation of the findings is limited and must be done with caution.

      We sincerely thank the reviewer for their thoughtful and constructive comments. We fully agree that when two independent variables (genotype and age) are being evaluated, the statistical analysis must appropriately account for both factors and their potential interaction. We appreciate the reviewers’ guidance in strengthening the statistical rigor of our study.

      Reviewer #2 (Public review):

      Weaknesses:

      (1) Sex needs to be considered as a biological variable, at a minimum in the reporting of the phenotypes observed in this manuscript, but also potentially by further experimentation. The only mention of sex I could find is that the authors reported the general protein SNO status in BAT is increased with age in male C57Bl/6J mice. Is this also true in female mice?

      We thank the reviewer for this insightful comment. In response, we examined whether aging affects Hsf1 and Adh5 transcript levels in wild-type female mice (3 months vs. 19 months). Our analysis did not reveal significant age-associated changes in the expression of either gene. These results have now been incorporated into the revised manuscript and are presented in Figure 4A.

      (2) It would be helpful to know the extent of ADH5 loss in the adipose tissue of knockout mice, either by mRNA or by immunoblotting for ADH5. It could also be helpful to know if ADH5 is deleted from the inguinal adipose tissue of these mice, especially since they seem to accumulate fat mass as they age (Figure 2B).

      We thank the reviewer for this suggestion. Indeed, we have previously measured ADH5 expression in both brown adipose tissue (BAT) and inguinal white adipose tissue (iWAT). These data were published in Cell Reports (PMID: 3478865).

      (3) For Figure 4D, it's not clear how these BAT samples were treated with HSF1A - was this done in vivo or ex vivo?

      We thank the reviewer for their thoughtful comment. We have now provided additional methodological details in the revised manuscript. In Figure 4D (current Figure 4E), BAT was collected from wild-type mice and cultured ex vivo as explants. The BAT explants were treated for 24 hours with HSF1A (an HSF1 activator; 20 µM). Following treatment, mRNA levels of the indicated genes were measured by RT-qPCR.

      (4) I didn't understand what was on the y-axis in Figure 5A, nor how it was measured.

      We apologize for not making these critical points clearer in the initial submission. Figure 5A shows the release profiles of HSF1A from collagen gels with nanoclay (Collagen–NC–HSF1A) and without nanoclay (Collagen–HSF1A), determined using an established standard curve method (Hu et al., PMID: 33225042).

      The concentration of HSF1A was quantified by UV–Vis spectroscopy. Briefly, a standard curve for HSF1A was generated by measuring the UV–Vis spectra of HSF1A at known concentrations (1.25, 2.5, 5, 10, and 20 µM) prepared in phosphate-buffered saline (PBS). Collagen gels with or without nanoclay were then fabricated to evaluate the release profile. At predetermined time points (1, 5, 9, 14, and 21 days), the PBS supernatant from each sample was collected and analyzed by UV–Vis spectroscopy. The amount of released HSF1A was calculated using the previously established standard curves. A brief description has now been included in the figure legend.

      (6) Figure 1B: What is the age of the positive (ADH5BKO) and negative (Adh5 fl) mice?

      We regret that we did not describe our results clearly in the first submission and have included detailed information in the revised manuscript.

      (7) Figure 1F: Can you clarify what I'm looking at in the P16ink4a panels? The red staining? Is the blue staining DAPI? This is also a problem in Figures 3C, 3D and 5G, and 5I. Figure 4B looks great - maybe this could be used as an example?

      We regret that we did not present results clearly in the first submission and have provided detailed information in these figures in the revised manuscript.

      (8) Figure 3B looks a bit odd. Can the approach to measuring IL-1β be clarified, and could the authors explain why they can't show units of mass for IL-1β levels?

      We have provided information in the revised manuscript.

      (9) What are the levels of nitric oxide synthase in the BAT of the aging model? Since protein S-nitrosylation is regulated by a balance of both, the attribution of greater protein S-nitrosylation to ADH5 is incomplete without determining nitric oxide synthase.

      We thank the reviewer for this thoughtful comment. In response, we have now included the analysis of iNOS transcript expression levels in the revised manuscript. These data are presented in Figure 1C.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (2) Presentation of metabolomics is not appropriate. The authors described, using color coding, the metabolites up- or downregulated in the experimental design. However, the current approach does not allow the reader to detect sample size, magnitude of changes, variability of the data, p-values, etc. This approach does not follow the standard practices of scientific rigor and should be modified. Metabolomic data could be uploaded as supplementary data, or a table with all necessary information to allow a full interpretation of the data should be provided.

      We have now provided the the metabolimic data in a table format as Figure 3I.

      (6) What are the levels of nitric oxide synthase in the BAT of the aging model? Since protein S-nitrosylation is regulated by a balance of both, the attribution of greater protein S-nitrosylation to ADH5 is incomplete without determining nitric oxide synthase.

      We thank the reviewer for their thoughtful comment. We have now included iNOS transcript levels expression level in the revised manuscript (Figure 1C).

      Minor Comments:

      (1) The conclusion of the abstract is somewhat vague. I suggest the authors rewrite it to better recapitulate what was found in the study.

      We thank the reviewers for this helpful suggestion. In response, we have revised the Abstract to improve the specificity and clarity of our conclusions.

      (2) In the introduction, the authors mention that an increased level of mitochondrial ROS activates UCP1. Given that the evidence for this statement is circumstantial and not supported by the current state-of-the-art (PMID: 28710335), where it is accepted that UCP1 activation diminishes ROS production, I suggest that the authors tone down this statement or at least acknowledge conflicting findings and interpretations.

      We thank the reviewer’s insight, we have included this important notion in the introduction.

      (3) Figure 2H - It is unclear what this figure (and statistical analysis) represents. Please, improve the description of the experiment and how the data were plotted to reach such a conclusion.

      We regret that we did not present results clearly in the first submission. The trend lines show the relationship between body weight and time on rotarod. The P value is the comparison of the slope of the line between Adh5 BKO mice and Adh5 fl/fl mice. The data implicate that the heavier the BKO mouse, the less time spent on the rotarod.

      (4) Figure 2M - The unit of LV thickness is missing. Please, provide it. In addition, I am missing the other cardiac parameters obtained from the echocardiogram.

      We have included this information in Figure 2M in the revised manuscript.

      (5) Figure 2G - I believe force is not the right unit for the grip strength test. Please, revise accordingly.

      We regret that we did not describe our results clearly in the first submission. We have corrected this unit in the revised figure.

      (6) Figure 3H - What is the unit when reporting mitochondrial area?

      We regret that we did not describe our results clearly in the first submission. We have added this information in the revised figure.

      (7) Is HFS1 also downregulated in iWAT?

      We thank the reviewer for this thoughtful comment. In response, we measured Hsf1 expression in iWAT from young and aged wild-type male mice. Our analysis did not reveal any significant age-associated changes in Hsf1 expression in iWAT. These results have now been included in the revised manuscript (Figure 4C).

      (8) Can the authors explain how HFS1 expression increases upon HSF1 activation? I understand ADH5 is controlled by HSF1, but what would control HSF1 itself? Off targets?

      We thank the reviewer for this insightful comment. At present, we do not have direct mechanistic evidence to definitively support this notion, and we cannot exclude the possibility of off-target effects of HSF1A. However, previous studies have reported that the HSF1 promoter contains heat shock elements (HSEs) in humans and HSE-like domains in mice. Based on this, we speculate that activated HSF1 may enhance its own transcription through an autoregulatory or positive feedback mechanism.

    1. eLife Assessment

      The study by Reed et al. provides fundamental findings defining the topological changes that occur during tumorigenesis. These compelling findings enhance the understanding of stable long-range connections among genes that reprogram cancer-related functions.

    2. Reviewer #1 (Public review):

      Summary:

      In their manuscript, Metz Reed and colleagues present an exceptionally thorough analysis of three-dimensional genome reorganization during breast cancer progression using the well-characterized MCF10 model system. The integration of high-resolution Micro-C contact maps with multi-omics profiling provides compelling insights into stage-specific dynamics of chromatin compartments, TAD boundaries, and looping events. The discovery that stable chromatin loops enable epigenetic reprogramming of cancer genes while structural changes selectively drive metastasis-associated pathways represents a significant conceptual advance. This work substantially deepens our understanding of genome topology in malignancy.

      Strengths:

      This work sets a benchmark for integrative 3D genomics in oncology. Its methodological sophistication and conceptual advances establish a new paradigm for studying nuclear architecture in disease.

      Comments on revised version:

      The authors made a significant effort to improve the manuscript. My comments were sufficiently addressed.

    3. Reviewer #2 (Public review):

      Using the MCF10 breast cancer progression sequence, the authors combined high-resolution Micro-C chromatin conformation capture with RNA-seq and ChIP-seq to depict the sequential reorganization of compartments, topologically associated domains (TADs), and long-range loops in benign, pre-tumor, and metastatic states, and coupled these three-dimensional changes with gene expression and enhancer activity. Four main findings were: (i) chromatin structure was largely quiescent, still limiting gene output differentiation, with upregulated sites being most significantly affected; (ii) enhancer-promoter contact strength covariated with transcriptional amplitude; (iii) 127 genes gained expression with increasing chromatin contact; and (iv) progression-related genes acquired altered histone markers in distal enhancers, which remained connected by stable loops. These conclusions are widely accepted and provide strong justification for the publication of this paper.

    4. Reviewer #3 (Public review):

      Summary:

      The authors tackle an important problem- that is defining the topological changes that occur during tumorigenesis. To study this, they use an established stepwise cell model of breast cancer. A strength of their study is a careful, robust differential analysis of topological features across each cell state that is presented clearly and rigorously. They define changes in compartmentalization, TAD structure and chromatin looping. Intriguingly, when the authors integrate differential gene expression with chromatin looping, they see that most differentially regulated genes are not involved in loop changes, suggesting that changes in promoter or enhancer chromatin marks may play a bigger role in regulating transcription than differential loops. The differential topology analysis and its integration with transcription is very well done- one of the best versions of this I have read in the 3D genome field! However, the paper is framed largely as a cancer biology study and it teaches us much less about this. I am worried that some of the trends for each topologic feature are not going to be consistent across the pre-malignant-malignant-metastatic spectrum and would like the authors to soften some of their claims a bit regarding how this clarifies our understanding of cancer evolution.

      Updated comments on revision:

      There are still some issues with this paper. First, it reads descriptively. It is a series of comparisons with limited biologic insight as changes are always seen in genomics and in this case, they're often not tied back to transcription or gene regulation in cancer. Cell lines do not represent cancer faithfully and in this case should not be argued to represent malignant transformation broadly. The authors did not really soften their language as much as I think required. I would caution the authors to further qualify their results in the context of a single, clonal cell line that has undergone stepwise transformation. This is not a patient cohort analysis or frank progression. This matters because there is likely to be much more noise, not pertinent to transformation, in a cell line model. It doesn't negate the validity of the study, but this language should be adjusted appropriately. It was nice to see the authors compare gene expression data from their model to the primary tumor data, however the limited overlap is concerning that at the least patterns of transcriptional regulation in their model are not faithful to primary tumors. If this is the case, it raises concern that the topological changes are also not generalizable to cancer.

      The authors declined a number of functional assays to validate their observations (which are purely correlative). And while I see that the burden of extra experiments may be beyond the scope of this study, they must soften their language to justify the observed relationships.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      This work sets a benchmark for integrative 3D genomics in oncology. Its methodological sophistication and conceptual advances establish a new paradigm for studying nuclear architecture in disease.

      We appreciate the very kind words.

      Weaknesses:

      Major Issues

      (1) Functional tests would strengthen the observed links between structure and gene changes. For example, the COL12A1 gene loop formation correlates with its increased expression. Disrupting this loop using CRISPR-dCas9 at chr6 position 75280 kb could prove whether the loop causes COL12A1 activation. Such experiments would turn strong correlations into clear mechanisms.

      We agree that targeted disruption of specific loops such as COL12A1 will be important for functional validation of the causal relationships between enhancer-promoter loop formation/dissipation and changes in gene expression. However, the intent of our current study was to profile changes in genome organization at a global scale to deduce general features of cancer progression-associated changes in genome organization, rather than to explore specific loop interactions. The current findings are a foundation for more targeted functional follow-up studies.

      (2) The H3K27ac looping idea needs deeper validation. Data suggests H3K27ac loss weakens loops without affecting CTCF. Testing how cohesin proteins interact with H3K27acmodified sites would clarify this process. Degron systems could rapidly remove H3K27ac to observe real-time effects. Also, the AP-1 motifs found at dynamic loop sites deserve functional tests. Knocking down AP-1 factors might show if they control loop formation.

      We agree that modulating histone modifications or transcription factors would provide insights into the underlying mechanisms driving the changes we observed. However, such studies utilizing degrons or small molecule inhibitors that globally knock down either H3K27ac or specific transcription factors are often difficult to interpret. For example, assessing the role of AP-1 factors, as suggested, would be complicated by the variety of AP-1 proteins. In addition, H3K27ac reduction could inhibit loop strength either directly (i.e. by reducing cohesin recruitment) or indirectly (i.e. by reducing gene expression which could in turn affect loop strength). Parsing out the exact relationships between these features will require extensive follow-up work and falls outside of the scope of the current study.

      (3) Connecting findings to patient data would boost clinical relevance. The MCF10 model is excellent for controlled studies. Checking if TAD boundary weakening occurs in actual patient metastases would show real-world importance. Comparing primary and metastatic tumor samples from the same patients could reveal new structural biomarkers. If tissue is scarce, testing cancer cells with added stroma cells might mimic tumor environment effects.

      We have leveraged publicly available datasets to link the observations from the progression model to clinical samples. Specifically, we have compared our datasets to chromatin organization data in non-cancerous mammary epithelial cells (HMEC), five cell lines representing distinct cancer subtypes ranging from less (luminal) to more aggressive (triple negative, TNBC), as well as tissue samples from TNBC patients with contralateral normal controls. We explored the conservation of both loops and TADs identified in the MCF10 progression system in each of these maps, paying particular attention to how features that are differential between MCF10 cells differ across other cancer cell types. We observe a high degree of conservation of static loops and TAD boundaries among the cancer samples, as well as some degree of cell-specific changes among loops and boundaries that change during MCF10 progression. These findings are included in Supplemental Figures 3 and 4 and are discussed on page 7.

      Minor Issues

      (1) Adding a clear definition for static loops would help readers. For example, state that static loops show less than 10 percent contact change across replicates.

      Static loops are defined as loops with a fold-change of 1.5 or more between any two MCF10 cell lines and an adjusted p-value of less than 0.025 considering change across biological and technical replicates. This definition is stated on page 6).

      (2) In the ABC model analysis, removing promoter regions from the enhancer list would focus results on true long-range interactions.

      The ABC model already excludes the promoter of each gene. Only self-promoters are excluded, whereas the model allows promoters of other genes to act as potential long-range enhancers of the target gene. We have added text to make this clear (see page 11).

      (3) Briefly noting why this study sees TAD weakening while other cancer types show different patterns would provide useful context.

      The biological reason for TAD weakening in the MCF10 model is not known, but neither the mechanism for boundary weakening nor the reason for apparently different behavior amongst cancers is known. We expanded the text on this discussion slightly, but we refrain from making any definitive claims. We do note that differences in the types of cancer studied or the methods used for detecting changes in TADs (i.e. different sensitivities and thresholds for detecting change) could be responsible (see page 15). We also mention that the loss of insulation at many TAD boundaries detected in our study are subtle changes in intensity that could be potentially missed if using methods tailored to find more drastic changes in TAD architecture.

      Reviewer #2 (Public review):

      While the conclusions are broadly supported, methodological and analytical refinements are required.

      We appreciate these comments.

      (1) Model representativeness. The long-term culture-adapted MCF10 genome harbours extensive aneuploidies and translocations. Validation of key COL12A1/WNT5A loop dynamics in an independent breast-cancer line (e.g., MDA-MB-231, T47D) or in patientderived organoids/PDX models would strengthen generalizability.

      Although the generation of Micro-C datasets in additional cell lines is outside of the scope of this study, we used publicly available Hi-C data from triple negative breast cancer (TNBC) progression and patient samples (Kim, Han & Chun et al. 2022) to assess generalizability of the MCF10 model findings. While these maps are lower resolution than the Micro-C maps used in our study, they are of sufficient depth to detect loops at a similar resolution (10 kb). We report these findings in Supplemental Figures 3 and 4 and discuss them on page 7.

      We find that chromatin loops and TAD boundaries detected across the MCF10 system are highly conserved across all other mammary epithelial lines studied. Chromatin loops that were more prominent in MCF10AT1 and MCF10CA1a lines were also significantly stronger in TNBC cells. Insulation score boundaries that were weakened in MCF10CA1a showed strong insulation across all cell lines in TNBC. These findings highlight that different model systems indeed have distinct profiles of structural change, just as they have distinct gene expression profiles.

      It is worth noting that direct comparison at individual loci is complicated by variations in gene expression profiles between the MCF10 model and the TNBC progression model; for example, COL12A1 is not significantly upregulated between normal and TNBC tissues in this study (unlike in the TCGA-BRCA data) and is downregulated between HMEC and TNBC cell lines. Regardless, our analysis provides some indication of conserved and divergent features in the various model systems.

      (2) The study remains purely correlative; no perturbation experiments are conducted to demonstrate causal roles of chromatin loops on gene expression. CRISPR interference (CRISPR-Cas9-KRAB/HDAC) or enhancer deletion/inversion should be applied to 3-5 pivotal loops (e.g., COL12A1, WNT5A) to test their impact on target-gene expression and cellular phenotypes (e.g., proliferation, migration).

      We agree that targeted disruption of specific loops such as COL12A1 will be important for understanding the causal relationships between enhancer-promoter loop formation/dissipation and changes in gene expression. However, the intent of our current study was to profile changes in genome organization at a global scale to deduce general features of cancer progression-associated changes in genome organization, rather than exploring specific loop interactions. The current findings are a foundation for more targeted follow-up functional studies.

      (3) The manuscript lacks integration with clinical datasets. Integrate TCGA-BRCA data to assess whether elevated COL12A1/WNT5A expression associates with overall survival (OS) or distant metastasis-free survival (DMFS)

      To assess clinical significance of specific loci, we have queried expression of all differentially expressed genes in the MCF10 progression system among TCGA-BRCA expression data. We summarize our findings in Supp. Fig. 5E and discuss them on page 8.

      We found that roughly 25% of genes that change in our model also change significantly in breast cancer, but only roughly half of those genes change in the same direction (i.e. up-regulated in MCF10CA1a vs MCF10A, and up-regulated in tumor vs normal samples). Interestingly, there was a higher degree of directional agreement between latechanging genes (i.e. genes that change in MCF10CA1a compared to MCF10A and MCF10AT1) than early-changing genes (i.e. genes that change in MCF10AT1 and MCF10CA1a compared to MCF10A).

      We have also explored the impact of select highlighted genes on overall survival (OS). We present these data in Supp. Fig. 6 and discuss it on page 8. While not all genes showcased in this study have a significant impact on overall survival, most trend in the same direction as their differential expression would suggest (i.e. genes more highly expressed in cancer vs tumor also have a hazard ratio above 1).

      Reviewer #3 (Public review):

      The differential topology analysis and its integration with transcription is very well done- one of the best versions of this I have read in the 3D genome field!

      We appreciate the reviewers’ endorsement.

      However, the paper is framed largely as a cancer biology study, and it teaches us much less about this. I am worried that some of the trends for each topologic feature are not going to be consistent across the pre-malignant-malignant-metastatic spectrum and would like the authors to soften some of their claims a bit regarding how this clarifies our understanding of cancer evolution.

      We agree that the strength of the study lies in its deep mapping of chromatin architecture and the landscape of enhancers and differentially expressed genes, which we hope to use to better understand the relationship between chromatin structure and gene expression, regardless of their cancer relevance. To better relate the findings in the progression system to cancer, we have added new data from direct comparisons of the MCF10 progression system with multiple patient-derived cancer cell lines and cancer tissues. These data are shown in Supp. Fig. 3 and 4 and discussed on p. 7. Regardless, we have softened the claims regarding cancer progression throughout the manuscript.

      Weaknesses:

      Major Concerns:

      (1) The integration of gene expression and chromatin loops is intriguing. The authors' differential analysis, however, omits consideration of genes that are on and simply further upregulated versus genes that transition on/off or off/on. It would be nice to see the authors break out looping patterns for these two different patterns of regulation, as it may be instructive regarding the rules for how EP loops govern transcription.

      To address different types of gene expression patterns, we analyzed 108 genes that went from an unexpressed or “off” state (2 or fewer read counts) in one cell line to an expressed “on” state (100 or more read counts) in another, and 111 genes that go from “on” to “high” (1000 or more read counts). We present these data in Supp. Fig. 8 and discuss the findings on page 9. While neither of these genes were enriched for differential loops, a large number overlap with loop anchors. We found a relationship between loop strength and gene expression levels; genes that are more strongly expressed are more likely to overlap with the anchor of a chromatin loop. All gene sets show similar strong trends at distal regulatory regions.

      (2) Given the paucity of differential loops at the majority of genes whose expression changes, the authors should examine chromatin subcompartments, as these may associate more with differential transcription.

      We present subcompartment analysis in Supp. Fig. 9. Our CALDER compartment calls are qualitative rather than quantitative, so to explore this we examined how compartments change genome-wide and at specific promoters. We show these data in Supp. Fig. 9 and discuss the findings on page 10-11. We see that between any two cell types, a majority of changes occur between closely related subcompartments, i.e. from A.2.2 to A.2.1 (1 step more A-like) or B.1.1 (1 step more B-like). The promoters of differentially expressed genes have minimal subcompartment changes, but genes that shift from on to off have larger changes. Differentially expressed genes with promoters that shift by multiple subcompartments have significant impacts on fold-change, but smaller shifts have minimal impacts on gene expression. In summary, small changes in subcompartments are very common and have little impact on gene expression, while larger changes are infrequent and correlate more strongly with changes in gene expression.

      (3) The authors could push their TAD analysis further by integrating it with transcription. Can they look at genes and their enhancers that span these altered boundaries to see if these shifts impact transcription?

      We provide this analysis in Supp. Fig. 9. We started, as suggested, by looking at genes with distal enhancers (as determined by the ABC model) that span a single TAD boundary. However, the number of genes that fit this definition was relatively small, so we expanded to look at any genes with promoters in the proximity (50kb) of differential insulation score boundaries, for which we saw the same trends with more robust signal. Our findings are shown in Supp. Fig. 9 and discussed on page 10. We found that genes near weakened boundaries are not enriched for differentially expressed genes, while those near strengthened boundaries are. Comparing the fold-change of genes near strengthened, weakened, and static boundaries showed a significant inverse correlation between boundary strength and gene expression, although effect sizes were small. These results show that changes in TAD boundary insulation have small but noticeable impacts on gene expression.

      (4) The progression of cancer critically goes from a benign -> pre-malignant -> malignant -> metastatic series of steps. The AT1 line is described as 'premalignant' and thus the authors' series omits a malignant line. While I think adding such a sample is an unreasonable request at this point (as it would have had to have been studied in 'batch' with these other samples), the authors should acknowledge that they omit this step and spend some time discussing the genetic, morphologic, and phenotypic features for their 3 conditions. The images in Figure 1S aren't particularly useful- they don't tell the reader that these cells are malignant/benign. The karyotypic data are intriguing but not fully analyzed, so it is hard to know what true phenotype these cells represent. For example, malignant means DCIS/invasive carcinoma - so then what does this pre-malignant cell model represent? The described alteration in the AT1 line is a Ras oncogene, so in some sense, the transition to this line really is just +/- Ras. The authors could spend some time thinking about the effects of Ras specifically on the 3D genome.

      We have expanded our discussion of the relevance of the MCF10 model on page 4, and the limitations of the model on page 17. The MCF10 progression model has been extensively used by many laboratories, and its properties have been discussed in detail (i.e. Polizzotti et al. 2012). Critically, the MCF10AT1 cell line is the product not only of Ras oncogene expression but then derived from a 100-day-old precancerous lesion that formed a squamous carcinoma in a mouse, and over this time it accumulated additional changes. The MCF10AT1 line is considered pre-malignant as it has accrued critical changes that prepare it for the metastatic transition, but it does not immediately form tumors when injected back into mice. Unlike the MCF10DCIS cell line which is malignant but not metastatic, the more aggressive MCF10CA1a is classified as both malignant and highly metastatic, forming tumors that quickly metastasize to the lungs in mouse xenografts. While both MCF10AT1 and MCF10CA1a are tumorigenic, we acknowledge the lack of a nonmetastatic malignant cell line in the discussion on page 17. We have also provided updated karyotype characterization of the cell lines used in this study in Supp. Fig. 1B and now include full composite karyotypes in the Methods section (page 18).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The reviewer’s recommendations are the same as their public review comments. See our response to the review comments above.

      Reviewer #2 (Recommendations for the authors):

      (1) If conditions permit, it is recommended that inclusion of primary human mammary epithelial cells (HMECs) to distinguish immortalisation-specific from malignancy-specific 3D changes.

      Micro-C data of equal resolution is not available for HMECs. We have, however, incorporated analysis of publicly available deeply sequenced Hi-C data of HMECs into several figures that explore the conservation of loops and TADs in these cells (Supp. Fig. 3 and 4).

      We find that chromatin loops and TAD boundaries detected across the MCF10 system are highly conserved across all other mammary epithelial lines studied. Chromatin loops that were more prominent in MCF10AT1 and MCF10CA1a lines were also significantly stronger in TNBC cells. Insulation score boundaries that were weakened in MCF10CA1a showed strong insulation across all cell lines in the TNBC system. These findings highlight that different model systems indeed have distinct profiles of structural change, just as they have distinct gene expression profiles.

      (2) The relationship between loop alterations and copy-number variations (CNVs) is not explored. If conditions permit, it is recommended that overlay differential loops with SNP/Indel/CNV data to exclude spurious differences arising from structural alterations.

      While we have not conducted an in-depth SNP analysis, we have clarified our discussion of the karyotype analysis on pages 21 and 23 and how we mitigated these effects when identifying differential loops between cell lines.

      (3) The horizontal and vertical coordinates of the diagram are difficult to view; it is recommended that the size of the text on the picture be adjusted to ensure that it is clear to read. Some of the text coordinates of the figure are labeled in gray; it is recommended that they be in black.

      The clarity of the figures has been improved.

      Reviewer #3 (Recommendations for the authors):

      I really like this paper. I think if the cancer focus can be down-emphasized (because I'm not fully clear what we've really learned about cancer), then it represents a nice dataset and a thoughtful, comprehensive analysis.

      We greatly appreciate the kind words and helpful feedback. The cancer focus has been toned down throughout the manuscript, as suggested.

      Minor Concerns:

      (1) The authors present a nice summary of the topological changes across samples. However, summary statistics can mask noise/bias and also don't fully convey the effect size of the reported changes. Highlighting individual loci and visualizing these would strengthen the paper and participate in maintaining a high standard for our genomic studies of topology, in which we summarize, but also provide representative examples. I would appreciate seeing more example plots at distinct loci (even if in the supplemental information).

      We have included several more example regions in Supp. Fig. 7 and 12, including four looped genes that change similarly between the MCF10 series and TCGA-BRCA data (2 stably looped genes and 2 differentially looped genes, 2 up-regulated and 2 downregulated), and six differentially looped and differentially expressed genes (3 which change in the same direction as the loops, and 3 which change in the opposite direction).

      (2) "To identify loops that changed significantly during cancer progression, we assessed changes in contact frequency among every loop in each cell type, correcting for karyotypic differences that result in differences in coverage between cell lines (see Methods)." The Methods section is not adequately explained. Also, could you go a bit deeper to define if these large-scale changes shift the 3D genome specifically? This is hard, but there may be some low-hanging fruit given the otherwise fairly isogenic features in your model.

      We have added more detail to the Methods section on pages 21 and 23 on how karyotypic abnormalities were included in our analysis and differential loop calling. A deeper analysis of how large-scale karyotypic changes affect chromatin organization (i.e. through the formation of neoloops and TADs through translocations) is indeed an attractive subject, but due to its complexity requires a separate dedicated study.

      (3) "Approximately half of chromatin loops featured some combination of active gene promoters and enhancers within 10kb of loop anchors". The authors have high-resolution topology data and should be more stringent; these features should have to overlap loop anchors or at least use a distance less than 10kb, which, in some sense, forfeits the advantages of high-resolution topology data.

      The threshold of 10kb was chosen for several specific reasons: First, the loop sizes detected here are large enough that this relatively large region still represents a small fraction of the loop span, and these regions are reasonably considered anchor-proximal. Second, the loops we detect are non-punctate, both in aggregate analysis (Figure 1H, bottom) and at individual loci (see example regions), showing increased contact frequency among several 5kb or 10kb bins. Therefore, adding 10kb to either side (2 pixels on 5kb maps and 1 pixel on 10kb maps) ensures that the full region of increased contact frequency is included. Finally, ultra-resolution Hi-C data has also shown that loops remain diffuse even with 1kb resolution maps (albeit they do get smaller than the 30kb used here) (Harris & Gu 2023). We have added a brief justification of this overlap size to the text on page 24.

      (4) "These results show that not only changes in either contact frequency and enhancer activity correlate with increased gene expression, but they also correlate with each other, suggesting a potentially linked functional role during enhancer-promoter communication." The authors could use this opportunity to disentangle the contributions of loops and chromatin modifications a bit more. The exceptions are of interest - e.g., loop is stable, gene expression changes or loop changes, gene expression does not. Highlighting exemplar cases for these exceptions (rather than just a genomics summary) would be nice to see.

      The additional example regions we have included in Supp. Fig. 7 and 12 now showcase a wider variety of scenarios; in addition to more examples of static loops with gene expression changes (Fig. 2, Supp. Fig. 7E-F) and differential loops with matching gene expression changes (Fig. 4, Supp. Fig. 7C-D, Supp. Fig. 12A-C), we now also feature examples of differential loops where gene expression changes in the opposite direction (i.e. a strengthened loop at a down-regulated gene, Supp. Fig. 12D-F).

    1. eLife Assessment

      This study reports a novel function for syntaxin 11, a specialized SNARE protein critical for the immune system whose mutations cause familial hemophagocytic lymphohistiocytosis type 4. The data convincingly show that depletion of STX11 impairs store-operated calcium entry in Jurkat T cells and that this defect is recapitulated in primary cells from a patient suffering from the disease; the authors further show that the syntaxin interacts with the pore subunit of the ORAI1 channel and propose that it primes the channel by promoting the assembly of multimers before activation by its endogenous ligand, the ER Ca2+ sensing protein STIM1. This is a conceptually important claim that challenges the prevailing view that all structural transitions in ORAI1 are STIM-driven. The data are high-quality and broadly consistent with the interpretation, but alternative mechanisms for the defects are not considered; additional work should rule out vesicular trafficking, discuss other mechanisms, and address methodological issues.

    2. Reviewer #1 (Public review):

      Summary:

      Patients with STX11 mutations develop familial hemophagocytic lymphohistiocytosis Type 4, a fatal immune disorder marked by defective T and NK cell cytotoxicity and cytokine storm. The conventional explanation attributes this to impaired cytotoxic granule release, but this has never fully accounted for the broader disease picture. This study proposes an alternative mechanism. The authors show that STX11 is required for store-operated calcium entry through ORAI1 channels, which are essential for both cytotoxic killing and NFAT-driven gene expression in T cells. In STX11-deficient cells, ORAI1 currents drop, NFAT nuclear translocation fails, IL-2 expression is suppressed, and degranulation is impaired. These defects are largely rescued by ionomycin or a constitutively active ORAI1 mutant, placing the primary lesion at calcium signaling rather than the fusion machinery. Mechanistically, STX11 binds the C-terminal tail of ORAI1 via its Habc domain and maintains ORAI1 in a state competent for productive assembly prior to STIM1-dependent gating, a step the authors call "priming."

      Strengths:

      The paper identifies a novel and disease-relevant role for STX11 in calcium channel regulation and raises the possibility of using channel agonists as a therapeutic strategy in the disease. The biochemical and functional data are of high quality and generally consistent with the interpretation. The proposal that a non-conventional syntaxin directly interacts with ion channels to prime its activation is novel and interesting.

      Weaknesses:

      For readers to appreciate the value of patient experiments derived from a single individual, the authors should quote prior studies showing that STX11 protein levels are abolished in all known human STX11 mutations. The priming model, while functionally well-supported, rests on indirect structural evidence, and the precise conformational transition involved remains to be defined. These are acknowledged limitations, but alternate mechanisms have not been explored and formally excluded. More direct evidence should be provided to exclude the possibility that STX11 could act as a conventional SNARE and sustain calcium fluxes by promoting the delivery of additional ORAI1 channels from vesicles.

    3. Reviewer #2 (Public review):

      Summary:

      Vig's lab delineates a critical role for STX11 in CRAC channel function, particularly in the context of the fatal immune disorder familial hemophagocytic lymphohistiocytosis type 4 (FHL4). They demonstrate that Syntaxin 11 directly binds and regulates Orai1, and that STX11 depletion abolishes CRAC currents and downstream signaling. Loss of STX11 reduces IL2 gene expression and impairs degranulation, both of which are rescued by the constitutively active Orai1 mutant H134S, whereas a gain‑of‑function mutant targeting the C‑terminus fails to restore these defects. The authors conclude that STX11 primes Orai1 for optimal local assembly that is independent of STIM1 yet required for CRAC channel gating.

      Strengths:

      This study is firmly grounded in disease biology and demonstrates that STX11 downregulation leads to profound functional defects. Using a comprehensive suite of methods and analyses, the authors interrogate the co-regulation of STX11 and Orai1 and present a near-complete view of STX11's modulatory role in CRAC channel function and downstream signaling pathways. The figures are clear, and the statistical analyses are rigorous and convincing.

      Weaknesses:

      The authors conclude that Syntaxin 11 directly binds Orai1. This conclusion is well supported by a multifaceted approach, including co-immunoprecipitation (co-IP), molecular dynamics simulations, co-localization/FRET assays, and targeted mutational analysis-all of which are thoroughly executed. While the interaction appears reasonably strong in co-IP experiments, the STX11-Orai1 interaction is comparatively weaker in pull-down assays, which the authors attribute to instability of the purified His-STX11 protein. A remaining gap is direct evidence of interaction in live cells; this is understandably challenging given that fluorescent tagging of STX11 is not feasible. Fully resolving this question lies beyond the scope of the present study and will require more advanced approaches to capture STX11 binding dynamics.

    4. Author response:

      eLife Assessment

      This study reports a novel function for syntaxin 11, a specialized SNARE protein critical for the immune system whose mutations cause familial hemophagocytic lymphohistiocytosis type 4. The data convincingly show that depletion of STX11 impairs store-operated calcium entry in Jurkat T cells and that this defect is recapitulated in primary cells from a patient suffering from the disease; the authors further show that the syntaxin interacts with the pore subunit of the ORAI1 channel and propose that it primes the channel by promoting the assembly of multimers before activation by its endogenous ligand, the ER Ca2+ sensing protein STIM1. This is a conceptually important claim that challenges the prevailing view that all structural transitions in ORAI1 are STIM-driven. The data are high-quality and broadly consistent with the interpretation, but alternative mechanisms for the defects are not considered; additional work should rule out vesicular trafficking, discuss other mechanisms, and address methodological issues.

      We thank the editor and reviewers for assessing our work. Although significant amount of data in this paper already rule out any potential defects in the vesicular trafficking of Orai1 in cells lacking STX11, we will still include the additional suggested experiments. In the revised version, we will include the various experiments that we had already performed to measure vesicular trafficking and ER-PM junctions in STX11 depleted cells. We will discuss any remaining alternate explanations, include missing methods, quantifications and calibrations, where applicable, and provide response to each of the reviewer’s comments.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      For readers to appreciate the value of patient experiments derived from a single individual, the authors should quote prior studies showing that STX11 protein levels are abolished in all known human STX11 mutations. The priming model, while functionally well-supported, rests on indirect structural evidence, and the precise conformational transition involved remains to be defined. These are acknowledged limitations, but alternate mechanisms have not been explored and formally excluded. More direct evidence should be provided to exclude the possibility that STX11 could act as a conventional SNARE and sustain calcium fluxes by promoting the delivery of additional ORAI1 channels from vesicles.

      In the revised version, we will include references for the prior STX11 human mutations that have been biochemically characterized till date (Bryceson, Rudd et al. 2007);(Muller, Chiang et al. 2014);(Macartney, Weitzman et al. 2011);(Marsh, Satake et al. 2010). As the reviewer has correctly pointed out, the STX11 protein levels were almost completely abolished in these studies. Therefore, the prior mutations are essentially comparable to the frameshift mutation characterized in this study, in terms of the mechanisms underlying the phenotypic defects reported here versus earlier. From a mechanistic point of view, we believe that our data from even a single FHLH4 patient, where STX11 levels were severely depleted, and additional knockdown studies across three different cell lines, are representative of all STX11 patients that have been reported thus far.

      Regarding the Reviewers’ concern that absence of STX11 as a conventional SNARE could affect Orai1 channel delivery from intracellular vesicles. We would like to point out the following:

      (1) In Miao et al. 2013 (Miao, Miner et al. 2013), Figure 3C-D, we conclusively showed that expression of a dominant negative mutant of NSF, a non-redundant protein in vesicle trafficking, impaired vesicle trafficking but did not impair SOCE. This experiment had essentially ruled out a role for vesicle trafficking in SOCE. In the same paper, we had also shown that Orai1 levels in the PM do not increase post-store depletion (Figure 3, figure supplement 2).

      (2) In this manuscript (Supplementary Figure 3B), we have shown that U2OS cells stably expressing Orai1-BBS-YFP have identical levels of Orai1 in the PM with and without STX11 depletion. This shows that the biosynthesis or delivery of Orai1 to the PM is not affected by STX11 depletion, another broadly classified member of the vesicle trafficking. The levels were also assessed in store-depleted U2OS cells but not included here because in Miao et al. 2013 we had already shown that levels of PM Orai1 are essentially equal in resting and store-depleted cells. In our revised submission, we will include the data from store-depleted cells in U2OS and also repeat this experiment in the other cell types used in this paper. In addition, in our revised submission, we will include three different vesicle trafficking assays performed in STX11 depleted cells.

      (3) Most importantly, in Figure 7I-J of this manuscript, we showed that calcium influx from a constitutively active mutant Orai1 (Orai H134S) is identical between STX11 depleted and scramble control cells. If wildtype Orai1 was indeed stuck in vesicles in STX11 depleted cells, then how would H134S Orai1 be able to rescue the defect in SOCE? In fact, the Orai1 mutant calcium flux assays were done using a 20X water objective, to visualize and confirm whether the expression of mutant and WT Orai1 was comparable in the PM. We will include the quantification of PM levels of Orai1 mutants w.r.t WT Orai1 in the revised paper.

      (4) We have generated and been using HEK293, U2OS and Jurkat cell lines that stably express fluorescently tagged Orai1 for most of our experiments (Miao, Miner et al. 2013); (Li, Miao et al. 2016);(Ramanagoudr-Bhojappa, Miao et al. 2021). In each case, we have never observed Orai1 in intracellular vesicles with or without store depletion. In all cases, it is constitutively and stably expressed in the PM.

      In summary, significant amount of data in this paper already rule out any potential reduction in the PM levels of Orai1 in cells lacking STX11. We will still do the additional experiments suggested by the Reviewer 1.

      Regarding STX11 induced precise conformational transition, we are trying to setup collaborations with scientists who might be able to visualize this in vivo.

      The readers should note that purification of isolated pore subunits of ion channels followed by crystallization or expression in membranes for cryo-EM is currently considered a gold standard in the analysis of ion channel pore subunits. However, we have shown that ion channels are dynamic macromolecular complexes, in vivo (Li, Miao et al. 2016), where synaptic proteins dynamically bind to induce conformational changes and affect their stoichiometry (Li, Miao et al. 2016). Please also see (Chorev, Baker et al. 2018) and (Dorwart, Wray et al. 2010). More advanced in vivo approaches therefore need to be developed to enable visualization of the dynamics of ion channel macromolecular complexes in the native environment. In the absence of such approaches, the structural insights obtained from detergent purified subunits will remain incomplete and biased.

      Reviewer #2 (Public review):

      Weaknesses:

      The authors conclude that Syntaxin 11 directly binds Orai1. This conclusion is well supported by a multifaceted approach, including co-immunoprecipitation (co-IP), molecular dynamics simulations, co-localization/FRET assays, and targeted mutational analysis-all of which are thoroughly executed. While the interaction appears reasonably strong in co-IP experiments, the STX11-Orai1 interaction is comparatively weaker in pull-down assays, which the authors attribute to instability of the purified His-STX11 protein. A remaining gap is direct evidence of interaction in live cells; this is understandably challenging given that fluorescent tagging of STX11 is not feasible. Fully resolving this question lies beyond the scope of the present study and will require more advanced approaches to capture STX11 binding dynamics.

      We thank the reviewer for acknowledging that the above studies will require standardization of advanced techniques which are beyond the scope of the present study. We plan to continue developing methods that will allow us to visualize the binding and unbinding of STX11 to Orai1 in vivo.

      References:

      Bryceson, Y. T., E. Rudd, C. Zheng, J. Edner, D. Ma, S. M. Wood, A. G. Bechensteen, J. J. Boelens, T. Celkan, R. A. Farah, K. Hultenby, J. Winiarski, P. A. Roche, M. Nordenskjold, J. I. Henter, E. O. Long and H. G. Ljunggren (2007). "Defective cytotoxic lymphocyte degranulation in syntaxin-11 deficient familial hemophagocytic lymphohistiocytosis 4 (FHL4) patients." Blood 110(6): 1906-1915.

      Chorev, D. S., L. A. Baker, D. Wu, V. Beilsten-Edmands, S. L. Rouse, T. Zeev-Ben-Mordehai, C. Jiko, F. Samsudin, C. Gerle, S. Khalid, A. G. Stewart, S. J. Matthews, K. Grunewald and C. V. Robinson (2018). "Protein assemblies ejected directly from native membranes yield complexes for mass spectrometry." Science 362(6416): 829-834.

      Dorwart, M. R., R. Wray, C. A. Brautigam, Y. Jiang and P. Blount (2010). "S. aureus MscL is a pentamer in vivo but of variable stoichiometries in vitro: implications for detergent-solubilized membrane proteins." PLoS Biol 8(12): e1000555.

      Li, P., Y. Miao, A. Dani and M. Vig (2016). "alpha-SNAP regulates dynamic, on-site assembly and calcium selectivity of Orai1 channels." Mol Biol Cell 27(16): 2542-2553.

      Macartney, C. A., S. Weitzman, S. M. Wood, D. Bansal, M. Steele, M. Meeths, M. Abdelhaleem and Y. T. Bryceson (2011). "Unusual functional manifestations of a novel STX11 frameshift mutation in two infants with familial hemophagocytic lymphohistiocytosis type 4 (FHL4)." Pediatr Blood Cancer 56(4): 654-657.

      Marsh, R. A., N. Satake, J. Biroschak, T. Jacobs, J. Johnson, M. B. Jordan, J. J. Bleesing, A. H. Filipovich and K. Zhang (2010). "STX11 mutations and clinical phenotypes of familial hemophagocytic lymphohistiocytosis in North America." Pediatr Blood Cancer 55(1): 134-140.

      Miao, Y., C. Miner, L. Zhang, P. I. Hanson, A. Dani and M. Vig (2013). "An essential and NSF independent role for alpha-SNAP in store-operated calcium entry." Elife 2: e00802.

      Muller, M. L., S. C. Chiang, M. Meeths, B. Tesi, M. Entesarian, D. Nilsson, S. M. Wood, M. Nordenskjold, J. I. Henter, A. Naqvi and Y. T. Bryceson (2014). "An N-Terminal Missense Mutation in STX11 Causative of FHL4 Abrogates Syntaxin-11 Binding to Munc18-2." Front Immunol 4: 515.

      Ramanagoudr-Bhojappa, R., Y. Miao and M. Vig (2021). "High affinity associations with alpha-SNAP enable calcium entry via Orai1 channels." PLoS One 16(10): e0258670.

    1. eLife Assessment

      This valuable study demonstrates that a multi-step differentiation programme in bacteria combining a bistable switch with two quorum-sensing systems is capable of generating autonomous and self-organized spatial patterns. The evidence for the core engineering system supporting patterning across several conditions is convincing, albeit incomplete for the stronger differentiation/maturation claims because the irreversibility of the proposed states is not consistently established, and some modelling and conceptual interpretation details require further clarification.

    2. Reviewer #1 (Public review):

      Summary:

      This paper by Boni and colleagues presents the engineering of a multi-step differentiation program in Escherichia coli based on synthetic gene circuits. The motivation behind the study was to engineer a system capable of undergoing differentiation in a step-wise manner without the presence of external spatial cues and without inducers added during the differentiation process. To achieve this, the authors created several synthetic gene circuits, one being a toggle switch, and the others being quorum-sensing-mediated gene expression modules. The outputs of the differentiation process are fluorescent proteins, which allowed the authors to quantify the behavior of the system using fluorescence intensity measurements. The authors additionally built a multi-component mathematical model which is able to reproduce the experimental data.

      The data presented are convincing and support the claims; the work is well executed.

      Strengths:

      (1) The differentiation process proceeds autonomously after the initial step in liquid culture in the presence of external inducers.

      (2) It is indeed a step-wise process.

      (3) The mathematical model predicts the outcome (% of green, blue and red FP-expressing cells in the population) when changing the initial ratio of green:blue FP-expressing cells.

      Weaknesses:

      (1) No spatial pattern emerges. There are some isolated colonies that turn on the downstream FPs, but I do not see a pattern, really. Nonetheless, some colonies do differentiate (i.e. they turn on additional FPs).

      (2) The mathematical model appears somewhat superfluous. While it can clearly reproduce the data, it is not used to make interesting predictions, changing parameters (and not initial conditions) that guide further experimental implementations.

      Future directions

      The utility of this differentiation process (e.g. in metabolic engineering or for the study of biofilm formation and antibiotic resistance) will become clearer once the FPs are substituted with functional proteins that exert an effect on the cells.

    3. Reviewer #2 (Public review):

      In this manuscript, the authors implement a three-step genetic programme in E. coli that converts an initially homogeneous population into spatially structured sender, receiver, and "matured" receiver colonies on agar without externally supplied positional information. They combine a TetR/LacI toggle switch for symmetry breaking, LuxI/LuxR quorum sensing for a paracrine signalling step, and CinI/CinR for an autocrine signalling-like maturation step, and complement the experiments with a mathematical model that qualitatively reproduces pattern formation over a range of initial conditions.

      While the article has many strengths such as a clear conceptual framing using Waddington landscapes, a modular and carefully optimised circuit design, thorough experimental characterisation of the toggle and quorum-sensing modules, integration of spatial modelling with experiments, and generally clear writing and figures, I think it will benefit the article to clarify the definition and stability of "differentiated" states, clarify several quantitative and modelling aspects, better explain how fitted curves and promoter engineering were done, and improve some figure design and wording to avoid ambiguity.

      Detailed comments below:

      (1) P5-8 / and more generally: A major concern is that producing a reporter output is not, by itself, differentiation. For a state to be credibly called "differentiated", it should be stable (self-maintained) over relevant timescales, ideally in the absence of the inducing context. As written, the manuscript sometimes seems to equate cell type with reporter expression. I strongly suggest adding a short subsection explicitly defining state versus output, and for each claimed state, stating whether it is stable/bistable or unstable/reversible, with evidence. Concretely, the authors should enumerate:<br /> a) Toggle-derived sender versus receiver: stable? under what conditions (inducer ranges, hysteresis window)?<br /> b) Paracrine-induced "red" receivers: is this a stable differentiated state, or a context-dependent induction requiring proximity to senders?<br /> c) "Mature" (yellow) state: does it persist after removal from the spatial signal field? If not, it should be described as an induced output programme rather than a mature lineage state.

      At present, later sections (and the "maturation" language) risk over-stating what is demonstrated.

      (2) Figure 2d: It is unclear whether this panel is intended to be qualitative (schematic/illustrative) or generated from quantitative data. The legend should explicitly state the origin (e.g., representative image, averaged data, simulation output, schematic) and, if quantitative, what was measured, how many replicates, and how the visualisation was constructed.

      (3) Figure 2e: The cross-sectional line is described as meant to be comparable, yet the leftmost plot appears to have a different slope from the others. The authors should explain whether this reflects a different scaling/normalisation, a different underlying dataset/condition, or simply a plotting artefact. If these are fitted trends, report the fit function (see also the comment on fitted lines below).

      (4) Around P7-8: (saddle/separatrix description): When describing the saddle or separatrix between the two valleys, it would be helpful to briefly connect this more directly to a quantitative dynamical-systems perspective: for instance, the intersection of nullclines and how nullcline geometry changes under IPTG/aTc induction. This will make the landscape picture more complete for readers familiar with the original genetic toggle switch work (Garder et al., 2000).

      (5) P9, lines 157-159: The current phrasing ("in absence of noise, the system would be fully deterministic... in living cells, however, stochastic bursts... change the trajectory") risks conflating predicting population-level percentages with predicting colony-level trajectories. It would help to clearly separate (i) the ability to predict the overall fraction of ON/OFF (green/blue) colonies from inducer conditions (which is largely deterministic at the population level) from (ii) the intrinsically stochastic choice of state made by any given founder cell and its colony.

      (6) P11, lines 193-195 (promoter engineering): The main text currently only refers to screening variants and choosing pLux76; I suggest briefly stating in the main text (not only in the supplement) what was changed (for example, promoter box variants, core promoter strength modifications) and what design criteria were used (reduced leakiness, increased dynamic range).

      (7) Use of fitted lines (Figures 2, 4, 5, 7): Wherever fitted curves are overlaid on data, the asuthors should indicate in the figure legend the explicit form of the fit as well as the fit equation/ parameters. As a reader, it is difficult to interpret what is empirical smoothing versus what is a mechanistic functional form.

      (8) P13, lines 232-235: The comparison between induction directly with C6-HSL and induction from sender colonies is qualitative ("significantly smaller range"). The authors should provide distances (for example, in mm) for the induction range in each case and, if possible, approximate total HSL amounts or concentrations, so that the reader can appreciate the magnitude of the difference.

      (9) P13, lines 259-262: The authors model the transition to the stationary phase via a monotonically decreasing sigmoid in time for biosynthetic capacity. What is the rationale or literature basis for this approach to model entry into the stationary phase? The authors should cite prior work and clarify why this form is appropriate here, versus alternatives (nutrient diffusion limitation, logistic growth with resource depletion, etc.).

      (10) Figure 6c: Are the areas of the plate shown in each column the same field of view across conditions/time, or are these simply representative regions selected per condition (possibly from different plates)? The caption/legend should clarify whether these are matched locations and how images were chosen.

      (11) Figure 7a: The combination of solid, dashed, and dash-dot arrows/lines is visually hard to read. I suggest replacing the dash-dot line with a fully dotted line or using different colours (if consistent with journal style) to improve readability.

      (12) Figure 7e and similar analyses: The authors should explain in the Methods and/or captions how "distance from sender colonies" is computed when multiple senders exist. Is the distance always measured to the nearest sender, and how are cases handled where a receiver is in the overlapping influence of several senders? This clarification is important for interpreting the fitted curves.

    4. Reviewer #3 (Public review):

      This manuscript presents an engineered 3-step circuit in E. coli that combines toggle-switch-based symmetry breaking with quorum-sensing interactions to generate colony-scale spatial patterns. The work is interesting as a synthetic circuit integration study and as a demonstration of self-organized patterning across physically separated colonies. The authors provided a compelling demonstration of the characterization/tuning of parts to guide the overall system engineering. A notable strength is the demonstration that a single circuit can generate a range of self-organized spatial patterns across separate colonies.

      However, I think the paper needs to tone down the extent to which the system demonstrates multi-step differentiation or morphogenesis, which is not critical for making the paper valuable. Only the first step of their circuit design (Figure 1), the toggle switch, generates stable alternative states. The latter steps are mainly signal-dependent reporter activation states layered on top of the blue receiver state, rather than true fate transitions. The authors explicitly state that red expression is added without replacing the blue identity, and they also acknowledge that red cells lose their identity upon restreaking unless they remain near sender cells. That substantially weakens the differentiation analogy and makes the Waddington framing too strong.

      A related concern is that the 3rd step does not introduce a new spatial organizing rule. The authors show that the second signal remains confined to cells already receiving the first signal, and explicitly conclude that it functions only as an autocrine cue rather than a second paracrine layer. As a result, the 3-step system seems more like an added local readout or maturation layer. Overall, the main 2-step outcome is sparse green sender colonies surrounded by red-expressing blue receivers, with distant receivers remaining blue. That is a valid engineered pattern, but it is still a local, threshold-response circuit architecture.

      The autonomy claim should be toned down and stated more precisely. The plate patterning occurs without externally imposed spatial gradients, which is a strength. However, by design, the overall system behavior depends strongly on pre-culture inducer conditions that set the sender:receiver ratio, and this externally imposed history is central to the final pattern. This property is tied to how the circuit is designed where steps 2 and 3 largely respond to symmetry breaking introduced in step 1, which is dependent on both history and initialization on the plate. In particular, currently the pattern formation process is quite variable (e.g. figure 5), depending on how different colonies flip the toggle switch, and consequently, how many become senders and how many become receivers. It would have been fascinating if they could also demonstrate the differentiation within individual colonies, leading to intra-colony patterns. This aspect should at least be discussed.

      The mathematical model is useful in guiding both the characterization of parts, modules and the overall system. However, the claims around its quantitative predictive power should also be made narrower. The simulations are built from multiple fitted and partly hand-tuned components, including toggle-switch response curves, colony-growth rules, diffusion, reporter-response functions, and activity decline. This supports a calibrated qualitative reconstruction of the observed patterns, but not a strong predictive or mechanistic validation.

      Other specific points:

      (1) Given the topic of the work, the authors should cite closely relevant studies in programming pattern formation, including:<br /> Cao et al, Cell 2016 Collective space-sensing coordinates pattern scaling in engineered bacteria<br /> Rajasekaran et al, Cell 2024 A programmable reaction-diffusion system for spatiotemporal cell signaling circuit design<br /> Lu et al, BioRxiv 2024 Discovery of interpretable patterning rules by integrating mechanistic modeling and deep learning

      (2) The model assumes identical diffusion coefficients for C6-HSL and C14-HSL despite their substantially different molecular sizes and hydrophobicities. This assumption could distort kinetic lag with differential diffusion in explaining the autocrine confinement of the third step. Its impact should at least be explored in the simulations.

      (3) The mCherry response parameters change significantly between the 2-step and 3-step systems. The authors acknowledged this change but did not provide a clear explanation.

      (4) The 3-step system is evaluated at only a single condition with no simulation comparison, in contrast to the systematic 11-condition validation of the 2-step system.

    1. eLife Assessment

      This important study provides detailed insights into the metabolic states of hemocyte populations across developmental stages and in both physiological and pathological contexts, including during immune challenge. The study provides convincing evidence by comparing the relative utilization of glycolysis and oxidative phosphorylation in Drosophila larval immune cells, and can have implications for metabolic programs that shape immune function in health and disease.

    2. Reviewer #1 (Public review):

      Summary:

      The metabolic profiles of immune cells under steady-state or immune-activated conditions remain poorly characterized. The authors find that embryonically derived hemocytes in Drosophila larvae predominantly utilize mitochondrial respiration to generate energy and exhibit minimal glycolysis rates under unchallenged conditions. Hemocytes developmentally elevate ATP production rates. Mitochondrial respiration drives metabolic activation in larval hemocytes. More specifically, lamellocytes exhibit unique metabolic activities, including enhanced trehalose catabolism and mitochondrial remodeling, required for their encapsulation response.

      Strengths:

      The study shows the metabolism that is most likely to operate in different immune cells in Drosophila during development and also during infection. This is related to mitochondrial organization and proliferation and/or differentiation state.

      Weaknesses:

      Even though there is a rigorous analysis of mitochondrial activity using the Sea Horse analyzer, the analysis of diverse mitochondrial activities in the different immune cell types across development and in infection could be carried out using microscopy. ROS, mitochondrial membrane potential, NADH/+ and FADH/+ levels in vivo are likely to give a more specific readout of change in cellular activities. The activities of mitochondrial fusion and fission need to be collectively tested to understand their role in development and also in infection. The relevance of the change in mitochondrial activity for development or immunity remains to be tested.

    3. Reviewer #2 (Public review):

      Summary:

      This study presents an analysis of the metabolism of Drosophila larval immune cells during development and activation. The authors compared the utilization of glycolysis and oxidative phosphorylation for energy metabolism. Although this topic has been widely discussed and well-studied in immune cell research, particularly in mammals, it has received little attention in insects. The authors demonstrated that quiescent and activated larval Drosophila immune cells predominantly use mitochondrial oxidative phosphorylation to produce energy. This finding is significant for the emerging field of insect immunometabolism research and is interesting in comparison to mammalian immunity, where immune cell activation is often associated with a shift toward greater reliance on glycolysis.

      Strengths:

      Using the Agilent Seahorse system, the authors developed and fine-tuned a method to measure the energy metabolism of Drosophila immune cells, obtaining high-quality, robust data. Through genetic manipulations targeting immune cells specifically, they analyzed metabolic changes in cells with different activations, going beyond developmental changes. They convincingly demonstrated ATP production, primarily in the mitochondria of immune cells, at various developmental stages and in various activated states. The results presented mostly support the conclusions drawn. This methodology and its results are valuable for further studies of insect immunometabolism. In a broader context, they are also valuable for comparing the metabolism of immune cells across different animal groups.

      Weaknesses:

      The genetic manipulations used were suitable for obtaining immune cells of various types and activation states, such as proliferation, differentiation, and immune activation. However, this method has limitations: the mixture of different cell types was always analyzed, and the specific type of interest was often a minority cell population. Had the other cells remained in their initial control state, the observed change in metabolism could have been primarily attributed to the desired cell type. However, the remaining cells that did not transform into the desired type were also usually influenced or activated in some way, making it difficult to determine to which group the observed change should be attributed. For example, consider the induction of lamellocyte differentiation using Hml>Hop[tum]. There are approximately 1,000 lamellocytes per larva, but according to Supplementary Figure 4, there are still about 5,000 Hml+ cells, and even these cells have activated Jak/Stat signaling. Therefore, it can be assumed that they are also activated. After a real infection, the proportion of lamellocytes is greater, but the remaining plasmatocytes are also activated. The authors should mention these limitations more clearly. However, as the authors correctly note, solving this problem will require single-cell approaches, which current technologies still limit. I see this as a problem when interpreting the proliferation effect. The crucial question is what percentage of the analyzed cells induced by Hml>Ras[V12] were actually in the division stage. Not all hemocytes are Hml+, so not all are induced. Of those that are induced, how many are in the division stage at the time of analysis? Meanwhile, those that were not dividing at that moment also had activated Ras, which triggers many processes besides division. Information on what percentage of the analyzed cells were dividing is missing. This information is important because the finding that dividing Drosophila immune cells primarily use mitochondria and oxidative phosphorylation to produce ATP contrasts with the debated significance of the Warburg effect in dividing mammalian cells. This finding would be significant, but unfortunately, it is not robustly supported by the presented data.

    4. Reviewer #3 (Public review):

      Summary :

      This study investigates the metabolic profiles of hemocytes across multiple stage/conditions and suggests that hemocytes act as regulators of metabolism rather than merely receivers of metabolic cues. The authors show that hemocytes rely primarily on mitochondrial respiration, which is further enhanced during proliferation in development or upon genetic manipulation of plasmatocytes, but not crystal cells.

      Metabolic respiration is also activated in lamellocytes, and this activation correlates with changes in mitochondrial morphology. The authors further attempt to identify mechanisms underlying this activation, proposing that mitochondrial fission may contribute to the ability of lamellocytes to encapsulate wasp eggs.

      Strengths:

      This work provides detailed and valuable insights into the metabolic phenotypes of hemocyte populations at different developmental stages and under both physiological and pathological conditions. The authors perform a longitudinal assessment of hemocyte metabolism and compare metabolic states across contexts.

      Importantly, they provide evidence that hemocytes regulate metabolism to perform essential immunological functions, such as wasp egg encapsulation. This reinforces the view that hemocytes are key regulators and communicators that adapt their metabolic programs according to developmental and environmental demands.

      Weaknesses:

      The results presented are insightful, although several controls and validations could strengthen the conclusions. It would be preferable to also include responder transgenes alone as a control for leakiness, and the scRNA-seq findings would benefit from in vivo validation.

      Some conclusions appear inconsistent or insufficiently supported. For instance, although mitochondrial respiration in plasmatocytes peaks at 96 h AEL, this increase is not accompanied by detectable mitochondrial rearrangement, which remains constant between 96 h AEL and 120 h AEL.

      In general, the authors should temper some statements or provide further data.

    1. eLife Assessment

      This manuscript makes a valuable contribution to the concept of fragility of meta-analyses via the so-called 'ellipse of insignificance for meta-analyses' (EOIMETA). The strength of evidence is convincing, supported primarily by an example of the fragility of meta-analyses in the association between Vitamin D supplementation and cancer mortality, but the approach could be applied in other meta-analytic contexts. The significance of the work could be enhanced with a more thorough assessment of the impact of between-study heterogeneity, additional case studies, and improved contextualization of the proposed approach in relation to other methods.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      This manuscript addresses an important methodological issue-the fragility of meta-analytic findings-by extending fragility concepts beyond trial-level analysis. The proposed EOIMETA framework provides a generalizable and analytically tractable approach that complements existing methods such as the traditional Fragility Index and Atal et al.'s algorithm. The findings are significant in showing that even large meta-analyses can be highly fragile, with results overturned by very small numbers of event recodings or additions. The evidence is clearly presented, supported by applications to vitamin D supplementation trials, and contributes meaningfully to ongoing debates about the robustness of meta-analytic evidence. Overall, the strength of evidence is moderate to strong.

      Strengths:

      (1) The manuscript tackles a highly relevant methodological question on the robustness of meta-analytic evidence.

      (2) EOIMETA represents an innovative extension of fragility concepts from single trials to meta-analyses.

      (3) The applications are clearly presented and highlight the potential importance of fragility considerations for evidence synthesis.

    3. Reviewer #3 (Public review):

      Summary and strengths:

      In this manuscript, Grimes presents an extension of Ellipse of Insignificant (EOI) and Region of Attainable Redaction (ROAR) metrics to meta-analysis setting as metrics for fragility and robustness evaluation of meta-analysis. The author applies these metrics to three meta-analyses of Vitamin D and cancer mortality, finding substantial fragility in their conclusions. Overall, I think extension/adaption is a conceptually valuable addition to meta-analysis evaluation, and the manuscript is generally well-written.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript addresses an important methodological issue-the fragility of meta-analytic findings-by extending fragility concepts beyond trial-level analysis. The proposed EOIMETA framework provides a generalizable and analytically tractable approach that complements existing methods such as the traditional Fragility Index and Atal et al.'s algorithm. The findings are significant in showing that even large meta-analyses can be highly fragile, with results overturned by very small numbers of event recodings or additions. The evidence is clearly presented, supported by applications to vitamin D supplementation trials, and contributes meaningfully to ongoing debates about the robustness of meta-analytic evidence. Overall, the strength of evidence is moderate to strong.

      Strengths:

      (1) The manuscript tackles a highly relevant methodological question on the robustness of meta-analytic evidence.

      (2) EOIMETA represents an innovative extension of fragility concepts from single trials to meta-analyses.

      (3) The applications are clearly presented and highlight the potential importance of fragility considerations for evidence synthesis.

      Reviewer #3 (Public review):

      (1) The manuscript would benefit from a clearer explanation of in what sense EOIMETA is generalizable. The author mentions this several times, but without a clear explanation of what they mean here.

      This is a point I was remiss not to better elucidate. With regards to generalisation, the text has been modified to explicitly state that generalisability in this context means no specific study dependence, just a net number of subjects required to flip a result. The text reads:

      “Atal's method is highly useful, but one possible objection is that it has the downside of non-generalisability, as it finds very specific combinations of trials and patients that would have to be re-coded (events classified as non-events and vice-versa) for results to become insignificant. For example, an Atal meta-analytic fragility of 4 pertains to a specific and often unique circumstance when 4 patients could be recoded from a specific study or combinations thereof to change outputs, but this does not generalise to any 4 patients in that meta-analysis. This makes this definition of meta-analytic fragility useful but not general, and perhaps less intuitive to interpret than a typical RCT fragility metric. In this work, we establish a generalizable meta-analytic fragility metric, based upon Ellipse of Insignificance (EOI) analysis for dichotomous outcome trials. This method creates a pool of events and non-events in both arms, adjusted for weighing, and answers the general question of how many patients would have to be effectively recoded in a meta-analysis for results to flip, without requiring specific study identification.”

      (2) The authors mentioned the proposed tools assume low between-study heterogeneity. Could the author illustrate mathematically in the paper how the between-study heterogeneity would influence the proposed measures? Moreover, the between-study heterogeneity is high in Zhang et al's 2022 study. It would be a good place to comment on the influence of such high heterogeneity on the results, and specifying a practical heterogeneity cutoff would better guide future users.

      This is a very fair observation, and I need to better explain myself here! So there are effectively two measures of heterogeneity considered in this work; the typical value from a meta-analysis and the measure of divergence between the crude and the inverse-variance weighed adjusted – when these differ my small amounts, one could conceivably use either measure. I’ve changed the text to better reflect this, including:

      “This modification in akin to pooled in a meta-analysis, and adjusts for study level heterogeneity. After this modification, a standard EOI analysis can then be applied to the vector . In addition, we can also employ ROAR analysis to the same vector, yielding the raw number of patients in either or both arm who could be added a given direction to change the result, and exact combination of control and experimental group redactions required to change the result from a significant finding to a null one. Caveats for implementation and interpretation are outlined in the discussion section.”

      (3) I think clarifying the concepts of "small effect", "fragile result", and "unreliable result" would be helpful for preventing misinterpretation by future users. I am concerned that the audience may be confusing these concepts. A small effect may be related to a fragile meta-analysis result. A fragile meta-analysis doesn't necessarily mean wrong/untrustworthy results. A fragile but precise estimate can still reflect a true effect, but whether that size of true effect is clinically meaningful is another question. Clarifying the effect magnitude, fragility, and reliability in the discussion would be helpful.

      This is an excellent suggestion – I’ve tried to do it with percentages, as in table 2, but these are minute in the case of the vitamin D trials, partially I suspect because they are extraordinarily weak. The Cohen’s H for these meta-analyses yields tiny values, which I think might be tied to the virtually negligible percentages we obtain for number needed to flip. With stronger data, it might be worth expanding this into a useful heuristic measure for robustness, though I don’t think vitamin D data as in this work is going to help us much. In light of the reviewer’s excellent comment, I added the following:

      In light of the reviewer’s excellent comment, I added lines 230-240 in the revised manuscript.

      (4) Comments on revisions:

      I am unable to find the author's responses to my previous round comments (Reviewer #3) in the revision package, though replies to the other reviewers are present. I will provide my updated feedback once these responses are available for review.

      My sincere apologies, I neglected the specific comments in error – this document should address them now, thank you again for giving this your time and consideration!

    1. eLife Assessment

      This is a valuable study presenting convincing data indicating that the bacterial GTPases EngA and ObgE enable single-step reconstitution of functional 50S ribosomal subunits under near-physiological conditions. The study elegantly bridges the gap between the non-physiological aspects of the previous two-step reconstitution method and the extract-dependent iSAT system to enable assembly of highly functional ribosomes under translation-compatible conditions. The reported findings represent substantial progress towards achieving a bottom-up reconstruction of the translation machinery from synthetic parts.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have provided new data and text that addresses all of the reviewers' comments on the previous versions in a wholly satisfactory way.]

      Summary:

      This study presents evidence that addition of the two GTPases EngA and ObgE to reactions comprised of rRNAs and total ribosomal proteins purified from native bacterial ribosomes can bypass the requirements for non-physiological temperature shifts and Mg+2 ion concentrations for in vitro reconstitution of functional E. coli ribosomes.

      Strengths:

      This advance allows ribosome reconstitution in a fully reconstituted protein synthesis system containing individually purified recombinant translation factors, with the reconstituted ribosomes substituting for native purified ribosomes to support protein synthesis. This represents a significant development in the long-term effort to produce synthetic cells.

    3. Reviewer #2 (Public review):

      This study has developed a single-step method to assemble active bacterial ribosomes under near-physiological conditions by using the GTPase factors EngA and ObgE. These factors eliminate the need for the traditional, harsh manipulations of temperature and magnesium levels. This integration is an important step toward the bottom-up construction of synthetic cells.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents evidence that addition of the two GTPases EngA and ObgE to reactions comprised of rRNAs and total ribosomal proteins purified from native bacterial ribosomes can bypass the requirements for non-physiological temperature shifts and Mg+2 ion concentrations for in vitro reconstitution of functional E. coli ribosomes.

      Strengths:

      This advance allows ribosome reconstitution in a fully reconstituted protein synthesis system containing individually purified recombinant translation factors, with the reconstituted ribosomes substituting for native purified ribosomes to support protein synthesis. This represents a significant development in the long-term effort to produce synthetic cells.

      Weaknesses:

      The authors carried out additional experiments indicating that ~60% of the reconstituted ribosomes are functional and that a significant proportion are capable of synthesizing GFP from the correct initiation codon to the correct stop codon, and also of producing an enzymatically active protein at appreciable levels. Their SDS-PAGE and MS analyses of N-terminally tagged GFP are also quite useful but did not assess the frequency of initiation at the wrong start codon, termination at the incorrect stop codon, or the frequency of frameshifting during elongation. This would require examining additional reporters designed to examine dependence on a Shine-Dalgarno sequence or the impact of an in-frame stop codon to assess the fidelity of initiation and termination events, respectively, and one with a programmed frameshift site to assess the elongation fidelity of their reconstituted ribosomes.

      In response to the reviewer’s comment, we expanded the MS analysis and performed additional analyses against amino acid sequences corresponding to all three reading frames (updated Supplementary Data 2). As a result, only a single peptide fragment likely derived from the +1 frame was detected, but its intensity was approximately 1/1000 of that of peptide fragments detected from the normal frame. No other out-of-frame peptides were detected, and no evidence of stop-codon readthrough was found. We consider that these results suggest that the kind of deterioration in ribosome function is not occurring in the reconstituted ribosomes. Because this analysis cannot completely rule out abnormal translation events such as initiation from internal start codons or termination at internal stop codons, we also added a statement acknowledging that further analyses will be required to examine all aspects of the translation reaction.

      Reconstitution studies in the past have succeeded by using all recombinant, individually purified RPs that, if successful here, would have eliminated the possibility that one or more unknown ribosome assembly factors that co-purify with native ribosomes was added to their reconstitution reactions.

      The issue raised by the reviewer was already added at the end of the Discussion in the previous revision. We fully agree with the reviewer’s point and we are currently continuing research in our laboratory aimed at achieving a more fundamental understanding of ribosome assembly.

      Reviewer #2 (Public review):

      This study has developed a single-step method to assemble active bacterial ribosomes under near-physiological conditions by using the GTPase factors EngA and ObgE. These factors eliminate the need for the traditional, harsh manipulations of temperature and magnesium levels. This integration is an important step toward the bottom-up construction of synthetic cells.

      Comments on revisions:

      The authors have addressed my concerns in the previous round of review.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors are urged to acknowledge that more sophisticated reporter assays would be required to compare the frequencies of errors occurring at each step of translation using their reconstituted versus native ribosomes.

      As described in our response to Reviewer #1, we performed additional MS analyses, updated Supplementary Data 2, and added a statement acknowledging the reviewer’s comment.

    1. eLife Assessment

      This study investigates how the HIV inhibitor lenacapavir influences capsid mechanics and interactions with the nuclear pore complex. It provides important insights into how drug-induced hyperstabilization of the viral shell can compromise its structural integrity during nuclear entry. The modeling is technically sophisticated, and the analyses provide convincing support for the mechanistic conclusions.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The paper from Hudait and Voth details a number of coarse-grained simulations as well as some experiments focused on the stability of HIV capsids in the presence of the drug lenacapavir. The authors find that LEN hyperstabilizes the capsid, making it fragile and prone to breaking inside the nuclear pore complex.

      Comments on previous round of revisions:

      I found that the authors addressed my concerns satisfactorily. The other reviewer raised a number of important points regarding the nuances of the model and the interpretation of the simulations, which the authors rebutted. I think the paper in its current form now is a worthwhile addition to the literature.

    3. Reviewer #3 (Public review):

      This is a technically sophisticated study that integrates coarse-grained modeling with live-cell imaging to address an important and timely question regarding HIV-1 capsid inhibition by lenacapavir.

      In summary, in my view, the manuscript represents a solid contribution to the field.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The paper from Hudait and Voth details a number of coarse-grained simulations as well as some experiments focused on the stability of HIV capsids in the presence of the drug lenacapavir. The authors find that LEN hyperstabilizes the capsid, making it fragile and prone to breaking inside the nuclear pore complex.

      I found the paper interesting. I have a few suggestions for clarification and/or improvement.

      (1) How directly comparable are the NPC-capsid and capsid-only simulations? A major result rests on the conclusion that the kinetics of rupture are faster inside the NPC, but are the numbers of LENs bound identical? Is the time really comparable, given that the simulations have different starting points? I'm not really doubting the result, but I think it could be made more rigorous/quantitative.

      (2) Related to the above, it is stated on page 12 that, based on the estimated free-energy barrier, pentamer dissociation should occur in ~10 us of CG time. But certainly, the simulations cover at least this length of time?

      (3) At first, I was surprised that even in a CG simulation, LEN would spontaneously bind to the correct site. But if I read the SI correctly, LEN was parameterized specifically to bind to hexamers and not pentamers. This is fine, but I think it's worth describing in the main text.

      Comments on revisions:

      I found that the authors addressed my concerns satisfactorily. The other reviewer raised a number of important points regarding the nuances of the model and the interpretation of the simulations, which the authors rebutted. I think the paper in its current form now is a worthwhile addition to the literature.

      Reviewer #3 (Public review):

      I have carefully reviewed the manuscript, the two referee reports, and the authors' detailed responses. I appreciate the substantial effort the authors have invested in addressing the reviewers' comments, and I also recognize the strength and ambition of the work. This is a technically sophisticated study that integrates coarse-grained modeling with live-cell imaging to address an important and timely question regarding HIV-1 capsid inhibition by lenacapavir.

      Embedded within Reviewer #2's report are several substantive points that warrant careful consideration, particularly with respect to framing, terminology, and engagement with the broader literature. I view my role here is to distinguish those issues from claims that I do not find to be supported.

      We thank Reviewer 3 for the positive assessment of our work.

      First, I do not agree with Reviewer #2's central assertion that the manuscript lacks novelty or fails to present meaningful new findings. While individual elements of the system studied herecapsid docking at the NPC, lenacapavir-induced capsid hyperstabilization, capsid rupture, and competition with FG- nucleoporins-have been observed previously, this work provides a coherent, mechanistic account of how these elements are coupled. In particular, the proposed sequence linking LEN-induced lattice hyperstabilization, preferential pentamer loss at the narrow end, NPC-induced mechanical stress, and failure of nuclear import represents a nontrivial integration that goes beyond prior phenomenological observations. I therefore do not view this work as redundant with existing literature.

      We thank Reviewer 3 for the positive assessment of our work.

      That said, Reviewer #2 is correct to note that the manuscript would benefit from broader and more explicit engagement with recent independent studies, including computational and hybrid modeling efforts that address capsid mechanics, nuclear entry, and LEN effects using different frameworks. While the authors' bottom-up coarse-grained approach is clearly distinct and, in many respects, more systematically derived, eLife readers would benefit from a clearer discussion of how the present results relate to, complement, or differ from these other approaches. I strongly encourage the authors to add a short discussion paragraph situating their work within this broader context, without disparaging alternative models.

      We have now added several sentences describing papers that use two other CG models that are of some relevance to our work at the beginning of the fourth paragraph of the Introduction, and we have also highlighted the distinguishing features of our work at the end of that paragraph.

      Second, I find that some mechanistic claims in the manuscript would benefit from more careful language distinguishing model-conditioned interpretation from de novo prediction. This applies in particular to discussions of LEN binding heterogeneity and stoichiometry, as well as to conclusions drawn from biased enhanced-sampling simulations. While I agree with the authors that parameterization does not invalidate mechanistic insight, it is important to be precise about what aspects of the behavior emerge from the simulations versus what is constrained by prior experimental knowledge. Modest tightening/revising of language (e.g., "suggests," "is consistent with," "within the model") would address this concern without weakening the scientific conclusions.

      We have revised and softened the language in several places as suggested. However, we do still asert that our overall CG modeling approach is quite rigorous. The use of limited “top down” information on LEN binding is not problematic and in fact warranted in this problem.

      Third, Reviewer #2 raises a legitimate semantic issue regarding the use of the term "elasticity." The manuscript infers changes in capsid mechanical response using heterogeneous elastic network models, which quantify effective stiffness and deformability rather than elasticity in the macroscopic materials sense. I recommend that the authors clarify this definition explicitly in the text to avoid confusion and unnecessary debate.

      We have now added a clarification at the end of the third paragraph of the subsection entitled “LEN binding to the capsid results in hyperstabilized lattice domains”. We have also added text in the second paragraph of the Discussion. Our view is that our perspective is more useful for this problem than a “macroscopic” perspective as the capsid is, in fact, a mesoscopic object and not a macroscopic one.

      Finally, I note that several of Reviewer #2's objections-particularly those asserting circular reasoning, misuse of enhanced sampling methods, or invalidity of coarse-grained predictions reflect a misunderstanding of contemporary bottom-up coarse-grained modeling rather than genuine methodological flaws. I do not believe these points require further rebuttal or revision beyond what the authors have already provided.

      We agree.

      In summary, in my view, the manuscript represents a solid contribution to the field, provided that the authors undertake a limited set of targeted revisions aimed at improving framing, clarity, and engagement with the broader literature. Addressing these points will strengthen the manuscript and ensure that its contributions are clearly and fairly communicated to the community.

      We have done so as suggested by the reviewer.

    1. eLife Assessment

      This valuable study examines the cleavage of motor neuron nucleoporins by proteases of enterovirus D68, a pathogen associated with acute flaccid myelitis. The evidence supporting the effects of EV-D68 proteases on nuclear import and export is generally solid, as is the independent examination of EV-D68 protease on spinal cord neuron toxicity. The specific conclusions related to RNA export were considered overstated relative to the data presented.

    2. Reviewer #1 (Public review):

      Summary:

      Zinn and colleagues investigated the role of proteases 2A and 3C of enterovirus D68 (EVD68), an emerging pathogen associated with outbreaks of acute flaccid myelitis (AFM), a polio-like disease, on the nucleocytoplasmic trafficking in different systems, including human neurons derived from pluripotent cells. They found that 2A specifically cleaved Nup98 and POM121. Using reporter proteins and RNA synthesis and trafficking assays in cells expressing viral proteases, they showed that 2A induces broad loss of the nuclear pore barrier function, but, surprisingly, the RNA export appears to be minimally affected. Since nucleocytoplasmic trafficking defects are known to be associated with neuropatologies, they propose a hypothesis that 2A-dependent cleavage of nucleoporins in motoneurons underlies the development of EVD68-induced AFM. They further show that a 2A-specific inhibitor increases the survival of human neurons differentiated from stem cells upon EVD68 infection.

      Strengths:

      Use of multiple methods to investigate the effect of 2A and 3C expression on nucleoporin cleavage and nucleocytoplasmic trafficking.

      Comments on revisions:

      The following issues remain unresolved:

      First, the authors still do not show representative images confirming specific nucleoporin degradation (Fig.1), which is the main focus of the work.

      Second, the conclusion that 2A-mediated degradation of the nucleo-cytoplasmic barrier does not affect export of the RNA from the nucleus is not supported by the presented data. The representative images shown in Fig 3C do not have the signal for GFP (like in Fig. 2), and therefore it is impossible to see if those cells indeed express EVD68 proteases.

      Moreover, to show RNA export, not only the decrease of nuclear EU signal should be quantified, but also the increase of the cytoplasmic signal. The diminishing of the nuclear staining may not necessarily reflect RNA export, but may well be explained by nuclease activity, all the more relevant in cells expressing 2A, where the nuclear-cytoplasmic barrier is disrupted and cytoplasmic nucleases may enter the nucleus.

      The same applies to images in Fig. 3D. There are no markers of infection; moreover, the experiment description indicates that EU labeling began at 24 h post-infection with an MOI of 5, i.e., essentially all cells should have been infected. This is difficult to believe as the replication cycle of most EVD68 strains in HeLa cells is no longer than 12 h, yet the images do not show any signs of CPE, and demonstrate a strong EU signal, inconsistent with the expected inhibition of nuclear transcription, a known attribute of enterovirus infections.

      The claim that nuclear transcription and RNA export remain unaffected in conditions of 2A-mediated disruption of the nucleo-cytoplasmic barrier is very strong and requires equally strong evidence.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the role of EV-D68 proteases 2A and 3C in nuclear pore complex (NPC) dysfunction and their contribution to motor neuron toxicity. The authors demonstrate that both proteases cleave only a limited number of nucleoporins, with 2A^pro showing the strongest impact by inhibiting nuclear import and export of proteins and disrupting NPC permeability without affecting RNA export. Importantly, treatment with the 2A^pro inhibitor telaprevir reduced neuronal cell death in a dose-dependent manner, achieving neuroprotection at concentrations below those required to inhibit viral replication. The study addresses a relevant mechanism underlying EV-D68-induced neuropathology and explores a potential therapeutic intervention.

    4. Reviewer #3 (Public review):

      Summary:

      The author showed expression of the viral proteases 2Apro and 3Cpro of EV-D68, which cleaved specific components of the nuclear pore complex (Nup98 and POM121 by 2Apro), and 2A but not 3C expression altered nuclear import and export. Similar nucleocytoplasmic transport deficits are observed in EV-D68-infected RD cells and iPSC-derived motor neurons (diMNs). 2A inhibitor telaprevir partially rescued the nucleocytoplasmic transport deficits and suppressed neuronal cell death after infection. While it's clear that 2A can cleave NPC proteins and affect nuclear transport, the link to neurotoxicity after EV-D68 infection is less convincing.

      This study opens up a very intriguing hypothesis: that EV-D68 2Apro could be directly responsible for motor neuron cell death, mediated by POM121 and possibly Nup98 cleavage, that ultimately results in paralysis known as acute flaccid myelitis. This hypothesis notably does run counter to other published data showing that human neuronal organoids derived from iPSCs can support productive EV-D68 infection for weeks without cell death and that EV-D68-infected mice can have paralysis prevented by depletion of CD8 T cells, still with EV-D68 infection of the spinal cord. However, even if 2Apro is not ultimately responsible for motor neurons dying in human infections, that does not exclude the possibility that cleavage of nups could still disrupt motor neuron function. Notably, most children with AFM have some amount of motor function return after their acute period of paralysis, but most still have some residual paralysis for years to life. It is possible that 2A pro could mediate the acute onset of weakness, while T cells killing neurons could determine the amount of long-term, residual paralysis.

      Strengths:

      The characterization of nuclear pore complex components that appear to be targets of both poliovirus and EV-D68 proteases is quite thorough and expansive, so this data set alone will be useful for reference to the field. And the process by which the authors narrowed their focus to EV-D68 2Apro reducing Nup98 and POM121 as consequential to both import and export of nuclear cargo but not RNA was technically impressive, thorough, and convincing. As will be detailed below, when the authors move from studying over-expressed proteases in transformed cell lines to studying actual virus infection in both transformed cell lines and iPSC-derived neurons, some of the data only indirectly support their conclusions; however, the quality of the experiments performed is still high. So even if the claim that 2Apro causes neurotoxicity is circumstantial, the data certainly are intriguing and certainly justify further study of the effects of EV-D68 2Apro on the NPC and how this impacts pathogenesis. This is a convincing start to an intriguing line of inquiry.

      Comments on revisions:

      The authors have returned a stronger revised manuscript, being responsive to most of the combined reviewers' comments. It was especially important to add the clarity and specificity that the data in this manuscript did not establish a direct link for 2Apro causing AFM. The authors have clarified this language adequately, such that it is appropriate to remove the "incomplete" portion of the short assessment as they have requested. Adding in experiments with EV-D68 virus infection to complement their work with recombinant proteases also strengthened their conclusions.

      There are still some areas where discrepancies remain, although these are minor and can mostly be acknowledged as limitations of their approach rather than needing more experiments, unless the authors choose to do the additional experiments. To try to make this understandable, I have copied from the rebuttal letter (*) original comment, (**) author's rebuttal, and (***) a reply to the rebuttal:

      (*)(2) Telaprevir was able to rescue nucleocytoplasmic transport in RD cells at low concentrations (Figure 4A). It is not shown if this correlates with its antiviral effect in RD cells, or could this correlate with inhibition of 2A cleavage of Nup98 or POM121, which is never measured.

      (**) In the aforementioned new experiment in Figure 4A, we have also included a dose-response curve for telaprevir showing its inhibition of POM121 and Nup98 cleavage.

      (***) Fig.4A is in diMN not RD cells. The EC50 of telaprevir could be very different in RD cell vs diMNs. This question remains unanswered.

      (*) (3) Building off of the prior point, the authors' claim that the neuroprotective effect of telaprevir is independent of its antiviral effect is not well-founded. Figure 4E (neuroprotection) was done with MOI 5, and Figure 4G (virus growth) was MOI 0.5. Telaprevir neuroprotection is not shown at MOI 0.5, nor is the neuroprotective effect correlated with inhibition of 2A cleavage of Nup98 or POM121.

      (**) The selection of MOIs for these two experiments was limited by technical considerations. If the viral growth curve were to be performed at MOI 5, it would be confounded by cell death. Further, a low MOI is required in order to allow multiple rounds of infection, and is therefore more sensitive for assaying the effect of telaprevir on viral replication. On the other hand, at MOI 0.5 diMN death is very gradual, and the neuroprotection assay we would have lacked the statistical power to determine whether a rescue of this small magnitude of toxicity is significant. The EC50 of telaprevir is not expected to vary at different MOIs.

      (***) This should be discussed in the Discussion as a limitation of the experiment.

      (**) We have also now correlated the inhibition of 2Apro cleavage of Nup98 and POM121 with the neuroprotective effect at comparable concentrations of telaprevir, as described above.

      (***) Unless you quantify this, my eye disagrees with you. In Fig.4A, cleavage of NUP98 is rescued by 3uM telaprevir, but that does not seem to be the case for POM121.

      Additionally, in Fig. 4D, why is only NLS but not NES is impaired in diMN? This should be discussed.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study examines the cleavage of motor neuron nucleoporins by proteases 2A and 3C of enterovirus D68, a pathogen associated with acute flaccid myelitis. The evidence supporting the effects of EV-D68 proteases on nuclear import and export is solid and confirms previous results on the specific targeting of nucleoporins by proteases from other enteroviruses. However, the claim that cleavage of nucleoporins by EV-D68 2A is neurotoxic, though intriguing, is incomplete, as the evidence is largely indirect.

      We appreciate that the reviewers highlighted multiple strengths of manuscript, including its detailed mechanistic dissection of the disrupted composition and function of the nuclear pore complex during EV-D68 infection, the finding that the viral 2A protease is toxic to motor neurons, and that several novel hypotheses on the pathogenesis of acute flaccid myelitis that are raised by our work.

      It appears that two independent eLife Assessments were made regarding the strength of evidence in our manuscript. The evidence supporting the impact of EV-D68 proteases on the NPC was felt to be solid.

      A second assessment was made as to whether our data support that “the cleavage of nucleoporins by EV-D68 2A is neurotoxic”. We would like to clarify that we did not intend to make this second claim in our manuscript and thought that we had been careful not to do so. In response to reviewer and editorial feedback, we have edited the text to improve the clarity on this issue. Although our data show that 2A<sup>pro</sup> is toxic to motor neurons, it cannot yet be determined whether this toxicity is mediated via 2A<sup>pro</sup>’s effects on the NPC. That is a logical hypothesis that arises from our manuscript, which we are testing through ongoing work that will require a significant volume of experiments that are outside the scope of the present study. We view this manuscript as an important first step towards a comprehensive understanding of the role of the 2A protease in the pathogenesis of AFM. Please see the response to point # 3 of Reviewer 2 below for a more detailed discussion of this issue and the changes we have made to the text in response. We respectfully request that a judgement on the role of nucleoporin cleavage as the mechanism of neurotoxicity not be included in the eLife Assessment.

      Also in response to reviewer feedback that our data was too reliant on the expression of recombinant viral proteins in isolation, we have added additional experiments extending our results into the context of live virus infection of cell lines and motor neurons. We feel that our revised manuscript has been improved as a result of the reviewers’ and editor’s input, and provides strong support for the following claims: (1) NPC composition and function is disrupted during EV-D68 infection, (2) 2A<sup>pro</sup> is primarily responsible for functional disruption, and (3) 2A<sup>pro</sup> is neurotoxic.

      We appreciate your review of this revised manuscript. Detailed responses to each of the reviewers’ comments are provided below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Zinn and colleagues investigated the role of proteases 2A and 3C of enterovirus D68 (EVD68), an emerging pathogen associated with outbreaks of acute flaccid myelitis (AFM), a polio-like disease, on the nucleocytoplasmic trafficking in different systems, including human neurons derived from pluripotent cells. They found that 2A specifically cleaved Nup98 and POM121. Using reporter proteins and RNA synthesis and trafficking assays in cells expressing viral proteases, they showed that 2A induces broad loss of the nuclear pore barrier function, but, surprisingly, the RNA export appears to be minimally affected. Since nucleocytoplasmic trafficking defects are known to be associated with neuropatologies, they propose a hypothesis that 2A-dependent cleavage of nucleoporins in motoneurons underlies the development of EVD68-induced AFM. They further show that a 2A-specific inhibitor increases the survival of human neurons differentiated from stem cells upon EVD68 infection.

      Strengths:

      Use of multiple methods to investigate the effect of 2A and 3C expression on nucleoporin cleavage and nucleocytoplasmic trafficking.

      We thank the reviewer for detailed and accurate review of our manuscript and recognition of these strengths.

      Weaknesses:

      Overall, the paper follows multiple others that extensively investigated the cleavage of nucleoporins by enterovirus 2As, so the results are of limited novelty. The hypothesis that infection of motoneurons is the cause of EVD68-induced neurological complications so far is supported by only one autopsy report. Other data suggest that infection of other cell types, such as astrocytes, and/or inflammatory cell infiltration in the CNS, are likely to be responsible for the symptoms. In any case, the claim that EVD68 is specifically neurotoxic because of the 2A-dependent cleavage of nucleoporins in neurons is unfounded, as the virus will be just as "toxic" for other infected cell types.

      While we agree that other papers have investigated this pathway in other enteroviruses, we note that our work is the first to do so in Enterovirus D68 and the most comprehensive study, in terms of the number of nucleoporins studied. As we reviewed in paragraph 5 of the introduction section, the activities of enterovirus proteases against specific nucleoporins varies from strain to strain, and is important to understand any strain-specific effects before determining whether this pathway is relevant to toxicity in AFM.

      The infection of motor neurons is strongly supported not only by the aforementioned autopsy data [1], but also by mouse model data demonstrating replication of EV-D68 within motor neurons in the anterior horn of the spinal cord.[2] There are also numerous reports of electromyography and nerve conduction studies from human AFM patients demonstrating that the site of pathology is the spinal motor neuron.[3-10]

      By contrast, infection of astrocytes has been demonstrated only in primary murine astrocyte cultures in which no neurons were present [11]. Therefore, while the available data suggest that EV-D68 infection of astrocytes is possible, in the in vivo context of human and mouse spinal cord, tropism to motor neurons appears to be preferential. The relative toxicity of neuron-autonomous vs non-autonomous processes such as glial dysfunction and inflammatory cell infiltration remain to be elucidated, and are not mutually exclusive.

      The paper also requires a more convincing presentation of the data.

      We are uncertain what other specific changes the reviewer would like to see based on this comment, but feel that the revisions have improved the presentation of the data.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the role of EV-D68 proteases 2A and 3C in nuclear pore complex (NPC) dysfunction and their contribution to motor neuron toxicity. The authors demonstrate that both proteases cleave only a limited number of nucleoporins, with 2A^pro showing the strongest impact by inhibiting nuclear import and export of proteins and disrupting NPC permeability without affecting RNA export. Importantly, treatment with the 2A^pro inhibitor telaprevir reduced neuronal cell death in a dose-dependent manner, achieving neuroprotection at concentrations below those required to inhibit viral replication. The study addresses a relevant mechanism underlying EV-D68-induced neuropathology and explores a potential therapeutic intervention.

      Strengths:

      (1) Provides significant mechanistic insight into how EV-D68 proteases alter NPC function and contribute to neuronal toxicity.

      (2) The use of recombinant 2A and 3C proteins allows clear dissection of the specific contribution of each protease.

      (3) Demonstrates a therapeutic effect of telaprevir, with neuroprotection independent of viral replication inhibition, adding translational value to the findings.

      (4) The topic is highly relevant given the association of EV-D68 with acute flaccid myelitis.

      We thank the reviewer for their insightful comments and recognition of these strengths in our study.

      Weaknesses:

      (1) Most experiments were performed with recombinant proteases, lacking validation in the context of viral infection, where both proteases act simultaneously.

      In response to this concern, we have added additional experiments in the context of viral infection. We show that POM121 and Nup98 are also cleaved in motor neurons infected with EV-D68 and that their cleavage is inhibited by telaprevir (Fig 4A). We also repeated the EU pulse-chase RNA export assay in EV-D68-infected RD cells and again found no effect on RNA export (Fig 3D-E).

      (2) The conclusion that RNA export is unaffected requires confirmation during actual infection.

      As above, we have repeated this experiment in EV-D68 RD cells, showing no effect of EV-D68 infection on RNA export.

      (3) The reduction of neurotoxicity by telaprevir does not fully demonstrate that the protective effect is solely mediated through NPC preservation; additional analyses of eIF4G cleavage, nucleoporin integrity, and stress granules are needed.

      We agree that while the evidence in our manuscript raises the hypothesis that telaprevir-mediated neuroprotection is mediated via NPC preservation, it does not fully demonstrate this to be the case. As discussed above, we have been careful to state only the following conclusions: (1) NPC composition and function is disrupted during EV-D68 infection, (2) 2A<sup>pro</sup> is primarily responsible for functional disruption, and (3) 2A<sup>pro</sup> is neurotoxic.

      Future work will determine the extent to which NPC dysfunction contributes to 2A<sup>pro</sup>-mediated motor neuron toxicity versus other potential targets of 2A<sup>pro</sup>, as suggested by the reviewer. This work is already underway in our lab and it is clear to us the additional experiments required will be extensive, likely 1-2 additional manuscripts. These experiments are therefore beyond the scope of the present study, which represents a key first step in this line of inquiry.

      We specifically acknowledged in the Discussion that “A significant limitation of our study, however, is that we cannot exclude potentially toxic effects of 2A<sup>pro</sup> on aspects of host neuronal biology aside from the NPC.” We have also made the following adjustments to the text to make it more clear that this remains an open question:

      Change the title to more clearly separate the effects of 2Apro on NPC function and motor neuron toxicity as independent events: “Enterovirus D68 2A protease causes nuclear pore complex dysfunction and independently contributes to motor neuron toxicity”

      In the abstract, shortened the following sentence: “We therefore sought to determine the impact of EV-D68 proteases on NPC composition and function” to avoid any implicit connection that a mechanistic link has been established between these two concepts. Neurotoxicity is now introduced later in the abstract by saying “Independently, we show…” instead of “We further show…”

      Removed language in the last paragraph of the Results section that may have been construed to suggest a mechanistic linkage: “Because similar deficits have been reported to contribute to neurotoxicity in neurodegenerative disease…” and simply stated “We next sought to determine the extent to which 2Apro activity independently contributes to motor neuron injury during EV-D68 infection.”

      Edited the opening sentence of the discussion, where it was ambiguous whether the word “their” was referring to the enterovirus protease (which was our intent) or to NPC disruption as the cause of motor neuron toxicity. We removed the discussion of toxicity from this paragraph entirely to remove such confusion.

      Edited the final paragraph of the discussion to include “We have also demonstrated that 2A<sup>pro</sup> activity contributes to nucleocytoplasmic transport dysfunction and separately to cell death in motor neurons infected with EV-D68”. We then go on to discuss the hypothesis that this toxicity might be mediated partially or entirely through NPC dysfunction, and propose that this be a focus of further study.

      (4) The study would be strengthened by including another 2A inhibitor (e.g., boceprevir) to confirm the specificity of telaprevir's protective effects.

      While we would like to be able to include multiple pharmacologic inhibitors of 2A<sup>pro</sup>, unfortunately telaprevir is the only known inhibitor of EV-D68 2A<sup>pro</sup>. The same study that identified telaprevir as an EV-D68 2A<sup>pro</sup> inhibitor also evaluated boceprevir and determined that its inhibitory activity against 2A<sup>pro</sup> is minimal [12].

      Reviewer #3 (Public review):

      Summary:

      The author showed expression of the viral proteases 2Apro and 3Cpro of EV-D68, which cleaved specific components of the nuclear pore complex (Nup98 and POM121 by 2Apro), and 2A but not 3C expression altered nuclear import and export. Similar nucleocytoplasmic transport deficits are observed in EV-D68-infected RD cells and iPSC-derived motor neurons (diMNs). 2A inhibitor telaprevir partially rescued the nucleocytoplasmic transport deficits and suppressed neuronal cell death after infection. While it's clear that 2A can cleave NPC proteins and affect nuclear transport, the link to neurotoxicity after EV-D68 infection is less convincing.

      This study opens up a very intriguing hypothesis: that EV-D68 2Apro could be directly responsible for motor neuron cell death, mediated by POM121 and possibly Nup98 cleavage, that ultimately results in paralysis known as acute flaccid myelitis. This hypothesis notably does run counter to other published data showing that human neuronal organoids derived from iPSCs can support productive EV-D68 infection for weeks without cell death and that EV-D68-infected mice can have paralysis prevented by depletion of CD8 T cells, still with EV-D68 infection of the spinal cord. However, even if 2Apro is not ultimately responsible for motor neurons dying in human infections, that does not exclude the possibility that cleavage of nups could still disrupt motor neuron function. Notably, most children with AFM have some amount of motor function return after their acute period of paralysis, but most still have some residual paralysis for years to life. It is possible that 2A pro could mediate the acute onset of weakness, while T cells killing neurons could determine the amount of long-term, residual paralysis.

      We thank the reviewer for their thoughtful comments. As discussed above, we agree that the present data demonstrate that 2A<sup>pro</sup> causes NPC dysfunction and is toxic in motor neurons, but has not proven that the mechanism of neurotoxicity is via NPC dysfunction.

      We appreciate the commentary on novel hypotheses opened by our work. Our recent thinking on this topic has been similar and we look forward to addressing these ideas further in future studies. Motor neuron dysfunction and motor neuron death may ultimately prove to have separate causes. The infection of motor neurons is likely the initiating event, with multiple downstream consequences which may be neuron-autonomous, or mediated by glial and inflammatory responses, or a mixture thereof.

      Strengths:

      The characterization of nuclear pore complex components that appear to be targets of both poliovirus and EV-D68 proteases is quite thorough and expansive, so this data set alone will be useful for reference to the field. And the process by which the authors narrowed their focus to EV-D68 2Apro reducing Nup98 and POM121 as consequential to both import and export of nuclear cargo but not RNA was technically impressive, thorough, and convincing. As will be detailed below, when the authors move from studying over-expressed proteases in transformed cell lines to studying actual virus infection in both transformed cell lines and iPSC-derived neurons, some of the data only indirectly support their conclusions; however, the quality of the experiments performed is still high. So even if the claim that 2Apro causes neurotoxicity is circumstantial, the data certainly are intriguing and certainly justify further study of the effects of EV-D68 2Apro on the NPC and how this impacts pathogenesis. This is a convincing start to an intriguing line of inquiry.

      We appreciate the reviewer’s recognition of our comprehensive evaluation of NPC disruption and our approach to arriving at a mechanistic understanding of this process. We agree with the reviewer’s viewpoint that the present study represents a beginning, rather than a conclusive end to this line of inquiry. For technical reasons, we were able to achieve more rigorous and mechanistic data in cell lines expressing recombinant proteins than in neurons infected with live virus. In response the reviewers’ comments, as described above, we have added additional experiments in this revision in which we further evaluate nucleoporin cleavage and RNA export during live virus infection, and performed these experiments in iPSC-derived neurons whenever it was technically feasible to do so.

      Weaknesses:

      This study falls a bit shy of actually showing that 2Apro effects are causing motor neuron toxicity because the evidence of this is fairly indirect. At points, the authors do admit these limitations, but at other times, they claim to have shown the link directly. The following are reasons why these claims are only indirectly supported:

      We agree that we have shown direct toxicity of 2A<sup>pro</sup> in motor neurons, but have not shown that the mechanism is via NPC dysfunction. We felt that we were careful to frame our conclusions as such. However, we have revised the text to improve the clarity on this point as described above.

      (1) Cleavage of Nup98 and POM121 after EV-D68 infection in RD cells and diMNs is never demonstrated.

      We have added data showing the cleavage of POM121 and Nup98 in EV-D68 infected diMNs (Figure 4A).

      (2) Telaprevir was able to rescue nucleocytoplasmic transport in RD cells at low concentrations (Figure 4A). It is not shown if this correlates with its antiviral effect in RD cells, or could this correlate with inhibition of 2A cleavage of Nup98 or POM121, which is never measured.

      In the aforementioned new experiment in Figure 4A, we have also included a dose-response curve for telaprevir showing its inhibition of POM121 and Nup98 cleavage.

      (3) Building off of the prior point, the authors' claim that the neuroprotective effect of telaprevir is independent of its antiviral effect is not well-founded. Figure 4E (neuroprotection) was done with MOI 5, and Figure 4G (virus growth) was MOI 0.5. Telaprevir neuroprotection is not shown at MOI 0.5, nor is the neuroprotective effect correlated with inhibition of 2A cleavage of Nup98 or POM121.

      The selection of MOIs for these two experiments was limited by technical considerations. If the viral growth curve were to be performed at MOI 5, it would be confounded by cell death. Further, a low MOI is required in order to allow multiple rounds of infection, and is therefore more sensitive for assaying the effect of telaprevir on viral replication. On the other hand, at MOI 0.5 diMN death is very gradual, and the neuroprotection assay we would have lacked the statistical power to determine whether a rescue of this small magnitude of toxicity is significant. The EC<sub>50</sub> of telaprevir is not expected to vary at different MOIs.

      We have also now correlated the inhibition of 2A<sup>pro</sup> cleavage of Nup98 and POM121 with the neuroprotective effect at comparable concentrations of telaprevir, as described above.

      (4) The use of mixed virus isolates only in the diMNs is problematic because different EV-D68 isolates are known to have drastically different effects on pathogenesis in mice. Since all initial data were generated with the MO isolate, adding the additional MD isolate to the diMN experiments actually adds uncertainty to the conclusions. It is not clear if the authors infected different cultures with the different isolates and combined the data or infected all cultures with a mixture of the two isolates. If the former, then the data should be reported separately to see the effect of each individual strain, which would be interesting to EV-D68 virologists. If the latter, then there is no way to know from these data whether one of the two isolates had increased fitness over the other and exerted a dominant effect. If the MD isolate overtook the MO isolate, from which all other data in this manuscript are derived, then we have much less of an idea how much the data from the first three figures supports the final figure.

      We apologize for the lack of clarity in describing this experiment. The MO/2014 and MD/2018 isolates were not mixed. These were performed in separate experiments, each with four biologically independent replicates. The original figure showed the mean and SEM for these 8 replicates together. To improve clarity, we separated each viral strain into its own panel of the figure. We have also increased the rigor of the statistical analysis in this experiment by using Cox proportional hazard regression instead of ANOVA.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Please consider both public reviews above and recommendations for the authors below. The general consensus among reviewers is that more evidence is needed to support the claim that 2A causes motor neuron toxicity during infection.

      Reviewer #1 (Recommendations for the authors):

      Most of the conclusions are made upon analysis of images, yet the images themselves are seldom shown. It is difficult to evaluate the validity of conclusions without seeing the material that was analyzed.

      (1) Figure 1. Representative Western blots should be shown.

      We considered including representative western blots in this already large figure, however the figure size and complexity became un-manageable because the figure summarizes the quantification of 246 Western blots. In the original submission, we uploaded a supporting data file that included complete un-cropped Western blots for all experiments, including ladders, loading controls, and clear labeling of the samples. We believe these data allow the reader to assess the quality and reliability of our Western blot experiments while maintaining the approachability of the figures and data presentation. We have also included these supporting data again in the revised manuscript.

      (2) Figure 3. Representative images should be shown. This is especially important for the ethynyl-uridine labeling experiment. It would be highly surprising that RNA transcription and processing would proceed normally in 2A-expressing cells on the background of a major redistribution of nuclear proteins. One possible explanation for that would be that cells that can be analyzed express a relatively small amount of 2A, which is known to be toxic, and thus may not fully represent the cellular changes upon infection. The results from bona fide infected cells would be much more convincing.

      Representative images have been added for the ethynyl-uridine pulse-chase experiment, and this experiment has been repeated in RD cells infected with EV-D68. Transfection of proteases or infection of the cells utilized the same protocols and timeframes upon which nucleoporin cleavage and disruption of protein transport were found to be present. The timepoint for all of these experiments was selected to precede the onset of toxicity, and the representative images demonstrate normal cellular morphology. We also selected for analysis only GFP+ cells with normal morphology, ensuring that only viable 2A<sup>pro</sup>-GFP-expressing cells were included in the analysis. The new experiments again showed no effect on RNA export. We were equally surprised as the reviewers by this outcome. However, as we note in the text, disruption of RNA export has not been uniformly present across all enteroviruses previously studied.

      (3) Figure 4 A-D. Similarly, representative images should be shown.

      We have added representative images for these experiments, which are now Fig 4B-E.

      (4) Figure 4G. The demonstration that the "neuroprotective" effect of 2A inhibitor is not related to the inhibition of viral replication requires a control showing that a similar inhibition of viral replication by an inhibitor with another target would not similarly diminish cell toxicity.

      Neuronal survival experiments showed inhibition of toxicity with concentrations of telaprevir as low as 0.3 uM, a concentration at which there was no significant effect on viral replication. Telaprevir had only a marginal inhibitory effect on viral replication at 10uM (achieving statistical significance in only one of two strains), and no consistent effect on replication at lower concentrations. Therefore, the suggested control experiment would not be possible, because the neuroprotective concentration of telaprevir does not inhibit viral replication

      Reviewer #2 (Recommendations for the authors):

      Major concerns:

      (1) Most of the experiments were performed with recombinant 2A and 3C proteins. While these experiments are highly informative for dissecting the role of each protease in NPC dysfunction, it would be important to also perform experiments in the context of infection. How are import and export processes affected when both proteases are present during infection? How is passive transport modified under these conditions?

      Thank you for this important comment. Please see the above discussion of additional experiments that we added utilizing live virus infection to complement the experiments that used recombinant proteins.

      (2) The results regarding RNA export in the presence of recombinant 2A and 3C proteases suggest that RNA export is not altered. It would be important to confirm this finding during infection.

      We agree that this is an important experiment, and have done so as described above.

      (3) While the background information suggests that NPC dysfunction contributes to neurotoxicity, the observed reduction of neurotoxicity by telaprevir does not demonstrate that this effect is solely due to the action of 2A on the NPC. It would be important to evaluate the integrity of eIF4G, nucleoporins, and stress granules during treatment.

      We agree that additional experiments would be required to determine the extent to which the toxicity of 2A<sup>pro</sup> is mediated through its effects on the NPC versus other potential targets. Please see above discussion for more details.

      (4) Including another 2A inhibitor (e.g., boceprevir) would strengthen the conclusions by confirming the results obtained with telaprevir.

      Please see above discussion of boceprevir

      Reviewer #3 (Recommendations for the authors):

      (1) Preferred ICTV nomenclature abbreviates rhinovirus as RV instead of HRV, so the authors should change their abbreviations appropriately. See Simmonds et al.

      Archives of Virology (2020) 165:793-797 https://doi.org/10.1007/s00705-019-04520-6

      We have updated these abbreviations accordingly.

      (2) There is no mention of Figures 1C and 1D in the text.

      These have been added in the appropriate locations.

      (3) In the section "2A protease alters nucleocytoplasmic trafficking of protein substrates" it would be very helpful to just directly state what each construct is meant to demonstrate. Along the lines of "NLS-tdTomato should be located in the nucleus, so seeing more signal in the cytoplasm would indicate a defect in nuclear import." And something equivalent for the other two constructs.

      Thank you for the suggestion. We have added descriptions of the use and interpretation for each construct.

      (4) The following sentence would be more accurate with the addition of "partially" because the effect is not returned to normal levels: "The mislocalization of NLS-tdTomato was partially rescued by 3μM telaprevir."

      We have edited this as recommended.

      (5) SNAP29 is probably a typo and meant to be CREB in the legend of Figure 1B.

      Thank you for catching this. We have corrected this to CREB.

      (6) "Panel A" should likely be "Panel E" in the Figure 4F legend.

      We have corrected this to refer to the appropriate panel, which has also been re-lettered due to the addition of new panels to this figure.

      (7) The authors should at least show representative Western blot data used to determine the data for Figure 1 in a supplemental figure.

      As discussed above, these Western blots were included as supplemental data in the original submission, and have also been included in the revised version.

      (8) As suggested in the public comments, if the diMNs were infected separately with the MO and MD strains of EV-D68, those data should be separated from each other and reported individually. In any case, whatever was done (combined virus inoculum or separate inocula) needs to be clarified.

      These data are now reported separately. Please see above discussion for details.

      References:

      (1) Vogt MR, Wright PF, Hickey WF, De Buysscher T, Boyd KL, Crowe JE, Jr. Enterovirus D68 in the Anterior Horn Cells of a Child with Acute Flaccid Myelitis. N Engl J Med. May 26 2022;386(21):2059-2060. doi:10.1056/NEJMc2118155

      (2) Hixon AM, Yu G, Leser JS, et al. A mouse model of paralytic myelitis caused by enterovirus D68. PLoS Pathog. Feb 2017;13(2):e1006199. doi:10.1371/journal.ppat.1006199

      (3) Andersen EW, Kornberg AJ, Freeman JL, Leventer RJ, Ryan MM. Acute flaccid myelitis in childhood: a retrospective cohort study. Eur J Neurol. Aug 2017;24(8):1077-1083. doi:10.1111/ene.13345

      (4) Elrick MJ, Gordon-Lipkin E, Crawford TO, et al. Clinical Subpopulations in a Sample of North American Children Diagnosed With Acute Flaccid Myelitis, 2012-2016. JAMA Pediatr. Feb 1 2018;173(2):134-139. doi:10.1001/jamapediatrics.2018.4890

      (5) Hovden IA, Pfeiffer HC. Electrodiagnostic findings in acute flaccid myelitis related to enterovirus D68. Muscle Nerve. Nov 2015;52(5):909-10. doi:10.1002/mus.24738

      (6) Knoester M, Helfferich J, Poelman R, et al. Twenty-Nine Cases of Enterovirus-D68 Associated Acute Flaccid Myelitis in Europe 2016; A Case Series and Epidemiologic Overview. Pediatr Infect Dis J. Jan 2018;38(1):16-21. doi:10.1097/INF.0000000000002188

      (7) Martin JA, Messacar K, Yang ML, et al. Outcomes of Colorado children with acute flaccid myelitis at 1 year. Neurology. Jul 11 2017;89(2):129-137. doi:10.1212/WNL.0000000000004081

      (8) Saltzman EB, Rancy SK, Sneag DB, Feinberg Md JH, Lange DJ, Wolfe SW. Nerve Transfers for Enterovirus D68-Associated Acute Flaccid Myelitis: A Case Series. Pediatr Neurol. Nov 2018;88:25-30. doi:10.1016/j.pediatrneurol.2018.07.018

      (9) Van Haren K, Ayscue P, Waubant E, et al. Acute Flaccid Myelitis of Unknown Etiology in California, 2012-2015. JAMA. Dec 22-29 2015;314(24):2663-71. doi:10.1001/jama.2015.17275

      (10) Natera-de Benito D, Berciano J, Garcia A, E MdL, Ortez C, Nascimento A. Acute Flaccid Myelitis With Early, Severe Compound Muscle Action Potential Amplitude Reduction: A 3-Year Follow-up of a Child Patient. J Clin Neuromuscul Dis. Dec 2018;20(2):100-101. doi:10.1097/CND.0000000000000217

      (11) Rosenfeld AB, Warren AL, Racaniello VR. Neurotropism of Enterovirus D68 Isolates Is Independent of Sialic Acid and Is Not a Recently Acquired Phenotype. Mbio. 2019;doi:10.1128/mBio

      (12) Musharrafieh R, Ma C, Zhang J, et al. Validating Enterovirus D68-2A(pro) as an Antiviral Drug Target and the Discovery of Telaprevir as a Potent D68-2A(pro) Inhibitor. J Virol. Jan 23 2019;doi:10.1128/JVI.02221-18

    1. eLife Assessment

      This study offers an important advance by extending an intuitive visualization tool that enables assessment of how dendritic and synaptic currents potentially shape neuronal output. The evidence supporting the tool's capabilities is convincing, with well-documented code, algorithmic innovation, and application to hippocampal pyramidal neurons. The work will be of interest to computational and systems neuroscientists seeking accessible methods to examine dendritic computations.

    2. Reviewer #1 (Public review):

      Summary

      Fogel & Ujfalussy report an extension of a visualization tool that was originally designed to enable an understanding of detailed biophysical neuron models. Named "extended currentscape", this new iteration enables visual assessment of individual currents across a neuron's spatially extended dendritic arbor with simultaneous readout of somatic currents and voltage. The overall aim was to permit a visually intuitive understanding for how a model neuron's inputs determine its output. This goal was worthwhile and the authors achieved it. Demonstrating the utility of extended currentscape, the authors leverage their models to generate interesting and detailed biophysical insights into widely studied neurophysiological phenomena with clear behavioral relevance. Overall, this study provides a valuable and well-characterized biophysical modeling resource to the neuroscience community.

      Strengths

      The authors significantly extended a previously published open-source biophysical modeling tool. Beyond providing important new capabilities, the potential impact of extended currentscape is boosted by its integration with preexisting resources in the field.

      In keeping with the authors' goal to provide an approachable platform with intuitive visualizations of how current flows through neurons, the manuscript is approachable to non-computationalists. In particular, a dedicated glossary and elegant illustrations in Figure 2 boost accessibility for biologists.

      Extended currentscape produces intriguing and detailed predictions spanning neurophysiological phenomena such as local dendritic spikes, complex spike generation, and feature selectivity (hippocampal place fields). By triggering analysis of modeled synaptic inputs on these events, the authors trace their origins from dendritic integration to synaptic input patterns.

      The authors cleverly apply a graph theoretical approach to efficiently model bidirectional current flow throughout a neuron's dendritic arbor. As a result, extended currentscape can run on a standard personal computer.

      The code is well-documented and freely available via GitHub.

      Weaknesses

      While extended currentscape meets its objective of modeling and illustrating the propagation of axial currents throughout a model neuron in great detail, it requires simulation and measurement of synaptic input currents. For this reason, there currently exists a very high technical barrier to conclusively test its intriguing predictions: simultaneous readout of synaptic inputs throughout a neuron's dendritic arbor. Mitigating this weakness, the authors propose a relatively more feasible alternative approach in Discussion: simultaneous voltage imaging of dendrites and their soma while estimating synaptic inputs from the distributions of voltage dynamics along individual dendritic branches.

    3. Reviewer #2 (Public review):

      The electrical activity of neurons and neuronal circuits is dictated by the concerted activity of multiple ionic currents. Because directly investigating these currents experimentally is not possible with current methods, researchers rely on biophysical models to develop hypotheses and intuitions about their dynamics. Models of neural activity produce large amounts of data that are hard to visualize and interpret. The currentscape technique helps visualize the contributions of currents to membrane potential activity, but it is limited to model neurons without spatial properties. The extended currentscape technique overcomes this limitation by tracking the contributions of the different currents from distant locations. This extension allows tracking not only the types of currents that contribute to the activity in a given location, but also visualizing the spatial region where the currents originate. The procedure is first illustrated in a simple setting that allows testing its validity in an intuitive situation where a cell with an apical trunk and two dendritic branches responds to synaptic inputs. The procedure is then applied to study the initiation of complex spike bursts in a model hippocampal place cell.

      The extended currentscape method represents a significant improvement over the original technique, which is already utilized by several research groups. By enabling the analysis of current contributions in spatially extended models, this technique provides a new lens for investigating neuronal and circuit dynamics and will be of use to the modeling community.

      Comments on revisions:

      The changes in Figure 2 greatly improved the manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Fogel & Ujfalussy report an extension of a visualization tool that was originally designed to enable an understanding of detailed biophysical neuron models. Named "extended currentscape", this new iteration enables visual assessment of individual currents across a neuron's spatially extended dendritic arbor with simultaneous readout of somatic currents and voltage. The overall aim was to permit a visually intuitive understanding for how a model neuron's inputs determine its output. This goal was worthwhile and the authors achieved it. Their manuscript makes two additional contributions of note: (1) a clever algorithmic approach to model the axial propagation of ionic currents (recursively traversing acyclic graph subsections) and (2) interesting, albeit not easily testable, insights into important neurophysiological phenomena such as complex spike generation and place field dynamics. Overall, this study provides a valuable and well-characterized biophysical modeling resource to the neuroscience community.

      Strengths:

      The authors significantly extended a previously published open-source biophysical modeling tool. Beyond providing important new capabilities, the potential impact of "extended currentscape" is boosted by its integration with preexisting resources in the field.

      The code is well-documented and freely available via GitHub.

      The author's clever portioning algorithm to relate dendritic/synaptic currents to somatic yielded multiple intriguing observations regarding when and why CA1 pyramidal neurons fire complex spikes versus single action potentials. This topic carries major implications for how the hippocampus represents and stores information about an animal's environment.

      Weaknesses:

      While extended currentscape is clearly a valuable contribution to the neuroscience community, this reviewer would argue that it is framed in a way that oversells its capabilities. The Abstract, Introduction, Results, and Methods all contain phrases implying that extended currentscape infers dendritic/synaptic currents contributing to somatic output., i.e. backwards inference of unknown inputs from a known output. This is not the case; inputs are simulated and then propagated through the model neuron using a clever partitioning algorithm that essentially traverses a biologically undirected graph structure by treating it like a time series of tiny directed graphs. This is an impressive solution, but it does not infer a neuron's input structure.

      We are sorry if our text could be interpreted as if we were inferring unobserved inputs from the known outputs. This was not intentional and we were unaware of the possibility of such interpretation.

      In fact, at the beginning of the Results, we started the description of the extended currentscape method by explicitly stating that we need to measure the input currents: “Our method … requires measuring the membrane and axial currents throughout the dendritic tree of a neuron (in every node of the circuit)”.

      To further clarify that our method starts with measuring the input currents, we made this information explicit already in the abstract (“Our approach relies on the iterative decomposition of the axial current flowing between neighbouring compartments in proportion to the underlying membrane currents measured in the model.”), and in the Introduction (“Even if the membrane currents are known, studying the impact of particular ion channels on the neuronal response in such a dynamical system under in vivo conditions is hindered by two major obstacles”). We also rewrote several parts of the text to remove any phrases that could imply the inference of the inputs (line 568). We believe that after clarifying this at the beginning of the paper, the readers will not misinterpret our descriptions later in the text.

      Because a directed acyclic graph architecture is shown in Figure 2, it is unintuitive that the authors can infer bidirectional current flow, e.g. Figure 3 showing current flowing from basal dendrites and axon to soma, and further towards the apical dendrites. This is explained in Methods, but difficult to parse from Results amidst lots of rather abstract jargon (target, reference, collision, compartment). Figure 2 would have presented an opportunity to clearly illustrate the author's portioning algorithm by (1) rooting it in the exact morphology of one of their multicompartmental model neurons and (2) illustrating that "target" and "reference" have arbitrary morphological meanings; they describe the direction of current flow which is reevaluated at each time step.

      We thank for this comment. We agree that the concepts introduced here to explain our method are rather abstract and could be difficult to understand. To help the reader we followed the instructions of Reviewer and redesigned Fig. 2 to provide a step by step explanation of the extended currentscape method. In particular,

      We used a simpler model where the structure of the graph can be directly related to the morphology of the model.

      We show that the target node can connect multiple subtrees with axial currents flowing in different directions. We explain that in this case the inward and the outward subtrees are pruned and partitioned separately.

      We provide a glossary in Table 1 to ensure that the readers can follow our description and do not get lost amidst lots of rather abstract jargon.

      We also clarified that although the target compartment is chosen arbitrarily by the user, it remains the same for all time points throughout the analysis.

      Analyses in Figure 7, C and D, are insightfully devised and illuminating. However, they could use some reconciliation with Figure 5 regarding initiation of individual APs versus CSBs within place fields.

      We thank the reviewer for the positive comments and also for pointing out the potential source of misunderstanding. We slightly changed the text at Fig 5 to emphasize that this is a single example trial, and we added the following sentence to the paragraph describing Fig 7CD: “Consequently, the somatic current dynamics before the iAP and the CSB presented in Fig 5Cc-Dd can be regarded as illustrative samples from a broad distribution, but the differences observed between them are not representative.}”

      The intriguing observations generated by extended currentscape also point to its main weakness, which the authors openly acknowledge: as of now, no experimental methods exist to conclusively tests its predictions.

      We agree with the Reviewer that not being able to apply our extended currentscape method to reveal the current types driving real neurons recorded in vivo is currently a weakness of our approach. However, we would like to emphasize that it may be feasible to use it to estimate the spatial distribution of the membrane currents driving the cell based on in vivo voltage imaging data, as we briefly outline in the discussion.

      Reviewer #2 (Public review):

      Summary

      The electrical activity of neurons and neuronal circuits is dictated by the concerted activity of multiple ionic currents. Because directly investigating these currents experimentally isn't possible with current methods, researchers rely on biophysical models to develop hypotheses and intuitions about their dynamics. Models of neural activity produce large amounts of data that is hard to visualize and interpret. The currentscape technique helps visualize the contributions of currents to membrane potential activity, but it's limited to model neurons without spatial properties. The extended currentscape technique overcomes this limitation by tracking the contributions of the different currents from distant locations. This extension allows tracking not only the types of currents that contribute to the activity in a given location, but also visualizing the spatial region where the currents originate. The method is applied to study the initiation of complex spike bursts in a model hippocampal place cell.

      Strengths.
>

      The visualization method introduced in this work represents a significant improvement over the original currentscape technique. The extended currentscape method enables investigation of the contributions of currents in spatially extended models of neurons and circuits. 
>

      Weaknesses.

      The case study is interesting and highlights the usefulness of the visualization method. A simpler case study may have been sufficient to exemplify the method, while also allowing readers to compare the visualizations against their own intuitions of how currents should flow in a simpler setting. 
>

      We thank the reviewer for this comment. In fact we had been also considering to include a simpler case study to illustrate the extended currentscape method in the original submission. In accordance with the comments from Reviewer 1, we now use a simple model to introduce the concepts in Figure 2 and provide a few examples where the reader can compare the results with their own intuition in simpler cases.

      Recommendations for the authors:

      Reviewing Editor Comments:

      (1) Model complexity vs. intuition/validation. The case study relies on a very complex CA1 model, making it difficult to build intuition about current flow and to validate the visualization. Inclusion of a simpler benchmark (e.g., soma plus a dendrite with two branches, fewer compartments) is recommended to demonstrate how the extended currentscape behaves in a more tractable setting.

      Inspired by the suggestions of the Reviewers, we modified Figure 2 and now first use a simple model with a soma and a dendrite with two branches to introduce the concepts of our analysis. We start with a few examples where the reader can compare the results with their own intuition in simpler cases.

      (2) Rationale and citations for input structure. The in vivo-like input design (untuned inhibition; 12 co-tuned excitatory clusters with large conductances; the goal of generating place fields) would benefit from a more explicit rationale and substantially more literature support. Alternative plausible scenarios (e.g., distributed co-tuned inputs and homosynaptic plasticity) should be articulated, and choices situated within the experimental literature on CA1 excitation/inhibition, including tuning and anti-tuning results.

      We extended the paragraph in the Results describing the input structure and added the most important references there. We added further references to the Methods section where we argue that “Reliable place cell tuning can be achieved by functional synaptic clustering without increased excitatory drive in the place field (Ujfalussy and Makara 2020) or via strong excitatory drive without input clustering (Grienberger et al., 2017, Ujfalussy and Makara, 2020). However, experimental data indicates that both of these mechanisms are present and contribute to the activity of place cells (Adoff et al., 2021,Tasciotti et al., 2025)” and “although interneurons can display spatial tuning, they typically have a broad tuning with low selectivity (Ego-Stengel et al., 2007, Dupret et al., 2013, Geiller et al., 2020). A weak disinhibition within the place field can also contribute to the selective firing of place cells (Geiller et al., 2022, Valero et al., 2022), this was not necessary for place cell activity in novel environments (Geiller et al., 2022) and the overall inhibitory input to place cells is largely untuned (Grienberger et al., 2017).”

      (3) Scope of PCA-based claims. The interpretations derived from the PCA analysis appear broader than warranted, given subcellular heterogeneity and the dominance of somatic action potential variance. These claims should be tempered with more explicit statements about what PCA can and cannot resolve in this context.

      We thank the Reviewer for the opportunity and encouragement to clarify this part of the text. We agree with the Editor and the Reviewers that the results of the PCA analysis can not be used to support claims regarding the presence or the absence of independent dendritic events. In fact, we aimed to use it as an illustration that global activity tends to dominate PCA analysis even when the “neuron is mainly driven by strong, functionally clustered synaptic inputs to a few dendritic branches”. We acknowledge that we did not formulate this point clearly in the original submission. Therefore we substantially rewrote this part of the Results and performed additional analysis to clarify that there is a substantial amount of soma-independent dendritic activity in our model that remains invisible for a PCA based analysis.

      Reviewer #1 (Recommendations for the authors):

      Major concerns:

      (1) Depolarization-inactivated K+ may be an important consideration to model burst-firing.

      Our current model includes 2 kinds of transient K+ channels that show inactivation after depolarization: a proximal and a distal type, as the original model in Jarsky et al., 2005. We now made this explicit in the main text (line 178).

      (2) Description of the in vivo-like model's excitatory and inhibitory input structure needs many more citations of biological studies to communicate rationale for the author's decisions, e.g. untuned inhibitory neurons, organization of a subset of excitatory inputs into 12 function synaptic clusters with co-tuned presynaptic neurons and outsized synaptic conductances. The goal is clearly to create CA1 pyramidal neurons with place fields, which would be helpful to state upfront. But additionally, (a) place fields could arise from homosynaptic potentiation of distributed co-tuned excitatory inputs (e.g., Bittner, et al. 2017 study describing BTSP made no assumptions) and (b) CA1 inhibitory interneurons can be spatially tuned (Ego-Stengel & Wilson, 2006; Wilent & Nitz, 2007; Geiller, et al. 2020) and even anti-tuned (Geiller, et al. 2021).

      We thank the Reviewer for pointing out the lack of appropriate references in this section. We made the following changes in the manuscript:

      (1) Stated explicitly that the goal was to create place cell activity.

      (2) Added references to the main text to justify our choices of the inputs (lines 234-241).

      (3) We included a longer rationale for the choice of synaptic clusters and the lack of inhibitory (anti-)tuning in the Methods section, describing the neuron model. In brief, Adoff et al., 2021 reported more clustering of excitatory inputs within the place field. In our model, the degree of clustering is somewhat larger than the clusters reported. Although inhibitory neurons can be tuned, their tuning is much weaker than that of place cells and seems to play only a minor role in the generation of place fields (Grienberger et al., 2017). The presence of inhibitory anti-tuning is controversial: although Geiller et al., 2021 reported weak (~10%) anti-tuning, they did not find it in novel environments, indicating that it is not needed for spatially selective activity (lines 628-646).

      (3) Interpretation of principal component-based analyses shown in Figure 4 could be toned down. As written in section "CSBs in the CA1 pyramidal neuron", it sounds like CA1 pyramidal neuron dendrites display minimal autonomous activity. However, PCA does not seem well-suited to address the heterogeneity of subcellular voltage dynamics over physiologically relevant timescales. Somatic action potentials, and their backpropagation/modulation of dendritic voltage, would of course explain a very large fraction of variance. However, if local dendritic events summate over fine timescales to initiate somatic firing, it is hard to imagine this important nuance being detected. On the other hand, it is hard to imagine single dendritic branches driving robust somatic firing except in the relatively extreme situation in which large numbers of synapses synchronously drive the same branch to initiate a local Ca2+ spike (Figure 3, A-C).

      We agree with the reviewer that PCA can not reveal the potential dendritic origin of somatic APs, and thus is not suitable to assess the role of local dendritic spikes in shaping the output of the cell. We wanted to highlight here that even in cells with excitable dendrites driven by strong, local input clusters, exhibiting frequent local dendritic spikes, the dendritic membrane potential dynamics will be dominated by global fluctuations with surprisingly little sign of local dynamics in the PCA components. As the reviewer also pointed out, this may not be surprising as local events either remain spatially restricted and thus contribute little to the overall variability of the dendritic Vm or they initiate somatic APs and will thus be counted as global events.

      To demonstrate the high propensity of local dendritic events, we analysed local Vm peaks in dendritic branches and found that ~7.6% of the peaks were not coupled to somatic APs.

      Although this number could seem low, we emphasize that most of the 92.4% of the dendritic peaks coupled to APs potentially reflect the backpropagation of the same somatic events to multiple dendritic sites. To confirm this, we performed an additional analysis measuring the spatial extent (number of branches involved) of the individual dendritic events. We found that 90% of the events remained local, restricted to a few dendritic branches, while 10% of the events were global, associated with BAPs and involving the majority of the dendritic tree. Interestingly, these global events dominate the PCA analysis and are responsible for >90% of the dendritic Vm peaks. These results are included in a new panel in Figure 4H.

      We conclude that, “this way, although only 10% of the dendritic Vm events were associated with bAPs, they were ~60-times larger than local events and they dominated the PCA analysis even in the presence of local regenerative dendritic events driven by strong, functionally clustered synaptic inputs.” We believe that this model and analysis could serve as an important benchmark for future experimental studies investigating the structure of membrane potential correlations in in vivo voltage imaging data (Lee et al., 2026).

      (4) One suggestion would be to display more data as shown in Figure 4F, with a longer X axis to clarify the temporal relationship between local dendritic spikes and the first somatic action potential.

      We added a few more examples including the CSBs presented in Fig8G-I as a new supplementary Figure S4. We also slightly extended the x-axis on this supplementary figure as the reviewer requested.

      If the models indicate that passively filtered EPSPs drive most somatic action potentials, as seems to be the case in Figure 5, then this would also be helpful to show as in Figure 4F.

      In Fig 5 we showed two examples of isolated APs. The first AP was indeed driven by passively filtered EPSPs. The second one was preceded and possibly caused by a dendritic spike, as highlighted by the black arrowhead labelled c in Fig. 5Cc. We further analysed the currents driving iAPs in Fig 7B and C, and found that there is considerable heterogeneity in the magnitude of the dendritic Na currents driving the soma before action potentials. Figure 8 and Figure S3 (now Fig. S5) show further examples for iAPs driven either by passively filtered EPSPs or dendritic spikes. We also included these examples in the new supplementary Figure S4.

      (5) Another suggestion would be to use one-hot vectors containing onset times of different event types, since this would divorce the amplitude/duration of events from their influence over total variance.

      In this paper our goal was to illustrate the ability of the extended currentscape method to reveal the origin of the axial currents driving neuronal activity. In Fig. 4, our primary intention was to characterize the membrane potential response of the model in a way that is easily comparable with experimental data. To further quantify the frequency of local events, we added a new panel showing the spatial extent of dendritic events (Fig. 4H). To make our model more comparable with recent publications, we also calculated two additional metrics used to evaluate the relationship between somatic and dendritic activity (Fig 4I-J). We hope that these additional analyses help the reader to characterize the prevalence and impact of local dendritic events on somatic activity.

      (6) From section "Input conditions for complex spike burst generation", paragraph 2: "Note that synapse density, the ion channel mechanisms and the input statistics are identical for tuft and oblique branches,...". The authors should justify this parameterization given the numerous known differences between tuft and oblique branches in both of these regards and acknowledge accompanying interpretational caveats.

      We agree with the reviewer that experimental data demonstrated several significant differences between the tuft and oblique branches regarding both the inputs they receive and the way they process it. However, in the present paper we chose not to include these differences for several reasons:

      Here we aimed to focus on the abilities of the dendritic currentscape methods and use CSBs as a case study to illustrate how dendritic currentscape can reveal the membrane currents underlying complex neuronal responses.

      Currently there is no CA1PN model that would be able to reproduce all data regarding tuft and oblique integration and would be able to fire calcium spikes. We only wanted to make minimal modifications to the existing CA1PN model to make it capable of generating Ca-spikes and CSBs. We are currently working towards developing and extensively testing a new model, examining the role of these regional differences in CSB generation.

      Although there is information regarding input statistics and dendritic physiology in the literature, many of the relevant parameters are underconstrained. We wanted to avoid overfitting by keeping the model simple.

      By maintaining identical inputs and ion channel distribution we can distinctly highlight the special role of tuft morphology in CSB generation. Altering the inputs or the ion channel density for the tuft would make the interpretation more ambiguous, and elucidating the specific role of the different factors in CSB generation is the subject of future investigations.

      In sum, although we acknowledge that our model does not reflect the full complexity of CA1 PNs and its inputs, we regard this simplicity as a useful feature of the model. We added a section discussing potential future extensions of the model and highlighting interpretational caveats in the discussion (lines 482-490).

      (7) Given the debate in the field regarding the level of functional autonomy present in dendrites, the authors' finding that dendritic voltage largely tracks that of the soma (though see concern above re: PCA), and their access to specific currents, the authors have an important opportunity investigate the divergence between Ca2+ and voltage sensors as reporters of dendritic activity.

      For instance, why have some studies reported relatively common isolated dendritic Ca2+ transients in CA1 pyramidal neurons while other studies, including voltage imaging studies, have reported the opposite?

      We thank the Reviewer for the opportunity to highlight a few important points regarding functional autonomy of dendrites based on the analysis of our model. We would like to first note that only parallel calcium and voltage imaging studies will be able to ultimately resolve this debate. Nevertheless, below we briefly summarize our take on this issue.

      (1) In general, most Ca2+ imaging studies found that soma-independent dendritic events are rare. "Isolated dendritic transients (no coincident somatic event; see fig. S6, C and D, for example) were overall rare. Isolated apical dendritic Ca2+ transients, which have not previously been reported in CA1PNs, were larger and more frequent than those observed in basal dendrites." (O’Hare et al., 2022). "Activity in the ... basal dendrites ... along the track but outside of the place field was rarely observed” (Sheffield and Dombeck, 2014) and “overall, isolated dendritic transients were similar in size but occurred far less frequently than coincident dendrite-soma transients”, or “data indicate that spatially reliable dendritic firing was almost exclusively yoked to somatic tuning, likely reflecting strong backpropagation of burst firing during traversals of the somatic PF” (Rolotti et al., 2022). Consistent with this observation, a dendritic Vm peak chosen randomly from any branch has ~93% probability to be related to a bAP in our model. However, it is also true that ~90% of events in the model are local events, simply because isolated events involve ~60-times fewer branches (1.8 on average) than events associated with bAPs (114 branches) in the model. If the spatial extent of typical local events are also similarly small in real neurons as in the model, then even rare occurrences of dendritic events may reveal substantial dendritic independence. We added a section quantifying the functional autonomy of dendrites in the model in the main text, around Fig 4H.

      (2) Ca2+ indicators are slower and nonlinear and thus they are somewhat unreliable reporters of dendritic voltage events, especially in distal dendrites (Wu et al., 2026; Gonzalez et al., 2026). To illustrate this, we calculated three metrics in our model that were also reported in recent dendritic Ca2+ imaging studies (Rolotti et al., 2022, Sheffield et al., 2014, 2017). First, we calculated the fraction of bAPs detected in a branch (called dendrite-soma coupling in Rolotti et al., 2022, see their Fig. 2C) as a function of the distance of the branch from the soma (our new Fig. 4I). In the Ca2+ imaging data, this was essentially constant ~30% between distances 5-100 µm from the soma. In contrast, the fraction of bAPs detected in the model was 100% in this range as bAPs propagation failures did not occur before µ100 µm. This is also consistent with a recent voltage imaging study showing that even low-transmission bAPs reliably propagate to the proximal dendrites (Lee et al., 2026, Fig 3G). The low and distance independent dendrite-soma coupling reported by Rolotti et al. can only be reconciled with the known biophysics of neurons if the recorded calcium signal is unreliable reporter of the underlying voltage. Indeed, it has been reported that Ca signals associated with bAPs can be absent in some dendritic branches (Landau et al., 2022) or that local, nonlinear Ca signals can appear in the absence of local regenerative voltage response (Weber et al., 2016, Tran-Van-Minh et al., 2016) and that the Ca signals are highly variable across cells (Eltes et al., 2019).

      Second, we calculated the fraction of local events as a function of the distance from the soma (our Fig 4J; see also Fig. 2F in Rolotti et al.). When averaged across all branches, this was somewhat lower in the model (18%) than in the data (38%) which, again, could be explained by the low reliability of detecting global voltage events in all compartments based on the calcium signal.

      Third, the range of branch-spike-prevalence (BSP) values in our model (0.5-0.9; Fig. 4H) seem consistent with that reported (0.4-0.8) at first (Fig 4C of Sheffield et al., 2014; Fig 2 of Sheffield et al., 2017). However, we note that there are several important differences: for technical reasons, Sheffield et al. reported BSP for place field traversals and not for individual events, and they measured Ca2+ dynamics in the basal dendrites. Since bAPs are almost always present in all basal dendrites in the model (basal BSP > 0.9 for all events with somatic spikes) and place field traversals were always accompanied by somatic APs, BSP for basal dendrites would be nearly 1 in the model. Thus, the lower BSP values reported by Sheffield et al. could be explained by the limited reliability of the Ca2+ indicators in reporting regenerative voltage events in neuronal processes.

      We briefly discussed these differences in the Discussion (lines 474-478).

      (3) Finally, to our knowledge, there are 3 relevant in vivo voltage imaging studies in CA1 PNs. Liao et al., 2024 found that in induced place cells the tuning of dendritic events (presumably local or back-propagating Na-spike) was similar to the somatic tuning, which is consistent with our model where dendritic activity and tuning is dominated by bAPs. However, they did not acquire simultaneous signals from the dendrites and the soma so they could not study the independence of the dendritic events. Lee et al. (2026) found that only 10% of the dendritic events are not associated with a somatic spike, which is lower than the number of independent events in the model. However, the events they found were generated in the distal apical trunk (their Fig 3D) and they could not record from the most distal branches where most of the isolated events were generated in our model. Gonzalez et al., 2026 measured voltage and calcium in selected locations within the dendritic tree, and could not reliably estimate the fraction of isolated events throughout the cell. (Gonzalez et al, 2024 measured voltage only in single spines and soma, but did not quantify independent dendritic events; Wong-Campos et al., 2023 measured dendritic integration and bAPs in L23 branches; Wu et al. 2026 recorded in CA2 neurons.)

      We added a paragraph in the discussion comparing the level of functional autonomy present in the model dendrites to recent Ca- and voltage-imaging studies (lines 467-474).

      Minor concerns:

      (1) Abstract:

      There is a need to explain what currentscape is - even at the cost of not invoking its name. To a reader not familiar with currentscape, the abstract is extremely difficult to understand.

      We reworded the title and the abstract to make them more accessible to readers not familiar with the term currentscape.

      (2) "Currentscape analysis of place field dynamics" section:

      It would be helpful to emphasize upfront that dendritic determinants of individual somatic APs versus CSBs will be discussed separately. Since somatic action potentials are discussed before CSBs, I found this section initially confusing as I attributed those findings to CSBs until reading the next paragraph.

      We added a sentence to clarify that we analysed subthreshold responses, APs and CSBs separately.

      (3) Bottom of p2 discussing mixed literature on what drives CSBs in CA1 PCs:

      Overall accurate and useful point, but an important nuance is glossed over which misportrays state of field. References ex vivo studies that fail to drive CSBs with somatic current injection and in vivo study successfully doing so. These aren't really conflicting results. In vivo current injection co-occurs with spontaneous synaptic input, which is high in CA1 and results in PCs that are significantly depolarized at rest relative to those in acute slices. Bittner 2017 ex vivo results are consistent with this: CSBs driven by Cs+-based internal solution to block K+ channels (partially, using strategy of purposefully high series resistance). Similar situation in vivo given that A-type K+ channels are inactivated by depol. Resulting increase in input resistance lowers input threshold to CSB. This is clarified in Results, p.5: "Under in vivo-like synaptic input conditions (see below and Methods), dendritic Ca2+-spikes could also be evoked by somatic current injection (Fig. S1E), as in Bittner et al. (2015).", which makes p. 2 feel especially awkward.

      We agree with the Reviewer that these are not necessarily conflicting results. We rephrased this section, emphasizing that the role of the different input pathways in the initiation of CSBs are not clear.

      (4) Abbreviating "pyramidal neuron" with PC is confusing:

      PC often means place cell. The authors could change this, such that PC refers to "pyramidal cell", or else use PN as an abbreviation. It is important to avoid confusion, especially because place cell dynamics feature prominently in the manuscript.

      Thanks for the suggestion. We replaced PC with PN throughout the manuscript.

      (5) Only apical dendritic parameters are described in section 2 of Results, but the full morphology is shown in Figure 3B with basal currents shown in panels C and F. Some clarification is needed - either what currents were considered for basal dendrites and why, or else why basal dendritic current parameters were not considered for this simulation using apical dendritic current injection but nonetheless examining basal dendritic currents.

      We clarified in the text that the original model contained a standard set of Na and K channels (line 178).

      (6) Clarify "i" and "s" in the Figure 3C legend - "intrinsic" and "synaptic" white letterings are small/hard to see in the bottom subpanels.

      We now spell out intrinsic and synaptic in the Figure and increased the contrast of the letterings.

      (7) Regarding the computational benefit of recursively decomposing axial currents along an adaptively truncated acyclic graph, it would be useful to (a) include a supplemental figure benchmarking this approach to standard approaches to quantify the described gain in computational efficiency and (b) describe computing hardware in the Methods.

      We included an estimated benefit of the pruning process (line 758) as well as the utilised computing hardware and the simulation times in the Methods (line 776).

      Reviewer #2 (Recommendations for the authors):

      The manuscript is in great shape, it is well organized, and the figures are gorgeous. I believe that the extended currentscape is a great extension of the original currentscape method. In particular, the possibility of partitioning currents by the spatial location of their sources is a great addition. 
>

      Recommendations:

      (1) The method is applied in the context of an interesting case study that highlights its usefulness. However, the model in the study is so complex that it is difficult to develop an intuition of how currents should be flowing, and this makes it hard to intuitively validate the visualization method. I think that applying the extended currentscape in a simpler model - maybe a soma with a dendrite with two branches, fewer compartments - would be instrumental in developing this intuition. 
>

      We now first use a simple model with a soma and a dendrite with two branches to introduce the concepts in Figure 2 and provide a few examples where the reader can compare the results with their own intuition in simpler cases. We also added the currentscape analysis of a standard, two-compartmental model from Pinsky and Rinzel, 1994 as Supplementary Figure 1.

      (2) I found a number of typos and minor stylistic details you may want to fix in a revised version of the manuscript.

      (a) Abstractine, line 12. I believe the word "recursive" is a bit technical at this point. It's meaning in this context becomes clear after ones goes through the details of the algorithm (Figure 2). 
>

      We replaced the word “recursive” with “iterative”. We hope that this will make the abstract clearer for the readers. In fact, we realized that the word iterative is a better description of the algorithm, so we replaced the “recursive” with “iterative” consistently throughout the manuscript.

      (b) Figure 1, caption."Since we included the capacitive current, the magnitude of the inward and the outward currents is identical (Kirchhoff's law)."This sentence can be confusing. If the inward and outward currents are the same, the membrane potential doesn't change. I believe that you are including the capacitive current in the inward (or outward) currents.

      Indeed, we included the capacitive current in the inward or outward currents. We changed the text to clarify this.

      (c) Lines 92-93. I do not fully understand this sentence. Are you making an assumption? What does 'continuos flow of axial current' mean?
>

      By ‘continuous flow of axial current’ we meant a spatially continuous stream of axial currents flowing from the reference to the target. To clarify this, we added the explanatory sentence: “i.e., if the axial current is not blocked or reversed between the reference and the target.”

      (d) Equation (1.) Why summing axial currents over j? Is this for the case of a branching point?

      The compartment could be 1) part of a continuous segment of dendritic branch, where axial currents can flow from the distal and the proximal direction (sum over 2); 2) It can be a branch point with 3 axial currents; 3) or it can be a leaf compartment with only one axial current, in which case the summation is not relevant. We clarified this in the text.

      (e) Figure 2, caption. Typo. "When the axial currents flows…" Should it be 'current'? - Figure 3, caption. Typo in (C) "Extended currentscape" 
>

      Corrected.

      (f) Figure 4. I cannot see the grey lines or the dotted lines mentioned in the caption. 
>

      We added an arrow highlighting the gray and the dotted lines in the figure.

      (g) Figure 5, caption. "Red boxes highlight regions analyzed in panels B-D."Because this is a spatially extended model, region may be confused with spatial location, but you are highlighting a temporal interval.
>

      We rephrased the caption referring to temporal intervals now.

      (h) Line 341. This is a numerical experiment, correct? 
>

      We clarified in the text and added that it was indeed a simulation experiment.

      (i) Line 349. Should it be 'distributions'? 
>

      Corrected

      (j) Line 422. Typo. Missing space 'in vivousing'
>

      Corrected

      (k) Line 537. "Preprocessing membrane…" I found this entire subsection a bit confusing and hard to read.

      We rephrased this subsection to clarify it and facilitate reading.

    1. eLife Assessment

      This important study reports characterisation of hepatocyte molecular pathways affected by a glycyrrhizin derivative in both in vivo and in vitro mouse models of alcohol-associated liver disease. The authors show convincing evidence indicating that IPP delta isomerase 1 (Idi1) is an intermediate in these pharmacological effects, via the binding of the glycyrrhizin derivative to an upstream regulator of Idi1, HSD11B1, although significant questions remain about some of the experiments and analyses provided. The findings would be of interest to immunologists and pharmacologists interested in liver inflammation and its amelioration.

    2. Reviewer #1 (Public review):

      Summary:

      In this article by Xiao et al. the authors aimed to identify the precise targets by which magnesium isoglycyrrhizinate (MgIG) functions to improve liver injury in response to ethanol treatment. The authors found through a series of in-vivo and molecular approaches that MgIG treatment attenuates alcohol-induced liver injury through a potential SREBP2-IdI1 axis. The revised manuscript adds to a previous set of literature showing MgIG improves liver function across a variety of etiologies, and also provides mechanistic insight into its mechanism of action. All major weaknesses were addressed in the revised submission.

      Strengths:

      (1) The authors use a combination of approaches from both in-vivo mouse models to in-vitro approaches with AML12 hepatocytes to support the notion that MgIG does improve liver function in response to ethanol treatment.

      (2) The authors use both knockdown and overexpression approaches, in-vivo and in-vitro, to support most of the claims provided.

      (3) Identification of HSD11B1 as the protein target of MgIG, as well as confirmation of direct protein-protein interactions between HSD11B1/SREBP2/IDI1 is novel.

      Weaknesses:

      The authors addressed all my concerns.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigated magnesium isoglycyrrhizinate (MgIG)'s hepatoprotective actions in chronic-binge alcohol-associated liver disease (ALD) mouse models and ethanol/palmitic acid-challenged AML-12 hepatocytes. They found that MgIG markedly attenuated alcohol-induced liver injury, evidenced by ameliorated histological damage, reduced hepatic steatosis, and normalized liver-to-body weight ratios. RNA sequencing identified isopentenyl diphosphate delta isomerase 1 (IDI1) as a key downstream effector. Hepatocyte-specific genetic manipulations confirmed that MgIG modulates the SREBP2-IDI1 axis. The mechanistic studies suggested that MgIG could directly target HSD11B1 and modulate the HSD11B1-SREBP2-IDI1 axis to attenuate ALD. This manuscript is of interest to the research field of ALD.

      Strengths:

      The authors have performed both in vivo and in vitro studies to demonstrate the action of magnesium isoglycyrrhizinate on hepatocytes and an animal model of alcohol-associated liver disease.

      Original comment (1):

      In Supplemental Figure 1A, all the treatment arms (A-control, MgIG-25 mg/kg, MgIG-50 mg/kg) showed body weight loss compared to the untreated controls. However, Figure 1E showed body weight gain in the treatment arms (A-control and MgIG-25 mg/kg), why? In Supplemental Figure 1A, the mice with MgIG (25 mg/kg) showed the lowest body weight, compared to either A-control or MgIG (50 mg/kg) treatment. Can the authors explain why MgIG (25 mg/kg) causes bodyweight loss more than MgIG (50 mg/kg)? What about the other parameters (ALT, ALS, NAS, etc.) for the mice with MgIG (50 mg/kg)?

      Author's response:

      We agree that this observation does not strictly follow a dose-dependent pattern. In vivo responses to pharmacological interventions, particularly in metabolic and liver disease models, are not always linear. The relatively greater body weight reduction observed in the 25 mg/kg group may be influenced by inter-individual variability, differences in metabolic adaptation, or sample size-related variation. Importantly, these differences in body weight were not statistically significant. Therefore, we selected the 50 mg/kg dose for subsequent animal experiments, as it demonstrated more consistent and stable improvements across multiple parameters, including body weight, ALT, AST, TG, and TC.

      New comment:

      My first question: All the treatment arms (A-control, MgIG-25 mg/kg, MgIG-50 mg/kg) showed significant body weight loss compared to the untreated controls (Supplemental Figure 1A), but the body weight significantly increased in the treatment arms (A-control and MgIG-50 mg/kg) compared to the untreated controls (Figure 1E). Why?

      My second question: Mice with MgIG (25 mg/kg) showed the lowest body weight, compared to either A-control or MgIG (50 mg/kg) treatment. According to the authors' explanation, the MgIG (25 mg/kg) caused bodyweight loss are attributed to inter-individual variability, differences in metabolic adaptation, or sample size-related variation. Did these differences happen in MgIG (25 mg/kg) only? or in all other groups? The mouse group assignment should be randomized; however, a large variation in bodyweight was seen in MgIG (25 mg/kg) group. It is not convincing for the author to select MgIG (50 mg/kg) group for subsequent animal experiments, because of a large variation in MgIG (25 mg/kg) group, and because that MgIG (50 mg/kg) group demonstrated more consistent and stable improvements across multiple parameters. The author should reanalyze and compare all the raw data between MgIG (50 mg/kg) group and MgIG (25 mg/kg) group, and address the issues being pointed out and justify rationale for the animal group assignment.

      Original comment (2):

      IL-6 is a key pro-inflammatory cytokine significantly involved in ALD, acting as a marker of ALD severity. Can the authors explain why MgIG 1.0 mg/ml shows higher IL-6 gene expression than MgIG (0.1-0.5 mg/ml)? Same question for the mRNA levels of lipid metabolic enzymes Acc1 and Scd1.

      Author's response:

      Thank you for this important comment. We agree that IL-6, as well as lipid metabolism-related genes such as Acc1 and Scd1, are key indicators in ALD. The relatively higher expression observed at 1.0 mg/mL MgIG compared to lower concentrations (0.1-0.5 mg/mL) may be related to experimental constraints associated with the MgIG formulation used in this study. Specifically, to maintain consistency with our in vivo experiments, we used a clinically available liquid formulation of MgIG (5 mg/mL), which is approved for intravenous administration in China. Due to its relatively low stock concentration, achieving higher working concentrations (e.g., 1.0 mg/mL) in vitro required a larger volume of the MgIG solution, thereby proportionally reducing the volume of culture medium. This reduction in effective culture conditions may adversely affect hepatocyte viability and function. Supporting this, our CCK-8 and LDH assays indicated that higher MgIG concentrations were associated with subtle cytotoxicity or impaired cell status.

      New comment:

      The author's response did not answer my question. If the authors believe it could be experimental constraints associated with the MgIG formulation, then it is questionable for this MgIG formulation used in all other associated experiments. The experiments, at least those the MgIG formulation associated experiments, need to be repeated.

      Original comment (3):

      For the qPCR results of Hsd11b1 knockdown (siRNA) and Hsd11b1 overexpression (plasmid) in AML-12 cells (Figure 5B), what is the description for the gene expression level (Y axis)? Fold changes versus GAPDH? Hsd11b1 overexpression showed non-efficiency (20-23, units on Y axis), even lower than the Hsd11b1 knockdown (above 50, units on Y axis). The authors need to explain this. For the plasmid-based Hsd11b1 overexpression, why does the scramble control inhibit Hsd11b1 gene expression (less than 2, units on the Y axis)? Again, this needs to be explained.

      Author's response:

      Thank you for this important comment, and we apologize for the lack of clarity in the Y-axis labeling, which may have led to misunderstanding.

      As shown in Figures 5A and 5B, we have revised the Y-axis description to clearly indicate that gene expression levels are presented as relative expression normalized to GAPDH (fold change relative to the control group).

      New comment:

      The author explained the relative expression was normalized to GAPDH (fold change), but they did not answer my question. My question is for Figure 5B. in Figure 5B (left, Hsd11b1-KD), scramble control showed over 100 (unit), however, in Figure 5B (right, Hsd11b1-OE), scramble control showed only 0.5-1 (unit). The data seemed that authors used same scramble control for both KD and OE? If yes, they should provide more details of the KD and OE experiments and explain why this happened. If they used plasmid for OE control, they also need to clarify it. In addition, qPCR is not a good assay to show the success of KD or OE, Western blotting should be done as convincing data to show the success of KD or OE.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) A few of the claims made are not supported by the references provided. For instance, line 76 states MgIG has hepatoprotective properties and improved liver function, but the reference provided is in the context of myocardial fibrosis.

      Thank you for the correction. We have made the revision on page 4, line74.

      (2) MgIG is clinically used for the treatment of liver inflammatory disease in China and Japan. In the first line of the abstract, the authors noted that MgIG is clinically approved for ALD. In which countries is MgIG approved for clinical utility in this space?

      Thank you for this important comment. MgIG has been recommended for the treatment of alcoholic liver disease (ALD) in Chinese clinical guidelines (2018). We have clarified this point in the manuscript (Page 5, Line 79-80).

      (3) Serum TGs are not an indicator of liver function. Alterations in serum TGs can occur despite changes in liver function.

      Thank you for this important comment. We fully agree that serum triglycerides (TGs) are not a direct indicator of liver function. ALT and AST are more appropriate markers for hepatocellular injury, whereas TG and TC primarily reflect systemic and hepatic lipid metabolism status. We have made the necessary revisions as suggested on page 12, lines 285-288

      (4) There are discrepancies in the results section and the figure legends. For example, line 302 states Idil is upregulated in alcohol fed mice relative to the control group. The figure legend states that the comparison for Figure 2A is that of ALD+MgIG and ALD only.

      We thank the reviewer for the valuable suggestion. Accordingly, we have revised the legend for Figures 2A and 2B as suggested.

      (5) Oil Red O staining provided does not appear to be consistent with the quantification in Figure 1D. ORO is nonspecific and can be highly subjective. The representative image in Figure 1C appears to have a much greater than 30% ORO (+) area.

      Thank you for this insightful comment. We acknowledge that Oil Red O (ORO) staining can be influenced by background signal and may appear subjective in representative images. In our quantification, only well-defined lipid droplets with strong positive staining were included, while diffuse background staining (e.g., light reddish hue) was excluded. This may explain the apparent discrepancy between the representative image and the quantified ORO-positive area. To further strengthen the reliability of our findings, we additionally measured hepatic triglyceride (TG) and total cholesterol (TC) levels. These biochemical assays yielded results consistent with the ORO quantification, thereby supporting our conclusion regarding lipid accumulation. Please refer to page12, lines 285-288. As requested, we have added the required information to Figures 1G.

      (6) The connection between Idil expression in response to EtOH/PA treatment in AML12 cells with viability and apoptosis isn't entirely clear. MgIG treatment completely reduces Idi1 expression in response to EtOH/PA, but only moderate changes, at best, are observed in viability and apoptosis. This suggests the primary mechanism related to MgIG treatment may not be via Idi1.

      Thank you very much. We agree that although MgIG almost completely reverses Idi1 expression induced by EtOH/PA, the improvements in cell viability and apoptosis are only moderate, suggesting a potential discrepancy between these observations. This may indicate that Idi1 functions as a permissive factor, rather than the sole mediator, in this pathological process. In other words, while modulation of Idi1 contributes to the protective effects of MgIG, additional pathways are likely involved in mediating its overall impact on hepatocyte viability and apoptosis. We have clarified this point in the revised manuscript (Page 12, Lines 325–335), stating that MgIG exerts its protective effects against ethanol-induced hepatocellular injury, at least in part, through the regulation of Idi1.

      (7) The nile red stained images also do not appear representative with its quantification. Several claims about more or less lipid accumulation across these studies are not supported by clear differences in nile red.

      Thanks a lot. We acknowledge that Nile Red staining can be influenced by imaging conditions and may appear less distinct in representative images, which could affect visual interpretation. To minimize subjectivity, all images were analyzed using a consistent and standardized thresholding method across groups. We agree that the visual differences in Nile Red staining alone may not be sufficiently pronounced to fully support the quantitative conclusions. Therefore, to strengthen the reliability of our findings, we have included additional biochemical measurements, including serum TG and TC levels, as well as hepatic TG and TC content. These independent assays consistently support the observed changes in lipid accumulation. The corresponding data have been added to the revised manuscript (page 12, lines 285-288)

      (8) The authors make a comment that Hsd11b1 expression is quite low in AML12 cells. So why did the authors choose to knockdown Hsd11b1 in this model?

      Thank you for this important comment. Although the basal expression of Hsd11b1 in untreated AML-12 cells is relatively low, we observed that it is inducible upon EtOH/PA stimulation, indicating its functional relevance under stress conditions. Therefore, knockdown experiments were performed to assess its contribution to EtOH/PA-induced hepatocellular injury. We have clarified this point in the revised manuscript (page 15, lines 281-382).

      (9) Line 380 - the claim that MGIG weakens the interaction between HSD11b1 and SREBP2 cannot be made solely based on one Western blot.

      Thank you for this important comment. We agree that the conclusion that MgIG weakens the interaction between HSD11B1 and SREBP2 should not be based solely on a single co-IP/Western blot experiment. In the revised manuscript, we have therefore toned down this statement to more appropriately reflect the data. Specifically, we now describe this result as a preliminary observation suggesting a potential modulation of the interaction, rather than a definitive conclusion. Please refer to Page 15, line 391.

      (10) It's not clear what the numbers represent on top of the Western blots. Are these averages over the course of three independent experiments?

      Thank you for this helpful comment. We apologize for the lack of clarity in the original figure presentation. The numbers shown above the Western blot bands represent the densitometric quantification of protein expression normalized to GAPDH, calculated from three independent experiments. However, this information was not clearly specified in the original figure, which may have led to confusion. To address this concern, we have now revised the manuscript by explicitly clarifying the meaning of these values in the figure legends. In addition, we have added bar graphs showing the quantified results from three independent experiments for Figures S3A, S4D, S6B, and S8H to improve transparency and data presentation.

      (11) The claim in line 382 that knockdown of Hsd11b1 resulted in accumulation of pSREBP2 is not supported by the data provided in Figure 6D.

      Thank you for pointing out this issue. We sincerely apologize for the incorrect description in the original manuscript. This was a wording error. We have made the revision on page 15, line394-396.

      (12) None of the images provided in Figure 6E support the claims stated in the results. Activation of SREBP2 leads to nuclear translocation and subsequent induction of genes involved in cholesterol biosynthesis and uptake. Manipulation of Hsd11b1 via OE or KD does not show any nuclear localization with DAPI.

      Thank you for this important comment. We agree that the original description was not sufficiently clear, which may have led to misunderstanding of the results. To clarify, Figure 6E includes two experimental contexts. Under basal (physiological) conditions in AML-12 cells, manipulation of Hsd11b1 (overexpression or knockdown) does not significantly affect the subcellular distribution of SREBP2. However, under EtOH/PA-induced stress conditions, Hsd11b1 overexpression promotes both nuclear and cytoplasmic levels of SREBP2, whereas Hsd11b1 knockdown reduces SREBP2 expression in both compartments. We have made the revision on page 16, line399.

      (13) The entire manuscript is focused on this axis of MgIG-Hsd11b1-Srebp2, but no Srebp2 transcriptional targets are ever measured.

      We sincerely appreciate this great suggestion. We have made the necessary revisions as suggested on page 12, lines 285-288, line 292 by adding the mRNA changes of Lcn2 and Ldlr, which are SREBP2 target genes. As requested, we have added the required information to Figures 1F and 1H.

      (14) Acc1 and Scd1 are Srebp1 targets, not Srebp2.

      Thank you for this important comment. We agree that Acc1 and Scd1 are well-established downstream target genes of SREBP1 rather than SREBP2. To better support our proposed SREBP2-related mechanism, we further examined canonical SREBP2 downstream target genes, including Lcn2 and Ldlr. The results are consistent with activation of SREBP2 signaling in our model. These data have now been included in the revised manuscript (Page 12, Lines 285–288 and 292; Figures 1F and 1H).

      (15) A major weakness of this manuscript is the lack of studies providing quantitative assessments of Srebp2 activation and true liver lipid measurements.

      Thank you for this important comment. We acknowledge the concern regarding the lack of direct quantitative assessment of SREBP2 activation in the original version of the manuscript. To address this limitation, we have strengthened the evidence supporting SREBP2 activation using multiple complementary approaches. Specifically, we assessed the expression of canonical SREBP2 downstream target genes (Page 12, Lines 285–288 and 292; Figures 1F and 1H), together with Western blot analysis (Figure 6D) and immunofluorescence staining (Figure 6F), which collectively support activation of SREBP2 signaling in the EtOH/PA-induced ALD model.

      In addition, to provide a more comprehensive evaluation of hepatic lipid accumulation, we measured serum TG and TC levels, as well as hepatic TG and TC content. These biochemical analyses further confirm the presence of significant lipid accumulation in our model. We have made the necessary revisions as suggested on page 12, lines 285-288 (Figure 1G).

      Reviewer #2 (Public review):

      (1) In Supplemental Figure 1A, all the treatment arms (A-control, MgIG-25 mg/kg, MgIG-50 mg/kg) showed body weight loss compared to the untreated controls. However, Figure 1E showed body weight gain in the treatment arms (A-control and MgIG-25 mg/kg), why? In Supplemental Figure 1A, the mice with MgIG (25 mg/kg) showed the lowest body weight, compared to either A-control or MgIG (50 mg/kg) treatment. Can the authors explain why MgIG (25 mg/kg) causes bodyweight loss more than MgIG (50 mg/kg)? What about the other parameters (ALT, ALS, NAS, etc.) for the mice with MgIG (50 mg/kg)?

      We agree that this observation does not strictly follow a dose-dependent pattern. In vivo responses to pharmacological interventions, particularly in metabolic and liver disease models, are not always linear. The relatively greater body weight reduction observed in the 25 mg/kg group may be influenced by inter-individual variability, differences in metabolic adaptation, or sample size–related variation. Importantly, these differences in body weight were not statistically significant. Therefore, we selected the 50 mg/kg dose for subsequent animal experiments, as it demonstrated more consistent and stable improvements across multiple parameters, including body weight, ALT, AST, TG, and TC.

      (2) IL-6 is a key pro-inflammatory cytokine significantly involved in ALD, acting as a marker of ALD severity. Can the authors explain why MgIG 1.0 mg/ml shows higher IL-6 gene expression than MgIG (0.1-0.5 mg/ml)? Same question for the mRNA levels of lipid metabolic enzymes Acc1 and Scd1.

      Thank you for this important comment. We agree that IL-6, as well as lipid metabolism–related genes such as Acc1 and Scd1, are key indicators in ALD. The relatively higher expression observed at 1.0 mg/mL MgIG compared to lower concentrations (0.1–0.5 mg/mL) may be related to experimental constraints associated with the MgIG formulation used in this study.

      Specifically, to maintain consistency with our in vivo experiments, we used a clinically available liquid formulation of MgIG (5 mg/mL), which is approved for intravenous administration in China. Due to its relatively low stock concentration, achieving higher working concentrations (e.g., 1.0 mg/mL) in vitro required a larger volume of the MgIG solution, thereby proportionally reducing the volume of culture medium. This reduction in effective culture conditions may adversely affect hepatocyte viability and function.

      Supporting this, our CCK-8 and LDH assays indicated that higher MgIG concentrations were associated with subtle cytotoxicity or impaired cell status.

      (3) For the qPCR results of Hsd11b1 knockdown (siRNA) and Hsd11b1 overexpression (plasmid) in AML-12 cells (Figure 5B), what is the description for the gene expression level (Y axis)? Fold changes versus GAPDH? Hsd11b1 overexpression showed non-efficiency (20-23, units on Y axis), even lower than the Hsd11b1 knockdown (above 50, units on Y axis). The authors need to explain this. For the plasmid-based Hsd11b1 overexpression, why does the scramble control inhibit Hsd11b1 gene expression (less than 2, units on the Y axis)? Again, this needs to be explained.

      Thank you for this important comment, and we apologize for the lack of clarity in the Y-axis labeling, which may have led to misunderstanding.

      As shown in Figures 5A and 5B, we have revised the Y-axis description to clearly indicate that gene expression levels are presented as relative expression normalized to GAPDH (fold change relative to the control group).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Use terms that show directionality to help the readers comprehend the data. For instance, Line 295 states MgIG treatment also modulated the expression.... In reality, MgIG treatment reduced the expression of those genes relative to ethanol-fed control mice.

      Thank you very much for this precious suggestion. We have thoroughly revised this part as ‘In line with the observed histological and physiological improvements, MgIG treatment also reduced the expression of genes involved in lipid synthesis metabolism (Srebp1, Srebp2, Acc1, and Scd1, Lcn2, and Ldlr), inflammation (Tnf-α and Il-6), and pro-apoptosis (Bax) while restored the level of anti-apoptotic gene (Bcl2) in the liver tissue of EtOH mice (Fig. 1G-1H).’. Please refer to page 12, lines 290-294.

      (2) Oil Red O staining is subjective and nonspecific. The authors make a claim that serum TGs are an indicator of liver function; however, measurement of hepatic TGs would be a better measure here and more consistent with the ORO staining.

      We sincerely appreciate this great suggestion. We have made the necessary revisions as suggested on page 12, lines 285-288 as ‘Notably, significant differences were observed between the EtOH group and the MgIG-treated (EtOH+M) group in serum levels of liver enzymes (ALT and AST), serum lipid parameters (TG and TC), as well as Liver TG and TC contents—-key indicators of liver function and lipid metabolism.’. As requested, we have added the required information to Figures 1G.

      (3) The focus of the paper is on this SREBP2 axis. However, in Figure 1, the authors do not show any SREBP2 target genes. This would be helpful in interpreting SREBP2 activity. Further, hepatic free cholesterol levels would also strengthen these data.

      We sincerely appreciate this great suggestion. We have made the necessary revisions as suggested on page 12, lines 285-288, line 292 by adding the mRNA changes of Lcn2 and Ldlr, which are SREBP2 target genes. As requested, we have added the required information to Figures 1F and 1H.

      (4) Labels showing directionality on the volcano plots in Figures 2A, B would be of great help here. It's unclear which groups are on the left or right.

      Thank you very much! The authors have revised Figures 2A-C as requested. Please refer to the new version of Figures 2A-C.

      (5) Ensure consistency in what is written in the results and the figure legends. See Figure 2 volcano plots for examples. The volcano plot in Figure 2B has no figure legend.

      We thank the reviewer for the valuable suggestion. Accordingly, we have revised the legend for Figures 2B as suggested.

      (6) Ensure consistency in the nomenclature. In some cases, the authors use ALD+MgIG, and in others, they just use MgIG. My recommendation would be to use Ctrl, EtOH, EtOH+M.

      We sincerely appreciate this great suggestion. We have made the necessary revisions as suggested on page 6, lines 111-112, page 11, line 280 and page 12, line 282, 284, 293, 298, 301.

      (7) The gene enrichment analysis in Figure 2C should also include some text about directionality, either in the figure or the figure legend. Upregulated DEGs in the MgIG group? It's unclear.

      We thank the reviewer for the valuable suggestion. Accordingly, we have revised the legend for Figures 2C as suggested.

      (8) The authors should consider shuffling the order of some of the figures for better transitions from one panel to the next. For instance, Figure 3B, C shows cell viability responses before showing the siRNA and OE are effective in knocking down and overexpressing their protein of interest.

      We thank the reviewer for the valuable suggestion. Accordingly, we have revised the legend for Figures 3B and 3C as suggested.

      (9) The authors need to be consistent in the colors that are used in the figures. It's incredibly hard to follow, as presented.

      We appreciate the reviewer's comment regarding color consistency. In response, we have carefully revised all figures to ensure consistent use of colors across the manuscript. The updated versions are shown in Figures 3, 6, and 7.

      (10) For Nile Red staining, multiple images at a lower objective need to be shown and/or cellular triglycerides and cholesterol levels should be quantified.

      We appreciate the reviewer's insightful comment regarding the Nile Red staining. In response, we have quantified triglyceride and total cholesterol levels in the cell supernatant, which are now presented on page 12, line 285-287 and Figures 2F. Furthermore, we have included additional Nile Red staining images at a lower objective in Supplementary Figures 2D, 3B, 4C to better illustrate the lipid droplet distribution.

      (11) Line 362 refers to Figure 4 when it should refer to Figure 5.

      Thank you very much! The authors have revised on page 14, line 364.

      (12) qPCR should be performed on canonical Srebp2 targets throughout the manuscript to tie in the MgG treatment with changes in sterol sensing and Srebp2.

      Thank you for your valuable suggestion. The results are now included on page 12, lines 292 and 311, and the corresponding data in Figures 1H and 2G have been enhanced accordingly.

      Reviewer #2 (Recommendations for the authors):

      (1) The statement, figure labeling, and figure legend for Figure 1A-C are confused. The MgIG dosing on the X-axis for Figure 2D is missing.

      Thank you for the correction. We have revised this problem. Please refer to the new version of Figure 1A-C and Figure 2D.

      (2) Figure 3E is not well described in the main text and figure legend. What are those numbers on top of the blotting bands? It was guessed that the numbers were the mean for each group. But where is the SD or SE for each group? It is hard to tell the statistical significance without showing SD or SE. The same question applies to Figure 5E, Figure 56C-6D, and Figure 7G.

      We sincerely appreciate this great suggestion. We have made the necessary revisions as suggested on page 13, lines 317-322. As suggested, we have added the required information to Figures S3A, S4D, S6B and S8H.

    1. eLife Assessment

      This valuable study advances our understanding of how organisms respond to chronic oxidative stress. Using the nematode C. elegans, the authors identified key neuronal signaling molecules and their receptors that are required for stress signaling and survival. The evidence supporting the conclusions is solid, including rigorous genetics, stress response analysis, and transcriptional profiling. This research will be of broad interest to neuroscientists and researchers working in the field of oxidative stress regulation.

    2. Reviewer #2 (Public review):

      In this paper, Biswas et al. describe the role of acetylcholine (ACh) signaling in protection against chronic oxidative stress in C. elegans. They showed that disruption of ACh signaling in either unc-17 mutant or gar-3 mutants led to sensitivity to toxicity caused by chronic paraquat (PQ) treatment. Using RNA seq, they found that approximately 70% of the genes induced by chronic PQ exposure in wild type failed to upregulate in these mutants. The overexpression of gar-3 selectively in cholinergic neurons was sufficient to promote protection against chronic PQ exposure in an ACh-dependent manner. The study points to a previously undescribed role for ACh signaling in providing organism-wide protection from chronic oxidative stress likely through the transcriptional regulation of numerous oxidative stress-response genes. The paper is well-written, and the data are robust. While the study identifies the muscarinic ACh receptor gar-3 as an important regulator of the response to PQ, the specific neurons in which gar-3 functions were not unambiguously identified, and the sources of ACh that regulate GAR-3 signaling and the identities of the tissues targeted by gar-3 remain unknown.

      Comments on revisions:

      No further comments.

    3. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The researchers aimed to identify which neurotransmitter pathways are required for animals to withstand chronic oxidative stress. This work thus has important implications for disease processes that are caused/linked to oxidative stress. This work identified specific neurotransmitters and receptors that coordinate stress resilience, both prior to and during stress exposure. Further, the authors identified specific transcriptional programs coordinated by neurotransmission that may provide stress resistance.

      Strengths:

      The manuscript is very clearly written with a well-formulated rationale. Standard C. elegans genetic analysis and rescue experiments were performed to identify key regulators of the chronic oxidative stress response. These findings were enhanced by transcriptional profiling that identified differentially expressed genes that likely affect survival when animals are exposed to stress.

      Weaknesses:

      Where the gar-3 promoter drives expression was not discussed in the context of the rescue experiments in Fig 7.

      Comments on revisions:

      This issue has now been appropriately addressed in the revision.

      We thank the reviewer for their time and constructive feedback.

      Reviewer #2 (Public review):

      In this paper, Biswas et al. describe the role of acetylcholine (ACh) signaling in protection against chronic oxidative stress in C. elegans. They showed that disruption of ACh signaling in either unc17 mutant or gar-3 mutants led to sensitivity to toxicity caused by chronic paraquat (PQ) treatment. Using RNA seq, they found that approximately 70% of the genes induced by chronic PQ exposure in wild type failed to upregulate in these mutants. The overexpression of gar-3 selectively in cholinergic neurons was sufficient to promote protection against chronic PQ exposure in an AChdependent manner. The study points to a previously undescribed role for ACh signaling in providing organism-wide protection from chronic oxidative stress likely through the transcriptional regulation of numerous oxidative stress-response genes. The paper is well-written, and the data are robust, though some conclusions seem preliminary and are not fully support the current data (see below). While the study identifies the muscarinic ACh receptor gar-3 as an important regulator of the response to PQ, the specific neurons in which gar-3 functions were not unambiguously identified, and the sources of ACh that regulate GAR-3 signaling and the identities of the tissues targeted by gar-3 were not addressed.

      Comments on revisions:

      The authors addressed my comments adequately in their revised submission. Please include representative images to accompany the quantification of the new results presented in Fig S4A.

      We thank the reviewer for their time and constructive feedback. We now include representative images as requested.

    1. eLife Assessment

      In this valuable study, the authors conducted an impressive amount of atomistic simulations with a glycosylated HIV-1 envelope glycoprotein (Env) trimer in a realistic asymmetric lipid bilayer. The aim was to probe how Env transmembrane domain, cytoplasmic tail, and membrane environment influence ectodomain orientation and antibody epitope exposure. The simulations convincingly show that ectodomain motion is dominated by tilting relative to the membrane and explicitly demonstrate the role of membrane asymmetry in modulating the protein conformation and orientation. Additional analyses of the authors' deposited MD trajectories could serve as invaluable extensions of this work to probe, for example, for exposure of cryptic epitopes and potential allosteric coupling.

    2. Reviewer #1 (Public review):

      Summary:

      In the manuscript "Conformational Variability of HIV-1 Env Trimer and Viral Vulnerability", the authors study the fully glycosylated HIV-1 Env protein using an all-atom forcefield. It combines long all-atom simulations of Env in a realistic asymmetric bilayer with careful data analysis. This work clarifies how the CT domain modulates the overall conformation of the Env ectodomain and characterizes different MPER-TMD conformations. The authors also carefully analyze the accessibility of different antibodies to the Env protein.

      Strengths:

      This paper is state-of-the-art given the scale of the system and the sophistication of the methods. The biological question is important, the methodology is rigorous, and the results will interest a broad elife audience. The authors also establish strong connections to previous literature and acknowledge the limitations of the CT-truncated protein construct, which enhances the manuscript's relevance to the community.

    3. Reviewer #2 (Public review):

      In this work, the authors elucidate how a viral surface protein behaves in a membrane environment and how its large-scale motions influence the exposure of antibody-binding sites. Using long-timescale, all-atom molecular dynamics simulations of a fully glycosylated, full-length protein embedded in a virus-like membrane, the study systematically examines the coupling between ectodomain motion, transmembrane orientation, membrane interactions, and epitope accessibility. Multiple model variants differing in cleavage state, initial transmembrane configuration, and presence of the cytoplasmic tail are compared to identify general features of protein-membrane dynamics relevant to antibody recognition.

      A major strength of this study is the scope and ambition of the simulations. The authors perform multiple microsecond-scale simulations of a highly complex, biologically realistic system that includes the full ectodomain, transmembrane region, cytoplasmic tail, glycans, and a heterogeneous membrane. The finding that the ectodomain explores a wide range of tilt angles while the transmembrane region remains more constrained, with limited correlation between the two, offers useful conceptual insight into how global motions may be accommodated without large rearrangements at the membrane anchor. The explicit consideration of membrane and glycan steric effects on antibody accessibility further strengthens the study.

      The main limitations relate to sampling and model dependence inherent to simulations of this size and complexity. The analysis of antibody accessibility is based on geometric and steric criteria, which do not capture potential conformational adaptations of antibodies or membrane remodeling during binding; the authors have appropriately noted this as a limitation.

      In the revised manuscript, the authors have addressed all previously raised concerns. Time series plots of the tilt angles have been added, figure captions and visual encodings have been clarified, quantitative descriptions of angular distributions have been strengthened, and the distance metric for MPER exposure is now accompanied by temporal data. The overall presentation is substantially improved, and the conclusions are well supported by the data as presented.

    4. Reviewer #3 (Public review):

      Summary:

      This study uses large-scale all-atom molecular dynamics simulations to examine the conformational plasticity of the HIV-1 envelope glycoprotein glycoprotein (Env) in a membrane context, with particular emphasis on how the transmembrane domain (TMD), cytoplasmic tail (CT), protomer cleavage, and membrane environment influence ectodomain orientation and antibody epitope exposure. By comparing Env constructs with and without the CT, explicitly modeling glycosylation, and embedding Env in an asymmetric lipid bilayer, the authors aim to provide an integrated view of how membrane-proximal regions and lipid interactions shape Env antigenicity, including epitopes targeted by MPER-directed antibodies.

      Strengths:

      The authors have made a genuine effort to address the concerns raised in the first round of review, and the revised manuscript is substantively improved. The addition of dynamical cross-correlation maps, expanded citation of prior computational work, clarification of the membrane composition rationale, data deposition to Zenodo, and the new discussion contextualizing the independence of ectodomain and TMD motions are all welcome. Several scientifically interesting aspects of the work merit highlighting before the remaining concerns are addressed.

      A key strength of this work remains the scope, scale, and realism of the simulation systems. The authors construct a very large, nearly complete-Env-scale model that includes a glycosylated Env trimer embedded in an asymmetric bilayer, enabling analysis of membrane-protein interactions that are difficult to capture experimentally. The inclusion of specific glycans at reported sites, and the focus on constructs with and without the CT or cleavage, are well motivated by existing biological and structural data.

      The observation that R696 orientation and its interacting partners give rise to asymmetric protomer conformations and distinct TMD tilts is a notable finding. The statement that interactions between R696 and lipid headgroups or CT residues can be strong enough to introduce a kink into the TMD is well-supported by representative snapshots and consistent with prior isolated-TMD simulations. The use of two initialization depths ("high" and "low") to probe R696 leaflet preference is methodologically interesting and the authors' interpretation - that there is a slight bias toward cytoplasmic leaflet interactions, but that these contacts could be highly dynamic over the course of viral entry - is appropriately cautious. It would be valuable to explicitly frame this as a hypothesis with testable predictions that future experimental or enhanced-sampling work could address. Similarly, the equilibration-driven kinking of the TMD core, consistent with prior isolated-TMD studies, represents a useful validation that extends those earlier observations to the intact trimeric context.

      The simulations reveal substantial tilting motions of the ectodomain relative to the membrane, with angles spanning roughly 0-30{degree sign} (and up to ~40{degree sign} in some analyses), while the ectodomain itself remains relatively rigid. This framing, that much of Env's conformational variability arises from rigid-body tilting rather than large internal rearrangements, is an important conceptual contribution. The authors also provide interesting observations regarding asymmetric bilayer deformations, including localized thinning and altered lipid headgroup interactions near the TMD and CT, which suggest a reciprocal coupling between Env and the surrounding membrane.

      The analysis of antibody-relevant epitopes across the prefusion state, including the V1/V2 and V3 loops, the CD4 binding site, and the MPER, is another strength. The study makes effective use of existing experimental knowledge in this context, for example by focusing on specific glycans known to occlude antibody binding, to motivate and interpret the simulations.

      Finally, the revised discussion provides more context that situates the study's findings and discrepancies within the broader literature, strengthening the manuscript's clarity and interpretability.

      Weaknesses:

      The revised work is much improved, but still includes substantive issues with writing including organization, such as paragraph run-ons, and citation issues. Improving these would help readers make the most of this important study.

      The revised Introduction now includes a paragraph summarizing prior MD work, which is an improvement. However, the paragraph remains structured around the limitations and setup of previous studies (e.g., "early studies were constrained by limited computational resources", short trajectory lengths, isolated constructs) rather than their findings. Readers benefit most from understanding what those studies showed - and where the present work confirms, extends, or diverges from those results. The current framing inadvertently positions prior work as deficient scaffolding rather than as independent data points converging on shared conclusions. The Introduction could be revised to briefly summarize the key biological conclusions from prior MD studies alongside their technical context, which could then be revisited in their appropriate place alongside key results.

      The authors have verified that PDB entries are cited at first mention, and this is noted. However, a recurring issue remains: key literature-supported conclusions appear in the Results and Discussion sections without accompanying citations at each point of use. Passages that summarize experimental or computational findings - particularly those used to validate or contextualize the authors' own results - require citation at every point of claim, not only at first introduction of a reference. This is not a minor stylistic preference. Downstream readers, systematic reviewers, and automated tools that map literature to claims (e.g., scite) rely on co-occurrence of claims and citations within the same passage. A citation appearing several paragraphs earlier does not carry attribution forward. As a practical example: the statement that "MPER-targeting antibodies bind effectively only after the gp120-gp41 trimer undergoes major conformational rearrangements toward a fusion-intermediate or post-fusion state (Frey et al., 2008; Alam et al., 2009; Chen et al., 2014; Lee et al., 2016)", which is appropriate. That same standard of inline attribution should be applied throughout - including in Results and Discussion subsections where prior experimental findings are mentioned without citation.

      Additionally, cited literature should be framed to highlight convergence with the authors' conclusions, not primarily to limitations of previous studies. Where prior studies independently support a finding, this should be stated explicitly. Independent replication across methods and systems is one of the strongest arguments for ground truth; treating it as such would improve the manuscript's scientific standing.

      Finally, the dynamical cross-correlation maps assess ectodomain-TMD coupling, and the authors appropriately acknowledge that microsecond simulations capture only the closed ground state. However, the revised manuscript does not address the question raised in the first review regarding CT-TMD and CT-ectodomain correlations. The Results section states that "very weak correlations between the ectodomain and the TMD" were found, but it is not clear whether the CT was included in this analysis or whether analogous correlation maps for CT-TMD and CT-ectodomain pairs were computed for the full-length systems. Additional analyses of the authors' deposited MD trajectories-such as probing for exposure of cryptic epitopes and potential allosteric coupling-could serve as valuable extensions of this work.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "Conformational Variability of HIV-1 Env Trimer and Viral Vulnerability", the authors study the fully glycosylated HIV-1 Env protein using an all-atom forcefield. It combines long all-atom simulations of Env in a realistic asymmetric bilayer with careful data analysis. This work clarifies how the CT domain modulates the overall conformation of the Env ectodomain and characterizes different MPER-TMD conformations. The authors also carefully analyze the accessibility of different antibodies to the Env protein.

      Strengths:

      This paper is state-of-the-art, given the scale of the system and the sophistication of the methods. The biological question is important, the methodology is rigorous, and the results will interest a broad audience.

      Weaknesses:

      The manuscript lacks a discussion of previous studies. The authors should consider addressing or comparing their work with the following points:

      (1) Tilting of the Env ectodomain has also been reported in previous experimental and theoretical work: https://doi.org/10.1101/2025.03.26.645577

      (2) A previous all-atom simulation study has characterized the conformational heterogeneity of the MPER-TMD domain: https://doi.org/10.1021/jacs.5c15421

      (3) Experimental studies have shown that MPER-directed antibodies recognize the prehairpin intermediate rather than the prefusion state: https://doi.org/10.1073/pnas.1807259115

      (4) How does the CT domain modulate the accessibility of these antibodies studied? The authors are in a strong position to compare their results with the following experimental study: https://doi.org/10.1126/science.aaa9804

      Based on the Reviewer’s comments and suggestions, we have added a discussion related to each previous study mentioned above.

      (1) Tilting of the Env ectodomain has also been reported in previous experimental and theoretical work: https://doi.org/10.1101/2025.03.26.645577

      At the end of the third paragraph (originally the second paragraph) in the Discussion section we added:

      “Shehata et al. also built a model of full-length gp120–gp41 trimer embedded in a lipid bilayer and performed all-atom simulations, in which a tilting motion of the ectodomain was observed. Based on the analysis of accessible surface area using different probe radii, they reported that antibody epitopes on the ectodomain are largely shielded by glycans, while the MPER epitope is mainly occluded by the membrane with tilt angles above 30° required to achieve greater MPER exposure (Shehata et al., 2025).”

      (2) A previous all-atom simulation study has characterized the conformational heterogeneity of the MPER-TMD domain: https://doi.org/10.1021/jacs.5c15421

      In the middle of the first paragraph in the Discussion section we added:

      “This is consistent with the all-atom simulations of MPER–TMD–CT and MPER–TMD in an asymmetric membrane conducted by Majumder et al., which likewise show multiple different conformational states of MPER and TMD (Majumder et al., 2025).”

      (3) Experimental studies have shown that MPER-directed antibodies recognize the prehairpin intermediate rather than the prefusion state: https://doi.org/10.1073/pnas.1807259115

      The paper mentioned by the Reviewer mainly reports the NMR structure of the MPER and TMD. In this study, the authors experimentally examined a series of MPER mutations to assess whether alterations in the MPER affect epitope accessibility in other regions of the Env ectodomain. This study did not investigate whether MPER-directed antibodies recognize the prehairpin intermediate. Instead, it cited prior studies (Frey et al.; 2008, Alam et al., 2009; and Chen et al., 2014) reporting that MPER-directed antibodies target the prehairpin intermediate conformation. We have already cited two of them (Alam et al., 2009 and Chen et al., 2014) in the original preprint, and we have now added the third one (Frey et al., 2008) in the revised manuscript.

      In the middle of the third paragraph (originally the second paragraph) in the Discussion section we added:

      “This is consistent with experiment studies indicating that MPER-targeting antibodies bind effectively only after the gp120–gp41 trimer undergoes major conformational rearrangements toward a fusion-intermediate or post-fusion state (Frey et al., 2008; Alam et al., 2009; Chen et al., 2014; Lee et al., 2016).”

      (4) How does the CT domain modulate the accessibility of these antibodies studied? The authors are in a strong position to compare their results with the following experimental study: https://doi.org/10.1126/science.aaa9804

      At the beginning of the second paragraph in the Discussion section we added:

      “Comparison of the full-length and CT-truncated systems shows that the primary difference arises from changes in the lipid bilayer, particularly in the exoplasmic leaflet, whereas differences in protein conformation and dynamics are less evident. Previous experimental studies have reported that mutations of the TMD residue and CT truncation can substantially affect antigenicity of ectodomain (Edwards et al., 2002; Chen et al., 2015; Dev et al., 2016). However, the ectodomain remains relatively rigid in our simulations for both full-length and CT-truncated systems. It is unclear whether this behavior reflects insufficient conformational sampling or artifacts associated with the model structures. Structural information for the CT is very limited, and the NMR structure (PDB ID: 7LOI) was the only available CT structure at the time the simulation systems were constructed. As a result, the extent to which this structure represents the native CT conformation remains uncertain. Additional experimental structural characterization of the CT will be important for achieving a more complete understanding of its functional role.”

      Reviewer #1 (Recommendations for the authors):

      A minor point: The RMSD values in Figure 3-figure supplement 1, seem a little too small. Please check the units.

      Figure 3-figure supplement 1 shows the RMSD of the ectodomain. Prior to RMSD calculation, the snapshots extracted from each trajectory were aligned to the initial structure using the ectodomain as the reference to avoid falsely high RMSD values arising from different orientations of the ectodomain. The relatively small RMSD values therefore reflect the intrinsic structural stability of the ectodomain, indicating that its internal conformation remains stable even though it undergoes substantial tilting motions.

      Reviewer #2 (Public review):

      Summary:

      In this work, the authors aim to elucidate how a viral surface protein behaves in a membrane environment and how its large-scale motions influence the exposure of antibody-binding sites. Using long-timescale, all-atom molecular dynamics simulations of a fully glycosylated, full-length protein embedded in a virus-like membrane, the study systematically examines the coupling between ectodomain motion, transmembrane orientation, membrane interactions, and epitope accessibility. By comparing multiple model variants that differ in cleavage state, initial transmembrane configuration, and presence of the cytoplasmic tail, the authors aim to identify general features of protein-membrane dynamics relevant to antibody recognition.

      Strengths:

      A major strength of this study is the scope and ambition of the simulations. The authors perform multiple microsecond-scale simulations of a highly complex, biologically realistic system that includes the full ectodomain, transmembrane region, cytoplasmic tail, glycans, and a heterogeneous membrane. Such simulations remain technically challenging, and the work represents a substantial computational and methodological effort.

      The analysis provides a clear and intuitive description of large-scale protein motions relative to the membrane, including ectodomain tilting and transmembrane orientation. The finding that the ectodomain explores a wide range of tilt angles while the transmembrane region remains more constrained, with limited correlation between the two, offers useful conceptual insight into how global motions may be accommodated without large rearrangements at the membrane anchor.

      Another strength is the explicit consideration of membrane and glycan steric effects on antibody accessibility. By evaluating multiple classes of antibodies targeting distinct regions of the protein, the study highlights how membrane proximity and glycan dynamics can differentially influence access to different epitopes. This comparative approach helps place the results in a broader immunological context and may be useful for readers interested in antibody recognition or vaccine design.

      Overall, the results are internally consistent across multiple simulations and model variants, and the conclusions are generally well aligned with the data presented.

      Weaknesses:

      The main limitations of the study relate to sampling and model dependence, which are inherent challenges for simulations of this size and complexity. Although the simulations are long by current standards, individual trajectories explore only portions of the available conformational space, and several conclusions rely on pooling data across a limited number of replicas. This makes it difficult to fully assess the robustness of some quantitative trends, particularly for rare events such as specific epitope accessibility states.

      In addition, several aspects of the model construction, including the treatment of missing regions, loop rebuilding, and initial configuration choices, are necessarily approximate. While these approaches are reasonable and well motivated, the extent to which some conclusions depend on these modeling choices is not always fully clear from the current presentation.

      Finally, the analysis of antibody accessibility is based on geometric and steric criteria, which provide a useful first-order approximation but do not capture potential conformational adaptations of antibodies or membrane remodeling during binding. As a result, the accessibility results should be interpreted primarily as model-based predictions rather than definitive statements about binding competence.

      Despite these limitations, the study provides a valuable and carefully executed contribution, and its datasets and analytical framework are likely to be useful to others interested in protein-membrane interactions and antibody recognition.

      Based on the Reviewer’s comments, we have revised the Discussion section to emphasize the limitation related to model construction and analysis of antibody accessibility.

      In the middle of the second paragraph in the Discussion section we added:

      “Similar limitations apply to other modeled regions where structural information is incomplete, including missing loops in the ectodomain, the cleavage site and heptad repeat 2 where two PDB structures (IDs: 6B0N and 7LOI) were merged. These regions introduce additional uncertainty, and the extent to which they influence the interpretation of our results remains an open question.”

      In the middle of the third paragraph (originally the second paragraph) in the Discussion section we added:

      “In addition, this analysis is based on geometric and steric criteria without accounting for potential conformational adaptations of gp120–gp41, antibodies, or the membrane; therefore, the calculated frequency of antibody accessibility should be interpreted as an approximation rather than a definitive indicator of binding competence.”

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 45-47: The phrase "A major breakthrough was the design of ..." may be confusing. The gp140 trimer refers to a naturally occurring form of the HIV envelope protein rather than a structure designed de novo. If this statement refers to the development of a specific experimental construct or model system, this should be clarified to avoid misunderstanding.

      We have revised the sentence to clarify that the statement refers to soluble gp140 trimer constructs developed to stabilize the prefusion Env ectodomain for structural and immunological studies.

      At the beginning of the second paragraph in the Introduction section, we have modified the following:

      “A major advance was the development of soluble gp140 trimers, composing gp120 and the ectodomain portion of gp41, designed to stabilize the prefusion Env trimer for structural and immunological characterization.”

      (2) Figure 1A: The figure displays a model structure lacking the cytoplasmic tail. Given that the full-length model is central to the study, the authors may wish to explain why the truncated structure is shown here or consider displaying the full-length model to better reflect the complete system analyzed.

      We have combined Figure 1 and Figure 1—figure supplements 1 to show both full-length and CT-truncated models in one figure. We have also added an explanation of why the CT-truncated model was used as the primary system for analysis.

      In the middle of the third paragraph in the Introduction section we added:

      “However, structural information for the CT remains limited, leading to uncertainty in its conformational organization. To reduce potential bias arising from this uncertainty, we also generated a CT-truncated model and used it as the primary system for analysis (Figure 1, Figure 1—figure supplements 1).”

      We have modified Figure 1

      We removed Figure 1—figure supplements 1

      (3) Line 106: The probability distributions of θEC and θTM are cited in support of the statement that the angles "typically range from ... with occasional tilting." Providing explicit quantitative measures (for example, means, percentiles, or fractions of time spent in different angular regimes) would strengthen this claim.

      We have revised the text to explicitly indicate that only 0.7‰ of the sampled θ<sub>EC</sub> values are greater than 40°.

      In the middle of the first paragraph in the subsection “The ectodomain maintains a rigid internal structure and tilts independently of the TMD” we have modified the following:

      “Across trajectories, θ<sub>EC</sub> typically ranges from 0° to 40°, with only 0.7‰ exceeding 40°”.

      (4) Figure 2: The meaning of the contour lines is not clearly explained. If these represent probability density estimates of angular values over the trajectory, this should be stated explicitly. In addition, because the angles may evolve over time, it would be helpful to clarify how temporal drift is accounted for in the contour representation.

      We have clarified in both the main text and the figure caption that the contour lines in Figure 2B represent the joint probability density of the ectodomain and TMD tilt angles. We have also added Figure 2—figure supplements 5–8 showing the temporal evolution of the ectodomain and TMD tilt angles.

      In the middle of the first paragraph in the subsection “The ectodomain maintains a rigid internal structure and tilts independently of the TMD” we have modified the following:

      “The temporal evolution of θ<sub>EC</sub> and θ<sub>TM</sub> is additionally shown in Figure 2—figure supplements 5–8. For the CT-truncated systems, the joint probability densities of θ<sub>EC</sub> and θ<sub>TM</sub> calculated from the final 0.5 µs of each trajectory are shown in Figure 2B, while those for the full-length systems are shown in Figure 2—figure supplement 9.”

      In the caption of Figure 2 we have modified the following:

      “(B) Probability densities of ectodomain and TMD tilt angles, calculated from CT-truncated systems with various initial configurations.”

      We have added Figure 2—figure supplements 5–8.

      We have modified the following:

      “The original Figure 2—figure supplements 5 has been renumbered as Figure 2—figure supplements 9.”

      (5) Figure 2 (supplements): Some datasets are shown using scatter plots, while others are presented as contour plots. Using a consistent visualization style across panels or clearly explaining the rationale for the different representations would improve clarity.

      The contour plots in Figure 2B and Figure 2—figure supplements 9 show the joint distribution of the ectodomain and TMD tilt angles during the final 0.5 µs of each trajectory, whereas the scatter plots in Figure 2—figure supplements 1–4 illustrate the variations of the tilt angles across different time intervals. Each 1-µs trajectory was divided into four 0.25-µs intervals, indicated by light gray, dark gray, black, and red respectively, as shown in the legends of Figure 2—figure supplements 1–4. We have clarified in the main text that the multi-colored scatter plots are intended to demonstrate that large conformational changes predominantly occurred during the first 0.5 µs of each trajectory.

      In the middle of the first paragraph in the subsection “The ectodomain maintains a rigid internal structure and tilts independently of the TMD” we have modified the following:

      “Each 1-µs trajectory is divided into four consecutive 0.25-µs intervals, and data points from each interval are distinguished by four different colors (Figure 2—figure supplements 1–4). The variations of θ<sub>EC</sub> and θ<sub>TM</sub> over time show that large conformational changes predominantly occurred during the first 0.5 µs, followed by convergence of the θ<sub>EC</sub> and θ<sub>TM</sub> distributions during the second 0.5 µs in most trajectories.”

      (6) As noted in Line 97, θEC and θTM tilt independently. In this context, presenting time series plots of θEC and θTM separately would be highly informative. Such plots would help readers distinguish between equilibration behavior, drift from initial conditions, and equilibrium fluctuations.

      We have added Figure 2—figure supplements 5–8 showing the temporal evolution of the ectodomain and TMD tilt angles, as noted in our response to comment (4).

      (7) Figure 3A: It is not immediately clear which panels correspond to top views and which correspond to side views. Explicitly labeling these views in the figure or caption would reduce ambiguity.

      We have added labels in Figure 3A to clearly denote the top-view and side-view panels.

      (8) Figure 3B: The description "...by solid and transparent colors..." is ambiguous, as it is unclear whether this refers to color intensity or transparency. The caption would benefit from explicitly stating the visual encoding used (for example, darker/lighter colors or left/right bars).

      We have revised the figure caption to clarify which boxes correspond to cleaved systems and which correspond to uncleaved systems.

      In the caption of Figure 3 we have modified the following:

      “For each residue, the distribution from cleaved systems is shown in dark color (left), and that from uncleaved systems is shown in light color (right).”

      (9) Figure 4H: The definition of "frequency" expressed as a percentage is unclear. If this represents the fraction of snapshots in which two atoms fall within a specified distance range, this should be stated explicitly. The authors should also clarify whether the reported quantity is a probability or a rate, and ensure that the units and terminology are consistent.

      We have revised the figure caption to clarify that the frequency represents the fraction of snapshots in which the heavy atoms of a TMD residue and the interacting component are within 5 Å.

      In the caption of Figure 4 we have modified the following:

      “For each TMD residue–interacting component pair, the frequency represents the fraction of snapshots in which the heavy atoms of the TMD residue and the corresponding component are within 5 Å. Bar shading reflects this fraction, with fully filled bars indicating 100% and empty bars indicating 0%.”

      (10) Line 170: The manuscript describes a "rapid rearrangement" of the transmembrane domain at early simulation times. It would be helpful to clarify whether this regime is considered equilibration and whether it is excluded from subsequent analyses. Plotting time series of the relevant tilting angles and transmembrane rearrangement metrics could help address this point.

      We have clarified that the TMD underwent conformational changes early in the equilibration stage to enable R696 to interact with lipid headgroups, ions, or CT residues, and these interactions were largely maintained throughout the production stage. The time series of TMD tilting angles are now shown in Figure 2—figure supplements 5–8. Notably, the TMD exhibits heterogeneous conformational changes, including tilting, bending, and partial loss of helical structure. Therefore, no single metric or limited set of metrics can comprehensively capture the full extent of TMD conformational variability.

      In the middle of the first paragraph in the subsection “The energetically unfavorable R696 in the hydrophobic core results in asymmetric, kinked TMD conformations and disrupts membrane integrity” we have modified the following:

      “Early in the equilibration stage, the TMD rapidly rearranged to allow R696 residues to interact with more favorable partners, including negatively charged lipid headgroups from either leaflet, ions and water molecules diffusing into the bilayer center, as well as polar and positively charged groups in the CT when present. Once the interactions between R696 residues and their binding partners (lipid headgroup, ions or CT residues) were established, they remained stable with minimal changes throughout the production stage.”

      (11) Line 213: As with earlier sections, time series plots of θEC and θTM, similar to those shown in Figure 3-figure supplement 1, would greatly aid interpretation by showing whether these angles drift or fluctuate around stable values.

      The time series of θ<sub>EC</sub> and θ<sub>TM</sub> are now shown in Figure 2—figure supplements 5–8. Line 213 refers to the conformational variability of the MPER. For the same reason discussed in our response to comment (10), the MPER exhibits even greater conformational heterogeneity than the TMD, and therefore cannot be adequately described by a single or small set of geometric metrics such as tilt or bending angles.

      (12) Lines 216-222: The term "trajectories" may be misleading in this context. It is unclear whether the differences discussed arise from different trajectories of the same system or from different systems altogether. Clarifying this distinction would improve interoperability.

      In this paragraph, we describe MPER conformational variations observed across all trajectories from all systems. A preceding sentence has been modified to emphasize that all trajectories from all systems are included. In addition, we have clarified which specific trajectory is referred to when discussing each example.

      At the beginning of the first paragraph in the subsection “MPER adopts diverse conformations, and its exposure depends on both MPER and TMD conformations” we have modified the following:

      “…, and a wide variety of conformations were sampled across all trajectories from all systems.”

      “Such conformation and orientation were maintained in some trajectories such as CL<sup>ΔCT</sup>3 (the third trajectory of the cleaved, CT-truncated system with the low TMD position, Figure 4—figure supplement 2C). In other trajectories, such as CL<sup>CT</sup>1, the helix-turn-helix MPER in one protomer shifted into a horizontal orientation parallel to the membrane surface (Figure 4—figure supplement 6A). In UL<sup>ΔCT</sup>1, the entire MPER adopted a more vertical arrangement, with both MPER-N and MPER-C tilted outward (Figure 4E, Figure 4—figure supplement 4A). We also observed in UH<sup>ΔCT</sup>3 and UL<sup>ΔCT</sup>3 that the HR2 helix in the ectodomain, MPER, and TMD merged into a continuous long helix (Figure 4C, F, Figure 4—figure supplement 3C, 4C). In addition, loss of helical structure within the MPER was common, particularly in the MPER-C region, which often transitioned to a random coil.”

      (13) Lines 280 and 287: Similar concerns apply to the use of the term "trajectories." If observations differ primarily between systems rather than between trajectories within a system, revising the wording accordingly would avoid confusion.

      We have revised the text to clarify that all trajectories from all systems are considered collectively.

      In the middle of the second paragraph in the subsection “Ectodomain epitopes are conditionally accessible, whereas MPER epitopes are virtually inaccessible in the closed prefusion state” we have modified the following:

      “When considering all trajectories from all systems collectively, approximately half of them exhibited at least one protomer with >35% accessibility (Supplementary file 1–Supplementary Table 2).”

      (14) Figure 5B: Providing a time series of the distance dF673, at least in the Supporting Information, would help assess sampling and equilibration. Such plots would complement the probability distributions and increase confidence in the reported trends.

      We have added Figure 5—figure supplement 1 showing the time series of the distance d<sub>F673</sub> to complement the probability distribution in Figure 5B.

      In the middle of the second paragraph in the subsection “MPER adopts diverse conformations, and its exposure depends on both MPER and TMD conformations”, we have modified the following:

      “In the initial ‘low’ and ‘high’ TMD configurations, dF673 was 6.1 Å and 9.1 Å, respectively, but across simulations it spanned a wide range from -15 Å to 20 Å (Figure 5A, B, Figure 5—figure supplement 1).”

      We have added Figure 5—figure supplement 1.

      Reviewer #3 (Public review):

      Summary:

      This study uses large-scale all-atom molecular dynamics simulations to examine the conformational plasticity of the HIV-1 envelope glycoprotein (Env) in a membrane context, with particular emphasis on how the transmembrane domain (TMD), cytoplasmic tail (CT), and membrane environment influence ectodomain orientation and antibody epitope exposure. By comparing Env constructs with and without the CT, explicitly modeling glycosylation, and embedding Env in an asymmetric lipid bilayer, the authors aim to provide an integrated view of how membrane-proximal regions and lipid interactions shape Env antigenicity, including epitopes targeted by MPER-directed antibodies.

      Strengths:

      A key strength of this work is the scope and realism of the simulation systems. The authors construct a very large, nearly complete Env-scale model that includes a glycosylated Env trimer embedded in an asymmetric bilayer, enabling analysis of membrane-protein interactions that are difficult to capture experimentally. The inclusion of specific glycans at reported sites, and the focus on constructs with and without the CT, are well motivated by existing biological and structural data.

      The simulations reveal substantial tilting motions of the ectodomain relative to the membrane, with angles spanning roughly 0-30° (and up to ~50° in some analyses), while the ectodomain itself remains relatively rigid. This framing, that much of Env's conformational variability arises from rigid-body tilting rather than large internal rearrangements, is an important conceptual contribution. The authors also provide interesting observations regarding asymmetric bilayer deformations, including localized thinning and altered lipid headgroup interactions near the TMD and CT, which suggest a reciprocal coupling between Env and the surrounding membrane.

      The analysis of antibody-relevant epitopes across the prefusion state, including the V1/V2 and V3 loops, the CD4 binding site, and the MPER, is another strength. The study makes effective use of existing experimental knowledge in this context, for example, by focusing on specific glycans known to occlude antibody binding, to motivate and interpret the simulations.

      Weaknesses:

      While the simulations are technically impressive, the manuscript would benefit from more explicit cross-validation against prior experimental and computational work throughout the Results and Discussion, and better framing in the introduction. Many of the reported behaviors, such as ectodomain tilting, TMD kinking, lipid interactions at helix boundaries, and aspects of membrane deformation, have been described previously in a range of MD studies of HIV Env and related constructs (e.g., PMC2730987, PMC2980712, PMC4254001, PMC4040535, PMC6035291, PMC12665260, PMID: 33882664, PMC11975376). Clearly situating the present results relative to these studies would strengthen the paper by clarifying where the simulations reproduce established behavior and where they extend it to more complete or realistic systems.

      A related limitation is that the work remains largely descriptive with respect to conformational coupling. Numerous experimental studies have demonstrated functional and conformational coupling between the TMD, CT, and the antigenic surface, with effects on Env stability, infectivity, and antibody binding (e.g., PMC4701381, PMC4304640, PMC5085267). In this context, the statement that ectodomain and TMD tilting motions are independent is a strong conclusion that is not fully supported by the analyses presented, particularly given the authors' acknowledgment that multiple independent simulations are required to adequately sample conformational space. More direct analyses of coupling, rather than correlations inferred from individual trajectories, would help align the simulations with the existing experimental literature. Given the scale of these simulations, a more thorough analysis of coupling could be this paper's most seminal contribution to the field.

      The choice of membrane composition also warrants deeper discussion. The manuscript states that it relies on a plasma membrane model derived from a prior simulation-based study, which itself is based on host plasma membrane (PMID: 35167752), but experimental analyses have shown that HIV virions differ substantially from host plasma membranes (e.g., PMC46679, PMC1413831, PMC10663554, PMC5039752, PMC6881329). In particular, virions are depleted in PC, PE, and PI, and enriched in phosphatidylserine, sphingomyelins, and cholesterol. These differences are likely to influence bilayer thickness, rigidity, and lipid-protein interactions and, therefore, may affect the generality of the conclusions regarding Env dynamics and antigenicity. Notably, the citation provided for membrane composition is a laboratory self-citation, a secondary source, rather than a primary experimental study on plasma membrane composition.

      Finally, there are pervasive issues with citation and methodological clarity. Several structural models are referred to only by PDB ID without citation, and in at least one case, a structure described as cryo-EM is in fact an NMR-derived model. Statements regarding residue flexibility, missing regions in structures, and comparisons to prior dynamics studies are often presented without appropriate references. The Methods section also lacks sufficient detail for a system of this size and complexity, limiting readers' ability to assess robustness or reproducibility.

      With stronger integration of prior experimental and computational literature, this work has the potential to serve as a valuable reference for how Env behaves in a realistic, glycosylated, membrane-embedded context. The simulation framework itself is well-suited for future studies incorporating mutations, strain variation, antibodies, inhibitors, or receptor and co-receptor engagement. In its current form, the primary contribution of the study is to consolidate and extend existing observations within a single, large-scale model, providing a useful platform for future mechanistic investigations.

      Following the Reviewer’s comments and suggestions, we have revised the manuscript accordingly.

      While the simulations are technically impressive, the manuscript would benefit from more explicit cross-validation against prior experimental and computational work throughout the Results and Discussion, and better framing in the introduction. Many of the reported behaviors, such as ectodomain tilting, TMD kinking, lipid interactions at helix boundaries, and aspects of membrane deformation, have been described previously in a range of MD studies of HIV Env and related constructs (e.g., PMC2730987, PMC2980712, PMC4254001, PMC4040535, PMC6035291, PMC12665260, PMID: 33882664, PMC11975376). Clearly situating the present results relative to these studies would strengthen the paper by clarifying where the simulations reproduce established behavior and where they extend it to more complete or realistic systems.

      We have added a summary of the prior computational studies in the Introduction section.

      At the beginning of the third paragraph in the Introduction section we added:

      “Molecular dynamics (MD) simulations have been employed to investigate the stability and conformational properties of monomeric and trimeric helical TMD in both aqueous and lipid bilayer environments since late 2000s (Kim et al., 2009; Gangupomu et al., 2010; Baker et al., 2014; Baker et al., 2014; Hollingsworth et al., 2018). Early studies were constrained by limited computational resources and therefore the simulation times are relatively short. Subsequent work employed metadynamics to probe rare events (Gangupomu et al., 2010; Baker et al., 2014), and simulations performed on Anton supercomputers extended sampling to multi-microsecond time scale (Baker et al., 2014). Piai and coworkers determined the NMR structure of a construct comprising the MPER, TMD, and CT, and carried out MD simulations to access the structural stability of the trimeric MPER–TMD–CT complex (Piai et al., 2021). Majumder et al. subsequently simulated the same MPER–TMD–CT complex and applied a machine learning-based approach to classify its conformational ensemble (Majumder et al., 2025). Maillie et al. combined conventional MD, steered MD, and coarse-grained simulations to examine interactions between MPER-targeting antibodies and membrane lipids (Maillie et al., 2025). In addition, MD simulations have been extensively applied to the well-studied ectodomain. Despite these advances, it remains challenging to investigate the gp120–gp41 trimer as an intact entity considering its structural complexity.”

      We have also added a discussion of previous MD simulation studies to the Result section regarding interactions of the TMD residue R696 with ions and lipid headgroups.

      At the end of the first paragraph in the subsection “The energetically unfavorable R696 in the hydrophobic core results in asymmetric, kinked TMD conformations and disrupts membrane integrity”

      “Previously, Kim et al. reported that the inter-chain interactions between protonated R696 gradually diminished over a short simulation time (23 ns), leading to increased crossing angles and reduced bundle length (Kim et al., 2009). Gangupomu et. al and Baker et. al observed that R696 snorkeled toward either exoplasmic or endoplasmic headgroups in simulations of the TMD monomer, resulting in TMD tilting and membrane thinning due to water penetration and lipid headgroups interacting with R696 (Gangupomu et al., 2010; Baker et al., 2014; Baker et al., 2014). These observations are consistent with our finding. Hollingsworth et. al also reported membrane thinning; however, they attributed this effect to interfacial interactions of R683 and R707 with both leaflets and proposed that R696 only interacted with water and ions permeating into the center of the TMD timer (Hollingsworth et al., 2018).”

      A related limitation is that the work remains largely descriptive with respect to conformational coupling. Numerous experimental studies have demonstrated functional and conformational coupling between the TMD, CT, and the antigenic surface, with effects on Env stability, infectivity, and antibody binding (e.g., PMC4701381, PMC4304640, PMC5085267). In this context, the statement that ectodomain and TMD tilting motions are independent is a strong conclusion that is not fully supported by the analyses presented, particularly given the authors' acknowledgment that multiple independent simulations are required to adequately sample conformational space. More direct analyses of coupling, rather than correlations inferred from individual trajectories, would help align the simulations with the existing experimental literature. Given the scale of these simulations, a more thorough analysis of coupling could be this paper's most seminal contribution to the field.

      We have added a discussion of the coupling between TMD, CT and Env antigenicity, and the independent motion of ectodomain and TMD in our simulation.

      In the middle of the second paragraph in the Discussion section

      “Our analysis of the ectodomain and TMD coupling indicates that the motions of these two domains are largely independent. This observation does not contradict experimental studies demonstrating functional coupling between the TMD, CT, and the antigenic profiles of Env (Chen et al., 2015; Dev et al., 2016). Munro et al. proposed that unliganded Env is intrinsically dynamic, transitioning among three distinct prefusion conformations: a closed ground state (predominant), a transient state, and a CD4-/co-receptor-stabilized state. Both laboratory-adapted and clinically isolated strains can spontaneously transition among these three states, although their relative occupancies differ (Munro et al., 2014). It is therefore possible that TMD mutations or CT truncation also alter the equilibrium distribution among three states, thereby affecting the epitope exposure, particularly for epitopes that are occluded in the closed ground state while exposed in the CD4-/co-receptor-stabilized state. However, transition among three states occur on millisecond-to-second timescales. Our simulations on microsecond timescales primarily capture conformational variations within the closed ground state and suggest that the MPER acts as a hinge, providing substantial flexibility that enables the ectodomain and TMD to move independently while Env remains in the closed ground state.”

      We have also calculated the dynamical cross-correlation maps showing very weak correlations between the ectodomain and the TMD.

      At the end of the first paragraph in the subsection “The ectodomain maintains a rigid internal structure and tilts independently of the TMD”

      “We also calculated the dynamical cross-correlation maps (Ichiye et al., 1991) of Cα atoms for all systems using CPPTRAJ (Roe et al., 2013). The results indicate only very weak correlations between the ectodomain and the TMD (Figure 2—figure supplements 10–13).”

      We have added Figure 2—figure supplements 10–13.

      The choice of membrane composition also warrants deeper discussion. The manuscript states that it relies on a plasma membrane model derived from a prior simulation-based study, which itself is based on host plasma membrane (PMID: 35167752), but experimental analyses have shown that HIV virions differ substantially from host plasma membranes (e.g., PMC46679, PMC1413831, PMC10663554, PMC5039752, PMC6881329). In particular, virions are depleted in PC, PE, and PI, and enriched in phosphatidylserine, sphingomyelins, and cholesterol. These differences are likely to influence bilayer thickness, rigidity, and lipid-protein interactions and, therefore, may affect the generality of the conclusions regarding Env dynamics and antigenicity. Notably, the citation provided for membrane composition is a laboratory self-citation, a secondary source, rather than a primary experimental study on plasma membrane composition.

      We have added references to primary experimental studies on plasma membrane composition (van Meer et al., 2008; Sampaio et al., 2011), as well as the prior simulation study proposing the lipid and cholesterol distributions (Ingolfsson et al., 2014).

      At the beginning of the Membrane subsection in the Materials and methods section

      We have modified the following:

      The full-length and CT-truncated gp120–gp41 models were embedded into an asymmetric lipid bilayer with the lipid composition corresponding to a mammalian plasma membrane (van Meer et al., 2008; Sampaio et al., 2011; Ingolfsson et al., 2014; Pogozheva et al., 2022),

      We have also clarified the limitations associated with the choice of lipid composition and emphasized the need to investigate its influence in future studies.

      At the end of the second paragraph in the Discussion section we added:

      “In addition to the limitations inherent to protein structure modeling, the choice of lipid composition remains an open question. In this work, we selected an asymmetric mammalian plasma membrane because it is one of the 18 complex biomembrane systems we previously studied (Pogozheva et al., 2022), and among them, it provides the closest available approximation to the HIV membrane. Nevertheless, experimental studies have reported differences in lipid composition between HIV virions and the host plasma membrane (Aloia et al., 1993; Brugger et al., 2006; Huarte et al., 2016; Mucksch et al., 2019; Tomishige et al., 2023). Although we do not anticipate that our main conclusions regarding Env domain motions and MPER flexibility would change substantially, evaluating the influence of lipid composition represents an important direction for future work.”

      Finally, there are pervasive issues with citation and methodological clarity. Several structural models are referred to only by PDB ID without citation, and in at least one case, a structure described as cryo-EM is in fact an NMR-derived model. Statements regarding residue flexibility, missing regions in structures, and comparisons to prior dynamics studies are often presented without appropriate references. The Methods section also lacks sufficient detail for a system of this size and complexity, limiting readers' ability to assess robustness or reproducibility.

      We have corrected the error in which PDB structure 7LOI was described as a cryo-EM structure; it is in fact an NMR structure. We have also verified that all PDB structures are properly cited at their first occurrence in the manuscript.

      We have clarified that the modeling of palmitoylation sites, glycans and lipid bilayers are done in an automated fashion by different modules in CHARMM-GUI, and added Supplementary file 1–Supplementary Table 8 showing the simulation settings for equilibration and production stages.

      At the end of the subsection “Modeling of full-length gp120–gp41 trimer” we have modified the following:

      “Two mutations (S764C and S837C) were introduced in the CT to restore the palmitoylation sites, and lipid tails oriented towards the hydrophobic core of the bilayer were then attached to the palmitoylation sites using the PDB Manipulation module in CHARMM-GUI (Jo et al., 2008; Jo et al., 2014; Park et al., 2023) (Figure 1D).”

      At the end of the subsection “Glycosylation” we added:

      “The select glycan sequences were represented in the Glycan Reader Sequence format (Jo et al., 2011; Park et al., 2017) and added to the corresponding glycosylation sites using the Glycan Reader & Modeler graphical interface.”

      In the middle of the subsection “Membrane” we added:

      “Membrane systems were constructed using CHARMM-GUI Membrane Builder, which provides a user-friendly graphical interface for selecting lipid types and defining their numbers in each leaflet (Jo et al., 2007; Jo et al., 2009; Wu et al., 2014; Lee et al., 2016; Lee et al., 2019).”

      In the middle of the subsection “Simulation details” we added:

      We have modified the following:

      “Positional and dihedral restraints were applied to proteins, glycans, and lipids, with force constants progressively reduced over successive intervals (Supplementary file 1–Supplementary Table 8).”

      We added Supplementary file 1–Supplementary Table 8.

      Reviewer #3 (Recommendations for the authors):

      Major concerns:

      (1) Strengthen analysis of conformational coupling: Consider analyses that more directly assess coupling between the TMD/CT and ectodomain, such as residue-residue correlation networks, comparisons to smFRET-defined conformational states, or data-driven (e.g., machine learning-based) trajectory analyses. Machine-learning analysis would be particularly helpful in understanding otherwise elusive allosteric networks that could govern large-scale behavior. Discuss how, due to the apparent local minima that occur after ~0.5 us, enhanced sampling methods might be employed to better cover the Env conformational landscape.

      We have calculated the dynamical cross-correlation maps showing very weak correlations between the ectodomain and the TMD.

      At the end of the first paragraph in the subsection “The ectodomain maintains a rigid internal structure and tilts independently of the TMD”

      “We also calculated the dynamical cross-correlation maps (Ichiye et al., 1991) of Cα atoms for all systems using CPPTRAJ (Roe et al., 2013). The results indicate only very weak correlations between the ectodomain and the TMD (Figure 2—figure supplements 10–13).”

      We added Figure 2—figure supplements 10–13.

      We have also noted in the Discussion section that enhanced sampling methods could be employed to better explore the conformational landscape of Env trimer, including fluctuations within the closed state as well as transitions among the closed ground, transient and CD4/co-receptor-stabilized states proposed in the previous experimental study (Munro et al., 2014).

      In the middle of the second paragraph in the Discussion section we added:

      “Enhanced sampling methods could be applied to more thoroughly explore the conformational landscape, including not only variations within the closed ground state but also transitions among the closed ground, transient and CD4-/co-receptor-stabilized states.”

      (2) Qualify strong independence claims: Rephrase or further support statements asserting independence of ectodomain and TMD motions, particularly in light of known experimental evidence for coupling (PMC4701381, PMC4304640, PMC5085267).

      In addition to adding the dynamical cross-correlation maps showing very weak correlations between the ectodomain and the TMD, we have added a discussion of the coupling between TMD, CT, and Env antigenicity, and the independent motion of ectodomain and TMD in our simulation.

      In the middle of the second paragraph in the Discussion section we added:

      “Our analysis of the ectodomain and TMD coupling indicates that the motions of these two domains are largely independent. This observation does not contradict experimental studies demonstrating functional coupling between the TMD, CT, and the antigenic profiles of Env (Chen et al., 2015; Dev et al., 2016). Munro et al. proposed that unliganded Env is intrinsically dynamic, transitioning among three distinct prefusion conformations: a closed ground state (predominant), a transient state, and a CD4-/co-receptor-stabilized state. Both laboratory-adapted and clinically isolated strains can spontaneously transition among these three states, although their relative occupancies differ (Munro et al., 2014). It is therefore possible that TMD mutations or CT truncation also alter the equilibrium distribution among three states, thereby affecting the epitope exposure, particularly for epitopes that are occluded in the closed ground state while exposed in the CD4-/co-receptor-stabilized state. However, transition among three states occur on millisecond-to-second timescales. Our simulations on microsecond timescales primarily capture conformational variations within the closed ground state and suggest that the MPER acts as a hinge, providing substantial flexibility that enables the ectodomain and TMD to move independently while Env remains in the closed ground state.”

      (3) Clarify membrane composition assumptions: Provide a clearer rationale for the chosen lipid composition, and explicitly discuss how differences between host plasma membranes and HIV virions (e.g., PS, sphingomyelin, and cholesterol enrichment) may affect the conclusions.

      We have clarified the limitations associated with the choice of lipid composition and emphasized the need to investigate its influence in future studies.

      At the end of the second paragraph in the Discussion section we added:

      “In addition to the limitations inherent to protein structure modeling, the choice of lipid composition remains an open question. In this work, we selected an asymmetric mammalian plasma membrane because it is one of the 18 complex biomembrane systems we previously studied (Pogozheva et al., 2022), and among them, it provides the closest available approximation to the HIV membrane. Nevertheless, experimental studies have reported differences in lipid composition between HIV virions and the host plasma membrane (Aloia et al., 1993; Brugger et al., 2006; Huarte et al., 2016; Mucksch et al., 2019; Tomishige et al., 2023). Although we do not anticipate that our main conclusions regarding Env domain motions and MPER flexibility would change substantially, evaluating the influence of lipid composition represents an important direction for future work.”

      (4) Address citation and reference issues: Replace PDB-only references with proper citations, correct mischaracterizations of structure determination methods, and ensure all supplementary citations are fully referenced.

      We have corrected the error in which PDB structure 7LOI was described as a cryo-EM structure; it is in fact an NMR structure. We have also verified that all PDB structures are properly cited at their first occurrence in the manuscript.

      (5) Expand the Methods section: Provide additional detail on system construction, glycan modeling, lipid asymmetry, equilibration, sampling, and limitations, including a discussion of potential benefits of enhanced-sampling approaches.

      We have clarified that the modeling of palmitoylation sites, glycans and lipid bilayers are done in an automated fashion by different modules in CHARMM-GUI, and added Supplementary file 1–Supplementary Table 8 showing the simulation settings for equilibration and production stages.

      At the end of the subsection “Modeling of full-length gp120–gp41 trimer” we have modified the following:

      “Two mutations (S764C and S837C) were introduced in the CT to restore the palmitoylation sites, and lipid tails oriented towards the hydrophobic core of the bilayer were then attached to the palmitoylation sites using the PDB Manipulation module in CHARMM-GUI (Jo et al., 2008; Jo et al., 2014; Park et al., 2023) (Figure 1D).”

      At the end of the subsection “Glycosylation” we added:

      “The select glycan sequences were represented in the Glycan Reader Sequence format (Jo et al., 2011; Park et al., 2017) and added to the corresponding glycosylation sites using the Glycan Reader & Modeler graphical interface.”

      In the middle of the subsection “Membrane” we added:

      “Membrane systems were constructed using CHARMM-GUI Membrane Builder, which provides a user-friendly graphical interface for selecting lipid types and defining their numbers in each leaflet (Jo et al., 2007; Jo et al., 2009; Wu et al., 2014; Lee et al., 2016; Lee et al., 2019).”

      In the middle of the subsection “Simulation details” we have modified the following:

      “Positional and dihedral restraints were applied to proteins, glycans, and lipids, with force constants progressively reduced over successive intervals (Supplementary file 1–Supplementary Table 8).”

      We added Supplementary file 1–Supplementary Table 8.

      The discussion of potential benefits of enhanced-sampling approaches is included in our response to major concern (1).

      (6) Data availability: In addition to code, deposit all MD trajectories for re-analysis. The scale of this simulation was likely costly (GPU time), and so data availability is imperative.

      We have deposit MD simulation trajectories to Zenodo.

      At the end of the section “Data availability” we added:

      “The simulation trajectories can be found at https://doi.org/10.5281/zenodo.18853902, https://doi.org/10.5281/zenodo.18854615, and https://doi.org/10.5281/zenodo.18854639.”

      Minor:

      (1) Stylistic: Suggested to revise Figure 1 to provide a clearer overview of all constructs with consistent nomenclature (e.g., "full-length" versus "ΔCT") and explicit domain boundaries. With a better overview figure, the current figures could comprise the Figure 1 associated with Figures 1 and 2.

      We have combined Figure 1 and Figure 1—figure supplement 1 to show both full-length and CT-truncated models in one figure.

      We have modified Figure 1.

      We have removed Figure 1—figure supplements 1.

      (2) Explicitly cross-validate against prior studies: Integrate comparisons to existing MD simulations and experimental studies (e.g., PMC2730987, PMC2980712, PMC4254001, PMC4040535, PMC6035291, PMC4701381, PMC5085267) directly into the Results and Discussion.

      We have added discussion of previous MD simulation studies to the Result section regarding interactions of the TMD residue R696 with ions and lipid headgroups.

      At the end of the first paragraph in the subsection “The energetically unfavorable R696 in the hydrophobic core results in asymmetric, kinked TMD conformations and disrupts membrane integrity” we have modified the following:

      “Previously, Kim et al. reported that the inter-chain interactions between protonated R696 gradually diminished over a short simulation time (23 ns), leading to increased crossing angles and reduced bundle length (Kim et al., 2009). Gangupomu et. al and Baker et. al observed that R696 snorkeled toward either exoplasmic or endoplasmic headgroups in simulations of the TMD monomer, resulting in TMD tilting and membrane thinning due to water penetration and lipid headgroups interacting with R696 (Gangupomu et al., 2010; Baker et al., 2014; Baker et al., 2014). These observations are consistent with our finding. Hollingsworth et. al also reported membrane thinning; however, they attributed this effect to interfacial interactions of R683 and R707 with both leaflets and proposed that R696 only interacted with water and ions permeating into the center of the TMD timer (Hollingsworth et al., 2018).”

      The discussion of PMC4701381 and PMC5085267 is included in our response to major concern (2).

      (3) "In the cryo-EM structure (PDB ID: 7LOI)": This is an NMR model and lacks citation.

      We have corrected this error and added the citation at the first occurrence of PDB ID: 7LOI in the Result section.

      In the middle of the first paragraph in the subsection “The energetically unfavorable R696 in the hydrophobic core results in asymmetric, kinked TMD conformations and disrupts membrane integrity” we have modified the following:

      “In the NMR structure (PDB ID: 7LOI) (Piai et al., 2021),”

      (4) "Higher RMSF values were observed in the residues missing from the cryo-EM structure": This is lacking citation, as there are multiple cryo-EM structures and several dynamics studies using NMR.

      The missing residues here specifically refer to those absent in the cryo-EM structure (PDB ID: 6B0N) used for model building, rather than all cryo-EM structures in the PDB. We have revised the text to clarify this distinction.

      In the middle of the second paragraph in the subsection “The ectodomain maintains a rigid internal structure and tilts independently of the TMD” we have modified th following:

      “Higher RMSF values were observed in the residues missing from the cryo-EM structure (PDB ID: 6B0N) (Sarkar et al., 2018), which was used for the ectodomain in model building (these missing residues are highlighted in red in Figure 1A, B),”

    1. eLife Assessment

      This study provides fundamental insights by demonstrating that the Nanog mRNA coding sequence (CDS) and 3′UTR domains are spatially segregated and functionally distinct in pluripotent stem cells and blastocysts, with 3′UTR-enriched border cells primarily influencing morphogenesis and CDS-enriched inner cells largely regulating transcription and epigenetic programs. The work opens a novel conceptual avenue for understanding how separable mRNA domains can differentially control cell behavior and differentiation. However, the evidence is incomplete, as key aspects of the molecular nature, biogenesis, and precise characterization of the separated 3′UTR and CDS RNA species, as well as causal links between their perturbation and the observed phenotypes (e.g., via rescue and deeper characterization of 3′UTR elements), remain to be fully established.

    2. Reviewer #1 (Public review):

      Summary:

      There is evidence that some genes encode mRNAs from which separate processed transcripts may arise, separating the coding sequence (CDS) from the 3'-UTR, and with both mRNA elements remaining stable in the cell. However, the functional consequences of these mRNA fragments have not been firmly established. In the manuscript by Yang et al., the authors probe the mRNA domain architecture of Nanog in the context of embryonic stem cell colonies and blastocysts. The authors detect spatial separation of Nanog CDS-containing mRNA from abundant Nanog 3'-UTR RNAs depending on the cell position in 2D embryonic stem cell colonies or in blastocysts.

      Strengths:

      The phenotypic analyses of the Nanog mRNA hold promise for revealing distinct roles for the Nanog encoded protein and a separate RNA encompassing the Nanog 3'-UTR.

      Weaknesses:

      There are a number of questions about the molecular nature of the mRNA species that the authors should address in order for the results to be firmly established, as noted below.

      (1) It is not clear how the authors verified that their probes are specific for Nanog CDS or 3'-UTR regions. Especially for the 3'-UTR probe, it is confusing why colonies show green only regions, suggesting only the CDS is present. I would expect the CDS and 3'-UTR probes to colocalize in the interior cells. Is it possible that the 3'-UTR probe is targeting another RNA?

      (2) It would help for the authors to include a graphic similar to Figure 3, Figure Supplement 1A, that diagrams the location of the CDS and 3'-UTR probes (this should also be done for Oct4 and Sox2). This graphic could also show all potential polyadenylation signals.

      (3) I think, based on the fluorescence patterns, there is evidence that the signal for the Nanog 3'-UTR probe is nuclear (images with DAPI staining), but this is not commented on that I could find. This should be discussed, as nuclear retention has implications for the noncoding function of the 3'-UTR fragment.

      (4) Figure 2, Figure Supplement 1A needs a better explanation. It's not clear how the reads map to the different regions of the Nanog mature mRNA. The authors should show examples at different ratios of CDS to 3'-UTR. Do the reads have a sharp boundary at the junction of where the isolated 3'-UTR is thought to occur?

      (5) I looked in the Zenbu browser at human NANOG CAGE mapping in the FANTOM5 dataset. I could not see evidence for substantial capping of a 3'-UTR fragment when filtering for embryonic cell types. Given the strong signal for the 3'-UTR in border cells, I would expect to see evidence for capping if the RNA were indeed capped. This suggests that if it exists, it is likely uncapped and (as noted in point 3) is likely nuclear retained.

      (6) Are there predicted polyadenylation signals near the end of the CDS that would generate a short 3'-UTR, and are these signals conserved across mammals?

      (7) It would help to see a zoomed-in view of the region targeted by one of the guide RNAs in the 3'-UTR, and where that site is relative to the polyadenylation signal. Is the polyadenylation signal upstream, i.e., CDS proximal?

      (8) A final note, the use of green and red together will be challenging for those who are colorblind. Providing a different false color palette would be helpful.

      I am refraining from comments on the cell biology and morphological insights, as they are remote from my core expertise.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript shows that the coding sequence (CDS) and 3' untranslated region (3'UTR) of mRNA transcripts from the Nanog gene have distinct expression patterns and functions. In both human and mouse embryonic stem cells colonies and blastocysts, these domains are spatially segregated, with 3'UTR-enriched cells occupying the borders and CDS-enriched cells residing in the interior. CDS mRNA expression is correlated with the expected regulation of transcription and epigenetics associated with the Nanog protein. Interestingly, expression of the 3'UTR appears to play an independent role in cell behavior and colony morphogenesis. Indeed, deletion of the 3'UTR causes specific defects in cell spreading and protrusive activity, with alteration in the localization of adhesion and cytoskeleton-associated proteins. Remarkably, a large proportion of those defects are rescued upon ROCK inhibition. Deletion of either Nanog CDS or 3'UTR leads to distinct modifications in the differentiation competence.

      Strengths:

      The independent role of 3'UTR mRNA domains, although identified in neurosciences a couple of years ago, is a novel and exciting field relatively unexplored in early development.

      The manuscript offers a multilayer series of experiments, in ES cells colony, blastocysts, and embryoid bodies, including imaging, -omics, genetic and pharmacological challenges, and differentiation experiments, thereby unveiling very convincingly the role of Nanog 3'UTR in morphogenesis.

      Weaknesses:

      The pathways leading to the generation of those distinct transcript domains are unknown. Although the functional differential roles are well demonstrated, whether the expression patterns are a cause or a consequence of the cells' localisation in the embryo remains to be explored.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, Yang et al reported distinct functions of the protein-coding sequence (CDS) and the 3' untranslated region (UTR) in the Nanog mRNA in pluripotent stem cells. They first observed different localization patterns for the CDS and 3' UTR in embryonic stem cells and in blastocyst embryos, and this pattern correlates with cell populations in different pluripotent states based on single-cell sequencing data. To characterize the potentially distinct functions of these regions, the authors generated knockout (KO) cell lines in which either the CDS or the 3' UTR was genetically ablated. These deletions led to different phenotypes in multiple assays. These results provided evidence that the CDS and 3' UTR of an mRNA could have distinct functions. Although these results are potentially interesting, several questions need to be addressed before the validity of their conclusion can be confirmed.

      Strengths:

      This study provides evidence for distinct functions of the protein-coding sequence and 3' untranslated region of an mRNA in pluripotent stem cells. The concept could be more broadly applied.

      Weaknesses:

      The initial observation (distinct localization of CDS and 3' UTRs) and the causal relationship between the KO and phenotype need further validation.

      Major points:

      (1) The authors showed distinct localization patterns of the CDS and 3' UTRs in human and mouse ESCs and blastocysts, and the overlap between their signals was minimal (Figure 1). Does this mean that the CDS and 3' UTR RNAs exist separately? For example, in cells that only showed signals for 3' UTRs, do these RNAs only contain 3' UTRs and lack CDS? Was this confirmed by RNA-seq experiments? If so, how are they generated (i.e., by transcription from a novel promoter or partial degradation of the full-length mRNAs)? This is a key question. Without a clear characterization of these RNAs, the rest of the study cannot be substantiated.

      (2) To confirm that the phenotypes of CDS or 3' UTR KO cells were caused by the deleted regions instead of other artifacts, rescue experiments should be performed.

      (3) As over-expression of the 3' UTR showed a phenotype, important regions within it should be identified, and also the possibility that the 3' UTR contains open reading frame(s) and is translated should be tested.

    5. Author response:

      eLife Assessment

      This study provides fundamental insights by demonstrating that the Nanog mRNA coding sequence (CDS) and 3′UTR domains are spatially segregated and functionally distinct in pluripotent stem cells and blastocysts, with 3′UTR-enriched border cells primarily influencing morphogenesis and CDS-enriched inner cells largely regulating transcription and epigenetic programs. The work opens a novel conceptual avenue for understanding how separable mRNA domains can differentially control cell behavior and differentiation. However, the evidence is incomplete, as key aspects of the molecular nature, biogenesis, and precise characterization of the separated 3′UTR and CDS RNA species, as well as causal links between their perturbation and the observed phenotypes (e.g., via rescue and deeper characterization of 3′UTR elements), remain to be fully established.

      We thank the editors and the three reviewers for their careful and constructive engagement with our manuscript. We greatly appreciate the reviewers’ recognition of the conceptual significance of the study and their thoughtful suggestions for strengthening the mechanistic and molecular characterization of the work. We have carefully considered all points raised and outline below the revisions planned for the revised manuscript.

      The phenomenon of differential CDS and 3’UTR expression is not unique to Nanog. Independent 3’UTR and CDS expression and differential CDS/3’UTR usage has been observed across multiple genes, tissues, and developmental contexts, including genome-wide (Mercer et al., 2011) and transcriptome scale studies (Kocabas et al., 2025, Ji et al., 2021). Prior studies have proposed that isolated 3’UTRs may arise through regulated RNA processing pathways coupled to exonucleolytic degradation and, in some cases, recapping mechanisms (Malka et al, 2017, Haberman et al., 2024). While the precise molecular mechanisms underlying isolated Nanog CDS and 3’UTR generation remain unresolved, our observations (contained here) support regulated RNA processing models. Our original submission included a brief discussion of this topic; however the revised manuscript will include substantially expanded analyses and discussion of the generation of isolated Nanog CDS and 3’UTR species.

      The revised manuscript will address the major concerns regarding:

      (1) The molecular nature, biogenesis, and precise characterization of the separated 3′UTR and CDS mRNA species

      (2) The causal relationship between perturbation of these RNA species and the observed phenotypes, including additional rescue experiments and deeper computational characterization of putative, functional 3′UTR elements.

      Specifically:

      (A) New supplementary analyses and schematics designed to further clarify the conceptual and mechanistic framework of the study, including:

      (i) Computational examination of the Nanog 3’UTR across all reading frames for open reading frames (ORFs).

      (ii) As suggested by Reviewers 1 and 3, single cell traces of Nanog mRNA expression from the full-length mESC dataset used in this study, illustrating distinct transcript isoforms and CDS/3’UTR expression patterns across individual cells, complementing the color-coded tSNE analyses currently presented in Fig. 2.

      (iii) Expanded schematic model and analyses addressing possible mechanisms underlying the generation of isolated Nanog CDS and 3’UTR enriched RNA species, including transcript architecture, predicted RNA structural barriers, and exonucleolytic processing models.

      (iv) Expanded discussion of the predominantly nuclear localization of the Nanog 3’UTR signal and its implications for transcript biogenesis, processing, and potential noncoding functions.

      (B) Correction of all minor labeling errors.

      (C) Additional experimental analyses, including:

      - Expansion of Nanog 3’UTR overexpression and rescue experiments to include cell spreading assays.

      - Expanded analysis of the effects of ROCK pathway inhibitors on colony morphology and cytoskeletal organization.

      - Examination of the ability of ROCK inhibition to restore normal embryoid body formation.

      Collectively, these planned revisions are intended to strengthen the mechanistic framing, molecular characterization, and broader significance of the study while clarifying the interpretation and scope of the conclusions.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      There is evidence that some genes encode mRNAs from which separate processed transcripts may arise, separating the coding sequence (CDS) from the 3'-UTR, and with both mRNA elements remaining stable in the cell. However, the functional consequences of these mRNA fragments have not been firmly established. In the manuscript by Yang et al., the authors probe the mRNA domain architecture of Nanog in the context of embryonic stem cell colonies and blastocysts. The authors detect spatial separation of Nanog CDS-containing mRNA from abundant Nanog 3'-UTR RNAs depending on the cell position in 2D embryonic stem cell colonies or in blastocysts.

      Strengths:

      The phenotypic analyses of the Nanog mRNA hold promise for revealing distinct roles for the Nanog encoded protein and a separate RNA encompassing the Nanog 3'-UTR.

      Weaknesses:

      There are a number of questions about the molecular nature of the mRNA species that the authors should address in order for the results to be firmly established, as noted below.

      (1) It is not clear how the authors verified that their probes are specific for Nanog CDS or 3'-UTR regions. Especially for the 3'-UTR probe, it is confusing why colonies show green only regions, suggesting only the CDS is present. I would expect the CDS and 3'-UTR probes to colocalize in the interior cells. Is it possible that the 3'-UTR probe is targeting another RNA?

      We thank the reviewer for raising the important question of probe specificity. We realize that the data that underlying this concern is the absence of colocalizing between CDS and 3’UTR probes in colony border cells.

      The absence of CDS/3’UTR colocalization in colony border cells is not due to probe failure but instead reflects the principal observation underlying the study. If Nanog CDS and 3’UTR sequences were present exclusively as intact full-length transcripts in a strict stoichiometric ratio, Nanog positive cells would be expected to be positive for both probes (appearing yellow). Instead, border cells exhibit strong 3’UTR signal with minimal or absent CDS signal, while adjacent interior cells show the opposite pattern.

      The fact that both probes robustly detect signal within the same sample but in spatially distinct cell populations, argues that both probes are functional and that the observed differential localization reflects genuine biological differences in levels of transcript components.

      The CDS probe targets ~300 bp within the coding region, while the 3’UTR probe targets ~300 bp within the proximal region of the Nanog 3’UTR. Hybridization specificity was validated as described in the Methods and in our previous studies (Kocabas et al 2015; Ji et al 2021), including negative controls. We additionally now provide a supplemental figure (New Figure 1-figure supplement 2A), highlighting that the Nanog 3’UTR and CDS probes label cell populations distinct from each other, further indicating their specificity.

      In addition, full-length scRNA seq datasets from both mouse and human ESCs demonstrate differential CDS/3’UTR expression patterns for Nanog and many other genes. To further clarify this point, the revised manuscript will include single cell transcript traces from mESCs illustrating the distinct Nanog isoforms detected across individual cells (New Figure 2-figure supplement 1A)

      (2) It would help for the authors to include a graphic similar to Figure 3, Figure Supplement 1A, that diagrams the location of the CDS and 3'-UTR probes (this should also be done for Oct4 and Sox2). This graphic could also show all potential polyadenylation signals.

      We agree that additional schematic clarification would improve readability. The revised manuscript will include schematics showing the locations of the CDS and 3’UTR probes for Nanog, Sox2 and Oct4 (New Fig. 1- figure supplement 1A).

      (3) I think, based on the fluorescence patterns, there is evidence that the signal for the Nanog 3'-UTR probe is nuclear (images with DAPI staining), but this is not commented on that I could find. This should be discussed, as nuclear retention has implications for the noncoding function of the 3'-UTR fragment.

      The reviewer is correct that the Nanog 3’UTR signal mostly nuclear. Whie this was noted in (the original) Figure 1-figure supplement 2A, we agree that it is possible that mechanistic and functional implications were not sufficiently discussed in the original manuscript. The revised manuscript will include expanded discussion of the relationship between nuclear localization transcript processing, and potential noncoding functions of isolated Nanog 3’UTR species

      (4) Figure 2, Figure Supplement 1A needs a better explanation. It's not clear how the reads map to the different regions of the Nanog mature mRNA. The authors should show examples at different ratios of CDS to 3'-UTR. Do the reads have a sharp boundary at the junction of where the isolated 3'-UTR is thought to occur?

      We thank the reviewer for this suggestion. The revised manuscript will include new single cell read maps across the Nanog locus from full length mESC scRNA-seq datasets (New Figure 2-figure supplement 1A), illustrating distinct CDS enriched and 3’UTR enriched transcript isoforms across individual cells.

      These analyses indicate that some CDS dominant transcripts contain 3’UTR sequence, while many appear to contain little or no detectable 3’UTR sequence. Conversely, many 3’UTR enriched transcripts contain only minimal or truncated CDS sequence. Importantly full CDS and 3’UTR mRNA components are frequently not present in a strict 1:1 ratio, either within individual cells, or across cell populations.

      The revised manuscript will also include expanded supplementary analyses integrating transcript architecture, predicted RNA structural barriers, polyadenylation analysis, and single cell coverage patterns to further examine possible mechanisms underlying the generation of isolated Nanog CDS and 3’UTR species (New Figure 2-figure supplement 1B,C).

      (5) I looked in the Zenbu browser at human NANOG CAGE mapping in the FANTOM5 dataset. I could not see evidence for substantial capping of a 3'-UTR fragment when filtering for embryonic cell types. Given the strong signal for the 3'-UTR in border cells, I would expect to see evidence for capping if the RNA were indeed capped. This suggests that if it exists, it is likely uncapped and (as noted in point 3) is likely nuclear retained.

      Prior studies have reported isolated uncapped and recapped 3’UTR species in multiple systems (Malka et al, 2017; Haberman et al, 2024). We agree that the predominantly nuclear localization and lack of a strong CAGE signal for Nanog are important observations and will expand discussion of these points in the revised manuscript.

      (6) Are there predicted polyadenylation signals near the end of the CDS that would generate a short 3'-UTR, and are these signals conserved across mammals?

      Computational analysis of the mouse Nanog 3'UTR identifies a single canonical PAS (AATAAA) at position 1074, located at the 3’ end of the annotated 3’UTR and this terminal PAS is conserved across mammals. These analyses will be included as a supplementary figure and discussed further in the revised manuscript section addressing Nanog transcript biogenesis.

      (7) It would help to see a zoomed-in view of the region targeted by one of the guide RNAs in the 3'-UTR, and where that site is relative to the polyadenylation signal. Is the polyadenylation signal upstream, i.e., CDS proximal?

      This will be provided in the revised manuscript (New Figure 2-figure supplement 1C,i) Two guide RNAs were used to generate the Nanog 3’UTR deletions. The downstream guide is upstream of the terminal polyadenylation signal at nt 1074 to preserve polyadenylation of the remaining Nanog CDS containing transcript.

      Consistent with this, all Nanog 3’UTR knockout lines retain normal Nanog protein levels. The revised manuscript will include supplementary schematics showing guide RNA positions relative to the CDS, 3’UTR probes, and terminal PAS.

      (8) A final note, the use of green and red together will be challenging for those who are colorblind. Providing a different false color palette would be helpful. 

      We appreciate this attention to accessibly. The red/green color combination was chosen to provide the highest contrast between CDS and 3’UTR signals in the in situ hybridization experiments, which is important for visualizing their differential spatial localization. We will ensure that figure legends clearly indicate channel assignments throughout the manuscript.

      I am refraining from comments on the cell biology and morphological insights, as they are remote from my core expertise.

      Reviewer #2 (Public review):

      Summary:

      This manuscript shows that the coding sequence (CDS) and 3' untranslated region (3'UTR) of mRNA transcripts from the Nanog gene have distinct expression patterns and functions. In both human and mouse embryonic stem cells colonies and blastocysts, these domains are spatially segregated, with 3'UTR-enriched cells occupying the borders and CDS-enriched cells residing in the interior. CDS mRNA expression is correlated with the expected regulation of transcription and epigenetics associated with the Nanog protein. Interestingly, expression of the 3'UTR appears to play an independent role in cell behavior and colony morphogenesis. Indeed, deletion of the 3'UTR causes specific defects in cell spreading and protrusive activity, with alteration in the localization of adhesion and cytoskeleton-associated proteins. Remarkably, a large proportion of those defects are rescued upon ROCK inhibition. Deletion of either Nanog CDS or 3'UTR leads to distinct modifications in the differentiation competence.

      Strengths:

      The independent role of 3'UTR mRNA domains, although identified in neurosciences a couple of years ago, is a novel and exciting field relatively unexplored in early development.

      The manuscript offers a multilayer series of experiments, in ES cells colony, blastocysts, and embryoid bodies, including imaging, -omics, genetic and pharmacological challenges, and differentiation experiments, thereby unveiling very convincingly the role of Nanog 3'UTR in morphogenesis.

      Weaknesses:

      The pathways leading to the generation of those distinct transcript domains are unknown. Although the functional differential roles are well demonstrated whether the expression patterns are a cause or a consequence of the cells' localization in the embryo remains to be explored.

      We thank the reviewer for these thoughtful comments and for recognizing the potential significance of independent 3’UTR functions in early developmental systems.

      Regarding the mechanisms underlying generation of distinct CDS and 3’UTR transcript domains, the revised manuscript will include new supplementary analyses and schematic models addressing possible Nanog transcript processing pathways, as outlined above.

      We agree that the relation between spatial location and Nanog 3’UTR expression is an important question. Specifically, it remains unclear whether cells first acquire high Nanog 3’UTR expression and subsequently localize to the colony border or whether border position itself promotes high Nanog 3’UTR expression.

      Our current data suggest that both processes may contribute. Deletion of the Nanog 3’UTR does not prevent colonies from establishing border/interior pattern, indicating that high Nanog 3’UTR is not strictly required for border pattern itself. At the same time, Nanog 3’UTR overexpression and rescue experiments increased the likelihood of border localization, suggesting that elevated Nanog 3’UTR expression promotes behaviors associated with border occupancy.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Yang et al reported distinct functions of the protein-coding sequence (CDS) and the 3' untranslated region (UTR) in the Nanog mRNA in pluripotent stem cells. They first observed different localization patterns for the CDS and 3' UTR in embryonic stem cells and in blastocyst embryos, and this pattern correlates with cell populations in different pluripotent states based on single-cell sequencing data. To characterize the potentially distinct functions of these regions, the authors generated knockout (KO) cell lines in which either the CDS or the 3' UTR was genetically ablated. These deletions led to different phenotypes in multiple assays. These results provided evidence that the CDS and 3' UTR of an mRNA could have distinct functions. Although these results are potentially interesting, several questions need to be addressed before the validity of their conclusion can be confirmed.

      Strengths:

      This study provides evidence for distinct functions of the protein-coding sequence and 3' untranslated region of an mRNA in pluripotent stem cells. The concept could be more broadly applied.

      Weaknesses:

      The initial observation (distinct localization of CDS and 3' UTRs) and the causal relationship between the KO and phenotype need further validation.

      Major points:

      (1) The authors showed distinct localization patterns of the CDS and 3' UTRs in human and mouse ESCs and blastocysts, and the overlap between their signals was minimal (Figure 1). Does this mean that the CDS and 3' UTR RNAs exist separately? For example, in cells that only showed signals for 3' UTRs, do these RNAs only contain 3' UTRs and lack CDS? Was this confirmed by RNA-seq experiments? If so, how are they generated (i.e., by transcription from a novel promoter or partial degradation of the full-length mRNAs)? This is a key question. Without a clear characterization of these RNAs, the rest of the study cannot be substantiated.

      We thank the reviewer for raising this important question, which overlaps substantially with several key points raised by Reviewer #1 concerning the molecular nature and characterization of the Nanog CDS and 3’UTR species.

      Colony border cells exhibit strong Nanog 3’UTR signal with minimal detectable CDS signal, while adjacent interior cells show the reciprocal pattern. These observations strongly suggest the existence of distinct Nanog transcript species rather than exclusively full-length transcripts containing stoichiometric amounts of both CDS and 3’UTR sequence.

      This conclusion is independently supported by full-length Smart-seq2 scRNA seq datasets from both mouse and human ESCs, which provide transcript coverage across both CDS and 3’UTR regions.

      (2) To confirm that the phenotypes of CDS or 3' UTR KO cells were caused by the deleted regions instead of other artifacts, rescue experiments should be performed.

      Rescue experiments were included in the original submission (Fig. 4). The revised manuscript will expand these analyses to include cell spreading. We will also include additional ROCK pathway modulation experiments.

      (3) As over-expression of the 3' UTR showed a phenotype, important regions within it should be identified, and also the possibility that the 3' UTR contains open reading frame(s) and is translated should be tested.

      The revised manuscript will also include supplementary computational analyses of the Nanog 3’UTR, including open reading frame prediction, Kozak scoring, and evolutionary conservation analysis. (New Figure 2-figure supplement 1B). These analyses identify no evidence for strongly supported coding potential within the 3’UTR. Further, isolated Nanog 3’UTR transcripts are largely confined to the nucleus, making active translation unlikely.

      The revised manuscript will include new supplementary analyses addressing Nanog transcript structure and possible biogenesis mechanisms (New Figure 2-figure supplement 1C).

      References:

      ViennaRNA/RNA fold – Lorenz et al 2011 Algorithms Mol Biol 6:26- RNA Secondary Structure stem loop, minimum free energy (MFE) prediction

      NCBI BLASTP- Altschul et al (1990) J Mol Biol 215:403- ORF conservation, protein sequence similarity search

      NCBI Entrez/Biohthon- Cock et al (2009) Bioinformatics 25:1422- sequence retrieval

      PhastCons/UCSC multiz alignments- Siepel et al (2005) Genome Res 15:1034- evolutionary conservation scoring

      UCSC Genome Browser- Kent et al. (2002) Genome Res 12:996-1006- conservation track access

      Eaton et al (2020) Mol Cell 78:439- Stall model

      Brannan et al (2012) Genes Dev 26:2621-Stall model

      Addition to Methods.

      ORFs (≥10 amino acids) were identified in all three forward frames according to Kozak (1987). Evolutionary conservation was assessed by BLASTP (Altschul et al., 1990) against RefSeq proteins. Poly(A) signals were identified by pattern matching for canonical and non-canonical hexamers. Conserved sequence blocks were obtained from UCSC PhastCons tracks (Siepel et al., 2005). RNA secondary structures were predicted using ViennaRNA RNAfold (Lorenz et al., 2011) with a sliding 80-nt window. The stall model for isolated transcript generation follows Eaton et al. (2020).

    1. eLife Assessment

      There is a need for better and safer dengue virus live attenuated vaccines. This manuscript describes important findings that could lead to the design of a strongly immunogenic, tetravalent live attenuated vaccine for dengue, without the risk of causing antibody-dependent enhancement. However, the experimental evidence presented is incomplete since only constructions of one serotype were tested to prove the principle.

    2. Reviewer #1 (Public review):

      Summary:

      Dalben et al. grafted the fusion loop mature (FLM) modification, based on a previously reported D2-FLM, to another serotype DENV4, and adapted them to replicate in Vero cells for live attenuated vaccine (LAV) manufacturing while retaining favorable antigenic profiles, generating two new strains: D2-vFLM and D4-vFLM. Deep sequencing revealed adapted mutations at the junction of envelope domains I and II (EDI and EDII), and both D2-vFLM and D4-vFLM showed no evidence of ADE in the presence of FL-targeting Abs. Sera from D2-vFLM immunized mice displayed strong homotypic and reduced heterotypic neutralization compared to wild-type viruses, with minimal to no ADE potential in vitro. Moreover, D2-vFLM immunization completely protected AG129 mice from lethal challenge with mouse-adapted D220. They demonstrate that the FLM modification platform is transferable across serotypes and yields strains with favorable immunogenicity and reduced ADE risk. The FLM approach provides a promising path toward the development of a safer tetravalent DENV LAV.

      Strengths:

      The authors carried out a series of experiments to generate and characterize two new strains (D2-vFLM and D4-vFLM) of FLM-modified viruses, and showed their antigenic and immunogenic profiles. The observation that the FLM modification platform is transferable across serotypes and yields strains with favorable immunogenicity and reduced ADE risk is interesting.

      Weaknesses:

      However, one concern is the total number of mutations (including originally introduced and compensatory mutations) in this FLM vaccine platform, and it is not clear regarding the future directions for the proof-of-concept vaccine in this study.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, YR Dalben et al describe the generation of DENV2 and DENV4 strains with mutations in the fusion loop (FL) of the E protein and pre-membrane (prM) protein to limit potential antibody-dependent enhancement (ADE) resulting from vaccination with live-attenuated vaccines and adapted these strains for growth in Vero cells. They show that the DENV2 version D2-vFLM is immunogenic and generates neutralizing serum against DENV2 and DENV4 after 2 boosts and is protective against lethal challenge. Serum from D2-vFLM also showed no ADE against DENV4.

      Strengths:

      Overall, the paper is well written and presented, and the data presented support most of the conclusions made. Grafting D2-FLM mutations to DENV4 and adapting both to growth in Vero cells is a good step to show that this method could be used to generate production-level LAV. The growth and stability data are clear and well-conducted.

      Weaknesses:

      However, there are several weaknesses, mostly in regard to the immunogenicity data, that limit the overall impact. The FLM mutations were only grafted to DENV4 but not to the other Dengue serotypes. The authors acknowledge that this is a proof-of-concept, but generating mutants of the other serotypes would strengthen the idea that this could be used to develop a tetravalent LAV. Immunizations in mice were only performed for D2-vFLM but not D4-vFLM. Immunogenicity data for D4-vFLM would strengthen this work if it shows that it can be immunogenic, protective, and limit ADE, as is shown for D2-vFLM. ADE from D2-vFLM was only tested against DENV4; does it also limit ADE from the other serotypes? This would better show that these mutations do limit ADE across serotypes and not just a single one.

      Additionally, some of the immunization data likely need to be repeated:

      The authors should describe why they pooled the sera from the mice and whether they purified total IgG or not (Figure 5). They should also probably repeat the challenge experiment since it was 4 mice (D2) against 5 (D2-vFLM), and it is unclear if there is a statistical difference between the results obtained. It is not even mentioned in the Results section (D2 result vs D2-FLM), and thus unclear if using D2-FLM is an improvement in the way the data is currently presented.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dalben et al. grafted the fusion loop mature (FLM) modification, based on a previously reported D2-FLM, to another serotype DENV4, and adapted them to replicate in Vero cells for live attenuated vaccine (LAV) manufacturing while retaining favorable antigenic profiles, generating two new strains: D2-vFLM and D4-vFLM. Deep sequencing revealed adapted mutations at the junction of envelope domains I and II (EDI and EDII), and both D2-vFLM and D4-vFLM showed no evidence of ADE in the presence of FL-targeting Abs. Sera from D2-vFLM immunized mice displayed strong homotypic and reduced heterotypic neutralization compared to wild-type viruses, with minimal to no ADE potential in vitro. Moreover, D2-vFLM immunization completely protected AG129 mice from lethal challenge with mouse-adapted D220. They demonstrate that the FLM modification platform is transferable across serotypes and yields strains with favorable immunogenicity and reduced ADE risk. The FLM approach provides a promising path toward the development of a safer tetravalent DENV LAV.

      Strengths:

      The authors carried out a series of experiments to generate and characterize two new strains (D2-vFLM and D4-vFLM) of FLM-modified viruses, and showed their antigenic and immunogenic profiles. The observation that the FLM modification platform is transferable across serotypes and yields strains with favorable immunogenicity and reduced ADE risk is interesting.

      We thank reviewer 1 for the encouraging comments for our work.

      Weaknesses:

      However, one concern is the total number of mutations (including originally introduced and compensatory mutations) in this FLM vaccine platform, and it is not clear regarding the future directions for the proof-of-concept vaccine in this study.

      Author response table 1.

      We summarize the mutations in the FLM platform below.

      The maturation mutations are located at the furin cleavage site, which is buried within the membrane or virion. As a result, only five mutations are surface exposed, two of which are in the fusion loop region targeted for removal. Therefore, for a proof-of-concept study, the total number of mutations remains well within the genetic diversity observed among DENV genotypes.

      Compensatory mutations may affect overall DENV antigenicity. Notably, one such mutation, K204R, has been reported to alter antigenicity and could contribute to the improved safety profile of the vaccine. However, we have also shown that multiple adaptive pathways can support Vero cell adaptation, and our data indicate that K204R is not absolutely required for this process.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, YR Dalben et al describe the generation of DENV2 and DENV4 strains with mutations in the fusion loop (FL) of the E protein and pre-membrane (prM) protein to limit potential antibody-dependent enhancement (ADE) resulting from vaccination with live-attenuated vaccines and adapted these strains for growth in Vero cells. They show that the DENV2 version D2-vFLM is immunogenic and generates neutralizing serum against DENV2 and DENV4 after 2 boosts and is protective against lethal challenge. Serum from D2-vFLM also showed no ADE against DENV4.

      Strengths:

      Overall, the paper is well written and presented, and the data presented support most of the conclusions made. Grafting D2-FLM mutations to DENV4 and adapting both to growth in Vero cells is a good step to show that this method could be used to generate production-level LAV. The growth and stability data are clear and well-conducted.

      We thank reviewer 2 for the encouraging comments for our work.

      Weaknesses:

      However, there are several weaknesses, mostly in regard to the immunogenicity data, that limit the overall impact. The FLM mutations were only grafted to DENV4 but not to the other Dengue serotypes. The authors acknowledge that this is a proof-of-concept, but generating mutants of the other serotypes would strengthen the idea that this could be used to develop a tetravalent LAV.

      We selected DENV2 and DENV4 because they are the most genetically divergent. Currently, our data should support the FLM mutations that can be grafted on both DENV2 and DENV4, likely extend to their corresponding genotypes. We agree that the FLM mutations should be evaluated in additional serotypes. We also have promising preliminary data for FLM mutation grafting in DENV1 and are currently applying the same approach to DENV3. We hope to include these results, whether positive or negative, in the revised manuscript.

      Immunizations in mice were only performed for D2-vFLM but not D4-vFLM. Immunogenicity data for D4-vFLM would strengthen this work if it shows that it can be immunogenic, protective, and limit ADE, as is shown for D2-vFLM.

      We are currently immunizing AG129 mice with DV4 and D4-vFLM, followed by heterotypic challenge with D220. Because DENV vaccine-related hospitalization in clinical trials typically occurs 3 - 4 years after vaccination, we are cautious about whether this experimental design will fully capture the added safety benefit of the FLM mutations. We are also developing a passive immunization model in AG129 mice using diluted DENV4 serum to better mimic long-term waning antibody titers. We will include the future findings in the revised manuscript.

      ADE from D2-vFLM was only tested against DENV4; does it also limit ADE from the other serotypes? This would better show that these mutations do limit ADE across serotypes and not just a single one.

      We are trying to keep the scope of the paper within DENV2 and DENV4, however, we will perform ADE and neutralization assays for all four serotypes in the revised manuscript.

      Additionally, some of the immunization data likely need to be repeated:

      The authors should describe why they pooled the sera from the mice and whether they purified total IgG or not (Figure 5).

      We used pooled serum, consisting of equal volumes from each mouse, rather than purified IgG. In Figure 5, our goal was to show the overall increase in serum titer after each immunization using cheek-bleed samples from individual animals. Because the available sample volume was limited, we pooled the sera for this analysis. We also measured end-point serum titers for each individual animal.

      They should also probably repeat the challenge experiment since it was 4 mice (D2) against 5 (D2-vFLM), and it is unclear if there is a statistical difference between the results obtained. It is not even mentioned in the Results section (D2 result vs D2-FLM), and thus unclear if using D2-FLM is an improvement in the way the data is currently presented.

      This experiment was designed to determine whether D2-vFLM protects AG129 mice against homotypic challenge as effectively as DV2-WT. Although the sample size was small, the results support our conclusion. However, we agree with the reviewer that the study should include more animals, and we will increase the group size to n > 8 to 10 in the revised experiment.

    1. eLife Assessment

      The authors propose a "simplified" model for intrinsically bursting neurons with explicitly controllable parameterization of oscillatory dynamics. The evidence that the modeling approach is generally appropriate and practical for modeling rhythmic bursting neurons and neural circuits is currently incomplete. Based on what the authors present, this model appears to have limited neurobiological relevance and utility but may be useful as a controller for an artificial system, such as in neuro-robotics applications.

    2. Reviewer #1 (Public review):

      Summary:

      The authors present a simplified neural bursting model with explicitly controllable parameterization of oscillator dynamics designed for neural circuit modeling involved in rhythm generation.

      Strengths:

      (1) The purpose of the model and applied abstractions are well articulated and justified (2D model, independent parameter control).

      (2) Explicit control of burst duration, inter-burst interval, amplitude, resetting-behavior/entrainment. This allows modelers to focus on circuit interactions and is especially useful when details of intrinsic currents and bursting mechanisms are unknown. One could even imagine a scenario where this model would help identify predictions on key underlying burst generation mechanisms.

      (3) The model is well described and validated with simulations and comparisons to the base model and one alternative model.

      (4) Circuit-level validation is convincing, as it reproduces not only trivial examples.

      (5) The underlying mechanism in phase space is well reasoned and justified, extends previous work, e.g., by McKean, by improving usability.

      Weaknesses:

      (1) The paper heavily relies on numerical demonstrations but does not provide a formal analysis of stability, bifurcations, or entrainment. While appropriate for the intended purposes, a more formal footing could strengthen the model.

      (2) Lots of nice demonstrations are shown, but it is less clear how model parameterization was chosen, how behavior depends on parameterization, and in what parameter ranges certain behavior can be expected. A more detailed description of parameterization/exploration of parameter space would greatly benefit anyone using this model in the future.

      (3) Some claims on reproduction of prior locomotor CPG model and production of "more biologically realistic activity" by the presented model are overstated. The key feature of the locomotor CPG models cited was that they not only reproduced speed-dependent gait expression of intact mice, but also changes of gait expression after silencing/removal of specific commissural and long propriospinal interneurons (e.g., selective loss of trot after deleting of V0V; changes in gait expression and step-to-step variability after silencing of descending long-propriospinal neurons or ascending V3 LPNs). While likely (at least partially) feasible with the model formulation, the correspondence of these silencing/ablation of neuron classes has not been shown by the model. Importantly, though, it appears that authors didn't show how the model in general behaves under the influence of noise, which is key to reproducing LPN silencing.

    3. Reviewer #2 (Public review):

      Summary:

      The authors propose a reduced model for intrinsically bursting neurons. The model simply consists of exponential decay of an adaptation variable in a phenomenological silent phase, an exponential growth of that variable in an active phase, and imposed thresholds for jumps between these phases, with some add-ons to allow for effects such as input-dependence.

      Strengths:

      The model could be used as a controller for an artificial system that needs to switch between on and off states with separate control of state durations. It has some flexibility to allow for variable levels of the activity variable during the active phase. The authors show that the model can be tuned to capture phase response properties of neurons and patterns generated by small networks of neurons.

      Weaknesses:

      The proposed approach lacks biological relevance, practicality, and originality.

      (1) Biological relevance:

      Central pattern generators and other bursting neurons use specific physical principles to generate their bursts of activity. These principles place constraints on the tuning of these bursts, including relationships between active and silent phase durations and other properties. By discarding these relationships, the proposed model risks losing key constraints that affect performance in biologically relevant scenarios. The proposed model does not allow for the emergence of interesting dynamical phenomena, which occur naturally in neurons and neuronal networks.

      It is also important to note that spikes within bursts can be important and of interest. Biophysical models allow for easy extension to include spikes via fast sodium and potassium currents. The proposed model does not allow for such extensibility.

      Finally, as shown in the seminal early-2000s work of Izhikevich, building on fast-slow decomposition work by Rinzel and others, there is a wide variety of possible neuronal bursting patterns. At the very least, several of these have been observed in neuronal recordings. The authors' model is specific to square-wave bursting.

      (2) Practicality:

      The model makes use of various cut-off functions and other aspects that are implemented as rules. Combining rules with differential equations makes for an awkward modeling framework that is inconvenient to implement, conceptualize, and analyze (e.g., from a bifurcation perspective). Moreover, the authors add more and more adjustments to their basic framework to capture additional features, but these add-ons simply make the model more, and unnecessarily, complicated and awkward. It's worth noting that the authors argue for their model based on the idea that more biophysical models are difficult to tune, yet they compare their model to a biophysical one that they were able to tune to achieve the various patterns that they study. They do not give any indication of how easy or hard it was to tune their own model, nor do they compare simulation times between the two models. I do note that the biophysical model seems to have 22 parameters, whereas the simplified one has 21 in Table 2, which is essentially the same number. Finally, although the authors give some extensions of the model to match observed data, their model does not seem useful for predicting performance in never-before-tested scenarios.

      (3) Originality:

      As the authors note, the use of low-dimensional, specifically planar, neural models dates back to early authors such as FitzHugh and Nagumo. What the authors fail to acknowledge is that Rinzel, Terman, Kopell, and others did seminal work on neuronal activity, including phenomena such as post-inhibitory rebound and fast threshold modulation, using a relaxation oscillation framework, starting several decades ago. Their work included applications to central pattern generators (e.g., see Terman and collaborators on respiratory CPGs). It is astonishing that the authors don't seem to be aware of this work and do not mention it at all. Moreover, I don't see any advantage of the proposed framework over the earlier relaxation oscillator setting, where many important mechanistic principles have already been analyzed, including extensions to networks. On a related note, even through they propose a piecewise linear model, the authors do not cite the substantial existing work on piecewise linear models (e.g., Hahnloser, Neural Networks, 1998, for an early example; 2024 SIAM Review article by Coombes et al and references therein for much more) including work specifically on bursting, nor do they cite various other previous efforts to capture bursting with simplified models including work on piecewise linear maps by Aguirre et al.

    4. Reviewer #3 (Public review):

      This computational modeling study introduces the methodology of replacing bursting neurons in a model circuit with a simplified piecewise-linear model with an "active" and a "quiet" state representing, respectively, the burst of spikes and the inter-burst interval. The shape of the active state loosely represents the intra-burst firing rate. Because (piecewise) linear systems are explicitly solvable, the transitions from quiet to active and vice versa can be calculated explicitly to match exactly what a biophysically realistic model or a biological neuron does in different conditions. The base piecewise-linear model is built to represent a 2D biophysical neuron with a cubic v-nullcline. The simplicity of the model allows for matching the kinetics of more complex models with a tractable simplified set of equations, as exemplified by approximations of burst duration and amplitude, phase-response curves, entrainment, and, finally, mimicking the activities of two CPG circuit models using this simplified representation.

      Major comments

      (1) The use of piecewise linear approximations to explicitly estimate properties of biophysical neurons is a well-known and common technique. This study adds nothing to the technique in terms of novelty.

      (2) Although the model explicitly matches active and inactive durations of a circuit neuron, the dynamics are explicitly "clamped" by the user because the reduced model parameters explicitly depend on the input. There are cases where this is useful, for example, when we are interested in the dynamics of _other_ neurons (B, C, D, ...) within the context of activity, and we "clamp" the dynamics of neuron A. One should note that this is no better than having a look-up table. Effectively, to give a comparison, it is like using a sine wave to represent a pacemaker neuron and explicitly define its frequency at different input levels so that it responds "dynamically". However, the neuron is restricted to what the user puts in, and therefore, calling it a dynamical system is entirely wrong. I am afraid that the use of this crude tool is not described well enough in the manuscript to warn a naïve user not to fall for this trap.

      (3) The phase resetting curves are used incorrectly. PRCs are useful when the perturbation is weak (soft), which would demonstrate the nature of the vector field near the limit cycle and therefore inform us of the nature of its stability or instability. A hard PRC would always reset the cycle to the fixed offset from the perturbation phase and is therefore uninformative in understanding dynamics. (It is, however, useful experimentally in identifying which neurons are part of the CPG.) The authors clearly know that the dynamics of the system away from the limit cycle do not conserve those of a biophysical neuron. So what is the point?

      (4) I work on the STG, one of the systems exemplified here. Even in the small and relatively regular CPGs of the STG, the definition of the active and quiet parts of a burst is often less clear than what the authors suggest. Bursting neurons often do multiple bursts in a cycle, and therefore, substituting the burst envelope is a subjective matter. This is even more problematic in bursting neurons in the brain, where there is often no quiet period. This should be discussed.

    5. Author response:

      We thank the editors and reviewers for their time and feedback. We are encouraged by the feedback that the purpose and abstractions of the model are well articulated and justified, that the explicit control of bursting characteristics is useful, and that the circuit-level validations are convincing.

      Before responding to individual reviewer comments, we would like to address the framing in the current assessment that the model "appears to have limited neurobiological relevance and utility but may be useful as a controller for an artificial system, such as in neuro-robotics applications." We respectfully suggest that this framing understates the model's relevance to neuroscience. Specifically, a growing body of literature aims to understand biological motor control by building embodied simulations. Yet, these simulations either use overly simple artificial neural network (ANN) units without dynamics or computationally intensive biophysical ones that are difficult to train. Our model is not intended as a biophysical account of how individual neurons generate bursts at the level of ionic mechanisms or spikes that goal is already well served by the conductance-based and reduced biophysical models we cite. Rather, its contribution is to make intrinsic bursting dynamics readily incorporable into neural circuit models that can be used in complex settings, with parameters that map directly onto quantities that circuit-level neuroscience most often measures and tunes in models (burst duration, duty cycle, amplitude, shape, input dependence). Indeed, Reviewer #1 notes that: "The purpose of the model and applied abstractions are well articulated and justified [...] This allows modelers to focus on circuit interactions and is especially useful when details of intrinsic currents and bursting mechanisms are unknown. One could even imagine a scenario where this model would help identify predictions on key underlying burst generation mechanisms."

      We see our work as a neuroscience contribution as much as a neuro-robotics one. Bringing tractable, controllable bursting into this regime allows circuit modelers to study how intrinsic bursting interacts with circuit connectivity without committing to specific biophysical mechanisms, and it lets ANN-style models incorporate a class of dynamics that is biologically pervasive but currently underrepresented. We validated the model against two well-studied biological CPGs (the crustacean pyloric circuit and the mammalian locomotor circuit) precisely because the target use case is biological circuit modeling.

      While we remain committed to the belief that bringing bio-inspired neurons with interpretable intrinsic dynamics into ANN-style modeling of biological control systems is a useful contribution as an eLife Methods paper, the reviews have made clear that we have not situated our work clearly enough within the literature. In revision, we will sharpen this positioning in the Introduction and Discussion, and better situate the model relative to both the long tradition of non-spiking relaxation-oscillator and piecewise-linear modeling in neuroscience and also to current trends in simulated control.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Formal analysis

      The paper heavily relies on numerical demonstrations but does not provide a formal analysis of stability, bifurcations, or entrainment. While appropriate for the intended purposes, a more formal footing could strengthen the model.

      We agree that a formal dynamical-systems treatment would deepen the work, and we appreciate the reviewer's acknowledgment that the numerical-only approach may nevertheless be appropriate for the intended purposes. Because the model is hybrid (continuous dynamics combined with discrete switching rules), a full formal analysis is non-trivial, and we view it as a substantial follow-up rather than something to fold into the present manuscript. In revision, we will discuss more explicitly the opportunities such formal analysis presents.

      (2) Parameter tuning and parameter-space characterization

      It is less clear how model parameterization was chosen, how behavior depends on parameterization, and in what parameter ranges certain behavior can be expected.

      We agree that this would substantially improve usability, and we will expand this aspect of the paper. The revision will include: (a) more details describing how parameters maps onto observable features of the bursting waveform, (b) recommended parameter ranges and the qualitative behaviors expected at their boundaries, and (c) practical guidance for tuning the model to match observations or embed into circuits.

      (3) Locomotor CPG interneuron ablation and noise

      The correspondence of these silencing/ablation of neuron classes has not been shown by the model. Importantly, though, it appears that authors didn't show how the model in general behaves under the influence of noise.

      The reviewer is right that the cited work establishes validity of the circuit model in large part through silencing/ablation experiments, and we did not reproduce those experiments. We understand those gait expression phenomena to be arising from non-bursting interneuron activations and a robust solution found for connection weights between them. The half-center bursting neurons only see a time-varying input signal, and their response is well-characterized by the constant, pulse, and periodic analyses we perform. As such, we chose to reproduce a few key experiments to retain a focus on our simplified neuron model. We will rephrase the relevant passages to make this scope explicit and ensure that our reproduction claims are appropriately stated. We will also expand on how the model interfaces with noise together with the proposed parameter-space characterization.

      Reviewer #2 (Public review):

      (1) Biological relevance

      Central pattern generators and other bursting neurons use specific physical principles to generate their bursts of activity. These principles place constraints on the tuning of these bursts, including relationships between active and silent phase durations and other properties. By discarding these relationships, the proposed model risks losing key constraints that affect performance in biologically relevant scenarios.

      We agree that biophysical models impose constraints that arise from underlying mechanisms. For instance, as input alters the curved shape of nullcline-v in Figure 1, the active/quite phase durations and duty cycle change in constrained ways. The question seems to be if our model is too flexible for instance, making it too easy to achieve desired phase durations, duty cycles, and other input-dependent responses. We see this as a valuable feature of our model, not a bug. Firstly, even if our model may be expressive enough to achieve a variety of response profiles (as in Figure 3—figure supplement 3), the careful modeler will ensure matching to experimental observations. Moreover, in many circuit systems, the relevant biophysical details are often unknown for the specific neurons being modeled as noted by Reviewer #1, and the modelers' primary goal is to reproduce circuit-level activity. Such can be achieved easily with a simplified model, and also with a biophysical model as data becomes available. Finally, we should note that modelers can and do tune the parameters of biophysical models within determined ranges in order to achieve desired phase durations and duty cycles, relaxing constraints somewhat in order to reproduce appropriate activity.

      It is also important to note that spikes within bursts can be important and of interest. [...] The authors' model is specific to square-wave bursting.

      We agree that spikes are important and interesting in many settings, and we believe that biophysical models would be most appropriate in these cases. In many cases, too, some abstraction and simplification is desirable, and this would not necessarily detract from the model's biological relevance. As we discuss in our high-level comments, we aim to bring intrinsic bursting dynamics into the ANN-style modeling regime that typically neglects intrinsic dynamics altogether. While the simplified model may be limited in some ways, it is nevertheless useful for many common biologically relevant scenarios, as validated by our circuit experiments. Finally, we would note that many of the raised limitations (no intra-burst spike structure, restricted bursting class, abstracted constraints) are shared by the relaxation-oscillator and piecewise-linear traditions that the reviewer cites approvingly, which suggests that our model lies along a familiar abstraction continuum rather than outside it. In revision, we will explicitly acknowledge that the model captures a basic/regular form of bursting within a broader taxonomy, and clarify the conditions under which abstracting the biophysical constraints is appropriate.

      (2) Practicality

      The model makes use of various cut-off functions and other aspects that are implemented as rules. Combining rules with differential equations makes for an awkward modeling framework

      On the modeling framework, we would defend the hybrid formulation (rules + ODE) as our aim is to prioritize usability by modelers, not the simplicity or elegance of equations. While a "pure-ODE" Fitzhugh-Nagumo-style polynomial may seem simple and elegant—with dv/dt = av^3 + bv^2 + cv + d and a, b, c, d parameters as the reviewer has pointed out a lot of complexity can arise from this. Tuning these parameters is far from intuitive, as small changes can produce nonlinear effects and qualitative shifts in behavior. Achieving the right phase durations, input-dependent scaling, waveform amplitude and shape, phase delays, and other characteristics simultaneously to match experimental data is quite cumbersome in the elegant models, not to mention the biophysical models. In contrast, these characteristics are easy to control in our model, because we translate complex dynamical behavior from implicit to explicit and surface a set of interpretable and tunable parameters.

      The authors argue for their model based on the idea that more biophysical models are difficult to tune, yet they compare their model to a biophysical one that they were able to tune to achieve the various patterns that they study. They do not give any indication of how easy or hard it was to tune their own model [...] The biophysical model seems to have 22 parameters, whereas the simplified one has 21 in Table 2, which is essentially the same number.

      To clarify, we did not tune the biophysical model, but rather copied its parameters from the cited work. We will make this more explicit in the relevant Methods section.

      We could not simply specify or tune these parameters because they have complex biological priors that must be derived from experimental data for example, the membrane capacitance (20 pF), ionic conductance and reversal potentials (4.5 nS, -62.5 mV), and many gating kinetics parameters (slopes, midpoints, time constants for sigmoid/bell curves).

      It is often the case that such parameters must be estimated in specific preparations then reused and refined over many years. For instance, the biophysical model we compare to borrowed parameters from (Kim et al. 2022), which retuned time constants relative to (Danner et al. 2017), which altered NaP conductance from (Danner et al. 2016), which retuned duty cycles from (Molkov et al. 2015), which adapted from respiratory networks of (Rubin et al. 2008), which used gating kinetics parameters from (Butera et al. 1999). Similarly, the crustacean pyloric circuit model we compare to is from (Alonso and Marder 2020), which augmented the circuit and parameters of (Prinz et al. 2004), which sampled from a database of procedurally generated parameters from (Prinz et al. 2003), which developed parameter priors from the lobster STG experimental work of (Turrigiano et al. 1995). These brief descriptions of the multi-decade lineage of parameter sets omit the substantial parallel and preceding work related their development, but they suffice to demonstrate the incredible science and effort that goes into building biophysical models for particular circuits. Such data is often unavailable and such detail is often undesirable for different research goals, in which case our simplified model is a valuable and practical tool.

      The key parameters of our simplified model are observable quantities like active/quiet durations (in seconds), input-dependent duration scaling (as a fraction of intrinsic durations), input strength that induces tonic firing, etc. As such, tuning the bursting neuron parameters for circuit models was easy, with manual tuning from scratch taking less than 1 day. As Table 3 shows, the resulting parameters are often simple, elegant numbers and can be derived directly from observations. For instance, the pyloric PD active and quiet durations (200 ms and 800 ms, respectively) are set using the exact target values that (Alonso and Marder 2020) encode in their objective for a genetic algorithm to tune their model’s biophysical parameters (or rather, a subset of them for tractability).

      Thus, the 22-vs-21 comparison is not very informative, because the parameters are not comparable in kind. However, to make it easier to tune our model, we will revise the manuscript to include: (a) more details describing how parameters maps onto observable features of the bursting waveform, (b) recommended parameter ranges and the qualitative behaviors expected at their boundaries, and (c) practical guidance for tuning the model to match observations or embed into circuits.

      (3) Originality

      What the authors fail to acknowledge is that Rinzel, Terman, Kopell, and others did seminal work on neuronal activity [...] The authors do not cite the substantial existing work on piecewise linear models [...] I don't see any advantage of the proposed framework over the earlier relaxation oscillator setting, where many important mechanistic principles have already been analyzed, including extensions to networks.

      We thank the reviewer for these pointers and apologize for the gap in our literature coverage. While we had cited McKean, FitzHugh-Nagumo, Izhikevich, et al. as representative examples of different model classes, we agree that the broader relaxation-oscillator and piecewise-linear traditions deserve more comprehensive treatment including Rinzel, Terman, Kopell, et al. on relaxation-oscillators; and Hahnloser, Coombes, Aguirre, et al. on piecewise-linear models. We will expand the related work discussion and clarify how our contribution is novel and valuable.

      To be clear, we do not claim to be the first to use piecewise-linear models for neurons. Our intended contribution is the specific construction a rectangular limit cycle whose horizontal/vertical decoupling permits a closed-form mapping from interpretable parameters to burst features and the demonstration that this construction integrates cleanly into firing-rate circuit models of biological CPGs, which we believe will provide realism for more complex models with learned components.

      Moreover, in contrast to many other relaxation-oscillator models including the elegant Fitzhugh-Nagumo-style model we discussed above, our model is not aimed at establishing mechanistic principles or being simple enough to analyze formally. It is a practical tool that affords precise control of many bursting characteristics, which is important for closer alignment between firing-rate circuit models and biological activity. We will state this contribution more precisely in the revision so it is not conflated with a broader novelty claim.

      Reviewer #3 (Public review):

      (1) Novelty of piecewise-linear approximation

      The use of piecewise linear approximations to explicitly estimate properties of biophysical neurons is a well-known and common technique. This study adds nothing to the technique in terms of novelty.

      We agree that piecewise-linear approximations of neurons are not themselves novel, and we have not intended to claim otherwise: We cite the McKean model as a direct predecessor and, prompted by Reviewer #2, we will substantially expand citations to the relaxation-oscillator and piecewise-linear traditions (Rinzel, Terman, Kopell, Hahnloser, Coombes, Aguirre, et al.). Our intended contribution is not the use of piecewise-linear pieces per se but the specific construction: a rectangular limit cycle whose horizontal/vertical decoupling permits a closed-form, interpretable mapping from burst features (duration, duty cycle, amplitude, shape, input dependence) to dynamics, and clean integration into firing-rate circuit models of biological CPGs. We will revise the relevant passages so this contribution and the boundaries of our novelty claim are stated precisely.

      (2) Dynamical system mechanism

      This is no better than having a look-up table [...] The neuron is restricted to what the user puts in, and therefore, calling it a dynamical system is entirely wrong.

      We would like to take the opportunity to clarify this point, because the model's behavior is much richer than the lookup-table characterization suggests. The model is closed-loop: trajectories evolve through coupled state variables whose response to time-varying input depends on current state, not on a precomputed table of input-to-output values.

      Specifically:

      (a) The input represents the net time-varying synaptic drive, not a clamped voltage level;

      (b) The adaptation and voltage variables evolve according to coupled differential equations both on and off the limit cycle;

      (c) The duration and scale parameters only constrain active/quiet durations at input endpoints (-1, 0, +1), while the response at intermediate inputs is determined by the dynamics and other parameters such as the adaptation time constant, which can qualitatively reshape the constant-input response curve (Figure 3—supplement figure 3);

      (d) The response to a transient input depends on the current state for example, excitatory pulses early in the active phase have little effect, as in the biophysical model.

      This is a direct result of the simplified model using a similar limit cycle and nullcline structure as the biophysical model’s dynamical system (Figure 1).

      (3) PRC usage

      The phase resetting curves are used incorrectly. PRCs are useful when the perturbation is weak (soft) [...] A hard PRC would always reset the cycle to the fixed offset from the perturbation phase and is therefore uninformative in understanding dynamics.

      We appreciate this point and would like to clarify what we show and why. We present finite (non-infinitesimal) PRCs across a range of input strengths and signs, spanning both the "soft" (weak-perturbation) regime as well as the "hard" (strong-perturbation) regime, rather than focusing on the "hard" regime alone. Importantly, even in the strong-perturbation regime we do not see that pulses "always reset the cycle to the fixed offset from the perturbation phase". In Figure 4, we see that the active phase exhibits a non-resetting region whose size and location depend on parameters. This region governs entrainability and phase-locking offset, and is thus a key aspect of the neuron's dynamics. Moreover, the strong-perturbation regime is also biologically relevant in our circuit examples. For instance, the inhibitory connections within the pyloric CPG are strong enough to cause hard resets, and these resets shape the circuit-level dynamics we reproduce. We will revise the pulse-input section to state these points more explicitly so the rationale is clear for showing PRCs across a range of inputs.

      (4) Defining active/quiet phases

      The definition of the active and quiet parts of a burst is often less clear than what the authors suggest. Bursting neurons often do multiple bursts in a cycle, and therefore, substituting the burst envelope is a subjective matter. This is even more problematic in bursting neurons in the brain, where there is often no quiet period.

      We agree that waveform envelope can be subjective in some preparations, and we can add this caveat to the discussion.

      On neurons with no quiet period, we note that this behavior is in fact already supported in our model, as seen in Figure 3: under strong excitatory input, both the biophysical and simplified models enter a regime in which firing rate never reaches zero. As the model can generally be viewed as an abstract limit cycle that maps onto periodic waveforms through the firing function, the quiet phase need not correspond to literal silence.

      On more complex waveforms, we could imagine different firing functions that produce richer burst shapes including multi-peak bursts, but we have not tried this explicitly. Of course, for research questions concerned with irregular bursting or spike-to-burst transitions, a lower-level biophysical model would be more appropriate. In revision, we will expand on how the firing function could produce more complex burst shapes.

    1. eLife Assessment

      This study presents potentially important findings linking peripheral inflammation to the remodeling of perinodal adipose tissue and draining lymph nodes, suggesting a mechanism by which local tissue inflammation can reshape LN structure and metabolism. The idea is solid and supported by observations. However, the evidence remains incomplete in parts, as several conclusions rely on correlative weight and cellularity measurements, and macrophage involvement requires further validation.

    2. Reviewer #1 (Public review):

      The idea is super interesting, and the subsequent work is potentially significant because it links peripheral inflammation to remodelling of perinodal adipose tissue and draining lymph nodes. This suggests an antigen-independent manner by which local tissue inflammation can communicate with and reshape immune organ structure and tissue metabolism. However, the evidence is suggestive. For instance, many conclusions rely on correlational weight/cellularity relationships, models with confounders (spontaneous wounding; potentially systemic IMQ), and macrophage dependence inferred from a single pharmacologic approach without definitive depletion/lineage or tracer-based causal link.

      Major Comments:

      (1) "Wounding/fighting" evidence is confounding.

      Unless I am mistaken, a large part of the argument for inflammation-driven perinodal fat pad atrophy and LN expansion relies on spontaneous fighting injuries in co-housed CCR2-/- males, including animals "culled...due to excessive wounding." Because wound severity, duration, infection load, stress, and cage dynamics are uncontrolled, isn't it difficult to assign causality to "cutaneous inflammation"?

      (2) The "CCR2-independent macrophage" conclusion.

      The manuscript interprets persistence/accumulation of macrophages despite reduced inflammatory monocytes as CCR2-independent recruitment or local proliferation. However, CCR2 deficiency can alter immune baselines and long-term tissue remodelling. Perhaps consider bone marrow chimeras (WT to CCR2-/-, CCR2-/- to WT ????) or an inducible CCR2 deletion approach to separate developmental/systemic effects from acute inflammation-driven mechanisms. If "in situ proliferation" is proposed, include a direct readout (e.g., Ki67 in ATMs in the fat pad).

      (3) IMQ and systemic effects.

      The work relies on topical Aldara/imiquimod as an "inflammation without antigen" driver of distal LN/fat-pad remodelling. But IMQ is well known (and cited by the authors) to enter circulation and drive systemic responses, which could blur whether effects are truly draining-site specific vs systemic metabolic/inflammatory effects. It would be ideal to provide systemic context: plasma cytokines and/or metabolic readouts (e.g., circulating FFAs) to distinguish local vs systemic drivers.

      (4) Macrophage dependence is inferred from CSF1R inhibitor treatment.

      However, validation of macrophage depletion and specificity is incomplete. The manuscript uses AZD7507 (CSF1R inhibitor) and observes partial rescue of fat pad/LN phenotype while skin severity (PASI) is unaffected. But, to this reviewer, the data shown do not clearly quantify actual macrophage depletion efficiency in the target fat pad, and LN at endpoint, and CSF1R blockade can affect multiple myeloid populations. Therefore, show absolute macrophage counts (and likely other myeloid populations) in fat pad and LN with/without AZD7507 at the analysed timepoints, not only outcome weights. (The methods describe dosing but not endpoint depletion quantification??)

      (5) Fat pad atrophy/LN expansion is a correlation.

      The paper emphasises negative correlations between fat pad and LN weights/cellularity at baseline and with inflammation. But correlation does not establish whether fat pad lipolysis drives LN expansion, whether LN changes drive fat remodelling, or whether both reflect systemic mediators. Add tissue-level evidence distinguishing true adipocyte loss vs other contributors to "weight change" (e.g., oedema/fibrosis).

      (6) Evidence for "fatty acid donation" from fat pad to LN.

      The lipid data are described as "exemplary," and the inference that LN fatty acids originate from the fat pad is based on temporal ordering and relative abundance. This does not rule out plasma spillover, LN-intrinsic metabolism, or altered lymph flow.