1. Last 7 days
    1. Reviewer #1 (Public review):

      Summary:

      A prevailing view is that translation of 5' capped mRNAs, i.e. mRNAs that are translated via ribosome scanning, is inhibited by highly structured 5' untranslated regions (5' UTRs). Despite having a common, structured 5' UTR, the mRNAs produced by the SARS-CoV-2 virus are efficiently translated. In this study, the authors identified a DRACH motif in stem-loop 3 (SL3), suggesting a potential site of m6A methylation of A74 by the enzyme METTL3. Given that such m6A modifications are known to disrupt RNA structure formation, the authors tested the hypothesis that this may be the basis underlying the efficient translation of these mRNAs. Mutational approaches complemented by METTL3 siRNA knockdown were employed to support this hypothesis. Additional experiments showed that this is required for efficient association of a reporter mRNA with polysomes (indicative of active translation), and suggest that the 5' UTR is more highly structured when methylation is abrogated.

      Strengths:

      The data clearly indicate that N6 methylation of A74 is required for efficient translation of SARS-CoV-2 mRNAs.

      Weaknesses:

      While the evidence supports the authors' central hypothesis, there are two issues that should be addressed. The first is that all of the approaches are indirect. All of the evidence for the presence of mRNA structural elements is based on computational and genetic analyses. We now know that there is something there, but we still do not know what it is. The authors need to use a biochemical approach to actually map the structural elements of the 5' UTR and determine how such structure(s) are changed by loss of methylation. The second hinges on the assumption that these mRNAs are translated via canonical ribosome scanning. RNA viruses are well-known to use a variety of other mechanisms, e.g. internal ribosome entry signals and ribosome tethering, to promote efficient translation. Alternatives to ribosome scanning should be considered.

    2. Reviewer #2 (Public review):

      The study addresses the conundrum of how the mRNAs of SARS-CoV-2 are efficiently translated since the 5' leader, which has common elements for all the viral genes, is highly structured. The authors test the hypothesis that m6A modification at position 74 is key to this translation. First, the authors show convincingly that this site is modified. Then, with extensive transfected reporter experiments using luciferase assays as well as sucrose gradient sedimentation, this modification is shown to be key for efficient translation. While the mechanism of this effect is not entirely clear (see comments/suggestions below), the authors show that it is independent of YTH "reader" proteins and likely involves altered interactions between SL3 (which contains A74) and downstream elements in the UTR. These results are important because they both offer insight into the function of m6A in gene expression and suggest how they may be important for the translation of viral mRNA in particular.

      While the data on their own make the overall case that the m6A modification in the 5'UTR of the viral genes is important to their expression, there are several things worth considering that could refine the model and make it more convincing.

      It is not clear whether putative uORF translation, particularly translation of the uORF that begins with a CUG codon at position 59 in the 5'UTR (as shown in Finkel et al., Nature 2020), would be impacted by this modification (as it includes the putative m6A site at position 74). It is also worth considering whether SL3 melting by translation of this uORF would alter the proposed mechanism.

      It is a bit unclear why the A74T mutant was put in the longer construct while the C75G mutant was put in a shorter construct. While not essential, the mechanistic arguments would be stronger if the same construct had been used to compare the mutations.

      While the authors show that "global depletion of m6A modification does not grossly alter translation efficiency" in a general sense (page 11), it would be of interest to know whether any host mRNAs with 5'UTR m6A (i.e., ACTA2 and COX8A, mentioned in this study) are affected by the mechanism here (i.e. run them in the luciferase assay).

      The authors show that the YTH "reader" proteins have a very small inhibitory effect (1.6-fold) on the translation of the viral mRNA with 5'UTR m6A. However, it remains unclear how important this is or whether it is generally true for host mRNAs with this modification.

      It is reassuring to see controls for changes in RNA levels in the supplemental material. The RNAs were generally stable under the experimental parameters explored, which would rule out RNA-decay-based mechanisms of m6A regulation. However, it should be noted that mRNA level experiments appear to have been done at 24 h while luciferase measurements were done at 48 h (as noted on p. 23, gene expression vs luciferase activity). It is not clear whether any RNA decay phenotypes would be apparent at 24 h.

      The authors use the term "ribosome profiling" (for example, on page 10), but it would appear the experiment performed is actually "polysome profiling" or "sucrose gradient sedimentation" since it did not involve ribosome footprinting.

    3. Reviewer #3 (Public review):

      Aly et al investigate the potential for a single N6-methyladenosine RNA modification in the context of the 5' UTR sequence of SARS-CoV-2 to regulate translation of a downstream luciferase reporter transfected into cells. They show using meRIP (m6A RNA IP) that this site is methylated in the plasmid-driven transcript, and convincingly show it mediates reporter translational efficiency using knockdown of the m6A methyltransferase METTL3 and mutation of the modified UTR site together with analysis of the transcript's association with polyribosomes. They suggest that the benefit to translation conferred by the modification is through its effect on the secondary structure of the 5' UTR, based on an RT-PCR-based assay in control and METTL3 knockdown cells linking RT processivity to translation (luciferase) output. They also extend their conclusions to two cellular mRNA 5' UTRs, also reported to contain a single m6A modification, and show METTL3-dependent changes in RNA structure stability, hinting at a broader significance of this mechanism of m6A control of gene expression.

      The conclusions of the paper are mostly well supported by the data presented, though validation of knockdown of METTL3 (and reader proteins) is absent.

      A major limitation of the work is the exclusive use of the reductionist artificial reporter system in uninfected cells. Though the 5' UTR site they identify is methylated in the context of a transcript generated in the nucleus (where the m6A installing complex is mainly localized, and believed to act exclusively in uninfected cells), how frequently this site is modified, if at all, on viral RNAs generated within cytoplasmic membrane-bound replication organelles. Similarly, whether the translation regulation by a single m6A modification identified here occurs within the context of an infected cell, in which there are many changes to the RNA and translational regulatory landscape, also remains to be tested.

      How this work can be reconciled with others that have concluded either little potential for translational regulation by 5' UTR modification (Guca et al 2024; PMID: 38244546) or that an eIF3-mediated mechanism is responsible (Meyer et al, 2015 PMID: 26593424) is not addressed in the discussion.

    1. It is valuable to have educators who can discuss with students the social and cultural context of code switching and the significance of encountering multiple linguistic varieties at school.

      possible example/evidence

    2. Educators should work to create courses that teach Black English and design assignments in other courses that encourage all students to use Black English in written, spoken, and other forms. When such courses and assignments are engaged with the community, then the justice of the actual speakers of Black English is brought to bear.

      Institutions should make courses that teach Black English, making assignments that support all students in using Black English in all it's forms written, and spoken. When these assignments and courses are put into communities, there is supported justice for speakers of Black English.

    3. the melody and rhythm of a speaker's voice may mark a speaker as African American even if all other aspects of the speaker's language sound standardized. Rhythmic patterns are often preserved in speakers who do not use many of the more socially stigmatized lexical, phonological, and grammatical features of AAE.

      Potential subtopic - How to tell the difference between speakers by the use of rhythmic patterns, lexical, phonological, and grammatical features of AAE.

    4. Many of the phonological features of AAE are shared by other varieties of American English, particularly Southern American English. Some features that are common to African American English also appear as features in the speech of younger Southern American English speakers; yet many listeners can determine if a speaker is African American after hearing just a short speech sample

      Phonlogical features of AAE are combined with varieties of American English like Southern American English. Southern American English speakers share common traits with African American English, but the different speakers can be differentiated between each other.

    5. Current research on AAE continues to build on knowledge of specific social groups and African American communities as well as the educational implications for AAE use on language and literacy skill acquisition.

      possible example/evidence

    6. The AAE lexicon is perhaps the feature that is best known to the general American public. The lexicon is emphasized and frequently represented in popular culture. There are also other often-unnoticed lexical differences that may have an effect on a child's understanding and success in the classroom

      Possible topic/theme - The use of SE speech between African American children vs. White and middle class children.

    7. Black English usage varies by the age, gender, region, and social class of the speaker. Most sociolinguistic studies do not examine every given feature of Black English, nor could they, and it is difficult to make cross-study comparisons of feature use over space, time, and demographic group.

      possible example/evidence

    8. As approaches to Black English have become more diasporic in nature, definitions of Black English have now been expanded to include varieties of African and Caribbean English in particular as well as individuals who come into contact with blacks of the diaspora and acquire some of their language patterns.

      With Black English spreading into more areas, it is also expanding with more African and Caribbean cultured English, as well as individuals coming into contact picking up on each other language pattern.

    9. The U.S. government, particularly the U.S. Department of Education, funded seminal studies of Black English in the 1960s and 1970s in order to address academic inequality. Black English was examined in contrast to SAE.

      There was movements made in Black communities for studies in Black English, but nothing was improving.

    10. The term Ebonics was widely adopted by educators and the general public following a movement in Oakland, California, in 1996–1997 that was designed to help teachers use Black English as a way to help students acquire the language of school instruction and assessment.

      Educators made efforts to help students with understanding each other, but it eventually was prevented from being done.

    11. A variety of terms have been used to describe English as spoken by African Americans in the United States, Black English (BE) including Ebonics, and African American Vernacular English. African American English (AAE) and African American Language (AAL) are the most encompassing terms used by educators and linguists to refer to all varieties of English used by speakers where African Americans live or historically have lived.

      Black English (BE), Ebonics, African American Vernacular English (AAVE), African American English (AAE), African American Language (AAL)

    1. I am your opus, I am your valuable,    The pure gold baby

      the author believes that she is something to be watched, perceived, and ultimately, that she serves as a muse to the audience

    2. O my enemy.

      i sort of imagine her saying this to herself. is she her own enemy? everything up unto this point is about herself; her face, her foot, her skin. i would presume she is talking about herself in this, maybe she is her own worse enemy.

    1. In the ensuing years her work attracted the attention of a multitude of readers, who saw in her singular verse an attempt to catalogue despair, violent emotion,

      Important info

    1. Though each speaker of AAVE is unique, and age, status, and contextual differences affect its use, some characteristics of AAVE have been detected across speakers and regions. Many of these characteristics reflect their African heritage, as well as their heritage in the American South. These features include distinctive phonology, or pronunciation; distinctive lexicon, or vocabulary; and distinctive syntax, especially regarding use of verb tenses and the copula (the “to be” verbs).

      potential example/evidence

    2. European-American linguist William Labov was probably the first linguist to thoughtfully analyze the grammar of AAVE in his 1965 article and more thoroughly in his 1972 book (see “Further Readings”). In both his article and his book, he urged readers to recognize AAVE as a respected variety of English, with its own distinctive grammatical rules.

      potential example/evidence

    3. During the Harlem Renaissance, many authors began to use AAVE in their writing and to show fond appreciation for it. For instance, Arna Bontemps not only used his native Creole dialect in his personal correspondence to intimate friends such as Langston Hughes, but also studied dialects in order to use authentic AAVE in his writings so that he would neither sound stereotypical nor be incomprehensible to speakers of SE.

      potential example/evidence

    4. Even after the Civil War and the all-too-brief period of Reconstruction, most writings by African Americans continued to be chiefly in SE, not in AAVE. In fact, European-American Southerners were more likely than African Americans to use a form of AAVE, typically through nostalgic stories of the good old days of plantation slavery. Typically, these narratives used highly stereotyped written representations of the phonics and syntax of AAVE, intended to portray the speakers as not only poorly educated but also unintelligent.
    5. not all African Americans speak AAVE as a native language form, and not all speakers of AAVE are African Americans. Even among native speakers of AAVE, most African-American writers use SE in their formal writing (e.g., for publication or for business audiences), whether or not they do so in their informal writing (e.g., when texting, tweeting, e-mailing, or corresponding with family or friends). The tradition of African-American writers using SE goes back to the earliest days of African-American written literature.

      potential subtopic- not all African Americans speak AAVE, and not all speakers of AAVE are African Americans

    1. Lingering Question- I understand that the bilingual brain does not need a different approach when teaching reading but what is the process of a teacher teaching reading to a student who speaks different languages and how can teachers fully support those students?

    2. ut the evidence is also clear that children need to learn the sound-symbol spelling system in order to be successful readers, "binding letters and sounds to meaning," as Pugh says. (In the world of reading instruction, this is known as "orthographic mapping.")

      Instructional Implications- This something that could directly influence classroom instruction or literacy practices.

    3. One difference is that students learning to read in a new language need additional oral language support so that they will understand the words and text being used to teach them to read.

      Instructional Implications- This is something that can could directly influence classroom instruction or literacy practices.

    4. one approach to teaching reading is right, and another is wrong. Whether it’s whole language vs. phonics, balanced literacy vs. science of reading, or bilingual brain vs. monolingual brain, with too few exceptions, lines get drawn and cleavages remain deep. Mixing in and perpetuating misinformation complicate things further — and needlessly.

      Key ideas- Teacher need to know that there are many ways to teach reading.

    5. (In the world of reading instruction, this is known as "orthographic mapping.") But the evidence does not support stopping there. As students learn the sound-symbol spelling system, they must also have oral and experiential exposure to develop their language much further — particularly, but not exclusively, vocabulary — and knowledge. As literacy skills develop, that exposure will include reading (and writing).

      Surprise findings- This has expanded my knowledge on the sound-symbol spelling system.

    6. But the evidence is also clear that children need to learn the sound-symbol spelling system in order to be successful readers, "binding letters and sounds to meaning," as Pugh says.

      Key Ideas- Teachers should know that children need to learn symbol spelling system in order to be successful readers.

    7. No, it does not. I contacted Dr. Kenneth Pugh (Yale University and Haskins Lab), an internationally-recognized cognitive neuroscientist specializing in the neuroscience of reading in first and second languages and who chaired a recent conference session on this very topic.

      Surprise findings- This cleared up a misconception I had about bilingual speakers learning to read.

    8. However, cross-language brain research confirms that learning to read is based on cognitive universals, specifically, that phonological development makes possible binding letters and sounds to meaning, which is foundational for learning to read in any language. Our understanding might change as the research evolves, but at the moment, in my opinion, there is nothing about ‘bilingual brain’ differences that suggests distinct or alternative pathways to literacy learning and best practice.

      Key ideas- This is something teachers should know when teaching students to learn because learning to read is based on cognitive universals and that reading is not something people are born with.

    9. This research supports the idea that what is true about teaching reading to monolinguals is also true for students learning to read in a language they are simultaneously learning to speak and understand.

      Key Ideas- All teachers should know that teaching reading to monolinguals is also true for students learning to read in a language they are simultaneously learning to speak and understand.

    10. Brain studies (opens in new window) and classroom studies (opens in new window) reveal that in order to learn to read, a person must connect (or “bind” is the term used in the neurolinguistic literature) the oral sounds in words to the letters that represent those sounds, then connect that connection to the words' meanings.

      Key Ideas- Every teacher should understand this when teaching reading.

    11. In my view, there is presently no evidence that how we teach reading should be in any way different based on brain differences between bilingual and monolingual learners.

      Surprising Findings- I did not know that teaching reading would not be different for bilingual and monolingual learners.

    Annotators

    1. Wonderful initiative that clearly fills a regulatory transparency gap in Europe. This would also satisfy the 'common knowledge requirement' of the U.S. GRAS framework.

      Do you plan to submit to a journal for peer-review? Minor modification (to state its suitability to U.S. premarket consultation review or GRAS notification) and scientific journal peer-review (by qualified experts in the field of cell-cultured meat food safety) would go a long way towards reaching 'consensus amongst qualified scientists', the second GRAS requirement.

      Feedback appreciated either here or in private at sewalt.bioconsulting@gmail.com. Cheers - Vince Sewalt

    1. a system away.

      one repeatable system away.

      Join us at the Solo Consultant Summit, and take the first steps towards your own scalable, relationship-led & client-converting system.

    2. This is NOT for you if you're looking for get-rich-quick tactics, you're not willing to put in the work to build a system, or you're selling to consumers (not businesses). No judgment, just different playbooks.

      If you're looking for ...this summit is not the right environment for you.

    3. Launch week Scholarship + cart open September 17 September 14 Applications open. Free summit ticket registers you to apply. September 17 Cart opens for the Booked Out Club. Scholarship winner announced live. September 17 — Bonus Day Live training: How to make 2027 your most profitable year. Brainstorm Hour: bring your business questions. Scholarship winner announced. Live 2027 Planning Training Brainstorm Hour Scholarship Announcement

      I'd remove this section entirely and somehow incorporate the info into the How it works section.

    4. September 14 Keynote kick-off. Relationship-led pipeline trainings go live. Available for 48 hours. Live inquiry form tear-downs with real feedbackwith Kendall Cherry. September 15 Scalable Lead Generation trainings go live. Available for 48 hours. Live consultant panel: what's working in 2026 with a panlist of solo consultants that stay booked out. September 16 Sales Skills That Convert trainings go live. Available for 48 hours. Networking party hosted by Liz Batsche of Well Hosted. Summit Bingo with prizes. Community connections that last past the event. September 17 Bonus Day for Booked Out Collective Applicants. Includes live training, scholarship announcement, and Brainstorm Hour.

      Can you split these out more so each of the components don't get lost?

    5. $100K+ in contracts using

      It might just be me, but that number doesn't seem super impressive in itself. What's impressive is the experiment that you ran... and clearly needed some solid strategies behind it. So I'd skip the 100k number entirely and lead stratight with the 94k experiement in 30 days.

    6. Sarah Noel Block After 15 years in corporate marketing, Sarah launched her own consulting business and signed $100K+ in contracts using the exact systems she now teaches. She ran a 30-day experiment that closed $94K in new business, created the Booked Out in Six framework, and has helped 50+ solo consultants build predictable referral and visibility engines. She built this summit so you can hear directly from the consultants who are doing the work, closing the deals, and building the systems you need.

      I'd consider switching this to first person so it comes across more personal

    7. Building a referral engine that doesn't depend on luck Follow-up frameworks that feel like hospitality, not hustle Turning warm connections into booked consultations

      I would turn this into a to-the-pont blurb rather than a bullet point list. The actual speaker topics inside these 3 tracks are going to take care of those specific bullet points and ideas.

    8. rk you love.

      This needs a transition to the next section. Something like... The free Solo Consultant Summit will set you up with a long list of actionable steps you can take today towards your own repeatable system.

    9. Every quarter feels like starting from zero, and the feast-or-famine cycle is exhausting.

      I would make this a standalone statement and format it differently so it really hits home.

    10. Sound familiar?

      I know short and sharp can work, but I'd give a little more of an intro before dropping this list on readers. ;-) Something like: Finding you here checking out this event gives me a pretty good indication of what your solo consultant reality looks like right now:

    1. Analysis of dialectic differences within the sample revealed that 80.8% of participants spoke MAE. Of the remaining participants, 8.0% spoke AAE, another 7.7% of participants spoke Southern American English, and the final 3.5% spoke other dialects

      The DELV included children who spoke several English dialects, helping ensure the test works fairly across different ways of speaking.

    2. The following classifications were examined within the demographic analysis of the sample: age, sex, race/ethnicity, parental education level, and region.

      The study included children from different backgrounds to reduce bias and make the assessment more inclusive.

    3. The authors based their standardization sample on this data in an attempt to closely mirror the percentiles across demographic categories

      Researchers made the sample similar to the U.S. population so the test results would be fair and representative

    4. A sample of over 900 participants was used in the standardization of the DELV-NR edition.

      The DELV was standardized using a large group of children to help ensure the test is reliable and accurate.

    5. items needed to be non-contrastive, resulting in minimal differences regardless of the dialect of English spoken by the test-takers.

      The DELV uses language features that are common across English dialects so children are not judged unfairly because of the way they speak.

    6. items needed to both differentiate between typically developing children and impaired children

      The questions were carefully chosen to identify true language disorders while recognizing normal language development.

    7. A group of 1370 test-takers from across the USA, including 450 speakers of MAE, 800 speakers of AAE, and 120 individuals who spoke other dialects, participated in refining the assessment material

      Researchers tested the DELV with a large and diverse group of children to make sure the assessment was fair and accurate.

    8. In 2005, DELV-NR became the first standardized assessment tool aimed specifically at assessing children who speak AAE and other variations of English

      The DELV-NR was an important step toward fair language testing because it recognizes dialect differences instead of treating them as language disorders

    9. The problem of using MAE as the norm against which to compare the speech and language of AAE speakers

      Using Mainstream American English as the standard unfairly judged children who spoke African American English, leading to inaccurate evaluations

    10. Historically, research has acknowledged that African American and other minority children are disproportionately represented in special education

      Minority children have often been placed in special education at higher rates, sometimes because of bias in language assessments rather than actual disabilities.

    11. Overall language performance is reported as a composite standard score and percentile rank.

      The DELV combines scores from several language areas to give an overall picture of a child's language abilities and compare them with other children of the same age.

    12. It is designed for children who speak English as their first spoken and primary language

      The assessment is intended for native English-speaking children, including those who speak different English dialects

    13. The DELV-NR is normed for children aged 4 years, 0 months through 9 years, 11 months

      The test is specifically designed and standardized for children between the ages of 4 and 9 years old.

    14. These items measure the child’s overall ability to use social communication and appropriately respond to questions and situations.

      The DELV measures how well children communicate, understand language, and use it in everyday social interactions

    15. The subtest items were selected on the basis of representing syntactic features of English that are universal among all of its dialects.

      The test focuses on language skills that are common across all English dialects to reduce cultural and linguistic bias

    16. The DELV-NR assesses the language domains of syntax, pragmatics, semantics, and phonology in children regardless of their dialect of English

      The DELV evaluates multiple areas of language while treating all English dialects fairly.

    17. Deficits in the use of non-contrastive features of English would be indicative of a true disorder rather than a dialectical difference.

      The DELV helps distinguish between a language disorder and a normal dialect difference so children are not misdiagnosed.

    18. in African American English (AAE), it is not obligatory to include “is.”

      This example shows that dialect differences are not mistakes—they are valid grammatical patterns within that dialect

    19. The DELV-NR strives toward being a culturally unbiased assessment

      The DELV-NR is designed to avoid cultural and language bias by recognizing that different English dialects follow different language rules.

    20. he DELV-NR is a diagnostic test that provides a standardized approach to distinguishing between speech and language differences versus delays and/or disorders in children.

      The DELV-NR helps professionals tell the difference between a normal language difference caused by dialect and an actual speech or language disorder.

    21. The Diagnostic Evaluation of Language Variation (DELV) is a group of tests designed to assess speech and language disorders in children who speak a dialect of English other than Mainstream American English (MAE

      The DELV helps evaluate children who speak different English dialects without unfairly judging their dialect as incorrect.

    1. Very nice work on improving CFPS. A few questions for you: 1) Is there a reason you didn't try doing at DoE for varying the components to try to identify which parameters, NTPs, AAs, genetic background etc. were of greatest weight of impact and how they might interact with each other? Could be an interesting way to continue to improve the system.

      2) Do all of these trends hold for other proteins besides GFP? Would be nice to test a few other proteins on your idealized lysate to test for universality by testing a couple other target proteins with different folding and biophysical properties.

      I also noticed a typo on the y-axis of the graph in panel A in Figure 2. The "," is off by a digit in all the numbers, e.g. 10,0000 instead of 100,000.

    1. Open-weight AI is having its Kubernetes moment. Let's not ruin it.
      • Paradigm Shift in Open-Weight AI
        • Open-weight AI models are undergoing an infrastructure and adoption inflection point similar to Kubernetes' emergence in cloud-native computing.
        • Open models are evolving from experimental open-source artifacts into enterprise-grade standards, threatening proprietary AI incumbents.
      • Political & Regulatory Friction
        • Big AI labs and legacy closed-source vendors are lobbying governments to restrict or outright ban open-weight models under the guise of national security and risk mitigation.
        • Attempts to target specific foreign or Chinese open-weight models face technical infeasibility, creating pressure for broader open-source AI regulations.
      • Commercial Realignment
        • Enterprise infrastructure is rapidly standardizing around self-hosted, fine-tuned open-weight architectures to avoid vendor lock-in and control operating costs.
        • The open ecosystem is building a robust stack—from orchestration and serving frameworks to fine-tuning tools—replicating the open cloud-native blueprint.

      Hacker News Discussion

      • Regulatory Infeasibility & Feasibility Concerns
        • Commenters point out that distinguishing "American" from "Chinese" or foreign open-weight models by inspecting weights alone is technically impossible, as weights are purely mathematical numbers.
        • Enforcing restrictions by origin would force regulators toward blanket bans on all open-weight models or mandatory, restrictive DRM/licensing protection systems.
      • Regulatory Capture & Lobbying
        • Users argue that proprietary AI labs (such as OpenAI and Anthropic) are using foreign threat narratives to push for open-source AI bans, aiming to eliminate zero-marginal-cost open-weight competitors.
      • Technical Evasion & Distillation
        • Community members note that even if specific model weights were banned, trivial adjustments—such as fine-tuning, architecture tweaks, or layer shifts—would alter checksums and bypass simple detection.
      • First Amendment & Legal Precedents
        • Participants draw parallels to the historical "Crypto Wars" and software-as-speech legal precedents, questioning whether banning model weight distribution would withstand constitutional scrutiny in US courts.
    1. Por otro lado, con respecto a la confusión que generaron en Montoneros y en otros sectores las huelgas de 1975, la verdad es que la organización no puso toda la carne en las coordinadoras regionales. Puso los huevos en distintas canastas: en el Partido Auténtico, en la militarización de la JP, en diversas políticas que fueron desarrollándose en ese momento. En julio de 1975, la clase obrera peronista le hizo el primer paro de toda su historia a un gobierno peronista; y sobre ese episodio Montoneros interpretó que los trabajadores se estaban independizando de su identidad y que, por ese motivo y en ese avance sostenido hacia la revolución o hacia la lucha política de masas, encontrarían en Montoneros su nueva identidad.

      !

    1. Relevance

      It is important that you stay on topic with your sources. It is easy to have your sources expand out into other topics related to the one you are writing about, but getting carried away makes the writing process harder.

    1. The finding that structural-token information is represented by attention heads early in the network, whereas 3D contact prediction emerges in heads much later in the network, is fascinating.

      In Fig. 3, the sequence-only head-ablation heatmap appears to show that L0H7 is already the most influential head by KL-divergence magnitude, even without structural-token input (although its influence clearly increases substantially when structure tokens are provided). Could L0H7 therefore be a more general early routing or integration hub, rather than (or in addition to) a head specifically dedicated to ingesting structural-token information? Maybe additional controls like masked sequence tokens with structural tokens might help distinguish these possibilities.

      More generally, it would be interesting to see the heads ranked by ablation KL divergence in each condition. By eye, several heads in the final layers also seem prominent, particularly under S+St. Do these overlap with the contact-predicting heads described in the text? If so, do you interpret them as reading out structural representations inferred from sequence under S and then augmented by explicit structural input under S+St?

      It would also be interesting to rank by the difference between KL div, as it seems L0H7 would be first, followed by some of the layer 44 or 47 heads. Maybe the diff would yield some heads that don't show up in individual heatmaps.

  2. accessmedicine-mhmedical-com.ezproxy.lib.torontomu.ca accessmedicine-mhmedical-com.ezproxy.lib.torontomu.ca
    1. After delivery of the firstborn, one clamp is placed near the neonate, and another is placed nearer the placenta. Until the last fetus is delivered, each cord must remain clamped to prevent fetal hypovolemia and anemia caused by blood leaving the placenta via anastomoses and then through an unclamped cord. Cord blood is generally not collected until after delivery of all fetuses. After the second neonate is delivered, two plastic clamps are placed on the placenta’s cord to differentiate it from the first. In higher-order deliveries, color-tagged or alphabetically labeled clamps can be simpler than adding additional clamps. This same practice holds for cesarean delivery. At this time, evidence is insufficient to recommend for or against delayed umbilical cord clamping in multifetal gestations (American College of Obstetricians and Gynecologists, 2020a).

      After delivery of the firstborn, one clamp is placed near the neonate, and another is placed nearer the placenta. Until the last fetus is delivered, each cord must remain clamped to prevent fetal hypovolemia and anemia caused by blood leaving the placenta via anastomoses and then through an unclamped cord. Cord blood is generally not collected until after delivery of all fetuses. After the second neonate is delivered, two plastic clamps are placed on the placenta’s cord to differentiate it from the first. In higher-order deliveries, color-tagged or alphabetically labeled clamps can be simpler than adding additional clamps. This same practice holds for cesarean delivery. At this time, evidence is insufficient to recommend for or against delayed umbilical cord clamping in multifetal gestations (American College of Obstetricians and Gynecologists, 2020a).

    1. I regularly use this “Renaissance” wax to protect paintings and, above all, very fragile inscriptions. After it dries (10 minutes) and is lightly buffed, the surface becomes very slippery and is no longer susceptible to friction or fingerprints. It also helps preserve nickel plating that is beginning to deteriorate on levers or axles, for example. It’s a truly reliable product.

      via u/Tall_Garage_7398

    1. Direct bargaining—“I’ll give you this if you do that”—can be useful in a negotiation, but it’s also something to be cautious about. Wheeler told a possibly apocryphal anecdote about Richard Nixon and Henry Kissinger in which Kissinger watched Nixon use a treat to coax his Irish setter off a chair in the Oval Office. Kissinger supposedly told Nixon that the lesson the dog had learned was not that he should stay off the chair—but rather that if he got on the chair, he’d get a treat for getting off it. (In some tellings, the problem behavior was chewing the rug.)

      but be careful as it doesn't have intrinsic motivation

    1. 59.6% of BH-flagged responses map to at least one PHQ-9, GAD-7, or PROMIS Global Health item.

      Think i mentioned this before, but really curious about the other 40% that do not map onto these instruments. What are these responses about? Would be helpful to know in the case that we do develop a structured data collection process.

    2. The five patients below are drawn from, and are the strongest examples within, this smaller repeat-CSAT group.

      This is such a cool way to display the qualitative data!

    3. Below is a single ordinal composite that collapses valence, temporality, and impact_specificity into an “improvement arc”,

      I might define the three variables (valence, temporality, and impact_specificity) in a bit more detail!

    4. This section asks a different question: is there evidence of improvement over time? The Improvement Arc composite below looks for a before / after narrative within a single response, a patient’s own retrospective account of change, abundant across the sample (n ≈ 1,009) but not, strictly speaking, two separate observations of the same person.

      Think this points to the need for structured data collection as well, which can help standardize future analyses.

    5. % mental-health dx

      Curious whether there are more granular dx codes for these MH conditions to map onto, similar to how you map onto the specific SDOH z codes below

    6. These are rows with genuine BH signal that didn’t map cleanly to any specific construct, a residual worth keeping in mind when reading construct-level prevalence as a share of “all BH signal.”

      I'd be curious to know what types of responses fall under the 10.1%!

    7. Each flagged response is scored against eight named constructs spanning depression, anxiety, overwhelm, coping and functioning, hope / hopelessness, loneliness / isolation, caregiver burden, and general distress.

      Sorry if i missed this, but how are constructs defined? based on a theoretical framework, LLM categorization, or something else? Might be explicit here

    8. Across these dimensions, the impact evidence above skews clearly positive.

      We don't create some sort of composite score within "impact evidence", right? ie., a response coded as relief improvement and attributed to solace is scored greater than a response coded as distress and vague attribution to solace?

      I guess it might be hard to do that since some aren't necessarily ordinal. Just trying to think about some ways to further quantify "impact"

    9. For one, patients who submit CSATs are very likely unrepresentative of the entire population of Solace patients, particularly those who experience severe symptions of behavioral health challenges.

      Yes- I think this is an important point to consider as we think about the type of impact we can make on those who would need it the MOST. There's obviously a response bias here. This points well to your recommendation that we should find a way to systematically collect more BH info from patients in a less-burdensome way

    10. All 13,512 positive CSATs (from 8,853 unique patients) where scored in an initial pass to identify BH signal.

      Could you look at negative csats? do we have these? curious whether this might paint a more rounded picture of Solace’s impact, instead of solely relying on the positive csats

    1. Lesson 1.1: Introduction to Post-Acute Workflows (Mentee Preceptor Alignment)¶

      This lesson is a simple online meeting with the student(s) and preceptor(s) to go over the course. No learning materials. Course overview.

    1. eLife assessment

      In this manuscript, Rademacher and colleagues examined the effect of a chemogenetic approach on the integrity of the dopamine system in mice with chronically stimulating dopamine neurons. These findings are important: (1) This approach led to an axon-first degeneration over a time course (2–4 weeks) that is suitable for experimental investigation; (2) The finding that direct excitation of dopaminergic neurons causes differential degeneration sheds light on dopaminergic neuron selective vulnerability mechanisms. Overall, the strength of the evidence is solid, but the behavior experiments that do not include a CNO control provide incomplete support for the findings.

    2. Reviewer #1 (Public Review):

      Summary:

      In this manuscript, the authors investigated the effect of chronic activation of dopamine neurons using chemogenetics. Using Gq-DREADDs, the authors chronically activated midbrain dopamine neurons and observed that these neurons, particularly their axons, exhibit increased vulnerability and degeneration, resembling the pathological symptoms of Parkinson's disease. Baseline calcium levels in midbrain dopamine neurons were also significantly elevated following the chronic activation. Lastly, to identify cellular and circuit-level changes in response to dopaminergic neuronal degeneration caused by chronic activation, the authors employed spatial genomics (Visium) and revealed comprehensive changes in gene expression in the mouse model subjected to chronic activation. In conclusion, this study presents novel data on the consequences of chronic hyperactivation of midbrain dopamine neurons.

      Strengths:

      This study provides direct evidence that the chronic activation of dopamine neurons is toxic and gives rise to neurodegeneration. In addition, the authors achieved the chronic activation of dopamine neurons using water application of clozapine-N-oxide (CNO), a method not commonly employed by researchers. This approach may offer new insights into pathophysiological alterations of dopamine neurons in Parkinson's disease. The authors also utilized state-of-the-art spatial gene expression analysis, which can provide valuable information for other researchers studying dopamine neurons. Although the authors did not elucidate the mechanisms underlying dopaminergic neuronal and axonal death, they presented a substantial number of intriguing ideas in their discussion, which are worth further investigation.

      Weaknesses:

      Many claims raised in this paper are only partially supported by the experimental results. So, additional data are necessary to strengthen the claims. The effects of chronic activation of dopamine neurons are intriguing; however, this paper does not go beyond reporting phenomena. It lacks a comprehensive explanation for the degeneration of dopamine neurons and their axons. While the authors proposed possible mechanisms for the degeneration in their discussion, such as differentially expressed genes, these remain experimentally unexplored.

    3. Reviewer #2 (Public Review):<br /> <br /> Summary:

      Rademacher et al. present a paper showing that chronic chemogenetic excitation of dopaminergic neurons in the mouse midbrain results in differential degeneration of axons and somas across distinct regions (SNc vs VTA). These findings are important. This mouse model also has the advantage of showing a axon-first degeneration over an experimentally-useful time course (2-4 weeks). 2. The findings that direct excitation of dopaminergic neurons causes differential degeneration sheds light on the mechanisms of dopaminergic neuron selective vulnerability. The evidence that activation of dopaminergic neurons causes degeneration and alters mRNA expression is convincing, as the authors use both vehicle and CNO control groups, but the evidence that chronic dopaminergic activation alters circadian rhythm and motor behavior is incomplete as the authors did not run a CNO-control condition in these experiments.

      Strengths:<br /> This is an exciting and important paper.<br /> The paper compares mouse transcriptomics with human patient data.<br /> It shows that selective degeneration can occur across the midbrain dopaminergic neurons even in the absence of a genetic, prion, or toxin neurodegeneration mechanism.

      Weaknesses:

      Major concerns:

      (1) The lack of a CNO-positive, DREADD-negative control group in the behavioral experiments is the main limitation in interpreting the behavioral data. Without knowing whether CNO on its own has an impact on circadian rhythm or motor activity, the certainty that dopaminergic hyperactivity is causing these effects is lacking.

      (2) One of the most exciting things about this paper is that the SNc degenerates more strongly than the VTA when both regions are, in theory, excited to the same extent. However, it is not perfectly clear that both regions respond to CNO to the same extent. The electrophysiological data showing CNO responsiveness is only conducted in the SNc. If the VTA response is significantly reduced vs the SNc response, then the selectivity of the SNc degeneration could just be because the SNc was more hyperactive than the VTA. Electrophysiology experiments comparing the VTA and SNc response to CNO could support the idea that the SNc has substantial intrinsic vulnerability factors compared to the VTA.

      (3) The mice have access to a running wheel for the circadian rhythm experiments. Running has been shown to alter the dopaminergic system (Bastioli et al., 2022) and so the authors should clarify whether the histology, electrophysiology, fiber photometry, and transcriptomics data are conducted on mice that have been running or sedentary.

    4. Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Rademacher and colleagues examined the effect on the integrity of the dopamine system in mice of chronically stimulating dopamine neurons using a chemogenetic approach. They find that one to two weeks of constant exposure to the chemogenetic activator CNO leads to a decrease in the density of tyrosine hydroxylase staining in striatal brain sections and to a small reduction of the global population of tyrosine hydroxylase positive neurons in the ventral midbrain. They also report alterations in gene expression in both regions using a spatial transcriptomics approach. Globally, the work is well done and valuable and some of the conclusions are interesting. However, the conceptual advance is perhaps a bit limited in the sense that there is extensive previous work in the literature showing that excessive depolarization of multiple types of neurons associated with intracellular calcium elevations promotes neuronal degeneration. The present work adds to this by showing evidence of a similar phenomenon in dopamine neurons. In terms of the mechanisms explaining the neuronal loss observed after 2 to 4 weeks of chemogenetic activation, it would be important to consider that dopamine neurons are known from a lot of previous literature to undergo a decrease in firing through a depolarization-block mechanism when chronically depolarized. Is it possible that such a phenomenon explains much of the results observed in the present study? It would be important to consider this in the manuscript. The relevance to Parkinson's disease (PD) is also not totally clear because there is not a lot of previous solid evidence showing that the firing of dopamine neurons is increased in PD, either in human subjects or in mouse models of the disease. As such, it is not clear if the present work is really modelling something that could happen in PD in humans.

      Comments on the introduction:

      The introduction cites a 1990 paper from the lab of Anthony Grace as support of the fact that DA neurons increase their firing rate in PD models. However, in this 1990 paper, the authors stated that: "With respect to DA cell activity, depletions of up to 96% of striatal DA did not result in substantial alterations in the proportion of DA neurons active, their mean firing rate, or their firing pattern. Increases in these parameters only occurred when striatal DA depletions exceeded 96%." Such results argue that an increase in firing rate is most likely to be a consequence of the almost complete loss of dopamine neurons rather than an initial driver of neuronal loss. The present introduction would thus benefit from being revised to clarify the overriding hypothesis and rationale in relation to PD and better represent the findings of the paper by Hollerman and Grace.

      It would be good that the introduction refers to some of the literature on the links between excessive neuronal activity, calcium, and neurodegeneration. There is a large literature on this and referring to it would help frame the work and its novelty in a broader context.

      Comments on the results section:

      The running wheel results of Figure 1 suggest that the CNO treatment caused a brief increase in running on the first day after which there was a strong decrease during the subsequent days in the active phase. This observation is also in line with the appearance of a depolarization block.

      The authors examined many basic electrophysiological parameters of recorded dopamine neurons in acute brain slices. However, it is surprising that they did not report the resting membrane potential, or the input resistance. It would be important that this be added because these two parameters provide key information on the basal excitability of the recorded neurons. They would also allow us to obtain insight into the possibility that the neurons are chronically depolarized and thus in depolarization block.

      It is great that the authors quantified not only TH levels but also the levels of mCherry, co-expressed with the chemogenetic receptor. This could in principle help to distinguish between TH downregulation and true loss of dopamine neuron cell bodies. However, the approach used here has a major caveat in that the number of mCherry-positive dopamine neurons depends on the proportion of dopamine neurons that were infected and expressed the DREADD and this could very well vary between different mice. It is very unlikely that the virus injection allowed to infect 100% of the neurons in the VTA and SNc. This could for example explain in part the mismatch between the number of VTA dopamine neurons counted in panel 2G when comparing TH and mCherry counts. Also, I see that the mCherry counts were not provided at the 2-week time point. If the mCherry had been expressed genetically by crossing the DAT-Cre mice with a floxed fluorescent reported mice, the interpretation would have been simpler. In this context, I am not convinced of the benefit of the mCherry quantifications. The authors should consider either removing these results from the final manuscript or discussing this important limitation.

      Although the authors conclude that there is a global decrease in the number of dopamine neurons after 4 weeks of CNO treatment, the post-hoc tests failed to confirm that the decrease in dopamine number was significant in the SNc, the region most relevant to Parkinson's. This could be due to the fact that only a small number of mice were tested. A "n" of just 4 or 5 mice is very small for a stereological counting experiment. As such, this experiment was clearly underpowered at the statistical level. Also, the choice of the image used to illustrate this in panel 2G should be reconsidered: the image suggests that a very large loss of dopamine neurons occurred in the SNc and this is not what the numbers show. A more representative image should be used.

      In Figure 3, the authors attempt to compare intracellular calcium levels in dopamine neurons using GCaMP6 fluorescence. Because this calcium indicator is not quantitative (unlike ratiometric sensors such as Fura2), it is usually used to quantify relative changes in intracellular calcium. The present use of this probe to compare absolute values is unusual and the validity of this approach is unclear. This limitation needs to be discussed. The authors also need to refer in the text to the difference between panels D and E of this figure. It is surprising that the fluctuations in calcium levels were not quantified. I guess the hypothesis was that there should be more or larger fluctuations in the mice treated with CNO if the CNO treatment led to increased firing. This needs to be clarified.

      Although the spatial transcriptomic results are intriguing and certainly a great way to start thinking about how the CNO treatment could lead to the loss of dopamine neurons, the presented results, the focussing of some broad classes of differentially expressed genes and on some specific examples, do not really suggest any clear mechanism of neurodegeneration. It would perhaps be useful for the authors to use the obtained data to validate that a state of chronic depolarization was indeed induced by the chronic CNO treatment. Were genes classically linked to increased activity like cfos or bdnf elevated in the SNc or VTA dopamine neurons? In the striatum, the authors report that the levels of DARP32, a gene whose levels are linked to dopamine levels, are unchanged. Does this mean that there were no major changes in dopamine levels in the striatum of these mice?

      The usefulness of comparing the transcriptome of human PD SNc or VTA sections to that of the present mouse model should be better explained. In the human tissues, the transcriptome reflects the state of the tissue many years after extensive loss of dopamine neurons. It is expected that there will be few if any SNc neurons left in such sections. In comparison, the mice after 7 days of CNO treatment do not appear to have lost any dopamine neurons. As such, how can the two extremely different conditions be reasonably compared?

      Comments on the discussion:

      In the discussion, the authors state that their calcium photometry results support a central role of calcium in activity-induced neurodegeneration. This conclusion, although plausible because of the very broad pre-existing literature linking calcium elevation (such as in excitotoxicity) to neuronal loss, should be toned down a bit as no causal relationship was established in the experiments that were carried out in the present study.

      In the discussion, the authors discuss some of the parallel changes in gene expression detected in the mouse model and in the human tissues. Because few if any dopamine neurons are expected to remain in the SNc of the human tissues used, this sort of comparison has important conceptual limitations and these need to be clearly addressed.

      A major limitation of the present discussion is that it does not discuss the possibility that the observed phenotypes are caused by the induction of a chronic state of depolarization block by the chronic CNO treatment. I encourage the authors to consider and discuss this hypothesis. Also, the authors need to discuss the fact that previous work was only able to detect an increase in the firing rate of dopamine neurons after more than 95% loss of dopamine neurons. As such, the authors need to clearly discuss the relevance of the present model to PD. Are changes in firing rate a driver of neuronal loss in PD, as the authors try to make the case here, or are such changes only a secondary consequence of extensive neuronal loss (for example because a major loss of dopamine would lead to reduced D2 autoreceptor activation in the remaining neurons, and to reduced autoreceptor-mediated negative feedback on firing). This needs to be discussed.

      There is a very large, multi-decade literature on calcium elevation and its effects on neuronal loss in many different types of neurons. The authors should discuss their findings in this context and refer to some of this previous work. In a nutshell, the observations of the present manuscript could be summarized by stating that the chronic membrane depolarization induced by the CNO treatment is likely to induce a chronic elevation of intracellular calcium and this is then likely to activate some of the well-known calcium-dependent cell death mechanisms. Whether such cell death is linked in any way to PD is not really demonstrated by the present results.

      The authors are encouraged to perform a thorough revision of the discussion to address all of these issues, discuss the major limitations of the present model, and refer to the broad pre-existing literature linking membrane depolarization, calcium, and neuronal loss in many neuronal cell types.

    5. Author response:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, the authors investigated the effect of chronic activation of dopamine neurons using chemogenetics. Using Gq-DREADDs, the authors chronically activated midbrain dopamine neurons and observed that these neurons, particularly their axons, exhibit increased vulnerability and degeneration, resembling the pathological symptoms of Parkinson's disease. Baseline calcium levels in midbrain dopamine neurons were also significantly elevated following the chronic activation. Lastly, to identify cellular and circuit-level changes in response to dopaminergic neuronal degeneration caused by chronic activation, the authors employed spatial genomics (Visium) and revealed comprehensive changes in gene expression in the mouse model subjected to chronic activation. In conclusion, this study presents novel data on the consequences of chronic hyperactivation of midbrain dopamine neurons.

      Strengths:

      This study provides direct evidence that the chronic activation of dopamine neurons is toxic and gives rise to neurodegeneration. In addition, the authors achieved the chronic activation of dopamine neurons using water application of clozapine-N-oxide (CNO), a method not commonly employed by researchers. This approach may offer new insights into pathophysiological alterations of dopamine neurons in Parkinson's disease. The authors also utilized state-of-the-art spatial gene expression analysis, which can provide valuable information for other researchers studying dopamine neurons. Although the authors did not elucidate the mechanisms underlying dopaminergic neuronal and axonal death, they presented a substantial number of intriguing ideas in their discussion, which are worth further investigation.

      We thank the reviewer for these positive comments.

      Weaknesses:

      Many claims raised in this paper are only partially supported by the experimental results. So, additional data are necessary to strengthen the claims. The effects of chronic activation of dopamine neurons are intriguing; however, this paper does not go beyond reporting phenomena. It lacks a comprehensive explanation for the degeneration of dopamine neurons and their axons. While the authors proposed possible mechanisms for the degeneration in their discussion, such as differentially expressed genes, these remain experimentally unexplored.

      We thank the reviewer for this review. We do believe that the manuscript has a mechanistic component, as the central experiments involve direct manipulation of neuronal activity, and we show an increase in calcium levels and gene expression changes in dopamine neurons that coincide with the degeneration. However, we agree that deeper mechanistic investigation would strengthen the conclusions of the paper. We have planned several important revisions, including the addition of CNO behavioral controls, manipulation of intracellular calcium using isradipine, additional transcriptomics experiments and further validation of findings. We anticipate that these additions will significantly bolster the conclusions of the paper.

      Reviewer #2 (Public Review):

      Summary:

      Rademacher et al. present a paper showing that chronic chemogenetic excitation of dopaminergic neurons in the mouse midbrain results in differential degeneration of axons and somas across distinct regions (SNc vs VTA). These findings are important. This mouse model also has the advantage of showing a axon-first degeneration over an experimentally-useful time course (2-4 weeks). 2. The findings that direct excitation of dopaminergic neurons causes differential degeneration sheds light on the mechanisms of dopaminergic neuron selective vulnerability. The evidence that activation of dopaminergic neurons causes degeneration and alters mRNA expression is convincing, as the authors use both vehicle and CNO control groups, but the evidence that chronic dopaminergic activation alters circadian rhythm and motor behavior is incomplete as the authors did not run a CNO-control condition in these experiments.

      Strengths:

      This is an exciting and important paper.

      The paper compares mouse transcriptomics with human patient data.

      It shows that selective degeneration can occur across the midbrain dopaminergic neurons even in the absence of a genetic, prion, or toxin neurodegeneration mechanism.

      We thank the reviewer for these insightful comments.

      Weaknesses:

      Major concerns:

      (1) The lack of a CNO-positive, DREADD-negative control group in the behavioral experiments is the main limitation in interpreting the behavioral data. Without knowing whether CNO on its own has an impact on circadian rhythm or motor activity, the certainty that dopaminergic hyperactivity is causing these effects is lacking.

      This is an important point. Although we show that CNO does not produce degeneration of DA neuron terminals, we do not exclude a contribution to the behavioral changes. We agree that this behavioral control is necessary, and will address it in revision with a CNO-only running wheel cohort.

      (2) One of the most exciting things about this paper is that the SNc degenerates more strongly than the VTA when both regions are, in theory, excited to the same extent. However, it is not perfectly clear that both regions respond to CNO to the same extent. The electrophysiological data showing CNO responsiveness is only conducted in the SNc. If the VTA response is significantly reduced vs the SNc response, then the selectivity of the SNc degeneration could just be because the SNc was more hyperactive than the VTA. Electrophysiology experiments comparing the VTA and SNc response to CNO could support the idea that the SNc has substantial intrinsic vulnerability factors compared to the VTA.

      We agree that additional electrophysiology conducted in the VTA dopamine neurons would meaningfully add to our understanding of the selective vulnerability in this model, and will complete these experiments in revision.

      (3) The mice have access to a running wheel for the circadian rhythm experiments. Running has been shown to alter the dopaminergic system (Bastioli et al., 2022) and so the authors should clarify whether the histology, electrophysiology, fiber photometry, and transcriptomics data are conducted on mice that have been running or sedentary.

      We will explicitly clarify which mice had access to a running wheel in our revision. Briefly, mice for histology, electrophysiology, and transcriptomics all had access to a running wheel during their treatment. The mice used for photometry underwent about 7 days of running wheel access approximately 3 weeks prior to the beginning of the experiment. The photometry headcaps sterically prevented mice from having access to a running wheel in their home cage.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Rademacher and colleagues examined the effect on the integrity of the dopamine system in mice of chronically stimulating dopamine neurons using a chemogenetic approach. They find that one to two weeks of constant exposure to the chemogenetic activator CNO leads to a decrease in the density of tyrosine hydroxylase staining in striatal brain sections and to a small reduction of the global population of tyrosine hydroxylase positive neurons in the ventral midbrain. They also report alterations in gene expression in both regions using a spatial transcriptomics approach. Globally, the work is well done and valuable and some of the conclusions are interesting. However, the conceptual advance is perhaps a bit limited in the sense that there is extensive previous work in the literature showing that excessive depolarization of multiple types of neurons associated with intracellular calcium elevations promotes neuronal degeneration. The present work adds to this by showing evidence of a similar phenomenon in dopamine neurons.

      We thank the reviewer for the careful and thoughtful review of our manuscript.

      While extensive depolarization and associated intracellular calcium elevations promotes degeneration generally, we emphasize that the process we describe is novel. Indeed, prior studies delivering chronic DREADDs to vulnerable neurons in models of Alzheimer’s disease did not report an increase in neurodegeneration, despite seeing changes in protein aggregation (e.g. Yuan and Grutzendler, J Neurosci 2016, PMID: 26758850; Hussaini et al., PLOS Bio 2020, PMID: 32822389). Further, a critical finding from our study is that in our paradigm, this stressor does not impact all dopamine neurons equally, as the SNc DA neurons are more vulnerable than the VTA, mirroring selective vulnerability characteristic of Parkinson’s disease. This is consistent with a large body of literature that SNc dopamine neurons are less capable of handling large energetic and calcium loads compared to neighboring VTA neurons, and the finding that chronically altered activity is sufficient to drive this preferential loss is novel.

      In addition, we are not aware of prior studies that have chronically activated DREADDs to produce neurodegeneration. Other studies have shown that acute excitotoxic stressors can produce neuronal degeneration, but the chronic increase in activity is central to our approach.

      In terms of the mechanisms explaining the neuronal loss observed after 2 to 4 weeks of chemogenetic activation, it would be important to consider that dopamine neurons are known from a lot of previous literature to undergo a decrease in firing through a depolarization-block mechanism when chronically depolarized. Is it possible that such a phenomenon explains much of the results observed in the present study? It would be important to consider this in the manuscript.

      As discussed in greater detail in the results section below, our data suggests this may not be a prominent feature in our model. However, we cannot rule out a contribution of depolarization block, and will expand on the discussion of this possibility in the revised manuscript.

      The relevance to Parkinson's disease (PD) is also not totally clear because there is not a lot of previous solid evidence showing that the firing of dopamine neurons is increased in PD, either in human subjects or in mouse models of the disease. As such, it is not clear if the present work is really modelling something that could happen in PD in humans.

      We completely agree that evidence of increased dopamine neuron activity from human PD patients is lacking and the existing data are difficult to interpret without human controls. However, as we outline in the manuscript, multiple lines of evidence suggest that the activity level of dopamine neurons almost certainly does change in PD. Therefore, it is very important that we understand how changes in the level of neural activity influence the degeneration of DA neurons. In this paper we examine the impact of increased activity. Increased activity may be compensatory after initial dopamine neuron loss, or may be an initial driver of death (Rademacher & Nakamura, Exp Neurol 2024, PMID: 38092187). Beyond what is already discussed in the manuscript, additional support for increased activity in PD models include:

      - Elevated firing rates in asymptomatic MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488)

      - Increased frequency of spontaneous firing in patient-derived iPSC dopamine neurons and primary mouse dopamine neurons that overexpress synuclein (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060)

      - Increased spontaneous firing in dopamine neurons of rats injected with synuclein preformed fibrils compared to sham (Tozzi et al., Brain 2021, PMID: 34297092)

      We will include and further discuss these important examples in our revision.

      Similarly, in future studies, it will also be important to study the impact of decreasing DA neuron activity. There will be additional levels of complexity to accurately model changes in PD, which may differ between subtypes of the disease, the disease stage, and the subtype of dopamine neuron. Our study models the possibility of chronically increased pacemaking, and interpretation of our results will be informed as we learn more about how the activity of DA neurons changes in humans in PD. We will discuss and elaborate on these important points in the revision.

      Comments on the introduction:

      The introduction cites a 1990 paper from the lab of Anthony Grace as support of the fact that DA neurons increase their firing rate in PD models. However, in this 1990 paper, the authors stated that: "With respect to DA cell activity, depletions of up to 96% of striatal DA did not result in substantial alterations in the proportion of DA neurons active, their mean firing rate, or their firing pattern. Increases in these parameters only occurred when striatal DA depletions exceeded 96%." Such results argue that an increase in firing rate is most likely to be a consequence of the almost complete loss of dopamine neurons rather than an initial driver of neuronal loss. The present introduction would thus benefit from being revised to clarify the overriding hypothesis and rationale in relation to PD and better represent the findings of the paper by Hollerman and Grace.

      We agree that the findings of Hollerman and Grace support compensatory changes in dopamine neuron activity in response to loss of dopamine neurons, rather than informing whether dopamine neuron loss can also be an initial driver of activity. We will clarify this point in our revision. In addition, the results of other studies on this point are mixed: a 50% reduction in dopamine neurons didn’t alter firing rate or bursting (Harden and Grace, J Neurosci 1995, PMID: 7666198; Bilbao et al, Brain Res 2006, PMID: 16574080), while a 40% loss was found to increase firing rate and bursting (Chen et al, Brain Res 2009. PMID: 19545547) and larger reductions alter burst firing (Hollerman & Grace, Brain Res 1990, PMID: 2126975; Stachowiak et al, J Neurosci 1987, PMID: 3110381). Importantly, even if compensatory, such late-stage increases in dopamine neuron activity may contribute to disease progression and drive a vicious cycle of degeneration in surviving neurons. In addition, we also don’t know how the threshold of dopamine neuron loss and altered activity may differ between mice and humans, and PD patients do not present with clinical symptoms until ~30-60% of nigral neurons are lost (Burke & O’Malley, Exp Neurol 2013, PMID: 22285449; Shulman et al, Annu Rev Pathol 2011, PMID: 21034221).

      Other lines of evidence support the potential role of hyperactivity in disease initiation, including increased activity before dopamine neuron loss in MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488), increased spontaneous firing in patient-derived iPSC dopamine neurons (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060), and increased activity observed in genetic models of PD (Bishop et al., J Neurophysiol 2010, PMID: 20926611; Regoni et al., Cell Death Dis 2020,  PMID: 33173027).

      It would be good that the introduction refers to some of the literature on the links between excessive neuronal activity, calcium, and neurodegeneration. There is a large literature on this and referring to it would help frame the work and its novelty in a broader context.

      We agree that a discussion of hyperactivity, calcium, and neurodegeneration would benefit the introduction. While we briefly discuss calcium and neurodegeneration in the discussion, we will expand on this literature in both the introduction and discussion sections. We will carefully review and contextualize our work within existing frameworks of calcium and neurodegeneration (e.g. Surmeier & Schumacker, J Biol Chem 2013, PMID: 23086948; Verma et al., Transl Neurodegener 2022, PMID: 35078537). We believe that the novelty of our study lies in 1) a chronic chemogenetic activation paradigm via drinking water, 2) demonstrating selective vulnerability of dopamine neurons as a result of altering their activity/excitability alone, and 3) comparing mouse and human spatial transcriptomics.

      Comments on the results section:

      The running wheel results of Figure 1 suggest that the CNO treatment caused a brief increase in running on the first day after which there was a strong decrease during the subsequent days in the active phase. This observation is also in line with the appearance of a depolarization block.

      The authors examined many basic electrophysiological parameters of recorded dopamine neurons in acute brain slices. However, it is surprising that they did not report the resting membrane potential, or the input resistance. It would be important that this be added because these two parameters provide key information on the basal excitability of the recorded neurons. They would also allow us to obtain insight into the possibility that the neurons are chronically depolarized and thus in depolarization block.

      We do report the input resistance in Supplemental Figure 1C, which was unchanged in CNO-treated animals compared to controls. We did not report the resting membrane potential because many of the DA neurons were spontaneously firing. However, we will report the initial membrane potential on first breaking into the cell for the whole cell recordings in the revision, which did not vary between groups. This is still influenced by action potential activity, but is the timepoint in the recording least impacted by dialyzing of the neuron by the internal solution. We observed increased spontaneous action potential activity ex vivo in slices from CNO-treated mice (Figure 1D), thus at least under these conditions these dopamine neurons are not in depolarization block. We also did not see strong evidence of changes in other intrinsic properties of the neurons with whole cell recordings (e.g. Figure S1C). Overall, our electrophysiology experiments are not consistent with the depolarization block model, at least not due to changes in the intrinsic properties of the neurons. Although our ex vivo findings cannot exclude a contribution of depolarization block in vivo, we do show that CNO-treated mice removed from their cages for open field testing continue to have a strong trend for increased activity for approximately 10 days (S1E).  This finding is also consistent with increased activity of the DA neurons. We will add discussion of these important considerations in the revision.

      It is great that the authors quantified not only TH levels but also the levels of mCherry, co-expressed with the chemogenetic receptor. This could in principle help to distinguish between TH downregulation and true loss of dopamine neuron cell bodies. However, the approach used here has a major caveat in that the number of mCherry-positive dopamine neurons depends on the proportion of dopamine neurons that were infected and expressed the DREADD and this could very well vary between different mice. It is very unlikely that the virus injection allowed to infect 100% of the neurons in the VTA and SNc. This could for example explain in part the mismatch between the number of VTA dopamine neurons counted in panel 2G when comparing TH and mCherry counts. Also, I see that the mCherry counts were not provided at the 2-week time point. If the mCherry had been expressed genetically by crossing the DAT-Cre mice with a floxed fluorescent reported mice, the interpretation would have been simpler. In this context, I am not convinced of the benefit of the mCherry quantifications. The authors should consider either removing these results from the final manuscript or discussing this important limitation.

      We thank the reviewer for this insightful comment, and we agree that this is a caveat of our mCherry quantification. Quantitation of the number of mCherry+ DA neurons specifically informs the impact on transduced DA neurons, and mCherry appears to be less susceptible to downregulation versus TH. As the reviewer points out, it carries the caveat that there is some variability between injections. Nonetheless, we believe that it conveys useful complementary data. As suggested, we will discuss this caveat in our revision. Note that mCherry was not quantified at the two-week timepoint because there is no loss of TH+ cells at that time.

      Although the authors conclude that there is a global decrease in the number of dopamine neurons after 4 weeks of CNO treatment, the post-hoc tests failed to confirm that the decrease in dopamine number was significant in the SNc, the region most relevant to Parkinson's. This could be due to the fact that only a small number of mice were tested. A "n" of just 4 or 5 mice is very small for a stereological counting experiment. As such, this experiment was clearly underpowered at the statistical level. Also, the choice of the image used to illustrate this in panel 2G should be reconsidered: the image suggests that a very large loss of dopamine neurons occurred in the SNc and this is not what the numbers show. A more representative image should be used.

      We agree that the stereology experiments were performed on relatively small numbers of animals. Combined with the small effect size, this may have contributed to the post-hoc tests showing a trend of p=0.1 for both the TH and mCherry dopamine cell counts in the SN at 4 weeks. As part of the planned experiments for our revision, we will perform an additional stereologic analysis to further assess the loss of SNc dopamine neurons. We will also review and ensure the images are representative.

      In Figure 3, the authors attempt to compare intracellular calcium levels in dopamine neurons using GCaMP6 fluorescence. Because this calcium indicator is not quantitative (unlike ratiometric sensors such as Fura2), it is usually used to quantify relative changes in intracellular calcium. The present use of this probe to compare absolute values is unusual and the validity of this approach is unclear. This limitation needs to be discussed. The authors also need to refer in the text to the difference between panels D and E of this figure. It is surprising that the fluctuations in calcium levels were not quantified. I guess the hypothesis was that there should be more or larger fluctuations in the mice treated with CNO if the CNO treatment led to increased firing. This needs to be clarified.

      We thank the reviewer for this comment. We understand that this method of comparing absolute values is unconventional. However, these animals were tested concurrently on the same system, and a clear effect on the absolute baseline was observed. We will include a caveat of this in our discussion. Panel D of this figure shows the raw, uncorrected photometry traces, whereas panel E shows the isosbestic corrected traces for the same recording. In panel E, the traces follow time in ascending order. We will also include frequency and amplitude data for these recordings.   

      Although the spatial transcriptomic results are intriguing and certainly a great way to start thinking about how the CNO treatment could lead to the loss of dopamine neurons, the presented results, the focusing of some broad classes of differentially expressed genes and on some specific examples, do not really suggest any clear mechanism of neurodegeneration. It would perhaps be useful for the authors to use the obtained data to validate that a state of chronic depolarization was indeed induced by the chronic CNO treatment. Were genes classically linked to increased activity like cfos or bdnf elevated in the SNc or VTA dopamine neurons? In the striatum, the authors report that the levels of DARP32, a gene whose levels are linked to dopamine levels, are unchanged. Does this mean that there were no major changes in dopamine levels in the striatum of these mice?

      We will review the expression of activity-related genes in our dataset, although we must keep in mind that these genes may behave differently in the context of chronic activation as opposed to acutely increased activity. We will also include experiments assessing striatal dopamine levels by HPLC in the revision.

      The usefulness of comparing the transcriptome of human PD SNc or VTA sections to that of the present mouse model should be better explained. In the human tissues, the transcriptome reflects the state of the tissue many years after extensive loss of dopamine neurons. It is expected that there will be few if any SNc neurons left in such sections. In comparison, the mice after 7 days of CNO treatment do not appear to have lost any dopamine neurons. As such, how can the two extremely different conditions be reasonably compared?

      Our mouse model and human PD progress over distinct timescales, as is the case with essentially all mouse models of neurodegenerative diseases. Nonetheless, in our view there is still great value in comparing gene expression changes in mouse models with those in human disease. It seems very likely that the same pathologic processes that drive degeneration early in the disease continue to drive degeneration later in the disease. Note that we have tried to address the discrepancy in time scales in part by comparing to early PD samples when there is more limited SNc DA neuron loss. Please note the numbers of DA neurons within the areas we have selected for sampling (Figure at right). Therefore, we can indeed use spatial transcriptomics to compare dopamine neurons from mice with initial degeneration and patients where degeneration is ongoing during their disease.

      Author response image 1.

      Violin plot of DA neuron proportions sampled within the vulnerable SNV (deconvoluted RCTD method used in unmasked tissue sections of the SNV).

      Control and early PD subjects.

      Comments on the discussion:

      In the discussion, the authors state that their calcium photometry results support a central role of calcium in activity-induced neurodegeneration. This conclusion, although plausible because of the very broad pre-existing literature linking calcium elevation (such as in excitotoxicity) to neuronal loss, should be toned down a bit as no causal relationship was established in the experiments that were carried out in the present study.

      Our model utilizes hM3Dq-DREADDs that function by increasing intracellular calcium to increase neuronal excitability, and our results show increased Ca2+ by fiber photometry and changes to Ca2+-related genes, strongly suggesting a causal relation and crucial role of calcium in the mechanism of degeneration. However, we agree that we have not experimentally proven this point, as we acknowledged in the text. Additionally, we have planned revision experiments involving chronic isradipine treatment to further test the role of calcium in the mechanism of degeneration in this model.

      In the discussion, the authors discuss some of the parallel changes in gene expression detected in the mouse model and in the human tissues. Because few if any dopamine neurons are expected to remain in the SNc of the human tissues used, this sort of comparison has important conceptual limitations and these need to be clearly addressed.

      As discussed, we can sample SN DA neurons in early PD (see figure above), and in our view there is great value for such comparisons. We agree that discussion of appropriate caveats is warranted and this will be clearly addressed in the revision.

      A major limitation of the present discussion is that it does not discuss the possibility that the observed phenotypes are caused by the induction of a chronic state of depolarization block by the chronic CNO treatment. I encourage the authors to consider and discuss this hypothesis.

      As discussed above, our analyses of DA neuron firing in slices and open field testing to date do not support a prominent contribution of depolarization block with chronic CNO treatment. However, we cannot rule out this hypothesis, therefore we will include additional electrophysiology experiments and add discussion of this important consideration.  

      Also, the authors need to discuss the fact that previous work was only able to detect an increase in the firing rate of dopamine neurons after more than 95% loss of dopamine neurons. As such, the authors need to clearly discuss the relevance of the present model to PD. Are changes in firing rate a driver of neuronal loss in PD, as the authors try to make the case here, or are such changes only a secondary consequence of extensive neuronal loss (for example because a major loss of dopamine would lead to reduced D2 autoreceptor activation in the remaining neurons, and to reduced autoreceptor-mediated negative feedback on firing). This needs to be discussed.

      As discussed above, while increases in dopamine neuron activity may be compensatory after loss of neurons, the precise percentage required to induce such compensatory changes is not defined in mice and varies between paradigms, and the threshold level is not known in humans. We also reiterate that a compensatory increase in activity could still promote the degeneration of critical surviving DA neurons, whose loss underlies the substantial decline in motor function that typically occurs over the course of PD. Moreover, there are also multiple lines of evidence to suggest that changes in activity can initiate and drive dopamine neuron degeneration (Rademacher & Nakamura, Exp Neurol 2024). For example, overexpression of synuclein can increase firing in cultured dopamine neurons (Dagra et al., NPJ Parkinsons Dis 2021, PMID: 34408150) while mice expressing mutant Parkin have higher mean firing rates (Regoni et al., Cell Death Dis 2020,  PMID: 33173027). Similarly, an increased firing rate has been reported in the MitoPark mouse model of PD at a time preceding DA neuron degeneration (Good et al., FASEB J 2011, PMID: 21233488). We also acknowledge that alterations to dopamine neuron activity are likely complex in PD, and that dopamine neuron health and function can be impacted not just by simple increases in activity, but also by changes in activity patterns and regularity. We will amend our discussion to include the important caveat of changes in activity occurring as compensation, as well as further evidence of changes in activity preceding dopamine neuron death.

      There is a very large, multi-decade literature on calcium elevation and its effects on neuronal loss in many different types of neurons. The authors should discuss their findings in this context and refer to some of this previous work. In a nutshell, the observations of the present manuscript could be summarized by stating that the chronic membrane depolarization induced by the CNO treatment is likely to induce a chronic elevation of intracellular calcium and this is then likely to activate some of the well-known calcium-dependent cell death mechanisms. Whether such cell death is linked in any way to PD is not really demonstrated by the present results. The authors are encouraged to perform a thorough revision of the discussion to address all of these issues, discuss the major limitations of the present model, and refer to the broad pre-existing literature linking membrane depolarization, calcium, and neuronal loss in many neuronal cell types.

      While our model demonstrates classic excitotoxic cell death pathways, we would like to emphasize both the chronic nature of our manipulation and the progressive changes observed, with increasing degeneration seen at 1, 2, and 4 weeks of hyperactivity in an axon-first manner. This is a unique aspect of our study, in contrast to much of the previous literature which has focused on shorter timescales. Thus, while we will revise the discussion to more comprehensively acknowledge previous studies of calcium-dependent neuron cell death, we believe we have made several new contributions that are not predicted by existing literature. We have shown that this chronic manipulation is specifically toxic to nigral dopamine neurons, and the data that VTA dopamine neurons continue to be resilient even at 4 weeks is interesting and disease-relevant. We therefore do not want to use findings from other neuron types to draw assumptions about DA neurons, which are a unique and very diverse population. We acknowledge that as with all preclinical models of PD, we cannot draw definitive conclusions about PD with this data. However, we reiterate that we strongly believe that drawing connections to human disease is important, as dopamine neuron activity is very likely altered in PD and a clearer understanding of how dopamine neuron survival is impacted by activity will provide insight into the mechanisms of PD.

    1. eLife Assessment

      This study addresses an important gap in drug discovery by delivering a rigorous, large-scale evaluation of widely used co-folding methods for predicting ligand-bound protein complexes and virtual screening. A key strength is the comprehensive benchmarking framework, which leverages structures and chemical compounds that were absent from the AI models training set, thereby providing particularly compelling and unbiased evidence of co-folding performance. The findings clearly delineate the complementary roles of deep learning-based co-folding and physics-based docking, offering practical guidance for their rational integration into drug discovery workflows. Overall, the conclusions are well supported by thorough analyses across a representative set of cases.

    2. Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      Strengths:

      Overall, this is a scientifically solid paper.

      The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Comments on revised version:

      The authors have adequately addressed my concerns.

    3. Reviewer #3 (Public review):

      Summary:

      Core conclusions are well-supported by data: co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods, provides an unbiased and rigorous benchmark dataset, which contains structures and compounds absent from the co-folding models training sets. Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment. The revised results clarify an intriguing finding: co-folding can predict correct ligand poses even when protein formations are mispredicted. The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      Weaknesses:

      The study identifies a major limitation of co-folding-failure to capture rare protein conformational changes, which deserve future investigation. The authors include uncalibrated Boltz-2 affinity data (addressing a prior comment) but note that large-scale free energy perturbation (FEP) comparisons are beyond their capabilities.

      Appraisal of Aims Achieved:

      The authors successfully achieved their primary aims and the results provide strong, well-supported evidence for their core conclusions. Key conclusions are grounded in the study's unbiased, training-set independent data, ensures the conclusions are not confounded by model memorization and are broadly applicable to the field's use of these co-folding models.

      Field Impact:

      This study provides a critical reality check for the field: co-folding models are powerful tools for pose prediction but are not yet standalone solutions for virtual screening, a key distinction that will prevent over-reliance on these models and guide more rational tool selection.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors conducted a comprehensive benchmarking and evaluation of co-folding platforms, including AlphaFold3, Boltz-2, Chai-1, and the docking algorithm Dock3.7, which employs a physics-based scoring function that incorporates van der Waals interactions, electrostatics, and ligand desolvation energies. The system of interest was the SARS-CoV-2 NSP3 macrodomain (Mac1), an increasingly popular antiviral target, and the ligand sets comprised 557 unseen ligand poses (keeping the training for these co-folding platforms in mind). Additionally, the authors investigated whether the co-folding models could distinguish true ligands from non-binding small molecules. The study is thorough, with extensive statistical support and consensus across multiple metrics (chemoinformatics for quantifying ligand similarity and efficacy). The questions that the authors aim to address are whether the co-folding models struggle with memorization, whether they can distinguish between a true and a false binder, whether they replicate experimental binding affinities and efficacy, and how they compare to the physics-based docking algorithm (Dock3.7).

      We thank Reviewer 1 for this thoughtful summary of our work.

      Strengths:

      Overall, this is a scientifically solid paper. The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment.

      Weaknesses:

      My main concern is that the study's aim is a bit unclear. Modern benchmarking studies comparing physics-based docking with deep learning-based co-folding approaches (e.g., AF3, Boltz-2, Chai-1, and others) are increasingly expected to go beyond aggregate performance metrics.

      Indeed, we have gone into several examples of failures and successes for each of these methods. As we are not developing these methods ourselves, we also think this dataset will be a valuable contribution for improving them further.

      In addition to rigorous dataset construction, transparent methodology, and appropriate statistical evaluation, high-impact benchmarks typically provide actionable guidance on when each method class is most appropriate, reflecting their distinct inductive biases and practical constraints. Failure-mode analyses that link performance differences to protein flexibility, ligand chemistry, or binding-site characteristics are particularly valuable, as they move comparisons beyond "scoreboard" assessments toward mechanistic understanding.

      Right now, we do not observe meaningful trends that separate the failure modes for any individual method. This is covered in Supplementary Figures 6 and 7.

      While full biological validation is not expected, qualitative interpretation grounded in physical and biological principles strengthens conclusions. Providing reproducible workflows or reference pipelines is not mandatory, but it is increasingly viewed as a best practice because it facilitates adoption and helps contextualize results for practitioners.

      We note that our code is available (https://github.com/jongbin99/Cofolding/) and all structural data will be publicly accessible in the PDB alongside publication (we only held it back only for “blinding” during peer review to avoid contamination with any new deep learning methods).

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Kim et al. evaluates the performance of three modern AI-based methods in predicting complex structures and binding affinities between proteins and chemical compounds. An honest 'prospective' evaluation is achieved by studying benchmark structures and chemical compounds that did not exist in the PDB at the time the AI structure prediction models (AlphaFold3, Chai-1, Boltz-2) were trained.

      Strengths:

      (1) The study addresses an important question in modern computational biology and drug discovery, and establishes the strengths and limitations of the three tools in solving various computational chemistry tasks, including compound pose prediction, active-inactive discrimination, and potency ranking.

      (2) The conclusions are based on examination of four separate targets and respective compound datasets, where for one of the targets, the authors also obtained numerous X-ray structures to serve as experimental answers for the binding pose prediction task.

      (3) The study reports relationships between structure prediction confidence, predicted energies (DOCK3.7), and affinity predictions (Boltz-2) with the geometric accuracy of compound pose prediction as well as the experimentally measured potency.

      (4) One of the key findings is the limited ability of co-folding methods to predict conformational rearrangements, which does not correlate with their ability to predict binding poses of the compounds inducing these rearrangements.

      (5) The findings could serve as useful guidelines for computational chemists in selecting appropriate software and scoring schemes for each task.

      We appreciate Reviewer 2’s summary of the novelty of the dataset and analysis.

      Weaknesses:

      While I consider this a solid study, several aspects would need to be addressed to make it really strong:

      (1) DOCK3.7 docking and scoring experiments were performed using one experimental structure of Mac1, selected from dozens of structures based on a criterion that is not sufficiently well justified. For sigma2 receptor, dopamine D4 receptor, and AmpC β-lactamase, it is not clear which structures or models were selected for docking at all. It is well known that geometry predictions, scoring, and active-inactive ROC AUCs are all strongly influenced by the selected structure. It would be important to attempt Mac1 docking using all available experimental Mac1 structures, or at least against representative structures in various conformations; it would also be quite insightful to compare results to docking of the same compound sets to AF3, Boltz-2 and Chai-1 predicted structures of Mac1. Same goes for the docking studies of sigma2, D4, and AmpC β-lactamase.

      In any program, a decision has to be made as to which template will be used for docking, we justified the choice in the methods:

      “We used this structure because the inhibitor (Z5014193706) was the most potent molecule with a structure determined around the same time as the ligands in this dataset were tested.”

      We stand by this as a reasonable assumption. Similarly, for sigma2, D4, and AmpC β-lactamase, the template was chosen in the respective papers:

      a) The σ2 receptor bound to cholesterol (PDB ID: 7MFI) was used in the docking calculations.

      - This structure was determined in the paper, the first structure of sigma2 and therefore a worthy template

      b) The D4 receptor campaign used PDB 5WIU

      - This was one of two D4 structures available and chosen because it was not bound to sodium

      c) For AmpC, the campaign used the structure in the Protein Data Bank (PDB) 1L2S

      - This maximizes comparisons to other docking studies that used the same receptor template.

      The major goal of this study is to compare different methods under reasonable (but perhaps as the reviewer points out, not optimal) conditions, not to optimize docking score.

      (2) For binding affinity predictions, as a control, authors should consider compound co-folding with an unrelated protein, or even with a pseudo-peptide that consists of a few random single amino acids - this would provide an honest baseline for such predictions.

      This suggestion would be valuable for understanding the performance for these methods from the perspective of ligand specificity (a valuable, but separate, goal). Surely this will generate some number or some prediction - but what would this baseline mean and how would it be relevant for drug discovery? Therefore, we do not think this suggestion is relevant for the issues being investigated in this manuscript.

      (3) ROC curves Figure 3 and elsewhere should be shown, and AUCs quantified/reported on a log or square-root scaled x-axis, to emphasize early enrichment, which is the area of practical significance for these predictions. For example, Figure 3A currently suggests that the pose prediction performance of AF3 exceeds that of Boltz-2 whereas the early enrichment is clearly better for Boltz-2.

      We agree with this, and added a semi-logAUC plot for Figure 3A. For Figure 5, we also generated a semi-logAUC plot to see early ligand enrichment clearly, added as Supplementary Figure 11. We added the text:

      “Considering its early enrichment performance, Boltz-2 Ligand ipTM was the strongest predictor of pose accuracy based on normalized logAUC (20.5% above random, Fig. 3a). In contrast, although Boltz-2 pIC50 showed poor overall discrimination, it overestimated its ability to enrich true positive poses at low false positive rates, despite having a weak early enrichment behavior”

      (4) 'Trained set' in figures and text should probably be 'training set'? Or otherwise explain this new term the first time it is introduced.

      Thank you for pointing out this for clarification. ‘Training set’ is the correct word, and we made changes appropriately across all figures and texts.

      (5) Figure 1 illustrates a projection onto the first two principal components of a space that apparently had only one (scalar) metric for each compound pair (% maximum common substructure or Tanimoto coefficient); the authors need to better explain the principle behind this analysis and visualization.

      This suggestion is valuable, since we often use PCA to reduce dimensionality for more complex features. For clarification, we actually have a full pairwise similarity matrix for all tested Mac1 compounds based on each of Tc and MCS%. PCA for each MCS% and Tc is a representation of each pairwise similarity matrix. We also made a change in Figure 1 caption to make this point clearer:

      “projection of compounds represented by their full pairwise similarity vectors (by ECFP-4 Tc and MCS%)”

      Reviewer #3 (Public review):

      Summary:

      This study's core conclusions are well-supported by data. It is shown that co-folding outperforms docking in known ligand pose/affinity prediction (validated by RMSD and IC₅₀ correlation), struggles with false-positive discrimination in virtual screens (lower AUC values), and is complementary to docking (non-correlated errors, distinct strengths in drug discovery stages).

      Strengths:

      (1) Unprecedented prospective design with 557 novel Mac1-ligand complexes ensures rigorous, independent evaluation of co-folding methods.

      (2) Comprehensive comparison of 3 co-folding tools (AlphaFold3, Chai-1, Boltz-2) with DOCK3.7 across diverse targets and metrics enables nuanced performance assessment.

      (3) The study clearly demonstrates complementary roles of co-folding (superior pose/affinity prediction for known ligands) and docking (better hit prioritization), and addresses deep learning memorization concerns via ligand similarity analysis.

      We thank Reviewer 3 for pointing out the unprecedented and comprehensive nature of our study

      Weaknesses:

      (1) Limited generalization to diverse protein families (e.g., no ion channels/transporters).

      We agree - we have not explored the entire proteome and these are important target classes that will surely be investigated by future studies. We focused on targets here where we had large number of X-ray crystal structures (Mac1) and affinity/inhibition measurements from docking (the other three targets).

      (2) Ambiguity in the mechanism underlying co-folding's failure to predict rare conformational changes.

      Again, we agree. We are not the developers of these methods. We observe that these methods do not predict conformational changes with high fidelity and this weakness is an area that co-folding methods will surely prioritize in the future.

      (3) Virtual screen comparison is unbalanced (docking-prioritized hit lists bias results).

      We acknowledge this in the results: “An important caveat is that the hit-lists were composed of molecules prioritized by docking in the first place, giving it an advantage on these particular sets.” and discussion: “Finally, comparing co-folding to docking based on hit-lists themselves selected by docking is arguably unfair to co-folding. Counter-balancing this is the inclusion, in each of the three hit lists, of molecules that had mediocre and poor docking scores intentionally selected to test the correlation between docking score and hit-rate. Here too, the correlation between co-folding score and likelihood to bind, what we sometimes call a “dock-response-curve” was no better than docking’s, often worse (SFig.11).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are suggestions for revisions:

      (1) The writing is at times obtuse and hard to follow.

      This happens sometimes when multiple authors are writing together. We apologize and are happy to respond to specific areas that can be streamlined to be easier to follow.

      (2) In the Results section, "A set of 557 previously unreported Mac1 ligand complexes", the authors have compared the ligand poses across different metrics such as Tc - a standard, highly effective method in chemo-informatics and MCS (maximum common substructures); these are standard metrics for quantifying the structural similarity between pairs of small molecules. This part of the analysis checks whether this is memorization; it is critical to compare the two metrics, but it is not sufficient to draw a conclusion.

      Thank you for pointing out about the structural similarity of molecules co-folded to those present in the training set (resolved as Mac1 complexes and deposited in PDB before training dates). We have conducted an analysis where we do a pairwise similarity comparison for all ligands present in the PDB (regardless of the target), by both Tc and MCS, and overlay the cluster of ligands we tested (Mac1, AmpC, sigma2, D4). This should show where our tested benchmark datasets lie in the chemical space covered in the entire PDB. Each cluster (around 500 to 1300 compounds per target system) is overlaid on the cluster of all ligands deposited in PDB (over 50,000 compounds), and each cluster was relatively diverse by both Tc and MCS.

      (3) In the "Co folding can accurately reproduce poses of ligands dissimilar to those trained." Subsection under Results, the authors' conclusions are hard to follow; they state that the co-folding models often mispredict or miss the alternative conformation, but they also predict poses that are distinct from the training set. What does that imply?

      Our interpretation is actually a somewhat unsettling one: co-folding gets the ligand pose right even when it gets the protein wrong, and even when the ligand is novel. This suggests the models may be anchoring on conserved pharmacophoric interactions (like the adenosine-mimicking purine scaffold) rather than truly modeling the physics of the full complex. We added to the results section:

      This result suggests that co-folding reliably recapitulates dominant ligand-binding interactions even in the absence of accurate protein conformational modeling, providing further support to the idea that they are learning specific interaction patterns rather than a deeper physics-based representation (Masters et al. 2025).

      (4) The Discussion section connects the results and conclusions, but it can be challenging to grasp the study's overall message.

      We think the final paragraph hits on three major points:

      - Co-folding accurately predicts ligand poses for known binders, but fails to capture conformational changes

      - Co-folding does not reliably distinguish true binders from false positives in virtual screening hit lists

      - Docking and co-folding are complementary rather than competing tools

      (5) The work is highly detailed and well executed, featuring thorough data analysis and statistical assessment. The value of the paper would be further enhanced by explaining how it differs from seemingly similar results reported in other studies, including the one cited in this manuscript (see https://www.biorxiv.org/content/10.64898/2025.12.04.692352v1).

      The Mac1 results are completely unique. However, the docking datasets are exactly the same as those analyzed in the Menon et al manuscript. We don’t think our results differs from conclusions of the Menon et al manuscript as we wrote: These observations are supported by a fascinating study on some of the same ligand sets as investigated here, using AlphaFold3, reaching similar conclusions (Menon et al. 2025).

      Reviewer #3 (Recommendations for the authors):

      (1) Expand target diversity to include ion channels, transporters, etc., beyond enzymes and GPCRs.

      (2) Investigate the cause of co-folding's failure in predicting rare conformational changes (e.g., adjust sampling, MSA inputs, or add experimental constraints).

      (3) Mitigate docking bias in virtual screens (e.g., re-analyze unbiased compound libraries).

      We addressed these three points in the public review above

      (4) Test Boltz-2's affinity predictions without linear calibration and compare with FEP.

      The data without linear calibration are included in the manuscript. Comparing such a large number of compounds with FEP is currently beyond our capabilities.

      (5) Conduct proof-of-concept to test co-folding-docking integration for better hit rates.

      We think this is well beyond the scope of this manuscript - but look forward to testing this idea in the future.

      We also got one community review that we respond to below:

      Summary

      This manuscript evaluates the performance of co-folding models when tasked with 1) the recapitulation of a large number of experimentally determined co-crystal structures of Mac1 with a series of Mac1 ligands and 2) the rescoring of hits to identify false positives originally derived from a set of large docking-based virtual screens. The evaluation leverages a dataset of crystal structures and affinity data from high-throughput crystallographic and biophysical screens, respectively. These data uniquely enable this report to focus on the ability of co-folding models to handle ligands, resulting in an analysis that is particularly timely given the wide adoption of co-folding models and the relative scarcity of such ligand-focused benchmarks among existing evaluations, which have primarily focused on protein structure prediction or binder design.

      Thank you for this thoughtful summary of our work

      Feedback

      The experiments and analyses in the manuscript are well thought-out and do not have any significant issues. There are a few high-level points that may improve the clarity and completeness of the results. Importantly, none of the suggested additional experiments will affect the conclusions of the paper, but rather help provide additional context for the results:

      The first section presents an exciting opportunity to frame the Mac1 ligands against ligands in the PDB more broadly. It would be informative to assess whether chemotypes that are easier or harder to predict accurately and confidently are over- or under-represented in the PDB as a whole. Note that this is not a recommendation that new scaffold similarity metrics be incorporated into the analysis, but rather that analyses similar to those already performed in the manuscript are performed using all ligands in the PDB. For example, PCA-based analyses similar to those in Fig. 1c could be used to examine Mac1 ligands in the context of all PDB ligands enabling questions such as whether similarity to a nearest PDB neighbor, cluster size in a Tc/MCS PCA space, or other frequency-based measures show any relationship with prediction vs. crystal structure RMSD. Such analyses could provide additional insight into how effectively models leverage ligand information present in the PDB overall, as opposed to biases arising specifically from scaffolds represented in Mac1 structures in the PDB, which are already well covered in the manuscript. The conclusion that Tc/MCS do not correlate with the ligand RMSDs for the ligands already associated with the Mac1 is well supported, and presumably suggests that a correlation would not exist against the backdrop of the PDB, but it would be interesting to see the data using analyses similar to those already done in the manuscript nonetheless.

      We are adding new figures in SFig.1 that consider how different clusters of ligands tested for our co-folding analysis are distributed across the chemical space in PDB. This is done by making a similarity comparison between every ligand in PDB and those tested in our analysis by Tc and MCS%, then plotting in PCA space for each metric. We are excited to see that each dataset covers a wide scope in PCA space, but at the same time, there are unexplored areas in the chemical space of PDB by co-folding.

      Similarly, even though the four proteins used in this manuscript are not themselves the primary focus of the analysis, it would be valuable to perform a high-level assessment of the precedent for each protein in the PDB (beyond the count of liganded structures in Table S6), either in protein sequence space (e.g., MSAs) or structural space (e.g., FoldSeek). An analysis like this would provide important context about whether any of the proteins in the study have close homologs with liganded structures in the PDB, or are generally overrepresented in the PDB. The fact that the AUC for L-pLDDT for AmpC is higher than σ2 and D4, for example, is notable given the relative abundance of liganded AmpC structures in the PDB (this raises potentially interesting questions related to where DOCK3.7 and AF3 actually place the ligands, given the orthosteric β-lactam binding pocket in AmpC, although this is outside of the scope of this manuscript).

      High-level assessment of the precedent for each protein in the PDB will definitely help to understand if proteins we used have close homologs with liganded structures in the PDB. Our Supplementary Table 6 covers the extent to which these liganded structures were available by cutoff dates for AF3, Chai-1 and Boltz-2. AmpC had more homologs than sigma2 and D4, and this may explain a better AUC for AF3 L-pLDDT specifically for this target.

      A discussion of the affinity probability results (`affinity_probability_binary`) from Boltz-2 is likely warranted in the second section in addition to the pIC50s that are already reported (`affinity_pred_value`). The former seems like it would be more applicable for section 2 of the manuscript, but both warrant inclusion—they should both be calculated by default when the affinity pipeline in Boltz-2 is turned on, so it wouldn't involve any more inference.

      As boltz-2 affinity module outputs both affinity probability binary output and affinity predicted value, we kept track of both metrics. So we tried re-ranking hit lists using both metrics. Where boltz-2 performed better (Sigma2, D4), binary probability values were more representative as a metric to differentiate true actives from non-binders. This was more clear in semi-logarithmic ROC plots. However, in AmpC, both Boltz-2 scoring metrics performed similarly. Such inconsistency in trend made it difficult to draw conclusions.

      Minor points

      A more detailed description of the experimental methods used to generate the ground-truth data in the introduction (even though these have been explained in prior works) would help orient the reader early on, and ground the benchmarking aspect of the story. In general, the abstract and introduction would benefit from a more cohesive through-line to tie the two complementary but orthogonal sections of the paper together.

      We will include a more thorough description alongside the PDB depositions. As for the two sections, we have tried to tie them together from the perspective of drug discovery workflows…

      The cutoffs in the "Co-folding can accurately reproduce..." section shift between 2.5 Å (from the ligand center of mass) and 2.0 Å. Is there a reason for this? Along similar lines, mentioning cutoffs for true positives/negatives when introducing the ROC analyses later on in the Mac1 section seems unnecessary since no cutoff should be necessary here.

      We used 2.5A distance to COM to just get at “broadly the correct binding site” for fast filtering and 2.0A RMSD because that is the broadly accepted standard in the field for “relatively correct binding pose”.

    1. Slack OAuth + token storage — slack-sender-service settles delivery, but the per-account "Connect Slack" install + token store is still net-new (and the service must be extended to resolve a token per account). Where do tokens live & who builds the install flow?

      Im not sure we want to do it as part of this task.

    2. so customer Slack routes through slack-sender-service instead.

      maybe need to call to slack from notification-service, in order to keep all things in one place. and user slack-sender-service as a provider

    3. dispatch audit

      i cant find another place to raise my thought about the audit: I think we need to change it to log what exact sent, show alert as row and drill down of users and channels (maybe adding status from providers). today it very hard to read it and not connect to mute channels / scheduler etc... It not part of this effort of course but I want to think like that when we build the system.

    4. 02 · CONTENT (render-context)Notification content arrives as a dynamic key/value object the realtime alert computed at trigger time — point-in-time-correct, no content fetch. Enrichment shrinks to recipient resolution (03) plus a rare fallback lookup via references[].

      I didn't understand what was meant here.

    1. Formulación

      antes de esta sección , debe ir una sección sobre ecuaciones diferenciales con difusión y saltos. Hay que presentar el lema de Itô con difusión y saltos, la fórmula de Itô con difusión y saltos y el teorema de existencia y unicidad para EDE con difusión y saltos. Tal vez bastaría con cambiar el nombre de la sección e introducir lo que hace falta

    2. . E[∫0TX(t)dW(t)]=0.

      antes de este tema, presentar las principales propiedades de la integral de Itô, entre ellas el lema de Itô y el hecho de que la integral respecto al movimiento browniano es una martingala

    3. continuo.

      aqui vamos a agregar la fórmula de Itô el caso general en una sola dimensión. También las operaciones del cálculo de Itô

    1. home

      This is a really helpful section overall. I think it leads well to an examination of social learning theories, which we focus a lot on in our criminology classes!

    1. The paper relies on several assertions about languages and classical texts which are highly questionable, but treated somewhat misleadingly as if they represent an academic consensus. For example, Jordanes in the 6th century (not named, but relied upon) indicated himself that he preferred classical written sources, and philogists today are critical of the old idea that he used oral traditions for his ideas about Scandinavia. His interest in speculating and putting together fantastic stories, including Bible-derived ideas about barbarians coming from islands in the north is also recognized. Furthermore, later Scandinavian migration stories all appear to be derived from his, whereas here they are treated here as separate traditions. See for example the detailed study by Arne Søby Christensen and the 1980s studies by Peter Heather. Concerning languages, the idea that Germanic already existed in Scandinavia before the Jastorf culture is an old tradition but by no means a field consensus! In reality there is no real evidence for this, and scholars like Kulikowski argue quite reasonably that such ideas would not exist without Jordanes. Normally by the way it is Southern Scandinavia which was traditionally associated with the pre-Goths, not Eastern Scandinavia. I find it worrying that the article is overall written in a confused and opaque way which seems to hide the normal interpretations of the relevant non-genetic evidence. For example, how can the article disprove a continental origin to Germanic if it does not even mention that this is the normal scholarly proposal? The article should at least mention the most widely held ideas, and ideally it should also test them. Perhaps connected to this the language used is sometimes so over-complicated and unnatural that it is ambiguous and arguably meaningless.

    1. Sutherland positioned himself against scholars and policymakers who emphasized eugenics as a way to improve American society

      This is a huge point that needs to be discussed in this book somewhere, as criminology has such a deep and problematic history associated with eugenics. Thinking specifically of Cesare Lombroso and the Criminal Man as an example of this.

    1. Pretrial Fairness Act

      I know there was a great deal of activism ahead of this Act. Would it be possible to talk a bit more about this example and link it to the inequalities literature in sociology/criminology? This would be a great place to talk about class and racial inequalities in the justice system and in incarceration! How do criminologists think about these inequalities? What does the Pretrial Fairness Act end up doing for residents and how does it address inequalities? If not included in this textbook, I think this is another place where external resources would be needed, paired with the book!

    1. Illinois

      In my experience, this chapter would fit more closely with our Criminal Justice System course (an upper-level undergrad course) rather than our criminology course. However, I think this is still useful as an introduction for students to think about the local criminal justice system. Whether offered in this book or using external resources, i would also want students to consider how criminologists think about/study the court system more broadly. That connection is not as present in the section.

    1. eLife Assessment

      This study identifies apoptotic retinal ganglion cells as a potential source of ATP-mediated activation of PANX1 channels that initiates developmental retinal Ca²⁺ waves and coordinates microglial activation and vascular outgrowth during postnatal maturation. The work is important because it proposes an integrative framework linking programmed cell death, spontaneous neural activity, immune responses, and angiogenesis into a self-regulating developmental loop. Although the mechanistic conclusions would benefit from complementary genetic validation, the study provides a convincing foundation for future investigations into the coordination of neural circuit development and tissue remodeling.

    2. Reviewer #1 (Public review):

      Summary:

      This study presents a potentially important integrative model linking spontaneous retinal waves, apoptosis, microglial activity, and vascular development during postnatal retinal maturation. Its significance lies in proposing a mechanistic framework that could reshape understanding of how neural activity and tissue remodeling are coordinated in the developing central nervous system. The evidence is strengthened by the use of multiple complementary techniques, including Ca++ imaging, high-throughput electrophysiology, transcriptomics, histology and pharmacology.

      Strengths:

      (1) Multimodal Validation: The authors correlate large-scale functional imaging (calcium imaging and MEA) with high-resolution structural and molecular data (scRNA-seq and IHC), providing strong topographical evidence for the "centrifugal expansion" pattern.

      (2) The primary significance lies in identifying apoptotic Retinal Ganglion Cells (RGCs) as the physiological "pacemakers" for stage II retinal waves. By linking programmed cell death directly to neural activity and subsequent angiogenesis, the authors propose a self-regulating developmental loop.

      Weaknesses:

      (1) While the PANX1 pharmacological data provides compelling functional support, extending these conclusions to the broader CNS may be premature. Additional direct mechanistic validation would further strengthen the claim of causality.

      (2) While the manuscript beautifully illustrates the co-occurrence of events during retinal development, strengthening the distinction between correlation and direct causation would enhance the impact of the findings.

      Appraisal of Aims and Conclusions:

      The authors successfully achieve their aim of presenting a cohesive, multi-layered framework for postnatal retinal maturation, aligning functional physiological data with structural and transcriptomic timelines. The data robustly supports the correlation between retinal waves, microglial activity, and vascular remodeling and also identifies apoptotic RGCs as the potential "pacemakers" of Stage II waves.

      Impact, Utility, and Community Asset:

      This work will significantly impact developmental neurobiology by reframing programmed cell death as an active, instructive driver of neural network patterning and angiogenesis, rather than a passive clearance process. Methodologically, the integration of large-scale MEA recordings and live calcium imaging with scRNA-seq sets an excellent benchmark for multimodal developmental studies. Furthermore, the transcriptomic datasets mapping microglial phenotypes and vascular remodeling will serve as a highly valuable reference repository for the broader visual neuroscience community.

      Additional Context for Readers:

      To fully appreciate this study, readers should view it through the lens of neurovascular unit assembly. While Stage II cholinergic waves are traditionally studied purely in the context of visual circuit refinement, this work adds vital context by showing they also regulate the surrounding metabolic ecosystem. It effectively demonstrates that early electrical activity, programmed cell death, and vascular scaffolding do not occur in isolation, but are deeply interdependent processes.

    3. Reviewer #2 (Public review):

      Summary:

      Savage et al. investigates the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Using scRNAseq, the authors classify autofluorescence cell clusters (ACCs) at the leading edge of vasculature outgrowth as Hmox-1+ microglia. From here they show microglia engulfment of apoptotic RGCs and the potential release of ATP may contribute to Ca2+ wave generation. The authors demonstrate these mechanisms through the use of two pharmacological to agents to either block the ATP release from Panx-1 or by blocking receptor binding to ATP. Furthermore, while previous studies have described the site of initiation of retinal Ca2+ waves as random, this study shows the initiation of Ca2+ waves are biased to the leading edge of vascular growth in the developing retina. To do this, the authors use a combination of wide-field Ca2+ imaging and multi-electrode arrays to pinpoint the sites of Ca2+ wave initiation in the developing retina.

      Strengths:

      Savage et al. uses a several techniques to interrogate these mechanisms, including single cell RNAseq, wide-field Ca2+ imaging, and multi-electrode arrays. With these experiments, this manuscript proposes several novel ideas, such as ATP as the Ca2+ wave initiating cue, and the localization the Ca2+ wave initiation to the leading edge of vascular growth.

      Weaknesses:

      The main limitation of this study is the reliance on only two pharmacological agents to test their central hypotheses. In future studies, these conclusions could be strengthened if they used genetic knockout models to perturb programmed cell death and/or ATP release (i.e. BAX-KO, Panx-1 KO).

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents a potentially important integrative model linking spontaneous retinal waves, apoptosis, microglial activity, and vascular development during postnatal retinal maturation. Its significance lies in proposing a mechanistic framework that could reshape understanding of how neural activity and tissue remodeling are coordinated in the developing central nervous system. The evidence is strengthened by the use of multiple complementary techniques, including Ca++ imaging, high-throughput electrophysiology, transcriptomics, histology, and pharmacology.

      Strengths:

      (1) Multimodal Validation: The authors correlate large-scale functional imaging (calcium imaging and MEA) with high-resolution structural and molecular data (scRNA-seq and IHC), providing strong topographical evidence for the "centrifugal expansion" pattern.

      (2) The primary significance lies in identifying apoptotic Retinal Ganglion Cells (RGCs) as the physiological "pacemakers" for stage II retinal waves. By linking programmed cell death directly to neural activity and subsequent angiogenesis, the authors propose a self-regulating developmental loop.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      (1) While the PANX1 pharmacological data provide compelling functional support, extending these conclusions to the broader CNS may be premature. Additional direct mechanistic validation would further strengthen the claim of causality.

      We agree with the reviewer that the conclusions would be greatly solidified with more direct mechanistic validation. However, we are unable to conduct more experimentation as the grant is finished and the Sernagor lab is in the process of being shutdown, after the unexpected passing of the PI.

      In order to make clearer that this mechanism was found in retinal tissue, not CNS, we have moved any mention of the implications of our work to a broader CNS mechanism to the discussion section. We have also added text into the discussion highlighting the need for more mechanistic investigation to uncover the full extent of the developmental processes described herein, see Line 413.

      (2) While the manuscript beautifully illustrates the co-occurrence of events during retinal development, strengthening the distinction between correlation and direct causation would enhance the impact of the findings.

      We have been clear to only present our findings as correlational as we were unable to fully explore the causational nature within the mechanisms presented. In the discussion, we have used published evidence and experimental papers to bolster our understanding of the causal aspects of this research. We have also included sections of text to address what experimentation is be required to examine the causal interactions more directly, see Line 413.

      Reviewer #2 (Public review):

      Summary:

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Using scRNAseq, the authors classify autofluorescence cell clusters (ACCs) at the leading edge of vasculature outgrowth as Hmox-1+ microglia. From here, they show microglia engulfment of apoptotic RGCs, and the potential release of ATP may contribute to Ca2+ wave generation. The authors demonstrate these mechanisms through the use of two pharmacological agents to either block the ATP release from Panx-1 or block receptor binding to ATP. Furthermore, while previous studies have described the site of initiation of retinal Ca2+ waves as random, this study shows that the initiation of Ca2+ waves is biased to the leading edge of vascular growth in the developing retina. To do this, the authors use a combination of wide-field Ca2+ imaging and multi-electrode arrays to pinpoint the sites of Ca2+ wave initiation in the developing retina.

      Strengths:

      The authors use several techniques to interrogate these mechanisms, including single-cell RNAseq, wide-field Ca2+ imaging, and multi-electrode arrays. With these experiments, this manuscript proposes several novel ideas, such as ATP as the Ca2+ wave-initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      The main weakness of the manuscript is the overreliance on only two pharmacological agents to test the central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP release (i.e., BAX-KO, Panx-1 KO).

      We thank the reviewer for their insightful suggestions for further experimentation to bolster the research. Initially, we utilised pharmacological interventions as they provided acute and quick answering of the research question. At the outset of the research, we were not certain that purinergic release through PANX-1 channels was the mediator for the developmental mechanisms described. We tested a wide variety of specific agonists and blockers before seeing any profound effects on wave generation. These agonists and antagonists have been used before and are proven to deliver reliable results. In addition, since the ACCs had never been reported before we were unsure if a knockout animal would display the same anatomical phenotype. Furthermore, it is known that knockout mouse lines, especially connexin and hemichannel pores, do not lose function but rather have other isoforms or compensation mechanisms which can substitute the original function. For the retina, for example, it was shown that Cx36 can functionally replace Cx45 after Cx45 KO (Frank et al, 2010).

      We agree that while direct mechanistic validation would significantly reinforce the arguments, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      In order to address the omission of mechanistic validation in the paper we have added text into the discussion highlighting the need deeper investigation in the causality of the developmental processes described herein, see Line 413.

      M. Frank et al., Neuronal connexin-36 can functionally replace connexin-45 in mouse retina but not in the developing heart, J. Cell Sci. 123, 3605 (2010).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      General and major comments

      (A) Introduction

      (1) The introduction is currently quite extensive. I recommend streamlining the background information to more directly frame the study's core objectives.

      We have reduced the background information contained in the introduction to better align with the direct outputs of the study. However, as this research paper examines the interactions of multiple complex developmental processes a relatively in-depth introduction is needed to inform the reader of the salient points.

      (2) To improve clarity, it would be highly beneficial to conclude the introduction with a sequential summary of key observations in the order they are presented in the study. This would provide a clearer roadmap for the reader.

      We have reworked the end of the introduction to better align with the key observations in the order they are presented through the figures.

      (B) Results

      (3) The characterization of ACCs (Apoptotic Cell Clusters) would be more effective if separated from the description of their spatiotemporal occurrence with SVPs. Consider moving the microscopic description to a dedicated section or merging it with the subsequent paragraph on molecular identification.

      We thank the reviewer for their suggestion to reorganise the results section. However, we believe that highlighting the integrated nature of the ACC positioning and development of the vascular plexus is an important stepping off point for the reader. It allows us to highlight the original serendipitous discovery of the ACCs and their highly cohesive role in bridging multiple developmental processes.

      (4) Please explicitly state the specific retinal developmental stages (e.g., P3-P6) within the section regarding RNA sequencing, as this timing is critical for interpreting the transcriptomic data.

      We have amended the text to specify the postnatal days used for RNA sequencing, see Line 123

      (5) Regarding the scRNA-Seq data: were ACC-negative samples (isolated via FACS) also processed? A direct comparison between ACC+ and ACC- sampled microglia would significantly strengthen the claim that microglia are specifically attracted to ACCs. If these data are available, they would make an elegant and compelling addition to the manuscript.

      We thank the reviewer for this important suggestion and agree that direct comparison of ACC+ and ACC− microglia would further strengthen the study. Unfortunately, ACC− populations were not processed for scRNA-seq in the current study because of the prioritisation of the rare ACC+ population. Nevertheless, several independent observations support the conclusion that microglia are preferentially associated with ACCs, including: (i) the enrichment of microglia within ACC-containing regions observed histologically, (ii) the spatial proximity analyses shown in Figure 3, and (iii) the distinct transcriptional profile of ACC-associated microglia identified by scRNA-seq.

      We have also added a section to the discussion to highlight the need for a direct comparison of the ACC+ and ACC- transcriptomics profiles in future work, see Line 344.

      (6) For the Ca2+ -imaging experiments, please briefly describe the staining protocol and specify which cell types (e.g., RGCs) were labeled within the Results text to assist the reader's immediate understanding.

      We have added a short description of the labelling technique in the results, see Line 228

      (7) The manuscript notes that MEA waves are evident in the graphs of Figure 7, but the raw wave data or representative traces are not shown. Including these (similar to those of the imaging waves) would provide necessary visual verification of the physiological phenomena described.

      We have added a supplementary figure 3 which details stage 2 retinal waves recorded using MEAs.

      (C) Interpretations and Logic

      (8) The finding ' ...Wholemount staining revealed a broad centro-peripheral gradient of apoptosis; however, this apoptotic annulus was positioned more peripherally than the ACCs, SVP, and Hmox1-positive microglia (Figure 4B)' seems to contradict the HMOX1/Yo-Pro-1 stained microglia. It is not clear whether the authors aim to prove that this particular set of microglia phagocytise RGCs, or another set that lines up better with dying cells and does not show up on the HMOX-1 label. I believe the authors intend to show the time difference of the two events - cells dying and HMOX-1 microglia appear at the site later. I believe the logic is good; it may need a sentence pointing this out at the end of this paragraph.

      We have added a statement in Line 182 which clarifies our intent to show that the Hmox1 microglia phagocytose the dying RGCs after they initiate apoptotic mechanisms.

      (9) If the authors intend to demonstrate a temporal lag between cell death and the appearance of HMOX1+ microglia as evidence of causality, a concluding sentence to this effect would greatly clarify the logic of this paragraph.

      We have added a concluding sentence to the paragraph in Line 190 which indicates a causal link between RGC cell death and appearance of hmox1 positive microglia.

      (D) Figures and Presentation

      (10) The blood vessel staining in the final panel of Figure 1B is currently quite faint. Increasing the brightness/contrast for this panel would allow the reader to better appreciate the underlying architecture.

      We have updated the panel in Figure 1B to match the brightness of the others of that series.

      (11) Given that the peripherality of events in Figure 7 suggests a specific sequence, the authors should consider adding a summary timeline (P3-P6). A plot using curves (mean or median values), color-coded to match the corresponding events, would provide a much-needed visual synthesis of the data.

      We agree with reviewer that the D1/2 metrics would benefit from more clarification to show the timeline of development more clearly. We have added another panel to Figure 7, which shows mean/standard deviation plots for each developmental measure using D1/2 as timelines. This allows the reader to better compare the progression of centrifugal spread more clearly.

      (12) Please ensure that graph labels and axis titles are uniform in size across all figures. e.g., the labels in Figure 6 and several other graphs are currently too small to be legible in the PDF; these should be enlarged for better accessibility.

      We have fixed the labels and axis titles to maintain readability across the paper

      Minor comments

      (1) In line 125, there is a missing closing parenthesis after the reference to Figure 1C.

      We have fixed this error

      (2) The specific algorithm used for the unsupervised cluster analysis has not been identified in the text. Please specify whether k-means, Louvain, or another method was employed to ensure reproducibility.

      We have reworked the section detailing the cluster analysis to make clear we used Louvain-based clustering, see line 475.

      (3) While it is appropriate to leave comprehensive technical details for the Methods section, a brief conceptual explanation of the D1, D2, and D3 metrics should be included in the Results text to aid general comprehension.

      We have added a brief description of the D1/2 and D1/3 metrics when they are first mentioned in the results section, see line 243.

      (4) In the Figure 7 schematic, the representation of the starburst amacrine cell should be revised to more accurately reflect its well-characterized morphology (e.g., thin primary and gradually thickening higher-order dendrites).

      We have changed the SAC representation to better match the characteristics of that cell type.

      (5) Throughout the manuscript (e.g., in lines 293-294), it is claimed that RGC apoptosis promotes the expression of PANX-1 hemichannels. While the data effectively demonstrate the release of purinergic molecules (e.g., ATP) via PANX-1 from dying cells, the evidence for an actual upregulation or increase in PANX-1 protein/mRNA levels is not explicitly shown. Please clarify whether the findings suggest increased activity of existing channels or a true increase in expression. If the latter is not empirically supported, the phrasing should be adjusted to reflect functional activation rather than de novo expression.

      We agree with the author that our research shows a functional increase in PANX-1 and we have adjusted the language to match. In the introduction and discussion, we provide published evidence and experimental papers which describe the upregulation of the PANX-1 molecule in dying RGCs.

      Reviewer #2 (Recommendations for the authors):

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Furthermore, the authors demonstrate the initiation of Ca2+ waves occurs at the leading edge of vascular growth in the developing retina. This manuscript proposes several novel ideas, such as ATP as the Ca2+ wave initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth. The main weakness of the manuscript is the overreliance on only two pharmacological agents to test their central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP releases (i.e., BAX-KO, Panx-1 KO). In addition, the following comments should also be addressed:

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Specific comments:

      (1) Line 128: Why would ACCs be involved in SVP guidance if they are trailing the leading edge of the vasculature? Would it be the other way around, where the leading edge of the vasculature would be trailing the ACCs?

      At this point in the paper we are suggesting that the highly stereotyped position of the ACCs under the leading edge of the SVP indicates that they have a mechanistic involvement in SVP growth. Not that they are the direct cause of the expansion. As the paper progresses, we make clear that contrary to our original hypotheses which state the ACCs may cause or control the integrated development of the retina, they are a hallmark of the apoptotic RGCs in the periphery which are the chemogenic beacons for vascular growth being ‘decommissioned’ by the microglia which fine-tune vascular growth and create the ACCs.

      (2) Line 137: Please state in the text and figure legend, at what age ACCs were isolated from the retina.

      We have added the relevant information to Line 123

      (3) Line 154: Why didn't RGCs form their own cluster? Why do the RBPMS+ cells appear across the entire dataset (Figure 2G)? Have previous investigations also shown engulfed cell transcriptomes appearing in the microglia clusters using scRNAseq?

      Previous transcriptomic studies have demonstrated that phagocytic microglia can encapsulate transcripts originating from neurons and other neural cell types, which are detectable by RNA‑seq despite not belonging to a common microglial genetic signature. For example, Solga et al. showed that CNS microglia contain neuronal and oligodendrocyte‑specific mRNAs that localise within microglia but are not translated. This research group interpreted that this RNA is acquired through phagocytosis or macropinocytosis of surrounding neural cells (Solga et al., 2015). Similarly, in zebrafish, synapse‑engulfing microglia identified in situ display neuronal and synaptic gene expression in single‑cell RNA‑seq profiles, consistent with engulfed neuronal material contributing to the detected transcriptome (Sliva et al., 2021). In line with these observations, and given our FACS strategy enriching autofluorescent ACCs rather than intact RGCs, we interpret the widespread Rbpms expression across ACC‑associated clusters as an expected consequence of microglial engulfment of apoptotic RGCs, rather than evidence for a distinct population of viable RGCs that failed to form a separate cluster.

      Solga, A.C., Pong, W.W., Walker, J., Wylie, T., Magrini, V., Apicelli, A.J., Griffith, M., Griffith, O.L., Kohsaka, S., Wu, G.F. and Brody, D.L., 2015. RNA‐sequencing reveals oligodendrocyte and neuronal transcripts in microglia relevant to central nervous system disease. Glia, 63(4), pp.531-548.

      Silva, N.J., Dorman, L.C., Vainchtein, I.D., Horneck, N.C. and Molofsky, A.V., 2021. In situ and transcriptomic identification of microglia in synapse-rich regions of the developing zebrafish brain. Nature communications, 12(1), p.5916.

      We have incorporated this information into the discussion in Line 329

      (4) Figure 4C-H: The data in the figure would be strengthened if the authors added quantification for their co-localization images.

      We thank the reviewer for this important suggestion and agree that quantification of the co-localisation would further strengthen the study. Unfortunately, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      (5) Figure 5A-I: The error bars are quite large. The figure legend says they represent SEM, but are you sure they don't represent standard deviation (SD)?

      The large SEM bars reflect substantial biological variability across retinas and across individual microglia, which is expected for morphometric measures such as circularity, perimeter, branch number and total skeleton length, as well as for counts of rare cell populations (Hmox1+ microglia, YO‑PRO‑1+ cells and double‑positive cells). Importantly, despite this variability, the effects of probenecid and PSB‑0739 on microglial morphology and on the frequencies of apoptotic and double‑positive cells remain statistically robust in our non‑parametric ANOVA and post‑hoc tests, as indicated by the reported P‑values.”

      For panels where the distributions were clearly non‑Gaussian, we used non‑parametric statistics (Kruskal‑Wallis ANOVA), reporting medians and 95% confidence intervals, and we retained SEM in Figure 5 for consistency with the original plotting routine while clarifying this choice in the legend and Methods.

      (6) Figure 5: What are these measurements made at? Please add the age to the results section and the figure legend.

      We have added the postnatal day of the animals used.

      (7) Line 236-239: There is no mention of the use of Probenecid in the text and in Figure 6B. This treatment of Ca2+ waves should be mentioned before line 245.

      We have added a brief description of probenecid application in Line 227.

      (8) Figure 6: The order of this figure and the results section may be clearer if panels C-F were switched with panels G-J.

      We thank the reviewer for their suggests to improve the flow of the results section and figure 6. We have swapped the panels as suggested and amended the text to fit the new flow.

      (9) For the discussion section: Why does probenecid only affect the wave initiation at the P3 timepoint, but not later timepoints, P4-P6.

      We have added some text in the discussion to explain our findings of differential effects of PANX-1 blockade across the P3-6 timeline. See Line 404.

      Editorial revisions:

      (1) Line 92-96: Citation needed.

      We have added appropriate citations for this section.

      (2) Figure 1F: Please consider changing the color scheme from green/red to green/magenta for colorblind readers.

      We have changed the image LUT

      (3) Figure 3A: This figure may benefit from separating out each channel separately (Iba1 and ACC) and then having a "merge" panel.

      We have separated out the panels in Figure 3A

      (4) Figure 4E, H: There is no label for the immunomarkers in Figure 4H. There is no inset in Figure 4E as mentioned in the figure legend. However, it appears that Figure 4H is a magnified image of Figure 4G and not Figure 4E.

      We have fixed the error in the figure legend and included an inset indicator in Figure 4G. We have added immunomarkers to Figure 4H

      (5) Line 198: "SAC" should be "SACs".

      We have fixed this error

    1. eLife Assessment

      This convincing contribution addresses a question of practical importance: when collecting tilt-series data, what is the optimal angular step size between successive tilt images? The work provides valuable practical insights into cryo-ET data acquisition by demonstrating that balancing two competing demands - sufficient dose per individual tilt image and fine angular sampling - is essential to achieve high-quality tomographic reconstructions. They demonstrate that tilt-series acquired with finer increments (1-3 degrees) yield superior alignment accuracy and improved template-matching performance.

    2. Reviewer #1 (Public review):

      This work addresses a question of practical importance that had never been systematically analysed in the cryo-ET field: when collecting tilt-series data, what is the optimal angular step size between successive tilt images? Due to the upper limit in electron exposure (100 - 150 e⁻/Ų), this question is important, since finer angular sampling improves attainable reconstruction resolution (Crowther criterion) but reduces the signal-to-noise ratio of each individual image, potentially compromising both image quality and the ability to computationally align successive frames. To address this, the authors designed a thorough benchmarking study comparing five tilt increments (1{degree sign}, 2{degree sign}, 3{degree sign}, 5{degree sign}, and 10{degree sign}) while keeping the total dose and tilt range constant. They evaluated the consequences at every stage of the cryo-ET workflow - from raw image quality and tilt-series alignment, through template matching for ribosome detection, to high-resolution subtomogram averaging - with the goal of providing the community with an evidence-based recommendation for data acquisition.

      The manuscript is well written, and the experimental design is carefully thought out. The work provides valuable practical insights into cryo-ET data acquisition by demonstrating that balancing two competing demands - sufficient dose per individual tilt image and fine angular sampling - is essential to achieve high-quality tomographic reconstructions. The identification of a practical optimum at 3{degree sign} tilt increment is the key contribution of the work. It will be interesting to see in the future whether this optimum shifts for smaller molecular targets, and how emerging tilt interpolation strategies such as cryoTIGER may interact with the choice of experimental angular increment.

      Comments on revised version.

      Well done! I really like the manuscript and from my point of view it's an excellent piece of work and super useful for the community. Thank you so much for the meticulous work!

    3. Reviewer #2 (Public review):

      The determination of macromolecular structures directly within their native cellular environment is becoming increasingly routine, making standardized data collection strategies essential. In this manuscript, Tuijtel et al. provide a timely and valuable contribution by benchmarking key acquisition parameters and establishing practical guidelines for in situ cryo-electron tomography (cryo-ET). Critically, the authors present a systematic framework for optimizing data collection to achieve the highest attainable resolution.

      Using Dictyostelium cells as a model system, the authors generate multiple datasets at a constant total dose while varying the tilt increment. They demonstrate that tilt-series acquired with finer increments (1-3 degrees) yield superior alignment accuracy and improved template-matching performance, resulting in higher-quality reconstructions than those collected with coarser increments (5 degrees or above). Furthermore, the authors show that for subtomogram averaging, a 3-degree tilt increment outperforms all other conditions tested, particularly after per-particle refinement as implemented in M.

      Comments on revised version.

      The authors have addressed all my concerns, and I have no further issues.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work addresses a question of practical importance that had never been systematically analysed in the cryo-ET field: when collecting tilt-series data, what is the optimal angular step size between successive tilt images? Due to the upper limit in electron exposure (100 - 150 e<sup>-</sup>/Å<sup>2</sup>), this question is important, since finer angular sampling improves attainable reconstruction resolution (Crowther criterion) but reduces the signal-to-noise ratio of each individual image, potentially compromising both image quality and the ability to computationally align successive frames. To address this, the authors designed a thorough benchmarking study comparing five tilt increments (1°, 2°, 3°, 5°, and 10°) while keeping the total dose and tilt range constant. They evaluated the consequences at every stage of the cryo-ET workflow - from raw image quality and tilt-series alignment, through template matching for ribosome detection, to high-resolution subtomogram averaging - with the goal of providing the community with an evidence-based recommendation for data acquisition.

      The manuscript is well written, and the experimental design is carefully thought out. The work provides valuable practical insights into cryo-ET data acquisition by demonstrating that balancing two competing demands - sufficient dose per individual tilt image and fine angular sampling - is essential to achieve high-quality tomographic reconstructions. The identification of a practical optimum at 3° tilt increment is the key contribution of the work. It will be interesting to see in the future whether this optimum shifts for smaller molecular targets, and how emerging tilt interpolation strategies such as cryoTIGER may interact with the choice of experimental angular increment.

      The conclusions of this paper are mostly well supported by data, but some aspects of data analysis need to be clarified and/or extended, including:

      (1) Line 109: The authors state that the tilt range was kept at ± 60° relative to the lamella plane. Assuming a typical lamella pre-tilt of ~10°, the absolute stage tilt would approach its mechanical limit. Two clarifications would be appreciated: (a) What was the average pre-tilt across all lamellae? (b) How many dark tilt images, if any, were excluded during tomogram reconstruction?

      We thank the reviewer for asking for further clarification. For all our datasets, the pre-tilt of the stage was +8° with the lamella untilted under the e-beam, resulting in a tilt range of -52° to + 68°, thereby not reaching the mechanical limit, which is 70° for our microscope stage.

      Regarding “dark tilt images”, for most datasets, we did not need to remove many tilt images. However, we now noticed notably more absence of images from higher tilt values for the 1° dataset (see SFig 1). When analysing further, we noticed that for this dataset, we did not actively remove many images prior to tomogram reconstruction, but rather that they were not acquired in the first place by SerialEM. During acquisition, SerialEM performs various safeguarding checks that can abort the acquisition of a tilt series (or of a single branch). As this seems predominantly a problem for the 1-degree tilt-increment dataset, we have decided to add this to the manuscript as follows, including the figure as new SFig 1.

      In the main text:

      “For most datasets, image acquisition was largely complete, with the exception of the 1-degree dataset, which showed a markedly higher proportion of missing images at high tilt angles (SFig. 1). Closer inspection revealed that many of these images were not acquired, as SerialEM applies built-in safeguards (e.g. autofocus inconsistency or insufficient image counts) that can abort a tilt-series branch before completion.”

      (2) Line 148: "When analysing tomographic volumes, we found that tomograms from data with a smaller increment displayed higher SNR values (see Fig. 2B)." It would be helpful to specify which comparisons are statistically meaningful (e.g. Mann-Whitney U test?). While the difference between 1° and 2° appears pronounced, the differences between 2°, 3°, and 5° seem minimal. From my point of view, reporting the mean SNR values +/- standard deviations for each condition would already indicate some significance. Furthermore, since SNR is expected to depend on lamella thickness, it should be clarified whether the average lamella thickness is comparable across the five datasets.

      We have now calculated the mean and standard deviation of the tomogram SNR, as follows:

      Author response table 1.

      Furthermore, we performed a statistical significance test. Kruskal-Wallis test confirmed significant differences in SNR across tilt increments (H=270.97, p<0.001). Pairwise Mann-Whitney U tests with Bonferroni correction revealed significant differences between all pairs except 2° and 3° (p=0.093), suggesting these two conditions indeed yield comparable SNR.

      Lastly, we have now added data regarding local lamella thickness for all tilt-series used in the study, as displayed in SFig. 4.

      We incorporated this in the manuscript as follows.

      In the main text:

      “When analysing tomographic volumes, we found that tomograms from data with a smaller increment displayed higher SNR values (see Fig. 2B and Supplementary Note), whilst showing a similar lamella thickness distribution (see SFig. 4A).”

      and:

      “Firstly, we selected ca. 20 tomograms per condition, based on tomogram content and local lamella thickness [31] (for more details, see Methods and SFig. 4B).”

      As a supplementary note:

      “As the tomogram SNR distribution of particularly the 2° and 3° dataset showed similar SNR distributions, we performed formal significance testing for the data in this panel (see Fig. 2B). Kruskal-Wallis test confirmed significant differences across conditions (H=270.97, p<0.001); pairwise Mann-Whitney U tests with Bonferroni correction revealed all pairs were significantly different except 2° vs. 3° (p=0.093), indicating comparable SNR for these two tilt increments.”

      And, adding the test in the Methods:

      “To quantify differences in signal-to-noise ratio (SNR) across tilt increment conditions, a non-parametric Kruskal-Wallis test was performed as an omnibus test of the null hypothesis that all groups are drawn from the same distribution. Because SNR distributions were not assumed to be normal, and sample sizes differed across conditions, non-parametric tests were used throughout. Following the omnibus test, all 10 pairwise comparisons between conditions were assessed using two-sided Mann-Whitney U tests. To control for multiple comparisons, raw p-values were adjusted using the Bonferroni correction (multiplied by the number of comparisons, n=10, capped at 1.0). Statistical significance was defined as a Bonferroni-corrected p-value below 0.05. All analyses were performed in Python using the scipy.stats module.”

      (3) Line 167: "Indeed, the variation in maximum resolution correlates with lamella thickness across all datasets (see Fig. 2F)." The reported R<sup>2</sup> values of 0.30 (1°), 0.38 (2°), 0.66 (3°), 0.61 (5°), and 0.60 (10°) reveal a notably weak linear relationship for the finer tilt increments. It is also difficult to assess whether the lamella thickness distributions are comparable across conditions from the current figures - visually, the 1° dataset appears to be based on thinner lamellae, while the 10° dataset appears to include thicker samples. A histogram of lamella thickness distributions for each condition, provided as supplementary material, would greatly aid interpretation. Given this thickness dependency, reporting mean +/- standard deviation of lamella thickness per condition is highly appreciated.

      We have added the full lamella thickness distribution per dataset now in SFig. 4.

      The apparent weaker relationship between resolution fit and local lamella thickness for the 1 dataset seems to be largely apparent to the few very thin data points in this data (for more clarity, see the same data plotted separately in Author response image 1). We speculate that this is due to even less signal in these very thin and very low-dose images.

      Author response image 1.

      (4) Figure 4: It should be specified which tomogram subsets were used for the Rosenthal-Henderson analysis, whether lamella thickness was taken into account in the subset selection, and whether ribosomes too close to the lamella edges were excluded. Finally, linear fits should be displayed across the full x-axis range for all tilt increments to facilitate direct visual comparison.

      We have described the process of tomogram subset selection in detail in the Methods section Template matching and 3D classification. To further add clarity, we have incorporated the local lamella distribution plots for the full data, as well as specifically for the tomograms subjected to TM and STA in SFig. 4.

      Regarding the linear fits, we respectfully disagree with this suggestion. Displaying the linear fits only over the range used for their calculation avoids implying that the linear relationship extends beyond the measured data, and in our view produces a clearer figure.

      (5) General: Were ribosomes located at the lamella edges excluded from the analysis? As demonstrated in the authors' own prior work (Tuijtel et al., Science Advances, 2024), Ga-FIB milling induces structural damage at the lamella surfaces. To exclude the influence on the STA results, particles near the lamella edges should be removed prior to analysis, and the criteria for this exclusion should be stated explicitly.

      We have not excluded any ribosomes from close to the surface. As we still treated all data the same for each condition, we anticipate that the results of the comparison reported here still hold true.

      The aim of the authors was to provide the cryo-ET community with an evidence-based recommendation for the choice of tilt increment, and they largely succeeded in this goal. The identification of 3° as a practical optimum - balancing sufficient dose per tilt image for effective per-particle refinement with fine enough angular sampling for accurate tilt-series alignment - is well supported by the data and consistent across the multiple quality metrics employed. The conclusion that coarser increments (5° and 10°) compromise tomogram quality, template matching accuracy, and STA resolution is robust and clearly demonstrated. However, the conclusion rests entirely on a single biological system using ribosomes as the sole molecular target, which are exceptionally favourable due to their abundance, size, and electron contrast. Whether the identified optimum holds for smaller, lower-abundance, or lower-contrast targets remains an open question.

      In future, it would be particularly interesting to test whether emerging tilt interpolation strategies, such as cryoTIGER, which is particularly intriguing, can effectively compensate for coarser experimental angular sampling in post-processing. Here, the optimal experimental increment may shift, and the interaction between these two approaches represents a promising direction for future work. More broadly, as cryo-ET datasets grow larger and public repositories expand, the practical tradeoffs between acquisition time, data storage, and structural quality identified here will become increasingly relevant to the field.

      We agree with the reviewer and thank them for this positive assessment. An interesting note to the use of cryoTIGER in particular is that it uses already aligned tilt-series as an input, and it therefore is unlikely to overcome severe alignment issues associated with large tilt-increments.

      Reviewer #2 (Public review):

      The determination of macromolecular structures directly within their native cellular environment is becoming increasingly routine, making standardized data collection strategies essential. In this manuscript, Tuijtel et al. provide a timely and valuable contribution by benchmarking key acquisition parameters and establishing practical guidelines for in situ cryo-electron tomography (cryo-ET). Critically, the authors present a systematic framework for optimizing data collection to achieve the highest attainable resolution.

      Using Dictyostelium cells as a model system, the authors generate multiple datasets at a constant total dose while varying the tilt increment. They demonstrate that tilt-series acquired with finer increments (1-3 degrees) yield superior alignment accuracy and improved template-matching performance, resulting in higher-quality reconstructions than those collected with coarser increments (5 degrees or above). Furthermore, the authors show that for subtomogram averaging, a 3-degree tilt increment outperforms all other conditions tested, particularly after per-particle refinement as implemented in M.

      Overall, the manuscript is clearly written, and the conclusions are well supported by the data presented. I have no major concerns. There are some minor points that the authors should address, including:

      (1) The phrase "electron optical density distribution" (line 31, Introduction) should be revised to "electrostatic potential" or "Coulomb potential distribution," which more accurately reflects what is measured in cryo-EM/ET.

      We thank the reviewer for this correction and have adjusted it in the text:

      “It captures the 3-dimensional (3D) electrostatic potential of the specimen under scrutiny and enables the structural analysis of macromolecular complexes within their native context.”

      (2) The authors state that the maximum tolerable electron dose is approximately 100-150 e<sup>-</sup>/Å<sup>2</sup> (line 34, Introduction). This is an oversimplification, as bacterial specimens, for example, have been shown to tolerate doses of 200 e<sup>-</sup>/Å<sup>2</sup> or higher (see Breigel et al., PNAS, 2009; https://www.pnas.org/doi/10.1073/pnas.0905181106#T1). The statement should be revised to reflect this variability.

      We adjusted this statement to now read:

      “One of these is the maximum electron dose that can be applied to biological specimens before irreversible damage occurs, which is about 100-150 e-/Å2 for most eukaryotic cells.”

      (3) Lines 56-57: The authors do not cite their own prior work benchmarking tilt-series acquisition strategies on in vitro samples. This earlier study provides important context and should be referenced and briefly discussed.

      We assume the reviewer meant this study: Turonova et al., Nat. Comms. (2020). We have now added and discussed this reference as follows:

      “Since accumulated radiation dose progressively degrades high-resolution information, this motivated the development of the dose-symmetric tilt scheme, which prioritizes acquisition of low-tilt images early to better preserve high-resolution information [11, 16].”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 159: "Surprisingly though, the resolution to which the CTF was fitted was similar for all conditions, despite an 8-fold increase in dose (see Fig. 2B, E)." The reference to Figure 2B at this point is unclear.

      We thank the reviewer for pointing this out, we have removed the reference to panel B.

      (2) Line 162: "As the data shown in Fig. 2D-F pertains to images of the untilted specimen, ..." For clarity, this should also be stated explicitly in the figure caption.

      We have added this to the figure legend; “both estimated with Gctf on projection images of the sample at effective zero-tilt position.”

      (3) Figure 3: These are compelling results, and the 3D classification outcomes provide an excellent visual representation of the quantitative data shown in Figure 3. Including (some or all) of the initial five classes in the figure would further strengthen this already convincing presentation. Additionally, applying a uniform extraction threshold (e.g., z-score of 3.75 or 5) across all tilt increments would facilitate a more direct comparison. But this is really a minor remark, the authors and the editors may judge if the current presentation is already sufficient.

      We have adjusted Figure 3 according to the reviewer’s recommendation:

      We referenced this in the main text as:

      “To ensure similar data processing strategies for all conditions, extraction thresholds were also lowered for the 1-, 2- and 3-degree conditions, and 3D classification was performed to filter out the junk particles (see Fig. 3 A and SFig. 8). “

      In order to directly compare the particle extraction, we have already carried out such a uniform extraction threshold, with a z-score threshold of 5 (apart for the 10-degree data, where this was not possible). This extraction was then used for the 3D classification that led to the TM analysis and further STA investigations.

      Reviewer #2 (Recommendations for the authors):

      (1) Supplementary Figure 1: To improve accessibility for a broader readership, the authors should annotate or highlight the key organelles and protein complexes visible in the tomographic slices.

      We thank the reviewer for this suggestion, but adding arrowheads made this figure too crowded in our opinion. Furthermore, for most of the panels, the mentioned features of interest is centred in the image panel, which should make identification straightforward.

      (2) 'In situ' should be in italics throughout the text.

      We have changed this.

    1. "Graph Representation Of property ValuEs".

      Interesting I seem to be approaching it from the other end of the telescope

      I go with the alternative, the greatest gift to mankind offerred by Javascript: class free objects

      eschew completeness, care about coherence and vistas

    1. eLife Assessment

      This is a valuable study on the metabolic adaptations upon succinate dehydrogenase loss in cancer cells. If confirmed, this study will offer some therapeutic vulnerabilities in treating SDH-deficient cancer. However, the evidence supporting the authors' claim is incomplete and would benefit from additional experimental evidence. This study will be of broad interest for cancer biologists focusing on metabolism.

    2. Reviewer #1 (Public review):

      Different studies have proposed distinct mechanisms by which succinate dehydrogenase (SDH)-deficient cells escape aspartate limitation, highlighting metabolic heterogeneity across experimental systems. In this study, the authors address these previously conflicting observations by longitudinally tracking the adaptation of multiple SDHB-knockout clones derived from the same parental cell line.

      The authors identify two distinct adaptive mechanisms: complex I suppression with predominantly GOT1-dependent aspartate synthesis, and preservation of complex I activity with increased PC-GOT2-dependent aspartate synthesis. They further define shared and unique dependencies associated with these adaptive states, providing a rationale for potential therapeutic targeting strategies.

      Overall, this is a strong study in cancer metabolism, integrating complementary longitudinal and mechanistic approaches, including long-term adaptation, isotope tracing, genetic perturbation, metabolomics, and functional cell growth assays. Although the study provides substantial mechanistic insight, several limitations remain.

      (1) MPC is proposed as a shared dependency of both adaptive states. Testing whether MPC inhibition suppresses SDH-deficient tumor growth in vivo would substantially strengthen the therapeutic relevance.

      (2) The distinction between complex I-intact and complex I-suppressed states is based mainly on the expression of two complex I subunits and the oxygen consumption. More direct assays of complex I activity or assembly are needed. Early-passage SDHB-knockout cells should also be included as controls in the OCR experiments.

      (3) The two adaptive states appear to rely differentially on glucose- versus glutamine-derived aspartate synthesis. Testing the sensitivity of EP and LP clones to glucose or glutamine deprivation would further support this metabolic distinction.

      (4) Since SDH is described as a tumor suppressor, the authors should clarify why SDHB loss initially inhibits hPheo1 cell proliferation.

      (5) The study focuses on SDHB loss, and it remains unclear whether similar adaptive mechanisms arise following loss of other SDH subunits, including SDHA, SDHC, or SDHD, across different biological contexts. This limitation should be discussed explicitly.

    3. Reviewer #2 (Public review):

      In the manuscript entitled "Adaptive plasticity of aspartate metabolism in succinate dehydrogenase-deficient cancer cells," Sokolov et al. delineate the metabolic adaptations that succinate dehydrogenase (SDH)-deficient cancer cells undergo over time to overcome the initial aspartate limitation. To do so, the authors generated five clonal osteosarcoma SDH subunit B (SDHB) knockout cell lines using the CRISPR/Cas9 system and compared the proliferation rates of early- and late-passage cells, revealing that the latter rewired central carbon metabolism to increase aspartate levels and therefore replicate faster than their early-passage counterparts. Using a series of pharmacological and/or genetic interventions, the authors show that this rewiring can occur via two different routes: either through reduced Complex I (CI) activity, whereby glutamine is channelled towards aspartate synthesis via reductive carboxylation, or through a metabolic rewiring in which aspartate is produced from glucose via the PC-GOT2 pathway while CI activity is preserved. The CI-suppression-independent route depends on PC expression, as evidenced by an analysis of DepMap cell-line data, in which higher PC expression is associated with decreased SDH dependency. Moreover, they find that other consequences of aspartate deprivation observed in SDH-deficient cells, including impaired pyrimidine synthesis, replication stress, and DNA damage, are ameliorated in late-passage cells.

      Overall, this study is interesting because it disentangles the different metabolic rewiring routes that SDH-deficient cells can undergo to reverse aspartate limitation and sheds light on previously reported, seemingly contradictory results in the field. However, the study's major premise requires further validation, and important controls are missing, diminishing the overall strength of the conclusions.

      Major points:

      (1) The main conclusion that two separate routes allow SDH-deficient cells to overcome aspartate limitation, defined by their CI-activity status, is not convincingly proven. Indeed, to show this dichotomous behaviour, the authors performed Western blots for two CI subunits and determined the basal oxygen consumption rate. However, these assays are insufficient to demonstrate that LP clones 2 and 3 maintain functional CI, in contrast to LP clone 1. Moreover, it is not ruled out that these clones show dysfunction in ETC complexes other than CI. To assess these points, the activities of all individual ETC complexes should be carefully measured, for instance, by Seahorse assay after permeabilization. Furthermore, given the complex nature of CI, a reduction in two subunits does not necessarily reflect a reduction in its assembly. Therefore, CI assembly should be assessed directly by BN-PAGE analysis of isolated mitochondria.

      (2) It is difficult to reconcile why the authors used an NDUFA8 KO in clone 2 EP to mimic the physiological long-term CI-suppression-dependent adaptation. Indeed, this approach seems to represent an extreme scenario of Complex I loss that may induce non-physiological adaptations that override the effects of SDH KO. To assess the distinct metabolic fluxes between the two proposed routes, it would be advisable to use a more physiological model and instead compare the tracing data from LP clone 2 with those from LP clone 1, which exhibits a "natural" CI-suppressed state. Does clone 1 LP show similar metabolic changes to A8KO, including increased reductive carboxylation?

      (3) It is unclear whether the loss of Complex I at late passage is an intrinsic progression of osteosarcoma cells rather than a feature specific to SDH-deficient cells. A proper comparison between SDHB-deficient cells and WT cells, both at early and late passage, should be carried out. This is essential to fully understand the adaptive trajectories of SDH-deficient cells. This comparison is essential to identify the baseline metabolic hardware of the osteosarcoma cells. Indeed, the authors state that "While wild-type 143B cells synthesize most aspartate from glutamine via oxidative TCA cycling and GOT2 activity, ..." (Page 7, third paragraph), but these data are not included in the manuscript and would represent an important control for assessing the observed metabolic changes in comparison with the wild-type context.

      (4) The data showing that the PC-GOT2 pathway is mainly driven by enhanced PC activity are not fully convincing, as PC activity seems to be equally important for maintaining aspartate levels in the NDUFA8 KO compared with clone 2 LP. Moreover, the extracted expression data from DepMap suggest that increased PC expression might not be transcriptionally regulated, as only a slight association between PC mRNA levels and SDH dependency was observed. Are PC mRNA levels increased in clone 2 LP? If not, PC might be regulated post-transcriptionally. To test this, the nascent translation of PC could be assessed.

    4. Author response:

      We sincerely thank the reviewers for their time, helpful critiques and overall positive evaluation of our work. We plan to investigate the points raised by the reviewers and look forward to submitting a revised manuscript that addresses concerns about therapeutic relevance, functional distinctions between the CI-intact and CI-suppressed adapted states, and PC regulation in this system.

    1. eLife Assessment

      This valuable study documents, for the first time, the degree of preference for contralateral versus ipsilateral forelimb movements in the motor cortex of mice, revealing subtly distinct profiles across cortical areas, with the orofacial region showing comparatively little limb selectivity relative to the forelimb motor areas. The experiments and analyses are, in general, solid and well-designed, and the data support the paper's conclusions.

    2. Reviewer #1 (Public review):

      Summary:

      This study describes motor cortical activity patterns during food handling in mice, investigating whether the hand/s used is reflected in distinct neural activity. The experiments focus on forelimb M1 and M2 (fM1, fM2) and an oral-manual region LOM. The main findings are that fM1 and fM2 have largely similar relationships with forelimb control, and LOM neurons are more broadly tuned. These conclusions are reached using a variety of analyses spanning straightforward firing rate analyses, selectivity metrics, PCA, and GLM decoding methods to assess tuning generalizability. The study's significance is strengthened by including analyses of bimanual control, and in this sphere, there are descriptive data and analyses that aficionados of cortical control of dexterous behaviors will find useful. The use of unimanual control is useful as a point of comparison here, but less novel overall. There are a number of places where the descriptions of what is being analyzed, what is being concluded, and data reporting should be strengthened and clarified. Additionally, the study could be greatly improved by consolidating figures and the analyses shown, since many are redundant. Many of the analyses need clearer reporting of means and effect sizes in the text, rather than just statistical outcomes. Overall, at this juncture, the study presents analyses of a unique dataset that may seed future investigations of mechanisms of bimanual coordination.

      Strengths:

      There are relatively few studies that compare neural activity across bimanual and unimanual control. This study uses a naturalistic food handling task to explore neural relationships to forelimb kinematics under these conditions. The uniqueness of the task and analysis target is a strength of the study.

      The authors remain fairly conservative and make few strong claims in the study, which may be warranted given the diversity of tuning profiles they observed.

      Weaknesses:

      There are a number of statistical tests that were accompanied by too little information to critically evaluate. Means and effect sizes needed to be better reported; some details of analyses were difficult to parse, making the strength of the conclusions difficult to evaluate.

    3. Reviewer #2 (Public review):

      Summary:

      Barrett et al. examine how neural activity in the mouse motor cortex varies when a movement is performed with the ipsilateral or contralateral forelimb. First, they train animals to grasp and manipulate a pellet of food with either the left forepaw, the right forepaw, or both. Next, they measure activity in the primary and secondary forelimb motor areas (fl-M1 and fl-M2) and in the classical tongue-jaw area (tj-M1 / LOM). While responses in the forelimb areas are diverse, with some neurons preferring ipsilateral or bilateral movements, a plurality of cells prefer the contralateral limb. In LOM, by contrast, little limb selectivity is observed. At the neural population level, structure is preserved across conditions in LOM, but not in the forelimb areas. Finally, paw position can be decoded from activity in all three areas, and the LOM decoder generalized across limbs.

      Strengths:

      While previous studies in macaques have compared motor cortical activity during movement (and perturbation) of the contralateral and ipsilateral arms, no analogous work has been undertaken in rodents. This paper closes this knowledge gap by showing, for the first time, moderate-to-strong lateralization in the forelimb motor cortical areas of mice transporting grasped food pellets to the mouth, and a relative absence of lateralization in the classical tongue-jaw area. On the whole, I think this is a solid paper that reports novel observations of interest to the motor systems community.

      Weaknesses:

      The central question posed is whether cortical activity depends on the effector(s) used (ipsi forelimb, contra forelimb, or both). The corresponding hypotheses (Figure 1) are somewhat coarse-grained and are not mutually exclusive. One might expect to see condition-independent, limb-selective, and uni-/bimanual-selective signals in motor cortex (though their magnitudes could differ substantially), and to find these signals intermingled at the level of single neurons. The authors may wish to consider setting up a more focused question. For example, can bimanual responses be explained as a sum of the unimanual responses from the left and right limbs?

      In the area usually identified as tongue-jaw motor cortex (here referred to as LOM), unit and population activity look quite similar for ipsilateral, contralateral, and bilateral forelimb reaches. The most parsimonious explanation is that the activity is related mostly to mouth and tongue movements, rather than limb movements. Systematic mapping studies with microstimulation in the rat (Neafsy et al., Brain Res. Rev. 1986) and optogenetic stimulation in the mouse (Mayrhofer et al., Neuron 2019) tend to support the idea that tjM1/LOM is specialized for control of the tongue and mouth. Thus, I'm not entirely convinced that it "encodes ingestion-related forelimb parameters necessary for oromanual coordination." The authors could say more about this issue: what specific limb-related parameters do they think are encoded, why would these parameters be effector-independent, what evidence for this encoding is presented here, and how can limb- and mouth-related components be distinguished? The problem could potentially be addressed experimentally, as well, by delivering food pellets directly to the mouth while preventing manipulation with the paws, but this experiment isn't strictly necessary.

      Because the corticospinal tract is strongly lateralized, cortical activity presumably has a smaller effect on ipsilateral than contralateral motor output. Somatosensory feedback should also be relatively lateralized for the forelimb areas. The authors could say a bit more about this issue and how it relates to their data and conclusions in the Discussion.

      An important limitation of the behavioral task is that it involves only a single stereotyped movement for each limb, instead of multiple directions, speeds, or loads. This issue and its consequences for the analyses (especially those in Figures 7-10) and conclusions could be discussed.

    4. Reviewer #3 (Public review):

      Summary:

      Barrett et al. compare the responses of different parts of the mouse primary and secondary motor cortex in the context of a task where the animals manipulate and eat food using either or both hands. They find that roughly half the activity is conserved when reaching with one hand vs. the other hand, or with both. Similarity of activity was somewhat higher in the "lateral oral and manual" (LOM) part of the motor cortex, consistent with notions of a more generalized oromanual function there.

      Strengths:

      This work aims at addressing two worthwhile questions in a mouse model of motor control: (1) what specializations do we have for controlling feeding movements, and (2) how are the arms and hands coordinated with one another? The authors develop a simple but innovative apparatus to block either hand during food handling, track the behavior at high temporal fidelity, and record a sizable neural dataset. The analyses come from numerous angles to take good advantage of the data, and succeed in showing multiple lines of evidence for greater invariance in LOM than in the forelimb parts of M1 and M2.

      Weaknesses:

      There are several limitations of the current study. Most importantly, the behavior presents an inherent challenge: there is only one type of movement for each of the three conditions (contra hand, ipsi hand, and bimanual). This is entirely reasonable from the perspective that this is the ethological behavior when feeding, but it limits what analyses are possible. In particular, it precludes disentangling the neural relationship with many correlated aspects of behavior, and limits identifying population-level features of the neural activity meaningfully. This means that there are a number of alternative possible sources of the neuron-level area differences found here, and the population-level features may not be reliable. Second, the behavior tracking was used at a relatively coarse level, and thus the relationships to various behavioral variables were left less distinguishable than they might have been. Finally, there may be an issue with the coordinates of what is being called forelimb M1 here, which may include some hindlimb M1.

    5. Author response:

      Reviewer #1 (Public review):

      Summary:

      This study describes motor cortical activity patterns during food handling in mice, investigating whether the hand/s used is reflected in distinct neural activity. The experiments focus on forelimb M1 and M2 (fM1, fM2) and an oral-manual region LOM. The main findings are that fM1 and fM2 have largely similar relationships with forelimb control, and LOM neurons are more broadly tuned. These conclusions are reached using a variety of analyses spanning straightforward firing rate analyses, selectivity metrics, PCA, and GLM decoding methods to assess tuning generalizability. The study's significance is strengthened by including analyses of bimanual control, and in this sphere, there are descriptive data and analyses that aficionados of cortical control of dexterous behaviors will find useful. The use of unimanual control is useful as a point of comparison here, but less novel overall. There are a number of places where the descriptions of what is being analyzed, what is being concluded, and data reporting should be strengthened and clarified. Additionally, the study could be greatly improved by consolidating figures and the analyses shown, since many are redundant. Many of the analyses need clearer reporting of means and effect sizes in the text, rather than just statistical outcomes. Overall, at this juncture, the study presents analyses of a unique dataset that may seed future investigations of mechanisms of bimanual coordination.

      Strengths:

      There are relatively few studies that compare neural activity across bimanual and unimanual control. This study uses a naturalistic food handling task to explore neural relationships to forelimb kinematics under these conditions. The uniqueness of the task and analysis target is a strength of the study.

      The authors remain fairly conservative and make few strong claims in the study, which may be warranted given the diversity of tuning profiles they observed.

      Weaknesses:

      There are a number of statistical tests that were accompanied by too little information to critically evaluate. Means and effect sizes needed to be better reported; some details of analyses were difficult to parse, making the strength of the conclusions difficult to evaluate.

      We will improve these aspects of the presentation and reporting of statistical analyses in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Barrett et al. examine how neural activity in the mouse motor cortex varies when a movement is performed with the ipsilateral or contralateral forelimb. First, they train animals to grasp and manipulate a pellet of food with either the left forepaw, the right forepaw, or both. Next, they measure activity in the primary and secondary forelimb motor areas (fl-M1 and fl-M2) and in the classical tongue-jaw area (tj-M1 / LOM). While responses in the forelimb areas are diverse, with some neurons preferring ipsilateral or bilateral movements, a plurality of cells prefer the contralateral limb. In LOM, by contrast, little limb selectivity is observed. At the neural population level, structure is preserved across conditions in LOM, but not in the forelimb areas. Finally, paw position can be decoded from activity in all three areas, and the LOM decoder generalized across limbs.

      Strengths:

      While previous studies in macaques have compared motor cortical activity during movement (and perturbation) of the contralateral and ipsilateral arms, no analogous work has been undertaken in rodents. This paper closes this knowledge gap by showing, for the first time, moderate-to-strong lateralization in the forelimb motor cortical areas of mice transporting grasped food pellets to the mouth, and a relative absence of lateralization in the classical tongue-jaw area. On the whole, I think this is a solid paper that reports novel observations of interest to the motor systems community.

      Weaknesses:

      The central question posed is whether cortical activity depends on the effector(s) used (ipsi forelimb, contra forelimb, or both). The corresponding hypotheses (Figure 1) are somewhat coarse-grained and are not mutually exclusive. One might expect to see condition-independent, limb-selective, and uni-/bimanual-selective signals in motor cortex (though their magnitudes could differ substantially), and to find these signals intermingled at the level of single neurons. The authors may wish to consider setting up a more focused question. For example, can bimanual responses be explained as a sum of the unimanual responses from the left and right limbs?

      In the revised manuscript, we will clarify that the possibilities illustrated in figure 1 are not intended as mutually exclusive. We indeed find all of these signals intermingled. In terms of a more focused question amenable to hypothesis testing, this can be expressed as: for each dimension (laterality vs manuality) are the activity patterns in each area closer to those predicted by invariance or dependence, as compared to the other areas? The various statistical analyses in the paper all essentially boil down to testing this question. Broadly, the answer is yes: we see activity closer to the invariant prediction LOM, and activity closer to the dependent prediction in fl-M1 and fl-M2. Testing whether bimanual activity can be explained as a simple linear sum of the left and right unimanual might provide additional insight into this question, and this analysis will be presented in the revised manuscript.

      In the area usually identified as tongue-jaw motor cortex (here referred to as LOM), unit and population activity look quite similar for ipsilateral, contralateral, and bilateral forelimb reaches. The most parsimonious explanation is that the activity is related mostly to mouth and tongue movements, rather than limb movements. Systematic mapping studies with microstimulation in the rat (Neafsy et al., Brain Res. Rev. 1986) and optogenetic stimulation in the mouse (Mayrhofer et al., Neuron 2019) tend to support the idea that tjM1/LOM is specialized for control of the tongue and mouth. Thus, I'm not entirely convinced that it "encodes ingestion-related forelimb parameters necessary for oromanual coordination." The authors could say more about this issue: what specific limb-related parameters do they think are encoded, why would these parameters be effector-independent, what evidence for this encoding is presented here, and how can limb- and mouth-related components be distinguished? The problem could potentially be addressed experimentally, as well, by delivering food pellets directly to the mouth while preventing manipulation with the paws, but this experiment isn't strictly necessary.

      The issue of orofacial movement confounds is an important one that we made a point of addressing in the discussion. There are three main points that we believe cast doubt on this as the most likely explanation for the effector-invariant representation in LOM.

      First, while we do not disagree that LOM has an important role in tongue and jaw control, there is plenty of evidence from mapping and behavioural studies (which we cite in the introduction) that it also plays a role in forelimb motor control as well.

      Second, while we cannot observe all orofacial movements, we have previously shown that the jaw is less active when the hands and LOM are most active (Barrett et al., 2024). Conversely, LOM firing is much lower during chewing, when the tongue and jaw are very active.

      Finally, the correlation between LOM firing and forelimb movements is not merely a coarse-grained one on the timescale of active manipulation vs passive holding phases. LOM firing closely tracks the position of the forelimb(s) on fast timescales and with near-zero lag, giving better decoding than from fl-M1 or fl-M2, as we have shown here and previously (Barrett et al., 2022). If we assume that LOM only encodes orofacial movements, then this result implies that orofacial movements correlate with forelimb position better than fl-M1 or fl-M2 firing correlates with forelimb position.

      The revised manuscript will include an expanded discussion to clarify these and related points.

      Because the corticospinal tract is strongly lateralized, cortical activity presumably has a smaller effect on ipsilateral than contralateral motor output. Somatosensory feedback should also be relatively lateralized for the forelimb areas. The authors could say a bit more about this issue and how it relates to their data and conclusions in the Discussion.

      We will discuss this in the revised manuscript.

      An important limitation of the behavioral task is that it involves only a single stereotyped movement for each limb, instead of multiple directions, speeds, or loads. This issue and its consequences for the analyses (especially those in Figures 7-10) and conclusions could be discussed.

      This limitation applies to the analyses relating to the transport-to-mouth movement (Figures 3-7). The population correlation structure and decoding analyses (Figures 8-10) consider activity throughout the full duration of food handling, which involves a much greater variety of movements (Barrett et al., 2020). Indeed, this was a major motivation for including these analyses. The revised manuscript will clarify this point.

      Reviewer #3 (Public review):

      Summary:

      Barrett et al. compare the responses of different parts of the mouse primary and secondary motor cortex in the context of a task where the animals manipulate and eat food using either or both hands. They find that roughly half the activity is conserved when reaching with one hand vs. the other hand, or with both. Similarity of activity was somewhat higher in the "lateral oral and manual" (LOM) part of the motor cortex, consistent with notions of a more generalized oromanual function there.

      Strengths:

      This work aims at addressing two worthwhile questions in a mouse model of motor control: (1) what specializations do we have for controlling feeding movements, and (2) how are the arms and hands coordinated with one another? The authors develop a simple but innovative apparatus to block either hand during food handling, track the behavior at high temporal fidelity, and record a sizable neural dataset. The analyses come from numerous angles to take good advantage of the data, and succeed in showing multiple lines of evidence for greater invariance in LOM than in the forelimb parts of M1 and M2.

      Weaknesses:

      There are several limitations of the current study. Most importantly, the behavior presents an inherent challenge: there is only one type of movement for each of the three conditions (contra hand, ipsi hand, and bimanual).

      See our response to reviewer #2 above regarding the variety of movements. We agree that this a limitation for the unit-level and PCA analyses, but one that is alleviated by the population correlation and decoding analyses, which relate to complex ongoing movements.

      This is entirely reasonable from the perspective that this is the ethological behavior when feeding, but it limits what analyses are possible. In particular, it precludes disentangling the neural relationship with many correlated aspects of behavior, and limits identifying population-level features of the neural activity meaningfully. This means that there are a number of alternative possible sources of the neuron-level area differences found here, and the population-level features may not be reliable.

      Important behavioural confounds include orofacial movements, non-specific movement initiation signals, and arousal. Orofacial movements we have discussed above in our response to Reviewer 2. Movement initiation signals would likely be transient and well-timed to movement onset, hence this may explain some of the effector-independent activity in fl-M1 and fl-M2 (consistent with e.g. (Kaufman et al., 2016)). However, we do not believe this to be the case in LOM as its activity is delayed and sustained relative to movement initiation. Regarding arousal, the mouse is actively engaged in consuming the food even when the hands are stationary, so there is no a priori reason to believe that arousal varies rapidly during the behaviour. Consistent with this, measurements of noradrenergic activity in the locus coeruleus during consumption suggest that arousal varies on slow timescales, on the order of seconds to tens of seconds (Sciolino et al., 2022). Such slow variation in arousal would not explain the rapid but condition-invariant changes in firing in any of the cortical areas studied here. The updated manuscript will include more detailed discussion of these points.

      Second, the behavior tracking was used at a relatively coarse level, and thus the relationships to various behavioral variables were left less distinguishable than they might have been.

      Behavior tracking was performed with kilohertz temporal resolution and submillimeter spatial resolution.

      Finally, there may be an issue with the coordinates of what is being called forelimb M1 here, which may include some hindlimb M1.

      Our recording coordinates are based on the territory of corticospinal neurons retrogradely labeled from C6 spinal cord, medial to any layer 4 labelling, as reported in our previous study (Yamawaki et al., 2021). Thus we are confident in calling this area forelimb M1.

      References:

      Barrett, J. M., Martin, M. E., Gao, M., Druzinsky, R. E., Miri, A., & Shepherd, G. M. G. (2024). Hand-jaw coordination as mice handle food is organized around intrinsic structure-function relationships. The Journal of Neuroscience, 44(42), e0856242024.

      Barrett, J. M., Martin, M. E., & Shepherd, G. M. G. (2022). Manipulation-specific cortical activity as mice handle food. Current Biology, 32(22), 4842-4853.e6.

      Barrett, J. M., Tapies, M. G. R., & Shepherd, G. M. G. (2020). Manual dexterity of mice during food-handling involves the thumb and a set of fast basic movements. PLOS ONE, 15(1), e0226774.

      Kaufman, M. T., Seely, J. S., Sussillo, D., Ryu, S. I., Shenoy, K. V., & Churchland, M. M. (2016). The Largest Response Component in the Motor Cortex Reflects Movement Timing but Not Movement Type. eNeuro, 3(4).

      Sciolino, N. R., Hsiang, M., Mazzone, C. M., Wilson, L. R., Plummer, N. W., Amin, J., Smith, K. G., McGee, C. A., Fry, S. A., Yang, C. X., Powell, J. M., Bruchas, M. R., Kravitz, A. V., Cushman, J. D., Krashes, M. J., Cui, G., & Jensen, P. (2022). Natural locus coeruleus dynamics during feeding. Science Advances, 8(33), eabn9134.

      Yamawaki, N., Raineri Tapies, M. G., Stults, A., Smith, G. A., & Shepherd, G. M. (2021). Circuit organization of the excitatory sensorimotor loop through hand/forelimb S1 and M1. eLife, 10, e66836.

    1. Message Body Template Examples

      We also support nested event fields, as shown in this example. { "summary": "Backup policy '{{data.config.name}}' was {{data.config.action}} by {{actor.identifier}}", "workload": "{{data.workload.type}}", "tenantId": "{{data.tenant.id}}", "policyId": "{{data.policy.id}}", "policyType": "{{data.policy.type}}", "status": "{{data.config.status}}", "schedule": "{{data.config.schedule}}", "retention": "{{data.config.retention}}", "occurredAt": "{{timestamp}}", "rawEvent": {{event}} }

      event will be sent like: { "summary": "Backup policy 'Daily Exchange Backup' was UPDATED by admin@contoso.com", "workload": "M365", "tenantId": "tenant_12345", "policyId": "policy_12345", "policyType": "Exchange Mailboxes", "status": "ENABLED", "schedule": "daily at 02:00 UTC", "retention": "90 days", "occurredAt": "2026-07-13T09:30:00Z", "rawEvent": { ...the full event JSON from above... } }

      I would add this kind of exsamples.

    2. You can nest group nodes to build complex filters, and you can negate any node so that it matches the events that do not meet the condition.

      nest group nodes are not supported

    1. eLife Assessment

      This manuscript employs cryo-EM, mutational analysis, and biochemical assays to explore the molecular basis by which glutamine promotes filamentation and regulates the activity of human glutamine synthetase (hGS) by stabilizing interactions between hGS decamers. The combined structural and biochemical data supporting the proposed mechanism are solid, although the evidence for a higher-order filamentous architecture and the precise assignment of the interface density would benefit from additional support. This work will be of particular interest and useful to groups interested in understanding the molecular basis of nutrient sensing, cellular metabolism, and structural regulation of enzyme activity.

    2. Reviewer #1 (Public review):

      Summary:

      The study is methodologically solid and introduces a compelling regulatory model. However, several mechanistic aspects and interpretations require clarification or additional experimental support to strengthen the conclusions.

      Strengths:

      (1) The manuscript presents a compelling structural and biochemical analysis of human glutamine synthetase, offering novel insights into product-induced filamentation.

      (2) The combination of cryo-EM, mutational analysis, and molecular dynamics provides a multifaceted view of filament assembly and enzyme regulation.

      (3) The contrast between human and E. coli GS filamentation mechanisms highlights a potentially unique mode of metabolic feedback in higher organisms.

      Comment on revised version.

      The authors have addressed all of my comments and concerns. The revisions have substantially improved the quality of the manuscript. I have no further questions or concerns.

    3. Reviewer #2 (Public review):

      Major concern 1: The manuscript does not clearly establish a bona fide GS filament state.

      The authors repeatedly refer to GS "filaments," but the data presented appear to support primarily a di-decameric assembly rather than a well-defined filamentous polymer.

      A two-decamer reconstruction can define a putative inter-decamer interface, but it cannot by itself demonstrate propagation of a repeating filament geometry. To establish a bona fide filament, the authors should provide evidence for a reproducible one-dimensional assembly, such as at least three consecutive repeating units or equivalent quantitative evidence that the same inter-decamer transform propagates along an assembly axis.

      In the current manuscript, many of the supporting 2D classifications appear to contain at most two adjacent GS decamers. This is particularly evident in the time-resolved cryo-EM datasets shown in Supplementary Figures 9-10, where I do not see convincing 2D classes corresponding to filaments. The same concern applies to other datasets, including Supplementary Figures 2, 6, and 11, where the apparent assemblies are primarily two-decamer particles.

      Moreover, many of the selected "filament" classes show only one well-resolved GS decamer, while the neighboring decamer density is blurred. This suggests substantial variability in the relative position and/or orientation of adjacent decamers. Such heterogeneity is difficult to reconcile with a stable repeating filament geometry.

      Therefore, the authors should explicitly define what they mean by "filament." If their evidence supports only a di-decameric or short oligomeric assembly, the terminology should be changed accordingly throughout the manuscript.

      Symmetry concern

      Given the low quality and heterogeneity of the 2D classifications for the putative "filament" classes, the use of D5 symmetry requires stronger justification. The current reconstructions primarily show the result after applying D5 symmetry to a two-decamer assembly. The authors should show reconstructions of the same particle sets processed under C1, C5, and D5 symmetry, and explain why D5 symmetry is justified.

      This is particularly important because the claimed interface density and ligand interpretation are sensitive to symmetry averaging. Without showing how the reconstruction behaves under less restrictive symmetry assumptions, it is difficult to determine whether the final D5 map reflects a true biological assembly or a symmetry-imposed interpretation.

      Filament abundance and physiological relevance

      Even under the authors' broad classification criteria, the filament-like population appears to be a minor species. In some datasets, especially Supplementary Figure 10, the apparent filament fraction is very low, approximately 2-10%. This raises a major concern: if GS filaments are rare even under high-concentration cryo-EM conditions, are they expected to form to a meaningful extent under physiological conditions?

      The authors propose a concentration-dependent assembly mechanism. If so, the relevance of GS filamentation in the lower-concentration cellular environment becomes even less clear. The authors should quantify filament abundance as a function of GS concentration and glutamine concentration, ideally under conditions closer to physiological ranges.

      K52/C53 interface mutations

      The authors use K52 and C53 as filament-interface residues, but the mechanistic contribution of these residues to filament assembly remains insufficiently explained. Why should K52A or C53A disrupt filament formation? Is the effect due to loss of a specific side-chain contact, altered local electrostatics, reduced crosslinker accessibility/reactivity, local structural destabilization, or nonspecific disruption of the interface?

      The manuscript states that the interface is "concentration dependent and driven primarily by electrostatic interactions," but the data presented before that statement do not clearly establish this. The authors should explicitly identify the interacting electrostatic partners and provide structural or biochemical evidence supporting this interpretation.

      Functional linkage between filamentation and kinetics is weak.

      The authors should establish the oligomeric state of GS under the actual assay conditions. In particular, what is the filament fraction during the Figure 2F / Supplementary Figure 15 kinetic assays? Is the change in KM ammonia quantitatively correlated with filament abundance?

      This is currently unclear. The direct comparison between decamer and 2-decamer fractions does not robustly show a functional difference, and the later glutamine-addition assays are interpreted as filament-mediated without directly demonstrating the filament fraction under the same assay conditions.

      In Figure 2F, WT, K52A, and C53A already show different ammonia-dependent kinetic parameters in the absence of added glutamine. K52A and C53A appear to have lower basal kcat/KM ammonia and higher KM ammonia than WT even without glutamine. The authors should explain why these interface mutants already alter basal ammonia kinetics. Without such an explanation, K52A and C53A cannot be treated as clean controls that selectively disrupt glutamine-stabilized filamentation.

      Major concern 2: The interface density is not convincingly assigned to glutamine.

      The second foundational issue is the assignment of the interface density to glutamine. At present, the evidence is not sufficient to support the conclusion that glutamine is the ligand at this interface.

      The local density at the interface appears weak and likely has lower local resolution than the reported global resolution. The current density could represent a low-occupancy or symmetry-averaged amino-acid-like density rather than a confidently assigned glutamine molecule.

      Ligand pose and hydrogen bonding.

      The proposed glutamine pose also requires more rigorous validation. The authors state that glutamine forms hydrogen bonds with interface residues, including K52, C53, and E55. These hydrogen bonds should be shown explicitly in a figure, with distances listed.

      The proposed interaction involving C53 appears unusual and should be justified chemically and geometrically.

      Glutamate has not been excluded.

      The largest problem is that the authors do not adequately consider glutamate as an alternative ligand. They compare the density with phosphate and ATP/ADP, but this is not sufficient. Glutamate is present at high concentration during turnover, and it is chemically and structurally very similar to glutamine. Given the limited local density and possible orientational averaging, distinguishing glutamine from glutamate from the current cryo-EM density alone is not justified.

      The authors should report or estimate the concentrations of glutamate and glutamine at the vitrification time point used for the high-resolution turnover-filament reconstruction. If glutamate is present at a much higher concentration than glutamine, the authors must explain why the interface density should be assigned to glutamine rather than glutamate.

      The authors should fit both glutamine and glutamate into the interface density using the same validation criteria and compare the results. Stronger support would come from direct structural experiments, such as cryo-EM structures of GS incubated separately with glutamate and glutamine under controlled conditions.

      Unless stronger evidence is provided, the claim that "glutamine binds to the filament interface" cannot be made.

      Specific comments

      Interface assembly statement:<br /> "These data suggest that the formation of the interface is concentration dependent and driven primarily by electrostatic interactions."

      What specific data support "concentration dependent" at this point in the manuscript? Which residues or chemical groups are proposed to form the electrostatic interactions? The authors should provide a more explicit explanation.

      Line 149-150:<br /> "In both scenarios, any signal is likely to be averaged out and experiments with symmetry expansion and focused classification did not yield any convincing density."

      Please show these analyses. Negative results are important here because they bear directly on the reliability of the interface interpretation.

      "Glutamine stabilizes larger GS filaments":<br /> What does "larger" mean? Longer filaments, more decamers per filament, or larger diameter? The authors should define this quantitatively, preferably by reporting filament-length distributions or the number of decamers per assembly.

      Filament classification:<br /> The criteria used to classify particles or 2D classes as "filament" are not sufficiently clear. The authors should provide the full 2D classification results for each time-resolved dataset, including selected and discarded classes, particle numbers, and objective selection criteria. Some selected and discarded classes appear visually similar, especially in Supplementary Figures 9-10.

      R298A decamer:<br /> The R298A mutant is presented as a turnover-decamer structure, not a filament structure. The authors should clarify whether R298A forms filament-like particles under comparable turnover conditions. If R298A does not form filaments, this should be reported and explained. If filament-like particles were present but excluded during processing, the authors should provide their abundance and justify why only the decameric form was analyzed. This point matters because R298A is used to connect E305-loop disorder with the proposed filament-associated mechanism, although R298A is a loop-stabilization mutant rather than a filament-interface mutant.

      Line 231-233:<br /> "a reaction time that should yield a high concentration of product due to the higher enzyme concentration than previous experiments"

      What is the estimated product concentration at vitrification? What concentration range qualifies as "high"? The authors should provide a quantitative estimate.

      Glutamine hydrogen bonds:<br /> The proposed hydrogen bonds linking glutamine to K52, C53, and E55 should be shown explicitly with atom identities and distances.

      Glutamate comparison:

      What is the glutamate concentration in the same sample? Given that glutamate is chemically similar to glutamine and likely present at high concentration, why is the interface density not glutamate? The authors should compare glutamine and glutamate fitting using the same validation criteria.

      Line 248-253:<br /> The speculation that apo filaments may arise from high GS concentration or residual glutamine should be moved to the Discussion. In the Results, this reads as an ad hoc explanation rather than a result directly supported by data.

      Actual assay-state oligomeric distribution:<br /> What is the filament fraction under the actual kinetic assay conditions? Is the KM ammonia change quantitatively correlated with filament abundance?

      Figure 2F:<br /> Why do WT, K52A, and C53A differ in basal ammonia-dependent activity even without added glutamine? The authors should explain whether these mutations alter intrinsic ammonia kinetics independent of filamentation.

      Supplementary Figure 15 / Figure 2F:<br /> Please clarify the relationship between Figure 2F and Supplementary Figure 15. The kinetic constants in Figure 2F appear to depend on global fitting of progress curves shown in Supplementary Figure 15. The authors should provide replicate-level raw progress curves, between-replicate variability, fitting residuals, and individual fitted parameters.

      Supplementary Figure 19:<br /> Supplementary Figure 19 should be presented consistently with Supplementary Figure 18, including the corresponding 2D classification results.

      In summary, although the revised manuscript improves the presentation of cryo-EM map processing, the two foundational claims remain unresolved. The current data establish, at most, a di-decameric or filament-like GS assembly, but not a rigorously defined filamentous polymer. In addition, the interface density is not convincingly assigned to glutamine, particularly because glutamate has not been excluded as the most relevant alternative ligand. Since the proposed negative-feedback mechanism depends directly on these two points, the current evidence does not support the strength of the title, abstract, or mechanistic conclusions.

    4. Reviewer #3 (Public review):

      In this manuscript, the authors propose a product-dependent negative-feedback mechanism of human glutamine synthetase, whereby the product glutamine facilitates filament formation, leading to reduced catalytic specificity for ammonia. Using time-resolved cryo-EM, the authors demonstrate filament formation under product-rich conditions. Multiple high-quality structures, including decameric and di-decameric assemblies, were resolved under different biochemical states and combined with MD simulations, revealing that the conformational space of the active site loop is critical for the GS catalysis. The study also includes extensive steady-state kinetic assays, supporting the view that glutamine regulates GS assembly and its catalytic activity. Overall, this is a detailed and comprehensive study. However, I would advise that a few points be addressed and clarified.

      Comments on revised version.

      The revision addresses several reviewer concerns: the authors add sharpened maps, ligand-density panels, symmetry expansion/focused classification, biochemical blank/substrate/TCEP controls, and E305-loop focused classification. The E305-loop part is stronger now: turnover decamer recovers partial E-flap density in few classes, while turnover filament does not.

      My only remaining comment is that - as also the authors agree on the need to integrate density with biochemical data and that local resolution/averaging complicates modeling - I would advise softening the claim regarding glutamine from "glutamine binds" to "density consistent with glutamine/product-associated density". In general, it would be best to avoid overstating atomic certainty at the filament interface and the safest framing is the observed interface density is compatible with glutamine but not independently conclusive.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study is methodologically solid and introduces a compelling regulatory model. However, several mechanistic aspects and interpretations require clarification or additional experimental support to strengthen the conclusions.

      Strengths:

      (1) The manuscript presents a compelling structural and biochemical analysis of human glutamine synthetase, offering novel insights into product-induced filamentation.

      (2) The combination of cryo-EM, mutational analysis, and molecular dynamics provides a multifaceted view of filament assembly and enzyme regulation.

      (3) The contrast between human and E. coli GS filamentation mechanisms highlights a potentially unique mode of metabolic feedback in higher organisms.

      Weaknesses:

      (1) The mechanism underlying spontaneous di-decamer formation in the absence of glutamine is insufficiently explored and lacks quantitative biophysical validation.

      (2) Claims of decamer-only behavior in mutants rely solely on negative-stain EM and are not supported by orthogonal solution-based methods.

      We thank the reviewer for the summary and noting of the strengths. We agree that the evolutionary divergence of metabolic feedback in GS homologs is a fruitful avenue for future studies. With regard to the weaknesses, the di-decamer in the absence of glutamine only forms under high (higher than physiological) concentrations of enzyme. Our primary evidence for the mutant behavior was the lack of crosslinking (Figure 1E), with supplementary support from the negative stain. In the revised version we will soften the language to say “reduced” rather than “did not support” filament formation.

      Reviewer #2 (Public review):

      The authors set out to resolve the high-resolution structure of a glutamine synthetase (GS) decamer using cryo-EM, investigate glutamine binding at the decamer interface, and validate structural observations through biochemical assays of ATP hydrolysis linked to enzyme activity. Their work sits at the intersection of structural and functional biology, aiming to bridge atomic-level details with biological mechanisms - a goal with clear relevance to researchers studying enzyme catalysis and metabolic regulation.

      Strengths and weaknesses of methods and results:

      A key strength of the study lies in its use of cryo-EM, a technique well-suited for resolving large, dynamic macromolecular complexes like the GS decamer. The reported resolutions (down to 2.15 Å) initially suggest the potential for detailed structural insights, such as side-chain interactions and ligand density. However, several methodological limitations significantly undermine the reliability of the results:

      (1) Cryo-EM data processing: The absence of critical details about B-factor sharpening - a standard step to enhance map interpretability - is a major concern. For high-resolution maps (<3 Å), sharpening is typically applied to resolve side-chain features, yet the submitted maps (e.g., those in Figures 1D, 2D, and supplementary figures) appear unprocessed, with density quality inconsistent with the claimed resolutions. This makes it difficult to evaluate whether observed features (e.g., glutamine binding) are genuine or artifacts of unsharpened data.

      (2) Modeling and density consistency: The structural models, particularly for glutamine binding at the decamer interface, do not align with the reported resolution. The maps shown in Figure 2D and Supplementary Figure S7 lack sufficient density to confidently place glutamine or even surrounding residues, conflicting with claims of 2.15 Å resolution. Additionally, fitting a non-symmetric ligand (glutamine) into a symmetry-refined map requires justification, as symmetry constraints may distort ligand placement.

      (3) Biochemical assay controls: While the enzyme activity assays aim to link structure to function, they lack essential controls (e.g., blank reactions without GS or substrates, substrate omission tests) to confirm that ATP hydrolysis is GS-dependent. The use of TCEP, a reducing agent, is also not paired with experiments to rule out unintended effects on the PK/LDH system, further limiting confidence in activity measurements.

      Achievement of aims and support for conclusions:

      The study falls short of convincingly achieving its goals. The claimed high-resolution structural details (e.g., side-chain densities, ligand binding) are not supported by the provided maps, which lack sharpening and show inconsistencies in density quality. Similarly, the biochemical data do not robustly validate the structural claims due to missing controls. As a result, the evidence is insufficient to confirm glutamine binding at the decamer interface or the functional relevance of the observed structural features.

      Likely impact and utility:

      If these methodological gaps are addressed, the work could make a meaningful contribution to the field. A well-resolved GS decamer structure would advance understanding of enzyme assembly and ligand recognition, while validated biochemical assays would strengthen the link between structure and function. Improved data processing and clearer reporting of validation steps would also make the structural data more reliable for the community, providing a resource for future studies on GS or related enzymes.

      We disagree with the reviewer’s overall assessment.

      With regard to sharpening and resolution: we examined sharpened maps and in a revised version will present additional supplementary figures showing these maps side by side. We note that the resolutions reported are global and that the most interesting features are, of course, in the periphery and subject to conformational and compositional heterogeneity. We will include supplementary figures of core side chain densities that are more like what are expected by the reviewer in the revision. With regard to modeling: the apo filament and turnover filament datasets were handled nearly identically. The additional density is therefore likely not artefactual to the symmetry operator - however, the lower resolution in this region noted by the reviewer is worthy of further exploration. The maps are public and we think this is the most plausible interpretation of the density, which we based primarily on the biochemical data and will include more speculation in the version.

      With regard to the biochemical controls: we point the reviewer to Figure S1, which shows that omission of ammonia or glutamate in the wild-type (tagless) system removes any coupling of the reactions. We will perform the additional controls to publication quality in the revised version along with the TCEP control. We note that the reducing agent is present across all experiments, ruling out an effect on any specific result. The inclusion of TCEP is also very standard in other published uses of the Coupled ATPase assay (e.g. PMID: 31778111 and PMID: 32483380 by our first author)

      Additional context:

      Cryo-EM has transformed structural biology by enabling high-resolution analysis of large complexes, but its success hinges on rigorous data processing and validation steps that are critical to ensuring reproducibility. The challenges highlighted here are not unique to this study; they reflect broader issues in the field where incomplete reporting of methods can obscure the reliability of results. By addressing these points, the authors would not only strengthen their current work but also set a positive example for transparent and rigorous structural biology research.

      All the data is public and the reviewer or anyone is free to reinterpret the maps and models - and we encourage that rather than just an interpretation of our static figures. In addition, we will upload the raw micrograph data for the apo filament and turnover filament datasets to EMPIAR prior to submitting the revision.

      Reviewer #3 (Public review):

      In this manuscript, the authors propose a product-dependent negative-feedback mechanism of human glutamine synthetase, whereby the product glutamine facilitates filament formation, leading to reduced catalytic specificity for ammonia. Using time-resolved cryo-EM, the authors demonstrate filament formation under product-rich conditions. Multiple high-quality structures, including decameric and di-decameric assemblies, were resolved under different biochemical states and combined with MD simulations, revealing that the conformational space of the active site loop is critical for the GS catalysis. The study also includes extensive steady-state kinetic assays, supporting the view that glutamine regulates GS assembly and its catalytic activity. Overall, this is a detailed and comprehensive study. However, I would advise that a few points be addressed and clarified.

      (1) In Figure 2D and Supplementary Figure 7, the extra density observed between the two decamers does not appear to have the defining features of a glutamine. A less defined density may be expected given the nature of the complex, but even though mutagenesis assays were performed to support this assignment, none of these results constitutes direct and conclusive evidence for glutamine binding at this site. I would thus suggest showing the density maps at multiple contour thresholds to allow readers to also better evaluate the various small molecules under turnover conditions that cannot be well fitted based on this density map, helping to provide a more balanced interpretation of the results.

      (2) On the same point regarding the density for the enzyme under turnover conditions, more details should be provided about the symmetry expansion and classification performed, and also show the approximate ratio of reconstructions that include this density. Did you try symmetry expansion followed by focused classification, especially on the interface region?

      (3) The interface between the two decamers of the model needs to be double-checked and reassigned, especially for the residues surrounding the fitted glutamine. For example, the side chain of the Lys residue shown in the attached figure is most likely modeled incorrectly.

      We thank the reviewer for the feedback. As noted above, we will include supplemental figures that show maps at multiple thresholds and sharpening schemes. We noted in the manuscript and above that our interpretation here is based on integrating biochemical evidence alongside the density and will make that even more clear in the revised manuscript. The filaments +/- the putative glutamine density were processed nearly identically, but we will attempt various schemes of focused classification/symmetry expansion in the revision as well. However, we point out that there is extensive averaging there that makes modeling a bit trickier than expected given the global resolution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Limitation to Di-decamer Formation:

      Could the authors clarify why hGS, when visualized by cryo-EM, predominantly forms di-decamers rather than extended filaments? Since glutamine bridges the two decameric rings, one would expect this to promote further polymerization. It remains unclear why longer filaments are not observed under the conditions used. Moreover, data in Figure 1B-1D are not mixed with Gln, and the mechanism of the apo-form filament formation is not clearly discussed. If K52 and C53 are in charge of filamentation, it is supposed to form a long filament, not a stack of 2 decamers. The kinetics of wild type and mutations, including K52A and C53A, are different. This confused me as K52 and C53 don't participate in the reactions of glutamine synthesis. Does data reduce catalytic efficiency with K52A or C53A mutation suggest that di-decameric GS exhibits a greater catalytic turnover rate than the pure decameric GS?

      We thank the reviewer for pointing this out and it was indeed filament length that was a point of curiosity during the study. Prior to preprint, we repeated the freezing conditions under identical turnover conditions and with high protein concentration as reported and indeed found much longer filaments - pointing to the capacity of the system to form larger complexes. However, we only captured a couple of ‘screening’ images and did not collect a full second dataset. Therefore, we hypothesize that filament length may be stochastic based on the subtle differences of individual grid vitrification based on the assumption that decamers within a filament can freely and quickly exchange. However, as shown in the time-resolved cryoEM experiment, the fraction of particles that are characterized as participating in a filament form (length-agnostic metric), does not appear to be subject to individual grid vitrification conditions but rather by experimental conditions (concentration of reaction-derived glutamine).

      The filament formation in apo state was not further explored because the concentrations required to achieve filament formation in this case were supraphysiological.

      We have edited the discussion to emphasize this point more clearly:

      “While the enzyme concentrations to achieve robust filamentation in the absence of glutamine are much higher than observed in cells, the protein concentrations used in our time-resolved cryoEM experiments where filamentation is correlated with accumulation of glutamine are within the range of intracellular GS concentrations in S. cerevisiae (Engel et al. 2025) and human cell lines (Wiśniewski et al. 2014).”

      Regarding the ability of apo-GS to form filaments - we identified that these residues were important based on their structural location at the decamer: decamer interface (Figure 1D; away from the active site as pointed out) and because their individual mutation to alanine attenuated the ability to form higher order filaments (Figure 1E). Therefore, these mutants were crucial controls in the steady-state kinetic experiments reported (Figure 2E). Here, we used exogenous glutamine to seed/stabilize filaments because we identified glutamine as serving this function and, crucially, because we did not observe glutamine occupancy in the active site under turnover conditions (which would suggest an orthosteric feedback inhibition mechanism; Supplementary Figure 10). Under these conditions, the wild-type, filament-competent protein displayed a ~3-fold K<sub>M, ammonia</sub> increase with glutamine addition compared to no glutamine, which when considered in the context of the greater E-305 flap conformational heterogeneity speaks to a model of allosteric feedback inhibition. Importantly, K52A and C53A show no difference in K<sub>M, ammonia</sub> between the glutamine and no glutamine condition, suggesting that attenuation of filament formation at this interface via mutation, eliminates the kinetic deficit. Therefore, K52A and C53A are not product-inhibited in the same manner as wild-type GS.

      We have clarified the discussion to emphasize this result:

      “Importantly, point mutations of the interfacial residues do not show a K<sub>M, ammonia</sub> defect in the presence of glutamine, indicative of the importance of the filament form for product feedback.”

      (2) Origin of Di-decamer Formation in the Absence of Glutamine:

      While the manuscript demonstrates glutamine-stabilized filamentation, the spontaneous formation of di-decamers under apo conditions is not mechanistically explained. The observation of 10-mer, 20-mer, and 40-mer species in Figure 1B should be validated against molecular weight standards or through SEC-MALS. The inference of higher-order oligomers based solely on migration is insufficient. Additional characterization (e.g., SEC-MALS, AUC, or mass photometry) would clarify whether these assemblies are biologically relevant or incidental.

      We thank the reviewer for pointing out the low precision of preparatory size exclusion chromatography assignments of GS molecular weight filament depicted in Figure 1. We have included calibration standards and assignment in Supplementary Figure 1 and updated Figure 1 to include the ambiguity of these assignments in panel B. It is important to note that GS has historically been underestimated in size via these methods and was originally assigned as an octamer for which there was previous consensus (PMID: 10708854). The low concentration requirements of mass photometry preclude its use for this purpose and we are not in a position to do SEC-MALS or AUC for this. Hopefully, the negative stain, crosslinking, and cryo-EM results are sufficient to indicate that we have correlated signals with the correct species!

      (3) Validation of Interface Mutants as Decamer-only Species:

      K52A and C53A mutants are used to disrupt di-decamer formation and are shown by negative-stain EM to exist as decamers. While supportive, this is qualitative. The inclusion of quantitative biophysical data (e.g., SEC-MALS or mass photometry) would more convincingly demonstrate that these mutants do not transiently assemble into higher-order oligomers. Furthermore, molecular measurements describing the spatial relationship of interface residues - such as the distance between K52 and E55 or between C53 residues of opposing decamers - would aid interpretation. The use of the term "adjacent" (line 530) is vague and should be made more precise.

      We thank the reviewer for the thoughtful comments. Beyond negative-stain EM we also performed a biochemical validation of the filament interface through bi-functional crosslinking based on the premise that the new filament interface, as defined by the apo-filament structure, presented new/unique pairs of nucleophilic amino acid R-groups in close proximity. We used Bis-sulfosuccinimidyl glutarate (BSG) or bis-maleimoethane (BMOE) to covalently link adjacent primary amines and sulfhydryls respectively (Figure 1E). This experiment defines two key principles of GS filament formation in the absence of glutamine:

      (1) It is concentration dependent. In the wild-type case there is a protein-dependent increase on crosslinking efficiency for both crosslinkers.

      (2) It is dependent on C53 and K52. Mutation of C53 or K52 significantly attenuated crosslinking efficiency.

      To make sure that these results are more prominent, we have now included a table of these intersubunit distances between epsilon amine groups of lysines and gamma sulfhydroxyl of cysteines groups based on the apo-filament structure and labeled this as either participating in the filament interface or not. Furthermore, in line with multiple reviewers comments, we have updated Figure 1D to include a sharpened representation of the map that shows strong side chain density for the amino acid side chains to further support these reported side chain distance measurements.

      We thank the reviewer for pointing out the low precision of the SEC chromatogram interpretation of Figure 1B and the figure has been amended to show filaments of variable length, instead of defined length. We also included a calibration curve to Supplemental Figure 1B and estimated molecular weights. While these estimates of size are lower than ground truth it is important to note that GS has historically displayed smaller than predicted molecular weights via size exclusion chromatography and analytical ultracentrifugation where initial characterization papers defined the oligomeric state as an octamer rather than decamer (PMID: 10708854). These studies and the present indicate potential adherence to resin and/or other factors about the shape of GS that lead to longer retention. Lastly, these SEC procedures were performed as a preparative step rather than for analytical purposes, so resolution was not the ultimate goal.

      (4) Terminology: "Scarless" hGS:

      The term "scarless human glutamine synthetase" is unconventional and potentially confusing. If it refers to the wild-type sequence lacking N- or C-terminal tags or mutations, I recommend using the term "native hGS" for clarity.

      We usually reserve “native” for proteins isolated from the original species and not recombinantly expressed (as here). So we will leave the term scarless in the document.

      (5) Helical Parameters of Filament Assembly:

      The manuscript states a ~26{degree sign} rotation (clockwise or counterclockwise?) between decamers in the filament, yet does not describe how this value was derived. Given that hGS filaments form helices, this parameter could be assessed via helical reconstruction. Is it possible the actual helical twist is ~30{degree sign}, implying 12 stacked decamers per full turn? Please elaborate on how the rotational angle was determined.

      Rotation was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was measured. Helical reconstruction was not pursued in this work owing to the typically short filaments observed in micrographs and the relative ease by which a 20-mer species could be selected via traditional 2D and 3D classification/reconstruction methods.

      We have added to the methods the following to better illustrate this measurement:

      “Decamer: Decamer rotation across the filament interface was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was depicted in Figure 1D.”

      (6) Time-Resolved Cryo-EM and Filament Growth:

      The use of time-resolved cryo-EM is innovative; however, the accessible timescales are relatively short. I suggest complementing this approach with techniques such as dynamic light scattering (DLS) or mass photometry, which allow extended real-time monitoring of filament assembly over longer durations (e.g., hours). These methods can also provide higher temporal resolution and particle size distributions.

      We thank the reviewer for this suggestion and agree that understanding the kinetics of filament formation is critical. While DLS and mass photometry are excellent for monitoring assembly over hours, our data indicates that GS filament formation occurs on a much faster timescale.

      As shown in Figure 2C, when we added ATP and Glutamine directly to GS and vitrified the sample after only 5 minutes, the majority of particles had already formed filaments, indicating that the interaction had reached saturation. This contrasts with our time-resolved experiment, where the kinetics of filament formation were likely rate-limited by the enzymatic generation of glutamine rather than the assembly process itself.

      Consequently, we anticipate that filament assembly occurs on the order of seconds or less—a timescale we interpret as a necessary prerequisite for a rapid and effective cellular feedback mechanism. Therefore, we believe the current cryo-EM data accurately captures the biologically relevant window of assembly.

      (7) Cryo-EM Symmetry Imposition and Loop Flexibility:

      The use of D5 symmetry in cryo-EM reconstructions may obscure conformational heterogeneity in flexible elements, such as the E305 loop. Since the authors used MD simulations to characterize loop dynamics, it would strengthen the study to also perform symmetry expansion followed by non-uniform refinement and alignment-free 3D classification of individual subunits. This could provide experimental validation of the proposed conformational variability.

      We agree with the reviewer that symmetry enforcement can mask conformational heterogeneity, particularly for flexible elements like the E305 loop (the E-flap). To address this, we followed the reviewer’s suggestion and performed symmetry expansion on both the turnover decamer and filament consensus maps. This was followed by focused, alignment-free 3D classification on the asymmetric unit containing the E-flap.

      Our analysis revealed a clear distinction: while 4 out of 12 turnover decamer classes showed partial density for the E-flap (class 3, 7, 8, and 12)—consistent with the flexibility observed in our MD simulations—none of the turnover filament classes demonstrated similar density. To ensure a direct comparison, we utilized a C5-expanded turnover decamer map to maintain an identical asymmetric unit to the D5-expanded turnover filament map. We note here that the turnover decamer consensus volume is different from the deposited map for which no symmetry was applied.

      Despite the different initial symmetries (D5 for filaments vs. C5 for decamers), we utilized a C5-expanded decamer map to maintain an identical asymmetric unit. For transparency, we have uploaded this C5-refined consensus map and all resulting 3D classification maps to Zenodo. We agree that the text is now strengthened given this result and we have added the following to the main text:

      “The differential loop density between turnover-decamer and turnover-filament species was further supported by 3D classification of symmetry-expanded particles, which recovered partial E305-loop density in 4/12 turnover-decamer classes (C5; Supplemental Figure 18) compared to 0/12 classes for the turnover-filament (D5; Supplemental Figure 19).”

      (8) Crosslinking Gel Analysis (Figure 1E):

      The SDS-PAGE gels shown in Figure 1E have molecular weight ladders cropped. For proper interpretation, please include full ladders with size markers and labels. In addition, clarify whether the crosslinked samples were denatured in reducing buffer. Crosslinking efficiency and specificity using BMOE or BSG require verification under reducing conditions (e.g., DTT, β-mercaptoethanol, or TCEP) to confirm covalent linkage between decamers.

      We have added in Figure 1E molecular weight markers estimates to aid in gel interpretation and have included the uncropped gels in Supplementary Figure 1D that contain the full MW ladder. The methods were clarified to indicate that reducing reagent was used in both the crosslinking reaction and all SDS-PAGE samples.

      “Protein samples were diluted to concentrations noted in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.1 mM TCEP) and, reacted with crosslinker to a final concentration of 0.5 mM for 10 mins at room temperature followed by quench in 5X SDS-PAGE sample buffer (225 mM Tris pH 6.8, 50% glycerol, 0.05 % SDS, 0.2 mg/mL bromophenol blue, 1M DTT) supplemented with 100 mM of either NH<sub>4</sub>Cl (to quench BSG reactions only) or DTT (to quench BMOE). Protein concentrations were normalized after quench prior to SDS-PAGE analysis.”

      (9) Missing Reference for NADH-Coupled Assay:

      Line 678-679 refers to an NADH-coupled assay described "previously" without citing a source. Please provide a proper reference to ensure reproducibility.

      The original paper describing the implementation of a coupled-assay to measure ADP production from glutamine synthetase was written by Bennett Shapiro and Eric Stadtman in 1970 and has been included. We will note that the conditions of this assay have been much improved since this time with better buffers, commercially available reagents of combined lactate dehydrogenase and pyruvate kinase, and modern plate readers. We added the following reference:

      “Shapiro, B.M. and Stadtman, E.R., 1970. [130] Glutamine synthetase (Escherichia coli). In Methods in enzymology (Vol. 17, pp. 910-922). Academic Press.”

      (10) Unclear Description of the NADH Assay:

      The stability of NADH is influenced by pH and light exposure. Please specify the pH range used in the assay and whether precautions (e.g., light shielding) were taken. NADH autoxidation at high pH or degradation at low pH could impact assay reliability and should be addressed in the Methods section.

      For clarity and transparency the following text was added to the Methods section.

      “Stocks of ATP, NADH, and phosphoenolpyruvate were made in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.5 mM TCEP) and the pH was adjusted until it reached 7.5 on ice prior to aliquoting, flash freezing, and storage at -80°C in the dark. NADH was only exposed to light upon thawing and assay set-up and no appreciable change in absorbance of control experiments were noted.”

      (11) Ligand Density in Figure 2 and Supplementary Figure 7:

      The density attributed to glutamine, ADP, and phosphate appears broader than expected. Please include cross-correlation (CC) values, estimated occupancies, and Q-factors for ligand fitting. Varying the contour level to assess density consistency would clarify whether the observed volume represents multiple conformations, partial occupancy, or overfitting. A similar concern applies to the cysteine sidechain density.

      We have updated Supplemental Figures to include:

      (1) Globally refined map in comparison to locally refined map where both are sharpened per previous feedback.

      (2) Ligand placement now also include Q-scores and CC values

      We did not include multiple contour levels because these are included in the resolution representative Supplemental Figure and because alternative contours do not influence Q-scores. From this analysis it is apparent that phosphate and ADP are both worse fits to the density.

      Moreover, we have now included Supplementary Table 2 that includes all ligand validation statistics for the reader to evaluate the range of B-factor, CC values, and Q-scores for all ligands in all models.

      (12) Missing Ligand B-factors in Supplementary Table 1:

      The ligand refinement statistics in Supplementary Table 1 are incomplete. Please include B-factors and occupancy values for all ligands.

      We have updated the PDB depositions to include B-factors in .cif files that are now available. We have also included Supplementary Table 2 in the manuscript detailing the ligand statistics for all models including cross-correlation, Qscore, and Bfactor.

      (13) Style and Formatting Issues: format consistently throughout.

      (a) Kinetic Parameters: Please follow the IUPAC and IUBMB-recommended formatting:

      kcat should be italic with subscript.

      KM should be italic K with upright M.

      Use lowercase s<sup>-1</sup>, not uppercase S<sup>-1</sup>.

      Refer to:

      IUBMB enzyme nomenclature guidelines https://iubmb.org/wp-content/uploads/2021/01/Current_IUBMB_recommendations_on_enzyme_nome nclature.pdf

      IUPAC Green Book https://publications.iupac.org/books/gbook/green_book_2ed.pdf

      We have corrected the abbreviations according to the reviewers recommendations.

      (b) Inconsistent Terminology and Typography:

      cryo-EM vs. cryoEM are used inconsistently - standardize throughout.

      FSC 0.143 appears with and without subscript formatting-please unify.

      Line 123: "X-ray" should be capitalized.

      Line 266: CryoEM should cryoEM, lowercase "c"

      Line 571: "100 μg ml<sup>-1</sup>"-use superscript minus; ensure consistency with "mg ml<sup>-1</sup>".

      Lines 605, 606, 626, 628: MgCl<sub>2</sub>-ensure the <sub>2</sub> is subscripted throughout.

      Temperature units (lines 607, 630, 640): Write as "4 {degree sign}C" instead of "4C".

      Microliters (lines 660, 681, 697): Replace "uL" with "μL".

      Line 797: Use superscripts: K<sup>+</sup>, Cl<sup>-</sup>.

      Line 862: CO<sub>2</sub> should appear with subscript.

      We have made all terminology and typography consistent throughout based on these suggestions.

      Reviewer #2 (Recommendations for the authors):

      To strengthen the manuscript and address the methodological and interpretational gaps identified, we recommend the following revisions and additions:

      (1) Data processing and cryo-EM map quality

      (a) B-factor sharpening: Reprocess all cryo-EM maps using standardized B-factor sharpening workflows (e.g., the autoSharpen tool in cryoSPARC or similar methods) to enhance side-chain and ligand density visibility.

      We have updated main and supplementary figures to include sharpened maps. All maps were sharpened using the Autosharpen feature of Phenix, specifically, by half-maps. We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      (b) Document the specific parameters used (e.g., B-factor values, solvent content estimates) in the Methods section to improve transparency.

      See above regarding the additions made to the methods section.

      (c) Map replacement and reanalysis: Replace all figures and supplementary panels displaying raw (unsharpened) maps (e.g., Figures 1D, 2D, 3A/B, 4A, 5B, and Supplementary Figs. 2D, S7B, S10) with the newly sharpened versions. Reanalyze density features (e.g., glutamine binding sites, ATP triphosphate groups) using these revised maps and update results to reflect any changes in interpretation.

      As requested, we have updated figures with sharpened maps and found our original analyses to hold. In particular, we have included multiple metrics of ligand model scoring in Supplementary Figure 7B including Q-score and CC. Additionally, we have included all ligand model statistics in Supplementary Table 2.

      (2) Structural modeling and validation

      (a) Ligand fitting justification: Provide high-resolution ({less than or equal to}3 Å) density slices or side-chain density close-ups (e.g., for phenylalanine rings or glutamine-binding regions) to validate claims of atomic-level detail. For non-symmetric ligands (e.g., glutamine) fitted into symmetry-refined maps, explicitly describe how symmetry constraints were adjusted or applied during fitting (e.g., local symmetry refinement, manual adjustment of ligand orientation) and include validation metrics (e.g., cross-correlation scores, density fit plots) to support the placement.

      We have supplied 5 new supplementary figures to demonstrate the resolution of our sharpened cryo-EM maps (most notably Supplementary Figures 3, 4, 8, 12, 22 and panels in others) . Of particular note is the sharpened map features of R298A decamer under turnover conditions which demonstrates multiple instances of a ring density for aromatic residues.

      See discussion above regarding the placement of glutamine in the interface density and updated handling of symmetry during refinement.

      (b) Ligand density supplements: Include supplementary figures showing representative ligand-density fits (e.g., ATP, glutamine) with clear side-chain or functional group annotations, as is standard in structural biology publications.

      In our revision we have included the following updated figures and figure panels demonstrating ligand density into sharpened maps:

      Turnover Filament Glutamine Ligand: Figure 2D-E (updated representation) and Supplementary Figure 7 (new and updated representations).

      Turnover Filament ATP and Mg(II): Supplementary Figure 10 (updated representation)

      Turnover Decamer ADP and Mg(II): Supplementary Figure 5 (new figure panel)

      Turnover R298A ADP and Mg(II): Supplementary Figure 14 (new figure panel)

      (3) Biochemical assay rigor

      (a) Control experiments: Perform and report the following controls to strengthen enzyme activity claims:

      - A blank control (reaction mixture without GS, ammonia, or glutamate) to quantify background ATP hydrolysis.

      - Substrate omission controls (reactions lacking ammonia or glutamate) to confirm that ATP hydrolysis depends on both substrates and GS catalysis.

      - A TCEP effect control (compare ATP hydrolysis rates with and without TCEP) to rule out reducing agent interference with the PK/LDH coupled assay.

      We have provided blank, substrate omission, and TCEP controls in Supplementary Figure 1. These results demonstrate negligible ATP hydrolysis without complete substrate inclusion and do not indicate any impact from TCEP inclusion.

      (b) Direct activity validation: Consider supplementing the coupled assay with a more direct measure of GS activity (e.g., quantifying inorganic phosphate release via malachite green assay) to cross-validate results.

      On the merits of the PK/LDH coupled assay being used for >55 years to measure steady-state activity of glutamine synthetases and that it is a robust assay as supported by the additional control experiments presented above in Supplementary Figure 1 we have elected not to pursue tedious cross-validation with a non-continuous assay and believe our interpretation of the enzyme kinetic results hold.

      (4) Writing and presentation clarity

      (a) Methods detail: Expand the Methods section to explicitly describe:

      - Cryo-EM data processing steps, including B-factor sharpening parameters, map reconstruction workflows, and any post-processing (e.g., filtering, masking).

      - Criteria used to validate ligand fitting (e.g., density threshold values, manual vs. automated docking).

      See above the revisions made in response to critique from review #1 which we will briefly summarize here:

      We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      Focused masks are represented in Figure 2E, Supplementary Figure 18, and Supplementary Figure 19. The details around focused mask utilization are included in the revised figure captions and the following was included in the Methods.

      “Focused masks were generated in ChimeraX (v.1.7 and above). Focused refinement and 3D classification (3 Å filter resolution, PCA initialization) were performed in cryoSPARC. Strategy of class picking and refinement are noted in Supplementary Figures 18 and 19.”

      Map reconstruction workflows are present in the relevant Supplementary Figures. No post-processing steps beyond map sharpening in Phenix were carried out. In general, human GS represents a straightforward protein to reconstruction via cryo-EM.

      Ligand identification criteria was described throughout the results section. Supplementary Figure 7 was revised to show sharpened density for either globally refined or locally refined maps fit with all three products of the glutamine synthetase reaction (ADP, Pi, and glutamine) individually showing the best CC and Qscore for glutamine. Beyond Supplementary Figure 7 we also combined both biochemical experiments and cryoEM to make this ligand assignment supported by:

      (1) Time-resolved cryo-EM experiments that show increasing filament particles over reaction time (Figure 2B and Supplementary Figures 9 and 10)

      (2) Glutamine+ATP cryo-EM screening (Figure 2C) showing long filaments

      (3) Supplementary Table 2 showing reasonable ligand statistics for glutamine

      To clarify this in the Methods sections we include the following statement:

      “Ligands were placed with ISOLDE (v1.7) and those with >0.5 Qscore and supporting biochemical and/or literature precedent were built.”

      (b) Results framing: In the Results, clearly distinguish between observations supported by sharpened maps and preliminary/unvalidated features. Avoid over interpreting density in unprocessed maps (e.g., referring to "glutamine binding" in Figure 2D without noting current density limitations).

      We have updated our discussion of Figure 2D (and now also Figure 2E) to include discussion of only sharpened maps and noted current density limitations to the interpretation.

      (5) Data and material availability

      (a) Ensure all supporting data are publicly accessible:

      - Upload raw cryo-EM movies, particle stacks, and processed maps to the Electron Microscopy Data Bank (EMDB) with appropriate accession codes.

      - Deposit final atomic models in the Protein Data Bank (PDB) and reference these accession codes in the manuscript.

      We deposited maps and models with accession codes in advance of review. The PDB and EMDB codes are available in Supplementary Table 1.

      Furthermore, for the focused maps and focused classifications that were generated during the review, and for the benefit of not cluttering the PDB/EMDB, we have included these more specific analyses in Zenodo: 10.5281/zenodo.20298855.

      - Provide detailed protocols for biochemical assays (e.g., TCEP handling, enzyme purification) in the Methods or as supplementary information to enable reproducibility.

      See updates above to reviewer #1

      (b) By implementing these revisions, the manuscript will better align with eLife's standards for methodological rigor, transparency, and reproducibility, allowing readers to confidently evaluate the study's contributions to structural and functional biology.

      We agree!

      Reviewer #3 (Recommendations for the authors):

      (1) In line 252, it would be helpful to show negative-stain EM images for each SEC peak, further probing whether any peaks correspond to partially aggregated, as this could affect the measured Kcat and Km.

      We aren’t in a position to do this experiment. We routinely check for aggregation by noting Absorbance at 340nm for non-specific scattering indicative of aggregation and observed no evidence of aggregation in our fractions.

      (2) In Supplementary Figure 6, many of the classes in the "Selected Filament Classes" inset appear to be averages of closely spaced particles, which may bias the calculation and should be excluded. In the "Selected Decamer Classes", I would suggest removing the top-view particle classes, as these particles not only have significantly different ice penetration rates, but are also more difficult to distinguish in 2D classification.

      We agree and have provided an additional, more strenuous cutoff, analysis of the tr-cryo-EM data wherein only classes that show clearly aligned decamers are included and all top views are omitted (Additional Supplementary Figure 6). We are happy to say that even with the more strenuous cutoffs that our main conclusions that filaments increase with forward reaction time holds.

      (3) In line 445, "Figure 5A" should be corrected to "Figure 5B".

      We thank the reviewer for pointing this out and have made the correction.

    1. eLife Assessment

      This work presents important findings on quantifying gene coexpression from spatial omics. These quantification methods have been applied to gastruloid to describe how genes are spatialised. The description of the quantifying tools is characterized by exceptional evidence after a thorough revision.

    2. Reviewer #3 (Public review):

      Summary:

      Triandafillou and colleagues report a single-cell resolved spatial atlas of gene expression of 26 gastruloids. While previous work had analyzed either single-cell gene expression or spatially coarse-grained patterns of gene expression (van den Brink et al, 2020) the authors here use multiplexed sequential RNA FISH (seqFISH) to create the first gastruloid atlas which is simultaneously spatially and cellularly resolved. This atlas adds to a growing list of resources cataloging gastruloid development (see also Suppinger et al 2023).

      To analyze this dataset, the authors also describe a novel analytical framework. Their analysis centers around the 'L-score', which measures the degree to which pairs of genes are either coexpressed or mutually exclusive. While this metric is similar to calculating correlations in gene expressions, it has important differences (including that it can in principle be asymmetric, although the authors symmetrize much of their analysis). In addition to the gene-centric L-metric analysis, the authors also analyze cells in their dataset according to the cell type entropy (an information-theoretical measure of confidence in cell type assignment) and the 'exposure index' (a measure of the similarity of nearest cellular neighbors).

      Using this framework, the authors focus analysis of two major features of development. The first is the differentiation of the bipotent neuromesodermal progenitor (NMP) cells in the posterior of the gastruloid into either presomitic mesoderm (PSM) or spinal cord SC lineages. They use L-metric analysis to compare overlap in marker genes used to separate NMP, PSM, and SC fates. They highlight that L-metric analysis can recover spatial patterns of gene expression (without explicit spatial information) and discern subtle features of marker genes beyond simple binning of cell types (e.g. that Epha5 expression in anterior NMPs may predict future SC differentiation).

      The second is the formation of endothelial (spatial) clusters within the gastruloid. The authors highlight two subtypes of endothelial clusters: (1) smaller clusters within the somitic anterior region, and (2) larger clusters associated with endoderm. While the authors discern some subtle differences in gene expression between these two clusters, their different spatial patterns suggest a potential physiological difference that would not be captured in traditional droplet microfluidic-based scRNAseq pipelines.

      Overall, this manuscript is a sophisticated and technically sound study that will provide a valuable beachhead for future studies of developmental patterning in gastruloids and organoids.

      Strengths:

      The major strengths of this study are the overall technical sophistication of the data set and analysis, as well as its potential generalizability to other developmental systems (both in vitro and in vivo). The data are extensively analyzed and reasonably interpreted, and this atlas makes good use of the variability in gastruloid development to extract statistical structure of developmental processes. The L-score offers a parameter-free tool to analyze transcriptomic datasets that could overcome pitfalls of other approaches.

      Weaknesses:

      The major limitations of this study are the depth and novelty of the developmental processes studied. The authors provide very convincing proof-of-concept that their data set can recover known features of gastruloid development, including NMP differentiation and endothelial development. However, further analysis and/or investigation would be required to discover new principles of gastruloid development and patterning.

      Comments on revised manuscript:

      In their revised manuscript, Triandafillou et al have made substantial updates including analysis of variability with their 26 gastruloid datasets; formalization of the L-score (formerly L-metric) and clarification on its interpretation; and validation of their gastruloid samples (e.g. Hox gene colinearity). They have also clarified and sharpened language throughout the manuscript. With these additions further bolster the usefulness of this study as a resource for the gastruloid field, they do not provide major advances in understanding gastruloid development.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors performed seqFISH in 26 gastruloids and performed a variety of computational analyses on these novel spatial data sets. Whilst the data is valuable and the computational concepts useful (exposure index, L-metric, ...), the article falls short on novelty and is written using a very clunky language, often with contradictory conclusions.

      We thank the reviewer for their comments about the value of our data and computational concepts. We agree with the reviewer’s critical comments and have endeavored to address them all. We believe the resulting manuscript is greatly clarified and improved.

      Major issues:

      (1) The authors did well in explaining and detailing the provenance of data and the individual experiments performed. However, their 26 gastruloid data still constitute a very limited sampling from their total organoids: one experiment pooled 4 plates at an 80-94% success rate; 6 different aggregation experiments were done, making a total of 1843 gastruloids, sampled 26 (~1-2%). A simple IF stain of 2-3 markers in a bigger sample could have given a more accurate picture of specific domains of interest and their proximity. Regardless, more information should be given about the existing samples: variation across experimental batches, differences between 300-cell vs 100-cell gastruloids that were used.

      This omission was an oversight on our part and we thank the reviewer for catching it. We did the following to address this point:

      (1) Added date labels to Figure S1.2d (now S1.1a) so that the proportion correct for each separate experiment is clear:

      (2) We added the raw images of the gastruloids used in the study taken before fixation.

      (3) We segmented these images and quantified metrics of the masks to address differences across samples in morphology. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation. Elongation was measured as 1-(the ratio of the width and the height of the segmented gastruloid area); the code can be found here:

      https://github.com/arjunrajlaboratory/ImageAnalysisProject/blob/1b2f2119f77083c27f58f8c36b14 c48d97ea706c/workers/properties/blobs/blob_metrics_worker/entrypoint.py

      Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300).

      The literature also supports our assumption that combining experiments with different starting numbers of cells would not dramatically affect the results (Bennabi et al. 2024). In this paper they show that only 35 genes were differentially expressed between gastruloids formed from 100 cells and those formed from 300 cells (compared to 319 for those formed from 1200 cells and 336 for those formed from 50 cells, both compared to 300 cells). The same paper also demonstrates that the positioning of gene expression (as measured by IF staining) for several representative genes (Bra and Foxc1) does not significantly differ between gastruloids formed from 100 and 300 cells when normalized for overall size and AP axis length (as we’ve done in this paper as well).

      We have updated the text and revised Figure S1.1 to reflect these changes.

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. The experiment was performed 3 times on different days, so to ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].

      To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (see also Revised Figure S1.1a-c).

      (2) Language in the manuscript should be revised. Overall the manuscript is very long, descriptive and written "impressions and beliefs" are often not adequately justified and indeed can be contradictory, e.g. in Section 1: the title states "cell types' locations ...are consistent", a few sentences down we find "there was substantial variation" and "within range of what would be considered a 'morphologically normal' gastruloid". "quite consistent", "compelling patterning", "we don't believe"... these types of expressions are best avoided and replaced with data or used and bolstered with quantitative numbers such as percentages when a given cutoff is used. Another example: "location of each cell type relative to gastruloid morphology was quite consistent the posterior region ... mainly consisted in NMPs." Given T expression in the posterior, this result phrased as such appears quite inflated, in fact, looking at cell types in Figures S1, 2a/b/c, this reviewer would state they are all but consistent and indeed it takes sophisticated analyses to find a pattern (of sorts) beyond the coarse domains expected!

      We thank the reviewer for their careful reading of our paper and appreciate that the work would be strengthened by increasing the degree to which quantitative measures are used to justify the statement we make. We have made the following changes to the manuscript to address this criticism:

      (1) We more clearly delineate where we are making qualitative descriptions and have removed summary language (like ‘consistent’, ‘normal’, ‘variable’ etc.) from these sections. For example, the section the reviewer refers to originally read:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we characterized the organization of each by mapping where each cell type was found relative to other types and overall morphology. The approximate location of each cell type relative to gastruloid morphology was quite consistent: the posterior region, although highly variable in size (Figure S1.2a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]....”

      And now reads:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we first qualitatively examined where each cell type was found relative to other types and overall morphology. The posterior region, although variable in size (Figure S1.3a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]...”

      (2) We follow this qualitative description with a quantitative analysis of cell type proportion where we clearly state which variable aspects are statistically significant:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d).”

      (3) We added a summary paragraph at the conclusion of the results from the first two figures which clearly states which aspects of gastruloid organization we find to be variable and which are consistent, with statistical testing:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (3) Figure 6 is one of the most valuable parts of the work, as the authors use the battery of analyses developed to investigate the variable and not-so-robust endothelial clusters in gastruloids. However, this investigation is still very preliminary, and it should be further linked with known biology. It is still unclear what the unique organization of this cell type is (circularity isn't convincing) and whether any signalling cues of adjacent cells could explain it. Is there any evidence that more mature endodermal cell types are generated (like the suggested "liver") to give rise to endothelial cells? It would certainly be interesting to perform IF for this cell type together with mesodermal and endodermal markers to validate seqFISH predictions on a bigger sample.

      We appreciate the reviewer pointing out that the comparisons between different endothelial cell types was interesting, and agree that the clustering methods were insufficiently justified and that a more explicit consideration of the signaling context of the gastruloid could strengthen our findings.

      We have re-evaluated how we calculate differentially expressed genes. We restricted our analysis to only consider genes that are expressed at > 2 counts/cell in at least 50% of the subsets considered. The results are in shown in the revised Figure 6.

      We find that, as the reviewer suggested, some signaling genes are significantly differentially expressed. Specifically, Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in notch signaling in the posterior of the gastruloid. Tek, on the other hand, is more expressed in the somite-associated endothelial cells, and Tek has been annotated to be involved in retinoic acid signaling. These findings align with the reviewer’s observation that signaling from adjacent cells could explain or relate to differentially expressed genes.

      We also did a more thorough review of the literature, and found several papers that reported unique subsets of endothelial precursors, albeit in related systems. In [Rossi 2022] and [Rossi 2021] the authors find a population of endoderm-associated endothelial cells in gastruloids grown with a different protocol that involves Matrigel embedding, treatment with factors that promote blood development, and growth for 168 hours. In [Veenlveit 2020] the authors find a unique somite-associated population of endothelial cells in Trunk-Like Structures, which are similar to gastruloids but model later in development and have more physical organization with discrete somites. To address the reviewer’s request that we further link with known biology we have added the following to the text:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which show more tissue-like organization than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      Finally, we tested several methods of clustering and calculating circularity, and determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and any conclusions drawn from the text.

      (4) Figures 1c and 6b need statistical significance assessments.

      We thank the reviewer for pointing out that without significance testing these plots are difficult to interpret. We have removed plot 6b (see response above about removing the circularity assessments). For plot 1c we appreciate that it is difficult to interpret which cell types vary more than others in their occurrence without significance testing. To address this we did two things: we first calculated the coefficient of variation for the proportion of each cell type across samples:

      Author response image 1.

      To calculate significance, we first considered that since these values are proportions, they must sum to 1 and changes in one cell type will affect at least one other cell type within the same sample. To properly account for this when applying statistical tests, we calculated the CLR-transformed proportion and tested all pairs of cell types for significant variation. The results are shown in the Author response image 2:

      Author response image 2.

      We added a plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We also address said variation in the text:

      “We sought to quantify variability in cell type composition between gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centred log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.3a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.4).

      The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.4d).”

      (5) The article should include an analysis of Hox colinearity expression in these gastruloids as a validation of the system.

      We thank the reviewer for pointing out the importance of these genes in validating the gastruloid system and agree that assessing their expression specifically would help readers assess data quality.

      We analyzed the center of mass of expression along the AP axis for the Hox genes included in our panel (Hoxb6, Hoxc10, Hoxd1, Hoxb9, Hoxc8, Hoxc6, Hoxaas3, and Hoxb1). We highlighted these genes in Figure S1.3: their expression along the AP axis is consistent with previously reported expression in the tomoseq dataset from [van den Brink 2020]. A summary of the correlation coefficients for each individual gastruloid for all genes (blue) and the Hox genes (orange) is shown in the Figure S1.2. The Hox genes have similar correlation coefficients overall, although their variation is higher. This is likely due to differences in gastruloid pseudo-age; in future experiments we plan to include more Hox genes and use their expression to further classify gastruloids (see updated Figure S1.2).

      We have updated the text with these new results:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an ambitious and technically challenging spatial-transcriptomic atlas of 26 gastruloids using seqFISH. The authors introduce quantitative metrics (mixing score, exposure index, L-metric / scL-metric, spatial L-metric, triplets) to characterize spatial organization at multiple scales. The dataset is valuable, and several analyses are original, particularly the rank-based L-metric family for mutual exclusivity.

      Strengths:

      The authors generate one of the most detailed spatial transcriptomic datasets of gastruloids to date. They propose creative computational metrics (L-metric/scL-metric) to quantify mutual exclusivity of gene expression without predefined thresholds, and they explore organizational principles from single-cell topology to cluster-level structure. Many observations align well with known gastruloid biology, such as posterior robustness and anterior variability. The writing is generally clear, and the figures are rich.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      Several central claims rely on metrics whose computation and justification are insufficiently explained, making it difficult to assess how robust or interpretable the results are. Many choices in the analysis appear arbitrary or are insufficiently motivated (normalization schemes, choice of parameters such as the number of neighbors, the distance cutoffs, hierarchical clustering setup, and so on). The interpretations of spatial consistency, gene-program inference, and endothelial heterogeneity are plausible but might be stronger than the evidence currently supports.

      The manuscript would benefit from stronger benchmarking, quantification of uncertainty, and explicit controls for known artifacts in spatial transcriptomics (e.g., spillover, 2D slicing, cell type assignment entropy). The biological insights are promising, but since several depend on methodological assumptions that have not yet been demonstrated to be stable, they would benefit from clearer methodological explanation.

      We thank the reviewer for spending time to give constructive and actionable comments, and we believe the manuscript is greatly strengthened and more consistent and clear as a result of the changes suggested.

      The work is rich and could become a reference dataset. Then, clarifying and validating the quantitative methods will considerably strengthen the impact and reliability of the conclusions.

      Reviewer #3 (Public review):

      Summary:

      Triandafillou and colleagues report a single-cell resolved spatial atlas of gene expression of 26 gastruloids. While previous work had analyzed either single-cell gene expression or spatially coarse-grained patterns of gene expression (van den Brink et al, 2020), the authors here use multiplexed sequential RNA FISH (seqFISH) to create the first gastruloid atlas, which is simultaneously spatially and cellularly resolved. This atlas adds to a growing list of resources cataloging gastruloid development (see also Suppinger et al 2023).

      To analyze this dataset, the authors also describe a novel analytical framework. Their analysis centers around the 'L-metric', which measures the degree to which pairs of genes are either coexpressed or mutually exclusive. While this metric is similar to calculating correlations in gene expressions, it has important differences (including that it can, in principle, be asymmetric; although the authors symmetrize much of their analysis). In addition to the gene-centric L-metric analysis, the authors also analyze cells in their dataset according to the cell type entropy (an information-theoretical measure of confidence in cell type assignment) and the 'exposure index' (a measure of the similarity of nearest cellular neighbors).

      Using this framework, the authors focus their analysis on two major features of development. The first is the differentiation of the bipotent neuromesodermal progenitor (NMP) cells in the posterior of the gastruloid into either presomitic mesoderm (PSM) or spinal cord SC lineages. They use L-metric analysis to compare overlap in marker genes used to separate NMP, PSM, and SC fates. They highlight that L-metric analysis can recover spatial patterns of gene expression (without explicit spatial information) and discern subtle features of marker genes beyond simple binning of cell types (e.g., that Epha5 expression in anterior NMPs may predict future SC differentiation).

      The second is the formation of endothelial (spatial) clusters within the gastruloid. The authors highlight two subtypes of endothelial clusters: (1) smaller clusters within the somitic anterior region, and (2) larger clusters associated with endoderm. While the authors discern some subtle differences in gene expression between these two clusters, their different spatial patterns suggest a potential physiological difference that would not be captured in traditional droplet microfluidic-based scRNAseq pipelines.

      Overall, this manuscript is a sophisticated and technically sound study that will provide a valuable beachhead for future studies of developmental patterning in gastruloids and organoids.

      Strengths:

      The major strengths of this study are the overall technical sophistication of the data set and analysis, as well as its potential generalizability to other developmental systems (both in vitro and in vivo). The data are extensively analyzed and reasonably interpreted, and this atlas makes good use of the variability in gastruloid development to extract the statistical structure of developmental processes. The L-metric offers a parameter-free tool to analyze transcriptomic datasets that could overcome the pitfalls of other approaches.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      The major limitations of this study are the depth and novelty of the developmental processes studied. The authors provide very convincing proof-of-concept that their dataset can recover known features of gastruloid development, including NMP differentiation and endothelial development. However, further analysis and/or investigation would be required to discover new principles of gastruloid development and patterning.

      We agree that the developmental processes studied here are not inherently novel, and we hope that by showing sufficient overlap with different, less highly resolved methods we have created a convincing document that highlights the potential for this technique to be used to analyze other systems. We appreciate the reviewer’s comments and that the manuscript is improved after making the suggested changes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) S1.2 plates shown individually, but unclear from which experiment.

      The reviewer was right to point out this oversight — we have updated the figure (now S1.1a) with labels for the individual experiments:

      (2) Figure 2 could include a clearer indication of the types of triples/ doublets to make it even more informative.

      We thank the reviewer for pointing out that the types of triples were not clear — we’ve added a color key and more explanatory text to this figure in order which explicitly explains the type of triplets considered.

      (3) Figures should be presented in order. Figure 3c is before 3b, etc.

      We appreciate the reviewer’s attention to detail and have swapped these two panels so that their order in the figure reflects the order they are referenced in the text.

      (4) Figure 3 is interesting, and the L-metric appears useful to pinpoint crucial genes that, when expressed, indicate a type transition has occurred. It would be great to test this with another set of cell types besides NMPs/presomitic/spinal cord.

      We thank the reviewer for their interest in this biological transition, and agree that testing on another transition would be really interesting. There isn’t another set of cell types expected at this stage of gastruloid development that are predicted to have the same type of bifurcating differentiation. However, in an effort to address the spirit of this comment (that looking at other sets of cell types with the L-score would be interesting), we have used our analytical framework on a non-spatial, single-cell dataset from van den Brink et al 2020.

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (5) Figures 4 and 5 felt exploratory, and I would recommend combining them into a single figure highlighting the usefulness of L-metric and its spatial version.

      We thank the reviewer for the suggestion to merge the content of Figures 4 and 5. While we agree that they are thematically related, given the size of the heatmaps generated in the analyses, we were unable to combine them in a way that preserved the readability of the figure and stayed within the space constraints of the page size; thus, we have chosen to keep these as separate figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Quantification methods require clearer formalization and justification

      A key limitation is that the manuscript relies on several spatial metrics whose definitions are not sufficiently formalized.

      (a) To evaluate biological interpretations, the reader needs a precise description of:

      - How each metric is calculated (mixing score, exposure index, scL-metric, spatial L-metric).

      - Why specific choices were made (normalizations, parameter values, distance thresholds).

      - What the expected ranges and interpretations are.

      We agree that these elements are crucial to interpreting quantitative metrics and thank the reviewer for their close read of the work. The reviewer had many comments on the exposure/mixing values calculated for Figures 1 and 2, and for the L-metric (now called L-score) values calculated for Figures 3, 4, and 5. To address the reviewers concerns we have done the following:

      (1) Created a new, unified framework for calculating cell type exposure and mixing.

      (2) Rewritten the methods section for this section with a particular emphasis on including elements the reviewer suggested, including specifically outlining normalizations and what they account for, distances chosen and the biological rationale behind them, and the range of values expected for each measure:

      “Quantification of Cell Type Spatial Relationships: Exposure index

      To characterize the spatial organization of cell types, we computed two related metrics: a pairwise exposure index matrix capturing type-specific spatial relationships, and a scalar mixing index summarizing overall spatial integration. For each cell, we identified neighbours as all cells whose centroids fell within a specified radius r of the focal cell's centroid. We chose a value of r of 16 μm, which gave an average of 5-6 neighbors per cell. We chose this value as it captures local interactions, which was the primary goal of these analyses. For each ordered pair of cell types (s, t), we calculated the exposure index as the proportion of type s cells' neighbors that are type t:

      where Ns→t denotes the count of neighbour pairs in which the focal cell is type ‘s’ and the neighbour is type ‘t’, and Ns denotes the total number of neighbours across all type ‘s’ cells. Each row of the resulting exposure matrix sums to unity and represents a probability distribution over neighbour types for a given source type. To account for differences in cell type abundance, we normalized exposure indices relative to the expectation under random spatial arrangement:

      Where pt is the proportion of cells that are type ‘t’. Normalized values of zero indicate exposure consistent with random mixing, positive values indicate spatial attraction (co-localization), and negative values indicate spatial avoidance, with a minimum of −1 representing complete exclusion.”

      “Quantification of Cell Type Spatial Relationships: Mixing index

      To summarize overall spatial integration across all cell types, we computed a mixing index defined as the fraction of neighbour pairs involving different cell types:

      where N_cross is the number of neighbour pairs involving cells of different types and N_total is the total number of neighbour pairs. For normalization, we compared the observed mixing to the expectation under random spatial arrangement:

      where

      is the expected cross-type interaction rate given cell type proportions. Normalized values of zero indicate random spatial mixing, positive values indicate greater integration than expected (hyper-mixing), and negative values indicate spatial segregation.

      Exclusion of untyped cells. When computing the mixing index, cells lacking confident type assignments were optionally excluded from both the numerator and denominator, ensuring the metric reflects only spatial relationships among typed cells. These cells were retained in the exposure matrix to quantify how typed cells interact with unclassified cells.

      Statistical analysis of variance. To identify cell type pairs whose spatial relationships varied significantly across samples, we computed the variance in exposure indices across samples for each type pair. To account for the expected relationship between mean exposure and variance, we regressed log-variance against log-absolute-mean across all type pairs and computed residuals. Type pairs with residual variance exceeding the 97.5th percentile (two-tailed α = 0.05) were considered significantly variable, indicating spatial relationships that differ across samples beyond what is expected from sampling variation and composition differences.”

      (3) We have also rewritten the methods for how the L-score is calculated, adding emphasis to where we normalize, what ranges of values are expected, and what the interpretation of these values are:

      “Calculating the single-cell L-score

      The single-cell L-score (scL-score) was computed for each ordered gene pair (gene A, gene B) within a single gastruloid. The cell-by-gene expression matrix was filtered to retain only cells with at least 2 detected transcripts for both genes. Cells were sorted in descending order by gene A's expression values; gene A thus serves as the reference distribution, and the score is asymmetric with respect to gene order.

      Three reference distributions were constructed for gene B: (1) perfect coexpression, in which gene B's values were sorted in the same descending order as gene A; (2) perfect mutual exclusivity, in which gene B's values were sorted in ascending order; and (3) independence, in which every cell was assigned the mean expression value of gene B.

      Cumulative sums of expression values were computed for gene A, for the observed expression of gene B, and for each reference distribution. The cumulative sum of each gene B distribution (observed and reference) was then plotted against the cumulative sum of gene A. This cumulative-sum-versus-cumulative-sum representation captures how gene B's expression accumulates relative to gene A's: if gene B's expression is concentrated in the same high-expressing cells as gene A, gene B's cumulative curve rises steeply at first; if concentrated in opposite cells, the curve rises steeply at the end. The area under each curve was calculated using trapezoidal integration and normalized by the product of gene A's and gene B's total expression, yielding four normalized areas: A_observed (observed relationship), A_positive (perfect coexpression), A_negative (perfect mutual exclusivity), and A_uniform (independence). This normalization ensures that scores are comparable across gene pairs with different overall expression levels (Figure S3.2a,b). The scL-score was then defined as follows:

      If A_observed > A_uniform: scL-score = (A_observed − A_uniform) / (A_positive − A_uniform), yielding values in (0, 1].

      If A_observed = A_uniform: scL-score = 0.

      If A_observed < A_uniform: scL-score = −(A_observed − A_uniform) / (A_negative − A_uniform), yielding values in [−1, 0).

      A score of +1 indicates perfect coexpression, −1 indicates perfect mutual exclusivity, and 0 indicates independence.

      The scL-score was computed for all gene pairs in each gastruloid from the 05/07/2025 dataset (n = 18 gastruloids). The other two datasets (n = 8 gastruloids) were excluded due to lower transcript detection quality. To generate an average scL-score matrix, the analysis was restricted to a common set of 202 genes well-detected across all 18 gastruloids, and per-gastruloid matrices were averaged. Unless otherwise noted, a symmetrized scL-score was used: scL-score_sym(A, B) = [scL-score(A, B) + scL-score(B, A)] / 2.

      Calculating the spatial L-score

      The spatial L-score extends the scL-score to spatial regions. For each gene, a kernel density estimate (KDE) was fitted over all detected transcript spots and evaluated on a regular square grid spanning the gastruloid. Bin side length was set to twice the median nearest-neighbour distance between detected spots, calculated separately for each gastruloid. Spatial bins were ranked by KDE-derived density and processed identically to the scL-score calculation. Low-density bins were not filtered, as KDE smoothing produced non-zero density values throughout the imaging area. The spatial L-score was symmetrized as for the scL-score, except when displaying asymmetric heatmaps.

      Hierarchical clustering of L-score matrices

      Gene-gene distances were defined as Euclidean distances between L-score vectors. Agglomerative hierarchical clustering was performed using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean'). This approach operates on L-score vector differences rather than on pairwise L-score values directly, and therefore does not require the L-score itself to satisfy the properties of a mathematical distance metric; the Euclidean distance between L-score vectors is non-negative and symmetric by construction, satisfying the requirements of Ward's method. Heatmaps display pairwise L-score values, not vector distances.

      We applied this clustering procedure to the following gene sets:

      (1) A subset of NMP, presomitic mesoderm, and spinal cord marker genes in one gastruloid (n = 36 genes; Figure 3g).

      (2) All well-detected genes excluding cell cycle genes, averaged across all gastruloids (n = 166 genes; Figure 4 and Figure S4.3a).

      (3) All well-detected genes common to all gastruloids, averaged across gastruloids (n = 202 genes; Figures S4.1a, S4.3b, S5.1a).

      (4) All well-detected genes excluding cell cycle genes in one example gastruloid (n = 171 genes; Figures 5b, S5.2a).

      (5) All well-detected genes common between our seqFISH panel and those that were detected in > 3 cells in scRNA-seq data from [XXX] (n=207 genes; Figure S4.5a).

      (6) All genes detected in > 3 cells in scRNA-seq data from [XXX] (n=19075 genes, Figure S4.5c-f).

      In Figure S4.2a,b a transformation of the L-metric values was used to cluster genes. The scL-scores were averaged across gastruloids as described above, and then each pairwise scL-score was transformed to a distance-like value with(1 - scL)/2. The matrix was then symmetrized as described previously. Agglomerative hierarchical clustering was performed directly on this transformed gene-gene distance matrix using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean').”

      (4) We have added an illustrative figure about how the L-score is calculated which defines expected behaviour for several cases, gives a visual explanation of the process, and shows several extreme behaviours and their biological interpretation (see Revised Figure S3.2).

      (b) For example:

      - Mixing score: The normalization is unclear. Why only 1-nearest neighbor instead of k-NN? Why not consider existing spatial-autocorrelation metrics such as Moran's I, which would also apply to gene-level mixing?

      - Exposure index: The normalization makes the metric unbounded (e.g., exposure > 1 when local frequency > global frequency). It is unclear whether this behavior is intended. Since exposure to self is meaningful, the same metric could replace the mixing score and simplify the framework. The choice of k = 5 is not justified; parameter-free approaches like Delaunay triangulation could avoid arbitrary cutoffs. If k-nn is preferred, then the robustness of the score to change the k value should be studied.

      These issues make it difficult to interpret the magnitude of reported effects or compare them across studies.

      The reviewer makes an excellent point — we have completely overhauled this analysis in the following way to address the issues raised:

      (1) Created one unified metric (see points 1 and 2 above) which considers for every cell, the identity of its neighbors in a 16 um radius (on average 5 or 6 neighbors for each cell in each gastruloid). This value was chosen so that in most cases, the cells in the immediate vicinity of a cell were considered, but not those further out (i.e. the measure is sensitive to close interactions rather than far ones). We made a matrix of all interaction pairs for a given gastruloid, with diagonal elements representing within-type interactions and off-diagonal elements representing cross-type interactions. We have replaced the previous description with the following:

      “Several studies of gene expression in gastruloids have used pooled measurements to infer the AP axis-location of genes and cell types [3,7,13,24,27] and our data are largely consistent with these lower-resolution findings (Figure S1.3a). Yet it is obvious from individual gene staining [1,4,8,28] and our detailed 2D maps of cell identity and location that gastruloid organization is much more complex than the average order of cells along the AP axis. We thus needed an analytical method for quantifying spatial organization beyond distributions along the AP axis. To further characterize spatial organization, we sought to quantify the degree to which cells were mixed in each gastruloid, and how that mixing might vary between gastruloids. For each cell in each gastruloid, we counted the interactions between that cell and all its neighbours within a 16 μm radius (on average 5-6 neighbors per cell), and summarized all these interactions for all cells in the gastruloid in a matrix, normalizing each element by the frequency of the cell type considered to be the ‘neighbour’ in the interaction.”

      (2) To quantify overall mixing, we calculate the sum across types of the frequency of self interactions (normalized to the total interactions) and then take the inverse (1-M). We then normalize this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of that cell type).

      Because the density of cells is fairly consistent across the gastruloids, w is very close to p (the proportion of that type).

      is the expectation of cross-type rate with random mixing.

      Mixing ranges from -1 (totally segregated) to +1 (more mixed than random, i.e. there is attraction between unlike types). 0 is completely random, and negative values indicate that cell types within that gastruloid tend to cluster. We added the following to the text to explain this:

      “To quantify overall mixing, we calculated the sum (across types) of the frequency of self interactions, normalized to the total interactions) and then took the inverse. We normalized this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of the neighbouring cell type). This gave us, for each gastruloid, a value that we call the mixing index that ranged from -1 (totally segregated) to +1 (totally mixed with less frequent self-interactions than expected from chance). A mixing index of 0 indicates a random distribution, i.e., neighbour frequency is exactly what would be predicted by that cell type’s frequency alone.”

      We also edited the following description of the overall distribution of mixing indices:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). All gastruloids had a negative mixing index, indicating that they all, on average, had more like-cell type interactions than would be expected given random mixing of types. However, we note that there is a ~14% difference in the mixing index across gastruloids, meaning some variation in overall mixing is present.”

      (3) The exposure index for a given pair can be found from the off-diagonal elements of the interaction matrix, and the normalization means it represents relative overexposure/clustering (positive values) or underexposure/avoidance (negative values). The minimum value is -1 and the maximum is (1-pt)/pt. We changed the description of how the exposure index is calculated to reflect this unified method of quantification:

      We noted that the off-diagonal elements of the matrix we used to calculate the mixing index were informative about cell type-cell type interactions. Specifically they quantify the degree to which each cell type (source) is exposed to another cell type (neighbours). To assess the overall frequency of cell type-cell type interactions, we first pooled the data from all gastruloids together into one interaction matrix (Figure 2b). The measure can range from -1 (no interactions at all), with higher values indicating a greater frequency of being found in close proximity. It is asymmetric in that the exposure of cell type A to B may not be the same as the exposure of cell type B to A.

      We have updated all of the quantification in Figures 1 and 2 with these new measures.

      (2) L-metric: unclear justification and interoperability

      (a) First of all, even if it is not a major issue, the L-metric is not a "metric" at least in the mathematical sense of a metric since a metric is always positive. The L-metric is central to several major conclusions (gene exclusivity, modules, spatial organization), but its conceptual basis and computational steps need more justification.

      We thank the reviewer for pointing this out and have changed “L-metric” to “L-score” throughout. We have also endeavoured to clarify the conceptual basis and have fleshed out the various computational steps as outlined in more detail in our responses below.

      (b) Several steps (ranking, cumulative curves, area under the curve) are difficult to interpret biologically It is unclear why each transformation is required and how it responds to common scenarios (highly expressed genes, correlated vs mutually exclusive patterns).

      Since the metric is rank-based, two genes that are both highly expressed in all cells may show low L-metric despite being truly correlated.

      We appreciate the reviewer’s comments about both the interpretation of the scL-score calculation and how it behaves in common expression scenarios, particularly for genes that are broadly expressed across many cells. To address these points, we generated a set of simulated examples spanning five scenarios: ubiquitously expressed genes with similarly high average expression, ubiquitously expressed genes with differing average expression, ubiquitously expressed genes with similarly low average expression, genes coexpressed across a subset of cells rather than all cells, and genes generally expressed in opposite subsets of cells. For the first three simulations, we independently sampled two genes across 50 cells using Poisson distributions with mean expression set to 100 or 50 (to simulate a gene with high or low average expression, respectively), without any expression bias towards any subsets of cells. For the latter two simulations of dependent expression relationships, we first sampled gene 1 across 50 cells using a Poisson distribution with mean expression set to 4, then generated gene 2 from gene 1 by sampling from cell-specific Poisson distributions fitted to either generally match or oppose the transcript count obtained for gene 1 in that cell. These simulations show that genes can independently appear broadly coexpressed at the population level (regardless of average expression) simply by being ubiquitously expressed, yet still receive low scL-score values. The simulations of dependent coexpression or mutually exclusive expression relationships receive scL-score values near +1 and -1, respectively. These results align with the reviewer’s prediction, but they reflect why we designed the L-metric to follow a rank-based methodology since they preserve the specificity of the metric’s upper bound (+1) for detecting non-random coexpression relationships rather than chance coexpression relationships resulting from independently ubiquitous expression. The results of these simulations are depicted (See Revised Figure 3.3).

      We have made the following edits to the text to specifically address the case the reviewer raised about highly expressed genes:

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3.”

      (c) Interpretation of L-metric values is ambiguous

      What does 0 represent? Randoms? Co-expression? Is 1 the strongest exclusivity? The manuscript currently mixes "co-expression" and "mutual exclusivity" scales.

      We agree with the reviewer that clearly defining what values of the L-score mean is critical to understanding the text. We have added a more explicit discussion of this in the text (excerpted from the response to 2c):

      “A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes.”

      And made a figure representing visually how the L-score is calculated which shows the behaviour and biological interpretation of several extreme values and an example of how the L-score is calculated. See new Figure S3.2:

      We also edited figure captions where we referred to plots as ‘co-expression’ plots, since in some cases the plots showed genes that were mutually exclusive or not expressed together in most cells. Figure S3.1:

      “c. Spatial distribution of the expression of Nkx1-2 and Rfx4 in an example gastruloid.”

      In all other cases we checked, we used the term “co-expression” to mean the opposite of mutually exclusive, as outlined in the definition above.

      (d) Additional issues also limit interpretability - Benchmarking is missing.

      - No tests on synthetic datasets, negative controls, or curated examples.

      - Prior exclusivity methods (EEI, COTAN) routinely benchmark against ground truth; this is now standard.

      - The code for the L-metric seems to be missing in the repository.

      We appreciate this suggestion offered by the reviewer as benchmarking against prior exclusivity-oriented methods provides an important comparison for clarifying both where the scL-score agrees with existing approaches and where it offers distinct advantages. To address this, we explicitly compared the scL-score to the Exclusively Expressed Index (EEI), which is bounded below by 0 and increases with mutual exclusivity, and to the signed coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework, in which positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. We performed this comparison using seven simulated scenarios as well as four representative gene pairs from one gastruloid sample (2025-05-07_roi2). In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (high_high, high_low, low_low), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained low, and COEX was also positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.

      We then applied the same comparison to four gene pairs from one gastruloid sample (2025-05-07_roi2). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types. Pax6-Eogt, Rfx4-Eogt, and Nkx1-2-Rfx4 all showed opposing expression by scL-score, with values of -0.572 and -0.526, -0.970 and -0.955, and -0.537 and -0.615, respectively. EEI detected exclusivity most strongly for Rfx4-Eogt (0.171), but gave values of 0 or approximately 0 for the other two pairs, while COEX was negative for all three pairs (Pax6-Eogt: -0.235, Rfx4-Eogt: -0.350, and Nkx1-2-Rfx4: -0.106), consistent with opposing expression, with strongest signal for Rfx4-Eogt. By contrast, Cdx4-Cdx2 showed moderate levels of coexpression by scL-score (0.402 and 0.424) and EEI (0), but COEX indicated that they were not coexpressed (-0.331).

      These comparisons also clarify the practical advantage of the L-metric over the EEI and COTAN frameworks. EEI is based on binary zero/non-zero quantification and is therefore designed specifically to measure exclusivity rather than coexpression. COEX provides a signed value and, in our simulations, tracked both exclusivity and coexpression; however, on representative gene pairs from one gastruloid sample, scL and COEX diverged in magnitude for highly exclusive expression relationships (Rfx4-Eogt) and sign for a coexpression relationship (Cdx4-Cdx2), motivating our introduction of a signed measure based on the mutual exclusivity of expression with the scL-score (see New Figures S3.3 and S3.4).

      We have updated the text to address the reviewer’s concerns: we benchmark using simulations of commonly occurring scenarios (such as varying expression levels, degree of mutual exclusivity, and amount of noise present in the relationship between the two genes in question) as the reviewer suggested. We also provided a direct comparison to an earlier exclusivity measure (EEI):

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3 and Figure 3.4.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We compared the scL-score to two existing measures of exclusivity. The Exclusively Expressed Index (EEI) (Nakajima et al. 2021) is bounded below by 0 and increases with mutual exclusivity. EEI is computed from binary zero/non-zero quantification and is designed specifically to measure exclusivity but not coexpression. The coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework (Galfrè et al. 2021) can also be used to quantify relationships between genes: positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (but with varying relative expression levels), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained close to 0, and COEX was positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We calculated the scL-score, EEI, and COEX for four gene pairs from one gastruloid sample (2025-05-07_roi2). The scL-score consistently delineated gene pairs possessing opposing expression profiles, while EEI was not always able to measure those exclusivity patterns (Pax6-Eogt: scL-score=-0.572 and -0.526, EEI=0; Rfx4-Eogt: scL-score=-0.970 and -0.955, EEI=0.171; Nkx1-2-Rfx4: scL-score=-0.537 and -0.615, EEI~0). Only the scL-score was able to detect the coexpression pattern present between the positively associated expression profiles of Cdx4 and Cdx2 (scL-score=0.413, EEI=0, COEX -0.331) (shown visually in Figure S3.4a). Thus, while EEI was informative for measuring gene expression relationships characterized by mutual exclusivity, the scL-score more clearly separated positive, random, and mutually exclusive relationships on a single signed bounded scale. The COEX value trended in the opposite direction than expected, but this may be due to the fact that it cannot be calculated on single-gene pairs and necessarily uses information from the entire count table, which here only consisted of 6 genes. These comparisons combined with the simulations in Figure S3.3, clarify a conceptual advantage of the L-metric over the EEI and COTAN frameworks. In contrast, by leveraging the ranked structure of transcript counts across cells, the L-metric framework does not binarize expression and does not require fitting a parametric distribution. It can be calculated on single gene pairs, and is more sensitive to mutual exclusivity.”

      We also amended the Data and Code Availability section to include a specific reference to the L-metric package that was previously missing:

      “All code used to process the raw data and generate figures, as well as the processed data and figures can be found at the following link :

      https://www.dropbox.com/scl/fo/bchkqlbcjb8ub9m606did/AIudcWZaC566toXzb2L-jXc?rlkey=u0wgtkq8j erxoqb5ump599oip&dl=0

      Additional custom scripts used to process the raw seqFISH data can be found on GitHub:

      https://github.com/arjunrajlaboratory/NimbusImage/

      The code for calculating the L-score can be found on GitHub:

      https://github.com/arjunrajlaboratory/l-metric

      Images of all gastruloids generated for this study, as well as single-channel seqFISH images with segmentation and annotations are available here:

      https://app.nimbusimage.com/#/project/69d3fa8f1be4701f5fab6359 Raw seqFISH images are available upon request.”

      (e) Clustering using L-metric vectors

      In this study, the authors use hierarchical clustering to group genes according to the L-metric. This choice is reasonable: hierarchical clustering provides a natural representation of similarity relationships across multiple scales, and the L-metric captures a form of signed dissimilarity between genes. However, this approach raises an important issue. Standard hierarchical clustering methods typically assume a non-negative metric, whereas, as noted earlier, the L-metric can take negative values, meaning it does not strictly satisfy the requirements of a metric in the mathematical sense.

      To the best of our understanding from both the text and the source code, the authors address this issue by defining the distance between two genes A and B as the Euclidean distance between two vectors: the L-metric values from A to all other genes, and from B to all other genes. Although this procedure is mentioned in the manuscript, it is neither justified nor accompanied by any discussion of how such a distance should be interpreted. It is not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes. These two gene distances are not without overlap, but they are not the same.

      This choice magnifies the interpretability issue: readers must understand two layers of transformations. If the [-1,1] range poses problems for hierarchical clustering, simple transformations (e.g., 1 − L) or alternative clustering methods could avoid these issues.

      Given that the L-metric underlies major biological inferences (novel gene modules, spatial subclusters, endothelial states), clearer justification and benchmarking are essential. We think this can lead to more consistency in the spatial metrics.

      We really appreciate that the reviewer took the time to understand our proposed method thoroughly, and apologize for any confusion resulting from a lack of clarity in how it is calculated, and the language used to describe it. The reviewer is absolutely correct that it is not a metric in the mathematical sense; we have replaced the word ‘metric’ with the word ‘score’ throughout the text.

      The reviewer also raised concern about how the hierarchical clustering was performed, and they were absolutely correct about what the vectors represent—they are, as the reviewer states, “not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes”. The heatmaps in figures 3, 4, and 5 are intended to cluster genes that have similar expression patterns, i.e. similar L-score values with all other genes. The reviewer pointed out that this transformation wasn’t clear, so we have added an explicit explanation of what the vectors represent, as well as an explanation of why we were interested in how these vectors, which represent a ‘fingerprint’ of how the gene interacts with all other genes, clustered (because this is in the section discussing NMP differentiation we focus on a specific subset of genes here, but later apply to the entire panel):

      “For each gene annotated as belonging to any of the three cell types (NMP, PSM, or spinal cord), we calculated a vector of scL-score values with all other genes. Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      We also appreciate the reviewer’s suggestion that alternative transformations of the scL-score may improve clustering interpretability. We tried using the reviewer’s suggestion of doing a simple transform: we averaged the scL-score matrices across the 18 gastruloids using the shared gene panels, transformed each scL-score from the original [-1,1] scale to a [0,1] scale using (1-scL)/2, symmetrized the resulting matrix so that each gene pair was represented by a single value, and then performed hierarchical clustering directly on this gene-by-gene distance matrix using average linkage. We used this analysis to test whether a more direct distance-based approach would change the gene groupings recovered by our original clustering method.

      This alternative approach largely recovered the same cell type-associated groupings, but the separation between branches in the dendrogram was smaller, making fine-scale ordering harder to interpret. We measured this by evaluating the average cell type dispersion, measured in terms of additive branch length, which was 0.682 and 0.671 for the 166-gene and 202-gene panels, respectively. Since the branch separation on our original dendrograms was greater and thus representative of more robust groupings, we chose to continue using hierarchical clustering based on Euclidean distances between scL-score vectors.

      We comment on this alternative transformation we tried in the next section, when considering the clustering of the entire gene panel. We feel this is appropriate as the motivation for looking at the difference vector was derived from expected behaviour of a smaller set of genes, and as the reviewer pointed out it is not clear that that expectation should or would hold for the entire panel.

      “To generate the heatmap shown in Figure 4a and S4.1a, we used the same clustering method as described earlier with the Euclidean distance between scL-score vectors. However, we also tried clustering directly on the scL-scores themselves, by transforming each scL-score from the original [-1,1] scale to a [0,1] distance-like scale using a (1-scL)/2 mapping (Figure S4.3a,b). The results were largely consistent, however the cophenetic distance scale was relatively compressed when the transformed values were used (Figure S4.3a,b). This shallow structure implies that many branches are separated by only modest distances, so fine-scale ordering within the dendrogram should be interpreted more cautiously than the larger-scale cell type block structure. We chose to continue using Euclidean distance-based hierarchical clustering of scL-score profiles, where cell type grouping is observed alongside larger cophenetic separations between clusters” (See New Figure S4.3).

      (3) Claims of "remarkably consistent" spatial organization are stronger than the data currently support

      (a) The manuscript emphasizes reproducible organization across gastruloids, but several factors complicate this interpretation.

      We agree that the distinction between what is reproducible/invariant between gastruloids and what varies was not clear in the original manuscript. To address this, we have updated the text to more explicitly distinguish between the two categories. The other suggestions made by the reviewer to sharpen the quantitative measures to strengthen these claims was very helpful and we appreciate the thought put into them, and have used that framework (emphasizing the statistically significant variations and consistencies) in summarizing our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrate several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior:posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (b) Possible selection bias. Only elongated, QC-passing gastruloids were retained; 18/26 datasets remain. seqFISH runs with uneven housekeeping signals were excluded.

      We agree with the reviewer that our data are elongated, QC-passing gastruloids, although these represent two sources of variation (biological and technical respectively). Our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category, and we have attempted to signal this to readers by consistently including language like “morphologically normal” and “elongated”. We have further updated the language in the manuscript to emphasize this point:

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. To ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].”

      In regards to the exclusion of datasets, the only time 18 out of 26 were used was when calculating the averaged L scores for all genes in Figures 4 and 5. In this case we used all 18 gastruloids from the seqFISH run performed on 4/4/2025; this dataset had the highest spot counts due to protocol improvement between runs, and integrating the datasets with very different spot counts was problematic because a minimum expression level is needed to calculate L scores. We used all 26 samples for the spatial metrics calculated in Figures 1 and 2 (Figures 3 and 6 focus on specific gastruloids). We have added additional labels in Figures 1, 2, 3, and 6 to make clear when all 26 datasets are used and when only a subset is used.

      (c) Gastruloids are known to be variable; restricting to morphologically "normal" samples could inflate apparent regularity.

      We thank the reviewer for this observation and agree that the degree of variability among gastruloids is an important consideration. As stated in response to b), our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category. The rate of occurrence of ‘normal’ gastruloids in our hands is ~80% (Figure S1.1a). We agree that it would be interesting to consider how variations from this baseline affect cell type composition and arrangement, and while we make no claims about it in this paper, we have updated the introduction to highlight this point:

      “To address these gaps, and to create a systematic, high-resolution dataset of gene expression in gastruloids considered to be morphologically normal, we developed a spatially resolved, single-cell molecular map of the location, identity, and gene expression of cells within 26 individual gastruloids with normal morphologies. We found that despite some morphological variability within the qualitative category of elongated and polarized, “normal” gastruloids had largely reproducible cell type composition.”

      (d) Partial lack of statistical validation. The manuscript shows descriptive consistency but no formal tests across runs or batches (e.g., mixed-effects models, ICCs, leave-one-run-out validation).

      We agree with the reviewer that a quantitative comparison between batches is important. We have added the following to the text to address this point:

      “To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (See Revised Figure S1.1a-c).

      (e) 2D sampling limitations. Spatial metrics rely on a single imaging plane chosen as the "midplane," but z-position varies between gastruloids. AP projections, mixing, and triplet analyses could all be sensitive to z-plane choice. Prior work shows that 2D slices can underestimate distances and contacts by large margins (https://pmc.ncbi.nlm.nih.gov/articles/PMC5522766). These limitations should be acknowledged explicitly.

      We agree that sampling in 2D can limit the interpretation of our findings and we thank the reviewer for bringing up this important point. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      “Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.”

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2017].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (4) Uncertainty in cell-type assignment is not incorporated into spatial metrics

      Many spatial measurements depend directly on cell-type calls (exposure, triplets, mixing). However:

      (a) Anterior cell types have higher entropy in their marker-based scores (Figure S1.1a).

      We thank the reviewer for pointing out that several of the cell types in the anterior have high entropy — specifically cardiac mesoderm and paraxial mesoderm. However, we think there is nuance to this point; two of the other prominent anterior cell types (somite and endothelial) have low entropy scores overall, and that spinal cord/neural precursor cells, which are mostly posterior, have somewhat higher entropy; higher entropy scores are not exclusive to the anterior, nor is low entropy exclusive to the posterior. Inspired by the reviewer’s comments, we have re-analyzed our data to include this nuance (see response to point d) below.

      (b) Uncertain labels inflate apparent "mixing" or "disorder," because misclassifications randomly create mixed neighbors and triplets.

      We agree that uncertainty in labels could affect the interpretation of mixing. We appreciate these comments and the reviewer’s suggestions, and we have followed them in our response to point d) below.

      (c) Posterior cell types have low entropy, so comparisons between anterior vs posterior mixing may partly reflect label uncertainty, not biology.

      We thank the reviewer for bringing up this important caveat to our findings. We incorporated discussion of this in our text edits (see point d) below.

      (d) The authors should incorporate confidence measures (e.g., probability-weighted neighbors, entropy filtering, bootstrapping) to confirm that patterns hold independently of classification noise.

      We appreciate these suggestions and have chosen to use entropy filtering to assess whether the spatial organization we observe is highly sensitive to what values are considered ‘low’ entropy. We have updated the text (see below), and added Figure S2.2 to address the reviewer’s comments:

      “The contrast between organized posterior clustering and disorganized anterior mixing matches expectations based on literature that shows that self-organization mechanisms in gastruloids in the anterior vs. posterior are differentially sensitive to culture conditions, with somitic patterning requiring external matrix support [1,4,8,24], distinguishing it from the seemingly more autonomous organization observed in posterior cell types.”

      “One potential caveat to this finding is that differences in uncertainty in cell typing could be the primary driver of mixing and cell type interaction differences, both between the anterior and the posterior within an individual gastruloid, or overall between gastruloid. To control for this, we applied an entropy filter to our dataset. We filtered out cells that had entropy > 1.5 (see plot below for cutoff), which was chosen based on the distribution of entropy values for ‘none’ type cells, which effectively describe the upper limit of random transcript assignment (99.7% of ‘none’ type cells are removed with this filter, and about 50% of cardiac mesoderm cells and paraxial mesoderm cells, see Figure S2.2a). We first examined overall mixing; there was strong correlation between the per-gastruloid mixing indices before and after entropy filtering (Pearson r = 0.809, Figure S2.2b). Globally, mixing indices decreased with filtering, meaning that overall the cell types were more clustered. When we compared the absolute value of the change in mixing index pre and post-filtering to the proportion of each cell type, the only significant correlation was with cardiac mesoderm (Figure S2.2c). Exposure indices were overall quite similar after filtering, although the strength of somite-somite and somite-paraxial mesoderm interactions increased (Figure S2.2d).”

      “We also examined how entropy filtering might affect the exposure index, given that the mixing index is calculated from the exposure index of across cell types. In general, the magnitude of the exposure index values increased when more uncertain cells were excluded, but the directionality and relative ordering was not affected. Although the magnitude of change in the posterior cells types was less than the anterior cell types, the cross-cell type exposure values, particularly between paraxial mesoderm/endothelium and somites, doubled. From these results we conclude that mixing in the posterior is driven mainly by NMP/presomitic mesoderm interactions, and is overall lower than mixing in the anterior, which is driven by rarer cell types like cardiac mesoderm, paraxial mesoderm, and endothelium, being interspersed within somite cells” (See Figure S2.2).

      (e) This leads to reviewing the claim on endothelial heterogeneity, which strongly depend on spatial adjacency and gene exclusivity metrics.

      We have extensively considered claims of endothelial cell heterogeneity, and these are discussed in detail in response to the reviewer’s next point. We have also copied them here for the reviewer’s convenience:

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in the Author response image 3:

      Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although many of the somite-associated genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0). See Author response image 4.

      Author response image 4.

      (5) Endothelial "spatially dependent" gene expression may reflect spillover rather than intrinsic state

      The comparison between anterior-associated and posterior-associated endothelial nuclei suggests two transcriptional states. However, spatial adjacency confounds the interpretation:

      (a) seqFISH assigns transcripts to nuclei in dense tissue; partial-volume effects can mix RNA from neighboring endodermal or somitic cells.

      We thank the reviewer for their attention to detail and agree that a careful consideration of these points is important. We also note that given that some of our differentially expressed genes in endothelial cells are endoderm or somite genes, there indeed may be some transcript misassignment.

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript misassignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for none-typed cells at any dilation considered.

      (b) Endoderm and endothelium are closely intermixed (Figure S6.1d), and their gene expression co-localizes in KDE maps (Figure 5b-c).

      We agree with the reviewer and thank them for their close reading of the manuscript. We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      (c) Without explicitly quantifying spillover, differential expression between these two endothelial subsets cannot be confidently attributed to cell-intrinsic differences.

      (d) You may control for spatial proximity with any of the following:

      - Include adjacency index as a covariate in DE models.

      - Use scL-metric to test the mutual exclusivity of endothelial vs endoderm genes within the same nucleus.

      - Apply local permutation nulls: shuffle transcripts within local windows and recompute DE.

      - Restrict analysis to gastruloids that contain both endothelial subsets.

      We thank the reviewer for bringing up this important point, which we were eager to address. We agree with the reviewer’s point that by only considering gastruloids that contain both subsets of endothelial cells is the correct way to do the analysis. We were already only considering this case (specifically the values calculated in Figure 6 and for 1 gastruloid pictured in Figure 6a). We added an n=1 label to revise Figure 6d to emphasize this point.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although some genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0) (See Author response image 4).

      We have also included in the supplement larger images of some of the top differentially expressed genes, which more intuitively show the differential expression results (Endoderm enriched and Somite enriched).

      Finally, we have referenced several previously-reported instances in the literature where distinct subsets of endothelial precursors with unique gene expression programs were identified. Although in these cases 1) the embryo models were different (in [Rossi 2021, Rossi 2022] gastruloids made with a different protocol and treated with factors designed to promote blood development, and in [Veenlveit 2020] trunk-like structures) and 2) the methods were different (IF and 10x single-cell sequencing) this at least establishes a precedent for the observation of multiple types of endothelial precursors. In the case of [Veenlveit 2020] the authors specifically note that one subset is associated with somites, and we have updated the text to reflect these new results:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b,c). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Pecam1 also clustered uniquely in our L-metric analysis (Figure 4c), suggesting this differential expression is conserved across gastruloids. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behaviour is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6)

      (6) Interpretation of gene-program modules may be overstated

      Claims that the L-metric reveals "novel gene programs" should be softened:

      (a) The seqFISH panel is an approx. 200-gene marker-enriched panel, already biased toward known cell-type markers.

      This is true and we appreciate that this came through in the text since it’s important for the reader to understand the approach we took in this study.

      (b) Strong blocks in Figure 4a and S4.1a may reflect panel design rather than newly discovered programs.

      We agree with the reviewer that the panel design was not sufficiently highlighted, so we have made the following changes to the text to emphasize which patterns would be expected due to the genes we are probing for, and which findings were surprising given the known functional role of the gene.

      At the end of the section titled ‘The L-metric captures the spatial distribution of gene expression despite being calculated without spatial information’:

      “These analyses demonstrate that information contained within the hierarchical relationships between genes, determined by scL-score can reveal novel information about cell states within cell types, although we acknowledge that since cell type is determined by a limited panel of marker genes, results should be further functionally verified. scL-score analysis can also identify distinct spatial locations of cells in this cell state, all without explicit encoding of spatial information, but rather quantifying and clustering the degree to which genes are mutually exclusively expressed with one another.”

      At the end of the section titled ‘Clustering scL-metric vectors clearly resolves cell types and reveals novel genetic interactions’:

      “Finally, although the strong blocks we find in the heatmaps in Figures 4a and S4.1a largely reflect cell types, as is consistent with our panel design, we discovered some novel functions of genes in the panel through their location in the scL-score tree: although Tgfβ was initially included in our panel to generally detect inflammatory and growth signaling, clustering by expression patterns revealed its unique association with endothelial precursors.”

      To further address the concern that the generality of clustering is due to gene selection, we performed random gene drop-out and assessed how well cell types clustered as a function of the number of genes removed:

      Author response image 5.

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      Finally, we performed scL-score analysis on an unbiased, scRNA-seq dataset without any panel selection, and were able to show similar groupings of cell type markers:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (c) Cluster robustness is not assessed (bootstrap, stability).

      We appreciate the reviewer’s suggestion that cluster robustness should be assessed. To address this point, we performed a clustering-stability analysis on the common 202-gene panel that included cell cycle genes by asking to what extent hierarchical clustering could recapitulate cell type-based groupings of genes as the number of genes used for clustering was varied across progressively larger, randomly sampled panel subsets. For each resulting tree, we averaged cell types’ dispersion of genes across the tree using the framework described in our response to suggestion 6m and compared the observed value to a permutation-based distribution generated on the same tree.

      This analysis showed that the observed cell-type dispersion remained consistently lower than the corresponding permutation distribution across all subset sizes examined. In other words, genes assigned to the same annotated cell type remained closer together in the dendrogram than expected by chance even when clustering was performed on reduced random subsets of the panel. We also observed that dispersion values increased as larger gene subsets were included, which is expected as the clustering problem becomes more complex with increasing panel size; however, the separation between the permuted distribution of average cell type dispersion and the observed dispersion value increased as the panel subset size increased. Taken together, these results indicate that the cell type-resolved organization captured by the scL-score derived hierarchy is not dependent on one particular subset of genes, but is instead a stable property of the broader gene panel. New Figure S4.2 addressing cluster stability:

      We have added the following explanation in the text:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      (d) Agreement with cNMF (claimed in text) is not quantified (ARI, Jaccard, hypergeometric overlap).

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See Revised Figure S4.2 (now S4.4)).

      We have updated the text to reflect these quantitative comparisons:

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (e) Testing the scL-metric on larger, unbiased scRNA-seq datasets would help demonstrate generality.

      We agree with the reviewer that this would demonstrate generality, so we applied scL-score analysis to the scRNA-seq dataset from van den Brink 2020 — several gastruloids at the same stage of development as those used in this paper were pooled and sequenced. We added a figure, new Figure S4.5 with the results of this analysis, and a new section in the paper describing them:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (7) Minor Comments

      (a) We suggest including representative raw seqFISH images. The manuscript does not show raw images, which makes it difficult to evaluate the quality of the underlying data that all spatial analyses depend on. A figure showing raw fluorescence channels, detected spots, and nuclei segmentation masks for at least one anterior region, one posterior region, and one dense interface (e.g., endoderm-endothelial) would allow readers to assess spot intensity and background levels, signal-to-noise ratio, segmentation accuracy, and potential over/under-segmentation, channel cross-talk, and transcript crowding or dropouts in dense tissues. A small panel of raw images would improve the transferability to the spatial metrics.

      This is an excellent suggestion and we have included examples of the raw images for a representative gene for all hybridizations, the spots as determined by the spot-finding algorithm distributed with the seqFISH instrument, deconvolved spots, and nuclear segmentation. These images can be found in Revised Figure S1.1d.

      (b) The Introduction mentions 3D organization, which can confuse readers into thinking the seqFISH dataset is volumetric. The data shown and analyzed come from a single 2D plane per gastruloid, not from full 3D z-stacks. Since all spatial metrics rely on true adjacency, the manuscript should explicitly state early on that the dataset is 2D and briefly justify why a single plane is sufficient for the analyses.

      We appreciate the reviewer bringing up this point and we have removed references to 3D so as not to confuse readers.

      We also have included an extensive discussion of the 2D nature of the data. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2018].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (c) Results, first paragraph: "good agreement" along the AP axis should be quantified or defined.

      In the first paragraph we compare the peak in gene expression along the (length-normalized) AP axis of each gene with a similar but orthogonally measured dataset from another group (van den Brink 2020). The text specifically reads:

      “When we compared how gene expression varies along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section.”

      We have revised Figure S1.2 with a summary plot showing the distribution of correlation coefficients for all genes and for the Hox genes in our panel (which are known to be expressed sequentially along the AP axis).

      We have updated the text as follows:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      (d) Clarify what "greater cell type distinction" means when using marker panels, and what metric demonstrates improvement?

      By “greater cell type distinction” we meant that when using traditional clustering methods we were not able to individually resolve some cell types: there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together. Cluster labeling was performed by considering which genes showed up as being differentially expressed in each cluster using the same associations as were used to perform the cell type scoring with the marker gene panel. While we acknowledge that these results are subjective to clustering parameters, this is a general problem with clustering and not specific to this study. We have added additional clarification in the text and removed the phrase “greater cell type resolution” since it wasn’t clear what we were comparing to:

      “To profile the spatial organization of individual gastruloids, we assigned a cell type to each nucleus using a cell type scoring method with known marker genes. We compared these results to those obtained with unsupervised clustering. We found that although clustering did produce clusters, they were not strongly separated and we were not able to individually resolve some cell types we expected to find: specifically, there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together (Figure S3.6a). Therefore, we proceeded with the scoring-based method; see Methods for additional details. A representative gallery of typed gastruloids is shown in Figure 1b; the full dataset is in Figure S1.3. Because each cell received a cell type score for each type, we could use the entropy of the cell type score probability distribution to assess confidence of our cell type assignment: a cell that received a similar score for multiple cell types would have a high entropy distribution, while one which scored highly for one type and low for the rest would have low entropy. On average, the cell type entropies for most cells in a given cell type were low, with the exception of paraxial mesoderm and cardiac mesoderm cells, which had intermediate values (see additional discussion below). The generally low entropies indicate that most of our cell type assignments were high-confidence (Figure S1.4a).”

      (e) The variation in "none-typed" cells across gastruloids should be expanded and shown quantitatively.

      We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      (f) Figure S1.1b: explain the criteria used to decide which tissues "varied significantly." Showing the proportion of "None" per gastruloid would help.

      We thank the reviewer for pointing this out. We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      To justify the use of the phrase ‘varied significantly’ we have calculated the degree to which cell type pairs significantly co-vary (adjusted p value < 0.05) in their proportions and added the plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We address said variation in the text:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids profiled. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.32a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.41dc).”

      “The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.41dc).”

      (g) In Figure 1c, showing the variability of "None" cells is important.

      We made this adjustment and added the plot to Figure S1.3c

      (h) For Figure 1e, consider normalizing to T-expression proportion, not only AP length. This could clarify multimodality in the endoderm and spinal cord.

      We thank the reviewer for this suggestion to normalize by molecular as well as physical features. We agree this could be a useful projection of the data, since expression of T is used in many contexts to define the posterior of gastruloids.

      We incorporated T expression into the length normalization in the following way: we generated a cumulative distribution of all T spots along the AP axis, and when this value exceeded a threshold (specifically 90% of all spots) we defined this as the midpoint of the gastruloid, and linearly normalized space before and after it from 0-0.5 and 0.5-1 respectively. We then re-projected nuclei for each gastruloid onto this new coordinate system, and visualized in the same way as the main text figure (See Figure 1f).

      To assess whether this increased or decreased variability in position, we calculated how much the location of the peak of each cell type in each gastruloid differed from the peak position of that cell type in all samples pooled together. We compared this difference to a bootstrapped null drawn from the pooled distribution. Interestingly, we found that nearly all cell types had statistically significant variation (meaning the average distance from the mean fell outside the bootstrapped distribution in the positive direction), however the effect size for most cell types was very small.

      Notably, as the reviewer suggested, normalizing to T expression decreased the effect size of this variability in all cases but one, and particularly decreased spinal cord variability (despite not being a marker for spinal cord):

      We interpret these findings in the following way — that most of the variation observed in cell type arrangement along the AP axis is due to morphological and molecular variability (likely due to stochasticity in initial cell number and differences in developmental timing between gastruloids), and that once these factors are taken into account, for most cell types the effect size of variability is very small. However for some cell types, notably the location of the differentiation front, the position varies among gastruloids. We hypothesize that this may be due to the rhythmic nature of somite development. We have updated the text to reflect these changes.

      “Given that the proportions of cell types within each gastruloid were fairly consistent, to what degree did their spatial organization vary? We first projected each cell’s expression onto the AP axis and looked at the distribution of AP axis locations across gastruloids (Figure 1f, left-hand side). The most posterior cell types (NMP and presomitic mesoderm) showed wide distributions from 0-30% of the axis. Centred around 30%, spinal cord precursors and endoderm had distinctive peaks (clusters), the exact location of which varied between gastruloids. Using a threshold of T expression to define the midpoint of each gastruloid and uniformly length-normalizing each half caused these cell types to collapse into a single peak (Figure 1f, right-hand side). The differentiation front was similarly located in one peak (between 30-40% of the AP axis), and the variation in the location of this peak decreased with T-expression normalization. To quantify this variability, we bootstrapped a null distribution of cell type locations by pooling each cell type together across samples and using the resulting distribution to define a reference mean. When we compared how much peak variation there was among samples randomly drawn from this null to our observed data, we found that although nearly every cell type (excepting paraxial mesoderm) had more variability than expected by chance, the effect size of this variation was small (Figure S1.5a). Normalizing location to T-expression decreased variability in most cell types, most notably for differentiation front and spinal cord (Figure S1.5b).

      We also asked how the order of cell types along the AP axis varied between gastruloids. When we ranked the peaks shown in Figure 1f per gastruloid, we found that Kendall’s W, an overall measure of rank coherence across independent samples that spans from 0 (no agreement) to 1 (complete agreement), was 0.834 (Figure S1.5c). We found that the cell types most likely to swap rank order were spinal cord and endoderm, and paraxial mesoderm and endothelium (Figure S1.5d).”

      (i) Clarify "mixing coefficient values range from 0.29-0.58" (incomplete sentence). Section "Cell types' arrangement is consistent across gastruloids, but varies by type" second paragraph.

      We appreciate that there was some confusion here and we thank the reviewer for pointing it out. We have updated the text with the new values and a more specific interpretation to aid the reader and address the reviewer’s comment:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). While all values are negative, the range was large. This observation led us to conclude that while in all the gastruloids profiled cell types tended to cluster together, there was variation between gastruloids in the degree of coherent clustering between types.”

      (j) The paragraph claiming organization consistency may be too strong, given the lack of statistical validation. "Overall, these data speak to the consistency of gastruloid organization...".

      We agree that quantification of variability, which we only assessed qualitatively, would help readers better evaluate claims of organizational consistency, and we appreciate the reviewer pointing this out.

      We have taken several steps to add quantification, including:

      (1) Calculating the coefficient of variation for the proportion of each cell type

      (2) Quantifying co-variation and assessing statistical significance

      (3) Quantifying the variation in AP-axis location, including comparison to a bootstrapped null, significance testing, and a new form of normalization which decreases some of the variability.

      (4) Quantifying order along the AP axis and calculating Kendall’s W to quantify how concordant this ordering is across gastruloids

      (5) Overhauling the methods used to quantify local patterning and mixing into one unified metric

      (6) Calculating significance for variation in exposure index across samples

      To address the question of whether claims of organization consistency are appropriate, we have added the following summary paragraph at the end of the results for the first two figures, and have taken the reviewer’s suggestion of only discussing statistically significant or quantified results in listing both consistent and variable features. Now rather than making an argument about whether gastruloids are consistent or not, we merely provide the readers with our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (k) Figure 3c: gene set sizes differ substantially; proportion-based normalization may be more appropriate.

      We thank the reviewer for carefully noting the gene set differences. While the NMP-only and spinal cord-only gene sets each have 10 genes, PSM-only and NMP+PSM have 6 genes, and NMP+spinal cord has 5 genes. Given the relatively small N, we feel that normalization would likely introduce a layer of abstraction that would be more confusing for the reader, especially given the qualitative nature of the claims made about the shape of the plots in question. However, we agree this point is important, so we have now indicated the size of the gene sets on the plot (revised Figure 3b, previously c).

      (l) The purpose of the first two graphs in Figure 3c is unclear.

      We thank the reviewer for pointing out that this is unclear. The purpose of this visualization is to demonstrate correlation between genes shared between cell types and genes exclusive to only one of the cell types. The text reads:

      “Figure 3b shows the total expression of each gene group versus the NMP genes for all cells in all gastruloids that we typed as NMP, presomitic mesoderm, or spinal cord. As expected, there was a clear correlation between the mixed categories and NMP genes, supporting the notion of a continuous differentiation process.”

      We have added a correlation line to revised Figure 3b (former Figure 3c) to emphasize this point.

      (m) Clustering of all genes: quantify whether clusters align with cell types.

      We appreciate this suggestion offered by the reviewer as this analysis will allow readers to quantitatively assess the overlap between cell type-based groupings of genes and our scL-score-determined clusters for this dataset. To address this, we have used a bootstrapping approach to evaluate how well our scL-score-based hierarchies capture cell type-based groupings. Specifically, following hierarchical clustering of genes using the scL-score, for each cell type, we computed the average of the minimum cophenetic distance between each pair of genes associated with that cell type. After obtaining each cell type’s average, we took the mean of these averages which we refer to as the cell type dispersion for that tree. We note that the absolute value depends on the topology of the tree. To establish a reference for this measure, we bootstrapped a null distribution by preserving the same tree topology and randomly assigning genes to leaves and calculating the resulting dispersion. We performed this 10,000 times to create a reference null distribution. We compared the true dispersion value to this null, and computed a bootstrapped p-value. In both cases (with and without cell cycle genes), the observed average cell type dispersion is less than that for all permutations (p-value=0.0001, bootstrapping), indicating that genes associated with the same cell type were significantly more likely to cluster together on the observed scL-score hierarchy than would be expected under random clustering.

      We have updated the text to reflect these quantitative comparisons:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type-specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b).

      Author response image 6.

      (n) Clarify how the distance between genes is computed for clustering; the current method is hard to interpret.

      We appreciate the reviewer’s suggestion to clarify how distances between genes were computed for hierarchical clustering. To clarify this point, we have added the following to the text:

      “Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      (o) Quantify similarity between cNMF and L-metric clusters.

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2, now in updated Figure S4.4c). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See updated Figure S4.4 (formerly Figure S4.2)).

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (p) In Figure 3e, clarify what the orange lines represent.

      We have updated the text:

      “Per-cell expression scatterplots of the two pairs of genes shown in b). The y-axis of each is the per-cell expression of Eogt. The x-axis is the per-cell expression of Pax6 (left) or Rfx4 (right). R is Pearson’s r, scL is scL-score. Count data is shown in black; smoothed 2D densities are shown in orange.”

      (q) Ensure figures and panels follow the text order. For example, Figure S1.3a is referenced earlier than 1.1 and 1.2.

      We appreciate the reviewer’s attention to detail. We have changed the order of these figures so that their reference in the text follows their numeric order.

      (r) Typo in "by covariation with any other cell type (Figure S1.1e)" did you mean S1.1.c?

      We have updated the text with this change.

      (s) For circularity (Figure 6), justify the convex-hull-based measure; thin protrusions can distort interpretation. Consider the volume difference between the convex hull and the original shape.

      We thank the reviewer for this helpful suggestion. We tested the difference method suggested by the reviewer, as well as several other methods of clustering and calculating circularity. We determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and text.

      Summary

      This work delivers a rich spatial dataset and introduces creative computational tools. The main limitations lie not in the data but in the clarity, justification, and validation of the quantitative methods. We suggest strengthening these aspects by adding formal definitions, parameter justification, benchmarking, robustness tests, and controlled interpretations. This will improve the manuscript's impact and reproducibility of the methods described, besides making it easier to understand for the readers.

      We thank the reviewer for their kind assessment and also for their many insightful comments for improvement. We feel the revised manuscript is greatly improved because of them.

      Reviewer #3 (Recommendations for the authors):

      In my view, this manuscript is well-designed and clearly written, and supports all of the claims made. I have no suggestions for major revisions for this manuscript; rather, I would suggest the following as outstanding questions for future investigation:

      We thank the reviewer for a careful reading of our manuscript and the several interesting suggestions and useful references in the literature. Including the discussion of these ideas in the text (see below) has improved the flow and scope of the manuscript.

      (1) On the NMP fate, bifurcation has been studied extensively in gastruloids and related structures (see, for example, Underhill et al 2023; Bolondi et al 2024). Can this dataset from Triandafillou and colleagues reveal new regulatory hierarchies in this process? This seems possible in principle, but I was not able to reach this interpretation (for example, it seems the authors interpret the clustering in Figure 3H as reflecting spatial patterns rather than a regulatory hierarchy).

      We agree that this is an exciting implication of the work, but we feel that with the current panel (which was originally chosen primarily to type cells and not to infer regulatory structure) we would not be able to comment on this. However we feel that with a larger set of genes such inference may be possible, so we took your suggestion in point 3) and applied the L-metric to a previously existing single-cell dataset. We looked for possible regulatory structure, and found at least one interesting case where a transcription factor showed strong co-expression with two previously unconnected genes. While more specific analyses and experimental validation would be required to establish a direct relationship, we feel that this suggestion by the reviewer represents an important potential future application of this methodology, so we’ve included a new figure and explanatory text to address this possibility:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.

      (2) On the formation of endothelial clusters ('blood islands'), this process has been observed in gastruloids (Rossi et al 2022). The observation of endoderm-associated endothelium in gastruloids is interesting, but it was also not clear how the authors interpret this finding. Are there two separate endothelial differentiation paths captured here? Or is there one path, and only some migrate towards the endoderm? The authors seem to raise possibilities, and it was slightly unclear on my reading how they interpreted their findings.

      We thank the reviewer for pointing us to this paper. We think it’s especially interesting that we also see close association between endoderm and endothelial precursors, especially given that the protocol used in the referenced paper was designed to generate blood precursors (treatment with VEGF, bFGF and ascorbic acid). In light of the comments of other we re-did this analysis — Author response image 7 shows the updated list of differentially expressed genes.

      Author response image 7.

      Some of the spatially differentially expressed genes are linked to signalling, and likely reflect overall signalling differences between the anterior (where the somite-associated endothelial cells are) and the posterior (where the endoderm-associated endothelial cells are). For example, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA is higher in the anterior). Pecam1 is involved in adhesion and cell morphology, and perhaps is higher in cells interacting with endoderm due to tighter packing/association; the same could also be true of Cdh5.

      While none of these answer the reviewer’s questions about the origin of the cells (and whether it is common), we found another instance of differential endothelial populations in embryo models: in Veenvliet 2020, they find that some endothelial cells have a mesodermal (somitic) origin. Thus we may be seeing a similar phenomenon in our samples. We have updated the text to reflect these additional lines of evidence and to clarify how we think the two populations may differ (while acknowledging that we lack the tools to confidently assign cell of origin or functional differences with this technique):

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.”

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      (3) Finally, a broader question about the analytical framework. The authors emphasize that the L-metric is parameter-free; however, much of their analysis still appears to rely on baked-in priors about known marker genes for cell type assignment. Is there a way to extend their analysis to infer cell types directly from the structure of the L-metric? The hierarchical clustering in e.g., Figure 3G suggests something like this: the hierarchy mostly (but not exactly) follows the marker gene annotation. Does this suggest that the cell type labeling should be revisited?

      Once of the initial motivations behind the creation of the L-metric (now L-score) was to have a more reliable and specific way of quantifying the interaction between genes that are known in the literature to be cell type markers — with more sophisticated (and noisy) methods of analysis like single-cell RNA sequencing, we found that there was substantial variation in the specificity and ubiquity of so-called ‘marker genes’, and that their usefulness often depended on context. We appreciate that the reviewer raises this point as well, and we think that scL-score analysis, such as that exemplified in Figure S4.4, can help identify new marker genes, or at the very least distinguish the biological context in which a marker gene is useful.

    1. If you do not add any rules, all events are forwarded with the default LogScale ingest format.

      The event will be sent in raw format, not in LogScale format. This is properly described in the "Adding Splunk Integrations" section.